Multi-tenant data isolation method for large language model, gateway and computing device
Patent Information
- Application Number
- CN202610696361.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-29
AI Technical Summary
若其中一个租户在访问存储有多租户的数据对应的工具服务时,恶意访问其他租户的数据,会造成其他租户的数据泄露,产生数据安全问题
Smart Images

Figure CN122845166A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a multi-tenant data isolation method, gateway, and computing device for a large language model. Background Technology
[0002] With the widespread application of Large Language Models (LLMs), in enterprise-level Software as a Service (SaaS) scenarios, LLMs need to frequently access enterprise private data through the Model Context Protocol (MCP).
[0003] In a SaaS scenario, multiple tenants may need to use the same tool service. In this case, the tool service stores multi-tenant data. If one tenant, while accessing the tool service that stores multi-tenant data, maliciously accesses other tenants' data, it can lead to data leaks for those tenants and create data security issues. Summary of the Invention
[0004] To address the aforementioned issues, embodiments of this application provide a multi-tenant data isolation method, gateway, and computing device for large language models, which can achieve visibility isolation of multi-tenant data, defend against tenant injection attacks, and improve the security of data access.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide a multi-tenant data isolation method for a large language model. The method includes: intercepting client requests; parsing the tenant identity identifier corresponding to the request; when the request includes a tool call request, injecting isolation conditions into the request based on the tenant identity identifier to obtain a rewritten request, wherein the isolation conditions include the range of data accessed by the tenant that is restricted based on the tenant identity identifier; and forwarding the rewritten request to the server.
[0006] Based on this solution, by intercepting, parsing, and rewriting the request in three stages, tenant identity identification and dynamic limitation of access scope are completed before the request enters the server to execute logic. This effectively blocks the risk of cross-tenant data leakage caused by Prompt injection or permission configuration oversights. Compared to traditional solutions that isolate different tenants through independent containers, the data isolation method of this application embodiment does not rely on underlying resource isolation. It only uses semantic layer request rewriting and policy injection to achieve millisecond-level, fine-grained data boundary control in a shared model service architecture, significantly reducing operational complexity and resource overhead. In addition, compared to solutions that achieve isolation by modifying hard logic in server-side code, the data isolation method of this application embodiment is less invasive, more secure, and has stronger portability and policy flexibility, adaptable to different model service frameworks and API system forms.
[0007] In one possible implementation, when the request includes a tool call request, an isolation condition is injected into the request based on the tenant identity to obtain a rewritten request. This includes: parsing the SQL string in the request to obtain an abstract syntax tree (API); traversing the API and injecting isolation conditions into the WHERE clause of the API based on the tenant identity to obtain a modified API; and restoring the modified API back to an SQL string to obtain the rewritten request.
[0008] Based on this scheme, the SQL string is parsed into an AST to facilitate the search and identification of WHERE clause nodes. Simultaneously, the WHERE clause nodes are rewritten within the AST to embed isolation conditions. While conforming to SQL language specifications, this enables secure injection based on isolation conditions, thereby achieving "row-level data isolation".
[0009] In another possible implementation, the abstract syntax tree (API) is traversed, and isolation conditions are injected into the WHERE clauses of the API based on the tenant identity to obtain a modified API. This includes: traversing the API using a depth-first search to find the target node in the API and obtaining the valid alias of the target node, where the target node includes nodes that reference controlled tables, and the controlled tables are a set of target physical tables that need to be isolated; and injecting isolation conditions into the WHERE clauses of the target node based on the tenant identity and the valid alias of the target node to obtain the modified API.
[0010] Based on this solution, the security, accuracy, and reliability of subsequent isolation condition injection are achieved by combining the target node, valid alias, and context information. Even if a large model is induced by an attacker to generate a malicious full table scan command, the gateway can forcibly restrict it to the tenant scope, and the backend server does not need to be aware of the tenant logic.
[0011] In another possible implementation, a depth-first search is used to traverse the abstract syntax tree to find the target node and obtain its valid alias. This includes: inputting the root node of the abstract syntax tree; initializing the result set, which includes an empty result list for storing target node information; initiating a recursive traversal; if the node type in the abstract syntax tree is a table source node, comparing the current physical table corresponding to the current node with the target physical table in the controlled table to determine if the current physical table is in the controlled table; if the current physical table is in the controlled table, identifying the current node as the target node and determining the valid alias of the physical table referenced by the target node; and encapsulating the target node's pointer and valid alias into structured data and recording it in the result list.
[0012] Based on this solution, subsequent rewriting components can automatically perform isolation condition injection using only the result set without understanding the SQL business statements. This process decouples security policies from SQL parsing, significantly improving the maintainability and scalability of tenant isolation.
[0013] Another possible implementation includes: when the node type in the abstract syntax tree is a connection structure, recursively visiting the left and right subtrees; when the node type in the abstract syntax tree is a query statement, updating the context and recursively visiting the child nodes.
[0014] Based on this scheme, when traversing the nodes of the abstract syntax tree, the traversal path can be determined based on the node type to realize a differentiated traversal strategy, thereby accurately finding each target node and preventing omissions.
[0015] In another possible implementation, isolation conditions are injected into the WHERE clause of the target node based on the tenant identity, resulting in a modified abstract syntax tree. This includes: constructing a tenant isolation condition node based on the tenant identity and a valid alias; checking the attribute status of the WHERE clause in the target node; if the attribute status of the WHERE clause indicates the existence of the original condition, enclosing the original condition in parentheses to obtain a bracket node; constructing an AND node to obtain a new WHERE clause, where the left child of the AND node is the bracket node and the right child of the AND node is the tenant isolation condition node; and replacing the original WHERE clause with the new WHERE clause to obtain the modified abstract syntax tree.
[0016] Based on this scheme, a "logic grafting" algorithm is used to construct binary expressions, and the isolation conditions are encapsulated with the original conditions using bracket encapsulation technology to prevent logical escape. While adhering to SQL syntax standards, this approach balances security and reliability, enabling seamless injection and non-destructive mutation rewriting of tenant isolation conditions.
[0017] Another possible implementation includes: when the attribute state of the WHERE clause is that the original condition does not exist, assign the tenant isolation condition node to the WHERE clause to obtain the modified abstract syntax tree.
[0018] Based on this scheme, the constructed tenant isolation condition node is assigned to the WHERE attribute of the target node, which is equivalent to adding tenant filtering conditions directly to a query without conditions, thereby obtaining the modified AST.
[0019] Another possible implementation includes filtering and reorganizing the physical tool list based on tenant identity to generate a list of tools available to the tenant.
[0020] Based on this solution, the physical tool list can be dynamically filtered and permission-based, retaining only the tools that the tenant has authorized to call, thereby generating a list of tools that the tenant can call.
[0021] In another possible implementation, the following is also included: if the request includes a resource discovery request, generating a virtual view of available tools based on the list of tools available to the tenant, which hides the tools available to other tenants from the current tenant's client; and returning the virtual view to the client.
[0022] Based on this solution, resource discovery requests can be used to map tenant authorization policies to the entire set of physical resources in real time, generating a virtual view containing only the tools that tenants can access, thereby achieving visibility isolation.
[0023] Another possible implementation includes: intercepting the client's initial network connection request; upgrading the initial network connection to a long-lived connection; extracting the tenant identity corresponding to the initial network connection request; creating a session, with the session object holding a reference to the long-lived connection object; binding the tenant identity to the session; during the session's existence, the binding relationship between the tenant identity and the session is based on the initial binding at the time the session is established, and the binding relationship remains valid and persists throughout the entire lifecycle of the session.
[0024] Based on this solution, by creating and maintaining a session context that spans the entire request lifecycle, a strong binding between tenant identity and network connection is achieved, enabling all subsequent security policies to be executed transparently, automatically, and accurately based on tenant identity.
[0025] Secondly, embodiments of this application provide a multi-tenant data isolation gateway for a large language model. Applying the method provided by any possible implementation of the first aspect above, the multi-tenant data isolation gateway for a large language model includes: a multi-tenant connection module configured to: intercept client requests and parse the tenant identity identifier corresponding to the request; a semantic isolation execution module configured to: inject isolation conditions into the request based on the tenant identity identifier when the request includes a tool call request, thereby obtaining a rewritten request, wherein the isolation conditions include the data range of the tenant's access restricted based on the tenant identity identifier; and the multi-tenant connection module further configured to: forward the rewritten request to the server.
[0026] In one possible implementation, the virtual resource mapping module is configured to filter and reorganize the physical tool list based on the tenant's identity to generate a list of tools available to the tenant.
[0027] In another possible implementation, the virtual resource mapping module is also configured to: generate a virtual view of available tools based on the list of tools available to the tenant, which hides the tools available to other tenants from the current tenant's client, and return the virtual view to the client when the request includes a resource discovery request.
[0028] Thirdly, embodiments of this application also provide a computing device, including: a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; the processor is used to execute the program instructions to perform the method as described in any of the first aspects above.
[0029] Fourthly, embodiments of this application provide a chip for performing the methods described in any of the first aspects above.
[0030] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a computer, implement the method as described in any of the first aspects.
[0031] In a sixth aspect, embodiments of this application provide a program product including a computer program that, when executed by a processor, implements the method as described in any of the first aspects. Attached Figure Description
[0032] Figure 1 This application provides a first flowchart illustrating a multi-tenant data isolation method for a large language model, based on some embodiments thereof. Figure 2 A schematic diagram of a multi-tenant data isolation system architecture for a large language model is provided for some embodiments of this application; Figure 3A second flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application; Figure 4 A third flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application; Figure 5 A fourth flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application; Figure 6 A fifth flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application; Figure 7 A sixth flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application; Figure 8 The seventh flowchart of a multi-tenant data isolation method for a large language model provided in some embodiments of this application; Figure 9 An architecture diagram of a multi-tenant data isolation gateway for a large language model is provided for some embodiments of this application; Figure 10 This is a schematic diagram of a computing device provided for some embodiments of this application. Detailed Implementation
[0033] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. To facilitate a clear description of the technical solutions of the embodiments of this application, the use of terms such as "first," "second," etc., in the embodiments of this application is for illustrative purposes and to distinguish the objects being described. There is no particular order between them, nor does it indicate a specific limitation on the number of devices in the embodiments of this application, and they do not constitute any limitation on the embodiments of this application.
[0034] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0035] It should be noted that many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below.
[0036] The following explanations of the technical terms mentioned in the embodiments of this application are provided to facilitate understanding by those skilled in the art.
[0037] Model Context Protocol (MCP): An open protocol for standardizing communication between large language model clients and external data sources or tool servers. MCP is typically built on top of the mature JSON-RPC message format, ensuring efficient and reliable function calls and data exchange between different systems through clearly defined interface specifications and interaction flows.
[0038] MCP Client: The party initiating the request, typically the large language model host environment (such as Claude Desktop, IDE plugins, AI Agent runtime, etc.). The core responsibility of the MCP Client is to parse user intent, construct standardized MCP requests, and securely distribute them to the corresponding MCP server; it is also responsible for receiving, verifying, and parsing responses from the server, ultimately presenting the structured results to the user or handing them over to the subsequent inference module for processing.
[0039] MCP Server: The party that provides resources, which usually includes the specific execution logic (such as database connectors, file system accessors, API wrappers, etc.). The MCP Server can receive and verify the requests sent by the MCP Client.
[0040] Abstract Syntax Tree (AST): A tree-like data structure used to describe the syntactic structure of program source code. Its nodes represent syntactic components in the program (such as expressions, statements, declarations, etc.), and edges represent the nesting and compositional relationships between components. AST strips away irrelevant lexical details (such as whitespace and bracket style), retaining only the key semantic structure. It is the core intermediate representation for compiler front-ends and AI code understanding systems to perform semantic analysis, static checking, automatic refactoring, and cross-language conversion.
[0041] Multi-tenancy: A software architecture design approach that allows multiple independent software instances to run together on the same physical hardware device, thereby making full use of the underlying computing, storage and network resources.
[0042] Prompt injection attacks refer to attacks where attackers maliciously construct input text to induce large language models to deviate from their preset behavioral boundaries, executing unexpected commands or leaking sensitive system information. These attacks often exploit the model's over-reliance on natural language, concealing instructions in user prompts, obfuscating context, or hijacking thought processes to bypass security mechanisms. In the MCP architecture, if the client does not rigorously validate and sandbox the incoming prompt, it may be used to manipulate server-side call logic, leading to unauthorized access or data leakage.
[0043] The multi-tenant data isolation method for large language models provided in this application will be described below with reference to the accompanying drawings.
[0044] Figure 1 This is a first flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application.
[0045] Figure 2 This is a schematic diagram of a multi-tenant data isolation system architecture for a large language model, provided for some embodiments of this application.
[0046] Combination Figure 1 and Figure 2 As shown, in some embodiments, the multi-tenant data isolation method for large language models includes: Step S110: Intercept the client's request.
[0047] For example, the client can be an MCP Client. Requests from an MCP Client can include Discovery Requests and Execution Requests.
[0048] Resource discovery requests typically occur in the initial stage (discovery phase) of the interaction between the MCP Client and the MCP Server. Their purpose is to allow the MCP Client to proactively inquire about or "discover" what tools or resources the MCP Server offers.
[0049] In some examples, resource discovery requests can be tools / list requests, used to query a list of currently available MCPServer resources.
[0050] Tool invocation requests typically occur during the execution phase. Their purpose is to allow the MCP Client to request the MCP Server to perform a specific tool operation. For example, invoking the "Database Query" tool to run an SQL query, or invoking the "File Reader" tool to read a specific file.
[0051] In some examples, a tool call request can be a tools / call request, which requests the MCP Server to perform a specific tool operation.
[0052] For example, a request interception middleware can be integrated between the MCP Client and the MCP Server. This allows the middleware to intercept client requests, preventing unauthorized Prompt injection attacks from directly reaching the server, thereby improving system security and the strength of data isolation between tenants.
[0053] Step S120: Parse the tenant identity identifier corresponding to the request.
[0054] For example, a tenant identity can be a tenant ID, such as TenantID.
[0055] Alternatively, the tenant identity identifier may also include other identity information of the tenant, such as the tenant name, authentication token, API key, or claim field embedded in the JWT payload, which is not limited in this embodiment.
[0056] In this embodiment, after intercepting the client's request, the tenant identity identifier corresponding to the request can be parsed first, which makes it easier to determine the relevant permissions corresponding to the request and match the preset access policy and data isolation rules based on the tenant ID.
[0057] In some embodiments, an access and session layer can be set up between the MCP Client and the MCP Server. Steps S110 and S120 described above can be implemented in the access and session layer, thereby enabling the interception and processing of requests from the MCP Client to achieve security isolation, authentication, and other functions.
[0058] For example, a Multi-tenant Connection Manager (MCM) can be deployed in the access and session layer. The MCM can act as the communication hub of the system, intercepting and managing requests from MCP Clients, thereby providing basic support for multi-tenant security isolation.
[0059] Step S130: In the case of requests including tool call requests, the request is rewritten based on the tenant identity to inject isolation conditions and obtain the rewritten request.
[0060] For example, when a request includes a tool call request, the request can be rewritten based on the tenant's identity to inject isolation conditions, thereby adding constraints to the request, achieving deep cleaning of the data plane, and preventing tenants from accessing sensitive data or performing unauthorized operations.
[0061] For example, isolation conditions can be the range of data that a tenant is restricted to access based on their tenant identity, or the range of data that a tenant cannot access based on their tenant ID. In other words, isolation conditions can be used to restrict tenants and prevent them from accessing other tenants' data without authorization, thereby achieving data isolation between tenants.
[0062] In some examples, step S130 above can be performed at the semantic execution layer of the system to achieve semantic-level rewriting, thereby achieving row-level data isolation.
[0063] For example, the semantic execution layer can deploy a Semantic Isolation Engine (SIE). The SIE can dynamically parse the MCP Client's requests based on the tenant's identity and inject tenant-specific access constraint logic, thereby enabling row-level filtering and field-level anonymization of requests at the data level, ensuring that sensitive fields are only visible to authorized tenants.
[0064] Step S140: Forward the rewritten request to the server.
[0065] For example, after receiving the rewritten request, the system can forward the request with isolation conditions to the server, thereby preventing the request from accessing the wrong domain and limiting the client's tool calls to the data boundary to which the tenant belongs.
[0066] In some examples, after the semantic execution layer of the system rewrites the request, the rewritten request can be forwarded to the access and session layer first, and then the access and session layer forwards the rewritten request to the server to ensure the consistency and traceability of the tenant isolation strategy across the entire chain.
[0067] The multi-tenant data isolation method for large language models provided in this application effectively blocks the risk of cross-tenant data leakage caused by Prompt injection or oversights in permission configuration by completing tenant identification and dynamic limitation of access scope in the three stages of interception, parsing, and rewriting before the request enters the server to execute logic. Compared with the traditional solution of isolating different tenants through independent containers, the data isolation method of this application does not rely on underlying resource isolation. It can achieve millisecond-level, fine-grained data boundary control in a shared model service architecture by only using semantic layer request rewriting and policy injection, which significantly reduces the complexity of operation and maintenance and resource overhead. In addition, compared with the solution of achieving isolation by modifying the hard logic of server code, the data isolation method of this application is less invasive, more secure, and has stronger portability and policy flexibility, and can be adapted to different model service frameworks and API system forms.
[0068] It is understandable that, in other implementations, the multi-tenant data isolation system of the large language model described above can be deployed in a gateway or implemented through WebAssembly (Wasm). In this embodiment, it is not limited to a gateway system, but is only used as an example for illustration.
[0069] It should be noted that the multi-tenant data isolation method for large language models provided in this application embodiment can also be applied to the OpenAI Plugin protocol or Agent communication protocol. In this embodiment, it is not limited to the MCP protocol, but is only used as an example for illustration.
[0070] Figure 3 This is a second flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application.
[0071] Combination Figure 3 As shown, in some embodiments, the multi-tenant data isolation method for large language models provided in this application further includes: Step S151: Intercept the client's initial network connection request.
[0072] For example, a connection manager can be deployed in the access and session layer to receive and distribute TCP / HTTP connection requests from different tenants.
[0073] For example, at the access and session layer, when the connection manager receives an initial HTTPS connection request from the MCP Client, the connection manager can process the connection request.
[0074] Step S152: Upgrade the initial network connection to a persistent connection.
[0075] For example, the connection manager can upgrade a one-time HTTP connection into a persistent, full-duplex (WebSocket) or one-way (SSE) long-lived connection to establish a communication channel that conforms to the MCP protocol specification. In this way, the MCP client and the gateway can continuously exchange multiple requests and responses (such as subsequent tools / list and tools / call) over this single connection, which is the basis for MCP protocol interaction.
[0076] For example, the connection manager can manage persistent connections, enabling connection reuse through functions such as establishing, maintaining, and reusing persistent connections. This avoids deploying and maintaining a separate physical connection to the backend service for each tenant, thereby saving network and computing resources between the gateway and the backend service and reducing operational complexity.
[0077] Step S153: Extract the tenant identity identifier corresponding to the initial network connection request.
[0078] For example, in the connection establishment or the first data frame, the MCP Client typically carries identity credentials (such as a JWT token). The session context builder intercepts and parses the JWT, extracting the crucial Tenant ID. This enables the authentication of the MCP Client and obtains the tenant identity upon which all subsequent security and isolation policies rely, which is a primary prerequisite for achieving multi-tenant isolation.
[0079] Step S154: Create a session.
[0080] The session object holds a reference to the long-connection object.
[0081] For example, the session context builder can create a session object in memory. This session object internally holds a reference to the persistent connection. In this way, the physical connection at the network layer can be abstracted into a logical session at the business layer. This session object serves as a unified context carrier for the flow of this connection throughout the gateway, allowing subsequent processing components to easily obtain connection and identity information without having to repeatedly parse the original request.
[0082] Step S155: Bind the tenant identity to the session.
[0083] For example, binding a tenant identity to a session enables a permanent binding of the tenant identity to a long-lived connection. Subsequently, all business operations performed through this long-lived connection (such as resource discovery and tool invocation) automatically "carry" the tenant identity, providing directly usable identity credentials for subsequent VRM and SIE processing, thereby achieving "session coloring".
[0084] In some embodiments, after the tenant identity is bound to the long-lived connection, the session state can be synchronized to the Redis cache to achieve state sharing and consistency in a distributed environment.
[0085] Step S156: During the session, the binding relationship between the tenant identity and the session is based on the initial binding at the time the session is established, and the binding relationship remains valid and lasts throughout the entire lifecycle of the session.
[0086] The session duration refers to the complete time period from when the gateway receives the client's initial connection request, completes authentication, and successfully creates the session object, until the session is destroyed due to timeout, client-initiated disconnection, or server-forced termination. During this period, the session object remains active and can continuously receive and process various requests sent by the client.
[0087] Initial binding refers to the binding process that occurs when a session is first accessed in step S155.
[0088] The entire lifecycle refers to all processing stages a session undergoes from creation to destruction, including but not limited to the tool discovery stage, tool invocation stage, multi-turn dialogue stage, and resource access stage. In each of these stages, the gateway reads the same tenant identity from the session object when processing each client request to ensure that isolation policies at each stage are executed based on the same identity context.
[0089] For example, after completing session creation in step S155, the gateway writes the extracted tenant identity identifier into the identity field of the session object and sets a write-protection flag for this field. During the session, each time the gateway's request processing module processes a request for that session, it reads the identity field from the session object to obtain the tenant identity identifier, without needing to re-parse or verify the identity information from the request message. In this way, the isolation conditions in subsequent stages are all generated based on the same identity context, avoiding inconsistencies in the isolation strategy caused by changes in identity information between multiple requests. At the same time, because the identity field is write-protected, even if a man-in-the-middle attack or malicious request can tamper with the identity declaration in the request message, it cannot modify the tenant identity identifier already embedded in the session object. This builds an identity verification barrier independent of the transport layer at the session level, ensuring the continuous reliability of data isolation during multiple rounds of interaction.
[0090] In some examples, when the gateway detects an attempt to modify the session identity field, it can refuse to execute the operation and log the exception. The multi-tenant data isolation method for large language models provided in this application establishes and manages long connections to meet the real-time interaction requirements of the MCP protocol, and can reuse backend connections through connection pooling to improve resource efficiency. Furthermore, identity authentication and extraction are completed at the beginning of communication, laying the foundation for the entire multi-tenant security system. Simultaneously, by creating and maintaining a session context that spans the request lifecycle, a strong binding between tenant identity and network connection is achieved, enabling all subsequent security policies to be executed transparently, automatically, and accurately based on tenant identity.
[0091] Figure 4 This is a third flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application.
[0092] Combination Figure 2 and Figure 4 As shown, in some embodiments, the multi-tenant data isolation method for large language models provided in this application further includes: Step S160: Filter and reorganize the physical tool list based on the tenant's identity to generate a tool list available to the tenant.
[0093] For example, after intercepting the MCP Client's request, the physical tool list can be dynamically filtered and permission-trimmed based on the tenant's identity, retaining only the tools that the tenant has authorized to call, thereby generating a list of tools that the tenant can call.
[0094] In some embodiments, among the tool services provided by the MCP Server, the toolbox that can be authorized to the current tenant is often a tool item with a low probability of data leakage. For example, this tool item may only store the current tenant's data. Other unauthorized tool items are often tool items with a higher probability of data leakage. For example, this tool item may store data from multiple tenants.
[0095] After obtaining the list of tools available to a tenant, the corresponding controlled table for that tenant can be obtained. The controlled table is a set of target physical tables that require security processing, storing data shared by multiple tenants. To prevent data obfuscation and leakage caused by tenants exceeding their privileges and calling unauthorized tools, the system determines the controlled table simultaneously with the tool list. This facilitates subsequent request rewriting based on the controlled table during the request execution phase.
[0096] In some embodiments, step S160 described above can be performed at the policy control layer, which can be located between the access and session layer and the semantic execution layer. In this way, after the access and session layer intercepts the request and identifies the tenant, the policy control layer can determine the list of tools available to the tenant based on the tenant's identity and simultaneously generate a controlled table mapping relationship bound to it, ensuring that the semantic execution layer can accurately inject tenant-specific security isolation conditions when rewriting requests.
[0097] For example, the policy control layer may be equipped with a policy decision engine to implement the above step S160.
[0098] In some embodiments, the tool list and controlled table obtained in step S160 above can be transmitted to the semantic execution layer to provide a reference for subsequent target node lookup.
[0099] Please continue reading. Figure 4 The multi-tenant data isolation method for large language models provided in this application embodiment also includes: Step S170: In the case of a request that includes a resource discovery request, generate a virtual view of available tools based on the list of tools available to the tenant, which hides the tools available to other tenants from the current tenant's client.
[0100] For example, if a tenant initiates a resource discovery request, such as "list the data tools I can use," the system will only return the tool metadata after permission-trimmed settings, hiding the existence of unauthorized tools and thus generating a virtual view. The virtual view does not expose the underlying physical structure, nor does it disclose other tenants' tool configurations or data distribution information. In this way, the virtual view can shield the tools available to other tenants from the current tenant's client, thereby achieving visibility isolation and further strengthening the data boundaries between tenants. The virtual view only presents the logical set of tools from the tenant's perspective, and its metadata is anonymized and generalized, eliminating the possibility of inferring the shared table structure from tool names, descriptions, or parameters.
[0101] In some examples, the virtual view can be embedded with a tenant context identifier during the generation process, ensuring that all subsequent tool call requests carry an immutable tenant identity identifier. The dynamic permission mapping between the tenant identity identifier and the controlled table is linked in real time, automatically triggering field-level data filtering and row-level access control during the execution phase, thereby building a dual isolation barrier between the logical view and physical execution.
[0102] Step S180: Return the virtual view to the client.
[0103] For example, after obtaining the virtual view, the system can return the virtual view to the client for the tenant to use.
[0104] In some embodiments, steps S170 and S180 described above can be performed at the policy control layer.
[0105] For example, the policy control layer can be deployed with a Virtual Resource Mapper (VRM). The VRM enables resource visibility isolation during the tool discovery phase.
[0106] For example, the VRM can receive a complete list of physical resources to determine all the tool metadata that the MCP Server can provide.
[0107] In addition, the VRM can receive authorization policies from each tenant to determine the subset of tools that each tenant can access, and dynamically build virtual view metadata based on the policies.
[0108] For resource discovery requests, VRM can calculate the intersection of the tenant authorization policy corresponding to the request and the complete list of physical resources, thereby generating a new, filtered list of tools.
[0109] For the filtered tool list, VRM can generate a virtual view based on the tool list for tenants to view. VRM can construct a data response message containing the virtual view based on the JSON-RPC protocol format to achieve visibility isolation and form a data carrier returned to the client.
[0110] In the multi-tenant data isolation method for large language models provided in this application embodiment, based on resource discovery requests, the tenant authorization policy can be mapped to the entire set of physical resources in real time to generate a virtual view containing only the tools that can be accessed, thereby achieving visibility isolation.
[0111] Figure 5 This is a fourth flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application.
[0112] Combination Figure 1 , Figure 2 and Figure 5 As shown, in some embodiments, step S130 above can be implemented by the following method: Step S131: In the case of requests including tool call requests, parse the SQL string in the request to obtain an abstract syntax tree.
[0113] For example, because the Abstract Syntax Tree (AST) accurately reflects the semantic structure of SQL, the data access object, filtering conditions, and operation intent can be identified based on the AST node types and contextual relationships. This facilitates locating the position where injection isolation conditions need to be applied, thereby improving the efficiency and accuracy of rewriting.
[0114] For example, the semantic execution layer of the gateway can call the semantic processing interface to perform the above step S131.
[0115] For example, the semantic processing interface can be the Go language interface of ExecutionEngine, which may specifically include: Etype ExecutionEngine interface { / / Parse: String -> AST Parse(sql string) (ASTNode, error) / / Process: Integrating search and rewrite logic / / Input: AST root node, tenant ID / / Output: Modified AST Process(root ASTNode, tenantID string) (ASTNode, error) } The `Etype ExecutionEngine interface {` is used to define an interface type named `ExecutionEngine` to declare an interface that specifies which methods the execution engine must implement. The design of this interface balances abstraction and extensibility.
[0116] The `Parse(sql string) (ASTNode, error)` method takes an SQL string as input and returns an `ASTNode` and an `error`. The `Parse` method performs semantic parsing, transforming the unstructured raw SQL command string into a structured abstract syntax tree. If the SQL syntax is incorrect, an error message is returned via the `error` parameter; if parsing is successful, the root node of the AST (`ASTNode`) is returned, providing a foundation for subsequent deep analysis.
[0117] The `Process(root ASTNode, tenantID string) (ASTNode, error)` method takes an AST root node and a tenant ID string as input and returns the modified AST node and an error. This is the core flow of the security processing engine, integrating "lookup" and "rewrite" logic. Input: Receive the parsed AST and the identity of the current tenant.
[0118] Internal processing: Traverse the AST, locate the target node that needs to be isolated, and construct security constraints based on tenantID.
[0119] Output: Returns a new AST that has been safely injected with tenant isolation conditions (e.g., AND org_id='A'). If the procedure fails, it returns an error.
[0120] In some examples, if parsing fails during the process of converting the SQL string in the JSON-RPC request into an abstract syntax tree, the request can be circuit-broken to prevent the tenant from accessing the service.
[0121] In some embodiments, the semantic execution layer may be provided with an AST parser to implement the above step S131. For related explanations, please refer to the description of step S131, which will not be repeated here.
[0122] Step S132: Traverse the abstract syntax tree and inject isolation conditions into the WHERE clause of the abstract syntax tree based on the tenant identity to obtain the modified abstract syntax tree.
[0123] For example, when a SQL query is initiated by a tool invocation class request, the AST of the query statement (such as SELECT) typically contains various types of clause nodes. An AST parsed from a standard SQL query statement typically contains corresponding nodes for the following common clauses: SELECT clause: Contains a list of columns or expressions to query.
[0124] The FROM clause contains the data source for the query, such as a table, subquery, or JOIN expression.
[0125] The JOIN clause is part of the FROM clause and describes in detail the connection method and conditions between tables.
[0126] The WHERE clause defines the data filtering conditions.
[0127] In addition, AST may also include GROUP BY clause, HAVING clause, ORDER BY clause, LIMIT / OFFSET clause, etc., and no specific restrictions are imposed in this embodiment.
[0128] The WHERE clause is a native, standard position in the SQL language designed for row-level filtering. Injecting isolation conditions into the WHERE clause allows for seamless, secure, and unambiguous row-level data isolation, thus adhering to SQL semantics and ensuring its correctness and reliability.
[0129] For example, using a traversal approach can ensure that isolation conditions are embedded in all data filtering paths, avoiding the omission of conditions due to multi-level nested queries or union queries.
[0130] Step S133: Restore the modified abstract syntax tree to an SQL string to obtain the rewritten request.
[0131] For example, after rewriting the abstract syntax tree, in order to enable normal interaction between MCP Clients and MCP Servers, the rewritten abstract syntax tree needs to be reserialized into a string format conforming to the SQL standard to ensure syntactic validity and execution compatibility. In this way, the rewritten SQL string will carry the tenant isolation conditions, thereby automatically completing tenant data isolation at the database execution layer without modifying business logic or relying on the application layer to manually concatenate conditions. This mechanism not only ensures data security boundaries in multi-tenant scenarios but also maintains the integrity and readability of the original SQL semantics, while avoiding logical vulnerabilities or performance losses caused by incorrect condition concatenation in traditional solutions.
[0132] In the multi-tenant data isolation method for large language models provided in this embodiment, SQL strings are parsed into an Abstract Syntax Tree (AST) to facilitate the search and identification of WHERE clause nodes. Simultaneously, the WHERE clause nodes are rewritten within the AST to embed isolation conditions. While conforming to SQL language specifications, secure injection of isolation conditions is achieved, thus realizing "row-level data isolation." The rewritten AST is decompiled into a standard SQL string, preserving the semantic structure and execution logic of the original query while embedding unavoidable constraints at the syntax layer. This effectively prevents Prompt injection attacks and ensures that any query path is forcibly constrained by the tenant context. This precise injection mechanism based on a syntax tree does not rely on regular expression matching or string concatenation, fundamentally avoiding semantic misjudgment and escape risks.
[0133] In some embodiments, the SIE may be equipped with a decompiler to implement step S133 described above.
[0134] Figure 6 This is a fifth flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application.
[0135] Combination Figure 2 and Figure 6 As shown, in some embodiments, step S132 above can be implemented by the following method: Step S1321: Traverse the abstract syntax tree by depth-first search to find the target node in the abstract syntax tree and obtain the valid alias of the target node.
[0136] The target nodes include nodes that reference controlled tables, which are a set of target physical tables that need to be isolated from data.
[0137] For example, Depth-First Search (DFS) is an algorithm used to traverse and search data structures such as trees and graphs. The traversal strategy of DFS is to start from the root node, explore as deeply as possible along a branch, until a leaf node is reached or further expansion is not possible, then backtrack to the nearest unvisited node and continue exploring the remaining branches.
[0138] As a tree-like data structure, AST (Abstract Syntax Tree) supports structures such as nested queries (sub-queries) and multi-table joins (JOINs). The Depth-First Search (DFS) algorithm can delve deep into subqueries and automatically backtrack to the outer context after completing the subquery traversal to match the potentially multi-level nested syntax of SQL statements. In this way, the DFS algorithm can traverse all nodes in the AST, even if the target physical table corresponding to the target node is hidden in a deep nested query or JOIN structure, it can still be traversed and discovered, thus avoiding omissions and locating all target nodes referencing controlled tables.
[0139] For example, the final alias of the target node can be a string. The final alias represents the exact name of the target physical table referenced by the target node in the current SQL fragment. Obtaining the final alias of the target node is crucial for constructing the correct isolation conditions.
[0140] For example, if you need to generate an isolation condition such as o.org_id = 'A', you need to ensure that the 'o' in the isolation condition is consistent with the name used when other parts of the query reference the table, in order to avoid the SQL failing to execute due to syntax errors.
[0141] For example, the context information of the target node can be ParentContext or ParentCtx, which is a pointer to the SELECT statement node (SelectStmt) to which the current target node belongs. The context information of the target node is the basis for determining the location of the injection isolation condition.
[0142] For example, when encountering nested subqueries, the inner subquery and the outer main query are different SELECT statement contexts. When injecting isolation conditions, it is necessary to accurately inject the isolation conditions into the SELECT statement node containing the target node, rather than incorrectly injecting them into the parent or child context; otherwise, the isolation logic will fail or the SQL syntax will be abnormal.
[0143] In some embodiments, a controlled target recursive finder may be set in the SIE. This controlled target recursive finder can be used to implement the above step S1321, thereby recursively traversing the AST nodes, automatically identifying nested JOINs, subqueries and alias links, accurately locating the target node, and calculating a valid alias for the target node to provide structured input for the subsequent rewriting stage.
[0144] For example, a controlled target recursive finder can determine the valid alias and context information of a target node using the following methods: Injection point definition (the intermediate object connecting the controlled target recursive finder and the rewriter) type InjectionTarget struct { NodePtr interface{} / / AST node pointer FinalAlias string / / Calculated alias ParentCtx interface{} / / Parent context } The `type InjectionTarget struct {` defines a structure type named `InjectionTarget`. `InjectionTarget` means injection target; defining it as a structure allows it to encapsulate and transmit all the necessary information about a complete target node to be processed between the program's lookup and rewrite phases, acting as a data carrier.
[0145] `NodePtr interface{}` / / AST node pointer. This refers to the `NodePtr` field, of type `interface{}`, which represents an AST node pointer. This field stores a reference (pointer) to a specific node in the abstract syntax tree, that is, which syntax node in the AST the current "injection target" corresponds to (e.g., a `TableSourceNode` representing `FROMorders`). During rewriting, the rewriter needs to use this pointer to locate the specific object to be operated on.
[0146] FinalAlias string / / The calculated alias, referring to the FinalAlias field, is of type string and represents the calculated alias. This field stores a string, which is the valid reference name of the target physical table in the current SQL context. The calculation rule is: if an alias is defined for the table in the SQL (such as 'o'), then that alias is used; otherwise, the table name itself is used (such as 'orders'). This is crucial for constructing correct isolation conditions, ensuring that conditions like 'o.org_id='A' can be correctly associated with other parts of the query.
[0147] ParentCtx interface{} / / Parent context, referring to the ParentCtx field, of type interface{}, representing the parent context. This field stores a pointer to the parent syntax structure (usually the SELECT statement node) to which the current target node belongs. It defines the scope of "overriding," telling the rewriter which level and which query statement's WHERE clause should be modified to inject conditions, which is crucial for correctly handling nested subqueries.
[0148] In some embodiments, the controlled target recursive finder can implement step S1321 above using the following data structure: `package security`: This declares a package named `security`. In Go, the `package` keyword is used to organize code into modules. Here, the relevant structs are defined within the `security` package, indicating that they are components of the gateway security subsystem.
[0149] Data Structure 1: / / FinderConfig Controlled Target Recursive Finder Configuration type FinderConfig struct { / / Controlled list of confessions (Map structure for faster lookup) ProtectedTables map[string]bool } The Controlled Target Recursive Finder Configuration is used to configure the behavior of the Controlled Target Recursive Finder when the algorithm starts, and to provide the Controlled Target Recursive Finder with the strategy data required at runtime.
[0150] FinderConfig is a configuration container. When initializing a controlled target recursive finder instance, a whitelist of controlled tables of this type needs to be passed in to inform the controlled target recursive finder which tables are controlled tables to be monitored.
[0151] The `protectedTables map[string]bool` field represents the `ProtectedTables` field, and its type is `map[string]bool`. This is the core configuration driving the entire lookup process. It is a key-value map. The key is a string representing the physical table name (e.g., "orders"). The value is a boolean, usually true, indicating that the table is on the controlled list. This map acts as a "whitelist" of controlled tables for high-speed queries. When the controlled target recursive lookup traverses the AST and encounters a table, it can quickly query this map by table name to determine whether the table is a sensitive data table that needs to be located and processed. Sensitive data tables are the target physical tables that need data isolation. For example, `{"orders": true, "customers": true}` indicates that the `orders` and `customers` tables need row-level isolation.
[0152] Data Structure 2: / / InjectionTarget: Inject the target object / / Describes the location in the AST and the specific table alias to target for injection. type InjectionTarget struct { Node ASTNode / / Pointer to the hit AST node ParentContext SelectStmt / / The main statement context to which this node belongs FinalAlias string / / Core field: Valid alias (e.g., "o" or "orders") } The InjectionTarget structure is the output of the controlled target recursive finder algorithm, serving as a standardized data contract connecting the "find" and "rewrite" phases. Each InjectionTarget object precisely describes a target node to be processed.
[0153] type InjectionTarget struct { Defines a struct type named InjectionTarget.
[0154] In the Node ASTNode, the Node field, of type ASTNode, stores a reference to a specific AST node. It points to the syntax node (usually a TableSourceNode) that references the controlled table and is "hit" by the controlled target recursive lookup. It provides the precise coordinates of the modifications made in the syntax tree.
[0155] ParentContext In SelectStmt, the ParentContext field is of type... The `SelectStmt` field stores a pointer to the parent `SELECT` statement node, defining the scope of condition injection. Because the `WHERE` clause of an SQL statement is part of the `SELECT` statement, the rewriter must know which `SELECT` statement's `WHERE` condition (which could be the outer main query or the inner subquery) should be modified. This field ensures that, in nested queries, the condition is injected into the correct query level.
[0156] In the `FinalAlias string` directive, the `FinalAlias` field, of type string, stores the calculated valid alias. In SQL, a table may be assigned an alias (e.g., `orders AS o`), and subsequent conditions use the alias to reference the column (`o.id`). The rule for determining this field is: Aliases defined in the SQL are preferred; if not defined, the `TableName` itself is used. The rewriter will use this alias to construct the correct isolation conditions (e.g., `o.org_id = 'A'`), ensuring semantic consistency with other parts of the query.
[0157] Data Structure 3: / / TableSourceNode (the element in the FROM clause) type TableSourceNode struct { TableName string / / Physical table name Alias string / / Alias defined in SQL (may be empty) } TableSourceNode is a node type in the AST that represents the source of a base table within the FROM clause. One of the primary goals of a controlled target recursive lookup algorithm is to locate nodes of this type.
[0158] type TableSourceNode struct { Defines a structure type named TableSourceNode.
[0159] In `TableName string`, the `TableName` field (physical table name) is of type string. The `TableName` field records the physical table name in the database. The controlled target recursive finder compares this value with the key of `FinderConfig.ProtectedTables` to determine if the node is a target.
[0160] In the `Alias` string, the `Alias` field (alias) is of type string. The `Alias` field records the aliases explicitly defined for this table in the SQL statement. If no alias is defined, this field is an empty string (""). This field is the direct input for calculating `InjectionTarget.FinalAlias`.
[0161] Data Structure 4: / / JoinNode: Connects nodes (JOIN clause) type JoinNode struct { Left ASTNode / / Left table Right ASTNode / / Right table On Expr / / Connection condition } Here, JoinNode is the node type in the AST that represents the JOIN operation. Since JOIN can complicate the query structure, a controlled target recursive searcher needs to be able to recursively process such nodes.
[0162] type JoinNode struct { Defines a structure type named JoinNode.
[0163] The fields Left ASTNode, Right ASTNode, and OnExpr collectively describe a join operation. The Left and Right fields point to the AST nodes that join the left and right halves of the table, respectively. These can be TableSourceNodes, another JoinNode, or even a subquery. This recursive definition allows a controlled-target recursive finder to handle complex multi-table joins by traversing them. The On field is an expression node representing the join condition (e.g., ON left.id = right.id). During the lookup phase, the algorithm primarily focuses on the table structure and may not delve into this condition, but it is fully preserved as part of the AST.
[0164] Of the four data structures mentioned above, FinderConfig is the input, defining "what to look for." TableSourceNode and JoinNode are two key node types that the algorithm needs to identify and process when traversing the AST, representing "where to look." InjectionTarget is the output, encapsulating complete information about the found target (location, alias, context), answering "how to describe it after finding it," and is directly passed to the rewriter in the next stage. These four data structures constitute the data foundation of the controlled target recursive finder algorithm.
[0165] Step S1322: Based on the tenant identity identifier and the valid alias of the target node, inject isolation conditions into the WHERE clause of the target node to obtain the modified abstract syntax tree.
[0166] For example, when injecting isolation conditions, the isolation policy needs to be determined based on the tenant's identity. This allows for the generation of different isolation conditions for different tenant requests, matching the access permissions of each tenant. This ensures that each tenant can only access its authorized subset of data, achieving true multi-tenant data isolation. The generation of isolation conditions can be based on a preset isolation policy configuration, which may include the mapping relationship between tenants and data scopes, field-level permission rules, and dynamic filtering expressions.
[0167] In this embodiment, the references to the target physical table that need to be isolated can be located by finding the target node. Furthermore, for each target node, the isolation condition can be determined by its valid alias. The context information corresponding to the target node can determine which SELECT statement the isolation condition should be injected into, thus precisely anchoring the insertion position of the WHERE clause. The combination of the target node, valid alias, and context information ensures the security, accuracy, and reliability of subsequent isolation condition injection. Even if a large model is induced by an attacker to generate a malicious full table scan command, the gateway can forcibly restrict it to the tenant scope, and the backend server does not need to be aware of the tenant logic.
[0168] Figure 7 This is a sixth flowchart illustrating a multi-tenant data isolation method for a large language model provided in some embodiments of this application.
[0169] Combination Figure 2 and Figure 7 As shown, in some embodiments, step S1321 can be implemented by the following method: Step a1: Input the root node of the abstract syntax tree and initialize the result set.
[0170] The result set includes a list of results used to store information about the found target nodes. In other words, the result set stores pointers to the target nodes, context information, and valid aliases, thus providing a container for the found target nodes to facilitate the retrieval of relevant information during subsequent rewriting.
[0171] For example, depth-first search can start from the root node of the abstract syntax tree of the SQL query, initialize the result set, and obtain an empty list of results (Targets). At the same time, the syntax context can be set during the initialization process to prepare for recursive traversal.
[0172] Step a2: Start the recursive traversal.
[0173] For example, after completing the initialization work, the recursive function Visit can be started to begin traversal and thus begin searching for the target node.
[0174] Step a3: Determine the node type of the abstract syntax tree.
[0175] For example, the node types of an abstract syntax tree can include query statement (SELECT) nodes, join (JOIN) nodes, and table source (TABLE) nodes.
[0176] In this context, the SELECT node represents a complete block of query statements. The SELECT node is the "container" of the query, containing clauses such as FROM, WHERE, and SELECT lists. In nested queries, one SELECT node can contain another SELECT node (a subquery).
[0177] The JOIN node represents a join operation between tables (such as INNER JOIN, LEFT JOIN). The JOIN node itself does not contain business data, but points to the joined table (or subquery) through the Left and Right attributes, and describes the join condition through the On attribute.
[0178] The TABLE node represents the most basic data source in a query, namely a physical table (such as FROM orders). The TABLE node contains TableName (physical table name) and optional Alias information. This is the leaf node that carries the actual data and is also a type of node that needs to be focused on when looking up the target node.
[0179] For example, when traversing the nodes of an abstract syntax tree, the traversal path can be determined based on the node type to implement a differentiated traversal strategy, thereby accurately finding each target node and preventing omissions.
[0180] Step a4: If the node type in the abstract syntax tree is a table source node, compare the current physical table corresponding to the current node with the target physical table in the controlled table to determine whether the current physical table is in the controlled table.
[0181] For example, if the node type in the abstract syntax tree is a table source node, it means that the current node is the leaf node that carries the actual data. In this case, it needs to be compared with the list of controlled tables one by one.
[0182] If the current physical table corresponding to the current node can match the target physical table in the controlled table, such as if the current physical table is one of the target physical tables, then it means that the current physical table is located in the controlled table.
[0183] If the current physical table corresponding to the current node cannot match the target physical table in the controlled table, such as if the current physical table is not one of the target physical tables, then it means that the current physical table is not in the controlled table.
[0184] Step a5: If the node type in the abstract syntax tree is a connection node, recursively visit the left and right subtrees.
[0185] For example, when the current node is a JOIN node, it indicates that a table join has been encountered. In this case, the recursive function can maintain the current context and recursively traverse its left and right subtrees respectively to ensure that the tables on both sides of the current join node are checked.
[0186] Step a6: If the node type in the abstract syntax tree is a query statement node, update the context and recursively visit the child nodes.
[0187] For example, when the current node is a SELECT node, it indicates that a new query block has been encountered. This query block can be a main query or a subquery. In this case, the recursive algorithm needs to update the current context to the SELECT node and then recursively traverse its FROM clause to explore the data source within the query block.
[0188] Step a7: If the current physical table is in a controlled table, determine the current node as the target node and determine the valid alias of the physical table referenced by the target node.
[0189] For example, if the current physical table is in a controlled table, it indicates that the current node is a node requiring close monitoring. The data in the physical table referenced by this node may be at risk of leakage, necessitating isolation condition injection. In this case, the current node can be identified as the target node, and the valid alias of the physical table referenced by the target node can be determined.
[0190] In some examples, the valid alias can be obtained from the Alias attribute of the TABLE node. If Alias is empty, TableName can be used as the default alias; if Alias is not empty, the valid alias is Alias. The valid alias can serve as a key identifier for subsequent permission context construction and dynamic SQL rewriting, ensuring that isolation conditions are accurately bound to the corresponding data source.
[0191] Furthermore, if the current physical table is not in the controlled table, it means that the current node does not need to be focused on and can be ignored, and the search can continue to the next node.
[0192] Step a8: Encapsulate the pointer and valid alias of the target node into structured data and record it in the result list.
[0193] For example, the pointer to the target node, valid aliases, and other information can be packaged into a structured data object InjectionTarget and stored in the result list for querying and use during subsequent isolation condition injection.
[0194] For example, the structured data object InjectionTarget may also include context information, etc., which is not limited in this embodiment.
[0195] In the multi-tenant data isolation method for large language models provided in this application, starting from the root node of the AST, a depth-first search strategy is used to transform unstructured, potentially security-risk-laden original SQL queries into a structured, explicit, and securely rewritten result set. This allows subsequent rewriting components to automatically inject isolation conditions based solely on this result set, without understanding the SQL business statements. This process decouples security policies from SQL parsing, significantly improving the maintainability and scalability of tenant isolation.
[0196] In some embodiments, steps a1-a8 described above can be implemented using a controlled target recursive finder.
[0197] For example, a controlled target recursive finder can implement the above process in the following way: / / RecursiveFinder Controlled Target Recursive Finder Implementation type RecursiveFinder struct { Config FinderConfig Targets []InjectionTarget } / / Visit recursive access function func (f RecursiveFinder) Visit(node ASTNode, ctx SelectStmt) { if node == nil { return} switch n := node.(type) { / / Scenario A: Encountering a SELECT statement case SelectStmt: f.Visit(n.From, n) / / Update context / / Scenario B: Encountering a JOIN structure case JoinNode: f.Visit(n.Left, ctx) / / Preserve the context f.Visit(n.Right, ctx) / / Scenario C: Encountering a specific table node (the end of the recursion) case TableSourceNode: / / 1. Check the controlled table if f.Config.ProtectedTables[n.TableName] { / / 2. Calculate aliases effectiveAlias := n.TableName if n.Alias != "" { effectiveAlias = n.Alias } / / 3. Record the goal f.Targets = append(f.Targets, InjectionTarget{ Node: n, ParentContext: ctx, FinalAlias: effectiveAlias, }) } } } The above method defines a structure called RecursiveFinder, which represents the implementation of the "controlled target recursive finder".
[0198] For example, the `Config FinderConfig` field stores the finder's configuration. The `Targets[]InjectionTarget` field is a slice of type `InjectionTarget`. It serves as a result set, initially empty, used during recursive traversal to collect and store all found targets that need to be rewritten.
[0199] For example, func (f RecursiveFinder) Visit(node ASTNode, ctx `SelectStmt) {` defines a method of `RecursiveFinder` named `Visit`, which is a recursive function. The input parameters are: an AST node and a context pointer (ctx) to `SelectStmt` (the SELECT statement). This is the recursive entry point and core traversal function of the algorithm. Each call processes one AST node. The `ctx` parameter is the crucial "context," recording which SELECT statement's scope the currently processed node belongs to, which is essential for handling nested subqueries.
[0200] For example, `if node == nil { return}` means that if the passed-in node is nil (empty), then return directly.
[0201] For example, `switch n := node.(type) {` means using a switch statement to assert the type of the variable `node`. This allows the code to jump to different case branches and execute the corresponding logic based on the actual type of the variable `node`.
[0202] For example, in the above method, three scenarios are divided according to different nodes, and the corresponding execution logic is defined in each scenario, thereby realizing the above steps a4-a6.
[0203] For example, in scenario C, by checking the controlled table, calculating aliases, and recording targets, the target node can be determined and its related information can be obtained. Finally, the key information is encapsulated and output to form the final output of the finder.
[0204] In some embodiments, a predicate injection rewriter can be set in the SIE, and the above step S1322 can be implemented by the predicate injection rewriter.
[0205] Figure 8 is a schematic diagram of the seventh process of a multi-tenant data isolation method for a large language model provided in some embodiments of this application.
[0206] Combination Figure 2 and Figure 8 As shown, in some embodiments, step S1322 can be implemented by the following method: Step b1: Construct tenant isolation condition nodes based on tenant identity identifiers and valid aliases.
[0207] For example, the tenant identity identifier TenantID, the effective alias FinalAlias, and the context information together constitute the three core pieces of information in the tenant isolation condition node.
[0208] For example, the predicate injection rewriter can construct a binary expression based on the tenant identity identifier TenantID and valid alias in the above information: [Alias].[IsolationColumn] = [TenantID], and use this constraint predicate as a tenant isolation condition node, laying the foundation for the subsequent injection of isolation conditions.
[0209] Step b2: Check the attribute status of the WHERE clause in the target node.
[0210] For example, by checking the attribute status of the WHERE clause in the target node, it is possible to determine whether the original condition exists in the target node, thereby determining the specific grafting method of the tenant isolation condition node, that is, determining the specific injection method of the isolation condition, so as to correctly handle any form of SQL statement and ensure the correctness of SQL syntax.
[0211] Step b3: If the attribute status of the WHERE clause is that the original condition does not exist, assign the tenant isolation condition node to the WHERE clause to obtain the modified abstract syntax tree.
[0212] For example, assigning the constructed tenant isolation condition node to the WHERE attribute of the target node is equivalent to adding a tenant filter condition directly to a query without conditions, thereby obtaining the modified AST.
[0213] For example, SELECT FROM orders is rewritten as SELECT FROM orders WHERE o.org_id ='T-1001'.
[0214] Step b4: If the attribute status of the WHERE clause is that the original condition exists, wrap the original condition in parentheses to obtain the parenthesis node.
[0215] For example, by creating a new parenthesis node (ParenExpr), the entire original WHERE condition subtree (such as the node representing A OR B) can be used as its Expr child node, thus forming the (A OR B) structure. In this way, the parenthesis can be used to forcibly increase the computational priority of the original conditions, wrapping them into an indivisible logical whole, preventing logic escape attacks, and ensuring the safety of subsequent injection.
[0216] It's important to note that if the conditions are concatenated directly without enclosing them in parentheses, and the original conditions contain OR logic, the added AND isolation condition will be split by the OR due to operator precedence. This results in the isolation condition only constraining a portion of the branches rather than the whole, leading to a logic escape vulnerability. This solution uses parentheses to force the original conditions to be treated as a whole, ensuring that the constraint effect of the AND node covers all branches of the original conditions.
[0217] Step b5: Construct the node to obtain the new WHERE clause.
[0218] In this context, the left child of the AND node is the bracket node, and the right child of the AND node is the tenant isolation condition node.
[0219] For example, when injecting a tenant isolation condition node into a WHERE clause with existing conditions, a new binary expression node (BinaryExpr) can be created, with its operator (Op) set to "AND". Then, the left child of this new AND node points to the bracket node (i.e., (A OR B)), and its right child points to the tenant isolation condition node (i.e., o.org_id = 'T-1001'). This allows for a new WHERE clause that satisfies SQL syntax requirements without losing the original conditions.
[0220] Step b6: Replace the original WHERE clause with the new WHERE clause to obtain the modified abstract syntax tree.
[0221] For example, the newly grafted WHERE clause is reassigned back to the WHERE attribute of the target node, completing the AST reconstruction. In this way, the entire SQL statement achieves reliable row-level data isolation while remaining syntactically correct.
[0222] Step b7: Output the modified abstract syntax tree.
[0223] For example, after obtaining the modified abstract syntax tree, it can be used as the output of the predicate injection rewriter so that the abstract syntax tree can be restored to SQL statements and returned to the connection manager.
[0224] In this embodiment, a "logic grafting" algorithm is used to construct a binary expression, and the isolation condition is encapsulated with the original condition using bracket encapsulation technology to prevent logical escape. While adhering to SQL syntax standards, this approach balances security and reliability, enabling seamless injection and non-destructive mutation rewriting of tenant isolation conditions.
[0225] In some embodiments, steps b1-b7 above can be implemented using the following data structure: Data Structure 5: / / RewriterConfig Rewrite Engine Configuration type RewriterConfig struct { TenantID string / / Tenant ID of the current session IsolationColumn string / / The name of the isolation field, such as "org_id" } For example, the data structure described above defines a structure named RewriterConfig, representing "rewrite engine configuration". This structure is a container for the rewriter's input parameters and policies. It encapsulates the core business parameters necessary for performing secure rewriting: TenantID string: Stores the identity identifier of the tenant currently making the request. This is the source of the value used to construct data filtering conditions.
[0226] `IsolationColumn` (string): Stores the database column name (e.g., "org_id") that implements row-level data isolation. This is a predefined field name shared by all controlled tables to distinguish tenants, and it is the source of the column used to construct data filtering conditions.
[0227] Data Structures VI: / / BinaryExpr binary expression node (core component of AST) / / Used to construct "Left OP Right" structures, such as "a = 1" or "x AND y". type BinaryExpr struct { Left ASTNode Op string / / Operators: "=", "AND", "OR" Right ASTNode } For example, the data structure six above defines a structure named BinaryExpr, representing a "binary expression node". This is the core syntactic unit in the AST used to represent an operator connecting two sub-expressions. It is a component for constructing and modifying query conditions.
[0228] Left ASTNode and Right ASTNode point to the expression child nodes on the left and right sides of the operator, respectively. They can also be complex AST nodes themselves.
[0229] Op string: Represents the specific operator. In the context of this algorithm, it is mainly responsible for constructing comparison conditions and connection logic conditions, in order to construct isolation condition nodes.
[0230] Data Structures 7: / / ParenExpr bracket expression node (critical security component) / / Used to construct the "(Expr)" structure / / Function: Forces higher precedence of the inner expression. type ParenExpr struct { Expr ASTNode } For example, the data structure 7 above defines a structure named ParentExpr, which represents a "bracket expression node". This is a safety component that represents a pair of parentheses () at the AST level, used to enforce higher precedence of the expression inside it.
[0231] Here, `Expr ASTNode` points to an expression child node enclosed in parentheses. This ensures that no matter how complex the original condition is, it remains logically an indivisible whole before being ANDed with the new tenant condition, thus completely eliminating the risk of logical escape caused by operator precedence.
[0232] Data Structures 8: / / ColumnExpr column reference node type ColumnExpr struct { TableAlias string / / "o" ColumnName string / / "org_id" } For example, the data structure above defines a structure named ColumnExpr, which represents a "column reference node". This structure represents a reference to a column of a specific table in the AST, and it is the part on the left side of the tenant filter condition.
[0233] Here, `TableAlias` (string): The alias for the table. This comes from the `FinalAlias` calculated by the lookup tool (e.g., "o"), ensuring consistency with references to the table in other parts of the SQL statement.
[0234] ColumnName (string): Column name. This comes from RewriterConfig.IsolationColumn (e.g., "org_id"). In the rewrite algorithm, a ColumnExpr node (e.g., {TableAlias: "o", ColumnName: "org_id"}) will be used as the Left child of the BinaryExpr node to represent the o.org_id part.
[0235] Data Structure Nine: / / ValueExpr constant value node type ValueExpr struct { Value string / / "T-1001" } For example, the data structure nine above defines a structure named ValueExpr, representing a "constant value node". This structure represents a constant value in the AST. It is the part on the right side of the tenant filter condition.
[0236] Here, Value string: stores a constant value in string form. In the rewrite algorithm, this value comes from RewriterConfig.TenantID (e.g., "T-1001"). A ValueExpr node (e.g., {Value: "T-1001"}) will be used as the Right child node of the BinaryExpr node to represent the 'T-1001' part.
[0237] In the data structures described above, the various data structures can work together to inject isolation conditions into the target node through operations such as input, construction of atomic conditions, secure encapsulation, and logical grafting, thereby obtaining a complete abstract syntax tree that conforms to the grammatical conditions, which facilitates the subsequent decompilation process.
[0238] In some embodiments, steps b1-b7 above can be implemented using the following methods: / / InjectPredicate executes the main logic for predicate injection / / stmt: SELECT statement node to be modified / / alias: Table aliases recognized by the finder / / config: Contains tenant ID and other configurations func InjectPredicate(stmt *SelectStmt, alias string, configRewriterConfig) { / / 1. Construct isolation constraints: alias.org_id = 'T001' constraint := &BinaryExpr{ Left: &ColumnExpr{TableAlias: alias, ColumnName:config.IsolationColumn}, Op: "=", Right: &ValueExpr{Value: config.TenantID}, } / / 2. Injection logic judgment if stmt.Where == nil { / / [Scenario 1] Original SQL has no WHERE clause / / Action: Direct Mount stmt.Where = constraint } else { / / [Scenario 2] The original SQL already has a WHERE clause (which may contain complex OR logic) / / Action: Encapsulation in brackets + Logical AND merging / / Step 2.1: Encapsulate the original conditions / / Change "status=1 OR status=2" to "(status=1 OR status=2)" safeOldWhere := &ParenExpr{ Expr: stmt.Where, } / / Step 2.2: Reconstructing the new tree / / Change to "(...) AND alias.org_id = 'T001'" newWhere := &BinaryExpr{ Left: safeOldWhere, Op: "AND", Right: constraint, } / / Step 2.3: Replace pointers stmt.Where = newWhere } } In the above method, the predicate injection rewriter uses bracket encapsulation to ensure that the newly injected tenant isolation condition performs a correct logical AND operation with the original condition, regardless of the complexity of the original WHERE condition. This avoids the risk of bypassing tenant filtering by constructing specific OR logic, thus preventing Prompt injection attacks. Before database execution, the predicate injection rewriter injects a filtering condition based on the current tenant ID into the WHERE clause, thereby achieving row-level data isolation in the result set. The predicate injection rewriter modifies the client request and is completely transparent to the standard backend database service, requiring no modification to the backend business code. Thus, at the abstract syntax tree level, the predicate injection rewriter can safely and enforce row-level tenant data isolation conditions into a SELECT query.
[0239] For example, the following case illustrates the process of injecting isolation conditions: Case scenario: Complex logical combination queries.
[0240] Tenant scenario: Tenant T001 attempts to query orders, and the original request contains OR logic.
[0241] Input (Original SQL): SELECT count( ) FROM orders AS o WHERE status = 'paid' OR amount >1000 Intermediate process (AST Processing): 1. The finder identifies the table "orders", also known as "o".
[0242] 2. The injector detected the presence of a WHERE clause.
[0243] 3. The injector executes a bracketed encapsulation: (status = 'paid' OR amount > 1000).
[0244] 4. Add constraint to the injector: AND o.org_id = 'T001'.
[0245] Output (Safe SQL): SELECT count( ) FROM orders AS o WHERE (status = 'paid' OR amount >1000) AND o.org_id = 'T001'.
[0246] For example, the following describes, in conjunction with specific methods, how the multi-tenant data isolation method for large language models provided in this application integrates the finder and rewriter into the main function by calling the previously designed RecursiveFinder.Visit and InjectPredicate functions to form a complete call chain: package execution_engine import ( "mcp / security" / / Refers to the previously defined finder and rewriter "mcp / parser" ) / / ProcessSQLRequest main entry point for the execution engine / / rawSQL: raw SQL / / tenantID: Current tenant / / policy: Isolation strategy func ProcessSQLRequest(rawSQL string, tenantID string, policySecurityPolicy) (string, error) { / / 1. [AST Construction] Parsing SQL astRoot, err := parser.Parse(rawSQL) if err != nil { return "", NewError("Invalid SQL Syntax") } / / 2. [Target Location] Initialize the recursive searcher finder := &security.RecursiveFinder{ Config: security.FinderConfig{ ProtectedTables: policy.ProtectedTables, / / eg {"orders": true} }, } / / Perform a traversal to obtain all targets that need to be injected. / / Targets are of type []InjectionTarget targets := finder.Visit(astRoot, nil) / / 3. [Safe Injection] Iterate through all targets and perform rewriting. rewriter := security.RewriterConfig{ TenantID: tenantID, IsolationColumn: "org_id", } for _, target := range targets { / / Call the predicate injection algorithm / / target.ParentContext is the SelectStmt node / / target.FinalAlias is the calculated alias (e.g., "o"). security.InjectPredicate( target.ParentContext, target.FinalAlias, rewriter, ) } / / 4. [Decompile] Restore to string safeSQL := parser.Restore(astRoot) return safeSQL, nil } For example, in the above method, by means of AST construction, target node location, isolation condition injection, and decompilation restoration, isolation conditions are transparently and forcibly injected into each query in the case of shared database tables to achieve row-level data isolation. At the same time, by using bracket encapsulation technology and logic grafting technology at the AST level, the risk of attackers inducing large models to generate unauthorized queries through carefully constructed OR and other logic is fundamentally solved, thereby achieving defense against Prompt injection attacks.
[0247] Figure 9 This is an architecture diagram of a multi-tenant data isolation gateway for a large language model, provided for some embodiments of this application.
[0248] Combination Figure 9 As shown, in some embodiments, the multi-tenant data isolation gateway 100 for large language models includes a multi-tenant connection module 110 and a semantic isolation execution module 120.
[0249] The multi-tenant connection module 110 can intercept client requests and parse the tenant identity identifier corresponding to the request.
[0250] For example, the multi-tenant connection module 110 may include the multi-tenant connector described above. For related descriptions, please refer to the foregoing description, which will not be repeated here.
[0251] The semantic isolation execution module 120 can inject isolation conditions into the request based on the tenant identity identifier when the request includes a tool call class request, and obtain the rewritten request.
[0252] The semantic isolation execution module 120 may include the SIE described above. For related descriptions, please refer to the foregoing descriptions, which will not be repeated here.
[0253] For example, the multi-tenant connection module 110 can also forward the rewritten request to the server.
[0254] The multi-tenant data isolation gateway 100 for large language models provided in this application embodiment can intercept client requests and extract tenant identity identifiers through the multi-tenant connection module 110, and rewrite requests through the semantic isolation execution module 120 to inject security isolation conditions, thereby achieving visibility isolation based on the principle of least privilege and avoiding the possibility of the model supermarket calling unauthorized high-level tools.
[0255] In some embodiments, the multi-tenant data isolation gateway 100 for large language models further includes a virtual resource mapping module 130. The virtual resource mapping module 130 can filter and reorganize the physical tool list based on the tenant's identity to generate a tool list available to the tenant.
[0256] The virtual resource mapping module 130 may include the virtual resource mapper described above. For related descriptions, please refer to the foregoing descriptions, which will not be repeated here.
[0257] In this embodiment, the virtual resource mapping module 130 can filter and reorganize the list of tools available to tenants, so as to provide tenants with personalized and fine-grained tool access control and ensure that each tenant can only call authorized tool services.
[0258] In the multi-tenant data isolation gateway 100 for large language models provided in this application embodiment, a three-layer protocol proxy architecture is formed between the MCP Client and the MCP Server through a multi-tenant connection module 110, a virtual resource mapping module 130, and a semantic isolation execution module 120. This proxy architecture does not require changes to the server-side code, thus forming a transparent proxy architecture that is compatible with all standard MCP ecosystems. Logically, the system is divided into three layers from top to bottom: the access and session layer, the policy control layer, and the semantic execution layer. The multi-tenant connection module 110 is located in the access and session layer, the virtual resource mapping module 130 is located in the policy control layer, and the semantic isolation execution module 120 is located in the semantic execution layer. Furthermore, the semantic isolation execution module 120 uses predicate injection technology based on AST to solve the problem that regular expression matching cannot handle complex SQL. The semantic isolation execution module 120 also uses sandbox anchoring based on path normalization to prevent file operation tools from escaping.
[0259] Furthermore, the multi-tenant data isolation gateway 100 architecture for large language models provided in this application embodiment can be deployed in modern cloud-native or distributed environments to cope with a large number of tenants and high concurrency requests.
[0260] In some embodiments, the virtual resource mapping module 130 may also generate a virtual view of available tools based on a list of tools available to the tenant and return the virtual view to the client when the request includes a resource discovery request.
[0261] In this embodiment, the virtual resource mapping module 130 can generate a virtual view of the list of available tools and dynamically render it as a tenant-specific interactive interface, so that the client only perceives the authorized set of tools.
[0262] Figure 10 This is a schematic diagram of a computing device provided for some embodiments of this application.
[0263] like Figure 10 As shown, the computing device 1000 includes a processor 1001 and a memory 1002. Exemplarily, the computing device 1000 may also include a communications interface 1003 and a communications bus 1004.
[0264] The processor 1001, memory 1002, and communication interface 1003 communicate with each other via communication bus 1004. The communication interface 1003 may include a transmitter and receiver for communicating with other devices or communication networks, and may be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet interface (GE).
[0265] In some embodiments, the processor 1001 is used to execute program 1005, which may specifically perform the relevant steps in the above method embodiments. Specifically, program 1005 may include program code, which includes computer-executable instructions.
[0266] For example, processor 1001 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of this application. Computing device 1000 may include one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs. The CPU may be a single-core CPU or a multi-core CPU.
[0267] In some embodiments, memory 1002 is used to store program 1005. Memory 1002 may include high-speed random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device.
[0268] Specifically, program 1005 can be invoked by processor 1001 and executed by computing device 1000 using the above method.
[0269] Some embodiments of this application provide a computer-readable storage medium storing at least one executable instruction that, when executed on a computing device 1000, causes the computing device 1000 to perform the multi-tenant data isolation method for the large language model described above.
[0270] For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device.
[0271] Some embodiments of this application provide a chip system applied to a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via lines. The interface circuits are used to receive signals from the server's memory and send signals to the processors, the signals including computer instructions stored in the memory. When the server processor executes the computer instructions, the server performs various steps in the multi-tenant data isolation method for large language models shown in the above-described method embodiments.
[0272] The beneficial effects that the readable storage medium provided in some embodiments of this application can achieve can be referred to the beneficial effects in the corresponding inference task execution method provided above, and will not be repeated here.
[0273] It should be noted that, in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0274] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0275] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0276] For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0277] More specific examples of computer-readable media (a non-exhaustive list) include the following: electrical connections having one or more wires (electronic devices), portable computer disks (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM).
[0278] Furthermore, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory. It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof.
[0279] In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc. The above embodiments are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of this application should be included within the scope of protection of this application.
Claims
1. A multi-tenant data isolation method for a large language model, characterized in that, include: Intercept client requests; Parse the tenant identity identifier corresponding to the request; In the case where the request includes a tool call request, an isolation condition is injected into the request based on the tenant identity to obtain a rewritten request, wherein the isolation condition includes the range of restricted data access for the tenant as confirmed by the tenant identity. The rewritten request is forwarded to the server.
2. The multi-tenant data isolation method for large language models according to claim 1, characterized in that, In the case where the request includes a tool invocation request, injecting isolation conditions into the request based on the tenant identity identifier to obtain a rewritten request includes: In the case where the request includes a tool call request, the SQL string in the request is parsed to obtain an abstract syntax tree; Traverse the abstract syntax tree and inject isolation conditions into the WHERE clause of the abstract syntax tree based on the tenant identity to obtain the modified abstract syntax tree; The modified abstract syntax tree is restored to an SQL string to obtain the rewritten request.
3. The multi-tenant data isolation method for large language models according to claim 2, characterized in that, The process of traversing the abstract syntax tree and injecting isolation conditions into the WHERE clause of the abstract syntax tree based on the tenant identity to obtain a modified abstract syntax tree includes: The abstract syntax tree is traversed by depth-first search to find the target node in the abstract syntax tree and obtain the effective alias of the target node. The target node includes nodes that reference controlled tables, which are a set of target physical tables that need to be isolated from data. Based on the tenant identity identifier and the valid alias of the target node, isolation conditions are injected into the WHERE clause of the target node to obtain the modified abstract syntax tree.
4. The multi-tenant data isolation method for large language models according to claim 3, characterized in that, The process of traversing the abstract syntax tree using a depth-first search to find the target node in the abstract syntax tree and obtain the valid alias of the target node includes: Input the root node of the abstract syntax tree, initialize the result set, wherein the result set includes a result list for storing the target node information, and the result list is empty; Start the recursive traversal; When the node type in the abstract syntax tree is a table source node, the current physical table corresponding to the current node is compared with the target physical table in the controlled table to determine whether the current physical table is in the controlled table; If the current physical table is in the controlled table, the current node is determined as the target node and a valid alias for the physical table referenced by the target node is determined; The pointer to the target node and the valid alias are encapsulated as structured data and recorded in the result list.
5. The multi-tenant data isolation method for large language models according to claim 4, characterized in that, Also includes: When the node type in the abstract syntax tree is a connection structure, the left and right subtrees are recursively visited. When the node type in the abstract syntax tree is a query statement, the context is updated and the child nodes are recursively visited.
6. The multi-tenant data isolation method for large language models according to any one of claims 3-5, characterized in that, The process of injecting isolation conditions into the WHERE clause of the target node based on the tenant identity to obtain a modified abstract syntax tree includes: Construct a tenant isolation condition node based on the tenant identity identifier and the valid alias; Check the attribute status of the WHERE clause in the target node; If the attribute status of the WHERE clause is that the original condition exists, enclose the original condition in parentheses to obtain a parenthesis node; Construct an AND node to obtain a new WHERE clause, wherein the left child node of the AND node is the bracket node, and the right child node of the AND node is the tenant isolation condition node; The original WHERE clause is replaced with the new WHERE clause to obtain the modified abstract syntax tree.
7. The multi-tenant data isolation method for large language models according to claim 6, characterized in that, Also includes: If the attribute status of the WHERE clause is that the original condition does not exist, the tenant isolation condition node is assigned to the WHERE clause to obtain the modified abstract syntax tree.
8. The multi-tenant data isolation method for large language models according to any one of claims 1-7, characterized in that, Also includes: The physical tool list is filtered and reorganized based on the tenant's identity to generate a tool list available to the tenant.
9. The multi-tenant data isolation method for large language models according to claim 8, characterized in that, Also includes: In the case that the request includes a resource discovery request, a virtual view of available tools is generated based on the list of tools available to the tenant, which hides the tools available to other tenants from the current tenant's client; Return the virtual view to the client.
10. The multi-tenant data isolation method for large language models according to any one of claims 1-9, characterized in that, Also includes: Intercept the client's initial network connection request; Upgrade the initial network connection to a persistent connection; Extract the tenant identity identifier corresponding to the initial network connection request; Create a session, where the session object holds a reference to the long-connection object; Bind the tenant identity identifier to the session; During the duration of the session, the binding relationship between the tenant identity and the session is based on the initial binding at the time the session is established, and the binding relationship remains valid throughout the entire lifecycle of the session.
11. A multi-tenant data isolation gateway for a large language model, characterized in that, The method described in any one of claims 1-10, wherein the multi-tenant data isolation gateway for the large language model comprises: The multi-tenant connection module is configured to: intercept client requests and parse the tenant identity identifier corresponding to the request; The semantic isolation execution module is configured to: when the request includes a tool call request, inject isolation conditions into the request based on the tenant identity to obtain a rewritten request, wherein the isolation conditions include the range of restricted data accessed by the tenant as confirmed based on the tenant identity; The multi-tenant connection module is also configured to forward the rewritten request to the server.
12. The multi-tenant data isolation gateway for large language models according to claim 11, characterized in that, Also includes: The virtual resource mapping module is configured to filter and reorganize the physical tool list based on the tenant identity to generate a tool list available to the tenant.
13. The multi-tenant data isolation gateway for large language models according to claim 12, characterized in that, The virtual resource mapping module is also configured to: when the request includes a resource discovery request, generate a virtual view of available tools based on the list of tools available to the tenant, which hides the tools available to other tenants from the current tenant client, and return the virtual view to the client.
14. A computing device, characterized in that, The computing device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the computing device to apply the multi-tenant data isolation method for a large language model as described in any one of claims 1 to 10.