Methods and systems to mitigate corporate data leakage when querying large-scale language models.

QueryShield addresses corporate data leakage in LLM queries by semantically analyzing and paraphrasing queries using lightweight models, ensuring secure access to external LLMs while maintaining message intent, thus reducing sensitive information exposure.

JP2026090209APending Publication Date: 2026-06-02TATA CONSULTANCY SERVICES LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TATA CONSULTANCY SERVICES LTD
Filing Date
2025-10-31
Publication Date
2026-06-02

Smart Images

  • Figure 2026090209000001_ABST
    Figure 2026090209000001_ABST
Patent Text Reader

Abstract

This invention provides methods and systems to mitigate the leakage of corporate data when querying large-scale language models. [Solution] The QueryShield platform is a platform that companies can use to interact with external LLMs without leaking data through queries. It detects whether a query is leaking data and generates paraphrased queries to minimize data leakage while reducing its impact on semantics. Language models are selected from a set of lightweight model candidates identified and fine-tuned for this purpose using massive datasets and evaluated using multiple metrics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - References and Priority to Related Applications) This application claims priority to Indian Application No. 202421084709, filed on November 5, 2024.

[0002] The disclosure of this specification generally relates to the field of machine learning, and more particularly, to methods and systems for reducing the leakage of enterprise data in queries to large language models.

Background Art

[0003] The rapid progress of generative AI (Gen - AI), particularly large language models (LLMs), has significantly improved productivity across various industries. These models, which can understand and generate human - like text, greatly reduce the time required for tasks that conventionally required a great deal of human effort. This efficiency allows companies to improve throughput without sacrificing the quality of output. AI has emerged as a tool to enhance human capabilities, and by integrating AI, companies can maintain their competitiveness. Companies that have adopted AI have experienced a significant improvement in productivity compared to those that have not. This gap is further widening with the introduction of Gen - AI.

[0004] However, special investigation is needed regarding the privacy, security, and safety implications of Gen-AI. Because it is trained on massive datasets, it has been observed that sensitive details may unintentionally surface in the model's output. The accurate and consistent performance of LLM stems from its ability to store sparse training samples, but this poses a significant privacy threat if the datasets used to train these samples contain sensitive data. On the other hand, since humans are the most vulnerable in terms of security and privacy, data can be leaked to LLM through user queries. LLM service providers can use this interaction data for further model training, and therefore, if attacked, they could potentially leak the same sensitive data that was once sent as a query.

[0005] This risk is exacerbated if company employees, in an attempt to gain a competitive edge, promptly leak confidential company data to external LLM services such as Chat GPT or Google Gemini. Despite the confidentiality guarantees provided by LLM service providers, there have been instances of unintentional chat data leaks. This concern has led some companies to systematically ban chat models. Such restrictions have a serious impact on a company's competitiveness, especially if competent alternatives are not available within the organization. There is a growing need for privacy-protecting prompt solutions that not only protect against data breaches but also ensure that the utility provided by robust external LLMs like GPT-4o remains unaffected. [Overview of the Initiative]

[0006] Embodiments of this disclosure present technical improvements as solutions to one or more of the above-mentioned technical problems recognized by the inventors in conventional systems. For example, one embodiment provides a method for mitigating corporate data leakage in queries to large language models (LLMs). The method includes receiving input user queries associated with queries to multiple large language models (LLMs) using one or more hardware processors. Furthermore, the method includes the step of calculating a sensitive data leakage level associated with an input user query by prompting a corporate data leakage mitigation model, trained on a first set of specific instructions (T1) as a prefix for the input user query, using one or more hardware processors, wherein the sensitive data leakage level is classified as one of (i) high and (ii) low based on an associated threshold. Furthermore, the method includes generating multiple paraphrasing queries associated with an input user query using one or more hardware processors, which, when the sensitive data leakage level is higher than a predefined threshold, reduce the sensitive data leakage level by preserving the semantics of the input user query and prompting the trained enterprise data leakage mitigation model with a second set of specific instructions (T2) as a prefix to the user query via a trained enterprise data leakage mitigation model. Furthermore, the method includes simultaneously identifying the type of sensitive data leakage associated with the input user query by prompting the trained enterprise data leakage mitigation model with a third set of specific instructions (T3) as a prefix to the input user query in the trained enterprise data leakage mitigation model using one or more hardware processors. Finally, the method includes repeating the above steps using one or more hardware processors until an optimal paraphrasing query is generated in which the sensitive data leakage level is below a predefined threshold.

[0007] In another embodiment, a system is provided for mitigating corporate data leakage in queries to large-scale language models (LLMs). The system includes at least one memory for storing programmed instructions, one or more input / output (I / O) interfaces, and one or more hardware processors operably coupled to at least one memory, the one or more hardware processors being configured by programmed instructions to receive input user queries associated with queries to multiple large-scale language models (LLMs). Furthermore, the one or more hardware processors being configured by programmed instructions to calculate the level of sensitive data leakage associated with an input user query by prompting a corporate data leakage mitigation model trained on a first set of specific instructions (T1) as a prefix for the input user query, the level of sensitive data leakage being classified as either (i) high or (ii) low based on an associated threshold. Furthermore, one or more hardware processors are configured by programmed instructions to generate multiple paraphrasing queries associated with an input user query, which, when the sensitive data leakage level is higher than a predefined threshold, preserve the semantics of the input user query through a trained enterprise data leakage mitigation model and reduce the sensitive data leakage level by prompting the trained enterprise data leakage mitigation model with a second set of specific instructions (T2) as a prefix to the user query. Furthermore, one or more hardware processors are configured by programmed instructions to simultaneously identify the type of sensitive data leakage associated with the input user query by prompting the trained enterprise data leakage mitigation model based on a third set of specific instructions (T3) as a prefix to the input user query in the trained enterprise data leakage mitigation model. Finally, one or more hardware processors are configured by programmed instructions to repeat the above steps until an optimal paraphrasing query is generated in which the sensitive data leakage level is below a predefined threshold.

[0008] In yet another embodiment, a computer program product is provided that includes a non-temporary computer-readable medium embodying a computer program for mitigating corporate data leakage in queries to large-scale language models (LLMs). When the computer-readable program is executed on a computing device, it causes the computing device to receive input user queries associated with queries to multiple large-scale language models (LLMs). Furthermore, when the computer-readable program is executed on the computing device, it causes the computing device to calculate the level of sensitive data leakage associated with the input user query by prompting the computing device to calculate a corporate data leakage mitigation model trained on a first set of specific instructions (T1) as a prefix for the input user query, and the level of sensitive data leakage is classified as either (i) high or (ii) low based on the associated threshold. Furthermore, when the computer-readable program is executed on a computing device, if the level of sensitive data leakage is higher than a predefined threshold, the computing device generates multiple paraphrased queries associated with the input user query, which reduce the level of sensitive data leakage by preserving the semantics of the input user query and prompting the traininged corporate data leakage mitigation model using a second set of specific instructions (T2) as a prefix to the user query. Furthermore, when the computer-readable program is executed on a computing device, the computing device simultaneously identifies the type of sensitive data leakage associated with the input user query by prompting the computing device to use a traininged corporate data leakage mitigation model based on a third set of specific instructions (T3) as a prefix to the input user query in the traininged corporate data leakage mitigation model. Finally, when the computer-readable program is executed on a computing device, the computing device repeats the above steps until it generates an optimal paraphrased query in which the level of sensitive data leakage is below a predefined threshold.

[0009] Please understand that the above summary and the following detailed description are both illustrative and intended to further illustrate the embodiments of the present invention as described in the claims.

[0010] The accompanying drawings incorporated herein and constituting part thereof illustrate one or more embodiments and serve to illustrate these embodiments together with this specification. [Brief explanation of the drawing]

[0011] [Figure 1A] This is a functional block diagram of a system for mitigating corporate data leakage in queries to large language models (LLMs) according to several embodiments of the present disclosure. [Figure 1B] This figure shows the overall functional architecture of a system for mitigating corporate data leakage in queries to LLMs, according to several embodiments of this disclosure. [Figure 2] The flowcharts below illustrate processor implementation methods for mitigating corporate data leakage in queries to LLMs, according to several embodiments of this disclosure. [Figure 3] This figure shows the distribution of data leakage types to mitigate corporate data leakage in queries to LLMs, according to several embodiments of this disclosure. [Modes for carrying out the invention]

[0012] Exemplary embodiments will be described with reference to the accompanying drawings. In the drawings, the leftmost digit of the reference number identifies the drawing in which the reference number first appears. For convenience, the same reference number is used throughout the drawings to refer to the same or similar parts. While embodiments and features of the disclosed principle are described herein, modifications, alterations, and other embodiments are possible without departing from the scope of the disclosed embodiments.

[0013] Today, there is growing concern about the leakage of sensitive data when querying Large Language Models (LLMs). This risk is exacerbated if an organization's employees leak confidential company data through prompts to LLM services such as Chat-GPT or Google Gemini in an attempt to gain a competitive edge. Therefore, there is a growing need for privacy-protecting prompt solutions that not only protect against data breaches but also ensure that the utility provided by strong external LLMs like GPT-4o is not compromised. This is an example of the Private Inference (PI) problem in neural networks, where inference is performed on encrypted data. Encryption techniques such as Fully Homomorphic Encryption (FHE) and Secure Multi-Party Computation (MPC) have been employed to address this problem. However, the complexity of communication and computation in these methods makes performing inference on large language models impractical. Furthermore, encryption techniques need to be implemented on both the server and client (prompter) sides. Such solutions are impractical because external LLM providers like OpenAI (ChatGPT) do not welcome the execution of server-side code.

[0014] A direct solution is data sanitization, which detects textual portions that leak sensitive information. This method is limited by the fact that even common words can leak personal information when the context in which they are used changes. Therefore, a method is needed to analyze the potential for data leakage from queries as a whole. Furthermore, this analysis should be used to rephrase queries so that data leakage, if any, is minimized without affecting the semantic integrity of the message the query is trying to convey. To achieve this, a system is needed that can understand queries semantically while simultaneously understanding the concept of data leakage.

[0015] Private inference (PI) refers to the process of extracting predictions from a neural network while keeping its input confidential. Traditionally, this has been achieved using cryptographic techniques such as fully homomorphic encryption (FHE). FHE has significant communication overhead, and hybrid approaches that optimize solutions from both ML and FHE perspectives have been used to advance PI delivery. However, due to the large scale of LLM, even such optimizations were insufficient to achieve real-time PI. Therefore, the focus shifted to other natural language processing (NLP) techniques. The first such attempts involved the use of part-of-speech (POS) tagging, named entity recognition (NER), and personally identifiable information (PII) detection. Differential privacy (DP)-based techniques add noise to personal data, ensuring compelling falsification for LLM queries at the word, sentence, and document levels. Word-level implementations that add noise to word embeddings are limited by context-based data leakage. Sentence-level DP techniques introduce noise into sentence embeddings. These capture context-based data leakage, where words leak data depending on the context in which they are used.

[0016] To address the technical complexities of conventional methods, embodiments of this specification provide methods and systems for mitigating enterprise data leakage in queries to large-scale language models (LLMs). This disclosure provides QueryShield, a platform between an enterprise environment and any external LLM. QueryShield detects outgoing queries that leak sensitive data and paraphrases the queries to remove sensitive content. Queries that do not leak sensitive data are allowed to pass to the external LLM, and paraphrased versions of highly sensitive queries, along with the identified type of leakage, are fed back to the user, who can optionally edit and resubmit the query. Specific contributions of this disclosure include: evaluation of modern lightweight language models for the task of identifying and paraphrasing data leakage found in enterprise queries, particularly multitask encoder-decoder models fine-tuned using curriculum learning; a dataset of 1500 queries sent from an enterprise environment to an external LLM, labeled with data leakage sensitivity, and corresponding gold-standard human-paraphrased versions of highly sensitive queries; and a novel evaluation metric, Cross-Reference ROUGE, for evaluating semantic-preserving paraphrasing of highly sensitive queries.

[0017] Herein, referring more specifically to Figures 1A to 3, in which the same reference numerals indicate features corresponding throughout the figures, preferred embodiments are shown, and these embodiments are described in relation to the following exemplary systems and / or methods.

[0018] Figure 1A is a functional block diagram of a system 100 for mitigating corporate data leakage in queries to large language models, according to several embodiments of the present disclosure. System 100 includes, or otherwise communicates with, a hardware processor 102, at least one memory such as memory 104, and an input / output (I / O) interface 112. The hardware processor 102, memory 104, and I / O interface 112 can be coupled by a system bus such as a system bus 108 or a similar mechanism. In one embodiment, the hardware processor 102 may be one or more hardware processors.

[0019] The I / O interface 112 can include various software and hardware interfaces, such as a web interface and a graphical user interface. Furthermore, the I / O interface 112 can enable the system 100 to communicate with other devices, such as a web server or an external database.

[0020] The I / O interface 112 can facilitate multiple communications within a wide variety of network and protocol types, including, for example, wired networks such as local area networks (LANs) and cable networks, and wireless networks such as wireless LANs (WLANs), cellular networks, or satellite networks. For this purpose, the I / O interface 112 may include one or more ports for connecting multiple computing systems to each other or to another server computer. The I / O interface 112 may also include one or more ports for connecting multiple devices to each other or to another server.

[0021] One or more hardware processors 102 can be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, graphical processing units (GPUs), node machines, logic circuits, and / or devices that manipulate signals based on operational instructions. Among other capabilities, one or more hardware processors 102 are configured to fetch and execute computer-readable instructions stored in memory 104.

[0022] The memory 104 may include any computer-readable medium known in the art, such as volatile memory, including static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, including read-only memory (ROM), erasable programmable ROM, flash memory, hard disk, optical disk, video random access memory (VRAM), and magnetic tape. In one embodiment, the memory 104 includes a plurality of modules 106. The memory 104 also includes a data repository (or repository) 110 for storing data processed, received, and generated by the plurality of modules 106.

[0023] The plurality of modules 106 includes programs or coded instructions that complement the applications or functions executed by the system 100 for reducing the leakage of enterprise data in queries to large language models. The plurality of modules 106 can include, among other things, routines, programs, objects, components, and data structures that perform specific tasks or implement specific abstract data types. Also, the plurality of modules 106 can be used as signal processors, node machines, logic circuits, and / or any other device or component that operates on signals based on operating instructions. Further, the plurality of modules 106 can be used by hardware, by computer-readable instructions executed by one or more hardware processors 102, or by a combination of these. The plurality of modules 106 can include various sub-modules (not shown). The plurality of modules 106 can include computer-readable instructions that complement the applications or functions executed by the system 100 for reducing the leakage of enterprise data in queries to large language models.

[0024] The data repository (or repository) 110 can include a plurality of abstracted code elements for fine-tuning and data processed, received, or generated as a result of the execution of the plurality of modules within the module 106.

[0025] Although the data repository 110 is shown inside the system 100, in an alternative embodiment, the data repository 110 may be implemented outside the system 100, and it should be noted that the data repository 110 can be stored in a database (repository 110) communicatively coupled to the system 100. The data contained in such an external database can be updated periodically. For example, new data may be added to the database (not shown in FIG. 1A), and / or existing data may be modified, and / or useless data may be deleted from the database. In one example, the data can be stored in an external system such as a lightweight directory access protocol (LDAP directory) or a relational database management system (RDBMS).

[0026] The overall architecture of the system of FIG. 1A will be described with FIG. 1B. Referring now to FIG. 1B, the present disclosure detects outgoing queries that leak confidential data and rephrases them to remove the confidential content. Queries that do not leak confidential data are permitted to pass to the external LLM, while the rephrased versions of sensitive queries are fed back to the user, along with the type of leak identified as an explanation, and can optionally be edited and resent.

[0027] The operation of the components of the system 100 will be described with reference to the method steps depicted in FIG. 2.

[0028] Figure 2 is an exemplary flowchart illustrating Method 200 for mitigating enterprise data leakage in queries to large language models implemented by the systems of Figures 1A and 1B, according to several embodiments of the present disclosure. In one embodiment, System 100 includes one or more data storage devices or memories 104 operably coupled to one or more hardware processors 102, and is configured to store instructions for the execution of steps of Method 200 by one or more hardware processors 102. Hereinafter, the steps of Method 200 of the present disclosure will be described with reference to the components or blocks of System 100 as depicted in Figures 1A and 1B, and the steps of the flowchart as depicted in Figure 2. Method 200 can be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, functions, etc., that perform a particular function or implement a particular abstract data type.

[0029] Method 200 can also be implemented in a distributed computing environment in which functions are performed by remote processing devices linked via a communication network. The order in which Method 200 is described is not intended to be construed as limiting, and any number of the described method blocks can be combined in any order to implement Method 200 or an alternative method. Furthermore, Method 200 can be implemented with any suitable hardware, software, firmware, or a combination thereof.

[0030] Referring now to Figure 2, in step 202 of method 200, one or more hardware processors 102 are configured to receive input user queries associated with multiple LLM queries by programmed instructions.

[0031] In step 204 of Method 200, one or more hardware processors 102 are configured to calculate the sensitive data leakage level associated with an input user query by prompting a programmed instruction for an enterprise data leakage mitigation model trained on a first set of specific instructions (T1) as a prefix for the input user query, wherein the sensitive data leakage level is classified as one of (i) high and (ii) low based on the associated threshold.

[0032] Users are permitted to query relevant LLMs from among multiple LLMs using input user queries only if the level of sensitive data leakage is below a predefined threshold. LLMs can be internal or external.

[0033] In step 206 of Method 200, one or more hardware processors 102 are configured by programmed instructions to generate multiple paraphrased queries associated with an input user query, which preserve the semantics of the input user query and reduce the level of sensitive data leakage, by prompting a trained enterprise data leakage mitigation model via a trained enterprise data leakage mitigation model using a second set of specific instructions (T2) as a prefix for the user query when the level of sensitive data leakage is higher than a predefined threshold.

[0034] In step 208 of Method 200, one or more hardware processors 102 are configured to simultaneously identify the type of sensitive data breach associated with an input user query by prompting a trained enterprise data breach mitigation model based on a third set of specific instructions (T3) as a prefix to the user query in the trained enterprise data breach mitigation model using programmed instructions. The types of sensitive data breaches include, but are not limited to, personally identifiable information (PII), business relationship information, proprietary data, internal policies, strategic plans, and research and development information. The identified type of sensitive data breach associated with the input user query is communicated to the user / administrator, who can edit the query as needed and use it to update the data breach mitigation model.

[0035] For example, PII includes contact information such as personal names, email addresses, or addresses. Business relationship information includes customer or vendor names, contact information, relationship values, transaction information, and contract terms. Proprietary data is all kinds of internal confidential / private data of a company, including internal data and work products. In an IT company, this would be source code, software requirements, algorithms, and implementation details. In a hospital, this would be treatment details and research reports. Internal policies include internal policies and procedures, security protocols, internal audits, project management guidelines / data, and governance and compliance guidelines / data. Strategic plans include long-term strategies, product / service launch plans, merger / acquisition / partnership proposals, and marketing / sales strategies (detailed sales forecasts, campaign information, etc.). Research and development information includes the latest research initiatives, ideas, and unpublished intellectual property.

[0036] In step 210 of method 200, one or more hardware processors 102 are configured to repeat the above steps by programmed instructions until they generate an optimal paraphrase query in which the level of sensitive data leakage is below a predefined threshold.

[0037] The enterprise data loss mitigation model is obtained as follows: First, multiple user queries are acquired from multiple sources. These sources include queries generated by user groups, ChatGPT, and queries from multiple publicly available datasets. Furthermore, multiple training instances for tasks T1, T2, and T3 are acquired from the multiple sources. For example, T1 is for classifying the sensitive data loss level in multiple user queries, T2 is for generating paraphrased queries when the classified queries have a high sensitive data loss level, and T3 is for detecting the type of sensitive data loss level in relation to a predefined list of sensitive data loss levels. After acquiring the training instances, gold standard labels are acquired for the multiple training instances of T1, T2, and T3. The gold standard labels for the training instances of T1, T2, and T3 are acquired from multiple associated predefined annotations. After acquiring the gold standard labels, the language model, pre-trained for K epochs, is fine-tuned on the training instance of T1. Each of the multiple training instances of T1 contains input text paired with expected labeled output text derived from a set of predefined annotations. The multiple training instances of T1 are debiased using anomaly detection techniques. Furthermore, another language model is fine-tuned over the relevant multiple training instances of T1, and T3 is tuned for K epochs with a validation loss smaller than a predefined threshold. The validation loss is calculated for the validation set in each of the K epochs, and the optimal validation loss for K epochs is selected based on the predefined validation threshold. Each of the multiple training instances of T3 contains text of a high-sensitivity data leak level paired with an associated high-sensitivity data leak type based on a set of predefined annotations. The language model is further fine-tuned using the optimal validation loss over the relevant multiple training instances of T1, T2, and T3.Each of T2's multiple training instances includes paired high-security, data-leakage level text and paraphrased output text generated based on a set of predefined annotations. The overall fine-tuning process described above follows a curriculum learning strategy, with easier training instances used first, followed by more difficult instances. After fine-tuning the language model, the trained language model is evaluated using multiple metrics.

[0038] For example, T1 is evaluated using the recall and F1 score metrics. T3 is evaluated using the micro-average and macro-average F1 scores. T2 is evaluated by (i) calculating a cross-reference score by comparing multiple paraphrasing queries, input user queries, and gold standard paraphrasing queries; (ii) calculating named entity leakage (NEL) as the rate at which named entities in multiple paraphrasing queries occur as part of false positives; (iii) evaluating multiple paraphrasing queries using the CRR score and NEL to preserve the original semantics of query-q to the greatest extent possible; and (iv) being evaluated by the label "LOW" of T1.

[0039] To cover both of these aspects (leakage and intent) with a single metric, this disclosure utilizes an evaluation metric called Cross-Reference ROUGE (CRR), which compares the generated text to two references (the original query and the Gold Standard paraphrased query), unlike vanilla ROUGE which uses a single reference. To illustrate this metric, consider the monogram form of CRR1 (Equations 1-7). Let O, G, and R be the sets of monograms contained in the original query, the Gold Standard paraphrased query, and the paraphrased query generated by the model, respectively. Now referring to Equations 1 through 7, the aspects of leakage: TIFF2026090209000002.tif6150 captures the sensitive content of the original query, and any overlap between R and this sensitive content indicates excess leakage. Thus, such overlaps are a set of false positives (FPl) that should not exist in R (Equation 1). The remaining terms of R are considered true positive (Equation 2), Used to calculate JPEG2026090209000003.jpg6150 (Equation 3). Intent side aspect: JPEG2026090209000004.jpg6150 captures the acceptable intent of the original query, and if these terms are not present in R, it indicates an Intent Loss (No Intent). Therefore, these missing terms result in a False Negative (FN). i ) becomes (Equation 4). The remaining terms in JPEG2026090209000005.jpg6150 are considered true positive (Equation 5), This is used in the calculation of JPEG2026090209000006.jpg6150 (Equation 6). Finally, The JPEG2026090209000007.jpg6150 score (Equation 7) is calculated as the final metric. JPEG2026090209000008.jpg5985

[0040] For example, encoder-only models are used for Task T1, a binary classification task, and Task T3, a multi-label multi-class classification task. Since Task T2 is a text generation task, an encoder-only model is not applicable. Based on previous research, Attn-BERT was used, which employs attention-weighted BERT representations of tokens in a query, concatenated with a [CLS] representation of the query. CLS stands for classification. The concatenated representations are passed to a softmax layer for final prediction. In the case of multi-label classification, each class label has a separate attention head, which is linked to its specific representation.

[0041] Table IA shows the prompts used for tasks T1 and T3 using a decoder-only model. Table IA JPEG2026090209000009.jpg154150

[0042] Similarly, Table IB shows the prompts used by the decoder-only model in task T2. Table IB JPEG2026090209000010.jpg144150

[0043] For example, a decoder-only model is used to solve all three tasks. For each task, a prompt was designed consisting of a detailed definition of data leakage categorized into six types and instructions to produce the desired output. In context-based learning, several demonstrations of the tasks are added as a few-shot example. For each query in the test set, eight of the most similar queries are selected from the training set and used as a few-shot example. At T2, only high-sensitivity training queries are selected, while at T1 and T3, four high-sensitivity training queries and four low-sensitivity training queries are selected. To identify the most similar queries from the training set, cosine similarity between text embeddings obtained using a sentence transformation model was used.

[0044] For example, an encoder-decoder model is used because, unlike an encoder-only model, it not only provides text generation capabilities, but is also reasonably sized and easy to fine-tune (unlike a large decoder-only model). Here, three tasks T1, T2, and T3 are formulated as text-to-text conversion tasks, and a single T5-based model is fine-tuned for all tasks. In each task, a specific instruction is appended before the query that constitutes the input text to the model. Table II shows the various instructions used for tasks T1, T2, and T3. The expected output also differs for each task. In T1, the output text is simply the data leakage level of the query, which is either HIGH or LOW. In T2, the output text is a paraphrased version of the input query that does not contain sensitive data and retains the original semantics as much as possible. In T3, the output text is a comma-separated list of data leakage types present in the input query. The consideration of a T5-based model was also necessary due to the constraint that it should be deployable even within organizations with limited hardware resources. Table II JPEG2026090209000011.jpg139150

[0045] Some examples of paraphrasing queries derived from this disclosure are shown below. JPEG2026090209000012.jpg139150

[0046] experiment:

[0047] Data Collection and Labeling: First, we investigated many publicly available datasets used for instruction tuning in LLM, such as OASST11 and ChatAlpaca 20K2. However, we observed that very few of the queries in these datasets were truly sensitive from an organizational perspective and fit the description of sensitivity. Therefore, we decided to create our own dataset (212).

[0048] Obtaining a collection of queries: In one embodiment, a set of 1500 queries is obtained by using three different strategies (214). • 600 queries were created semi-automatically (216). Multiple employees within the organization recorded an initial set of queries based on business requirements. Subsequently, ChatGPT was used as an assistant, generating similar additional queries using the human-generated queries as seeds. 300 querysets were again generated by ChatGPT, but this time specifying a particular type of data leak. • 600 querysets were randomly selected from the publicly available dataset ign_clean.

[0049] To obtain the gold standard label, each query in the dataset was manually annotated as follows: Task T1: A label (HIGH or LOW) indicating whether the query contains sensitive data from the organization's perspective. • Task T2: If the T1 label is HIGH, the paraphrased version of the query will not contain sensitive data, and the original semantics will be preserved as much as possible. Task T3: If the T1 label is HIGH, create a set of labels indicating the data leakage type described in the query. In T1, each query was annotated by two annotators, and the inter-annotator agreement score, based on Cohen's Kappa statistics, was 0.875. Disagreements were resolved through discussion. Of the 1500 queries, 464 were identified as highly sensitive queries from a data leakage perspective. These 464 queries were manually rephrased and added to the dataset with the T1 label set to "LOW" (T2 / T3 labels are NA), resulting in a final effective dataset size of 1964 queries. Figure 3 shows how the six data leakage types are distributed, and Table I shows some examples of these annotations. Table III JPEG2026090209000013.jpg214150

[0050] The experiment confirmed that the decoder-only model does not work well at T1. For T3, Attn-BERT is the best model for both micro and macro-F1. For T2, Mistral-7B-instruct is... JPEG2026090209000014.jpg6150, along with the two most important indicators in T2, JPEG2026090209000015.jpg is the best in terms of 6150.

[0051] Results and Analysis: Table 2 shows the overall evaluation results for all tasks from the perspective of all metrics. In T1, T5-base_CL showed the best performance, followed closely by Attn-BERT. In T1, decoder-only models did not show good results. In T3, Attn-BERT was the best model in both micro and macro-F1. For T2, Mistral-7B-instruct, JPEG2026090209000016.jpg6150, along with the two most important indicators in T2, JPEG2026090209000017.jpg is the best in terms of 6150. Table III shows several examples of paraphrasing. Overall, T5-base_CL is the best model across all three tasks. Furthermore, T5-base_CL's T1 performance was observed to be consistently high across all six data leakage types. Table IV JPEG2026090209000018.jpg98150

[0052] Ablation Analysis: Ablation analysis was performed on T5-base_CL to evaluate two design choices: curriculum learning and multitask learning. The results showed that T1 and T3 performance was significantly affected when curriculum learning and multitask learning were absent. For T2, the advantages of these two design choices were not very decisive, with multitask learning being particularly prominent. However, the model trained only on T2 showed... JPEG2026090209000019.jpg11150 and In terms of JPEG2026090209000020.jpg11150, it can be observed that both lag behind T5-base_CL.

[0053] Deployment Scenario: QueryShield (this disclosure) includes all three models, namely AttnBERT, T5-base_CL, and Mistral-7B-Instruct, and is configured by the system administrator considering (i) accuracy, (ii) inference time per query, (iii) and the fine-tuning capability that T5-base_CL can use to fine-tune using incremental training data from user feedback. The default recommendation for best end-to-end accuracy is to use T5-base_CL for T1, Mistral-7B-Instruct for T2, and Attn-BERT for T3.

[0054] Long Queries: Where Mistral excels over T5 is its longer context window. Therefore, for queries longer than 512 tokens, the Mistral model is preferable for paraphrasing. In T1 / T3 using T5-base_CL and Attn-BERT, when encountering a long query, it is first split into multiple chunks, and inference is performed for each chunk. If any of these chunks is found to be sensitive, T1 predicts HIGH for the entire query, while T3 predicts the sum of the predicted leak types for all chunks.

[0055] This specification describes the subject matter of the present invention in such a way that those skilled in the art can carry out and utilize the embodiments. The scope of the embodiments of the subject matter is defined by the claims and may include other modifications that will be conjured up by those skilled in the art. Such other modifications shall be within the scope of the claims if they have similar elements that do not differ from the language of the claims, or if they include equivalent elements that differ only slightly from the language of the claims.

[0056] Embodiments of this disclosure address the unresolved problem of mitigating corporate data leaks in queries to large-scale language models. This disclosure provides a balance between access to external LLMs and the potential risk of corporate data leaks. The QueryShield platform of this disclosure sits between any external LLM and the corporate environment to not only detect sensitive data leaks in queries but also paraphrase the original query to eliminate potential data leaks. This disclosure considers several lightweight language models as part of QueryShield, enabling them to be hosted on-premises with limited hardware resources. These models are evaluated using a manually annotated dataset of 1500 queries for the tasks of detecting sensitive data leaks, paraphrasing sensitive queries, and identifying the type of data leak. Furthermore, compared to traditional methods that consider individual words for detection, this disclosure considers the entire query during detection, paraphrasing, and identification.

[0057] The scope of protection extends beyond such programs to computer-readable means having messages internally, and such computer-readable storage means includes program code means for performing one or more steps of the method when the program is executed on a server or mobile device or any suitable programmable device. The hardware device can be any type of programmable device, including any type of computer, such as a server or personal computer, or the like, or any combination thereof. The device may also include means that can be hardware means such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or a combination of hardware and software means, such as an ASIC and an FPGA, or at least one memory having at least one microprocessor and internally disposed software modules. Thus, the means may include both hardware and software means. Embodiments of the method described herein can be implemented in hardware and software. The device may also include software means. Alternatively, embodiments may be implemented using different hardware devices, such as multiple CPUs, GPUs and edge computing devices.

[0058] Embodiments described herein may comprise hardware and software elements. Software-implemented embodiments include, but are not limited to, firmware, resident software, and microcode. Functions performed by the various modules described herein may be implemented by other modules or combinations of other modules. In this specification, a computer-usable medium or computer-readable medium may be any device capable of storing, communicating, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The illustrated steps are set up to illustrate the illustrated exemplary embodiments, and it should be expected that the way in which specific functions are performed may change due to ongoing technological development. These examples are presented herein for illustrative purposes only and are not limiting. Furthermore, boundaries of functional units are arbitrarily defined herein for the convenience of explanation. Alternative boundaries may be defined as long as the specified functions and their relationships are adequately performed. Alternatives (including equivalents, extensions, modifications, and deviations of those described herein) will be apparent to those skilled in the art based on the teachings contained herein. Such alternatives are within the scope of the disclosed embodiments. Furthermore, the terms “equipped,” “possessed,” “incorporated,” and “included,” as well as other similar forms, are intended to be semantically equivalent and open-ended, in that the one or more items following any one of these terms do not mean an exhaustive list of such one or more items, nor do they mean that the list is limited to only the one or more items listed. Also note that, as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural anaphora unless the context clearly indicates otherwise. In addition, one or more computer-readable storage media may be used when carrying out embodiments consistent with the present disclosure. Computer-readable storage media means any type of physical memory capable of storing information or data that is readable by a processor.Accordingly, a computer-readable storage medium can store instructions for execution by one or more processors, including instructions for causing a processor to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible objects and exclude carrier waves and transient signals, i.e., non-transient. Examples of embodiments include random-access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, and any other known physical storage medium.

[0059] This disclosure and examples are merely illustrative, and the true scope of the disclosed embodiments shall be indicated by the appended claims.

Claims

1. A processor implementation method (200), The steps include (202) receiving input user queries associated with queries to multiple large language models (LLMs) by one or more hardware processors, (204) A step of calculating the sensitive data leak level associated with the input user query by prompting an enterprise data leak mitigation model trained on a first set of specific instructions (T1) as a prefix for the input user query using one or more hardware processors, wherein the sensitive data leak level is classified as one of (i) high and (ii) low based on associated thresholds, (206) A step of generating a plurality of paraphrased queries associated with the input user query by the one or more hardware processors, which, when the sensitive data leakage level is higher than a predefined threshold, reduces the sensitive data leakage level by preserving the semantics of the input user query and prompting the trained enterprise data leakage mitigation model, which is trained using a second set of specific instructions (T2) as a prefix to the user query, via the trained enterprise data leakage mitigation model, The steps include: simultaneously identifying the type of sensitive data breach associated with the input user query by prompting the trained enterprise data breach mitigation model based on a third set of specific instructions (T3) as a prefix to the input user query in the trained enterprise data breach mitigation model using one or more hardware processors (208); The step (210) is repeated by the one or more hardware processors until an optimal paraphrase query is generated in which the confidential data leakage level is less than a predefined threshold, Processor implementation methods, including those mentioned above.

2. The processor implementation method according to claim 1, wherein the user is permitted to query a relevant LLM from among the plurality of LLMs using the input user query only if the confidential data leakage level is below a predefined threshold.

3. The processor implementation method according to claim 1, wherein the identified type of confidential data leak associated with the input user query is presented to the user and further used to update the data leak mitigation model.

4. The aforementioned corporate data leakage mitigation model is A step of receiving multiple user queries from multiple sources, wherein the multiple sources include queries generated by user groups, ChatGPT, and queries from multiple publicly available datasets. A step of obtaining multiple training instances for tasks T1, T2, and T3 from the multiple sources, wherein T1 is for classifying the sensitive data leakage levels in the multiple user queries, T2 is for generating paraphrased queries when the classified queries have a high sensitive data leakage level, and T3 is for detecting the type of sensitive data leakage level with respect to a predefined list of sensitive data leakage levels, A step of obtaining gold standard labels for a plurality of training instances of T1, T2, and T3, wherein the gold standard labels for the training instances of T1, T2, and T3 are obtained from a plurality of associated predefined annotations, A step of fine-tuning a language model pre-trained in the training instance of T1 for K epochs, wherein each of the plurality of training instances of T1 includes input text paired with expected labeled output text obtained from a plurality of predefined annotations, and the plurality of training instances of T1 are debiased using anomaly detection techniques. A step of fine-tuning a language model for K epochs with a validation loss smaller than a predefined threshold, using a plurality of related training instances of T1 and T3, wherein the validation loss is calculated for a set of validations in each of the K epochs, and an optimal validation loss is selected over the K epochs based on a predefined validation threshold, and each of the plurality of training instances of T3 includes paired high-security data leakage level text and labeled output text generated based on a plurality of predefined annotations, A step of fine-tuning the language model with the optimal validation loss using a plurality of associated training instances of T1, T2, and T3, wherein each of the plurality of training instances of T2 includes the paired high-security data leakage level text and paraphrased output text generated based on a plurality of predefined annotations, The steps include: performing the final training of the language model for K epochs using the training instances T1, T2, and T3; A step of evaluating the trained language model using multiple metrics, The processor implementation method according to claim 1, obtained by...

5. The method according to claim 4, wherein T1 is evaluated using recall and F1 score metrics, and T3 is evaluated using micro-average and macro-average F1 scores.

6. The aforementioned T2 is (i) Calculate a cross-reference score by comparing multiple paraphrasing queries, input user queries, and gold standard paraphrasing queries. (ii) Calculate named entity leakage (NEL) as the proportion of named entities in the multiple paraphrasing queries that occur as part of false positives. (iii) Using the CRR score and the NEL, evaluate multiple paraphrased queries to preserve the original maximum semantics of query-q. and (iv) the accuracy of the label "LOW" of T1, A processor implementation method according to claim 4, as evaluated by [the specified method].

7. System (100), The system comprises at least one memory (104) for storing programmed instructions, one or more input / output (I / O) interfaces (112), and one or more hardware processors (102) operably coupled to the at least one memory (104), The one or more hardware processors (102) described above, in accordance with the programmed instructions, The steps include receiving an input user query associated with queries to multiple large language models (LLMs), A step of calculating the sensitive data leak level associated with the input user query by prompting an enterprise data leak mitigation model trained on a first set of specific instructions (T1) as a prefix for the input user query, wherein the sensitive data leak level is classified as one of (i) high and (ii) low based on an associated threshold, The steps include generating a set of paraphrased queries associated with the input user query, which, when the level of sensitive data leakage is higher than a predefined threshold, reduce the level of sensitive data leakage by preserving the semantics of the input user query and prompting the trained enterprise data leakage mitigation model with a second set of specific instructions (T2) as a prefix to the user query, The steps include simultaneously identifying the type of sensitive data breach associated with the input user query by prompting the trained corporate data breach mitigation model based on a third set of specific instructions (T3) as a prefix to the input user query in the trained corporate data breach mitigation model, The above steps are repeated until an optimal paraphrasing query is generated in which the level of confidential data leakage is below a predefined threshold, A system (100) configured to perform the following.

8. The system according to claim 7, wherein the user is permitted to query a relevant LLM from among the plurality of LLMs using the input user query only if the level of confidential data leakage is below a predefined threshold.

9. The system according to claim 7, wherein the identified type of confidential data leak associated with the input user query is presented to the user.

10. The aforementioned corporate data leakage mitigation model is A step of receiving multiple user queries from multiple sources, wherein the multiple sources include queries generated by user groups, ChatGPT, and queries from multiple publicly available datasets. A step of obtaining multiple training instances for tasks T1, T2, and T3 from the multiple sources, wherein T1 is for classifying the sensitive data leakage levels in the multiple user queries, T2 is for generating paraphrased queries when the classified queries have a high sensitive data leakage level, and T3 is for detecting the type of sensitive data leakage level with respect to a predefined list of sensitive data leakage levels, A step of obtaining gold standard labels for a plurality of training instances of T1, T2, and T3, wherein the gold standard labels for the training instances of T1, T2, and T3 are obtained from a plurality of associated predefined annotations, A step of fine-tuning a language model pre-trained in the training instance of T1 for K epochs, wherein each of the plurality of training instances of T1 includes input text paired with expected labeled output text obtained from a plurality of predefined annotations, and the plurality of training instances of T1 are debiased using anomaly detection techniques. A step of fine-tuning a language model for K epochs with a validation loss smaller than a predefined threshold, using a plurality of related training instances of T1 and T3, wherein the validation loss is calculated for a set of validations in each of the K epochs, and an optimal validation loss is selected over the K epochs based on a predefined validation threshold, and each of the plurality of training instances of T3 includes paired high-security data leakage level text and labeled output text generated based on a plurality of predefined annotations, A step of fine-tuning the language model with the optimal validation loss using a plurality of associated training instances of T1, T2, and T3, wherein each of the plurality of training instances of T2 includes the paired high-security data leakage level text and paraphrased output text generated based on a plurality of predefined annotations, The steps include: performing the final training of the language model for K epochs using the training instances T1, T2, and T3; A step of evaluating the trained language model using multiple metrics, The system according to claim 7, obtained by

11. The system according to claim 10, wherein T1 is evaluated using recall and F1 score metrics, and T3 is evaluated using micro-average and macro-average F1 scores.

12. The aforementioned T2 is (i) Calculate a cross-reference score by comparing multiple paraphrasing queries, input user queries, and gold standard paraphrasing queries. (ii) Calculate named entity leakage (NEL) as the proportion of named entities in the multiple paraphrasing queries that occur as part of false positives. (iii) Using the CRR score and the NEL, evaluate multiple paraphrased queries to preserve the original maximum semantics of query-q. and (iv) the accuracy of the label "LOW" of T1, The system according to claim 10, as evaluated by [the relevant authority].

13. One or more non-temporary machine-readable information storage media containing one or more instructions, wherein the instructions, when executed by one or more hardware processors, The steps include receiving an input user query associated with queries to multiple large language models (LLMs), A step of calculating the sensitive data leak level associated with the input user query by prompting an enterprise data leak mitigation model trained on a first set of specific instructions (T1) as a prefix for the input user query, wherein the sensitive data leak level is classified as one of (i) high and (ii) low based on an associated threshold, The steps include generating a set of paraphrased queries associated with the input user query, which, when the level of sensitive data leakage is higher than a predefined threshold, reduce the level of sensitive data leakage by preserving the semantics of the input user query and prompting the trained enterprise data leakage mitigation model with a second set of specific instructions (T2) as a prefix to the user query, The steps include simultaneously identifying the type of sensitive data breach associated with the input user query by prompting the trained corporate data breach mitigation model based on a third set of specific instructions (T3) as a prefix to the input user query in the trained corporate data breach mitigation model, The above steps are repeated until an optimal paraphrasing query is generated in which the level of confidential data leakage is below a predefined threshold, One or more non-temporary machine-readable information storage media that perform the following actions.

14. The one or more non-temporary machine-readable information storage media according to claim 13, wherein the user is permitted to query a relevant LLM from among the plurality of LLMs using the input user query only if the level of confidential data leakage is below a predefined threshold, and the identified type of confidential data leakage associated with the input user query is presented to the user.

15. The aforementioned corporate data leakage mitigation model is A step of receiving multiple user queries from multiple sources, wherein the multiple sources include queries generated by user groups, ChatGPT, and queries from multiple publicly available datasets. A step of obtaining multiple training instances for tasks T1, T2, and T3 from the multiple sources, wherein T1 is for classifying the sensitive data leakage levels in the multiple user queries, T2 is for generating paraphrased queries when the classified queries have a high sensitive data leakage level, and T3 is for detecting the type of sensitive data leakage level with respect to a predefined list of sensitive data leakage levels, A step of obtaining gold standard labels for a plurality of training instances of T1, T2, and T3, wherein the gold standard labels for the training instances of T1, T2, and T3 are obtained from a plurality of associated predefined annotations, A step of fine-tuning a language model pre-trained in the training instance of T1 for K epochs, wherein each of the plurality of training instances of T1 includes input text paired with expected labeled output text obtained from a plurality of predefined annotations, and the plurality of training instances of T1 are debiased using anomaly detection techniques. A step of fine-tuning a language model for K epochs with a validation loss smaller than a predefined threshold, using a plurality of related training instances of T1 and T3, wherein the validation loss is calculated for a set of validations in each of the K epochs, and an optimal validation loss is selected over the K epochs based on a predefined validation threshold, and each of the plurality of training instances of T3 includes paired high-security data leakage level text and labeled output text generated based on a plurality of predefined annotations, A step of fine-tuning the language model with the optimal validation loss using a plurality of associated training instances of T1, T2, and T3, wherein each of the plurality of training instances of T2 includes the paired high-security data leakage level text and paraphrased output text generated based on a plurality of predefined annotations, The steps include: performing the final training of the language model for K epochs using the training instances T1, T2, and T3; A step of evaluating the trained language model using multiple metrics, One or more non-temporary machine-readable media according to claim 13, obtained by...