Data processing system and method for masking sensitive data

The data processing system addresses data security challenges by masking sensitive information within user queries, ensuring accurate and secure interactions with third-party tools by generating multiple prompts and validating responses, thus maintaining data integrity and confidentiality.

JP7911467B2Active Publication Date: 2026-08-26ACCENTURE GLOBAL SERVICES LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024147536
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-08-31
Filing Date
2024-08-29
Publication Date
2026-08-26
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

Existing data management systems face challenges in protecting confidential information from unauthorized access and exposure during interactions with third-party platforms, often leading to data leakage and inaccurate responses due to traditional security measures like encryption and access control, which fail to maintain data integrity and confidentiality.

Method used

A data processing system and method that masks sensitive information within user queries before interaction with third-party tools, using techniques like anonymization and contextual replacement to generate multiple prompts, aggregate responses, and validate results against confidential information, ensuring accurate and secure data retrieval.

Benefits of technology

Maintains data security and privacy while preserving the accuracy of responses by masking sensitive information, reducing computation time, and preventing hallucinations in third-party model outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007911467000001
    Figure 0007911467000001
  • Figure 0007911467000002
    Figure 0007911467000002
  • Figure 0007911467000003
    Figure 0007911467000003
Patent Text Reader

Abstract

To provide a method, data processing system and computer-readable storage medium for masking sensitive data for fetching information from foundation models.SOLUTION: The method comprises: receiving from a user a query pertaining to a request for information; generating prompts based on the query from the user, by masking sensitive information in the query; receiving responses from foundation models in response to inputting the prompts; generating a common result set based on the responses; generating a response by validating the common result set with the sensitive information and the query; generating a user response by supplementing the response with the sensitive information; and providing the user response to the user in response to the query.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of data processing, and more particularly to a data processing system and method for masking confidential data.

Background Art

[0002] In modern data management systems, it is of utmost importance to ensure the security and confidentiality of confidential information during interactions with third-party platforms such as the underlying model. Protecting confidential data from unauthorized access or exposure while facilitating the process of seamlessly acquiring and presenting data is becoming increasingly emphasized. What stands out in addressing such issues is the continuous need for a robust framework that can protect confidential information and ensure its integrity and confidentiality are maintained throughout various interactions between users and platforms.

Summary of the Invention

Means for Solving the Problems

[0003] The present disclosure generally relates to the field of data processing, and more particularly to a data processing system and method for masking confidential data to fetch information from the underlying model.

[0004] The implementation of this disclosure generally focuses on masking sensitive data while retrieving information. Sensitive information is identified and masked before the information request reaches a third-party tool, ensuring that enterprise data remains secure within the enterprise environment. Advantageously, sensitive data is masked in a way that preserves its contextual and relational relevance. This ensures that external tools can provide accurate and appropriate responses without gaining access to the sensitive information. The external tools may be implemented as external foundational models. By leveraging masked sensitive data for information retrieval, the required computation time is reduced without compromising accuracy. As a result, the proposed systems and methods transform the field of foundational models by protecting sensitive data from external / third-party tools implementing the foundational model without compromising the accuracy of the model output.

[0005] In general, the innovative aspects of the subject matter described herein provide a data processing system. The data processing system includes at least one processor and at least one non-temporary processor-readable medium for storing instructions to be executed by the at least one processor. The at least one processor is configured to receive queries from a user regarding requests for information, generate a plurality of prompts based on the user queries, the plurality of prompts being generated by masking sensitive information in the queries, receive a plurality of responses from a plurality of underlying models in response to inputting the plurality of prompts, generate a common result set using the sensitive information and queries, generate user responses by supplementing the responses with sensitive information, and provide user responses to the user in response to queries.

[0006] This disclosure further describes a method for masking sensitive data during information acquisition. This disclosure further describes a non-temporary processor-readable storage medium coupled to one or more processors, in which instructions are stored, and when executed by one or more processors, the instructions cause one or more processors to perform operations in the manner described in this specification.

[0007] Naturally, the methods disclosed herein may include any combination of aspects and features described in this specification. In other words, the methods disclosed herein are not limited to the combinations of aspects and features specifically described in this specification, but may include any combination of aspects and features shown.

[0008] Details of one or more implementations of this disclosure are described in the accompanying drawings and the following description. Other features and advantages of this disclosure will become apparent from this description and drawings, as well as from the claims.

[0009] Various embodiments of this disclosure are described with reference to the following drawings.

[0010] The same reference number and name in various drawings refer to the same component. [Brief explanation of the drawing]

[0011] [Figure 1] This document provides an example environment that may be used to implement the disclosure. [Figure 2] This disclosure provides an example architecture of a data processing system that masks sensitive data during information acquisition, based on its implementation. [Figure 3] This diagram shows the process flow for masking sensitive data from the base model based on the implementation of this disclosure. [Figure 4] A block diagram illustrating an exemplary process flow for generating user responses from user queries using the implementation of this disclosure is shown. [Figure 5]This disclosure provides a detailed process flow for masking sensitive data during information acquisition, based on the implementation described herein. [Figure 6] This flowchart illustrates an exemplary method of implementing this disclosure. [Figure 7] This shows computer systems that can be used to implement knowledge systems. [Modes for carrying out the invention]

[0012] In the following description, various embodiments are shown in the figures of the accompanying drawings as examples, not as limitations. References to various embodiments in this disclosure do not necessarily refer to the same embodiment, but rather mean that there is at least one. Specific implementations and other details are taken up, but naturally, they are for illustrative purposes only. Those skilled in the art will see that other components and configurations may be used without departing from the scope of the claimed subject matter.

[0013] Any reference to “example” in this specification (such as “for example,” “an example of,” “by way of example,” or similar) shall be considered non-exclusive examples, whether explicitly stated or not.

[0014] The terms used herein generally have the common meaning in the art within the context of this disclosure and in the specific context in which each term is used. Substitute words and synonyms may be used for one or more of the terms discussed herein, and whether a particular term is detailed or discussed herein is not of particular importance. Synonyms are provided for certain terms. The inclusion of one or more synonyms does not preclude the use of other synonyms. The use of examples in any part of this specification, including examples of any terms discussed herein, is illustrative only and is not intended to further limit the scope and meaning of this disclosure or any of the examples provided. Similarly, this disclosure is not limited to the various embodiments shown herein.

[0015] Without any intention to limit the scope of this disclosure, examples of instruments, apparatus, systems, methods, and related results according to embodiments of this disclosure are given below. For the convenience of the reader, headings or subheadings may be used in the examples, but this is not intended to limit the scope of this disclosure. Unless otherwise defined, technical and scientific terms used in this specification have the meanings commonly understood by those skilled in the art in the relevant field of this disclosure. In case of any inconsistency, this document shall prevail, including definitions.

[0016] When the term "comprising" is used, it means "including, but not necessarily limited to," specifically indicating an unrestricted inclusion or attribution within the combination, group, series, and similar items described in that way.

[0017] The term "a" means "one or more" unless the context clearly indicates a single component.

[0018] "First", "second", etc. are labels for distinguishing components or blocks with similar names in other respects, but do not imply any order or numerical limitation.

[0019] "And / or" for two candidates means one or both of the candidates described (where "A and / or B" targets only A, only B, or both A and B together), and when presented with three or more candidates described, it means any one of the individual candidates, all candidates together, or some combination of candidates less than all candidates. When A through N are candidates, the phrase in the form "at least one of A...and N" means "and / or" for the candidates described (e.g., at least one A, at least one N, at least one A and at least one N, etc.).

[0020] Furthermore, it should be noted that in some alternative implementations, the functions / operations shown may occur in an order different from that shown in the drawings. For example, depending on the related functionality / operation, two steps disclosed or shown consecutively may actually be performed substantially simultaneously, or each step may sometimes be performed in reverse order.

[0021] To enable a full understanding of the embodiments, specific details are set forth in the following description. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these specific details. For example, the system may sometimes be shown in block diagrams so as not to obscure the embodiments with unnecessary details. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to prevent obscuring the exemplary embodiments.

[0022] This specification and the drawings are to be regarded in an illustrative rather than a restrictive sense. Nevertheless, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the invention as set forth in the claims.

[0023] The emergence of new third - party tools and platforms such as foundation models in the new era has brought both opportunities and challenges in protecting confidential information. Traditional tools and platforms utilize advanced algorithms and data - processing methods to enhance the accuracy and relevance of the acquired data. Despite such progress, the integration of confidential data with such tools brings inherent risks, including unauthorized access and data leakage, thereby necessitating robust security measures.

[0024] Traditional methods of protecting confidential information during interactions with any tool or platform such as a foundation model have typically relied on encryption and access - control mechanisms. However, such an approach often faces challenges in ensuring comprehensive protection against unauthorized access or inadvertent exposure. As the digital ecosystem evolves and data interactions become increasingly complex, a robust security framework that can effectively mask or conceal confidential information from external entities is needed.

[0025] Recent advancements in cybersecurity have once again highlighted the vulnerabilities inherent in traditional security measures when interfacing with external systems. To mitigate risks during the process of information acquisition and presentation and defend confidential data, it has become essential to incorporate advanced masking methods and secure communication protocols. Previously, masking of confidential information often involved methods such as deletion or replacement with general - purpose variables, resulting in what is commonly referred to as hallucination, where incorrect responses and untrustworthy data outputs occur, potentially losing the reliability of the acquired data.

[0026] Furthermore, the proliferation of interconnected digital platforms and the rise of the distributed computing paradigm have exacerbated the complexities surrounding data security and privacy. Seamless integration of secure masking methods must be ensured while maintaining the efficiency and accuracy of information retrieval operations. Addressing these challenges requires innovative approaches to strengthen the security posture of data interactions while adhering to applicable regulatory requirements and industry standards for data protection.

[0027] The implementation of the disclosed subject involves generating responses to user queries and generating multiple prompts by masking sensitive information within the queries. The generated multiple prompts are then used to aggregate responses from various underlying models. Next, a common result set is generated from the aggregated responses and validated against the user query (the original query received from the user) and sensitive information. Subsequently, the validated information is integrated to create a user response, which is then provided to the user as a consistent response to the user query. This approach maintains data security and privacy without compromising the accuracy and reliability of the model output.

[0028] Figure 1 shows an exemplary environment 100 that may be used to implement the present disclosure. In some examples, the exemplary environment 100 allows for the masking of sensitive information when retrieving information about user queries.

[0029] As shown in Figure 1, the exemplary environment 100 includes computing devices 102 and 104, a backend system 106, and a network 108. In some examples, computing devices 102 and 104 are used by individual users 110 and 112 to log in to and interact with a computing platform running applications according to the implementation of this disclosure. Examples of computing devices 102 and 104 may include desktop computing devices, smartphones, laptops, tablets, voice-enabled devices, and / or similar. Implementations of this disclosure may be implemented using any suitable type of computing device. In some examples, computing devices 102 and 104 may each include a web browser application running on the device, which may be used to display one or more web pages of the computing platform running the applications. In some examples, computing devices 102 and 104 may each display one or more graphical user interfaces (GUIs) that enable individual users 110 and 112 to interact with the computing platform.

[0030] In some examples, network 108 includes a local area network (LAN), a wide area network (WAN), the internet, or a combination thereof, and connects websites, computing devices 102 and 104, and a backend system 106. In some examples, network 108 may be accessed over wired and / or wireless communication links. For example, a computing device such as a smartphone may access network 108 using a cellular network.

[0031] In some examples, one or more of the backend systems 106 may be implemented as on-premises systems operated by an enterprise or third party engaged in cross-platform interaction and data management. In some examples, the backend systems 106 may be implemented as off-premises systems (e.g., cloud or on-demand) operated by an enterprise or a third party acting on behalf of the enterprise. In some examples, one or more of the backend systems 106 may be implemented in a cloud environment. For simplicity, the backend systems 106 shown in Figure 1 may be a cloud environment intended to represent various forms of servers, including web servers, application servers, proxy servers, network servers, server pools, and / or similar.

[0032] In some examples, the backend system 106 includes one or more data processing systems 114 that host components for information retrieval (e.g., knowledge embeddings). Furthermore, the data processing systems 114 receive requests from users 110 and 112 seeking services provided by the data processing systems 114 through individual computing devices 102 and 104. In response to received requests, the data processing systems 114 provide the requested services to computing devices 102 and 104 over the network 108. Requests received from users 110 and 112 through individual computing devices 102 and 104 may be user queries. The data processing systems 114 may be configured to allow the use of multiple underlying models while masking sensitive data from multiple underlying models, ensuring that accurate responses to user queries are generated without compromising data security. The interaction between the data processing systems 114 and users 110 and 112 may be inherently conversational, including conversational queries and conversational responses to those queries.

[0033] According to the implementation of this disclosure, the data processing system 114 may be configured to mask sensitive data while the user interacts with the underlying model, while ensuring accuracy in information retrieval. Numerous examples demonstrating the masking of sensitive data to ensure data security during information retrieval are described in detail, along with the following drawings.

[0034] Figure 2 shows an illustrative architecture 202 of a data processing system 114 that masks sensitive data during information acquisition, as implemented in this disclosure. In one example, as shown in Figure 2, the data processing system 114 receives one or more queries and generates content / responses for one or more queries. The responses may include, but are not limited to, text, images, audio, video, and / or similar, as responses to queries. One or more queries may include prompts to generate responses to queries about the world, or a particular system, or an enterprise, or a business, etc. The generated responses may include additional information, feedback / input, and / or similar, necessary to answer the queries.

[0035] The data processing system 114 includes a knowledge base 204, a user interface (UI) / user experience (UX) module 206, and a processing engine 208. The knowledge base 204 can be described as a structured repository or database associated with the data processing system 114. The knowledge base 204 may incorporate various knowledge representation schemes, such as ontologs, taxonomies, or semantic networks, to encode and organize information in a machine-understandable format, thereby enabling advanced retrieval, reasoning, and argumentation capabilities. Furthermore, the knowledge base 204 may leverage advanced techniques, including natural language processing, machine learning, and knowledge engineering methods, to enhance the processes of knowledge acquisition, updating, and improvement, ensuring continuous relevance and adaptability to evolving needs and circumstances.

[0036] In one embodiment of this disclosure, the knowledge base 204 stores user queries 210, masking guidelines 212, enterprise data 214, persona rules 216, responses 218, metadata 220, information about the enterprise or business, and additional information (not shown) about the data processing system 114. User queries 210 refer to one or more queries received from a user. User queries 210 may relate to requests to obtain answers or responses by processing information related to the enterprise. Furthermore, user queries 210 may include sensitive information.

[0037] Masking guidelines 212 may be described as guidelines or techniques relating to the sequential steps and decision-making processes involved in masking sensitive data within user queries 210. Examples of masking guidelines / techniques 212 include anonymization, tokenization, data scrambling, dynamic data masking (DDM), format-preserving encryption, static data masking, and similar methods.

[0038] Enterprise data 214 may include information encompassing data, facts, and insights obtained from various data sources related to the enterprise. Such information may be organized in a consistent manner to support decision-making, problem-solving, and system operation within the data processing system 114. Enterprise data 214 may be temporal, numerical, and / or textual.

[0039] Enterprise Data 214 includes enterprise information that is expressed across multiple hierarchies and incorporates domain terminology. Hierarchies refer to structured levels or layers within data about an organization, defining relationships and classifications from broader categories to specific details. Examples of hierarchies may include organizational structure levels (departments, teams, roles, etc.), product classifications (categories, subcategories, products, etc.), and geographical divisions (regions, countries, cities, etc.). Domain terminology refers to specialized language, vocabulary, or jargon specific to a particular field or industry that facilitates accurate communication and understanding within that context. Examples of domain terminology may include medical terminology (diagnosis, treatment, patient care, etc.), legal terminology (e.g., jurisdiction, litigation, contracts, etc.), and financial terminology (assets, liabilities, dividends, etc.).

[0040] In this context, enterprise data 214 may be stored across multiple sub-databases within the knowledge base 204, with each of the sub-databases relating to a domain or solution concerning the enterprise. For example, an HR database for the HR domain, a fraud database for the fraud prevention domain, a collection database, a contract financial management database, and so on. Each of the sub-databases may have a specific data table with relational interactions based on the specific data stored within each of the sub-databases.

[0041] Persona rule 216 refers to a rule regarding the accessibility of information about a user persona. For example, a developer may not necessarily need access to information related to the hiring of other employees, but an administrator or HR employee may have access to it. Response 218 refers to a response to a query generated by one of several foundational models or data processing systems 114. Response 218 may include answers to user queries. Metadata 220 can be described as descriptive information about data stored in the knowledge base 204, including user queries 210, masking guidelines 212, enterprise data 214, persona rules 216, and responses 218. Furthermore, each query and response may have relevant metadata that improves the accuracy of response 218 by providing context to the query or response.

[0042] Furthermore, knowledge base 204 includes additional data, such as raw data about the enterprise. This raw data may be used for training purposes. Typically, the raw data includes hierarchies, labels, categorical variables, and similar elements. Such components of the raw data provide insights into the data, which may then be used to implement specific masking methods based on each scenario for masking sensitive data.

[0043] The UI / UX module 206 can be defined as a module that designs and manages the user interface (UI) and user experience (UX) that users use to interact with the data processing system 114. The UI / UX module 206 may often utilize the principles of human-computer interaction (HCI) and graphic design to integrate various technologies and frameworks and optimize the visual layout, interactive elements, and overall usability.

[0044] In some examples, the UI / UX module 206 may represent one or more front-end components / interfaces 222a-222n of a chatbot that may run on one or more of the computing devices 102 and 104 to enable the receiving of user queries and the provision of one or more user responses to those user queries. In some examples, user queries may be received by a variety of modalities, including but not limited to question input to the chatbot, requests provided through a graphical user interface (GUI), email, and / or similar.

[0045] The processing engine 208 is configured to process queries received through the UI / UX module 206 using multiple foundational models. These foundational models can be described as general-purpose generative artificial intelligence (GAI) models, such as large-scale deep learning neural networks. A large-scale deep learning neural network may be trained using a broad range of generalized, unlabeled training data, and it may perform a number of common tasks. Examples of tasks may include generating text, generating images, conversing in natural language, generating video, generating audio, and / or similar. In some examples, applications may be built that interact with the foundational models. In some examples, multiple foundational models may be used to perform various functionalities of the application.

[0046] The foundational model may include, for example, a Large Language Model (LLM), which is a form of GAI that can be used to generate text for various use cases. In some examples, the LLM may be integrated with a digital assistant (e.g., a chatbot) to provide text responses to user input / queries in place of traditional rule-based systems. An LLM can be described as an advanced language model trained using deep learning methods on a massive amount of text data. Typically, an LLM generates human-like text and performs various natural language processing (NLP) tasks (e.g., translation, question answering, and / or similar). In some examples, LLM refers to a model that uses deep learning methods and has multiple parameters. LLMs typically capture complex patterns in language and produce text that is often indistinguishable from that written by a human. The generated text can be processed through deep learning architectures such as recurrent neural networks (RNNs), transformer models, and / or similar.

[0047] While the implementation of this disclosure is described in more detail with non-restrictive reference to LLM as an illustrative foundational model, it is believed that the implementation of this disclosure can be realized using any suitable foundational model, machine learning (ML) model, or artificial intelligence (AI) model. Such a model may generate content / responses based on any suitable modality (e.g., text, audio, images, video, and / or similar). In some examples, the responses may correspond to one or more tasks expressed by conversational queries.

[0048] In some examples, the foundation model may be provided by one or more third parties. In some examples, the foundation model may be provided via the data processing system 114. The foundation model receives requests / queries and provides responses to those queries. For example, a question / request for information may be received as a query through an Application Programming Interface (API).

[0049] The processing engine 208 includes one or more processors 224, specific modules 226, 228, a prompt generation module 230, an information acquisition module 232, a persona module 234, and a rule engine 236.

[0050] The processor 224 may include, for example, a microprocessor, a microcomputer, a microcontroller, a digital signal processor, a central processing unit, a state machine, a logic circuit, and / or any device that manipulates data or signals based on operational instructions. Among several capabilities, the processor 224 is configured to fetch and execute computer-readable instructions in memory coupled to the data processing system 114 in an operable manner in order to mask sensitive information.

[0051] The specific module 226 identifies sensitive information within the user query 210 and sensitive information related to the user query 210. The specific module 226 is trained using raw data before deployment. The raw data includes details such as hierarchy, labels, categorical variables, and similar, and the specific module 226 essentially understands the relationships between different data points in the raw data. The specific module 226 understands at least one of linear relationships between different data points or nonlinear relationships between different data points. Advantageously, by understanding nonlinear relationships, the specific module 226 can quickly and accurately identify the appropriate sensitive information. Here, the specific module 226 can utilize multiple relationship identification methods to identify sensitive information, including statistical correlation, Pearson correlation, ontology correlation, and similar methods. This allows the specific module 226 to gain a detailed understanding of the enterprise data 214. Once deployed, the specific module 226 identifies sensitive information using fuzzy matching.

[0052] Fuzzy matching refers to a technique used in the information retrieval and data validation process that allows for approximate matching, accepting variations such as misspellings, synonyms, or similar patterns, even if an exact match is not possible. For example, if enterprise data 214 contains the customer name "Smith & Co.", the identification module 226 can effectively integrate and manage customer information by matching this with similar entries such as "Smith and Company" or "Smith & Company" using the fuzzy matching method. The identification module 226 then provides the masking module 228 with the sensitive information identified from enterprise data 214, along with details indicating the relationships of the sensitive information.

[0053] The masking module 228 masks sensitive information identified by the specific module 226 within the user query 210. The masking module 228 masks sensitive information using one or more masking methods or masking guidelines 212. One or more masking methods / guidelines 212 are selected based on the identified sensitive information and its correlation and context with respect to the enterprise data 214.

[0054] The masking module 228 masks sensitive information using a variety of techniques, including but not limited to anonymizing sensitive data, replacing sensitive information with contextual and correlational variables, and obfuscating multiple ranges, in order to ensure data privacy while leveraging accurate responses from multiple underlying models. For example, if the identified sensitive information in user query 210 is the term "North America," the masking module 228 may replace this term with "first American region" so that the fact that the region is in the United States is preserved without introducing specificity that would affect privacy. In this way, the masking module 228 generates a masked query by masking the sensitive information in the query. The variables used in place of sensitive information maintain contextual and relational relevance and ensure the accuracy of the response generated in response to the masked query.

[0055] By masking sensitive information within the user query, the prompt generation module 230 generates multiple prompts to retrieve information from multiple base models. These multiple prompts are generated based on the masked query produced by the masking module 228. Advantageously, the prompt generation module 230 generates multiple prompts that repeat the user query in various natural language styles to verify the accuracy and completeness of the responses received from the multiple base models.

[0056] Typically, foundational models often produce hallucinations and contradictory answers, resulting in inaccurate responses and making them unsuitable for use. The functionality of the prompt generation module 230 provides a solution to such hallucinations by providing multiple prompts, so that multiple foundational models question themselves and the responses they provide are checked against each other. For example, if the user query is "What is the sales projection for California team?", the masked query may be "What is the sales projection for first US team?", and the prompt generation module 230 may generate relevant prompts such as "please provide sales projection for first US team", "first US team sales projection 2025 Q1", "sales projection US teams for FY25", and similar ones.

[0057] Upon receiving multiple prompts, the information acquisition module 232 retrieves appropriate information from the knowledge base 204 and multiple base models to generate a user response. In one embodiment, the information acquisition module 232 utilizes multiple base models to generate responses to multiple prompts. In some cases, all of the multiple prompts are provided to each of the multiple base models. In other cases, one or more of the multiple prompts may be provided to each of the multiple base models. During operation, any permutation or combination of multiple prompts and multiple base models may be implemented by the information acquisition module 232.

[0058] Furthermore, the information acquisition module 232 retrieves appropriate information from the enterprise data 214 to supplement multiple responses. The enterprise data 214 includes sensitive information related to the query. The enterprise data 214 may also include sensitive information relating to any enterprise, operation, service, facility, or similar. In one example, if the data processing system 114 is used by a government, the enterprise data 214 (and its extension, sensitive information) may relate to voter databases, population information, citizen details, and similar. In another example, if the data processing system 114 is used by a hospital, the enterprise data 214 (and its extension, sensitive information) may relate to medical records, healthcare information, identity information, and similar.

[0059] Since the enterprise data 214 includes hierarchy and domain terminology, when unmasking the response to generate a user response, the information acquisition module 232 maps the domain terminology to multiple foundational models. During preprocessing, the specific module 226 trains / tunes multiple foundational models based on the multiple domain terminology. The specific module 226 may train separate foundational models for each domain. For example, the specific module 226 may train / tune a first foundational model for the medical domain and a second foundational model for the aerospace domain. The specific module 226 may utilize one or more training methods to train / tune the foundational models. One or more training methods include supervised fine-tuning, feedback methods, reinforcement learning, human feedback-based reinforcement learning, and similar methods.

[0060] Furthermore, specific module 226 generates an index that includes mappings between multiple foundational models and the domain terminology / domains to which each of the foundational models is trained / tuned. In the previous example, the index may include mappings between the first foundational model and the medical domain, and between the second foundational model and the aerospace domain. The index can be implemented as at least one of the following: a vector index, a one-to-many index, a hash map index, or a one-to-one index. During response unmasking, information acquisition module 232 maps domain terminology to multiple foundational models based on the index.

[0061] Furthermore, the information retrieval module 232 receives multiple responses and generates a common result set. A common result set refers to a set of responses that are common among multiple responses received from multiple base models. For example, suppose the following responses are received for the query "what would be the annual revenue growth for this fiscal year?": +7%, +5%, +2%, -2%, +1%, +5%, +7%, +1%, +5%. Here, positive numbers represent predicted annual revenue growth, and negative numbers represent decline. The information retrieval module 232 may also identify a common result set based on the number of occurrences of a particular response. In this example, the response "+5%" is selected as the common result set because it is repeated the most times. A simple example is given to illustrate how a common result set is generated. Actual implementations may be more complex, especially based on the user query 210.

[0062] Subsequently, the information retrieval module 232 verifies the common results using confidential information. Once the confidential information is acquired, the information retrieval module 232 verifies the common results by determining whether the common result set is true in light of the confidential information. For example, if the common result set is "the first US team showed a 40% increase in sales," the information retrieval module 232 verifies the common result set using confidential information (actual sales figures and team / region names) to ensure its accuracy. Advantageously, the information retrieval module 232 verifies the common results using confidential information, ensuring accuracy and preventing hallucination.

[0063] In some cases, the common result set may not correlate with the confidential information. In such cases, the information retrieval module 232 regenerates the common result set. For example, if the common result set is "the first US team showed a 40% increase in sales," but during verification using confidential information (actual sales figures and team / region name), the information retrieval module 232 identifies that North America (confidential information corresponding to the first US team) actually showed a 10% increase in sales, the common result set may be regenerated by selecting a different set of responses from multiple responses that more accurately reflects the enterprise data 214 in light of the user query 210. Alternatively, in such cases, the information retrieval module 232 may locally modify the common result set. In the above example, the information retrieval module 232 may locally change the common result set to "the first US team showed a 10% increase in sales."

[0064] Furthermore, the information acquisition module 232 generates a response based on the validated common result set. The information acquisition module 232 generates a response by supplementing the (previously validated) common result set with sensitive information. Advantageously, supplementing multiple responses with sensitive information provides more context to the complete (supplemented) response and ensures accuracy. Furthermore, if the information acquisition module 232 determines that a response does not correlate with the sensitive information, it may identify an alternative response by validating the common result set based on the sensitive information.

[0065] Once a common result set is generated, the information acquisition module 232 generates a user response by supplementing the response with confidential information. Considering the example above, the user response might be, "The North American sales team showed a 10% increase in sales, reaching $25,000 in sales for FY24." The user response is provided to the user via the UI / UX module 206.

[0066] Persona module 234 validates the user, and only if validation is successful is the user response provided to the user. Persona module 234 is coupled to information retrieval module 232 for communication and provides information about user validation to information retrieval module 232. Persona module 234 fetches role information corresponding to the user from enterprise data 214. Persona module 234 then determines, based on persona rule 216, whether the user has access to information related to the user query. If it is determined that the user should not have access to the user response, persona module 234 sends a notification to the user informing them that they do not have access to the requested information. For example, if a developer tries to access employment information in a colleague's HR record, persona module 234 may send a notification saying, "you do not have access to this information."

[0067] The rule engine 236 is used to generate rules to prevent fraudulent activity in the enterprise. The rule engine 236 automatically identifies fraudulent patterns in the enterprise data 214, generates individual processor-readable masking rules to prevent fraudulent activity, and executes them. During preprocessing, the specific module 226 trains one or more classification models on the fraudulent activity data so that one or more classification models can identify patterns in the fraudulent activity in the data. One or more classification models may use any classification or pattern recognition method to identify patterns in the fraudulent activity data. The fraudulent activity data may include historical data related to fraud, real-time data related to fraud, enterprise data related to fraud from known fraud cases, general-purpose data related to fraud, and similar data. Patterns in fraudulent activity in the data may be represented as anomalies in the data.

[0068] Alternatively, a specific module 226 may train one or more classification models in real time on data related to fraudulent activity. Here, training may be performed periodically, for example, monthly, weekly, quarterly, or annually. Advantageously, by training one or more classification models in real time, the models are kept up-to-date with known fraudulent events, new fraudulent patterns, or techniques. This allows the classification models to constantly acquire information about new fraudulent techniques and identify patterns for which one or more masking rules should be generated for fraud prevention.

[0069] Since the rule engine 236 has access to the knowledge base 204, the rule engine 236 continuously processes the enterprise data 214. When the rule engine 236 identifies a fraudulent activity pattern, it generates one or more masking rules (logic) to counteract (prevent) that fraudulent activity pattern. Here, the rule engine 236 utilizes one or more classification models to identify fraudulent activity patterns in the enterprise data 214. Once the rule engine 236 identifies a fraudulent activity pattern, it translates the identified fraudulent pattern in the data into logic related to the fraudulent activity. Furthermore, to identify similar fraud that has occurred in the past, the logic related to the fraudulent activity may be mapped to past occurrences of fraudulent activity. Such countermeasure logic derived from similar past fraud may be used to generate countermeasure logic for current fraudulent activity.

[0070] The rule engine 236 may utilize one or more grid search methods to generate countermeasure logic, and therefore one or more masking rules. Grid search refers to processing data by intelligently applying a vast number of permutations and combinations to data in a hyperparameter search space. Examples of grid search methods include utilizing at least one of Bayesian optimization, stochastic optimization, or tree-structured Parzen estimator (TPE) methods. In some cases, the rule engine 236 may utilize a pre-trained foundational model that has been previously trained for logic generation to generate countermeasure logic, where the generation of countermeasure logic includes the generation of rule conditions and actions by the pre-trained foundational model. The rule engine 236 uses these rule conditions and actions to generate one or more masking rules.

[0071] Subsequently, the rule engine 236 executes one or more masking rules within the data processing system 114 to prevent fraudulent activity. For example, if a fraudulent pattern is found in foreign banking transactions, the rule engine 236 may generate countermeasure logic that flags future foreign banking transactions (those after the fraudulent pattern has been found in foreign banking transactions). In this context, the rule engine 236 may execute a masking rule that requires the user to provide additional identity verification for depositing or withdrawing money from abroad. Advantageously, by detecting fraudulent activity in a timely manner and automatically generating and executing masking rules, it is ensured that future fraudulent activity is prevented and that fraudsters cannot exploit the data processing system 114.

[0072] The data processing system 114 may be used to mask data while utilizing an underlying model in at least one domain, including revenue and probability forecasting, medical billing management, medical fraud detection, talent acquisition and management, banking, financial portfolio management, intellectual property management, government operations, social services management, defense services management, and similar.

[0073] Figure 3 shows a block diagram 300 illustrating the process flow for masking sensitive data from the base model using the implementation of this disclosure.

[0074] Block diagram 300 shows a query 302 being supplied to a secure environment 304. Query 302 refers to a user query provided by an enterprise user. The secure environment 304 refers to a controlled computing infrastructure configured to protect sensitive information from unauthorized access or exposure, particularly in interactions involving third-party tools such as the underlying model 312. Thus, queries or prompts provided to the underlying model 312 are masked to protect sensitive information, and furthermore, responses or model outputs received from the underlying model 312 do not contain any sensitive information. Sensitive information is included in the model output to generate the user response 310 before being provided to the user.

[0075] As shown in Figure 3, query 302 is input to the secure environment 304 via the data processing system 114. The processing engines 208 / 306 and knowledge bases 204 / 308 are interconnected for communication and included in the secure environment 304. The secure environment 304 may refer to the internal environment of the enterprise. Advantageously, sensitive data concerning the enterprise is kept within the secure environment 304 and masked from the underlying model 312.

[0076] Confidential information is masked by the processing engine 306. The processing engine 306 masks the confidential information by replacing it with variables that maintain contextual and relational relevance. In this way, the underlying model 312 provides an accurate response and prevents hallucination in the underlying model 312 by sharing contextual and relational relevance between the masked / protected confidential data.

[0077] In one example, if query 302 is "What market units are projected to have negative growth next quarter?", the query is provided to a secure environment 304. Within the secure environment 304, the processing engine 306 preprocesses the query to identify sensitive information. The sensitive data in this query may include "negative growth", "market units", and "quarter / year information". The processing engine 306 then masks the sensitive information to generate multiple responses, which are sent to the underlying model 312. The underlying model 312 generates a natural language response to query 302 and shares it with the processing engine 306. The processing engine 306 then extracts appropriate enterprise data 214 from the knowledge base 308 and supplements it in the response to generate a user response 310. The appropriate enterprise data 214 may also relate to sensitive data. In this way, user response 310 could be, "The Northeast, India, and ANZ market units are projected to have negative growth next quarter."

[0078] Figure 4 shows a block diagram 400 illustrating an exemplary process flow for generating user responses from user queries using an implementation of the present disclosure.

[0079] As shown in Figure 4, the data processing system 114 receives queries 402 / 302 / 210 from a user, processes them, and generates responses (singular or plural) 420 / 310 / 218. The user may be an employee, owner, merchant, vendor, and similar entity related to the enterprise. Furthermore, query 402 may be a request for information about the enterprise. In addition, the query can be composed (i.e., expressed) in any natural language recognized by the data processing system 114.

[0080] Subsequently, the data processing system 114 generates a series of prompts 404a, 404b, 404c, ... 404n based on the user query 402. These prompts 404a, 404b, 404c, ... 404n are generated by masking sensitive information within the query 402. During operation, the data processing system 114 replaces sensitive information in the query 402 with variables based on one or more masking criteria. A variable refers to a symbol or placeholder that represents a value or entity in at least one of the domains of mathematics, science, programming, and / or computing. Furthermore, the variables represent contextual and relational relationships to the sensitive information. In this way, the underlying models 406a, 406b, 406c, ... 406n accurately understand relational dependencies within the enterprise data 214, which leads to accurate responses that reduce hallucination.

[0081] During masking, the data processing system 114 identifies technical terms belonging to the data-specific structure and replaces them with explicit variable indicators consistent with the enterprise data 214. In this regard, the data processing system 114 also takes into account misspellings and alternative names. For example, if query 402 is "For North America what are the most dilutive market units to our revenue growth at fiscal year 2023?", the masked query may be expressed as "For MarketX, what are the most dilutive market units to our revenue growth at fiscal year ValueX?". This allows multiple foundational models that lack knowledge of the hierarchical details and structure in the enterprise data 214 to correctly interpret the relationship between sensitive information and query 402. For example, confidential information concerning the client group "South CMT" can be understood as a Client Group segment through multiple underlying models.

[0082] Multiple prompts 404a, 404b, 404c, ... 404n refer to specific input queries / statements provided to multiple underlying models to generate responses 408a, 408b, 408c, ... 408n. Multiple prompts 404a, 404b, 404c, ... 404n function as instructions that guide the models in processing and producing appropriate information based on the nature and structure of the prompts themselves. In operation, multiple prompts 404a, 404b, 404c, ... 404n contain multiple representations of query 402.

[0083] When multiple base models 406a, 406b, 406c, ... 406n are prompted with various versions of the same query, the likelihood of hallucination is greatly reduced. In one example, if query 402 is "Which market has the highest revenue projection?", the first prompt 404a could be "The highest projected revenue was made for which market?", and the second prompt 404b could be "In which market would I expect the highest revenue?". In another example, if query 402 is "Which client groups are responsible for driving the revenue growth of our three most impactful services in the North American market?", the first prompt may contain query 402 with an inaccurate filter, the second prompt may contain the accurate query 402, and the third prompt may contain query 402 with an inaccurate aggregation.

[0084] In response to the input of multiple prompts 404a, 404b, 404c, ... 404n, multiple base models 406a, 406b, 406c, ... 406n send multiple responses 408a, 408b, 408c, ... 408n to the data processing system 114. The data processing system 114 then generates a common result set based on the multiple responses 408a, 408b, 408c, ... 408n. The common result set is the set of responses that are most frequently observed among the multiple responses 408a, 408b, 408c, ... 408n received from the multiple base models 406a, 406b, 406c, ... 406n. In some cases, the common result set may also be validated using query 402.

[0085] Based on the common result set, enterprise information 410a, 410b, 410c, ... 410n for each of the multiple responses 408a, 408b, 408c, ... 408n of the common result set is extracted from the knowledge base 204 / 308 associated with the data processing system 114. The enterprise information may include confidential information. For example, if query 402 concerns revenue growth in various markets, and the first prompt includes query 402 with an inaccurate filter, the first information may concern revenue growth in the "Midwest RES, Northeast H&PS" region; if the second prompt includes an accurate query 402, the second information may concern revenue growth in the "South RES, Midwest PRD, Canada RES" region; and if the third prompt includes query 402 with an inaccurate aggregation, the third information may concern revenue growth in the "Canada H&PS, Canada RES, South RES" region.

[0086] The data processing system 114 generates a response 412 by validating a common result set using confidential information and query 402. In this context, the data processing system 114 executes each of the multiple responses 408a, 408b, 408c, ... 408n in a virtual environment by adding (i.e., including) confidential information within responses 408a, 408b, 408c, ... 408n. In this way, responses 408a, 408b, 408c, ... 408n are checked against query 402 and enterprise information 410a, 410b, 410c, ... 410n to ensure accuracy.

[0087] A virtual environment refers to a simulated / software-based environment that mimics real-world functionality. Virtual environments allow testing and deployment of actions that could occur in the real world without the need to implement them in the real world. This is advantageous because it allows for increased accuracy without compromising on the increased involvement of computing resources. Since this is done in a virtual environment, lessons learned from validation can be used to select and improve response 412. Thus, responses 408a, 408b, 408c, ... 408n are executed by validating response 412 against the data repository (knowledge base) and query 402.

[0088] In some cases, responses 408a, 408b, 408c, ... 408n are each provided with a score based on their individual accuracy when performed. Response 412 is selected based on which of the multiple responses 408a, 408b, 408c, ... 408n received the highest score compared to the others.

[0089] Subsequently, the data processing system 114 generates a user response 414 by supplementing response 412 with confidential information. The term "supplementing" refers to replacing data that has been masked or otherwise concealed with the original confidential data, or including confidential information in response 412, or adjusting response 412 in light of the confidential information. Based on verification, the data processing system 114 provides the user response 414 to the user.

[0090] Before providing a user response 414 to the user, the data processing system 114 verifies the user based on one or more predetermined criteria. These predetermined criteria refer to specific conditions or requirements set by the enterprise to verify the user's eligibility or authority. These predetermined criteria relate to the user's persona. A persona refers to a profile or characterization of a user based on attributes such as roles, preferences, behaviors, or other distinctive elements. Persona information includes roles, locations, interests, ethnicity, teams, groups, and similar.

[0091] The data processing system 114 is coupled to the knowledge bases 204 / 308 in a communicative manner, thus providing constant access to persona or role-based information for each user associated with the enterprise (and, as an extension, the data processing system 114). In this context, the data processing system 114 validates queries 402 and user responses 414 against one or more persona rules. Persona rules refer to guidelines or standards developed based on a user's role or profile within the enterprise, ensuring that queries 402 and user responses 414 adhere to specified standards or permissions. For example, persona rules may define which departments or hierarchical roles have access to specific enterprise-related information, such as employment records.

[0092] In this regard, the data processing system 114 shares the user response 414 with the user if the user is permitted to access appropriate information regarding query 402 and user response 422. Alternatively, if the user is not permitted to access the information, the data processing system 114 shares a notification with the user informing them that they cannot access the information. The level of facilitating sensitive data in the user response 414 should be determined by persona information so that the user's role / privileges are taken into consideration when generating the user response 414 from the selected response 412. Therefore, separate user responses corresponding to multiple personas may be generated from the same selected response 412.

[0093] Furthermore, the user response 414 may be provided to the user based on one or more masking rules. A masking rule refers to a specific rule that, when implemented, adds another level of security by masking sensitive data to protect it from malicious activity. In this context, the data processing system 114 identifies patterns in the enterprise data 214 associated with the data processing system 114. A pattern refers to a recognizable sequence or behavior that indicates malicious activity aimed at illegally obtaining sensitive information. These patterns typically involve repeated methods or means used to exploit vulnerabilities in security systems or protocols, enabling unauthorized access to sensitive data.

[0094] The data processing system 114 generates one or more masking rules based on identified patterns. For example, if it is identified that a fraudster is attempting to log in to a bank customer's account via a foreign banking transaction and attempt to empty it, a masking rule may be implemented to verify the user's identity based on biometric information or a security secret question. Advantageously, the data processing system 114 automatically executes one or more masking rules to ensure enhanced security. As a result, this saves costs and time, because the process of generating and executing masking rules is typically a very time-consuming process that requires manual intervention by domain-specific experts.

[0095] Figure 5 shows a detailed process flow 500 for masking sensitive data during information acquisition, as implemented in this disclosure.

[0096] As shown in Figure 5, the process flow 500 for masking sensitive data during information acquisition is implemented using two or more domains simultaneously. These domains, namely the secure enterprise domain 502 (also referred to above as the secure environment 304) and the external domain 504, enable the solution presented in this disclosure to function properly. By using two or more domains in accordance with the implementation of this disclosure, sensitive enterprise data is securely stored within the secure enterprise domain 502, leading to enhanced data protection measures and ensuring compliance with strict security protocols.

[0097] The data processing system 114 performs iterative steps concerning a secure enterprise domain 502, while multiple foundational models 312 / 408a-408n perform iterative steps concerning an external domain. In some cases, multiple foundational models 312 / 408a-408n are implemented as generative artificial intelligence (Genetic AI or Gen AI) models. In such cases, multiple foundational models 312 / 408a-408n may employ prompt engineering methods for accuracy. The term “prompt engineering” refers to the process of designing / making effective prompts / queries to elicit desired responses from an artificial intelligence system. This involves structuring prompts to optimize the understanding and accuracy of responses, and is often tailored to suit a specific task or application. Prompt engineering methods may be implemented as few-shot prompting methods.

[0098] Furthermore, Figure 5 also represents the subprocesses within the process flow 500, where one or more steps in the process flow 500 are part of at least one subprocess. These subprocesses are iterated as an input generalization subprocess, a data model design and implementation subprocess, an information retrieval supervision subprocess, a query post-processing subprocess, and a natural language response generation subprocess.

[0099] In step 506, the user provides a query to the data processing system 114 via the UI / UX module 206. In some cases, the user may provide the query via a chatbot hosted by the UI / UX module 206. As a result, under the input generalization subprocess of step 508, the data processing system 114 masks sensitive information in the query and prepares multiple prompts for multiple base models. During masking, to properly mask sensitive information, numerical values ​​may be supplemented with contextual variables, text values ​​may be replaced with generalized terms, and so on.

[0100] Multiple foundation models receive multiple prompts from the data processing system 114. In step 510, under the information retrieval supervision subprocess, at least one of the multiple foundation models generates multiple instances of information retrieval queries to retrieve appropriate enterprise data (or enterprise information) from the knowledge base 204. The knowledge base 204 may be implemented as at least one of the following: a vector database, a structured query language (SQL) database, a no-SQL (not only structured query language) database, and similar. The appropriate enterprise data may further be masked to protect sensitive information present within the enterprise data 214.

[0101] Furthermore, under the data model design and implementation subprocess, the information retrieval query is then provided to the data processing system 114, which further utilizes the information retrieval query to extract enterprise data from the knowledge base 204. In step 512, the data processing system 114 unmasks the information retrieval query to extract enterprise data 214 from the knowledge base 204. In some cases, the extracted enterprise data 214 is stored locally in the knowledge base 204 for further processing.

[0102] In step 514, under the query post-processing subprocess, the data processing system 114 masks the sensitive information of the query in order to generate multiple prompts related to the query. These prompts are written in a distinctive style and format relative to the query (the same query) in order to obtain multiple data points from multiple underlying models. This is advantageous as it helps to ensure the accuracy of the response. The data processing system 114 provides multiple prompts to multiple underlying models.

[0103] Subsequently, in step 516, under the natural language response generation subprocess, multiple foundational models generate multiple responses as responses to multiple prompts. The multiple responses are generated based, in particular, on queries, information retrieval queries, and masked enterprise data 214. The multiple foundational models provide the multiple responses to the data processing system 114.

[0104] In step 518, the data processing system 114 unmasks the response and shares it with the user via the UI / UX module 206.

[0105] Naturally, the data processing system 114 is often configured to engage in multiple conversations with the user across various dialogues. Thus, the process shown in Figure 5 will constantly repeat based on the interaction with the user.

[0106] Figure 6 is a flowchart illustrating an exemplary method 600 according to an implementation of the present disclosure. In some implementations, method 600 may be performed within a data processing system 114, as described in relation to Figure 2.

[0107] In step 602, a query is received from a user. The user may be an enterprise user. In some cases, the user may be an employee of the enterprise. In other cases, the user may be a merchant, customer, vendor, and similar related to the enterprise. The query concerns a request for information. The requested information may be about the enterprise. Furthermore, the query may contain or relate to confidential information about the enterprise.

[0108] In step 604, multiple prompts are generated based on the user query. These prompts are generated to fetch the appropriate response from multiple underlying models. Furthermore, the prompts are generated by masking sensitive information within the query. Sensitive information may relate to enterprise data 214 that, if disclosed or leaked, could harm individuals, organizations, or systems, potentially leading to privacy breaches, financial losses, or reputational damage. The data processing system 114 may generate multiple prompts in response to the query, each prompt being variably configured. This process ensures comprehensive coverage of possible responses tailored to diverse contextual interpretations of the query.

[0109] In step 606, multiple responses are received from multiple foundational models. These responses are received in response to multiple prompt inputs. The multiple foundational models may use deep learning methods such as transformer architectures to generate multiple responses by processing the obtained multiple prompts. The multiple foundational models may leverage pre-trained weights and contextual embeddings to generate diverse and contextually appropriate responses by adapting their outputs based on the semantic nuances embedded in each prompt.

[0110] In step 606, a common result set is generated based on multiple responses. This involves aggregating the responses and verifying their consistency and relevance to ensure that the final set comprehensively addresses diverse contextual interpretations of the query.

[0111] In step 610, a response is generated by validating a common result set using confidential information and queries. This validation involves supplementing the response with confidential information and verifying it against enterprise data 214 and the query (the original). If inconsistencies or inaccuracies are detected, the response is regenerated to maintain consistency and accuracy.

[0112] In step 612, a user response is generated by supplementing the response with sensitive information. In some cases, the sensitive information is directly appended to the response. In other cases, the response may be modified to be consistent with the sensitive information. Advantageously, this ensures a comprehensive and contextually appropriate interaction with the user while maintaining data integrity and confidentiality.

[0113] In step 614, a user response is provided to the user in response to the query. Advantageously, method 600 enables secure, accurate, efficient, and reliable information retrieval, thereby achieving improved accuracy.

[0114] Figure 7 shows a computer system 700 that may be used to implement the data processing system 114. More specifically, computing machines such as desktops, laptops, smartphones, tablets, and wearables that may be used to handle conversational interactions in the data processing system 114 may have the structure of the computer system 700. The computer system 700 may include additional components not shown, and some of the process components described may be removed and / or modified. In another example, the computer system 700 may be deployed on an external cloud platform such as the cloud, an in-house corporate cloud computing cluster, or the organization's computing resources, and / or similar.

[0115] The computer system 700 includes one or more processors 702, such as a central processing unit, an ASIC, or another type of processing circuit; input / output devices 704, such as a display, mouse, and keyboard; a network interface 706, such as a local area network (LAN), a wireless 802.11x LAN, a 3G or 4G mobile WAN, or a WiMax WAN; and a processor-readable medium 708. Each of these components may be coupled to a bus 710 to enable operation. The computer-readable medium 708 may be any suitable medium involved in providing instructions to the processor(s) 702 for execution. For example, the computer-readable medium 708 may be a non-transient or non-volatile medium, such as a magnetic disk or solid-state non-volatile memory, or a volatile medium, such as RAM. Instructions or modules stored on the computer-readable medium 708 may include machine-readable instructions 712, which are executed by the processor(s) 702 and cause the processor(s) 702 to perform the functions of the method and data processing system 114.

[0116] The data processing system 114 may be implemented as software stored on a non-temporary processor-readable medium and executed by the processor 702. For example, the computer-readable medium 708 may store the code for the data processing system 114 and an operating system 714 such as MAC OS, MS WINDOWS, UNIX, or LINUX. The operating system 714 may be multi-user, multi-processing, multi-tasking, multi-threading, real-time, and similar. For example, during runtime, the operating system 714 is running and the code for the data processing system 114 is executed by the processor(s) 702.

[0117] The computer system 700 may include data storage 716, which may include non-volatile data storage. The data storage 716 stores any data used or generated by the data processing system 114.

[0118] The network interface 706 connects the computer system 700 to an internal system, for example, via a LAN. Furthermore, the network interface 706 may connect the computer system 700 to the internet. For example, the computer system 700 may connect to a web browser and other external applications and systems via the network interface 706.

[0119] Some of these variations, along with examples, are described and illustrated in this specification. The terms, descriptions, and drawings used in this specification are illustrative only and not intended to be limiting. Numerous variations are possible within the intent and scope of the subject matter intended to be defined by the attached claims and their equivalents.

[0120] The implementation of this disclosure provides several technical improvements and addresses the shortcomings of the conventional base model. For example, the implementation of this disclosure provides accurate retrieval of information using the base model to respond to user queries. This accuracy directly leads to improved performance of the base model, while ensuring data privacy by masking sensitive information from the base model. This accuracy leads to efficient, actionable, and correct responses from the base model, drastically reducing the possibility of hallucination.

[0121] Furthermore, the present invention enhances the accuracy of responses generated by the underlying model by intelligently masking sensitive information while preserving contextual and relational nuances within the data. This approach eliminates the need for internal domain experts to manually monitor security measures, thereby optimizing computational efficiency. By preserving the relational context of the masked data, not only are more accurate responses ensured, but the occurrence of misleading information retrieval is reduced, thereby mitigating risks associated with data inaccuracies, such as hallucination. In addition, by intelligently protecting sensitive data, the invention defends the enterprise from potential misconduct and strengthens compliance with data security and regulatory standards.

[0122] The implementations and all functional operations described herein may be realized in digital electronic circuit configurations, or in computer software, firmware, or hardware including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. An implementation may be realized as one or more computer program products (i.e., one or more modules of computer program instructions encoded on a computer-readable medium to be executed by or control the operation of a data processing device). The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a configuration of material that provides machine-readable propagating signals, or one or more combinations thereof. The term “computing system” encompasses all devices, machines, and equipment that process data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an equipment may include code that creates the execution environment for the computer program in question (e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or any one or more appropriate combinations thereof). A propagated signal is an artificially generated signal (e.g., a mechanically generated electrical, optical, or electromagnetic signal) that is produced to encode information to be transmitted to a suitable receiver device.

[0123] Computer programs (also known as programs, software, software applications, scripts, or code) may be written in any suitable form of programming language, including compiled or interpreted languages, and may be deployed in any suitable form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs do not necessarily correspond to files in a file system. A program may be part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple collaborative files (e.g., multiple files containing one or more modules, subprograms, or parts of code). Computer programs may be deployed to run on one computer, or to run on multiple computers located in one place or distributed across multiple locations and interconnected by a communication network.

[0124] The processes and logic flows described herein may be executed by one or more programmable processors that run one or more computer programs that perform functions by acting on input data and generating outputs. The processes and logic flows may also be executed by dedicated logic circuit configurations (e.g., FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits)), and the device may also be implemented as such dedicated logic circuit configurations.

[0125] Processors suitable for executing computer programs include, for example, both general-purpose and dedicated microprocessors, as well as any one or more processors in any suitable type of digital computer. Generally, a processor receives instructions and data from read-only memory, random-access memory, or both. The components of a computer may include a processor that executes instructions and one or more memory devices that store instructions and data. Generally, a computer further includes one or more mass storage devices (e.g., magnetic, magneto-optical disks, or optical disks) that store data, or is operablely coupled to them to receive data, transfer data, or both. However, a computer does not have to have such devices. Furthermore, a computer may be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio player, a Global Positioning System (GPS) receiver). Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be complemented by or incorporated into dedicated logic circuit configurations.

[0126] Each implementation may be realized on a computer having a display device (e.g., a CRT (cathode ray tube), an LCD (liquid crystal display) monitor) that displays information to the user in order to provide user interaction, as well as a keyboard and pointing device (e.g., a mouse, trackball, touchpad) that allows the user to provide input to the computer. Other types of devices may also be used to provide user interaction. For example, the feedback provided to the user may be any appropriate form of sensory feedback (e.g., visual feedback, auditory feedback, haptic feedback), and user input may be received in any appropriate form, including acoustic, voice, or haptic input.

[0127] The implementation may be in a computing system including backend components (e.g., as a data server), a computing system including middleware components (e.g., an application server), and / or a computing system including frontend components (e.g., a client computer having a graphical user interface or a web browser that a user can use to interact with the implementation), or in any suitable combination of one or more such backend components, middleware components, or frontend components. The components of the system may be interconnected by digital data communication (e.g., a communication network) in any suitable form or medium. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs") such as the Internet.

[0128] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The client-server relationship arises from computer programs running on separate computers that have a client-server relationship with each other.

[0129] This specification contains numerous details, which should not be construed as limitations on the scope of this disclosure or the claims, but rather as describing features specific to particular implementations. Certain features described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can be implemented separately in multiple implementations, or in any suitable combination of components. Furthermore, each feature may be described above as operating in a particular combination, and may even be initially claimed as such; however, in some cases, one or more features of a claimed combination can be removed from that combination, and the claimed combination may be subject to the components of the combination or variations of the components of the combination.

[0130] Similarly, while the operations are shown in a specific order within the diagrams, this should not be interpreted as requiring that the operations be executed in the specific order or sequence shown, or that all operations shown be executed, in order to achieve the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the implementation described above should not be interpreted as requiring such separation in all implementations, and naturally, the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.

[0131] Several implementations are described. However, naturally, various modifications can be made without departing from the intent and scope of this disclosure. For example, various forms of the flow shown above may be used, in which steps are rearranged, added, or removed. Thus, other implementations are within the scope of the attached claims.

Claims

1. At least one processor, The following: Receiving queries from users regarding information requests, The process involves generating a plurality of prompts based on the query from the user, wherein the plurality of prompts are generated by masking sensitive information within the query. Receiving multiple responses from multiple base models in response to inputting the aforementioned multiple prompts, To generate a common result set based on the aforementioned multiple responses, A response is generated by verifying the common result set using the confidential information and the query. To generate a user response by supplementing the response using the confidential information, and Provide the user response to the user in response to the query. For the purpose of storing instructions executed by the at least one processor, and A data processing system that includes this.

2. The data processing system according to claim 1, wherein the query is composed of natural language.

3. The instruction for providing the user response is provided to at least one processor, Verify the user based on one or more predetermined criteria relating to the user's persona, Based on the above verification, the user response is provided to the user in response to the query. A data processing system according to claim 1, which enables the following.

4. The instruction for masking sensitive information in the query is provided to at least one processor, The data processing system according to claim 1, wherein the sensitive information in the query is replaced with a variable based on one or more masking criteria, and the variable represents contextual and relational relationships of the sensitive information.

5. The data processing system according to claim 1, wherein the at least one non-transient processor-readable medium includes a data repository associated with the data processing system, the information in the data repository is represented across multiple layers and incorporates domain-specific terminology.

6. The instruction for masking the confidential information in the query is provided to at least one processor, The data processing system according to claim 5, which maps the domain terminology to the plurality of base models.

7. The instruction for generating the common result set is provided to at least one processor, The execution of each of the responses in a virtual environment by adding the confidential information to the response, wherein the response is executed by verifying the response against the data repository and the query, Select the response based on the results of the verification. A data processing system according to claim 5, which causes the following:

8. The data processing system according to claim 1, wherein the instruction for providing the user response is based on one or more masking rules, the one or more masking rules cause the at least one processor to identify and mask the confidential information based on the one or more masking rules, and the one or more masking rules include rules relating to the masking of the confidential information.

9. The instruction for providing the user response based on the one or more masking rules is provided to at least one processor. Identifying patterns in data related to the data processing system, wherein the patterns relate to fraudulent activity, and the identification of such patterns To generate one or more masking rules based on the identified pattern, Executing processor-readable instructions relating to one or more of the aforementioned masking rules A data processing system according to claim 8, which enables the following.

10. The process involves the processor receiving a query from the user regarding the request for information, A step of generating a plurality of prompts by the processor based on the query from the user, wherein the plurality of prompts are generated by masking sensitive information in the query, The steps include: receiving multiple responses from multiple base models by the processor in response to the multiple prompts; The steps include: generating a common result set based on the plurality of responses using the processor; The steps include generating a response by the processor by verifying the common result set using the confidential information and the query, The steps include generating a user response by the processor by supplementing the response using the confidential information, The steps of providing the user response to the user via the processor in response to the query, and A method that can be executed by the processor, including the following.

11. The query is composed of natural language and is executable by the processor according to claim 10.

12. The step of providing the user response is: A step of verifying the user using the processor based on one or more predetermined criteria relating to the user's persona, Based on the verification, the processor provides the user response to the user in response to the query. A method that can be executed by the processor according to claim 10, including the above.

13. The step of masking sensitive information in the aforementioned query is: A processor-executable method according to claim 10, comprising the step of having the processor replace the sensitive information in the query with variables based on one or more masking criteria, wherein the variables represent contextual and relational relevance of the sensitive information.

14. The information in the data repository related to the processor is represented across multiple layers and incorporates domain-specific terminology, and the step of masking the sensitive information in the query is, The step of mapping the domain terminology to the plurality of base models using the processor. A method that can be executed by the processor according to claim 10, including the above.

15. The step of generating the aforementioned common result set is: A step of executing each of the responses by the processor in a virtual environment by adding the confidential information to the response, wherein the response is executed by verifying the response against the data repository and the query, A step in which the processor selects the response based on the results of the verification. A method executable by the processor according to claim 14, including the above.

16. A method executable by the processor according to claim 10, wherein the step of providing the user response is based on one or more masking rules, the one or more masking rules cause the processor to identify and mask the confidential information based on the one or more masking rules, and the one or more masking rules include rules relating to the masking of the confidential information.

17. The step of providing the user response based on one or more masking rules is: A step of identifying a pattern in data related to a data processing system using the processor, wherein the pattern relates to malicious activity, and the identifying step, The steps include generating one or more masking rules using the processor based on the identified pattern, The steps include: executing a processor-readable instruction relating to one or more masking rules by the processor; A method that can be executed by the processor according to claim 16, including the method described in claim 16.

18. A non-temporary processor-readable storage medium containing machine-readable instructions, wherein the machine-readable instructions are provided to the processor. Receiving queries from users regarding information requests, The process involves generating a plurality of prompts based on the query from the user, wherein the plurality of prompts are generated by masking sensitive information within the query, Receiving multiple responses from multiple base models in response to inputting the aforementioned multiple prompts, To generate a common result set based on the aforementioned multiple responses, The response is generated by verifying the common result set using the confidential information and the query, The user response is generated by supplementing the response using the confidential information, To provide the user response to the user in response to the query, A non-temporary processor-readable storage medium that enables this operation.

19. The machine-readable instructions for providing the user response are provided to the processor, Verify the user based on one or more predetermined criteria relating to the user's persona, Based on the above verification, the user response is provided to the user in response to the query. A non-temporary processor-readable storage medium according to claim 18, which causes the following:

20. The machine-readable instructions for masking sensitive information within the query are provided to the processor, A non-temporary processor-readable storage medium according to claim 18, wherein the sensitive information in the query is replaced with variables based on one or more masking criteria, the variables representing contextual and relational relationships of the sensitive information.

Citation Information

Patent Citations

  • Methods and apparatus for internet searching

    JP2013537332A

  • Personal information protection-based query processing service provider system

    JP2021503117A

  • Anticipating queries for interactive metrics based on usage

    US20220012244A1

  • Query validation with automated query modification

    US20230017396A1

  • Document anonymizer

    WO2008126149A1