Secure large language models
The user-specific LLM generator software application addresses LLM inadequacies in sensitive domains by employing PEFT to create secure, efficient models that adhere to user permissions, preventing leaks and optimizing computational resources.
Patent Information
- Application Number
- PCT/US2024/058217
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-03
- Publication Date
- 2025-09-25
AI Technical Summary
Conventional large language models (LLMs) are inadequate for sensitive domains due to information leaks, computational and monetary expenses of retraining, and inability to reason about exotic secret domains, leading to unauthorized information exposure.
A user-specific LLM generator software application employs Parameter-Efficient Fine-Tuning (PEFT) to adapt LLMs to user permissions, allowing secure interaction with restricted data silos by generating user-specific models with compositional fine-tuning and anomaly detection.
Enables secure and efficient generation of outputs within user-defined permissions, preventing information leaks and reducing computational costs while maintaining model performance.
Smart Images

Figure US2024058217_25092025_PF_FP_ABST
Abstract
Description
SECURE LARGE LANGUAGE MODELSRELATED APPLICATION
[0001] This patent claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 63 / 616,440, titled “SECURE LARGE LANGUAGE MODELS”, filed on December 29, 2023, which is hereby incorporated by reference herein in its entirety.GOVERNMENT SUPPORT
[0002] This invention was made with government support under W911NF-22-C-0060 awarded by the U.S. Army Research Office, CCF1231216 awarded by the National Science Foundation, and FA8750-19-2- 1000 awarded by the Air Force Office of Scientific Research. The government has certain rights in this invention.FIELD
[0003] The techniques described herein relate generally to machine learning and, more particularly, to secure large language models.BACKGROUND
[0004] Machine learning (ML) generally refers to the field of deploying computer algorithms (and / or associated hardware) that iteratively improve by using data in applications (e.g., real- world applications, simulated applications) to generate outputs and provide feedback of an evaluation of the outputs to the computer algorithms. Some ML models may generate outputs in response to prompts requesting information such as in facilitating data access to restricted access and / or otherwise secure data sources in various applications or industries (e.g., healthcare, information security, finance, national government). Typically, such ML models provide the outputs irrespective of data security considerations, which may be disadvantageous to the objectives of restricting access and / or otherwise securing data sources in such various applications or industries.SUMMARY
[0005] In accordance with the disclosed subject matter, systems, apparatus, articles of manufacture, and methods are provided for secure large language models.
[0006] Some embodiments relate to an example method for producing user specific machine learning (ML) models for execution in restricted access applications. The example methodcomprising: (A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data including at least one of information in at least one datastore of a plurality of datastores or data associated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pre-trained ML model to generate an output associated with the training data; (B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users correspond to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; and for at least one user of the one or more users: identifying, by using the one or more data associations, one or more of the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters.
[0007] Some embodiments relate to an example apparatus for producing user specific machine learning (ML) models for execution in restricted access applications. The example apparatus comprising: at least one memory storing processor executable instructions; and at least one hardware processor configured to execute the processor executable instructions to perform a method comprising: (A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data comprising at least one of information in at least one datastore of a plurality of datastores or data associated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pre-trained ML model to generate an output associated with the training data; (B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users corresponds to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; and for at least one user of the one or more users: identifying, by using the one or more data associations, one or more of the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters.
[0008] Some embodiments relate to at least one example computer-readable storage medium comprising processor executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform a method comprising: (A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data comprising at least one of information in at least one datastore of a plurality of datastores or data associated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pre-trained ML model to generate an output associated with the training data; (B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users corresponds to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; and for at least one user of the one or more users: identifying, by using the one or more data associations, one or more of the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters.
[0009] Some embodiments relate to an example system for producing user specific machine learning (ML) models for execution in restricted access applications. The example system comprising: at least one hardware processor; and at least one computer-readable storage medium storing processor executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method comprising: (A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data comprising at least one of information in at least one datastore of a plurality of datastores or data associated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pre-trained ML model to generate an output associated with the training data; (B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users corresponds to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; and for at least one user of the one or more users: identifying, by using the one or more data associations, one or moreof the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters.
[0010] The foregoing summary is not intended to be limiting. Moreover, various aspects of the present disclosure may be implemented alone or in combination with other aspects.BRIEF DESCRIPTION OF FIGURES
[0011] Various aspects and embodiments will be described with reference to the following figures. In the figures, each identical or nearly identical component that is illustrated in various figures is represented by a like reference character. For purposes of clarity, not every component may be labeled in every drawing. The drawings are not necessarily drawn to scale, with emphasis instead being placed on illustrating various aspects of the techniques and devices described herein.
[0012] FIG. 1 is an example system including a user specific machine learning model generator software application for generating user specific machine learning models, in accordance with some embodiments of the technology described herein.
[0013] FIG. 2A illustrates an example workflow to generate a machine learning model for a specific user, in accordance with some embodiments of the technology described herein.
[0014] FIG. 2B illustrates generating a machine learning output for the user of FIG. 2A at the highest permission level for the user, in accordance with some embodiments of the technology described herein.
[0015] FIG. 2C illustrates generating a machine learning output for the user of FIG. 2A at a reduced permission level for the user, in accordance with some embodiments of the technology described herein.
[0016] FIG. 3 illustrates an example workflow for generating an interaction specific machine learning model for at least two users, in accordance with some embodiments of the technology described herein.
[0017] FIG. 4 is a flowchart representative of an example process that may be performed and / or example machine-readable instructions that may be executed by processor circuitry to implement the user specific machine learning model generator software application of at least FIG. 1 to generate a user specific machine learning model, in accordance with some embodiments of the technology described herein.
[0018] FIG. 5 is a flowchart representative of an example process that may be performed and / or example machine-readable instructions that may be executed by processor circuitry to implement the user specific machine learning model generator software application of at least FIG. 1 to generate an interaction specific machine learning model, in accordance with some embodiments of the technology described herein.
[0019] FIG. 6 is a flowchart representative of an example process that may be performed and / or example machine-readable instructions that may be executed by processor circuitry to implement the user specific machine learning model generator software application of at least FIG. 1 to perform anomaly detection, in accordance with some embodiments of the technology described herein.
[0020] FIG. 7 is a flowchart representative of an example process that may be performed and / or example machine-readable instructions that may be executed by processor circuitry to implement the user specific machine learning model generator software application of at least FIG. 1 to perform data leak detection, in accordance with some embodiments of the technology described herein.
[0021] FIG. 8 is an example electronic platform structured to execute the machine-readable instructions of FIGS. 4, 5, 6, and / or 7 to implement the user specific machine learning model generator software application of at least FIG. 1, according to some embodiments.DETAILED DESCRIPTION
[0022] The inventors have developed techniques for configuring and / or generating a model, such as a machine learning model, specific to one or more users (e.g., human users, electronic users). An example use case for the techniques developed by the inventors is monitoring input (e.g., plain text, natural language text) for anomalies and / or data leaks in sensitive applications. Another example use case is generating output from machine learning models in sensitive applications.
[0023] Non-limiting examples of sensitive applications include accessing and / or processing data output from a generative machine learning model in restricted access domains. Any other sensitive application is contemplated. Non-limiting examples of restricted access domains include finance, government, healthcare, law, and national defense. The generative machine learning model may be a large language model (LLM). For example, a sensitive application can be providing a prompt to an LLM, which may be used as part of a chatbot or other software solution, to generate an output in connection with information that may be restricted to authorized users.
[0024] Machine learning (ML) refers to a field of artificial intelligence (Al) that involves creating, deploying, and / or using computer software that can learn to perform a task from data and whose performance on the task can improve when additional data is provided. Such computer software may use one or more machine learning models (e.g., neural networks, large language models, hidden Markov models, and other types of models, examples of which are provided herein).
[0025] A machine learning model may include parameters. Those parameters may be assigned values. For example, a neural network model may include millions of parameters or more (sometimes termed “weights”) and those parameters may be assigned values (e.g., the values of the weights). A machine learning model may be used to process an input to generate a respective output and to do so, the input and the ML model parameter values may be used to calculate the respective output. As such, the output depends on the input and on the values of the parameters of the ML model. For example, a neural network (e.g., an LLM) can generate output text based on an input prompt and values of the neural network’s parameters.
[0026] Accordingly, prior to using an ML model to process inputs to generate respective outputs, parameters of the ML models are assigned values. The process of assigning values to parameters of an ML model based on data (e.g., numerous examples of input-output pairs) is sometimes termed “training” and the data that is used to determine the parameter values to assign is sometimes termed “training data”. Determining values of ML model parameters from training data is sometimes referred to as “learning” or “estimating” ML model parameters. There are various techniques for training ML models including supervised, semisupervised, and unsupervised techniques for training ML models.
[0027] Machine learning models include discriminative ML models and generative ML models. Discriminative models may be trained to perform a classification task - the task of separating data points into different “classes”. For example, a discriminative model may be trained to determine whether a medical image of a patient indicates the presence of a cancer. In this way, this discriminative model separates images (the “data points” in this context) into two different classes: medical images that indicate the presence of a cancer and medical images that do not indicate the presence of the cancer. To this end, a discriminative model may be trained to perform the classification task by learning boundaries between classes among which the model is being trained to discriminate. Non-limiting examples of ML models that can be trained to operate as discriminative models include clustering models,decision trees, logistic regression models, neural networks, random forests, and support vector machines.
[0028] Unlike discriminative models, which may be able only to discriminate between different types of data, generative models may be trained to generate new examples of data. For example, instead of merely being able to separate medical images into two classes (e.g., medical images indicating presence of a cancer and medical images not indicating the presence of the cancer), a generative model may be trained and used to generate new examples of medical images (not present in the training data) that indicate the presence of the cancer. As another example, a generative model may be trained and used to generate new text examples, new image examples, new sound examples, new biological sequence (e.g., protein sequence) examples, etc. To this end, a generative model may learn the underlying statistical distribution of the data using examples of that data present in the training data and, in turn, use the learned representation of that data distribution to generate new data examples. The manner in which that distribution is utilized may depend on input to the generative ML model.
[0029] In some examples, a generative ML model may be trained to generate content in response to an input. For example, a generative ML model may be configured to generate text in response to an input textual prompt. A chatbot, such as ChatGPT, is one example of such a generative ML model because it is trained to generate natural language text responsive to text input, which may be a prompt from the user. As another example, a generative ML model may be configured to generate an image based on an input textual prompt. A text-to-image model, such as DALL-E, is another example of such a generative ML model because it is trained to generate digital images in response to text input, which may be a prompt from the user.
[0030] An example generative model may be an LLM. An LLM is a type of machine learning model (e.g., a deep-learning model) trained to conduct a probability distribution over words. Some LLMs may generate new combinations of natural language text in the form of natural-sounding language while some LLMs may generate other types of output such as new audio, images, video, and / or any combination(s) thereof. Some LLMs are trained to generate an output including natural language text by predicting the next most appropriate word to fill in a blank space in a sentence, phrase, paragraph, etc. LLMs are used in a variety of NLG tasks, such as content generation, language transformation (e.g., changing a tense of natural language text in a sentence, a paragraph, etc.), question answering, and text summarization.
[0031] Language models, such as LLMs, are growing in popularity and are becoming sophisticated day to day electronic and / or computer-based assistants to users. Yet, there are many sensitive domains where they cannot be applied because they can leak information. One example type of leak that LLMs suffer from is explicit, where a malicious user attempts to induce an LLM into revealing information it should not. Another example type of leak is implicit, where LLMs output information that should otherwise be kept sensitive. Conventional LLMs are not structured to prevent such leaks, which reduces the applicability of these LLMs to sensitive domains.
[0032] The inventors have recognized that conventional approaches to using an LLM in a sensitive domain are inadequate for many use cases. One conventional approach is fine retraining (e.g., fine-tuning) a trained LLM for a particular sensitive domain. For example, an LLM may be trained using general information such as information retrieved from the Internet. In such an example, the trained LLM may be retrained for a specific application, such as diagnosing a medical condition for a patient, by using information for the specific application, such as digital biomarkers for the patient. However, the inventors have recognized that retraining and fine-tuning are expensive (e.g., computationally expensive, monetarily expensive) and both require significant amounts of data (e.g., training data). Retraining or fine-tuning an LLM results in a separate LLM per domain (or per application), which can be expensive (e.g., computationally expensive, monetarily expensive) to store. By way of example in domains that are combinatorial, assume access to data associated with patient X, Y, and Z is desired. However, generating a separate LLM for each patient quickly becomes prohibitively expensive (e.g., as the number of patients is greater than 10, 100, 1000, etc., patients) and requires substantial amounts of data.
[0033] Another conventional approach to using LLMs in sensitive domains is Retrieval- Augmented Generation (RAG) where LLMs have access to an external database. However, information external to an LLM inherently limits the LLM. For example, if the domain that is being retrieved is unknown to the LLM, the LLM will not be able to retrieve data effectively from the domain. By way of example, an entity, such as a national defense entity, would not want to reveal to an LLM that a special access program (SAP) (e.g., a specific class of classified information that imposes safeguarding and access requirements to prevent unauthorized access) named X is about Y, just so that the LLM could apply RAG effectively. Moreover, LLMs are limited because they are not used to reason about these exotic secret domains. For example, it is as if there is a new topic that the LLM cannot know, but must read a quick description of and figure out.
[0034] Other conventional approaches to using LLMs in sensitive domains such as differential privacy can only provide statistical guarantees of privacy. Further, current approaches to differentially private LLMs significantly degrade the performance of models and cannot adapt to a complex classification scheme.
[0035] Another challenge hindering the application of LLMs in restricted access applications relates to vigilance. Human persons cannot effectively supervise a machine learning system that mostly works as not enough attention can be minded ensuring such systems always provide outputs in conformance with various privacy and / or security considerations. LLMs applied to sensitive information are such systems, because they mostly give a user the answer they seek but at times may provide the user with more than needed or expected. By way of example, if a user that is authorized for SAPs A and B queries an LLM, the LLM may reason about SAPs A, B, and C to provide an output to the user. For instance, the LLM may answer a prompt about A and, as part of the answer, may provide a related fact about C unintentionally. In another example, the LLM may use A and B to conjecture something that is true about C, effectively leaking information. Thus, even the most vigilant of users cannot use such systems.
[0036] To overcome the technical challenges of conventional approaches to deploying and / or executing LLMs in restricted access applications, the inventors have developed techniques for fine-tuning and / or retraining an LLM for every type and combination of secret / sensitive information available for a particular domain and such fine-tuned / retrained LLM may deeply understand the domain and its new internal logic. The techniques can adapt an LLM to generate outputs that every user using the LLM is authorized and / or permitted to know. For example, each data silo or information data source can be associated with a distinct finetuning, and users can only compose the collection of fine-tunings they are authorized to access, allowing a provably secure LLM configured using the authorized collection of finetunings to perform compositional tasks across data.
[0037] In some embodiments, a user may be a human user (e.g., a person). For example, a user may be a person using an electronic device to interact with an LLM, such as by generating a prompt to the LLM for the LLM to generate output or being presented with output from the LLM.
[0038] In some embodiments, a user may be a computer user (e.g., a machine, software). Non-limiting examples of a computer user include a robot, a machine such as a programmable processor executing computer readable instructions, and software such as an application programming interface (API) and / or a chatbot.
[0039] The inventors have developed techniques that deploy / execute an LLM that can interact with and / or process both large and small amounts of data. The techniques can be used to manage information silos that are related to and / or associated with one another. The techniques can be used to produce an LLM that adapts to the access privileges the user is communicating with, so that it produces responses appropriate for the interaction (e.g., the conversation) with the user and / or between the user and other user(s). The LLM can answer questions, prompts, etc., using information among different datasets that a user is permitted to access and not using information among other datasets that the user is not permitted to access.
[0040] The inventors have developed techniques for dynamically creating an LLM that imports the general-purpose knowledge that LLMs have from external sources (e.g., the Internet), combined with the knowledge from the information silos that the user wants to use, and information silos they want to avoid but know about. For example, such dynamically created LLMs can prevent users from accessing information they cannot (e.g., they are not permitted to access), provide them with maximal assistance in their domain, and assist them to communicate with users having lower privileges without leaks (e.g., data leaks).
[0041] The inventors have developed techniques for identifying anomalies in data streams. For example, using LLMs in environments with sensitive information is fraught with technical problems, such as models being convinced to reveal information they should not, to answer questions they should not, to reveal their training data and prompts, to run API calls they should not, etc. The inventors have recognized that it is a significant technical challenge to discern when a model is sharing information it should not.
[0042] The techniques developed by the inventors can overcome these technical problems of using LLMs in sensitive applications by processing input data into quantifiable metrics. For example, a pre-trained LLM can be executed to determine a metric of input data (e.g., an e- mail, a collaboration software message, a text message, audio transcription of a video call). In such an example, another LLM can be executed to compute that same metric but conditioned on secure / sensitive data. Furthering the example, the metrics can be compared to determine whether the data stream includes data leaked from a secure / sensitive data source. The techniques can be used to identify a source of the leaked data, such as one or more secure / sensitive datastores (e.g., databases).
[0043] Examples disclosed herein describe and illustrate such techniques but the disclosure is not so limited.
[0044] The techniques developed by the inventors include an LLM generator software application (e.g., a user specific LLM generator software application) that can configure,instantiate, generate, and / or execute secure LLMs, such as LLMs for execution in restricted access applications. In some disclosed examples, the LLM generator software application implements a compositional fine-tuning technique that enables effective RAG and secure application API calls. For example, the LLM generator software application can obtain and / or instantiate a pre-trained LLM, such as an LLM trained on public domain sources (e.g., the Internet). Furthering the example, the LLM generator software application can fine-tune the pre-trained LLM one information silo at a time to determine, generate, and / or output a set of parameters (e.g., ML parameters, LLM parameters) for a portion of the pre-trained LLM and corresponding to the information silo. For example, let MG be a general knowledge model (e.g., a pre-trained LLM) that was autoregressively trained on a very large dataset. Given N data silos { Si, S2, ... , SN } of proprietary information, N fine-tuned LLMs { Mi, M2, ... , MN } are created. In such an example, the LLM generator software application can restrict what parts of an LLM can be updated such that only a relatively small number of parameters may be updated. Beneficially, by updating only a relatively small number of parameters, the techniques disclosed herein improve computational efficiency and reduce computational expenditures in subsequent processing because fewer parameters are needed for configuring and / or executing the LLM.
[0045] Restricting the parts of an ML model, such as an LLM, that can be updated is referred to as Parameter-Efficient Fine-Tuning (PEFT), but any other type of fine-tuning technique is contemplated. The LLM generator software application can identify and / or extract the set of parameters for storage and subsequent processing. Beneficially, the set of parameters can be used to teach the LLM about the new domain (e.g., the information silo corresponding to the set of parameters) and how to perform RAG and / or API calls in that domain.
[0046] Non-limiting examples of a set of parameters include one or more modules (e.g., one or more ML modules, one or more LLM modules), one or more layers (e.g., ML layers, LLM layers), one or more weight values, and one or more hyperparameters.
[0047] Non-limiting examples of a module (e.g., an ML module, an LLM module) include one or more portions of a model, which can include one or more layers of the model. For example, the LLM generator software application can actuate over one or more layers of a pre-trained model substantially simultaneously such that the LLM generator software application is actuating over one or more modules of the pre-trained model.
[0048] As used herein “real time”, “substantially real time”, “substantially real-time”, and “substantially simultaneously” refer to occurrence in a near instantaneous manner recognizing there may be real-world delays for computing time, transmission, etc. Thus,unless otherwise specified, “real time”, “substantially real time”, “substantially real-time”, and “substantially simultaneously” refer to being within a 5-second time frame, a 1 -second time frame, a 0.5-second time frame, a 250-millisecond time frame, a 100-millisecond time frame, a 10-millisecond time frame, etc., of real time. For example, an event described herein occurring in “real time”, “substantially real time”, “substantially real-time” and “substantially simultaneously” is occurring within 5 seconds, within 1 second, within 0.5 seconds, within 250 milliseconds, within 100 milliseconds, within 10 milliseconds, etc.
[0049] Information silos can be implemented by datastores or other types of data repositories, which can include documents, information, and / or other data associated with a particular topic. By way of example in a healthcare application, a first information silo can be implemented by one or more first datastores that store data related to patients, a second information silo can be implemented by one or more second datastores that store data related to medical images (e.g., images from a magnetic resonant imaging machine, an ultrasound machine), a third information silo can be implemented by one or more third datastores that store data related to test results (e.g., results from laboratory blood tests), and so on. In this example, the LLM generator software application can generate a first set of parameters corresponding to the first information silo, a second set of parameters corresponding to the second information silo, a third set of parameters corresponding to the third information silo, and so on. For example, the LLM generator software application can generate a first adapter including the first set of parameters and which can be added to a pre-trained LLM to extend the capabilities of the pre-trained LLM to tasks associated with the first information silo.
[0050] In some disclosed examples, given a new user who has a given set of permissions (e.g., data access permissions, security permissions), the LLM generator software application can assemble, compile, and / or generate an LLM that is specific for the user in accordance with the given set of permissions. Such an LLM may be referred to as a user specific LLM. By way of example, the user, based on the set of permissions, may have access to an information silo about a particular listening platform, a unit, a mission, etc., related to a national defense entity. The LLM generator software application can map the set of permissions (and / or other inputs) to set(s) of parameters. The LLM generator software application can add, append, and / or integrate the set(s) of parameters into the LLM.
[0051] In some disclosed examples, at least some of the set of parameters can be used in a positive way and / or at least some of the set of parameters can be used in a negative way. By way of example, when the user is being assisted with an LLM in writing an email, the positive portion of the set of parameters can be used to generate information that can beshared with the email recipient. Furthering the example, the negative portion of the set of parameters can be used to censure, restrict, and / or identify information that cannot be shared with the email recipient. Beneficially, this user specific and / or otherwise custom LLM, composed of positive and negative portion(s), can assist the user generate outputs (e.g., draft an email) with improved efficiency, security, and / or privacy. Additional examples and benefits in connection with the techniques developed by the inventors are described below in connection with the figures.
[0052] In some disclosed examples, the LLM generator software application can detect anomalies in data streams. For example, the LLM generator software application can process input data (e.g., audio, text, images, video) into quantifiable metrics. In such an example, a pre-trained LLM can be executed to determine a metric of input data (e.g., an e-mail, a collaboration software message, a text message, audio transcription of a video call).Furthering the example, another LLM can be executed to compute that same metric but conditioned on secure / sensitive data. The metrics can be compared to determine whether the data stream includes data leaked from a secure / sensitive data source. In some embodiments, the LLM generator software application can be executed to identify a source of the leaked data, such as one or more secure / sensitive datastores (e.g., databases).
[0053] The techniques described herein may be implemented in any of numerous ways, as the techniques are not limited to any particular manner of implementation. Examples of details of implementation are provided herein solely for illustrative purposes. Furthermore, the techniques disclosed herein may be used individually or in any suitable combination, as aspects of the technology described herein are not limited to the use of any particular technique or combination of techniques.
[0054] Turning to the figures, the illustrated example of FIG. 1 is an example system 100 including a machine learning model generator software application 102 for generating models 104. The system 100 of this example is a machine learning model generation system. For example, the system 100 can be used to generate machine learning (ML) models, such as large language models (LLMs).
[0055] The machine learning model generator software application 102 of this example is a user specific large language model generator software application configured to generate an ML model, such as an LLM, to correspond to a specific user (e.g., a human user, a machine user). Additionally and / or alternatively, the machine learning model generator software application 102 may be any other type of custom machine learning model generator software application, such as an application specific LLM generator software application or aninteraction specific LLM generator software application. For example, the machine learning model generator software application 102 can be configured to generate a machine learning model, such as an LLM, to correspond to a specific application, industry, use case, and / or interaction (e.g., a conversation, a question-and-answer interaction, a prompt-and-answer interaction).
[0056] The machine learning model generator software application 102 of this example is implemented by a combination of hardware (e.g., one or more servers such as computer servers), software, and / or firmware. For example, the machine learning model generator software application 102 can be implemented by one or more programmable processors executing and / or instantiating computer-readable instructions. Non-limiting examples of programmable processors include central processing units (CPUs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), and graphics processing units (GPUs).
[0057] Alternatively, the machine learning model generator software application 102 may be implemented by hardware alone. For example, the machine learning model generator software application 102 may be implemented by one or more neural network (NN) processors, one or more application specific integrated circuits (ASICs), and / or one or more field programmable gate arrays (FPGAs).
[0058] In some embodiments, the models 104 can be ML models. Non-limiting examples of ML models include generative adversarial networks (GANs), LLMs, and neural networks (NNs). Any other type of ML model is contemplated.
[0059] Alternatively, one(s) of the models 104 may be any type of non-ML model. Nonlimiting examples of non-ML models include computer-implemented decision trees, Markov Chains, template-based generation models, and context-free grammars (CFGs). Any other type of model is contemplated. By way of example, outputs of the computer-implemented decision trees, the Markov Chains, the template-based generation models, and / or the CFGs may be natural language outputs. Markov Chains can be used to generate text by transitioning from one state to another based on a set of probability distributions. In this context, states often correspond to words or sequences of words, and the probability distributions are derived from the frequency of words or sequences in a training corpus. In template-based generation, text can be generated using predefined templates where specific slots are filled in based on certain rules or heuristics. This technique can be used in natural language generation (NLG) systems for tasks like report generation or automated messaging. CFGs are sets of recursive rewriting rules or productions. By applying these rules in a random or guided manner, it is possible to generate a wide variety of sentences.
[0060] The models 104 of FIG. 1 may be neural networks. The neural networks may be deep neural networks. A non-limiting example of a deep neural network is an LLM. For example, the models 104 can be LLMs as shown in FIG. 1. More specifically, the models 104 can be user specific LLMs and are identified in FIG. 1 as at least USER 1 LLM and USER N LLM.
[0061] The machine learning model generator software application 102 can obtain and / or receive a pre-trained model 106. For example, the machine learning model generator software application 102 can retrieve the pre-trained model 106 from one or more computer readable mediums, such as at least one memory and / or mass storage device(s).
[0062] Additionally and / or alternatively, the machine learning model generator software application 102 may retrieve the pre-trained model 106 from a network (not shown but the system 100 may include one or more networks). In some embodiments, the network may be implemented by any wired and / or wireless network(s) such as one or more cellular networks (e.g., 4G LTE cellular networks, 5G cellular networks, future generation 6G cellular networks, etc.), one or more data buses, one or more local area networks (LANs), one or more optical fiber networks, one or more private networks, one or more public networks, one or more satellite networks, one or more wireless local area networks (WLANs), etc., and / or any combination(s) thereof. For example, the network may be the Internet, but any other type of private and / or public network is contemplated.
[0063] In some embodiments, the pre-trained model 106 is a pre-trained ML model. The pretrained ML model may be a pre-trained LLM. As shown, the pre-trained model 106 is a pretrained LLM with a set of parameters 108. Alternatively, the pre-trained model 106 may be any other type of ML model.
[0064] In some embodiments, the set of parameters 108 is a set of model parameters. The model parameters may be a set of ML parameters. Non-limiting examples of ML parameters include one or more modules (e.g., ML modules), one or more layers (e.g., ML layers), and one or more weight values (e.g., ML weight values). Any other type of ML parameter is contemplated such as one or more hyperparameters (e.g., ML hyperparameters).
[0065] The set of ML parameters and / or, more generally, the set of model parameters, may be a set of LLM parameters. As shown, the set of parameters 108 is a set of LLM parameters.
[0066] Non-limiting examples of LLM parameters include one or more modules (e.g., LLM modules), one or more layers (e.g., LLM layers), and one or more weight values (e.g., LLM weight values). Any other type of LLM parameter is contemplated such as one or more hyperparameters (e.g., LLM hyperparameters).
[0067] As shown, the system 100 of FIG. 1 includes a plurality of users 109 with respective user data 110 (identified by USER 1 DATA, USER N DATA), a plurality of datastores 112 (identified by DATASTORE 1, DATASTORE N), and a plurality of application programming interfaces (APIs) 114 (identified by API 1, API N). For example, the machine learning model generator software application 102 may generate and / or produce user specific models for the plurality of users 109 and in accordance with their respective user data 110.
[0068] The user specific models of this example are the models 104, which can be deployed for execution in restricted access applications. Each of the user specific models can correspond to a respective one of a plurality of users 109 and generated in accordance with their respective user data 110. For example, USER 1 LLM can correspond to USER 1 and USER N LLM can correspond to USER N.
[0069] The user data 110 of the shown example can correspond to, be associated with, and / or represent users 110 that are authorized and / or verified to access sensitive information. Nonlimiting examples of such users 110 include human users and electronic users (e.g., machine users).
[0070] Non-limiting examples of human users include medical doctors authorized to access patient records, government employees authorized to access special access programs (SAPs) and / or other confidential information, and corporate personnel authorized to access nonpublic corporate financial information.
[0071] Non-limiting examples of electronic users include at least one model (e.g., at least one computer-based model, ML model, LLM model), at least one programmable processor executing computer readable instructions, at least one machine, at least one robot, at least one server, a chatbot, a software agent, an API, and at least one autonomous vehicle. Nonlimiting examples of autonomous vehicles include a land vehicle (e.g., an automobile), a marine vehicle (e.g., a boat, a ship), and an aerial vehicle (e.g., an aircraft, a drone, a rotorcraft).
[0072] Non-limiting examples of the user data 110 include a name of a user (e.g., a given and / or family name), credentials (e.g., access and / or login credentials such as an account user name and / or password), a geographical location of the user, a type and / or classification of a network used by the user, a multi-factor authentication (MFA) identifier (e.g., a token), data classification level(s) to which the user has access, an identifier of an electronic device associated with the user, and user preferences (e.g., a user’s inclination to request information pertaining to only a first or second datastore).
[0073] Non-limiting examples of an identifier of an electronic device include a type of the electronic device (e.g., a desktop computer, a cellular phone), an Internet Protocol (IP) address and / or IP port number assigned to the electronic device, and a media access control (MAC) address assigned to the electronic device.
[0074] Non-limiting examples of types of electronic devices include desktop computers, workstations, laptop computers, servers (e.g., computer servers), tablet computers, cellular phones (e.g., smartphones, mobile phones), and wearable devices (e.g., headphones, headsets (e.g., augmented reality and / or virtual reality (AR / VR) headsets), smartwatches, smart glasses, etc.).
[0075] The datastores 112 of the shown example can be implemented by any technology for storing data. For example, the datastores 112 can be implemented by a volatile memory (e.g., a Synchronous Dynamic Random Access Memory (SDRAM), a Dynamic Random Access Memory (DRAM), a RAMBUS Dynamic Random Access Memory (RDRAM), etc.) and / or a non-volatile memory (e.g., flash memory). The datastores 112 may additionally and / or alternatively be implemented by one or more double data rate (DDR) memories, such as DDR, DDR2, DDR3, DDR4, DDR5, mobile DDR (mDDR), etc. The datastores 112 may additionally and / or alternatively be implemented by one or more mass storage devices such as hard disk drive(s) (HDD(s)), compact disk (CD) drive(s), digital versatile disk (DVD) drive(s), solid-state disk (SSD) drive(s), etc. While in the illustrated example each of the datastores 112 is illustrated as a single datastore, the datastores 112 may be implemented by any number and / or type(s) of datastore.
[0076] The datastores 112 of the shown example can store data associated with varying degrees of permission level access. For example, the datastores 112 are shown as storing data including information (identified by INFO 1, INFO N) and documents. The information may include audio, images, text (e.g., plain text, natural language text), and / or video. The documents may be documents including images and / or text.
[0077] Furthering the example, a first one of the users 109 (e.g., USER 1) may be authorized to access some portions of a first one of the datastores 112 (e.g., DATASTORE 1) but not other portions of the first one of the datastores 112. For example, the first one of the users 109 may be authorized to access a first set of DOCUMENTS 1 but not a second set of DOCUMENTS 1. By way of another example, USER 1 may be authorized to access the first one of the datastores 112 (e.g., DATASTORE 1) but not a second one of the datastores 112 (e g., DATASTORE N).
[0078] In some embodiments, one(s) of the datastores 112 store information that at least one of (i) includes at least Sensitive Compartmented Information (SCI), (ii) is associated with one or more SAPs, or (iii) includes at least any other kind of sensitive or restricted access information.
[0079] Non-limiting examples of sensitive or restricted access information include information in applications or industries operating with data to which access may be restricted where such applications or industries include finance, healthcare, national government, etc. In some embodiments, sensitive or restricted access information can include information that needs to be segmented along one or more distinguishing parameters. For example, the sensitive or restricted access information can include first data related to child friendly content and second data related to mature adult content where users may have access to the first data but not the second data.
[0080] The information and / or documents of the datastores 112 may include data such as audio data, image data, text data, video data, and / or any combination(s) thereof. Additionally and / or alternatively, the datastores 112 may store any other type of data. Furthermore, the data stored in the datastores 112 in any data format. Non-limiting examples of data formats include a flat file, binary data, comma delimited data, tab delimited data, and structured query language (SQL) structures.
[0081] In some embodiments, the datastores 112 may include and / or be implemented by a database system, such as one or more databases, and / or include the database system. The term “database” as used herein means an organized body of related data, regardless of the manner in which the data or the organized body thereof is represented. For example, the organized body of related data may be in the form of one or more of a table, a log, a map, a grid, a packet, a datagram, a frame, a file, an e-mail, a message, a document, a report, a list or in any other form.
[0082] In the illustrated example, the first one of the datastores 112 (e.g., DATASTORE 1) includes and / or is implemented by at least one database identified by DATABASE 1. For example, the first one of the datastores 112 can store data in the form of one or more of a table, a log, a map, a grid, a packet, a datagram, a frame, a file, an e-mail, a message, a document, a report, a list or in any other form. Also shown is the second one of the datastores 112 (e.g., DATASTORE N) includes and / or is implemented by at least one database identified by DATABASE N.
[0083] The APIs 114 of the shown example can represent sets of definitions and / or protocols to build and integrate application software. For example, a first one of the APIs 114 (e.g.,API 1) can represent code (e.g., commands, functions, instructions, etc.) that, when invoked, can be used to access an external data domain (e.g., the Internet), an internal data domain (e.g., an intranet for an entity, such as a healthcare organization, a government office, etc.), and / or one(s) of the datastores 112. For example, the models 104 can be configured to invoke one(s) of the APIs 114 to effectuate Retrieval -Augmented Generation (RAG).
[0084] In example operation, the machine learning model generator software application 102 and / or, more generally, the system 100, can produce user specific LLMs, such as the models 104, for execution in applications, such as restricted access applications. The models 104 can be constructed using the pre-trained model 106 and one or more fine-tunings, and the one or more fine-tunings can be selected in accordance with a user’s user data.
[0085] In example operation, the machine learning model generator software application 102 can obtain the pre-trained model 106. For example, the machine learning model generator software application 102 can obtain the pre-trained model 106 from one or more networks and / or one or more computer readable mediums.
[0086] In example operation, the machine learning model generator software application 102 can instantiate an instance of at least one of pre-trained model, such as an instance of at least one of the pre-trained model 106. For example, the machine learning model generator software application 102 can instantiate the pre-trained model 106 by loading the set of parameters 108 into one or more computer readable mediums. In some such embodiments, the machine learning model generator software application 102 can store the set of parameters 108 into an LLM parameter datastore 118.
[0087] In example operation, the machine learning model generator software application 102 can (A) train the at least one of the pre-trained model 106, using training data to generate a first set of ML parameters 116. In some embodiments, the training data includes at least one of information in at least one datastore of the plurality of datastores 112 or data associated with at least one API of the plurality of APIs 114.
[0088] By way of example, the machine learning model generator software application 102 can train, such as by retraining and / or fine-tuning one or more portions (e.g., one or more modules, one or more layers) of the at least one pre-trained model 106 using information in at least datastore 1 and / or data associated with at least API 1 to generate the first set of ML parameters 116. Furthering the example, the machine learning model generator software application 102 can fine-tune the pre-trained model 106 using any combination of one or more fine-tuning techniques, such as Parameter-Efficient Fine-Tuning (PEFT) based finetuning techniques. For example, the machine learning model generator software application102 can fine-tune the pre-trained model 106 using any one or more of Low-Rank Adaptation (LoRA), prefix tuning, prompt tuning, adapter-based tuning, etc., to extract out the tuned parameters for subsequent storage in at least one datastore. In such an example, the tuned parameters can represent fine-tunings of the pre-trained model 106, and the tuned parameters can be stored in the LLM parameter datastore 118 as one or more sets of ML parameters (e.g., the first set of ML parameters 116).
[0089] In some embodiments, training the at least one pre-trained model includes using multiple pre-trained models (e.g., multiple LLMs). In some embodiments, the at least one pre-trained model includes multiple instances of the same model. For example, the at least one pre-trained model can include at least a first and second instance of the pre-trained model 106.
[0090] In some embodiments, at least one of the at least one pre-trained model is different from other model(s) of the at least one pre-trained model. For example, the at least one pretrained model can include the pre-trained model 106 and another pre-trained model different from the pre-trained model 106 shown in FIG. 1.
[0091] The machine learning model generator software application 102 can generate a set of model parameters (e.g., ML parameters, LLM parameters). For example, the machine learning model generator software application 102 can attach the existing (e.g., already generated) parameters of a first LLM to a second LLM. In another example, the machine learning model generator software application 102 can derive some portion of the first set of ML parameters 116 automatically and then attach them to an LLM without using the LLM to generate the first set of ML parameters 116. In yet another example, the machine learning model generator software application 102 can use a smaller LLM to learn the first set of ML parameters 116 and then attach them to a bigger variant of that LLM.
[0092] In some embodiments, the machine learning model generator software application 102 can identify and / or use one or more fine-tuning techniques based on information to be used for fine-tuning. By way of example for a government entity that uses a classification scheme, different parts of the scheme can have different impacts. A SCI, SAP, foreign government information (FGI), and dissemination control may each require a different PEFT- based technique. Further, depending on the nature of the information in the program (e.g., database, API, documents), the fine-tuning technique may also change. This allows the identification of different fine-tuning techniques for specific parts of the classification scheme, or any other type of restricted access schema.
[0093] In some embodiments, the first set of ML parameters 116 can include one or more modules (e.g., one or more additional modules), which include one or more weights (e.g., one or more weight values), to be added to, appended to, combined with, and / or integrated with the pre-trained model 106. In some embodiments, the first set of ML parameters 116 can include one or more layers (e.g., one or more additional layers), which include one or more weights (e.g., one or more weight values), to be added to, appended to, combined with, and / or integrated with the pre-trained model 106.
[0094] The machine learning model generator software application 102 can store the first set of ML parameters 116 in the LLM parameter datastore 118, which also includes the set of parameters 108 of the pre-trained model 106. Additionally and / or alternatively, the set of parameters 108 may be stored elsewhere (e.g., in a different datastore, in cloud storage).
[0095] At least one portion of the pre-trained model 106 can be configured with the first set of ML parameters 116 to generate an output associated with the information in datastore 1. For example, the machine learning model generator software application 102 can add and / or link the first set of ML parameters 116 to the one or more last layers, last modules, etc., of the pre-trained model 106. By way of example, an LLM that includes the first set of ML parameters 116 can be configured to perform RAG or API calls to datastore 1 while another LLM that includes a different set of parameters may be configured to perform RAG or API calls to datastore 1 and / or datastore N.
[0096] In some embodiments, the machine learning model generator software application 102 can generate the models 104 by configuring a portion of the pre-trained model 106 such that all the plurality of sets of parameters 120 are attached to the pre-trained model 106. In some such embodiments, the machine learning model generator software application 102 can zero out (or attenuate to the extent of being irrelevant) one or more parameters of the plurality of sets of parameters 120. For example, the machine learning model generator software application 102 can control a degree to which param eter(s) of the pre-trained model 106 and / or parameter(s) added to the pre-trained model 106 is / are attenuated.
[0097] In example operation, the machine learning model generator software application 102 can (B) generate one or more data associations of one or more users 109 and the first set of ML parameters 116 after determining that user data 110 of the one or more users 109 correspond to one or more portions of the training data (e.g., the information in the at least one datastore and / or the data associated with the at least one API). For example, the machine learning model generator software application 102 can generate one or more data associations of user 1 and user 2 and the first set of ML parameters 116 after determining that user data110 of user 1 and user 2, such as portion(s) of their respective user data (e.g., data classification level(s)), correspond to one or more aspects and / or characteristics of data in datastore 1 and / or of data associated with API 1. In some such embodiments, the machine learning model generator software application 102 can link and / or otherwise associate user 1 and user 2 with the first set of parameters 116 after determining that a data classification level of those users match and / or exceed a data classification level of one or more documents in datastore 1 and / or data associated with API 1.
[0098] In example operation, the machine learning model generator software application 102 can perform (A)-(B) as described above among different sets of training data to generate a plurality of sets of parameters 120, where each of the plurality of sets of parameters 120 correspond to a different one of the different sets of training data. For example, let MG be a general knowledge model that was autoregressively trained on a very large dataset. Given N data silos { Si, S2, ... , SN } of proprietary information, N fine-tunings { Mi, M2, ... , MN } are created. In this example, MG can correspond to the pre-trained model 106, the N data silos can correspond to the datastores 112, and the N fine-tunings can correspond to the sets of parameters 120. In such an example, Si can correspond to datastore 1 and Mi can correspond to the set of parameters 108 for the pre-trained model 106. In yet another example, S2 can correspond to datastore 2 and A / ? can correspond to the first set of parameters 116
[0099] By way of example, the machine learning model generator software application 102 can retrieve and / or instantiate additional instance(s) of the at least one pre-trained model 106 (e.g., re-instantiating the instance(s) of the at least one pre-trained model 106). The machine learning model generator software application 102 can train the at least one pre-trained model 106 using information in a different combination of the datastores 112, such as datastore 2 and / or datastore 3, and / or data associated with a different combination of the APIs 114, such as API 2 and / or API 3. The machine learning model generator software application 102 can train the at least one pre-trained model 106 to generate a second set of parameters 122 corresponding to the different combination of the datastores 112 and / or the APIs 114. The machine learning model generator software application 102 can generate one or more data associations of one or more users 109 and the second set of ML parameters 122 after determining that user data 110 of the one or more users 109 correspond to the information in the different combination of the datastores 112 and / or the different combination of the APIs 114.
[0100] In example operation, the machine learning model generator software application 102 can generate the models 104 using one(s) of the plurality of sets ofparameters 120. For example, for at least one user of one or more users 109, the machine learning model generator software application 102 can identify, by using the one or more data associations, one or more of the plurality of sets of parameters 120 that correspond to the at least one user. In some such embodiments, the machine learning model generator software application 102 can identify which one(s) of the plurality of sets of parameters 120 is / are applicable to a user, such as user 1, user 2, etc., based on the user data 110 for the user.
[0101] In example operation, for the at least one user of the one or more users 109, the machine learning model generator software application 102 can generate a trained ML model, such as one of the models 104, for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters. For example, the machine learning model generator software application 102 can assemble, compile, package, and / or otherwise generate an LLM (e.g., USER 1 LLM) for a user (e.g., USER 1) by configuring the one or more portions of the at least one pre-trained model 106 with one or more of the plurality of sets of parameters 120 that correspond to the user. In some such embodiments, the machine learning model generator software application 102 can attach at least some of the set of parameters 116 to the set of parameters 108.
[0102] As described in connection with FIG. 1 and / or other embodiment s) described herein, the machine learning model generator software application 102 and / or, more generally, the system 100, can generate an ML model, such as an LLM, that is customized and / or tailored to user data for a specific user. In this way, the machine learning model generator software application 102 can enable an LLM to be fine-tuned to know about multiple silos of information (e.g., multiple datastores of information that may be isolated and / or otherwise separated from one(s) of each other) and to avoid reasoning and / or speaking about other silos of information in a mathematically provably correct way that is guaranteed to never leak. For example, such fine-tuned LLMs may never leak because they may not be generally trained on multiple silos of information but instead fine-tuned to perform RAG or use API(s) on silo(s) of which a user is permitted to access. Thus, in some embodiments, such fine-tuned LLMs may not output information from an information silo that a user is not permitted to access because the fine-tuned LLMs may not be generally trained on such information.
[0103] FIG. 2A illustrates an example workflow 200 to generate a machine learning model for a specific user. In the example workflow 200, a user 202 (identified by USER 1) provides permission access data 204 to the machine learning model generator software application 102 of FIG. 1.
[0104] In some embodiments, the user 202 can be one of the users of FIG. 1 represented by the user data 110. For example, the user 202 can be a human user and / or an electronic user (e.g., an API, an agent, a chatbot).
[0105] In some embodiments, the permission access data 204 can be implemented by and / or correspond to the user data 110 of FIG. 1. For example, the permission access data 204 can be login and / or password credentials.
[0106] In the workflow 200 of the shown example, the machine learning model generator software application 102 determines which one(s) of a plurality of sets of parameters 120 stored in the LLM parameter datastore 118 of FIG. 1 is / are applicable to the user 202.
[0107] In the workflow 200, the machine learning model generator software application 102 can determine, based on the permission access data 204, that the user 202 is permitted and / or authorized to use an LLM with the set of parameters 108 and the first set of ML parameters 116. For example, the user 202 can be linked to the set of parameters 108 because the set of parameters 108 are the parameters for the pre-trained model 106 of FIG. 1. In some such embodiments, each user may be linked to the base parameters of the pre-trained model 106 such that any generated LLM for any user may include the set of parameters 108.
[0108] Further, the machine learning model generator software application 102 can determine that the user 202 can be linked to the first set of ML parameters 116 but not the second set of parameters 122. For example, the machine learning model generator software application 102 can determine that the first set of parameters 116 is associated with one or more first datastores that store information with the same or lower degree of permission access than that of the user 202. In such an example, the machine learning model generator software application 102 can determine that the second set of parameters 122 is associated with one or more second datastores that store information with a higher degree of permission access than that of the user 202. Accordingly, the machine learning model generator software application 102 can determine not to link the second set of parameters 122 to the user 202 because the second set of parameters 122 is associated with information that the user 202 is not permitted to access. Alternatively, the machine learning model generator software application 102 may determine to link the second set of parameters 122 to the user 202 but zero out and / or attenuate respective values of the second set of parameters 122 such that they do not contribute to output generated by a corresponding LLM.
[0109] In the workflow 200 of FIG. 2A, after determining which one(s) of the parameters 120 is / are applicable to the user 202, the machine learning model generatorsoftware application 102 can generate an LLM 206 (identified by USER 1 LLM) that corresponds to the user 202. For example, the machine learning model generator software application 102 can generate a user specific LLM that, when executed and / or instantiated, can generate outputs using only information the user 202 is permitted to access. The machine learning model generator software application 102 can generate the LLM 206 by configuring a pre-trained LLM, such as the pre-trained model 106, using the set of parameters 108 for the pre-trained LLM and the first set of ML parameters 116.
[0110] In some embodiments, the machine learning model generator software application 102 can generate an LLM for a user using a set of parameters that can represent a positive influence (or effect or impact) on an output and / or a negative influence (or effect or impact) on the output. For example, positive and negative influence parameters can be used asymmetrically in the LLM. Positive influence parameters can be parameters (e.g., module(s), layer(s), weight(s)) added to an LLM such as by adding the positive influence parameters to the LLM directly. However, negating a learned set of parameters does not remove the information they contain because indeed, the LLM can leak that very same piece of information. Instead, negative information can be used in the head of the LLM adversarially. Essentially, the machine learning model generator software application 102 can configure an LLM to use the positive influence parameters to know what to say, and then watch its own output and censor it with the negative influence parameters. The negative influence parameters because of this do not contribute to the reasoning of the LLM, which effectuates data privacy and / or security. For instance, it may be beneficial to not draw complex inferences based on information that the LLM should not be disclosing. Negative silos can also be aided by a Chain-of-Thought (CoT) technique where the LLM is equipped with a decoder that will not discuss negative information such as by re-examining and rereasoning about its output before providing it to the user.
[0111] In some embodiments, the machine learning model generator software application 102 can use the government in a box approach to train an LLM to combine finetunings, both positive and negative. The government in a box approach may refer to a model, such as an LLM, that is trained for a specific government purpose. Thus, the trained model may be aware of sensitive government information that a malicious entity may desire to access. The malicious entity may attempt to trick the trained model into outputting information that it should not be disclosing by asking the trained model specific questions about classification marking, SCI, and SAPs. The machine learning model generator software application 102 can thereby ask an LLM to produce questions about classification marking,SCI, or SAPs. Specifically, the machine learning model generator software application 102 must not ask generic questions, but adversarial questions that can only be answered with that specific access, and that a generic LLM cannot answer. The machine learning model generator software application 102 can use those question and answer pairs to train the user specific LLM that combines fine-tunings, both positive and negative.
[0112] In some embodiments, information silos, which may be implemented by one or more datastores as described herein, may also be arranged hierarchically. For example, SCIs are generally part of a SAP. This can be reflected in the structure of combinations of tunings. Related SCIs can be combined together and then be combined with fine-tunings for related SAPs in a hierarchical way (or an automatically learned manner). The machine learning model generator software application 102 can use this natural structure for silo combinations to improve the reasoning of the trained LLM.
[0113] FIG. 2B illustrates generating a machine learning output for the user 202 of FIG. 2A at the highest permission level for the user 202. In the shown example, the LLM 206 includes the set of parameters 108 and the first set of ML parameters 116. In FIG. 2B, the user 202 can provide an input (e.g., a prompt) to the LLM 206 such that an output of the LLM 206 is to conform to the highest permission level of the user 202.
[0114] FIG. 2C illustrates generating a machine learning output for the user 202 of FIG. 2A at a reduced permission level for the user 202. In the shown example, the LLM 206 includes the set of parameters 108 and the first set of ML parameters 116. In FIG. 2C, the user 202 can provide an input (e.g., a prompt) to the LLM 206 such that an output of the LLM 206 is to conform to a reduced permission level of the user 202. For example, the user 202 can instruct and / or otherwise cause the LLM 206 to generate an output using a portion of the parameters of the LLM 206, such as by using the set of parameters 108 and not the first set of ML parameters 116.
[0115] In some embodiments, such as those illustrated by the examples of FIGS. 2 A, 2B, and / or 2C, given a classification marking of an interaction (e.g., a question-and-answer interaction between the user 202 and the LLM 206, a prompt-and-answer interaction between the user 202 and the LLM 206, a conversation between two users via one or more LLMs) and the access permissions of a user, one can automatically determine the LLM that should help them. Every information silo that is shared in the conversation should be positive and every information silo that the user has access to but is not part of the conversation should be negative. This prevents inadvertent leaks. The LLM will avoid saying something that happens to have higher classification, like mentioning two innocuous terms that together are sensitive.This also prevents the user 202 from exploiting the system. For example, the machine learning model generator software application 102 may not plug in any silo-tuning (e.g., a set of parameters) that the user 202 themselves do not have access to. Thusly, the user 202 cannot ask a question about sensitive information that the user 202 does not have access to and get a revealing denial.
[0116] In some embodiments, the machine learning model generator software application 102 can detect anomalies in data streams. In FIGS. 2B and / or 2C, the data streams can be implemented by input (e.g., USER 1 INPUT) and / or the output (e.g., LLM OUTPUT) via the LLM 206. For example, the machine learning model generator software application 102 can detect whether the input and / or the output is representative of information that the user 202 is not authorized and / or permitted to access and / or otherwise have possession.
[0117] In some embodiments, the machine learning model generator software application 102 can detect anomalies in data streams by processing input data and / or output data into one or more quantifiable metrics. An example of a quantifiable metric is perplexity.
[0118] As used herein, “perplexity” may refer to a metric used in ML and natural language processing (NLP) to measure a degree to which a language model, such as an LLM, predicts the next character or word in a sequence. A quantification of perplexity can be implemented by a perplexity score.
[0119] Perplexity is directly related to a model’s uncertainty and unpredictability. For example, a lower perplexity score can indicate that an LLM is more confidence in its predictions and is better at accurately predicting the data on which the LLM is trained. A higher perplexity score can indicate that the LLM’s predictions are less accurate.Accordingly, an LLM with lower perplexity (e.g., lower perplexity score(s)) can indicate that the LLM is a better (e.g., more accurate) LLM than an LLM with a higher perplexity score. For example, an LLM having a perplexity score of 10 on a given text dataset can indicate that, on average, the LLM is as confused as if it had to choose uniformly and independently from 10 possibilities for each word.
[0120] In some embodiments, perplexity (e.g., a perplexity score, a perplexity value) is calculated and / or determined using cross-entropy (e.g., average cross-entropy), which in turn is calculated using the number of words in a dataset and the predicted probability of a word (e.g., a target word) as per the preceding context. In some such embodiments, the preceding context can be represented by a fixed-length sequence of words that precede the target word.
[0121] In some embodiments, cross-entropy may be determined by computing the cross-entropy loss between the predicted word probabilities and the actual word probabilities. Cross-entropy loss measures the difference between two probability distributions. For example, the predicted word probabilities can correspond to probabilities output from a pretrained LLM, such as the pre-trained LLM 106 of FIG. 1. In such an example, the actual word probabilities can correspond to probabilities output from the LLM 206 of FIGS. 2B and / or 2C. Accordingly, the cross-entropy loss can be determined by measuring the difference between the probabilities output from the pre-trained LLM 106 and the LLM 206.
[0122] A lower cross-entropy loss may indicate less deviation between outputs of the pre-trained LLM 106 and the LLM 206 such that expected behavior or trained understanding is observed and a detection of an anomaly in a given data stream is unlikely. A higher crossentropy loss may indicate greater deviation between outputs of the pre-trained LLM 106 and the LLM 206 such that a deviation from expected behavior or trained understanding is observed and a detection of an anomaly in a given data stream is likely.
[0123] In some embodiments, perplexity for an LLM may be represented as the exponential of the average negative log-likelihood of a sequence of words as represented by the example of Equation (1) below:
[0124] Equation (1),
[0125] where N is the total number of words in the sequence,(P(Wj |w1(w2, ... , Wj_x)) is the probability of the word wtgiven the preceding words w1(w2, ... , Wj_x, and log is the natural logarithm.
[0126] In some embodiments, a perplexity score for comparing two LLMs may be implemented by computing the natural exponent of the fine-tuned loss (e.g., the loss of the LLM 206) minus the pre-trained model loss (e.g., the loss of the pre-trained LLM 106) as represented by the example of Equation (2) below:
[0127] Equation (2),
[0128] where &uuis the perplexity score, the pre-trained model is fe, the fine-tuned model is fg>, the labels as z, and model loss as l(fe,z).
[0129] In some embodiments, the perplexity score can be determined using logits to determine how closely related any plain-text embedding is to a given composed model MT. For example, let MG be a general knowledge model (e.g., a pre-trained LLM) that was autoregressively trained on a very large dataset. Given N data silos { &, &, ... , SN } ofproprietary information, N fine-tuned LLMs { Mi, M2, ... , MN } are created. For example, Mi can correspond to the pre-trained LLM 106 configured with the first set of ML parameters 116, M can correspond to the pre-trained LLM 106 configured with the second set of ML parameters 122, and so on. By iteratively evaluating the perplexity of a plain-text statement ho, the machine learning model generator software application 102 can determine if ho E St, Vi. If the output logits from MG are used, then the perplexity will naturally skew towards MG. Advantageously, using plain text as the starting point for input, allows for this technique to be used to detect leaks of proprietary information from human-generated text as well.
[0130] In some embodiments, conversion from plain text to logits used to compute perplexity is implemented by tokenizing the plain-text, and then assigning a value of 1 to the corresponding index z of the logit vector X / vfor each token k and 0 for every other index in the logit vector.
[0131] XI= f (X / ) ' ==0n,1other=wi Mse )
[0132] In some embodiments, given fine-tunings to compose Mi, . . .nand input x, logit composition may be implemented by performing the complete forward pass for each fine-tuning independently to obtain logit probabilities. Then those logits from each composed model are compared in the last layer of the model by element-wise sum or maximum for each logit position to result in a one-dimensional logit vector. In some such embodiments, the perplexity score for a fine-tuning can be determined using the one-dimensional logit vector and can be compared against perplexity scores determined using other fine-tunings.
[0133] By way of example, the machine learning model generator software application 102 can receive and / or monitor (e.g., actively monitor, passively monitor) a data stream, such as the input and / or output of FIGS. 2B and / or 2C. The machine learning model generator software application 102 can process, using the LLM 206, the input into the output. The machine learning model generator software application 102 can determine a perplexity score for the output in accordance with Equation (1) and / or Equation (2) above. Additionally and / or alternatively, the machine learning model generator software application 102 can process, using the pre-trained LLM 106 of FIG. 1, the input into an output. The machine learning model generator software application 102 can determine a perplexity score for this output in accordance with Equation (1) and / or Equation (2) above.
[0134] Furthering the example, the machine learning model generator software application 102 can determine a cross-entropy loss by measuring the difference between the perplexity score for the LLM 206 and the pre-trained LLM 106 (e.g., in accordance withEquation (2) above). The machine learning model generator software application 102 can detect an existence of anomaly in the data stream by determining whether the cross-entropy loss satisfies a threshold. For example, the machine learning model generator software application 102 can detect an anomaly in the data stream by determining that the crossentropy loss is greater than a cross-entropy loss threshold value and the cross-entropy loss thereby satisfies the cross-entropy loss threshold. Alternatively, the machine learning model generator software application 102 may detect an anomaly in the data stream by determining that the cross-entropy loss is less than a cross-entropy loss threshold value and the crossentropy loss thereby satisfies the cross-entropy loss threshold.
[0135] The machine learning model generator software application 102 can determine, responsive to a detection of an anomaly, that the data stream includes data leaked from a secure / sensitive data source. For example, a lower perplexity may indicate a lack of access to certain information silos, whereas a higher perplexity may indicate access to the same silos. Correspondingly, a lower cross-entropy loss may indicate a lack of access to certain information silos, whereas a higher cross-entropy loss may indicate access to the same silos.
[0136] In some embodiments, the machine learning model generator software application 102 can determine a source of the leaked data in the data stream such that the leaked data can be attributed to one or more information silos that the user 202 may not have access. In some such embodiments, the machine learning model generator software application 102 can iteratively process a data stream using a plurality of fine-tunings into a respective perplexity score and determine, from the perplexity scores, which of the plurality of fine-tunings is more closely associated with the data stream. For example, the machine learning model generator software application 102 can identify an information silo as a source of information in the data stream by determining, using the perplexity scores, that a fine-tuning (e.g., one of the sets of parameters 120) trained on the information silo is more closely associated with the data stream than other fine-tunings.
[0137] By way of example, the machine learning model generator software application 102 can process, using a plurality of fine-tunings, a data stream, such as user 1 input of FIGS. 2B and / or 2C, into a respective perplexity score. For example, the machine learning model generator software application 102 can process the data stream, using the pretrained LLM 106 configured with the set of parameters 108, into a first perplexity score. The machine learning model generator software application 102 can process the data stream, using a second LLM configured with the set of parameters 108 and the first set of MLparameters 116, into a second perplexity score. The machine learning model generator software application 102 can process the data stream, using a third LLM configured with the set of parameters 108, the first set of ML parameters 116, and the second set of parameters 122, into a third perplexity score. In the above example, the set of parameters 108 can correspond to a first information silo, the first set of ML parameters 116 can correspond to a second information silo, and the second set of parameters 122 can correspond to a third information silo.
[0138] Furthering the example, the machine learning model generator software application 102 can determine that the data stream includes information from at least one of the first, second, or third information silos using the perplexity scores. For example, the first perplexity score may be 0.32, the second perplexity score may be 0.84, and the third perplexity score may be 0.95. In such an example, the machine learning model generator software application 102 can determine that the data stream includes information from the third information silo based on the third perplexity score being the highest of the perplexity scores.
[0139] FIG. 3 illustrates an example workflow 300 for generating an interaction specific machine learning model for at least two users. The workflow 300 of this example includes the machine learning model generator software application 102 and the LLM parameter datastore 118 of FIG. 1. The LLM parameter datastore 118 includes the set of parameters 108, the first set of ML parameters 116, and the second set of parameters 122 of FIG. 1.
[0140] Further shown in this example are a first user 302 and a second user 304. The first user 302 has and / or is associated with first permission data 306. The second user 304 has and / or is associated with second permission data 308. In some embodiments, the first permission data 306 and / or the second permission data 308 can be implemented by and / or correspond to the user data 110 of FIG. 1. For example, the first permission data 306 can correspond to USER 1 data of FIG. 1 and / or the second permission data 308 can correspond to USER 2 data (or USER N data).
[0141] In some embodiments, the first permission data 306 includes permission level(s) for the first user 302. For example, the first permission data 306 can include at least a first permission level assigned to the first user 302. In such an example, the first permission level can indicate that the first user 302 is authorized and / or permitted to access information associated with one(s) of the sets of parameters 120. For example, the first permission levelcan indicate that the first user 302 is authorized to access information associated with the set of parameters 108, the first set of ML parameters 116, and the second set of parameters 122.
[0142] In some embodiments, the second permission data 308 includes permission level(s) for the second user 304. For example, the second permission data 308 can include at least a second permission level assigned to the second user 304. In such an example, the second permission level can indicate that the second user 304 is authorized and / or permitted to access information associated with one(s) of the sets of parameters 120. For example, the second permission level can indicate that the second user 304 is authorized to access information associated with the set of parameters 108 and the first set of ML parameters 116.
[0143] In the illustrated example, the first user 302 is authorized to access one(s) of the datastores 112 of FIG. 1 as indicated by the first permission data 306. As shown, the first user 302 is permitted to access information silos corresponding to the set of parameters 108, the first set of ML parameters 116, and the second set of parameters 122.
[0144] In the illustrated example, the second user 304 is authorized to access one(s) of the datastores 112 of FIG. 1 as indicated by the second permission data 308. In some embodiments, the second user 304 can access different one(s) of the datastores 112 than the first user 302. As shown, the second user 304 is permitted to access information silos corresponding to the set of parameters 108 and the first set of ML parameters 116 but not the second set of parameters 122. For example, the permission level associated with the first user 302 is higher than the permission level of the second user 304 with respect to at least the second set of parameters 122 because the first user 302 is permitted to access information associated with the second set of parameters 122 while the second user 304 is not permitted to access information associated with the second set of parameters 122.
[0145] In the illustrated example of FIG. 3, the first user 302 and the second user 304 intend to have an electronic interaction (e.g., an electronic conversation) with each other such as by sending electronic messages via a messaging application. For example, the first user 302 and / or the second user 304 can provide electronic input to the machine learning model generator software application 102 to start an electronic conversation between the first user 304 and the second user 304. Additionally and / or alternatively, the first user 302 and the second user 304 can have an electronic interaction via audio (e.g., audio messages, a telephone call) and / or video (e.g., a video conference).
[0146] In the shown example, the first user 302 and the second user 304 effectuate an electronic interaction via an interaction specific LLM 310 (identified by INTERACTION LLM). For example, the machine learning model generator software application 102 canobtain the first permission data 306 and the second permission data 308 and, based on the permission data, can identify which one(s) of the plurality of sets of parameters 120 the users 302, 304 have in common. Based on the data associations of the users 302, 304 and the one(s) of the plurality of sets of parameters 120, the machine learning model generator software application 102 can generate the interaction specific LLM 310 to include the set of parameters 108 for the pre-trained model 106 and the first set of ML parameters 116.
[0147] In example operation, the interaction specific LLM 310 can effectuate an interaction, such as a conversation, between the users with the appropriate degree of permission access. For example, the interaction specific LLM 310 can be executed and / or run iteratively. The users 302, 304 can ask a question and retroactively ask, what is the lowest classification scheme at which they can use this answer. The users 302, 304 can set an agreed upon permission level of the interaction. For example, the users 302, 304 can agree to a permission level up to a permission level associated with the first set of ML parameters 116, to which both are authorized, or to a lower permission level to that of the set of parameters 108 if they desire a lower permission level for the interaction.
[0148] An example way to implement this schema is for the interaction specific LLM 310 to generate the answer, then have the interaction specific LLM 310 compute its likelihood (e.g., probability) for different classification markings (which is substantially faster than the interaction specific LLM 310 generating a new answer) and using a dynamic programming approach to find the lowest marking for that answer. In some embodiments, the efficiency of this algorithm depends on the properties of the marking scheme, such as having a strict order, a partial order, being encodable as a lattice, etc., with each progressively increasing level of marking requiring more sophisticated search algorithms. Similarly this can be used for interactions as well, such as the example shown in FIG. 3, to keep an interaction at a steady level.
[0149] By way of another example, one of the users 302, 304 can ask the interaction specific LLM 310 to briefly increase the classification level of the interaction before lowering it. In some embodiments, the interaction specific LLM 310 can generate an alert to one or both users 302, 304 when the minimum classification level of the interaction is likely changing by passively listening to the interaction with the credentials of each user and the target classification level.
[0150] In some embodiments, the machine learning model generator software application 102 detects an anomaly associated with a data stream between the users 302, 304 via the interaction specific LLM 310. For example, the machine learning model generatorsoftware application 102 can determine whether data, information, etc., in the interaction between the users 302, 304 contain data, information, etc., that one or both users 302, 304 is / are not authorized to be presented with and / or otherwise have access.
[0151] In some embodiments, the machine learning model generator software application 102 can detect an anomaly associated with a data stream by processing at least a portion of the data stream using one or more fine-tunings. At least a portion of the data stream may include text, such as plain text.
[0152] By way of example, the machine learning model generator software application 102 can process, using the pre-trained model 106 of FIG. 1, plain text from at least a portion of the data stream into a first model loss. The machine learning model generator software application 102 can tokenize the plain text to generate tokenized text. The machine learning model generator software application 102 can determine a logit vector using the tokenized text (e.g., process the tokenized text into a logit vector). The machine learning model generator software application 102 can process, using the pre-trained model 106, the logit vector to generate a first model loss associated with the pre-trained model 106 (e.g., a pre-trained model loss).
[0153] Furthering the example, the machine learning model generator software application 102 can process, using the interaction specific LLM 310, the logit vector to generate a second model loss associated with the interaction specific LLM 310 (e.g., a finetuned model loss). The machine learning model generator software application 102 can determine a perplexity score using the first model loss and the second model loss in accordance with Equation (2) above.
[0154] The machine learning model generator software application 102 can determine whether the perplexity score satisfies a threshold. For example, if the machine learning model generator software application 102 determines that the perplexity score is greater than a threshold value, the machine learning model generator software application 102 can detect an anomaly associated with the plain text. Otherwise, the machine learning model generator software application 102 may determine that the plain text is not anomalous. The machine learning model generator software application 102 may continue to perform anomaly detection on the data stream between the users 302, 304 by monitoring (e.g., passively monitoring) the data stream and processing the data stream portion(s) as described above.
[0155] In some embodiments, the machine learning model generator software application 102 can identify a source of a data stream between the users 302, 304. For example, the machine learning model generator software application 102 can process at leasta portion of plain text from the data stream to identify one or more data sources, such as a datastore and / or information silo, associated with the portion of plain text. In an example, association with the one or more data sources may refer to the plain text originating from the one or more data sources. For example, data may have been leaked from the one or more data sources such that one or both users 302, 304 were presented with the leaked data in a different setting or environment. In another example, association with the data source(s) may refer to an LLM trained on the data source(s) may generate an output that includes the plain text.
[0156] In some embodiments, the machine learning model generator software application 102 can identify a source associated with at least a portion of a data stream by processing at least a portion of the data stream using one or more fine-tunings. At least a portion of the data stream may include text, such as plain text.
[0157] By way of example, the machine learning model generator software application 102 can process, using the pre-trained model 106 of FIG. 1, plain text from at least a portion of the data stream into a first model loss. The first model loss may have a value of 0.51. The machine learning model generator software application 102 can tokenize the plain text to generate tokenized text. The machine learning model generator software application 102 can determine a logit vector using the tokenized text (e.g., process the tokenized text into a logit vector). The machine learning model generator software application 102 can process, using the pre-trained model 106, the logit vector to generate a first model loss associated with the pre-trained model 106 (e.g., a pre-trained model loss).
[0158] Furthering the example, the machine learning model generator software application 102 can process, using the interaction specific LLM 310, the logit vector to generate a second model loss of 0.89 associated with the interaction specific LLM 310 (e.g., a fine-tuned model loss). The machine learning model generator software application 102 can process, using another fine-tuning such as the pre-trained LLM 106 configured with the second set of parameters 122, the logit vector to generate a third model loss of 0.92 associated with this LLM.
[0159] Continuing the example, the machine learning model generator software application 102 can identify, from the model losses, which source is associated with the plain text. For example, the machine learning model generator software application 102 can determine that the plain text is associated with the second set of parameters 122 because the third model loss had the largest value. Having the largest value in this example can indicate that the plain text had the closest association to the second set of parameters 122, which weregenerated using a particular source of data, such as datastore N of FIG. 1, or portion(s) thereof (e.g., info N).
[0160] In some embodiments, the machine learning model generator software application 102 can identify, from perplexity scores, which source is associated with the plain text. For example, the machine learning model generator software application 102 can determine a first perplexity score based on the first and second model losses in accordance with Equation (2) above and a second perplexity score based on the first and third model losses in accordance with Equation (2) above. In such an example, the machine learning model generator software application 102 can identify that the plain text is more closely related to a data source associated with the first set of ML parameters 116 than the second set of parameters 122 when the first perplexity score is greater than the second perplexity score. In another example, the machine learning model generator software application 102 can identify that the plain text is more closely related to a data source associated with the second set of ML parameters 122 than the first set of parameters 116 when the second perplexity score is greater than the first perplexity score.
[0161] In some embodiments, the machine learning model generator software application 102 can determine that an LLM trained on datastore N, which led to the generation of the second set of parameters 122, is likely to be associated with the plain text. In an example, the plain text can be stored in datastore N. In another example, the LLM trained on datastore N can process input into output, which can include the plain text. In yet another example, the LLM trained on datastore N including the plain text can process input into output such that the output is generated at least in part based on the plain text.
[0162] In some embodiments, responsive to the identification of the source(s), such as one or more of the plurality of datastores 112 of FIG. 1 to which the plain text is associated, the machine learning model generator software application 102 can determine a likelihood that one or more permission levels associated with the identified source(s) conforms to the permission level of at least one user 302, 304. For example, in response to identifying datastore N as being associated with the plain text from the interaction between the users 302, 304, the machine learning model generator software application 102 can determine a permission level associated with datastore N. In such an example, the machine learning model generator software application 102 can determine whether a permission level of the first user 302 and / or the second user 304 conforms to the permission level of datastore N.
[0163] In some embodiments, the machine learning model generator software application 102 can generate an alert if the permission level of datastore N is higher than thepermission level of that of the first user 302 and / or the second user 304. For example, the 102 can alert, inform, and / or notify the first user 302 that the permission level of the first user 302 and / or the second user 304 is lower than the permission level associated with datastore N, which is implicated from data, information, etc., in the data stream. Responsive to the alert, the first user 302 can terminate the interaction.
[0164] FIGS. 4-7 are flowcharts 400, 500, 600, 700 representative of example processes to be performed and / or example machine-readable instructions that may be executed by processor circuitry to implement the machine learning model generator software application 102 of at least FIG. 1. Additionally and / or alternatively, block(s) of one(s) of the flowcharts 400, 500, 600, 700 of FIGS. 4, 5, 6, and / or 7 may be representative of state(s) of one or more hardware-implemented state machines, algorithm(s) that may be implemented by hardware alone such as an ASIC, etc., and / or any combination(s) thereof.
[0165] FIG. 4 is a flowchart 400 representative of an example process that may be performed and / or implemented using hardware logic and / or example machine-readable instructions that may be executed by processor circuitry to implement the machine learning model generator software application 102 of at least FIG. 1 to generate a user specific machine learning model.
[0166] The flowchart 400 of FIG. 4 begins at block 402, at which the machine learning model generator software application 102 may train at least one pre-trained machine learning (ML) model using training data to generate a set of ML parameters associated with the training data. For example, the machine learning model generator software application 102 may fine-tune the pre-trained model 106 using training data including at least one of information in at least one of the datastores 112 of FIG. 1 or data associated with at least one of the APIs 114 to generate a set of ML parameters, such as the first set of ML parameters 116 of FIG. 1. In some such embodiments, the set of ML parameters can correspond to at least one of the information in the at least one of the datastores 112 or the data associated with the at least one of the APIs 114.
[0167] At block 404, the machine learning model generator software application 102 may generate data association(s) of user(s) and the set of ML parameters after determining that user data of the user(s) correspond to the information in the datastore(s). For example, the machine learning model generator software application 102 may generate data association(s), such as by linking and / or tagging data (e.g., metadata, permission data), of one or more users (e.g., USER 1, USER 2, etc.) and the first set of ML parameters 116 after determining that portion(s) of user data of the one or more users (as indicated by their respective user data)correspond to the at least one of the information in the at least one of the datastores 112 or the data associated with the at least one of the APIs 114.
[0168] At block 406, the machine learning model generator software application 102 may determine whether to generate a set of ML parameters for other sets of training data. For example, the machine learning model generator software application 102 may determine to generate another set of ML parameters for another set of the datastores 112 not previously processed and / or for another set of the APIs 114.
[0169] If, at block 406, the machine learning model generator software application 102 determines to generate a set of ML parameters for other datastore(s), control returns to block 402. Otherwise, control proceeds to block 408.
[0170] At block 408, the machine learning model generator software application 102 may select at least one user to process. For example, the machine learning model generator software application 102 may select user 1 of FIG. 1 to process.
[0171] At block 410, the machine learning model generator software application 102 may identify one(s) of the sets of ML parameters that correspond to the at least one user. For example, the machine learning model generator software application 102 may identify at least the set of parameters 108 and the first set of ML parameters 116 as corresponding to user 1. In such an example, the machine learning model generator software application 102 may identify the first set of ML parameters 116 as corresponding to user 1 by determining that permission access data (e.g., the permission access data 306, 308 of FIG. 3) for user 1 is associated with the first set of ML parameters 116. Furthering the example, the permission access level and / or classification level indicated by the permission access data for user 1 may at least partially match the permission access level and / or classification level assigned to the first set of ML parameters 116.
[0172] At block 412, the machine learning model generator software application 102 may generate a trained ML model for the at least one user based on a combination of the at least one pre-trained model and the identified one(s) of the sets of ML parameters. For example, the machine learning model generator software application 102 may generate user 1 LLM of the models 104 of FIG. 1 for user 1 based on a combination of the pre-trained model 106 and the first set of ML parameters 116. In such an example, the machine learning model generator software application 102 may generate user 1 LLM by attaching, connecting, and / or otherwise combining at least the set of parameters 108 and the first set of ML parameters 116. In some such embodiments, the machine learning model generator software application 102 may generate user 1 LLM as an executable (e.g., an executable file) that can beinstantiated and / or executed to process input (e.g., audio, image, text, video) into output (e.g., audio, image, text, video). In some such embodiments, the executable may be stored in a datastore (e.g., the LLM parameter datastore 118) that is local to the machine learning model generator software application 102 (e.g., accessible via an electronic bus, on the same server) and / or separate from the machine learning model generator software application 102 (e.g., accessible via network(s)).
[0173] At block 414, the machine learning model generator software application 102 may determine whether to select additional user(s) to process. For example, the machine learning model generator software application 102 may determine to select user N of FIG. 1 to process.
[0174] If, at block 414, the machine learning model generator software application 102 determines to select additional user(s) to process, control returns to block 408. Otherwise, the example flowchart 400 of FIG. 4 concludes.
[0175] FIG. 5 is a flowchart 500 representative of an example process that may be performed and / or implemented using hardware logic and / or example machine-readable instructions that may be executed by processor circuitry to implement the machine learning model generator software application 102 of at least FIG. 1 to generate an interaction specific machine learning model. For example, the flowchart 500 of FIG. 5 may correspond to generating an interaction specific machine learning model as described above in connection with FIG. 3.
[0176] The flowchart 500 of FIG. 5 begins at block 502, at which the machine learning model generator software application 102 may obtain permission data for two or more users. For example, the machine learning model generator software application 102 may obtain the first permission data 306 and the second permission data 308 of FIG. 3.
[0177] At block 504, the machine learning model generator software application 102 may identify sets of machine learning (ML) parameters in accordance with the permission data. For example, the machine learning model generator software application 102 may identify one(s) of the plurality of sets of parameters 120 shown in FIG. 3 that correspond to the first user 302 and the second user 304 by using the first permission data 306 and the second permission data 308.
[0178] At block 506, the machine learning model generator software application 102 may generate an interaction specific ML model by configuring a pre-trained ML model using the identified sets of ML parameters. For example, the machine learning model generator software application 102 may generate the interaction specific LLM 310 by configuring thepre-trained model 106 of FIG. 1 using the set of parameters 108 and the first set of ML parameters 116.
[0179] At block 508, the machine learning model generator software application 102 may effectuate an interaction between the two or more users. For example, the interaction specific LLM 310 may obtain user input, prompts, etc., from one(s) of the users 302, 304 and generate responses.
[0180] At block 510, the machine learning model generator software application 102 may determine whether the permission level for the interaction has changed. For example, the machine learning model generator software application 102 may determine that the user input and / or responses indicate that the permission level for the interaction has changed, such as by generating responses using information from information silos at a higher permission level than the level set for the interaction. In some embodiments, the interaction specific LLM 310 may generate an alert informing one or both users 302, 304 that the permission level has changed. For example, the interaction specific LLM 310 may generate an alert informing one or both users 302, 304 that the permission level has been escalated to a higher permission level than the level initially set for the interaction.
[0181] If, at block 510, the machine learning model generator software application 102 determines that the permission level for the interaction has changed, control proceeds to block 512. At block 512, the machine learning model generator software application 102 may re-configure the interaction specific LLM in accordance with the changed permission level. For example, the machine learning model generator software application 102 may add another set of parameters to the interaction specific LLM 310, if available to both users 302, 304, to at least temporarily increase the permission level of the interaction. In such an example, the machine learning model generator software application 102 may further configure the pre-trained model 106 of FIG. 1 using the second set of ML parameters 122.
[0182] Additionally and / or alternatively, in some embodiments, the interaction specific LLM 310 may generate an alert informing one or both users 302, 304 that the permission level is unable to be changed. In response to re-configuring the interaction specific LLM in accordance with the changed permission level, control returns to block 508 to effectuate an interaction between the two or more users in accordance with the changed permission level.
[0183] If, at block 510, the machine learning model generator software application 102 determines that the permission level for the interaction has not changed, control proceeds to block 514. At block 514, the machine learning model generator software application 102 may determine whether the interaction ended. For example, the machine learning model generatorsoftware application 102 may determine whether one(s) of the users 302, 304 have terminated the interaction explicitly (e.g., one or both users exited from the interaction) or implicitly (e.g., a time period has elapsed since the last received user input).
[0184] If, at block 514, the machine learning model generator software application 102 determines that the interaction has not ended, control returns to block 508 to continue to effectuate an interaction between the two or more users. Otherwise, the example flowchart 500 of FIG. 5 concludes.
[0185] FIG. 6 is a flowchart 600 representative of an example process that may be performed and / or example machine-readable instructions that may be executed by processor circuitry to implement the machine learning model generator software application 102 of at least FIG. 1 to perform anomaly detection. The flowchart 600 of FIG. 6 begins at block 602 at which the machine learning model generator software application 102 may tokenize input from a user. For example, the machine learning model generator software application 102 may receive input from a data stream between the users 302, 304 of FIG. 3. In such an example, the input may be text, such as plain text. The machine learning model generator software application 102 may tokenize the plain text.
[0186] At block 604, the machine learning model generator software application 102 may determine a logit vector using the tokenized input. For example, the machine learning model generator software application 102 may process the plain text as input into one or more tokens as output.
[0187] At block 606, the machine learning model generator software application 102 may select a set of machine learning (ML) parameters associated with the user. For example, the machine learning model generator software application 102 may select the first set of ML parameters 116, which is associated with the first user 302 and the second user 304 as shown in FIG. 3.
[0188] At block 608, the machine learning model generator software application 102 may generate an ML model using a pre-trained ML model and the selected set of ML parameters. For example, the machine learning model generator software application 102 may assemble an LLM by configuring the pre-trained model 106 using the set of parameters 108 and the first set of ML parameters 116. The assembled LLM may be the interaction specific LLM 310 of FIG. 3.
[0189] At block 610, the machine learning model generator software application 102 may process the logit vector using the ML model to output a perplexity score. For example, the machine learning model generator software application 102 may process the logit vectorusing the pre-trained model 106 to generate a first model loss. The machine learning model generator software application 102 may process the logit vector using the interaction specific LLM 310 to generate a second model loss. The machine learning model generator software application 102 may determine a perplexity score based on the first and second model losses in accordance with Equation (2) above.
[0190] At block 612, the machine learning model generator software application 102 may determine whether the perplexity score indicates a detection of an anomaly. For example, the machine learning model generator software application 102 may determine that, when the perplexity score exceeds a perplexity score threshold, the input indicates an anomaly from expected model behavior. In another example, the machine learning model generator software application 102 may determine that, when the perplexity score is less than a perplexity score threshold, the input indicates that the input represents expected model behavior.
[0191] If, at block 612, the machine learning model generator software application 102 determines that the perplexity score does not indicate a detection of an anomaly, the example flowchart 600 of FIG. 6 concludes. Otherwise, control proceeds to block 614.
[0192] At block 614, the machine learning model generator software application 102 may generate an output representative of an anomaly detection associated with the input. For example, the machine learning model generator software application 102 may generate an alert and / or provide the alert to the first user 302 and / or the second user 304. In such an example, responsive to the alert, the machine learning model generator software application 102, the first user 302, and / or the second user 304 may terminate the interaction shown in FIG. 3. After generating the output at block 614, the example flowchart 600 of FIG. 6 concludes.
[0193] FIG. 7 is a flowchart 700 representative of an example process that may be performed and / or example machine-readable instructions that may be executed by processor circuitry to implement the machine learning model generator software application 102 of at least FIG. 1 to perform data leak detection. The flowchart 700 of FIG. 7 begins at block 702 at which the machine learning model generator software application 102 may tokenize input. For example, the machine learning model generator software application 102 may receive input from a data stream between the users 302, 304 of FIG. 3. In such an example, the input may be text, such as plain text. The machine learning model generator software application 102 may tokenize the plain text.
[0194] At block 704, the machine learning model generator software application 102 may determine a logit vector using the tokenized input. For example, the machine learning model generator software application 102 may process the plain text as input into one or more tokens as output.
[0195] At block 706, the machine learning model generator software application 102 may select a set of machine learning (ML) parameters. For example, the machine learning model generator software application 102 may select the first set of ML parameters 116.
[0196] At block 708, the machine learning model generator software application 102 may generate an ML model using a pre-trained ML model and the selected set of ML parameters. For example, the machine learning model generator software application 102 may assemble an LLM by configuring the pre-trained model 106 using the set of parameters 108 and the first set of ML parameters 116.
[0197] At block 710, the machine learning model generator software application 102 may process the logit vector using the ML model to output a perplexity score. For example, the machine learning model generator software application 102 may process the logit vector using the pre-trained model 106 to generate a first model loss. The machine learning model generator software application 102 may process the logit vector using the LLM configured with the set of parameters 108 and the first set of ML parameters 116 to generate a second model loss. The machine learning model generator software application 102 may determine a first perplexity score based on the first and second model losses in accordance with Equation (2) above.
[0198] At block 712, the machine learning model generator software application 102 may determine whether to select another set of ML parameters to process. For example, the machine learning model generator software application 102 may determine to select the second set of parameters 122 to process.
[0199] If, at block 712, the machine learning model generator software application 102 determines to select another set of ML parameters to process, control returns to block 708. At block 708, the machine learning model generator software application 102 may generate another ML model using the pre-trained ML model and the selected set of ML parameters. For example, the machine learning model generator software application 102 may assemble another LLM by configuring the pre-trained model 106 using the set of parameters 108, the first set of ML parameters 116, and the second set of parameters 122. At block 710, the machine learning model generator software application 102 may process the logit vector using the ML model to output a perplexity score. For example, the machine learning modelgenerator software application 102 may process the logit vector using the LLM configured with the set of parameters 108, the first set of ML parameters 116, and the second set of parameters 122 to generate a third model loss. The machine learning model generator software application 102 may determine a second perplexity score based on the first and third model losses in accordance with Equation (2) above.
[0200] If, at block 712, the machine learning model generator software application 102 determines not to select another set of ML parameters to process, control proceeds to block 714.
[0201] At block 714, the machine learning model generator software application 102 outputs an identification of which of the ML model(s) is associated with the input using the perplexity score(s). As an example, the first set of ML parameters 116 may be generated by training an LLM using a first data source, such as datastore 1 of FIG. 1, and the second set of parameters 122 may be generated by training an LLM using a second data source, such as datastore 2 of FIG. 1. Furthering the example, if the first perplexity score is greater than the second perplexity score, the machine learning model generator software application 102 can determine that the input is more closely related to the first data source than the second data source. Conversely, if the first perplexity score is less than the second perplexity score, the machine learning model generator software application 102 can determine that the input is more closely related to the second data source than the first data source. Accordingly, the machine learning model generator software application 102 may output (e.g., to the system 100, to one(s) of the users 302, 304) an identification of a data source, such as datastore 1 and / or datastore 2, to which the input is likely to be associated with. After outputting the identification at block 714, the example flowchart 700 of FIG. 7 concludes.
[0202] FIG. 8 is an example implementation of an electronic platform 800 structured to execute the machine-readable instructions of FIGS. 4, 5, 6, and / or 7 to implement the machine learning model generator software application 102 of FIGS. 1, 2A, and / or 3. It should be appreciated that FIG. 8 is intended neither to be a description of necessary components for an electronic and / or computing device to operate as the machine learning model generator software application 102, in accordance with the techniques described herein, nor a comprehensive depiction.
[0203] The electronic platform 800 of this example may be an electronic device, such as a handset device (e.g., a cellular network device, a smartphone, etc.), a desktop computer, a laptop computer, a tablet computer, a server (e.g., a computer server, a blade server, a rackmounted server, etc.), a wearable device (e.g., an augmented reality and / or virtual reality(AR / VR) device, a heads-up display (HUD) device, a smartwatch, smart glasses, smart goggles, etc.), a workstation, or any other type of computing and / or electronic device.
[0204] The electronic platform 800 of the illustrated example includes processor circuitry 802, which may be implemented by one or more programmable processors, one or more hardware-implemented state machines, one or more ASICs, etc., and / or any combination(s) thereof. For example, the one or more programmable processors may include one or more CPUs, one or more DSPs, one or more FPGAs, one or more GPUs, etc., and / or any combination(s) thereof.
[0205] The processor circuitry 802 includes processor memory 804, which may be volatile memory, such as random-access memory (RAM) of any type. The processor circuitry 802 of this example implements the machine learning model generator software application 102 of at least FIG. 1.
[0206] The processor circuitry 802 may execute machine-readable instructions 806 (identified by INSTRUCTIONS), which are stored in the processor memory 804, to implement the machine learning model generator software application 102 of at least FIG. 1. The machine-readable instructions 806 may include data representative of computerexecutable and / or machine-executable instructions implementing techniques that operate according to the techniques described herein. For example, the machine-readable instructions 806 may include data (e.g., code, embedded software (e.g., firmware), software, etc.) representative of the flowcharts 400, 500, 600, 700 of FIGS. 4, 5, 6, and / or 7, or portion(s) thereof.
[0207] The electronic platform 800 includes memory 808, which may include the instructions 806. The memory 808 of this example may be controlled by a memory controller 810. For example, the memory controller 810 may control reads, writes, and / or, more generally, access(es) to the memory 808 by other component(s) of the electronic platform 800. The memory 808 of this example may be implemented by volatile memory, non-volatile memory, etc., and / or any combination(s) thereof. For example, the volatile memory may include static random-access memory (SRAM), dynamic random-access memory (DRAM), cache memory (e.g., Level 1 (LI) cache memory, Level 2 (L2) cache memory, Level 3 (L3) cache memory, etc.), etc., and / or any combination(s) thereof. In some examples, the nonvolatile memory may include Flash memory, electrically erasable programmable read-only memory (EEPROM), magnetoresistive random-access memory (MRAM), ferroelectric random-access memory (FeRAM, F-RAM, or FRAM), etc., and / or any combination(s) thereof.
[0208] The electronic platform 800 includes input device(s) 812 to enable data and / or commands to be entered into the processor circuitry 802. For example, the input device(s) 812 may include an audio sensor, a camera (e.g., a still camera, a video camera, etc.), a keyboard, a microphone, a mouse, a touchscreen, a voice recognition system, etc., and / or any combination(s) thereof.
[0209] The electronic platform 800 includes output device(s) 814 to convey, display, and / or present information to a user (e.g., a human user, a machine user, etc.). For example, the output device(s) 814 may include one or more display devices, speakers, etc. The one or more display devices may include an augmented reality (AR) and / or virtual reality (VR) display, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot (QLED) display, a thin-film transistor (TFT) LCD, a touchscreen, etc., and / or any combination(s) thereof. The output device(s) 814 can be used, among other things, to generate, launch, and / or present a user interface. For example, the user interface may be generated and / or implemented by the output device(s) 814 for visual presentation of output and speakers or other sound generating devices for audible presentation of output.
[0210] The electronic platform 800 includes accelerators 816, which are hardware devices to which the processor circuitry 802 may offload compute tasks to accelerate their processing. For example, the accelerators 816 may include artificial intelligence / machine-learning (AI / ML) processors, ASICs, FPGAs, graphics processing units (GPUs), neural network (NN) processors, systems-on-chip (SoCs), vision processing units (VPUs), etc., and / or any combination(s) thereof. In some examples, the machine learning model generator software application 102 may be implemented by one(s) of the accelerators 816 instead of the processor circuitry 802. In some examples, the machine learning model generator software application 102 may be executed concurrently (e.g., in parallel, substantially in parallel, etc.) by the processor circuitry 802 and the accelerators 816. For example, the processor circuitry 802 and one(s) of the accelerators 816 may execute in parallel function(s) corresponding to the machine learning model generator software application 102.
[0211] The electronic platform 800 includes storage 818 to record and / or control access to data, such as the machine-readable instructions 806. In this example, the storage 818 implements the models 104 (identified by USER LLMs), the user data 110, one(s) of the datastores 112, one(s) of the APIs 114, and the LLM parameter datastore 118. The storage 818 may be implemented by one or more mass storage disks or devices, such as HDDs, SSDs, etc., and / or any combination(s) thereof.
[0212] The electronic platform 800 includes interface(s) 820 to effectuate exchange of data with external devices (e.g., computing and / or electronic devices of any kind) via a network 822. The interface(s) 820 of the illustrated example may be implemented by an interface device, such as network interface circuitry (e.g., a NIC, a smart NIC, etc.), a gateway, a router, a switch, etc., and / or any combination(s) thereof. The interface(s) 820 may implement any type of communication interface, such as BLUETOOTH®, a cellular telephone system (e.g., a 4G LTE interface, a 5G interface, a future generation 8G interface, etc.), an Ethernet interface, a near-field communication (NFC) interface, an optical disc interface (e.g., a Blu- ray disc drive, a Compact Disk (CD) drive, a Digital Versatile Disk (DVD) drive, etc.), an optical fiber interface, a satellite interface (e.g., a BLOS satellite interface, a LOS satellite interface, etc.), a Universal Serial Bus (USB) interface (e.g., USB Type-A, USB Type-B, USB TYPE-C™ or USB-C™, etc.), etc., and / or any combination(s) thereof.
[0213] The electronic platform 800 includes a power supply 824 to store energy and provide power to components of the electronic platform 800. The power supply 824 may be implemented by a power converter, such as an alternating current-to-direct-current (AC / DC) power converter, a direct current-to-direct current (DC / DC) power converter, etc., and / or any combination(s) thereof. For example, the power supply 824 may be powered by an external power source, such as an alternating current (AC) power source (e.g., an electrical grid), a direct current (DC) power source (e.g., a battery, a battery backup system, etc.), etc., and the power supply 824 may convert the AC input or the DC input into a suitable voltage for use by the electronic platform 800. In some examples, the power supply 824 may be a limited duration power source, such as a battery (e.g., a rechargeable battery such as a lithium-ion battery).
[0214] Component(s) of the electronic platform 800 may be in communication with one(s) of each other via a bus 826. For example, the bus 826 may be any type of computing and / or electrical bus, such as an I2C bus, a PCI bus, a PCIe bus, a SPI bus, and / or the like.
[0215] The network 822 may be implemented by any wired and / or wireless network(s) such as one or more cellular networks (e.g., 4G LTE cellular networks, 5G cellular networks, future generation 8G cellular networks, etc.), one or more data buses, one or more local area networks (LANs), one or more optical fiber networks, one or more private networks, one or more public networks, one or more wireless local area networks (WLANs), etc., and / or any combination(s) thereof. For example, the network 822 may be the Internet, but any other type of private and / or public network is contemplated.
[0216] The network 822 of the illustrated example facilitates communication between the interface(s) 820 and a central facility 828. The central facility 828 in this example may be an entity associated with one or more servers, such as one or more physical hardware servers and / or virtualizations of the one or more physical hardware servers. For example, the central facility 828 may be implemented by a public cloud provider, a private cloud provider, etc., and / or any combination(s) thereof. In this example, the central facility 828 may compile, generate, update, etc., the machine-readable instructions 806 and store the machine-readable instructions 806 for access (e.g., download) via the network 822. For example, the electronic platform 800 may transmit a request, via the interface(s) 820, to the central facility 828 for the machine-readable instructions 806 and receive the machine-readable instructions 806 from the central facility 828 via the network 822 in response to the request.
[0217] Additionally and / or alternatively, the interface(s) 820 may receive the machine- readable instructions 806 via non-transitory machine-readable storage media, such as an optical disc 830 (e.g., a Blu-ray disc, a CD, a DVD, etc.) or any other type of removable non- transitory machine-readable storage media such as a USB drive 832. For example, the optical disc 830 and / or the USB drive 832 may store the machine-readable instructions 806 thereon and provide the machine-readable instructions 806 to the electronic platform 800 via the interface(s) 820.
[0218] Advantageously, the machine learning model generator software application 102 of at least FIG. 1 and / or, more generally, the system 100 of FIG. 1, has many benefits. For example, the resulting models (e.g., the models 104 of FIG. 1) are resistant to known privilege escalation attacks. A user cannot exfiltrate information they do not have access to. Users cannot jailbreak or fool the resulting models into revealing information simply because that fine-tuning is not attached to the models. Users cannot poison the training data of the models because they do not have access to it, this happens automatically without human intervention following the classification scheme. Similarly, other members of an interaction cannot fool a user’s model with more permission than their own into revealing information through a prompt injection attack because above-classification data is negative in the model. The resulting models cannot reason about that negative information, it can only avoid it. The resulting models are also secure from side channel attacks because unused silos are kept separate from the main model.
[0219] Other example benefits include using a database of silo-tuned LLM parts where each part need only have access to one information silo at a time during training to avoid leakage while enabling efficient learning. The silo-tuned LLM part database (e.g., the LLMparameter datastore 118 of FIG. 1) can be used both in a positive way and a negative way, or a mixture of the two. The database of silo-tuned LLM parts can be used to assemble a userspecific LLM (e.g., the models 104 of FIG. 1). The database of silo-tuned LLM parts can also be used to assemble a user-specific interaction-specific LLM (e.g., the interaction specific LLM 310 of FIG. 3) with a target classification marking and / or level of permission access.
[0220] Other example benefits include, given a classification marking scheme, repeatedly using an LLM to make a hypothetical secret entity (e.g., a national government entity) in a box, which can then be used to select, train, and tune the specific combination of fine-tuning methods needed for an LLM to reason about that marking scheme. The specific combination of silo-tuned LLM parts and techniques can be learned based on this hypothetical secret entity in a box. The combination of silo-tuned LLM parts can reflect the known structure and relationships between information silos.
[0221] Yet other example benefits include the resulting LLMs described herein suggesting a mathematically sound lowest marking for an interaction, although users may confirm this. Generated LLMs as described herein can automatically provide a mathematically sound lowest classification for an answer that they provide. The techniques developed by the inventors can accommodate complex classification marking schemes, like that which defense agencies of national governments may use, even when no partial order exists between classification markings and, at times, swapping algorithms depending on the specific mathematical structure of the marking scheme.
[0222] Additionally, the techniques developed by the inventors are applicable to other types of models or language models, assuming the parameters of a model are compressed into a small region that can be added or removed. This may be equivalent to attenuating parts of the model, some of at an extreme level, for example by taking a linear combination which happens to have a zero weight in it.
[0223] Techniques operating according to the principles described herein may be implemented in any suitable manner. The processing and decision blocks of the flowcharts 400, 500, 600, 700 above represent steps and acts that may be included in algorithms that carry out these various processes. Algorithms derived from these processes may be implemented as software integrated with and directing the operation of one or more single- or multi-purpose processors, may be implemented as functionally equivalent circuits such as a DSP circuit or an ASIC, or may be implemented in any other suitable manner. It should be appreciated that the flowcharts 400, 500, 600, 700 included herein do not depict the syntax or operation of any particular circuit or of any particular programming language or type ofprogramming language. Rather, the flowcharts 400, 500, 600, 700 illustrate the functional information one skilled in the art may use to fabricate circuits or to implement computer software algorithms to perform the processing of a particular apparatus carrying out the types of techniques described herein. For example, the flowcharts 400, 500, 600, 700, or portion(s) thereof, may be implemented by hardware alone (e.g., one or more analog or digital circuits, one or more hardware-implemented state machines, etc., and / or any combination(s) thereof) that is configured or structured to carry out the various processes of the flowcharts 400, 500, 600, 700. In some examples, the flowcharts 400, 500, 600, 700, or portion(s) thereof, may be implemented by machine-executable instructions (e.g., machine-readable instructions, computer-readable instructions, computer-executable instructions, etc.) that, when executed by one or more single- or multi-purpose processors, carry out the various processes of the flowcharts 400, 500, 600, 700. It should also be appreciated that, unless otherwise indicated herein, the particular sequence of steps and / or acts described in each flowchart is merely illustrative of the algorithms that may be implemented and can be varied in implementations and embodiments of the principles described herein.
[0224] Accordingly, in some embodiments, the techniques described herein may be embodied in machine-executable instructions implemented as software, including as application software, system software, firmware, middleware, embedded code, or any other suitable type of computer code. Such machine-executable instructions may be generated, written, etc., using any of a number of suitable programming languages and / or programming or scripting tools, and also may be compiled as executable machine language code or intermediate code that is executed on a framework, virtual machine, or container.
[0225] When techniques described herein are embodied as machine-executable instructions, these machine-executable instructions may be implemented in any suitable manner, including as a number of functional facilities, each providing one or more operations to complete execution of algorithms operating according to these techniques. A “functional facility,” however instantiated, is a structural component of a computer system that, when integrated with and executed by one or more computers, causes the one or more computers to perform a specific operational role. A functional facility may be a portion of or an entire software element. For example, a functional facility may be implemented as a function of a process, or as a discrete process, or as any other suitable unit of processing. If techniques described herein are implemented as multiple functional facilities, each functional facility may be implemented in its own way; all need not be implemented the same way. Additionally, these functional facilities may be executed in parallel and / or serially, asappropriate, and may pass information between one another using a shared memory on the computer(s) on which they are executing, using a message passing protocol, or in any other suitable way.
[0226] Generally, functional facilities include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. Typically, the functionality of the functional facilities may be combined or distributed as desired in the systems in which they operate. In some implementations, one or more functional facilities carrying out techniques herein may together form a complete software package. These functional facilities may, in alternative embodiments, be adapted to interact with other, unrelated functional facilities and / or processes, to implement a software program application.
[0227] Some exemplary functional facilities have been described herein for carrying out one or more tasks. It should be appreciated, though, that the functional facilities and division of tasks described is merely illustrative of the type of functional facilities that may implement using the exemplary techniques described herein, and that embodiments are not limited to being implemented in any specific number, division, or type of functional facilities. In some implementations, all functionalities may be implemented in a single functional facility. It should also be appreciated that, in some implementations, some of the functional facilities described herein may be implemented together with or separately from others (e.g., as a single unit or separate units), or some of these functional facilities may not be implemented.
[0228] Machine-executable instructions (e.g., processor-executable instructions) implementing the techniques described herein (when implemented as one or more functional facilities or in any other manner) may, in some embodiments, be encoded on one or more computer-readable media, machine-readable media, etc., to provide functionality to the media. Computer-readable media, machine-readable media, etc., include magnetic media such as a hard disk drive, optical media such as a CD or a DVD, a persistent or non-persistent solid-state memory (e.g., Flash memory, Magnetic RAM, etc.), or any other suitable storage media. Such a computer-readable medium, a machine-readable medium, etc., may be implemented in any suitable manner. As used herein, the terms “computer-readable media” (also called “computer-readable storage media”), “computer-readable medium” (also called “computer-readable storage medium”), “machine-readable media” (also called “machine- readable storage media”), and “machine-readable medium” (also called “machine-readable storage medium”) refer to tangible storage media. Tangible storage media are non -transitory and have at least one physical, structural component. In a “computer-readable medium” and“machine-readable medium” as used herein, at least one physical, structural component has at least one physical property that may be altered in some way during a process of creating the medium with embedded information, a process of recording information thereon, or any other process of encoding the medium with information. For example, a magnetization state of a portion of a physical structure of a computer-readable medium, a machine-readable medium, etc., may be altered during a recording process.
[0229] Further, some techniques described above comprise acts of storing information (e.g., data and / or instructions) in certain ways for use by these techniques. In some implementations of these techniques — such as implementations where the techniques are implemented as machine-executable instructions — the information may be encoded on a computer-readable storage media. Where specific structures are described herein as advantageous formats in which to store this information, these structures may be used to impart a physical organization of the information when encoded on the storage medium. These advantageous structures may then provide functionality to the storage medium by affecting operations of one or more processors interacting with the information; for example, by increasing the efficiency of computer operations performed by the processor(s).
[0230] In some, but not all, implementations in which the techniques may be embodied as machine-executable instructions, these instructions may be executed on one or more suitable computing device(s) and / or electronic device(s) operating in any suitable computer and / or electronic system, or one or more computing devices (or one or more processors of one or more computing devices) and / or one or more electronic devices (or one or more processors of one or more electronic devices) may be programmed to execute the machine-executable instructions. A computing device, electronic device, or processor (e.g., processor circuitry) may be programmed to execute instructions when the instructions are stored in a manner accessible to the computing device, electronic device, or processor, such as in a data store (e.g., an on-chip cache or instruction register, a computer-readable storage medium and / or a machine-readable storage medium accessible via a bus, a computer-readable storage medium and / or a machine-readable storage medium accessible via one or more networks and accessible by the device / processor, etc.). Functional facilities comprising these machineexecutable instructions may be integrated with and direct the operation of a single multipurpose programmable digital computing device, a coordinated system of two or more multipurpose computing device sharing processing power and jointly carrying out the techniques described herein, a single computing device or coordinated system of computing device (colocated or geographically distributed) dedicated to executing the techniques described herein,one or more FPGAs for carrying out the techniques described herein, or any other suitable system.
[0231] Embodiments have been described where the techniques are implemented in circuitry and / or machine-executable instructions. It should be appreciated that some embodiments may be in the form of a method, of which at least one example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0232] Various aspects of the embodiments described above may be used alone, in combination, or in a variety of arrangements not specifically discussed in the embodiments described in the foregoing and is therefore not limited in its application to the details and arrangement of components set forth in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined in any manner with aspects described in other embodiments.
[0233] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both,” of the elements so conjoined, e.g., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, e.g., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B,” when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0234] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0235] As used herein in the specification and in the claims, the phrase, “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identifiedwithin the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently, “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, ,and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0236] Use of ordinal terms such as “first,” “second,” “third,” etc., in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another or the temporal order in which acts of a method are performed, but are used merely as labels to distinguish one claim element having a certain name from another element having a same name (but for use of the ordinal term) to distinguish the claim elements.
[0237] Also, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” “having,” “containing,” “involving,” and variations thereof herein, is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.
[0238] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0239] The word “exemplary” is used herein to mean serving as an example, instance, or illustration. Any embodiment, implementation, process, feature, etc., described herein as exemplary should therefore be understood to be an illustrative example and should not be understood to be a preferred or advantageous example unless otherwise indicated.
[0240] Having thus described several aspects of at least one embodiment, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be part of this disclosure and are intended to be within the spirit and scope of the principles described herein. Accordingly, the foregoing description and drawings are by way of example only.
Claims
CLAIMSWhat Is Claimed Is:
1. A method for producing user specific machine learning (ML) models for execution in restricted access applications, comprising:(A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data comprising at least one of information in at least one datastore of a plurality of datastores or data associated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pre-trained ML model to generate an output associated with the training data;(B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users corresponds to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; and for at least one user of the one or more users: identifying, by using the one or more data associations, one or more of the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters.
2. The method of claim 1, wherein the at least one pre-trained ML model is at least one pre-trained generative ML model.
3. The method of claim 2, wherein the at least one pre-trained generative ML model is at least one large language model.
4. The method of any one of claims 1-3, wherein the information in the at least one datastore at least one of (i) comprises Sensitive Compartmented Information, (ii) is associatedwith one or more Special Access Programs, or (iii) comprises sensitive or restricted access information.
5. The method of any one of claims 1-3, wherein the set of ML parameters comprises one or more ML modules respectively comprising one or more ML weight values.
6. The method of any one of claims 1-3, wherein the at least one pre-trained ML model comprises one or more first ML modules respectively comprising one or more first ML weight values, training the at least one pre-trained ML model using the training data comprises generating one or more second ML modules respectively comprising one or more second ML weight values, and the set of ML parameters comprises the one or more second ML modules and the one or more second ML weight values.
7. The method of claim 6, wherein the one or more first ML modules are different from the one or more second ML modules.
8. The method of claim 6, wherein training the at least one pre-trained ML model using the training data comprises changing a value of at least one of the one or more first ML weight values, and the set of ML parameters comprises the changed value of the at least one of the one or more first ML weight values.
9. The method of claim 6, wherein training the at least one pre-trained ML model using the training data comprises generating one or more ML modules and attaching the one or more ML modules to the at least one pre-trained ML model, and the set of ML parameters comprises the one or more ML modules.
10. The method of claim 6, wherein storing the plurality of sets of ML parameters comprises storing the one or more second ML modules and the one or more second ML weight values in a datastore different from the plurality of datastores.
11. The method of any one of claims 1-3, wherein training the at least one pre-trained ML model comprises instantiating the at least one pre-trained ML model, and performing (A)-(B) among the different sets of training data comprises re-instantiating the at least one pre-trained ML model.
12. The method of any one of claims 1-3, wherein one or more first ones of the plurality of datastores are associated with a first classification of restricted access, second ones of the plurality of datastores are associated with a second classification of restricted access, and the one or more first ones of the plurality of datastores are isolated from the one or more second ones of the plurality of datastores in accordance with the first and second classifications of restricted access.
13. The method of any one of claims 1-3, wherein training the at least one pre-trained ML model using the training data comprises training the at least one pre-trained ML model using a fine tuning technique.
14. The method of claim 13, further comprising determining a type of the fine tuning technique based on the information in the at least one datastore.
15. The method of any one of claims 1-3, wherein generating the trained ML model comprises executing an ML model using the user data and the plurality of sets of ML parameters as inputs to learn combinations of the at least one pre-trained ML model with one or more of the plurality of sets of ML parameters in connection with the user data.
16. The method of any one of claims 1-3, wherein identifying, by using the one or more data associations, the one or more of the plurality of sets of ML parameters that correspond to the at least one user comprises: determining a first permission level at which the at least one user is permitted to access data; identifying information in one or more of the plurality of datastores being associated with a second permission level, the second permission level being the same or lower than the first permission level; and identifying the one or more of the plurality of sets of ML parameters that correspond to the one or more of the plurality of datastores being associated with the second permission level.
17. The method of any one of claims 1-3, wherein the set of ML parameters comprises one or more first ML parameters representing information for use in generating the outputand one or more second ML parameters representing information not to be used for generating the output.
18. The method of any one of claims 1-3, further comprising executing the trained ML model, using an input by the at least one user, to generate an output from the trained ML model.
19. The method of claim 18, wherein generating the output from the trained ML model comprises obtaining data for generating the output by calling the at least one API.
20. The method of claim 18, further comprising determining a likelihood that the output from the trained ML model conforms to one or more permission levels.
21. The method of any one of claims 1-3, further comprising: processing input to identify one or more of the plurality of datastores to which the input is associated; and determining a likelihood that one or more permission levels associated with the identified one or more of the plurality of datastores conforms to the permission level of at least one user.
22. The method of claim 21, wherein processing the input comprises: tokenizing the input to generate tokenized text; determining a logit vector using the tokenized text; processing the logit vector using ones of the plurality of sets of ML parameters to generate a respective perplexity score; and identifying, using the perplexity scores, which of the plurality of sets of ML parameters is associated with the input.
23. The method of claim 21, wherein the input is plain text.
24. The method of any one of claims 1-3, wherein the one or more users comprise at least a first user associated with a first trained ML model and a second user associated with a second trained ML model, and the method further comprising effectuating an electronicinteraction between the first user and the second user by executing the first trained ML model for the first user and the second trained ML model for the second user.
25. The method of claim 24, wherein the first user and the second user have a first permission level, and the method further comprising receiving input from at least one of the first user or the second user to set a second permission level for the electronic interaction, the second permission level lower than the first permission level.
26. The method of claim 24, further comprising generating an alert representing a change in a permission level of the electronic interaction based on information in the electronic interaction, a first permission level of the first user, a second permission level of the second user, and a target permission level for the electronic interaction.
27. The method of claim 24, further comprising: generating an output with the first trained ML model; and providing the output to the second user after determining that the output corresponds to an agreed upon permission level of the electronic interaction by the first user and the second user.
28. The method of claim 27, further comprising generating an alert after determining that the output does not conform to the permission level of the second user.
29. The method of claim 27, wherein the output is a first output, and further comprising generating a second output that conforms to the permission level of the second user after determining that the first output does not conform to the permission level.
30. An apparatus for producing user specific machine learning (ML) models for execution in restricted access applications, comprising: at least one memory storing processor executable instructions; and at least one hardware processor configured to execute the processor executable instructions to perform a method comprising:(A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data comprising at least one of information in at least one datastore of a plurality of datastores or dataassociated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pretrained ML model to generate an output associated with the training data;(B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users corresponds to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; and for at least one user of the one or more users: identifying, by using the one or more data associations, one or more of the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters.
31. At least one non-transitory computer-readable storage medium comprising processor executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform a method comprising:(A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data comprising at least one of information in at least one datastore of a plurality of datastores or data associated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pre-trained ML model to generate an output associated with the training data;(B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users corresponds to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; andfor at least one user of the one or more users: identifying, by using the one or more data associations, one or more of the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained ML model and the identified one or more of the plurality of sets of ML parameters.
32. A system for producing user specific machine learning (ML) models for execution in restricted access applications, comprising: at least one hardware processor; and at least one computer-readable storage medium storing processor executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method comprising:(A) training at least one pre-trained ML model using training data to generate a set of ML parameters, the training data comprising at least one of information in at least one datastore of a plurality of datastores or data associated with at least one application programming interface (API), the set of ML parameters for configuring one or more portions of the at least one pretrained ML model to generate an output associated with the training data;(B) generating one or more data associations of one or more users and the set of ML parameters after determining that user data of the one or more users corresponds to one or more portions of the training data; performing (A)-(B) among different sets of training data to generate a plurality of sets of ML parameters, each of the plurality of sets of ML parameters corresponding to a different one of the different sets of training data; and for at least one user of the one or more users: identifying, by using the one or more data associations, one or more of the plurality of sets of ML parameters that correspond to the at least one user; and generating a trained ML model for the at least one user based on a combination of the at least one pre-trained MLmodel and the identified one or more of the plurality of sets of ML parameters.
Citation Information
Patent Citations
US202363616440P