Edge prompt personalization in large language models
The PEM on edge devices personalizes LLM queries using user profiles, optimizing query reformulation with MLP embeddings, addressing the challenge of uniform responses from LLMs and enhancing query accuracy without extensive user effort or resource consumption.
Patent Information
- Application Number
- PCT/GR2024/000002
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-17
AI Technical Summary
Existing large language models (LLMs) provide similar responses to all users, disregarding individual preferences and contexts, requiring users to meticulously craft queries for desired responses, which is time-consuming and labor-intensive.
A Personalized Edge Model (PEM) is deployed on edge devices to reformulate user queries based on user profiles, including demographic and interest data, using a Multi-Layer Perceptron (MLP) with derivative-free optimizations to generate prompt and data embeddings, optimizing query personalization without modifying LLM parameters.
Provides higher accuracy responses from LLMs by personalizing queries based on user profiles, reducing the need for extensive user input and resource-intensive training, while maintaining model agnosticism and minimizing training effort.
Smart Images

Figure GR2024000002_17072025_PF_FP_ABST
Abstract
Description
EDGE PROMPT PERSONALIZATION IN LARGE LANGUAGE MODELSTechnical Field
[0001] This disclosure relates io systems, methods, and computer-readable media for edge personalization of large language model (LLM) queries.Background
[0002] Various computing networks that implement machine learning, deep learning, and / or artificial intelligence (Al) can include a series of interconnected computing devices (e.g., cloud computers) and a number o f edge devices capable of performing a series of processing tasks. The edge devices in such computing networks can obtain data and provide the data to the cloud computers for processing.
[0003] In some cases, the edge devices can perform processing as part of a machine learning or deep learning system. For example, an edge device can implement processing as part o f a deep neural network (DNN) for any of a variety of use cases, such as image recognition, portrait mode photography, text prediction, user profiling, de-noising, camera enhancement, activity recognition, etc.
[0004] Various techniques can be implemented on the edge of a computing network to assist in processing queries for a large language model (LLM), This can include an edge device being configured to provide queries of a user to the LLM at the edge. Configuring an edge device to interact with the LLM can improve latency and privacy, as data may not need to be transmitted between the edge device and a remote cloud computing infrastructure.Summary
[0005] The present embodiments relate to systems, methods, and computer-readable media for edge prompt personalization for LLMs, Particularly, the present embodiments can allow for reformatting of queries / prompts to LLMs with the consideration of unique contextual information and preferences unique to each user and edge device, A PersonalizedEdge Model (PEM) can be deployed on an edge device such as a smartphone or web plugin, for example. The PEM can be designed to proactively reformulate user questions based on a user profile, including factors such as age, gender, interests, email communications, etc. This personalization can provide higher accuracy responses from the LLM based on the user profile specific to the user that provided the initial query.
[0006] to a first example embodiment, a method for personalizing large language mode! (LLM) queries using user-specific profile information at an edge device. The embodiments can include obtaining, by a personalized edge model (PEM) implemented on an edge device, a set of user profile information specific to a user and associated with the edge device.
[0007] The embodiments can also include training the PEM. Training the PEM can include generating one or more query embeddings for a test query using both input data and the set of user profile information. Training the PEM can also include transmitting the test query with the one or more query embeddings to a huge language model (LLM) implemented on one or more remote computing devices. Training the PEM can also include receiving, at the PEM, a response to the test query from the LLM. The response to the test query can be compared to a ground truth to determine an accuracy of the received response io the test query.
[0008] The embodiments can also include obtaining, at the PEM, an initial query. The embodiments can also include generating, by the PEM, a prompt embedding and / or a data embedding based on the set of user profi le information specific to the user to generate a persona li zed query' .
[0009] The embodiments can also include transmitting, by the PEM, the personalized query that includes the prompt embedding and / or the data embedding to the LLM. The embodiments can also include receiving, at the PEM, a response to the personalized query from the LLM.
[0010] In some instances, the set of user profi le information comprises data obtained at the edge de vice that inc ludes any of demographic information relating to the user, interests of the user, and tracked browser data specific to the user. In some instances, the training further comprises determining ah accuracy of the response to the test query based on a similarity between the response to the test query and the ground truth, updating the one or more embeddings for the test query based on the determined accuracy of the response, transm itting the updated embeddings to the LLM, and receiving an updated response from the LLM-
[0011] In some instances, the prompt embedding modifies text of the initial query based on the set of user profi le information, and the data embedding transforms the initial query based on the set of user profile information.
[0012] In some instances, the PEM comprises a Multi-Layer Perceptron (MLP) including a three-layer Fully Connected Network (FCN).
[0013] In some instances, the prompt embedding and / or the data embeddings modify the initial query without providing any sensitive user data to the LLM,
[0014] In another example embodiment, a system is provided. The system can include one or more computing devices implementing a large language model (LLM) and an edge computing device associated with a user, the edge device implementing a personalized edge model (PEM).
[0015] The edge de vice can be configured to obtain an initial query, generate any of a prompt embedding and a data embedding based on a set of user profile information, transmit any of the prompt embedding and the data embeddi ng to the LL M , and receive a response to the initial query from the LLM,
[0016] In some instances, the edge device is further configured to train the PEM by generating one or more query embeddings for a test query using both an input query and the set of user profile information, transmitting the test query with the one or more query embeddings to a large language model (LLM) implemented on one or more remote computing devices and receiving, at the PEM, a response to the test query from the LLM, The response to the test query can be compared to a ground truth response to determine an accuracy of the response to the test query.
[0017] In some instances, the training further comprises determining an accuracy of the response to the test query based on a similarity between the response to the test query and the ground truth, updating the one or more embeddings for the test query based on the determined accuracy of the response, transmitting the updated embeddings to the LLM, and receiving an updated response from the LLM,
[0018] In some instances, the prompt embedding modifies text of the initial query' based on the set of user profile information, and wherein the data embedding transforms the initial query based on the set of user profi le in format! on.
[0019] In some instances, the PEM comprises a Multi-Layer Perception (MLP) including a three-layer Fully Connected Network (FCN) configured to extract textual features in the initial query.
[0020] In some instances, the PEM further comprises an optimizer configured to generate tire prompt embeddings and / or the data embeddings,
[0021] In some instances, the set of user profi le information comprises data obtained at the edge device that includes any of demographic information relating to the user, interests ofthe user, and tracked browser data specific to the user, and the set of user profile information is stored as part of a database or table accessible to the edge device.
[0022] In another example embodiment, a computer-implemented method is provided. The embodiments can include obtaining, by a personalized edge model (PEM) implemented on an edge device, a set of user profile information specific to a user and associated with the edge device,
[0023] T he embodiments can also include training the PEM. Training the PEM can include generating one or more query embeddings for a test query using both input data and the set of user profile information. Training the PEM can also include transmitting the one or more query' embeddings to a large language model (LLM) implemented on one or more remote computing devices. Training the PEM can also include receiving, at the PEM, a response to the test query from the LLM. The response to the test query can be compared to a ground truth to determine an accuracy of the received response to the test query. Training the PEM can also inc lude determining an accuracy of the response to the test query based on a similarity between the response to the test query and the ground truth. Training the PEM can also include updating the one or more embeddings for the test query based on the determined accuracy of the response. Training the PEM can also include transmitting the test query and the updated embeddings to the LLM.
[0024] The embodiments can also include receiving an updated response from the LLM. The embodiments can also include obtaining, at the PEM, an initial query. The embodiments can also include generating, by the PEM, a prompt embedding and / or a data embedding based on the set of user profile inform ation specific to the user. The embodiments can also include transmitting, by the PEM, any of the prompt embedding and the data embedding to the LLM. The embodiments can also include receiving, at the PEM, a response to the initial query from the LLM.
[0025] In some instances, the PEM comprises a Multi-Layer Perceptron (MLP) including a three-layer Fully Connected Network (FCN) configured to extract textual features in the initial query.
[0026] In some instances, the PEM further comprises an optimizer configured to generate the prompt embeddings and / or the data embeddings.
[0027] In some instances, the set of user profile information comprises data obtained at the edge device that includes any of demographic information relating to the user, interests of the user, and tracked browser data specific to the user.
[0028] In some instances, the set of user profile information is stored as part of a database or table accessible to the edge device.
[0029] In some instances, the prompt embedding modifi es text of the initial query based on the set of user profile information, and the data embedding comprises a transformation of the initial query based on the set of user profile information.
[0030] In some instances, the prompt embedding and / or the data embeddings modify the initial query without providing any sensitive user data to the LLM.
[0031] This Summary' is provided to summarize some example embodiments, so as to provide a basic understanding of some aspects of the subject matter described in this document. Accordingly, it will be appreciated that the features described in this Summary are merely examples and should not be construed to narrow the scope or spirit of the subject matter described herein in any way. Unless otherwise stated, features described in the context of one example may be combined or used with features described in the context of one or more other examples. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following Detailed Description, Figures, and Claims.Bri ef Description of the Drawings
[0032] The above and other aspects of the disc losure, its nature, and various features will become more apparent upon consideration of the following detailed description, taken in conjunction with the accompanying drawings, in which like reference characters may refer to like parts throughout, and in which:
[0033] FIG, 1 is an example computing network including a series of interconnected computing devices and edge devices according to an embodiment.
[0034] FIG. 2 illustrates example queries to an LLM with no personalization according to an embodiment.10035] FIG. 3 illustrates an example set of responses to queries by an LLM with user personalization according to an embodiment.
[0036] FIG . 4 illustrates a flow process of an example method for personalizing large language model (LLM) queries using user-specific profile information at an edge device according to an embodiment.
[0037] FIG. 5A illustrates an example process for training a PEM according to an embodiment.
[0038] FIG. 5:B illustrates an example process for PEM edge inference according to an embodiment.
[0039] FIG. 6 illustrates an example process for training a PEM on the edge according to an embodiment.
[0040] FIG. 7 illustrates an example process for training a PEM on the edge according to an embodiment.
[0041] FIG. 8 is a block diagram of a special-purpose computer system according to an embodiment. Detailed Description
[0042] A large language model (LLM) is a type of machine learning model that can perform a variety of natural language processing (NLP) tasks such as generating and classifying text, answering questions in a conversational manner, and translating text from one language to another, for example. LLMs can have a number of values (parameters) that can change autonomously as the model learns.
[0043] LLMs have demonstrated robust capabilities across a range oflearning tasks, at least partially due to their capacity to process vast amounts of data. As a result, end-users operating edge devices are increasingly seeking to leverage the power of LLMs for a variety of personalized applications.
[0044] However, in practical applications, when posed with a question, LLMs often provide similar responses to all users, disregarding individual preferences and contexts, This may force users to meticulously craft their queries to elicit desired responses, which can be time-consuming and labor-intensive.
[0045] The present embodiments relate to edge prompt personalization for LLMs. The embodiments can allow for reformatting of queries / prompts to LLMs with the consideration of unique contextual information and preferences unique to each user and edge device. A Personalized Edge Model (PEM) can be deployed on an edge device such as a smartphone or web plugin, for example. The PEM can be designed to proactively reformulate user questions based on a user profile, including factors such as age. gender, interests, email communications, etc. For instance, the PEM can generate prompt embeddings and / or data embeddings that modify or append data to an initial query to the LLM that personalizes the query with data specific to the user. This personalization can provide higher accuracyresponses from the LLM based on the user profile specific to the user that provided the initial query.
[0046] The edge prompt personalization as described herein can be designed to be trained exclusively on user data, without any need to modify the parameters in the LLMs. Such edge personalization can provide several advantages, such as being model agnostic (being applicable to any LLM), providing high quality results from any LLM, providing high degrees of personalization for the query being learned and personalized, and / or minimal training effort needed for the LLM, as only the PEM is trained on edge data to significantly reduce the time and resources required for model implementation.
[0047] Further, various issues specific to LLMs are addressed by the methods and systems as described herein. For instance, LLMs generally can be very resourcefintensive in order to be loaded and executed in edge devices. Further, LLMs trained from general data may usually provide one-fit-all answers, other than personalized answers related to the user. Also, LLMs tend to provide an ambiguous answer when a question is not well formulated by users, i.e., an improper query / prompt.
[0048] LLMs are generally implemented on one or more interconnected computing devices, such as a cloud computing network. Edge devices can include computing devices associated with a user. FIG. 1 is an example computing network 100 including a series of interconnected computing devices 102 and edge devices 104, 106. As shown in FIG. 1, the series of interconnected computing devices 102 (or cloud computing network) can provide computing resources implementing a portion or the entirety of a LLM 108 as described herein. Further, edge devices can include devices associated with an end user (e.g., smartphone 104, computer 106). The edge devices 104, 106 can perform processing on the edge of the network, with each device 104, 106 implementing a PEM 1 10 as described herein.
[0049] The methodology of the present embodiments can include specifying the LLM as a service with edge personalization. In many cases, users interacting with other LLMs with little or no personalization can provide low quality results to queries. FIG. 2 is an illustration 200 of example queries to an LLM 206 with no personalization.
[0050] In FIG. 2, both a first user (computer science student 202) and a second user(e.g., zoologist 204) can each provide a query 208A-B of “What is python” to the LLM 206. However, the context in which each user 202, 204 is providing the query may differ. For example, the computer science student 202 can be providing a query relating to a softwarelanguage, while zoologist 204 can be providing the same query in the context of a snake. Without personality, the results provided by the LLM may not take such context into account, leading to responses that may be undesired and requiring detailed query drafting by each user 202, 204.
[0051] Various network architectures can be implemented when running language models oft edge devices. For instance, language models can be trained on the edge device. This can include the model being trained from scratch on one or more edge devices. The model and data can be preserved in the edge devices. However, the model performance can be constrained by the resources of the edge device, which can result in a lower model quality but higher personalization.
[0052] Another example network architecture can include LLMs as a cloud service. Such an architecture can have edge devices sending / receiving queries and results by communicating with LLMs on the cloud. However, such an architecture can have a greater model quality but little to no personalization for users.
[0053] Yet another architecture can include language models pre-trained on the cloud.This architecture can include pretraining a language model on the cloud, then transferring the model to the edge for fine-tuning. Such architectures can have a medium model quality and some personalization for users.
[0054] Various solutions can be used for implementing language models on the edge that consider user personalization to varying degrees. For instance, natural language processing (NLP) models can generally perform basic language processing tasks. Further, a hardware-friendly transformer-based architecture model can be adapted to provide language- based results to queries.
[0055] Some techniques can implement various amounts of user personalization. For instance, running language models on edge devices can provide some personalization at the cost of model quality due to resource constraints at the edge. Further, language models can be trained on the edge, such as hardware-aware transformers. Other techniques can also be adopted, such as using LLMs as a cloud service with cloud-edge communications, and language models pre-trained on the cloud with edge transfer. LLMs with edge personalization as described in the present embodiments can implement language models with user personalization that maximizes model quality and user personalization.
[0056] The present embodiments provide systems and methods to automatically personalize their prompts / queries to LLM's, without much additional effort from users to formulate the queries from users to get desired results.
[0057] In many instances, raw data may not be transferred from the edge to the cloud, and the system may be less strict for pulling information from the cloud to the edge. A Personalized Edge Model (PEM) can be deployed on the Edge device to proactively reformulate user queries based on the user profile (e,g„ a profile based on factors such as age, gender, interests, communications, etc.).
[0058] The personalized edge model can include a backbone model, such as a Multi- Layer Perceptron (MLP) including a three-layer fully connected network (FCN). The PEM can include an optimizer (e.g., derivative-free optimization, Bayesian Optimization, evolutionary algorithms). In some instances, the gradients may not be available as the loss may not be backpropagated to the LLMs, which can change the parameters in LLMs. The PEM can use derivative-free optimizations. More precisely, the optimizer can be implemented by Evolutionary Optimization and optimize Multi-Layer Perceptron (MLP) to produce appropriate embeddings.
[0059] The Personalized Edge Model (PEM) can be iteratively optimized via derivative-free optimizations, such as Evolutionary Optimization. The parameters in the PEM can be continuously updated unti l reaching the determined accuracy. For instance, the parameters can be represented as a real-valued vector, the PEM can be optimized by continually changing the values in the vector, until the response reaches the determined accuracy.
[0060] The present embodiments can provide results specific to each user based on a user profile. FIG. 3 is an illustration 300 of an example set of responses to queries by an LLM 310 with user personalization. As shown in FIG. 3, a first user 302 and a second user 304 can each provide an identical query 306A-B. Further, a PEM 308A-B on edge device on the edge 316 can provide user-specific profile information that can be used to curate the query to the LLM 310 operating in the cloud 312 and provide results 314A-B unique and specific to each user. For instance, if the first user 302 is a computer science student querying “What is python,’’ the user profile in PEM 308A can specify that the query is likely in a computer science context, and the result from LLM 310 can be based on a programming language, not the animal.
[0061] The edge device with the PEM as described herein can be trained using a training process. The training process can include the PEM taking profile data for the user and locally training data as inputs and building a Query Embedding to send to the LLMs on the cloud or otherwise implemented by one or more remote computing devices. The PEM can be optimized by calculating the loss with LLMLs answer P and local ground truth F. The PEM parameters can be updated continuously io minimize loss. This process can include using derivative-free optimizations. For instance, Evolutionary Optimization can include retaining parameters effective in loss reduction while updating less useful parameters. This dynamic adjustment can continue until the parameters effectively minimize the loss, indicating an optimized PEM.
[0062] Further, edge inference can be implemented for end risers. The PEM can re- formulate queries provided by a user using the user profile. The PEM can build a personalized query embedding and send the embedding(s) to the LLMs on the cloud. The LLMs can return a personalized answer to the user on the edge device.
[0063] FIG. 4 illustrates a flow process 400 of an example method for personalizing large language model (LLM) queries using user-specific profile information at an edge device. At 402, a PEM implemented on an edge device can obtain a set of user profile information specific to a user and associated with the edge device. The set of user profile information can include data obtained at the edge device that includes any of demographic information relating to the user, interests of the user, and tracked browser data specific to the user.
[0064] User profile data can be retrieved from various data sources and stored in different ways, such as via a database, table, file- structure, etc. An example database can store the basic demographic information of the user in a database table specifying parameters such as age, job, gender, country, city, civil status, etc. A file structure can save the tracked browser data, historical conversation records with Cloud LLMs, etc. The PEM can further include a multi-layer network, such as a Multi-Layer Perceptron (MLP) including a three- layer Fully Connected Network (FCN).
[0065] Al 404, the PEM can be trained. Training the PEM can include generating one or more query embeddings for a test query using both input data and the set of user profil e information (e.g., at 406). The FEM can initially create query embeddings based on user profile information to personalize the query as part of training foe PEM.
[0066] At 408, training the PEM can also include transmitting the test query with the one or more query embeddings to a large language model (LLM) implemented on one or more remote computing devices. The LLM can process the test query and query embeddings to provide a response to the query that incorporates the information included in the query embeddings to personalize the response based on the user profile information.
[0067] At 410, training the PEM can include receiving a response to the test query from the LLM. The response can be in the form of a text-based response or include other forms of media (e.g., images, video, audio, etc.). The response to the test query can be compared to a ground truth to determine the accuracy of the recei ved response to the test query.10068] In some instances, the training can include determining an accuracy of the response to the test query based on a similarity between the response to the test query and the ground truth. This can include identifying a number of similar elements between the response and the ground truth. Further, the one or more embeddings for the test query can be updated based on the determined accuracy of the response.
[0069] For instance, the Personalized Edge Model (PEM) can be iteratively optimized via derivative-free optimizations, e.g.. Evolutionary’ Optimization. The parameters in the PEM can be continuously updated until reaching the determined accuracy. For instance, the parameters can be represented as a real -valued vector, the PEM can be optimized via changing the values in the vector, until the response reaches the determined accuracy.
[0070] Training the PEM can also include transmitting the test query and the updated embeddings to the LLM and receiving an updated response from the LLM. This process can be performed over a number of iterations based on a determined accuracy of the responses provided by the LLM.
[0071] At 412, the PEM can obtain an initial query'. The initial query can include an Input provided by a user of the edge device, such as a query of “What is py thon?’' provided by the user.
[0072] At 414, the PEM can generate a prompt embedding and / or a data embedding based on the set of user profile information specific to the user. The prompt embedding can modify text of the initial query' based on the set of user profile information. The data embedding can append additional text to the initi al query based on the set of user profile information. Tire embeddings can be produced using a Multi-Layer Perceptron (MLP) and cun be represented as a real-valued vector.(0073] In some instances, the prompt embedding and / or the data embeddings can modify the initial query without providing any sensitive user data to the I.L.M. The user can select which types of private information can be included in building the embeddings.Therefore, the generated embeddings can contain only the information that users want to provide.
[0074] At 416, the PEM can transmit the personalized query that includes the prompt embedding and the data embedding to the ELM. The initial query car, be passed to PEM, and the PEM can optimize the initial input query to a personalized query and only transfer tire personalized query' to the ELM.
[0075] At 418, the PEM can receive a response to the personalized query from theELM. The ELM can generate a response to the initial query that takes the embeddings into account.(0076] As described herein, the PEM can be trained at the edge device using user profile information. FIG. 5 A illustrates an example process 500A for training a PEM. A PEM 508 on the edge 516 can obtain a set of user profile information 506 that can include various information types, such as a user’s age, job, interests, etc. The PEM 508 can generate a personalized query embedding 510 based on the user information 506.
[0077] The PEM 508 can transform a test query to a personalized query using the personalized query embedding 510 that can be transmitted to the LEM 512 on the cloud 518. the ELM 512 can provide a response 514 based on the personalized query that can be fed to the PEM 508. The PEM 508 can compare the response 514 with a ground truth 502 to determine an accuracy of the response as part of an optimizing process 504. The PEM 508 can modify subsequent queries to improve the accuracy of the responses from the ELM 512 over a number of iterations.
[0078] The PEM 508 can also provide an inference process on the edge. FIG. 5B illustrates an example process 500B for PEM edge inference. As shown in FIG. 5B, the PEM 508 can obtain user profile information 506 and generate personalized query embeddings 510 specific io the user. Example embeddings 510 can include a modification to the query that incorporates user profile data, such as to edit a query “What is Python?” to “What is Python in computer programming?” based on user profile information specifying the user as a computer programmer. The embeddings 510 can also include additional data sent with the query, such as additional text specifying that the user is a software programmer, for example.
[0079] The LLM 512 can obtain the personalized query and provide a response 514. The response 514 can be based on the query as well as the embeddings 510 provided by the PEM 508.
[0080] In some instances, the present embodiments can implement personalized prompt learning on the Edge. This can include generating optimal prompt and data embeddings on the edge considering the user profile. The query can be optimized to produce optimized predictions on local data for the user.
[0081] The PEM can obtain multiple sources of user profi le d ata and generate different embeddings based on the user data. FIG. 6 illustrates an example process 600 for training a PEM on the edge. For example, the PEM 606 can obtain user profile data 604A, user query data 604B of varying data types from any of a database, table, file directory, etc. For example, parameters for the user data 604A-B can include user age, job, interests, stored user data on the edge device, prior queries, web browser data, etc. The value “P” can include the user’s profile data chai-acterizing the user’s background or personal behaviors. The valueiUX” can include a set of initial queries (e.g., questions) for training the PEM, while value “YMcan include the weighted ground-truth answers for the queries.
[0082] The PEM 606 can process the user profile data 604A and user query 604B and generate embeddings 610, 612 on the edge 614. As noted above, tire prompt embedding 610 and data embeddings 612 can be generated as a function of the user-speci fic profile data and the query (e.g., 608). The prompt embedding 610 can the modify text of the initial query based on the set of user profile information, and the data embedding 612 can transform the initial query based on the set of user profile information. Further, an optimization process 622 can improve PEM 606 performance as described herein.
[0083] Given user profile data P and user query data X, the PEM fwcan generate the personal query fw(P, X), which can be passed to the LLM F to generate the answer F(fw(P,X)). The PEM can find the optimal parameters w in the parameter space W, which can minimize the loss value between the LLM answer F(fr(P, X)) and ground truth Y
[0084] Further, given a user profile data P and user query data X, the PEM can be considered as a query-rewriting function fw, which can generate the optimized Query Q.
[0085] The embeddings 610, 612 can be fed to the LLM 618 at the cloud 616. In response, the LLM 618 can provide a response 620 to the personalized query. The PEM 606 can further optimize personalized queries by comparing the response 620 with a ground truth response 602 using an optimization process as described herein.
[0086] In another example, the present embodiments can implement personalized edge in-context learning. In-context learning on the edge can be personalized by the user profile and target learning tasks. Further, optimal demonstrations can be selected based on the edge, which can be supervised or unsupervised. Further, optimal instructions (queries) can be used for in-context learning.
[0087] FIG\ 7 illustrates an example process 700 for training a PEM on the edge.
[0088] The user profile data in FIG. 7 can be similar or the same as that in FIG. 6. X is user query data (e.g.. questions), Y can include the ground truth answers of the queries. The values {xi, yd, i-0, .... n-1 can be a set of demonstration instances, which can be transferred to the PEM along with the user query Xn. while the ground truth answer yncan be conserved for calculating the loss.
[0089] In FIG. 7, the PEM 706 can obtain a number of demonstration instances 724, 726 and user profile data 704 either separately or as pairs of parameters. For instance, demonstration instances 724, 726 can be obtained at the PEM 706 from a database, table, or file structure storing various demonstration instances 724, 726.
[0090] The PEM 706 can perform both a demo selection 708 and instruction (query) formatting processes 710. The demo selection process can include extracting features from an input demonstration instance to identify its characteristics. This can include a natural language processing process to identify features in the demonstration instance. The instruction formatting process 710 can include formatting the query and generating embeddings (using query function 712) as described herein. For example, the instruction formatting process 710 can generate a prompt embedding that modifies the input query and a data embedding that transforms the query based on the user profile information.
[0091] The personalized query can be sent from the edge 714 to the LLM 718 at the cloud 716. The .LLM 718 can provide a response 720 that can be compared with a ground truth 702 and optimized as part of an optimization process 722,Computing System Overview
[0092] FIG. 8 is a block diagram of a special-purpose computer system 800 according to an embodiment. The computer-implemented methods and processes described herein maysimilarly be implemented by tangible, non-transitory computer readable storage mediums and / or computer-program products that direct a computer system to perform the actions of the methods and processes described herein. Each such computer-program product may comprise sets of instructions (e.g.. codes) embodied on a computer-readable medium thatdirects the processor of a computer system to perform corresponding operations. The instructions may be configured to ran in sequential order, or in paralie! (such as under different processing threads), or in a combination thereof.
[0093] Special-purpose computer system 800 comprises a computer 802, a monitor 804 coupled to computer 802, one or more additional user output devices 806 (optional) coupled io computer 802, one or more user input devices 808 (e.g., keyboard, mouse, track ball, touch screen) coupled to computer 802, an optional communications interface 810 coupled to computer 802, and a computer-program product including a tangible computer- readable storage medium 812 in or accessible to computer 802. Instructions stored on computer-readable storage medium 812 may direct system 800 to perform the methods and processes described herein. Computer 802 may include one or more processors 814 that communicate with a number of peripheral devices via a bus subsystem 816. These peripheral devices may include user output device(s) 806, user input device(s) 808, communications interface 810, and a storage subsystem, such as random-access memory (RAM) 818 and non- volatile storage drive 820 (e.g., disk drive, optical drive, solid state drive), which are forms of tangible computer-readable memory.
[0094] Computer-readable medium 812 may be loaded info random access memory 818, stored in non-volatile storage drive 820, or otherwise accessible to one or more components of computer 802. Each processor 814 may comprise a microprocessor, such as a microprocessor from Intel® or Advanced Micro Devices, Inc.®, or the like. To support computer-readable medium 812, the computer 802 runs an operating system that handles die communications between computer-readable medium 812 and the above-noted components, as well as the communications between the above -noted components in support of the computer-readable medium 812. Exemplary operating systems include Windows® or the like from Microsoft Corporation, Solaris® from Sun Microsystems, LINUX, UNIX, and tire like. In many embodiments and as described herein, the computer-program product may be an apparatus (e.g., a hard drive including case, read / write head, etc., a computer disc including case, a memory card including connector, case, etc.) that includes a computer-readable medium (e.g., a disk, a memory chip, etc.). In other embodiments, a computer-program product may comprise the instruction sets, or code modules, themselves, and be embodied on a computer-readable medium.
[0095] User input devices 808 include all possible types of devices and mechanisms to input information to computer system 802, These may include a keyboard, a keypad, amouse, a scanner, a digital drawing pad, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and other types of input devices. In various embodiments, user input devices 808 are typically embodied as acomputer mouse, a trackball, a track pad, a joystick, wireless remote, a drawing tablet, a voice command system. User input devices 808 typically allow a user to select objects, icons, text and the like that appear on the monitor 804 via a command such as a click of a button or the like. User output devices 806 include all possible types of devices and mechanisms to output information from computer 802. These may include a display (e.g., monitor 804), printers, non- visual displays such as audio output devices, etc.
[0096] Communications interface 810 provides an interface to other communication networks and devices and may serve as an interface to receive data from and transmit data to other systems, WANs and / or the Internet, via a wired or wireless communication network 822, In addition, communications interface 810 can include an underwater radio for transmitting and receiving data in an underwater network. Embodiments of communications interface 810 typically include an Ethernet card, a modem (telephone, satellite, cable, ISDN), a (asynchronous) digi tal subscriber line (DSL) unit, a FireWire® interface, a USB® interface, a wireless network adapter, and the like. For example, communications interface 810 may be coupled to a computer network, to a Fire Wire® bus, or the like. In other embodiments, communications interlace 810 may be physically integrated on the motherboard of computer 802, and / or may be a software program, or the like,
[0097] RAM 818 and non-volatile storage drive 820 are examples of tangible computer-readable media configured to store data such as computer-program product embodiments of the present invention, including execu table computer code, human-readable code, or the like. Other types of tangible computer-readable media include floppy disks, removable hard disks, optical storage media such as CD-ROMs, DVDs, bar codes, semiconductor memories such as flash memories, read-only-memories (ROMs), battery- backed volatile memories, networked storage devices, and the like, RAM 818 and non- volatile storage drive 820 may be configured to store the basic programming and data constructs that provide the functionality of various embodiments of the present invention, as described above.
[0098] Software instruction sets that provide the functionality of the present invention may be stored in computer-readable medium 812, RAM 818, and / or non-volatile storage drive 820. These instruction sets or code may be executed by the processors) 814. Computer-readable medium 812, RAM 818, and / or non-volatile storage drive 820 may also provide a repository to store data and data structures used in accordance with the present invention.RAM 818 and non-volatile storage dri ve 820 may include a number of memories including a main random- access memory (RAM) to store instructions and data during program execution and a read-only memory (ROM) in which fixed instructions are stored. RAM 818 and non- volatile storage drive 820 may include a file storage subsystem providing persistent (non- volatile) storage of program and / or data files. RAM 818 and non-volatile storage drive 820 may also include removable storage systems, such as removable flash memory.
[0099] Bus subsystem 816 provides a mechanism to allow the various components and subsystems of computer 802 communicate with each other as intended. Although bus subsystem 816 is shown schematically as a single bus, alternative embodiments of the bus subsystem may utilize multiple busses or communication paths within the computer 802.
[0100] For a firmware and / or software implementation, the methodologies may be implemented with modules (e,g., procedures, functions, and so on) that perform the functions described herein, Any machine-readable medium tangibly embodying instructions may be used in implementing the methodologies described herein. For example, software codes may be stored in a memory'. Memory may be implemented within the processor or external to the processor. As used herein the term “memory” refers to any type of long term, short term, volatile, nonvolatile, or other storage medium and is not to be limited to any particular type of memory’ or number of memories, or type of media upon which memory is stored .
[0101] Moreover, as disclosed herein, the term ‘'storage medium” may represent one or more memories for storing data, including read only memory' (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage mediums, optical storage mediums, flash memory devices andfor other machine-readable mediums for storing information. The term “machine-readable medium” includes but is not limited to portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage mediums capable of storing that contain or carry instruction(s) and / or data.Conclusion
[0102] Whereas many alterations and modifications of the present invention will no doubt become apparent to a person of ordinary skill in the art after having read the foregoing description, it is to be understood that the particular embodiments shown and described by way of illustration are in no way intended to be considered limiting.
[0103] Moreover, the processes described above, as well as any other aspects of the disclosure, may each be implemented by software, but may also be implemented in hardware, firmware, or any combination of software, hardware, and firmware. Instructions for performing these processes may also be embodied as machine- or computer-readable code recorded on a machine- or computer-readable medium. In some embodiments, the computer-readable medium may be a non-transitory computer-readable medium. Examples of such a non-transitory computer-readable medium include but are not limited to a read-only memory, a random-access memory, a flash memory, a CD-ROM, a DVD. a magnetic tape, a removable memory card, and optical data storage devices. In other embodiments, the computer-readable medium may be a transitory computer-readable medium. In such embodiments, the transitory computer-readable medium can be distributed over network-coupled computer systems so that the computer-readable code is stored and executed in a distributed fashion. For example, such a transitory computer-readable medium may be communicated from one electronic device to another electronic device using any suitable communications protocol. Such a transitory computer-readable medium may embody computer-readable code, instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and may include any information delivery media, A modulated data signal may be a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
[0104] It is to be understood that any or each mod ule of any one or more of any system, device, or server may be provided as a software construct, firmware construct, one or more hardware components, or a combination thereof, and may be described in the general context of computer-executable instructions, such as program modules, that may be executed by one or more computers or other devices. Generally, a program module may include one or more routines, programs, objects, components, and / or data structures that may perform one or more particular tasks or that may implement one or more particular abstract data types. It is also to be understood that the number, configuration, functionality, and interconnection of the modules of any one or more of any system, device, or server are merely illustrative, and that the number, configuration, functionality, and interconnection of existing modules may be modified or omitted, additional mod ules may be added, and foe interconnection of certain modules may be altered.
[0105] While there have been described systems, methods, and computer-readable media for enabling efficient control of a media application at a media electronic device by a user electronic device, it is to be understood that many changes may be made therein without departing from the spirit and scope of the disclosure. Insubstantial changes from the claimed subject matter as viewed by a person with ordinary skill in the art, now known or later devised, are expressly contemplated as being equi valently within the scope of the claims. Therefore, obvious substitutions now or later known to one with ordinary- skill in the an are defined to be within the scope of the defined elements.
[0106] Therefore, those skilled in the art will appreciate that the invention can be practiced by other than the described embodiments, which are presented for purposes of illustration rather than of limitation.
Claims
ClaimsWhat is Claimed is:
1. A method for personalizing large language model (LLM) queries using user-specific profile information at an edge device, the method comprising: obtaining, by a personalized edge model (PEM) implemented on an edge device, a set of user profile information specific to a user and associated with the edge device; training the PEM by: generating one or more query embeddings for a test query using both input data and the set of user profile information; transmitting the test query with the one or more query embeddings to a large language model (LLM) implemented on one or more remote computing devices; and receiving, at the PEM, a response to the test query from the LLM, wherein the response to the test query' is compared to a ground truth to determine an accuracy of the received response to the test query; obtaining, at the PEM, an initial query; generating, by the PEM, a prompt embedding and / or a data embedding based on the set of user profile information specific to the user io generate a personalized query’; transmitting, by the PEM, the personalized query that includes the prompt embedding and / or the data embedding to the LLM; and receiving, at the PEM, a response to the personalized query’ from the LLM.
2. The method of claim L wherein the sei of user profile information comprises data obtained at the edge device that includes any of demographic information relating to the user, interests of the user, and tracked browser data specific to the user.
3. The method of claim 1 , wherein the training further comprises: determining an accuracy of the response to the test query based on a similarity between the response io the test query and the ground truth; updating the one or more embeddings for the test query based on the determined accuracy of the response; transmitting the test query and the updated embeddings to the LLM; and receiving an updated response from the LLM.
4. The method of claim 1. wherein the prompt embedding modifies text of the initial query based on the set of user profile information, and wherein the data embedding transforms the initial query based on the set of user profile information.
5. The method of claim 1. wherein the PE M comprises a Multi-Layer Perceptron (MLP) including a three-layer Fully Connected Network (FCN).
6. llie method of claim 1 , wherein the prompt embedding and / or die data embeddings modify the initial query without providing any sensitive user data to the LLM.
7. A system comprising: one or more computing devices implementing a large language model (LLM); and an edge computing device associated with a user, the edge device implementing a personalized edge model (PEM), wherein the edge device is configured to: obtain an initial query; generate any of a prompt embedding, and a data embedding based on a set of user profile information; transmit any of the prompt embedding and the data embedding to the LLM; and receive a response to the initial query from the LLM.
8. The system of claim 7, wherein the edge device is further configured to train the PEM by: generating one or more query embeddings for a test query using both an input query and the set of user profi le information; transmitting the test query with the one or more query embeddings to a large language model (LLM) implemented on one or more remote computing devices; and receiving, at the PEM. a response to the test query from the LLM, wherein the response to the test query is compared to a ground truth response to determine an accuracy of the response to the test query'.
9. The system of claim 8, wherein the training further comprises:determining an accuracy of the response to the test query based on a similarity between the response to the test query and the ground truth; updating the one or more embeddings for the test query based on the determined accuracy of the response: transmitting the updated embeddings to the LLM; and receiving an updated response from the LLM,10. The system of claim 7, wherein the prompt embedding modifies text of the initial query based on the set of user profile information, and wherein the data embedding transforms the initial query based on the set of user profile information,1 1 , The system of claim 7, wherein the PEM comprises a Multi-Layer Perceptron (MLP) including a three-layer Fully Connected Network (FCN) configured to extract textual features in. the initial query.12, The system of claim 1 1 , wherein the PEM further comprises an optimizer configured to generate the prompt embeddings and / or the data embeddings.
13. The system of claim 7, wherein the sei of user profile information comprises data obtained at the edge device that includes any of demographic information relating to the user, interests of the user, and tracked browser data specific to the user, and wherein the set of user profile information Is stored as part of a database or table accessible to the edge de vice. 14, A computer-implemented method comprising: obtaining, by a personalized edge model (PEM) implemented on an edge device, a set of user profile information specific to a user and associated with the edge device; training the PEM by: generating one or more query embeddings for a test query using both input data and the set of user profile information; transmitting the one or more query embeddings to a large language model (LLM) implemented on one or more remote computing devices;receiving, at the PEM, a response to the test query from the LLM. wherein the response to the test query is compared to a ground truth to determine an accuracy of the received response to the test query; determining an accuracy of the response to the test query based on a similarity between the response to the test query and the ground truth; updating the one or more embeddings for the test query based on the determined accuracy of the response; transmitting the test query and the updated embeddings to the LLM; and receiving an updated response from the I.LM: obtaining, at the PEM, an initial query: generating, by the PEM, a prompt embedding and / or a data embedding based on the set of user profile information specific to the user; transmitting, by the PEM, any of the prompt embedding and the data embedding to the LLM; and receiving, at the PEM, a response to the initial query from the LLM15. The computer-implemented method of claim 14, wherein the PEM comprises a Multi- Layer Perceptron (MLP) including a three-layer Fully Connected Network (FCN) configured to extract textual features in the initial query.
16. The computer-implemented method of claim 15, wherein the PEM further comprises an optimizer configured to generate the prompt embeddings and / or the data embeddings.
17. The compu ter-implemented method of claim 14. wherein the set of user profi le information comprises data obtained at the edge device that includes any of demographic information relating to the user, interests of the user, and tracked browser data speci fic to the user.
18. The eo.mputer-implemented method of claim 14, wherei n the set of user profile information is stored as part of a database or table accessible to the edge device.
19. The computer-implemen ted method of claim 14, wherein the prompt embedding modifies text of the initial query based on the sei of user profile information, and wherein thedata embedding comprises a transformation of the initial query based on the set of user profile information.20, The computer-implemented method of claim 14, wherein the prompt embedding and / or the data embeddings modify the initial query without providing any sensitive user data to the UM.
Citation Information
Patent Citations
Machine-Learned Language Models Which Generate Intermediate Textual Analysis in Service of Contextual Text Generation
US20230177276A1