Generating and implementing contextual profiles in processing queries using base model
By iteratively analyzing multiple context profiles through a context generation system, and combining edge networks and cloud computing systems, the resource and latency issues of the basic model in generating contexts were resolved, achieving efficient and accurate response generation.
Patent Information
- Application Number
- CN202480020190.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-15
- Filing Date
- 2024-06-10
- Publication Date
- 2025-11-04
AI Technical Summary
The base model is limited by computational resources and latency requirements when generating contexts, and performs poorly, especially in environments with limited processing power and strict latency requirements, resulting in inaccurate responses and high costs.
By iteratively analyzing multiple context profiles through a context generation system, the system selects context profiles with high relevance scores and low costs to generate responses. This approach, combined with edge networks and cloud computing systems, optimizes resource utilization and latency management.
It improves the accuracy of responses and reduces the cost of generating responses, enabling efficient context generation within limited resources and latency.
Smart Images

Figure CN120898199A_ABST
Abstract
Description
BACKGROUND
[0001] In recent years, the proliferation and application of artificial intelligence (AI) and machine learning (ML) has increased significantly. Moreover, as services hosted by cloud computing systems become increasingly available to end users and other organizations, the accessibility of more complex and robust computing models, such as large language models (LLMs), has become more common. These base models can be trained to perform a wide variety of tasks, such as chatbots, providing answers to general questions, generating code and other programming scripts, and in some cases, providing specific information about a particular topic.
[0002] While base models and other base models provide useful tools for a wide variety of applications, base models have a number of limitations and drawbacks. For example, base models are often limited in providing information or processing tasks that involve specialized knowledge about a particular domain. In these and other applications, contextual retrieval can be very important to ensure that base models provide accurate and complete information. In traditional base models, the process of generating context often involves the utilization of large amounts of computing resources, and can be problematic for computing environments that have limited processing capabilities and / or strict latency requirements. These and other problems exist in generating context and utilizing base models in a wide variety of applications. BRIEF DESCRIPTION OF DRAWINGS
[0003] Figure 1 An example environment including a context extraction system for retrieving context to process queries is shown in accordance with one or more embodiments.
[0004] Figure 2 An example implementation of a context extraction system on a server device is shown in accordance with one or more embodiments.
[0005] Figure 3 An example workflow showing an example implementation of a context extraction system in accordance with one or more embodiments is shown.
[0006] Figure 4 Another example workflow showing an example implementation of a context extraction system in accordance with one or more embodiments is shown.
[0007] Figure 5 An example implementation of a context extraction system including one or more components implemented on an edge network is shown.
[0008] Figure 6 An example implementation of a context extraction system utilizing a compressed base model in accordance with one or more embodiments is shown.
[0009] Figure 7Another example implementation of a context extraction system including one or more components implemented on an edge network is shown.
[0010] Figure 8 A series of actions for extracting context and processing a query using a base model is shown in accordance with one or more embodiments.
[0011] Figure 9 Another series of actions for extracting context and processing a query using a base model is shown in accordance with one or more embodiments.
[0012] Figure 10 Certain components that can be included within a computer system are shown. DETAILED DESCRIPTION
[0013] This application relates to systems, methods, and computer-readable media for generating context for a base model, such as a large language model (LLM) or other artificial intelligence (AI) or machine learning (ML) system(s). The base model can prepare multiple context profiles for a particular query. Each context profile can have a different complexity and associated cost (e.g., latency and / or processing budget). The base model can prepare an initial analysis of the query based on each context profile within a particular latency budget. Responses based on different context profiles can be analyzed for relevance, and a relevance score can be generated for each response. If the response for one of the context profiles is within a threshold relevance, that context profile can be selected and a final response to the query is prepared. This can help improve the relevance of the response while reducing the cost of the response by the base model to the query.
[0014] In some embodiments, the base model is used to iteratively check context profiles until a response within a relevance threshold is prepared. For example, the base model can prepare a first context profile and prepare a first response based on the first context profile. If the first response is not within the relevance threshold, the base model can prepare a second context profile and prepare a second response based on the second context profile. If the second response is within the relevance threshold, the base model can prepare a final response based on the second context profile. If the second response is not within the relevance threshold, the base model can iteratively prepare additional context profiles until a response within the relevance threshold results from one of the context profiles. In some embodiments, subsequent context profiles have higher complexity and higher associated costs. In some embodiments, subsequent context profiles are prepared based on a predetermined pattern. In some embodiments, if the base model reaches a latency budget, the base model selects a context profile associated with a response having a higher relevance.
[0015] In some embodiments, the base model simultaneously examines the context profiles and selects the best response from the simultaneously examined responses. For example, the base model can confirm multiple context profiles and simultaneously prepare responses based on the context profiles. The base model can prepare a relevance score from the responses and select the context profile with the higher (or highest) relevance score. In some examples, the base model can prepare the context profiles to be analyzed within a latency budget.
[0016] In some embodiments, the base model maintains a set of different context profiles, their associated computation times, and possibly relevance scores. For example, as the base model processes queries using different context profiles, the base model can record the particular context profile, computation time, and relevance score. As the base model builds a database or collection of these context profiles, the base model can select a set of context profiles to analyze in response to a particular query. In this way, the base model can generate a response using the context profile most likely to fall within a relevance threshold while reducing the cost (e.g., latency and / or resource utilization) of the response. As used herein, “response cost” refers to a measure of latency and / or utilization of resources. For example, a high cost can refer to a large amount of latency and / or a large amount of computational resources utilized by the system(s). Alternatively, a low cost can refer to a low measure of latency and / or a small amount of computational resources utilized by the system(s).
[0017] In some embodiments, a method includes receiving a query for a base model. The base model can receive the query and extract a first context profile for the query using a language and a context database (or domain database) of the query. The first context profile can include first context content based on the language of the query. Using the first context profile, the context generation system can generate a first prompt for the base model. The first prompt includes a first concatenation of the query and the first context content. The context generation system can input the first prompt into the base model, resulting in a first response from the base model. The context generation system can generate a first relevance score or a relevance score for the first response based on a relevance of the first response to the query. The context generation system can extract a second context profile for the query using the language and the context database of the query. The second context profile can include second context content based on the language of the query. Using the second context profile, the context generation system can generate a second prompt for the base model. The second prompt includes a second concatenation of the query and the first context content. The context generation system can input the second prompt into the base model, resulting in a second response from the base model. The context generation system can generate a second relevance score or a relevance score for the second response based on a relevance of the second response to the query. The context generation system can select one of the first response or the second response based on the first relevance score and the second relevance score.
[0018] In some embodiments, a context generation system receives a query for a base model. The context generation system extracts a plurality of context profiles from a context database using the query. The context generation system generates a plurality of prompts for the base model. Each prompt of the plurality of prompts is generated using context from an associated context profile of the plurality of context profiles. The context generation system inputs the plurality of prompts into the base model to generate a plurality of responses to the query. The context generation system generates a plurality of relevance scores. Each relevance score of the plurality of relevance scores is associated with a response of the plurality of responses. The relevance score indicates a relevance of the response to the query. The context generation system selects a best response of the plurality of responses based on a higher relevance score of the plurality of relevance scores.
[0019] The context generation system provides a number of advantages and benefits over conventional systems and methods. For example, by analyzing a plurality of responses based on a plurality of context profiles, the context generation system improves accuracy of responses to queries relative to conventional systems. Indeed, analyzing a plurality of responses and iteratively determining which responses are more relevant enables the context generation system to determine which context profile provides a more accurate answer. In some embodiments, analyzing a plurality of responses allows the context generation system to generate over time a database of context profiles that have been found to provide relevant and cost-efficient results to requests.
[0020] In some examples, the context generation system can reduce the amount of processing resources used to generate responses to input queries by analyzing multiple responses based on multiple context profiles. For example, in one or more implementations, the context generation system iteratively evaluates the relevance of different context profiles to determine a context profile that can provide an accurate response that meets a relevance score threshold. Further, the context generation system can take into account latency budgets associated with generating and utilizing context profiles to generate and / or confirm a particular context profile that is both relevant and falls within a threshold latency budget.
[0021] In some examples, the context generation system can pre-select one or more context profiles based on a historical data set of executing multiple previous queries. For example, the context generation system can maintain a profile database of historical relevance scores for various context profiles. When the context generation system receives a query, the context generation system can compare the query and associated context to the profile database. The context generation system can determine the context profile by pre-selecting one or more context profiles that are predicted to generate a relevant response within a threshold latency. This pre-selection of context can significantly reduce the utilization of processing resources on the cloud and / or on the client device.
[0022] The context generation system can additionally be implemented in a flexible manner that facilitates offloading processing to different computing environments. For example, in one or more implementations described herein, the context generation system facilitates offloading context generation to an edge network, while actions that apply more robust base models can be performed on data centers on a cloud computing system. In one or more embodiments, context selection or generation can be performed in part on a client device and / or on an edge network, which provides faster latency and facilitates generating responses to input queries within a limited latency budget.
[0023] As shown in the foregoing discussion, the present disclosure utilizes a variety of terms to describe features and advantages of the context generation system. Additional details regarding the meaning of a number of these terms are now provided.
[0024] For example, as used herein, a “base model” refers to an AI model or ML model that is trained based on a large dataset to generate an output in response to an input. The base model can include a neural network with a large number of parameters (e.g., billions of parameters) that the base model can take into account when performing a task or otherwise generating an output based on an input. In one or more embodiments described herein, the base model is trained to generate a response to a query. In some implementations, the base model refers to a large language model (LLM). The base model is trained in pattern recognition and text prediction. For example, the base model can be trained to predict the next word of a particular sentence or phrase. In one or more implementations described herein, the base model specifically refers to an LLM, but other types of base models can be used when generating a response to an input query. Indeed, while one or more embodiments described herein relate to features associated with determining a context for an LLM, similar features can be applied to determining and / or generating a context for other types of base models.
[0025] As used herein, the term “compressed base model” or “compressed model” refers to a simplified version of a base model that has fewer parameters than the associated non-compressed model. For example, a compressed base model can be trained to perform the same or similar tasks or functions as a corresponding non-compressed base model (or simply “base model”) while using fewer parameters. For example, as will be discussed herein, a base model can prepare a response to a query that is not deep enough, not relevant enough, or not accurate enough, but can provide a relevance score that is used to predict a relevance score of a non-compressed version of the base model if a similar or identical input is provided to the non-compressed version of the base model.
[0026] As used herein, “context” can be information that can be used by a base model or other machine learning model to generate a more relevant or accurate response to a query by a guide model. Context information can include information related to a query that is not directly stated in the query. For example, in one or more embodiments described herein, context is information generated based on a text similarity measure between a query and a database of additional information (e.g., domain-specific information). As will be discussed in further detail below, context information can be identified through a wide variety of mechanisms. For example, context information can be identified using a wide variety of different text similarity measures. Determining a text similarity measure can involve comparing text of a query to text of one or more documents in a context database or a domain-specific database. Text similarity measures can be generated using any number of techniques. For example, a text similarity measure can refer to different types of similarity measures including, by way of example and not limitation, cosine similarity, Jaccard coefficient, L2 norm, L∞ norm, inverted file index, Hamming distance, any other text similarity measure, and combinations thereof. In some examples, context information can be generated using vector embeddings. In some examples, context information can include a context database and / or a plugin that uses a context database to generate.
[0027] As used herein, a “context profile” can be a particular combination of one or more techniques used to generate context. For example, a context profile can refer to context generated based on a particular text similarity measure. In some examples, a context profile can include a particular vector embedding format. In some examples, a context profile can include a particular plugin related to a domain of a context database. In some examples, a context profile can include a combination of techniques used to generate context, including one or more of a text similarity measure, a vector embedding, or a plugin. In some examples, a context profile can include multiple elements of techniques, such as multiple text similarity measures, multiple vector embedding formats, multiple plugins, and combinations thereof.
[0028] As used herein, an “edge network” or “edge data center” can interchangeably refer to an extension of a cloud computing system located on the periphery of the cloud computing system. An edge network can refer to a tier of one or more devices that provide connectivity to devices and / or services on data centers within a cloud computing system framework. An edge network can provide a plurality of cloud computing services on hardware with an associated validated configuration without requiring a client to communicate with internal components of the cloud computing infrastructure. In effect, an edge network provides a virtual access point that enables more direct communication with components of a cloud computing system as compared to another entry point of the cloud computing system, such as a public entry point.
[0029] Figure 1This is a representation of a computing system 100 according to at least one embodiment of the present disclosure. The computing system 100 may host or implement a base model 102. The base model 102 can be accessed via the Internet or other communication networks (e.g., 5G telecommunications networks, cloud computing networks, edge networks). As an example, the base model 102 may reside on a cloud computing network 101 accessible via the Internet. In some embodiments, some or all of the base model 102 may reside on an edge network of the cloud.
[0030] The base model 102 can be trained based on information from the domain database 115. In some embodiments, the entire domain database 115 used to train the base model 102 is implemented on the Internet and / or accessible via the Internet. In some embodiments, at least a portion of the domain database is stored on the same server and / or in the same location as the base model 102 and is accessible via a local network. In some embodiments, at least a portion of the domain database 115 is located at the edge network where the base model 102 resides.
[0031] Users can access the base model 102 through user device 108. Users can generate queries on user device 108. The base model 102 can receive queries over a network. The base model 102 can prepare responses to queries based on information in the domain database 115.
[0032] The computing system 100 includes one or more computing devices 103 that can host the context extraction system 110. In some embodiments, the context extraction system 110 receives queries from a user device 108. The context extraction system 110 can communicate with an underlying model 102. In some embodiments, the context extraction system 110 communicates directly with the underlying model 102. For example, the context extraction system 110 may reside on the same local network, edge network, and / or data center as the underlying model 102 within a cloud computing network 101. In some embodiments, the context extraction system 110 communicates with the underlying model 102 via the Internet. In some embodiments, the context extraction system 110 resides on the user device 108. For example, the context extraction system 110 may be part of an application on the user device 108. In some examples, the context extraction system 110 may reside at least partially on a cloud and / or edge network.
[0033] The context extraction system 110 can determine a context profile from the query. For example, the context extraction system 110 can generate a context profile based on content from the domain database 115 and text from the query. In some examples, the context extraction system 110 can select a context profile from a profile database of context profiles. The context extraction system 110 can use context information from the context profile and the query to generate a prompt for the base model 102. For example, the context extraction system 110 can generate a text concatenation of the query and the context information from the context profile to generate the prompt. The base model 102 can prepare a response to the query based on the input prompt and send the response back to the user device 108. Additional information related to the steps of the process described above will be discussed in further detail below.
[0034] As discussed in further detail herein, the context extraction system 110 can generate multiple context profiles. The context extraction system 110 can analyze the multiple context profiles and identify a best context profile for the base model 102 to utilize. For example, the context extraction system 110 can iteratively analyze the multiple context profiles until the context extraction system 110 identifies a context profile with a relevance or a context profile with a relevance score greater than or equal to a threshold relevance. The context extraction system 110 can select the best context profile and use the best context profile to prepare a prompt for the base model 102.
[0035] In some examples, the context extraction system 110 can analyze multiple context profiles (e.g., any number of context profiles) simultaneously in parallel. The context extraction system 110 can prepare a relevance score for each context profile. The context extraction system 110 can select a context profile with a higher relevance score (or a relevance score that exceeds a threshold and has a lower expected latency). The context extraction system 110 can then send a prompt to the base model 102 based on the best context profile.
[0036] In some embodiments, the context extraction system 110 analyzes the context profiles by the base model 102 running a prompt on each context profile. This can generate a complete answer to the query for each context profile. In some embodiments, the context extraction system 110 analyzes the context profiles by running a compressed base model prompt on the context profiles. The compressed base model can generate a relevance score based on a lower cost analysis of the query while still indicating which context profiles are best or which context profiles exceed a threshold.
[0037] In some embodiments, the context extraction system 110 includes a profile database. The profile database can include records of a plurality of context profiles associated with queries. When the context extraction system 110 receives a query from a user, the context extraction system 110 can determine which context profiles in the profile database can generate answers above a relevance threshold for the relevance. The context extraction system 110 can select one or more context profiles for the base model 102 to process. In this way, the context extraction system 110 can determine which context profiles the base model 102 uses to prepare a response to the query.
[0038] According to at least one embodiment of the present disclosure, the context extraction system 110 can use empirical data to determine which context profiles in the profile database can be most relevant to a particular query. For example, when the base model 102 answers a query using a particular context profile, the context extraction system 110 can record a relevance score associated with that context profile in the profile database. The context extraction system 110 can utilize the profile database to provide and / or recommend better context profiles to answer queries.
[0039] Figure 2 FIG. 2 is a representation of a computing system 200 according to at least one embodiment of the present disclosure. Each component of the computing system 200 can include software, hardware, or both. For example, a component can include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices, such as a client device or a server device. When executed by one or more processors, the computer-executable instructions of the computing system 200 can cause the computing device(s) to perform the methods described herein. Alternatively, a component can include hardware, such as a specialized processing device for performing a certain function or group of functions. Alternatively, a component of the computing system 200 can include a combination of computer-executable instructions and hardware.
[0040] Further, a component of the computing system 200 can be implemented, for example, as one or more operating systems, one or more standalone applications, one or more modules of an application, one or more plug-ins, one or more library functions or functions callable by other applications, and / or a cloud computing model. Thus, a component can be implemented as a standalone application, such as a desktop or mobile application. Further, a component can be implemented as one or more web-based applications hosted on a remote server. A component can also be implemented in a suite of mobile device applications or “apps.”
[0041] According to at least one embodiment of the present disclosure, computing system 200 includes a context extractor 210, a domain database 215 in communication with context extractor 210, and a base model 202 in communication with domain database 215 and context extractor 210. Context extractor 210 includes a context extraction system 214. Context extraction system 214 can extract a context from domain database 215 based on a query using one or more of a variety of context extraction techniques. For example, context extraction system 214 can utilize a text similarity measure 216, a vector embedding 218, a plugin 220, any other context extraction technique, and combinations thereof to generate a context for a query. In some embodiments, context extraction system 214 extracts a context profile including context information from a text similarity measure 216, a vector embedding 218, a plugin 220, any other context extraction technique, and combinations thereof.
[0042] Domain database 215 can be a database from which a context for a context profile is extracted. In some embodiments, base model 202 is trained by information in domain database 215. In some embodiments, domain database 215 refers to an entire domain-specific database used to train base model 202. In some embodiments, domain database 215 includes a subset of information used to train base model 202.
[0043] Domain database 215 can include context documents 222. Context documents 222 can be any type of document used by context extraction system 214 to generate a context. For example, context documents 222 can include documents accessible over the internet, documents saved locally at context extractor 210 and / or a local database, documents related to a context for base model 202, documents related to a focus of base model 202, if any, any other context document, and combinations thereof. In some embodiments, context extraction system 214 can apply a text similarity measure 216 to context documents 222 in domain database 215. Context extraction system 214 can generate a context profile using a context generated by a text similarity measure 216 analyzing a text similarity between a query and context documents 222.
[0044] The domain database 215 can include a plugin database 224. The plugin database 224 can include a database of any plugins that can be used by the base model 202. Plugins can be used to focus the base model 202 on a particular search and / or to provide context during the process of generating a context profile. For example, a plugin can include publicly available information for a particular website or a particular company. As an example, non-limiting example, a plugin can include information from the website YELP. When a user inputs a query asking for the best restaurants, the base model 202 can utilize information from YELP to prepare a response. When the context extractor 210 receives the query asking for the best restaurants, the context extraction system 214 can identify a plugin 220 from the plugin database 224 that confirms YELP. This can help focus the search of the base model 202 to the relevant website, thereby reducing the cost of the search of the base model 202.
[0045] The context extractor 210 can include a prompt generator 226. The prompt generator 226 can utilize the query and the context information to generate a prompt to be input into the base model 202. The prompt generator 226 can generate the prompt in any manner for use by the base model 202. For example, the prompt generator 226 can prepare a string concatenation of the query and the context information. In some examples, the prompt generator 226 can generate the prompt as a script in a scripting language or generate instructions for the base model 202. In some examples, the prompt generator 226 can prepare an entry for an entry form for the base model 202. In some examples, the prompt generator 226 can generate a prompt that is focused on and / or specific to a particular base model 202, such as using a particular syntax.
[0046] The context extractor 210 can include a relevance analyzer 228. The relevance analyzer 228 can analyze each context profile to determine a relevance score for the context profile. The relevance analyzer 228 can determine the relevance score for the context profile in any manner. For example, the relevance analyzer 228 can determine the relevance score for the context profile based on the details and / or content of the context profile compared to the query. In some examples, the relevance analyzer 228 can determine the relevance score for the context profile based on a result response from the base model 202 to the query. To analyze relevance using the context profile, the relevance analyzer 228 can analyze the context document 222 used by the text similarity measure 216 and / or the vector embedding 218 in the context profile. In some examples, to analyze relevance using the context profile, the relevance analyzer 228 can analyze the plugin database 224 using the plugin 220 identified by the context extraction system 214.
[0047] In some examples, the relevance analyzer 228 can determine a relevance score for a context profile based on an initial response to a query. For example, the compressed base model 230 can use a prompt from the prompt generator 226 to prepare an initial response to a query. The relevance analyzer 228 can analyze the initial response from the compressed base model 230 and determine how the response relates to the query.
[0048] The context extractor 210 can include a response selector 232. The response selector 232 can analyze the relevance scores from the relevance analyzer 228 to select a best context profile. In some embodiments, a context profile is selected based on the context profile having a higher relevance score (e.g., having a higher relevance score than other context profiles and / or a relevance score above a threshold, while having a lower latency than other context profiles). In some embodiments, the best context profile is the context profile with a relevance score above a relevance threshold with a lowest cost (e.g., a lowest expected latency budget). For example, multiple analyzed context profiles can have a relevance score greater than a relevance threshold. To reduce processing of the base model, the response selector 232 can select the context profile with the lowest cost.
[0049] According to at least one embodiment of the present disclosure, the context extractor 210 analyzes multiple context profiles for a single query. In some embodiments, the base model 202 has a latency budget. The latency budget can be an amount of time the base model 202 has to prepare a response to a query. The latency budget can include transmission time between a user device and the base model 202 over the internet or a local network. In some examples, the latency budget can include processing time for the base model 202. In some examples, the latency budget can include a specific amount of time for confirming and selecting a best context profile.
[0050] As discussed herein, the context extractor 210 can include a profile database 212. As the context extractor 210 prepares and analyzes context profiles, the context extractor 210 can save the resulting costs, context information, relevance scores, queries, prompts, any other information, and combinations thereof to the profile database 212. Over time, as the context extractor 210 processes queries for the base model 202, the profile database 212 can include multiple relevance scores for a particular query and / or context profile.
[0051] According to at least one embodiment of the present disclosure, upon receiving a query by the context extractor 210, the context extractor 210 determines which context profiles can have a higher likelihood of having a relevance score above a relevance threshold (and / or above other generated context profiles). For example, the context extractor 210 can analyze the profile database 212 to identify context profiles that have been previously used to answer related queries. The context extractor 210 can empirically determine which context profiles are associated with high relevance scores and / or low processing costs. In some embodiments, the context extractor 210 can empirically determine which context profiles are associated with a ratio of relevance score to processing cost.
[0052] To prepare context profiles for a particular query, the context extractor 210 can select queries from the profile database 212 that can be analyzed by the context extractor 210. In some embodiments, pre-selecting one or more context profiles to be analyzed by the context extractor 210 helps to reduce the processing cost of analyzing context profiles. In some embodiments, pre-selecting one or more of the context profiles to be analyzed by the context extractor 210 helps to identify the best context profile while reducing the total number of context profiles that are analyzed. For example, a context profile in the profile database 212 that has a low stored relevance score can not be analyzed because it is unlikely that the relevance score of the context profile will be above the threshold. In some examples, a context profile in the profile database 212 that has a stored relevance score close to the relevance threshold can be analyzed by the context extractor 210 to determine whether the relevance score is actually above the relevance threshold.
[0053] Figure 3 is a schematic representation of a context analysis system 334 according to at least one embodiment of the present disclosure. The context analysis system 334 can include a context extractor 310. The context extractor 310 can receive a query 336. The context extractor 310 can analyze the query 336 and use a context database 315 to prepare a plurality of context profiles 338. A prompt generator 326 can receive the query 336 and the context profiles 338 and prepare a plurality of prompts 340. A base model 302 can receive the prompts 340 and generate a response 342 for each prompt. Each response 342 can have an associated relevance score. A response selector 344 can select a best response 346 from the responses 342 and send the best response to a user.
[0054] The context analysis system 334 can generate and analyze multiple context profiles to generate the best response 346 having a cost above a relevance threshold. As discussed herein, the context analysis system 334 can analyze the context profiles 338 in any manner. For example, the context analysis system 334 can iteratively analyze the context profiles 338 until a context profile reaches a relevance score above a relevance threshold. In some examples, the context analysis system 334 can analyze the context profiles 338 in parallel and select the context profiles having a relevance score above the relevance threshold. In some examples, the context analysis system 334 can identify multiple responses 342 having a relevance score above the relevance threshold. The response selector 344 can select the best response 346 having the lowest cost among all responses 342 having a relevance score above the threshold relevance score.
[0055] According to at least one embodiment of the present disclosure, the context analysis system 334 can analyze the context profiles 338 in parallel. For example, the context extractor 310 can extract multiple context profiles for 336 and submit all of the context profiles to the prompt generator 326 in parallel. The prompt generator 326 can generate prompts 340 based on all of the context profiles 338 and send all of the prompts 340 to the base model 302. The base model 302 can analyze all of the prompts 340 and send all of the responses 342 to the response selector 344. The response selector 344 can analyze all of the responses 342 and select the best response 346 from the list of responses 342.
[0056] In some embodiments, the context extractor 310 prepares the context profiles 338 based on a latency budget. Each context profile 338 can have an associated cost, and the cost can include a latency cost. The context extractor 310 can prepare the context profiles 338 such that the processing of the context profiles 338 (including profile generation, prompt generation, base model 302 analysis, and response selection) can occur within the latency budget.
[0057] In some embodiments, the context analysis system 334 iteratively analyzes multiple sets of context profiles 338. For example, the context extractor 310 can generate a first set of context profiles 338. The context analysis system 334 can analyze the first set of context profiles 338. If the best response 346 cannot be selected, the context analysis system 334 can cause the context extractor 310 to generate a second set of context profiles 338, and the context analysis system 334 can analyze the second set of context profiles 338 to determine whether the best response 346 can be determined. This process can be repeated until the best response 346 can be selected.
[0058] As discussed herein, the context extractor 310 can be in communication with a profile database 312. The profile database 312 can include records of context profiles historically used, such as context profile A, context profile B, context profile C, and the like. The context profiles in the profile database 312 can include details about the context profile. For example, each context profile in the profile database 312 can include a technique for extracting a context, such as a similarity measure, a vector embedding, a plug-in, any other technique, and combinations thereof.
[0059] The profile database 312 can also include a cost associated with a particular context profile. In the illustrated embodiment, the cost is shown as a latency cost. For example, each context profile in the profile database 312 can have an associated latency cost for extracting a context. But it should be understood that the cost can be any type of cost, such as a processor capacity cost, a transmission bandwidth cost, a financial cost, any other cost, and combinations thereof.
[0060] The profile database 312 can also include a relevance score associated with each context profile. For example, each context profile can include a relevance score associated with a particular context extraction technique. In some examples, the relevance score can be an estimated relevance score. In some embodiments, the relevance score is based on previous uses of the particular context profile. For example, the context analysis system 334 can record each of the context profiles in the context profile 338 used to generate the best response 346 in the profile database 312 along with their respective relevance scores. The relevance scores listed in the profile database 312 can include an average relevance score for multiple uses of the associated context profile. In some embodiments, the relevance score is associated with a particular query, a type of query, a subject matter of a query, or other query-based metrics.
[0061] The profile database 312 can include a record of each context profile used when the context analysis system 334 analyzes a query. The context extractor 310 can select a context profile from the profile database 312 based on the expected relevance score in the profile database 312. In this way, the context extractor 310 can use historical relevance data for context profiles to determine a context profile that is most likely to result in a relevance score above a relevance threshold for a particular query 336.
[0062] Figure 4is a schematic representation of a contextual analysis system 434 according to at least one embodiment of the present disclosure. The contextual analysis system 434 can include a context extractor 410. The context extractor 410 can receive a query 436. The context extractor 410 can analyze the query 436 and use a context database 415 to prepare a context profile 438. A prompt generator 426 can receive the query 436 and the context profile 438 and prepare a prompt 440. The base model 402 can receive the prompt 440 and generate a response 442 to the prompt.
[0063] The response 442 can have an associated relevance score. A response selector 444 can examine the response 442 and determine 445 whether the response 442 is above a response threshold. If the response 442 is above the response threshold, the response selector 444 can submit the response 442 to the user. If the response 442 is not above the response threshold, the response selector 444 can request that the context extractor 410 generate a new context profile. The context extractor 410 can generate a new context profile 438, the prompt generator 426 can generate a new prompt 440, and the base model 402 can generate a new response 442. In this way, the contextual analysis system 434 can iteratively prepare responses to the query 436 until the response selector 444 identifies a response that is above the relevance threshold. When the response selector 444 identifies a response that is above the relevance threshold, the response selector 444 can select that response as the best response 446.
[0064] According to at least one embodiment of the present disclosure, the first context profile 438 generated by the context extractor 410 is a low-cost context profile 438. The cost of subsequent context profiles 438 can incrementally increase until the response selector 444 confirms a response 442 with a relevance score greater than the relevance threshold. This can facilitate identifying a context profile 438 with an overall lowest cost.
[0065] In some embodiments, the first context profile 438 generated by the context extractor 410 is not the lowest-cost context profile 438. The context profiles 438 can generate context profiles based on other factors. For example, the context profiles 438 can generate context profiles from a profile database. The context extractor 410 can select the first context profile 438 based on profile selection criteria, such as relevance scores, cost, or other profile selection criteria. In some examples, the context extractor 410 can select subsequent context profiles 438 that prioritize higher relevance scores, lower cost, specific context extraction techniques, any other metric, and combinations thereof.
[0066] In some embodiments, the context analysis system 434 iteratively analyzes the context profiles 438 until the response selector 444 identifies a first response 442 with a relevance score above a relevance threshold. In some embodiments, the context analysis system 434 iteratively analyzes the context profiles 438 until a latency budget has been exceeded. In some embodiments, the context analysis system 434 iteratively analyzes the context profiles until a latency budget has been exceeded, and if no response 442 has a relevance score above a relevance threshold, the context analysis system 434 can continue to iteratively analyze the context profiles 438 until the relevance threshold is satisfied.
[0067] Figure 5 is a schematic representation of a context analysis system 534 according to at least one embodiment of the present disclosure. The context analysis system 534 can include a context extractor 510. The context extractor 510 can receive a query 536. The context extractor 510 can analyze the query 536 and use a context database 515 to prepare a plurality of context profiles 538. A prompt generator 526 can receive the query 536 and the context profiles 538 and prepare a plurality of prompts 540. A base model 502 can receive the prompts 540 and generate a response 542 for each prompt. Each response 542 can have an associated relevance score. A response selector 544 can select a best response 546 of the responses 542 and send the best response to the user.
[0068] According to at least one embodiment of the present disclosure, at least a portion of the context analysis system 534 is offloaded to a cloud network or an edge network 548. In some cases, the processing of a query can take up valuable processing power on a user device. This can reduce the user device’s ability to perform other computational processes and / or consume the user device’s battery. Offloading at least a portion of the context analysis system 534 to the edge network 548 can help reduce the processing that occurs on the user device.
[0069] In the illustrated embodiment, the context selection by the context extractor 510 and the prompt generation by the prompt generator 526 are offloaded to the edge network 548. However, it should be understood that any portion of the context analysis system 534 can be offloaded to the edge network 548. For example, the context database 515 or at least a portion of the context database 515 can be stored on the edge network 548. In some examples, the base model 502 or at least a portion of the base model 502 can be located on the edge network 548. For example, an initial check portion of the base model 502 can be located on the edge network 548. In some examples, the response selector 544 can be located on the edge network 548. In some embodiments, the entire context analysis system 534 can be located on the edge network 548.
[0070] In some embodiments, one or more processes of the context analysis system 534 can be executed on an edge network of a cloud computing system, e.g., an edge of a telecommunications network. For example, determining the plurality of context profiles can be executed on a server device on an edge network of a fifth generation (5G) telecommunications environment. The base model can be implemented on a data center of a cloud computing system accessible via the edge network.
[0071] Figure 6 is a schematic representation of a context analysis system 634 according to at least one embodiment of the present disclosure. The context analysis system 634 can include a context extractor 610. The context extractor 610 can receive a query 636. The context extractor 610 can analyze the query 636 and prepare a plurality of context profiles 638 using a context database 615. A prompt generator 626 can receive the query 636 and the context profiles 638 and prepare a plurality of prompts 640.
[0072] According to at least one embodiment of the present disclosure, one or more compressed base models 630 receive the prompts 640. The compressed base model(s) 630 can prepare a plurality of initial responses 650 to the query 636. As discussed herein, the compressed base model(s) 630 can be versions of the base model 602 that utilize fewer connections and / or have lower complexity. The compressed base model(s) 630 can generate answers at a lower cost than the base model 602. In some embodiments, the compressed base model(s) 630 can generate an initial response 650 for each of the context profiles 638. The initial responses 650 can be less complete responses than the response that the base model 602 would generate, but are sufficiently complete or good for analysis purposes. In one or more embodiments, the environment includes a plurality of compressed base models 630 that can be iteratively run on the generated prompts before providing the query and context to the full base model 602.
[0073] The response selector 644 can analyze the initial responses 650 and determine a best initial response 652 from the initial responses 650. For example, the response selector 644 can generate an initial relevance score for each initial response 650. Based on the initial relevance scores, the response selector 644 can generate the best initial response 652. The response selector 644 can send the best initial response 652, including the associated context profile 638 and prompt 640. The base model 602 can generate a full response 654 to the query 636 and send the full response 654 to the user. In this way, the compressed base model(s) 630 can help reduce the processing load and / or cost of the context analysis system 634.
[0074] As discussed in further detail herein, the compression base model(s) 630 can be utilized as the context analysis system 634 iteratively and / or in parallel processes the context profiles 638. For example, as the context profiles 638 are iteratively analyzed, the compression base model(s) 630 can iteratively prepare initial responses 650 to the query. This can help reduce the cost of each iteration, allowing more iterations to be performed before the latency budget is reached. In some examples, the compression base model(s) 630 can analyze each of the context profiles 638 in parallel. This can allow the compression base model(s) 630 to analyze more context profiles 638 in parallel to determine the best response to the query.
[0075] Figure 7 is a schematic representation of a context analysis system 734 according to at least one embodiment of the present disclosure. The context analysis system 734 can include a context extractor 710. The context extractor 710 can receive a query 736. The context extractor 710 can analyze the query 736 and use a context database 715 to prepare a plurality of context profiles 738. A hint generator 726 can receive the query 736 and the context profiles 738 and prepare a plurality of hints 740.
[0076] According to at least one embodiment of the present disclosure, one or more compression base models 730 receive the hints 740. The compression base model(s) 730 can prepare a plurality of initial responses 750 to the query 736. As discussed herein, the compression base model(s) 730 can be versions of the base model 702 that utilize fewer connections and / or have lower complexity. The compression base model(s) 730 can generate answers with lower processing and latency costs than the base model 702. In some embodiments, the compression base model(s) 730 can generate an initial response 750 for each of the context profiles 738. The initial responses 750 can be less complete responses than the response that the base model 702 would generate, but are sufficiently complete or good enough for analysis purposes.
[0077] According to at least one embodiment of the present disclosure, at least a portion of the context analysis system 734 can be located and / or hosted on a remote network (e.g., an edge network 748 as illustrated). As discussed herein, offloading at least a portion of the context analysis system 734 on the edge network 748 can help reduce processing costs on the user device.
[0078] Any portion of the context analysis system 734 can be offloaded to the context analysis system 734. For example, the compression base model(s) 730 can be located on and / or hosted by the context analysis system 734. Offloading the compression base model(s) 730 to the edge network 748 can allow the context analysis system 734 to process more context profiles 738. In this way, the context analysis system 734 can generate context profiles 738 that are more relevant to the query 736.
[0079] Figure 8 to Figure 9 The text and examples corresponding to the above provide a number of different methods, systems, devices, and computer-readable media of a context analysis system. In addition to the foregoing, one or more embodiments can also be described in terms of a flow diagram including acts for implementing a Figure 8 to Figure 9 as shown. Figure 8 to Figure 9 More or fewer acts can be performed. Also, acts can be performed in a different order. Additionally, acts described herein can be repeated or performed in parallel with one another, or with different instances of the same or similar acts.
[0080] As described above, Figure 8 shows a flow diagram of a series of acts for generating a context for a base model, in accordance with one or more embodiments. While Figure 8 shows acts in accordance with one embodiment, alternative embodiments can omit, add to, reorder, and / or modify Figure 8 any of the acts shown. Figure 8 The acts can be performed as part of a method. Alternatively, a computer-readable medium can include instructions that, when executed by one or more processors, cause a computing device to perform the acts. In some embodiments, a system can perform the acts. Figure 8 The acts. Figure 8 The acts.
[0081] At 856, the context analysis system can receive an input query for a base model. The input query can include a request for a response from the base model. At 858, the context analysis system can determine a first context profile for the input query. The first context profile can be based on a language of the input query and a context database. The first context profile can be generated and / or determined based on a first comparison of textual similarity between the input query and content of the context database. At 860, the context analysis system can generate and provide a first prompt as input to the base model. The first prompt includes a first string of text based on the input query and the first context profile. At 862, the context analysis system can determine a first relevance score for a first response of the base model in response to the first prompt. The first relevance score indicates a measure of confidence that the first response is relevant to the input query.
[0082] At 864, the context analysis system can determine a second context profile for the input query. The second context profile is based on the language of the input query and a context database. The second context profile is based on a second comparison of textual similarity between the input query and the content of the context database. At 866, the context analysis system can generate and provide a second prompt as input to the base model. The second prompt includes a second text string based on the input query and the second context profile. At 868, the context analysis system can determine a second relevance score for the base model's second response to the second prompt. The second relevance score indicates a confidence measure of the relevance of the second response to the input query.
[0083] In some embodiments, at 870, the context analysis system selects either a first response or a second response based on a first relevance score and a second relevance score. As discussed herein, the context analysis system may select a response based on a higher relevance score. In some embodiments, the context analysis system selects a response based on the lowest cost. In some embodiments, the context analysis system selects a response based on the ratio of relevance score to cost.
[0084] Although the two context profiles and two responses are described Figure 8 However, it should be understood that the context analysis system can prepare more than two context profiles and more than two associated responses. In some embodiments, the number of context profiles analyzed is based on the latency budget for the response to the query. In some embodiments, the number of context profiles analyzed is based on the number of context profiles required to reach a specific relevance threshold. The number of context profiles analyzed can include any number of context profiles (e.g., tens, hundreds).
[0085] As mentioned above, Figure 9 A flowchart illustrating a series of actions for generating a context for a base model, according to one or more embodiments, is shown. Although Figure 9 Actions according to one embodiment are shown, but alternative embodiments may omit, add, reorder, and / or modify them. Figure 9 Any action shown. Figure 9 The action can be performed as part of a method. Alternatively, the computer-readable medium may include actions that cause a computing device to perform when executed by one or more processors. Figure 9 In some embodiments, the system can execute instructions for action. Figure 9 The action.
[0086] At 976, the contextual analysis system can receive an input query for the base model. The input query can include a request for a response from the base model. At 978, the contextual analysis system can determine, for the input query, a plurality of context profiles based on a language of the input query. As discussed herein, the context profiles can be generated based on a comparison of text similarity between the input query and content of the context database. For example, the context profiles can be generated based on at least one of a text similarity metric, a vector embedding (e.g., a vector embedding of the query compared to vector representations of the context database), or a plug-in of the context database. At 980, the contextual analysis system can generate and provide a plurality of prompts as input to the base model. The plurality of prompts includes a string of text based on the input query and the plurality of context profiles.
[0087] At 982, the contextual analysis system can determine, for the base model, a plurality of relevance scores responsive to the plurality of prompts. The plurality of relevance scores indicate a measure of confidence that a response from the plurality of responses is relevant to the input query. At 984, the contextual analysis system can select a best response of the plurality of responses based on a comparison of the relevance score of the best response to an additional relevance score of the plurality of relevance scores.
[0088] As discussed herein, in some embodiments, the contextual analysis system can offload at least a portion of the methods discussed herein. For example, the contextual analysis system can offload extracting the plurality of context profiles and generating the plurality of prompts to an edge network. In some examples, the contextual analysis system can offload selection of the context profiles to the edge network. Figure 9
[0089] According to at least one embodiment of the present disclosure, the contextual analysis system can apply the plurality of prompts to a compressed base model. The compressed base model can be located on an edge network. In some embodiments, the contextual analysis system can generate a plurality of initial relevance scores for a plurality of initial responses. The contextual analysis system can select the best response by selecting a best initial response from the plurality of initial responses based on the plurality of initial relevance scores.
[0090] In some embodiments, the contextual analysis system can determine to offload extracting the plurality of context profiles and generating the plurality of prompts based on a processor condition of a user device providing the query. For example, if the user device does not have the processing capacity to extract the plurality of context profiles and / or generate the plurality of prompts within a particular latency budget, the contextual analysis system can determine to offload these processes.
[0091] In some embodiments, the contextual analysis system can extract the contextual profile based on the empirical relationship in the contextual database. For example, after receiving the query, the contextual analysis system can check the contextual database based on the query. The contextual analysis system can confirm the empirical relationship of the contextual profile related to the query in the contextual database to select and extract the contextual profile based on the empirical relationship.
[0092] In some embodiments, the number of the plurality of contextual profiles is based on a latency budget for the base model. In some embodiments, the contextual analysis system iteratively inputs the plurality of prompts into the base model when generating the plurality of prompts. In some embodiments, the contextual analysis system inputs the plurality of prompts into the base model in parallel.
[0093] Embodiments of the present disclosure can include or utilize special-purpose or general-purpose computers including computer hardware, such as, for example, one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specifically, one or more of the processes described herein can be implemented at least in part as instructions embodied in a non-transitory computer- readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium (e.g., a memory), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0094] A computer-readable medium can be any available medium or means that can be accessed by a general purpose or special purpose computing system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Accordingly, embodiments of the present disclosure can include at least two distinct types of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0095] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives ("SSDs") (e.g., based on RAM), Flash memory, phase-change memory ("PCM"), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0096] A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmission media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0097] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) or vice versa. For example, computer-executable instructions or data structures received by way of network or data link can be buffered in RAM within a network interface module (e.g., a "NIC"), and then eventually transferred to computer system RAM and / or to non-volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can include computer system memory or memory on other computer system components.
[0098] Computer-executable instructions include, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed by a general purpose computer to transform the general purpose computer into a special purpose computer that is uniquely configured to perform elements of the present disclosure. Computer- executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0099] Those skilled in the art will appreciate that the disclosure can be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure can also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0100] Embodiments of the disclosure can also be implemented in a cloud computing environment. As used herein, the term "cloud computing" refers generally to the use of a shared pool of configurable computing resources (e.g., networks, servers, storage, processes, applications, etc.) to perform tasks for multiple users. For example, cloud computing can be employed to provide ubiquitous and on-demand access to the shared pool of configurable computing resources via the Internet. The shared pool of configurable computing resources can be rapidly provisioned and released with minimal management effort or service provider interaction.
[0101] A cloud computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and the like. The cloud computing model can also expose various service models such as, for example, Software as a Service ("SaaS"), Platform as a Service ("PaaS"), and Infrastructure as a Service ("IaaS"). The cloud computing model can also be deployed using different deployment models such as, for example, private cloud, community cloud, public cloud, hybrid cloud, and the like. Additionally, as used herein, the term "cloud computing environment" refers to an environment in which cloud computing is employed.
[0102] Figure 10 Certain components can be included in the computer system 1000, as shown. One or more computer system 1000 can be used to implement various devices, components, and systems described herein.
[0103] The computer system 1000 includes a processor 1001. The processor 1001 can be a general purpose single- or multi-chip microprocessor (e.g., an Advanced RISC (Artificially RISCed Machine) (ARM)), a special purpose microprocessor (e.g., a digital signal processor (DSP)), a microcontroller, a programmable gate array, etc. The processor 1001 can be referred to as a central processing unit (CPU). Although Figure 10 Only a single processor 1001 is shown in the computer system 1000, but in an alternative configuration, a combination of processors (e.g., an ARM and a DSP) could be used.
[0104] The computer system 1000 also includes a memory 1003 in electronic communication with the processor 1001. The memory 1003 can be any electronic, magnetic, optical, or other physical storage device that can store electronic information. Examples of memory 1003 include random-access memory (RAM), read-only memory (ROM), magnetic disk storage mediums, optical storage mediums, flash memory devices in RAM, on-board memory included with the processor, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) memory, registers, and the like, including combinations of them.
[0105] Instructions 1005 and data 1007 can be stored in the memory 1003. The instructions 1005 can be executable by the processor 1001 to implement some or all of the functionality disclosed herein. Executing the instructions 1005 can involve the use of the data 1007 that is stored in the memory 1003. Any of the various examples of modules and components described herein can be implemented, partially or entirely, as instructions 1005 stored in memory 1003 and executed by processor 1001. Any of the various examples of data described herein can be among the data 1007 stored in memory 1003 and used during execution of the instructions 1005 by processor 1001.
[0106] The computer system 1000 can also include one or more communication interfaces 1009 for communicating with other electronic devices. The communication interface(s) 1009 can be based on wired communication technology, wireless communication technology, or both. Some examples of communication interfaces 1009 include a Universal Serial Bus (USB), an Ethernet adapter, a wireless adapter that operates in accordance with an Institute of Electrical and Electronics Engineers (IEEE) 802.11 wireless communication protocol, Bluetooth®, and an infrared (IR) communication port. Wireless communication adapters and an IR communication port.
[0107] The computer system 1000 can also include one or more input devices 1011 and one or more output devices 1013. Some examples of input devices 1011 include a keyboard, mouse, microphone, remote control device, button, joystick, trackball, touchpad, and lightpen. Some examples of output devices 1013 include speakers and a printer. One specific type of output device that is typically included in a computer system 1000 is a display device 1015. The display device 1015 used with the embodiments disclosed herein can utilize any suitable image projection technology, such as liquid crystal display (LCD), light-emitting diode (LED), gas plasma, electroluminescence, or the like. A display controller 1017 can also be provided, in order to convert data 1007 stored in the memory 1003 into text, graphics, and / or moving images (as appropriate) shown on the display device 1015.
[0108] The various components of the computer system 1000 can be coupled together by one or more busses, which can include a power bus, a control signal bus, a status signal bus, a data bus, etc. For clarity, the various buses are illustrated in Figure 10 FIG. 10 as bus system 1019.
[0109] In the foregoing specification, the disclosure has been described with reference to specific embodiments thereof. Various embodiments and aspects of the disclosure are discussed in connection with details of the foregoing description. The various embodiments and aspects of the disclosure have been presented for purposes of illustration and description only. The above description and drawings are illustrative of various embodiments of the disclosure and are not intended to be limiting thereof.
[0110] The disclosure can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein can be performed with fewer or additional steps / actions, or in different order. Additionally, steps / actions described herein can be repeated or performed in parallel with each other or with different instances of the same or similar steps / actions. The scope of the disclosure is, therefore, indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1. A method in a computing environment (100) including a base model (102) for processing queries, the processing being based on a context generated for the query, the method comprising: Receive an input query (336), the input query (336) including a request for a response from the base model (102); For the input query (336), a first context profile is determined based on the language and context database (115) of the input query, the first context profile being generated based on a first comparison of text similarity between the input query and the content of the context database (115); Generate and provide a first prompt as input to the base model (102), the first prompt comprising a first text string based on the input query and the first context profile; For the first response of the base model (102) to the first prompt, a first relevance score is determined, the first relevance score indicating the confidence measure of the first response’s relevance to the input query (336); For the input query (336), a second context profile is determined based on the language of the input query (336) and the context database (115), the second context profile being based on a second comparison of text similarity between the input query (336) and the content of the context database (115); A second prompt is generated and provided as input to the base model (102), the second prompt comprising a second text string based on the input query and the second context profile; For the second response of the base model (102) to the second prompt, a second relevance score is determined, the second relevance score indicating the confidence measure of the second response in relation to the input query (336); as well as Choose one of the first response or the second response based on the first correlation score or the second correlation score.
2. The method of claim 1, wherein the second context profile is generated based on the first relevance score being less than a threshold relevance score.
3. The method of claim 1, wherein the first context profile and the second context profile are generated in parallel and provided as input to the base model, and wherein the first relevance score and the second relevance score are compared when the first response and the second response are generated.
4. The method of claim 1, wherein selecting the first response or the second response comprises: The first response is selected based on the fact that the first response has a higher relevance score than the second response.
5. The method of claim 4, wherein selecting the first response is further based on a latency budget of less than a threshold associated with generating the first response.
6. The method according to claim 1, Determining the first correlation score includes: The first prompt is provided as input to the compressed base model, which is a compressed version of the base model. Determining the second relevance score includes providing the second hint as input to the compressed base model.
7. The method of claim 6, wherein the compression infrastructure model is implemented on an edge network, and wherein the infrastructure model is implemented in a data center of a cloud computing system.
8. The method of claim 1, wherein determining the first context profile and determining the second context profile are performed on a server device on an edge network in a fifth-generation (5G) telecommunications environment, and wherein the underlying model is implemented on a data center of a cloud computing system accessible via the edge network.
9. A method in a computing environment (100) including a base model (102) for processing queries, the processing being based on a context generated for the query, the method comprising: Receive an input query (336), the input query (336) including a request for a response from the base model (102); For the input query (336), a plurality of context profiles are determined based on the language and context database (115) of the input query (336), the plurality of context profiles being generated based on a comparison of text similarity between the content of the input query (336) and the content of the context database (115); Generate and provide multiple prompts as input to the base model (102), the multiple prompts including text strings based on the input query (336) and the multiple context profiles; For the base model (102) in response to the multiple prompts, multiple relevance scores are determined, the multiple relevance scores indicating the confidence measure of the relevance of the responses from the multiple responses to the input query (336); as well as Based on a comparison of the relevance score of the response with an additional relevance score among the plurality of relevance scores, a response is selected from the plurality of responses.
10. The method of claim 9, wherein the plurality of context profiles are generated based on different types of text similarity metrics.
11. The method of claim 9, wherein the plurality of context profiles are generated based on a combination of text similarity metrics, vector embeddings, and plugins associated with the context database.
12. The method of claim 9, wherein the selection of the response is based on a comparison of the correlation score with the additional correlation score and on a latency budget less than a threshold latency associated with generating a first response associated with the correlation score.
13. The method of claim 9, further comprising generating a set of context profiles, the set of context profiles including the plurality of context profiles and estimated relevance scores associated with corresponding context profiles in the plurality of context profiles, the estimated relevance scores being based on historical relevance score data, the historical relevance score data being based on generating and providing the plurality of prompts as input to the base model.
14. The method of claim 9, wherein determining the plurality of correlation scores comprises: The aforementioned prompts are provided as input to a compressed base model, which is a compressed version of the base model.
15. The method of claim 14, wherein the compression underlying model is implemented on an edge network, and wherein the underlying model is implemented in a data center of a cloud computing system.
16. The method of claim 9, wherein determining the plurality of context profiles is performed on a server device on an edge network in a fifth-generation (5G) telecommunications environment, and wherein the underlying model is implemented on a data center of a cloud computing system accessible via the edge network.
17. A system comprising: At least one processor; A memory that communicates electronically with the at least one processor; as well as Instructions stored in the memory, which can be executed by the at least one processor to: Receive an input query (336), the input query (336) including a request for a response from the base model (102); For the input query (336), a first context profile is determined based on the language and context database (115) of the input query, the first context profile being based on the input query and the context database. (115) is generated by a first comparison of textual similarity between the contents; Generate and provide a first prompt as input to the base model (102), the first prompt comprising a first text string based on the input query and the first context profile; For the first response of the base model (102) to the first prompt, a first relevance score is determined, the first relevance score indicating the confidence measure of the first response’s relevance to the input query (336); For the input query (336), a second context profile is determined based on the language of the input query (336) and the context database (115), the second context profile being based on a second comparison of text similarity between the input query (336) and the content of the context database (115); A second prompt is generated and provided as input to the base model (102), the second prompt comprising a second text string based on the input query and the second context profile; For the second response of the base model (102) to the second prompt, a second relevance score is determined, the second relevance score indicating the confidence measure of the second response in relation to the input query (336); as well as Choose one of the first response or the second response based on the first correlation score or the second correlation score.
18. The system of claim 17, wherein the second context profile is generated based on the first relevance score being less than a threshold relevance score.
19. The system of claim 17, wherein selecting the first response or the second response comprises: The first response is selected based on the first response having a higher relevance score than the second response, and the selection of the first response is also based on the latency associated with generating the first response being less than a threshold latency budget.
20. The system according to claim 17, Determining the first correlation score includes: The first prompt is provided as input to the compression base model. Determining the second relevance score includes providing the second hint as input to the compressed base model.