Method and apparatus for processing user request, and device and medium
By utilizing pre-verified reference prompt words and sample libraries, accurate prompt words and responses can be directly generated, solving the problem of content recognition in small-sample scenarios in machine learning and achieving more efficient and accurate model responses.
Patent Information
- Application Number
- PCT/CN2024/084226
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-10-02
AI Technical Summary
In machine learning, the few-sample problem makes it difficult for models to learn patterns. Existing methods such as meta-learning, transfer learning, data augmentation, and self-supervised learning are costly and complex, making it difficult to effectively solve the content recognition problem in few-sample scenarios in practical applications.
By leveraging pre-verified reference prompt words and sample libraries and directly reusing the contextual learning capabilities of the language model, accurate prompt words and responses can be generated, avoiding model training and improving content recognition capabilities in low-sample scenarios.
It reduces the complexity of model training and fine-tuning, improves the response accuracy and generalization ability of machine learning models in small sample scenarios, and adapts to dynamic changes and uncertainties.
Smart Images

Figure CN2024084226_02102025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and medium for processing user requests Technical Field
[0001] Exemplary implementations of the present disclosure generally relate to computer technology, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for processing user requests. Background Art
[0002] The few-shot problem refers to machine learning problems in which only a few samples are available for model training. For example, in business, only a few pieces of data are often collected for a certain category, making it difficult for a model to learn patterns. Because traditional machine learning algorithms, especially deep learning algorithms, typically require a large amount of labeled training data to learn patterns and extract features, the few-shot problem poses a challenge to building machine learning models. Therefore, it is desirable to address the few-shot scenario.
[0003] Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for processing a user request is provided. In this method, in response to receiving a user request, a reference prompt word matching the user request is determined; a set of reference samples matching the user request is determined, wherein the reference samples in the set of reference samples include a reference user request and a reference response to the reference user request, and the task type specified by the user request is the same as the reference task type specified by the reference user request; and based on the user request, the reference prompt word, and the set of reference samples, a prompt word for executing the user request is generated.
[0005] In a second aspect of the present disclosure, a device for processing a user request is provided. The device includes: a reference prompt word determination module configured to, in response to receiving a user request, determine a reference prompt word that matches the user request; a reference sample determination module configured to determine a set of reference samples that match the user request, wherein the reference samples in the set of reference samples include a reference user request and a reference response to the reference user request, and the task type specified by the user request is the same as the reference task type specified by the reference user request; and a prompt word generation module configured to generate a prompt word for executing the user request based on the user request, the reference prompt word, and the set of reference samples. The device also includes other modules configured to implement other steps in the above method.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] FIG1 illustrates a schematic diagram of an example environment in which implementations of the present disclosure can be implemented;
[0011] FIG2 shows a schematic diagram for processing a user request according to some implementations of the present disclosure;
[0012] FIG3 illustrates an example data structure of a reference sample library according to some implementations of the present disclosure;
[0013] FIG4 shows a schematic diagram of searching for a set of reference samples according to some implementations of the present disclosure;
[0014] FIG5 illustrates an example data structure of a reference prompt vocabulary according to some implementations of the present disclosure;
[0015] FIG6 shows a schematic diagram of generating a response based on a prompt word according to some implementations of the present disclosure;
[0016] FIG7 shows another schematic diagram of generating a response based on a prompt word according to some implementations of the present disclosure;
[0017] FIG8 shows a flowchart of a method for processing a user request according to some implementations of the present disclosure;
[0018] FIG9 shows a block diagram of an apparatus for processing a user request according to some implementations of the present disclosure; and
[0019] FIG10 illustrates a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION
[0020] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the implementation of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.
[0022] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0023] It is understandable that before using the technical solutions disclosed in each implementation of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0024] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0025] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0027] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.
[0028] Sample Environment
[0029] As mentioned above, in a few-sample scenario, it is difficult for the model to learn the patterns therein. For the sake of convenience, the environment of Figure 1 is used as an example to describe the few-sample problem in machine learning. Figure 1 shows a schematic diagram of an example environment 100 in which the implementation of the present disclosure can be implemented. In this example environment 100, a user 110 can input a user request 120 into a machine learning model 130. Based on the user request 120, the machine learning model 130 can return a response 140 to the user 110.
[0030] In some exemplary implementations, the user request 120 may include a request for text classification, object detection, speech recognition, etc. The machine learning model 130 returns a corresponding response 140 based on the type of the user request 120 .
[0031] In some exemplary implementations, the machine learning model 130 may be trained using samples of performing different tasks, and may perform a variety of tasks. In the following, a language model will be described as an example of the machine learning model 130. For example, the machine learning model 130 may perform Task 1 and Task 2 respectively based on different user requests received. Task 1 involves a large sample. During the training phase, the machine learning model 130 can extract features and learn rules based on a large amount of labeled data. Therefore, the machine learning model 130 performs better when performing Task 1. Task 2 involves a small sample. During the training phase, it is difficult for the machine learning model 130 to learn the patterns therein. Therefore, the performance of the machine learning model 130 when performing Task 2 is poor.
[0032] Traditionally, methods for solving few-shot learning include meta-learning, transfer learning, data augmentation, and self-supervised learning.
[0033] The idea behind meta-learning is to design algorithms that understand the learning process itself, allowing models to quickly and effectively adapt to new tasks. Typical meta-learning methods include Model-Agnostic Meta-Learning (MAML). While meta-learning can enable rapid adaptation to new tasks with minimal training examples, designing and implementing effective meta-learning algorithms can be complex. Some meta-learning algorithms require specialized training settings, which can be difficult to implement in practice.
[0034] Transfer learning is the process of applying knowledge learned in a related field (source field) to another different but related field (target field) to solve problems. When dealing with few-sample problems, this method first pre-trains a model on big data and then fine-tunes the model on the few-sample task. Although transfer learning can use information from the source task to improve the performance of the target task, if the correlation between the source task and the target task is not high, the effect of transfer learning may not be ideal. In addition, the pre-trained model may be overly complex and unable to achieve efficiency optimization on a specific task. And generally speaking, in few-sample scenarios, the correlation between the target task and the source task is difficult to predict in advance.
[0035] Data augmentation increases the amount of data by making small changes to the original data, such as rotating, scaling, or cropping images. While data augmentation improves model generalization performance by expanding the training dataset, it requires manual design and selection of appropriate data transformations, which requires specialized knowledge and can be time-consuming. Furthermore, in some cases, excessive data augmentation can introduce noise, negatively impacting model performance.
[0036] Self-supervised learning uses prediction tasks (such as predicting the next frame or word) to allow the model to self-generate labels for learning, thus avoiding reliance on large amounts of manually labeled data. While self-supervised learning can fully utilize unlabeled data and learn by self-proposing prediction tasks, it requires designing appropriate prediction tasks to drive the model to learn useful representations. This design process can require expertise and experience. Generally speaking, training models for different few-shot scenarios is required to adapt to these tasks, which takes a long time.
[0037] Based on the traditional approaches to solving the few-shot problem, we can see that most existing solutions require continued model training or elaborate data or model design, which is costly. For few-shot tasks and content recognition problems with highly variable and unpredictable data, adjusting the model for each new type of few-shot task is unacceptable.
[0038] In-context learning (ICL) is a learning method that refers to the acquisition of knowledge and skills in a specific environment or situation. This method emphasizes the acquisition and mastery of knowledge from practical situations that are closely related to the learner's actual life environment and experience. In the fields of artificial intelligence and machine learning, ICL focuses on enabling models to understand and use contextual information. This is achieved by allowing the model to pay attention to the environment, background, or other related information of the input information when processing it. For example, in natural language processing, the meaning of various words and phrases may be affected by the context. ICL enables the model to understand and adapt to different contexts, thereby more accurately understanding and predicting language.
[0039] "Interactive Closed Learning (ICL)" refers to the ability to learn from experience and practice in specific contexts or environments. For example, language models, because they have more parameters, they possess greater tolerance and complexity, which are required to understand and handle complex tasks. This includes understanding and utilizing contextual information, or the ability to achieve ICL. Specifically, the following factors contribute to a model's ICL capabilities.
[0040] Deeper network layers give models the ability to learn and generalize complex, abstract features and patterns. Language models are composed of more neural network layers, which enables them to learn and represent more complex and abstract features and patterns. This ability enables them to understand and use more complex contextual information. More parameters also give models the ability to learn and generalize complex, abstract, and abstract data. With more parameters, language models can learn more diverse data representations. When processing contextual data, this allows them to learn how to use context to better understand the task. In addition, adaptability to small sample sizes also gives models the ability to learn and generalize complex, abstract, and abstract data. Due to their depth and breadth, language models can efficiently learn and generalize from small samples, which is key to contextual learning. In this case, the hope is to leverage the model's ICL capabilities to provide more accurate responses.
[0041] Overview of solving the few-shot problem
[0042] To at least partially address the shortcomings of the existing technology, a method for processing user requests is proposed according to an exemplary implementation of the present disclosure. Based on the above considerations, the present disclosure proposes directly reusing the learning capabilities of the language model (ICL) to transform the conventional content recognition problem into a system module for retrieval, context learning, and recognition, thereby solving the problem of few-sample content recognition.
[0043] 2 , which illustrates a schematic diagram 200 for processing a user request according to some exemplary implementations of the present disclosure. As shown in FIG2 , in response to receiving a user request 210, a reference prompt word 232 matching the user request 210 is determined. A set of reference samples 222 matching the user request 210 is determined. The reference samples in the set of reference samples 222 include a reference user request and a reference response to the reference user request. The task type specified by the user request 210 is the same as the reference task type specified by the reference user request. Based on the user request 210, the reference prompt word 232, and the set of reference samples 222, a prompt word 240 for executing the user request 210 is generated.
[0044] Using the exemplary implementations of this disclosure, previously verified accurate and reliable reference prompts and reference samples can be directly utilized. Without requiring model training, by determining a reference prompt and a set of reference samples that match the user request, a prompt word can be generated to execute the user request. Furthermore, more accurate responses can be obtained based on the prompt word, thus reducing the complexity of training and fine-tuning the machine learning model and enabling the machine learning model to output more accurate responses.
[0045] Detailed process of solving the few-shot problem
[0046] In some exemplary implementations, the technical solution according to an example implementation of the present disclosure can be called only in a few-sample scenario. Specifically, in response to determining that the number of samples associated with the task type is less than a predetermined threshold, the method for processing user requests proposed by the exemplary implementation of the present disclosure is executed. When performing tasks related to few samples, since it is difficult for the machine learning model to learn patterns in a few samples (for example, less than 10 or other amounts of data), this method can be used to solve the few-sample problem. The predetermined threshold here can be 5, 10, etc., and this application does not limit this. Using the exemplary implementation of the present disclosure, this method can be executed when performing tasks related to small samples, thereby improving the processing capability of the machine learning model by generating more accurate prompt words.
[0047] Continuing with reference to FIG2 , in some exemplary implementations, a reference prompt word 232 can be selected from a reference prompt word library 230 comprising a plurality of sample prompt words. The reference prompt word library 230 is generated based on prompt words associated with the task type. The prompt words associated with the task type here are verified to be correct and effective. For different few-sample tasks, the corresponding reference prompt words can be the same or different. Generally speaking, an efficient prompt word can be designed for each type of task to achieve the best problem-solving effect on that type of task.
[0048] The configuration of prompt words can, for example, include prompt word screening and verification steps. During this step, since multiple prompt words may exist for each task type, manual screening and verification are required to confirm the prompt word template used for each task type. For example, each reference prompt word can correspond to a task type, which can include, for example, identifying the type of school referenced in a text (e.g., elementary school, middle school, university, etc.) or identifying the type of object in an image.
[0049] In some exemplary implementations, a set of reference samples 222 can be determined from a reference sample library 220 including a plurality of reference samples, and the reference sample library 220 is generated based on samples associated with a task type. The samples associated with the task type here are samples that have been verified to be correct and valid. The reference sample library 220 will be described below with reference to FIG3 , which shows an example data structure 300 of the reference sample library 220 according to some implementations of the present disclosure. As shown in FIG3 , the example data structure 300 may include a task type 310 and reference samples 320. Exemplarily, the task type 310 may include types such as text classification, image classification, and speech classification. The reference samples 320 corresponding to the text classification task may include samples that classify different texts into different school types, for example, the following classifications may be included:
[0050] “1.Education-Elementary:some children are learning words and pronunciation…”. “2.Education-Middle:students are conducting experiment in lab to obtain oxygen by heating potassium permanganate…”. “3.Education-College:freshmen are so excited when they come into their dreamed university…”.
[0051] It should be understood that although the specific task of text classification is described above using English as an example of a natural language, alternatively and / or additionally, the text may be written in other languages such as Chinese, French, Japanese, etc. In some exemplary implementations, the reference samples 320 corresponding to the image classification task may include samples that classify different images as different animal types, for example, some images may be classified as dogs, and other images may be classified as cats, etc.
[0052] The purpose of the reference sample library 220 is to provide online content management and retrieval, maintaining the library's size within a manageable range and including as many key tasks as possible. In some exemplary implementations, when designing the reference sample library 220, key tasks can be selected first, that is, the key tasks to be learned by the machine learning model. These tasks should be practical problems that the model will encounter in future processing. These tasks can be customized or automatically stored in the library.
[0053] Next, for each key task, you can collect some samples that represent that task. These samples should contain enough information to solve the task, but the number should be much smaller than traditional large-scale datasets. You can create these samples yourself or obtain them from online tasks after manual review. Then, according to the requirements of the input model, each sample is converted into an appropriate form. For example, if you need to identify the type of school, you need to include at least the text and the corresponding category as basic information, and the text needs to be properly preprocessed to make it consistent with the input model format.
[0054] Using the exemplary implementation of the present disclosure, without requiring model training, more accurate prompt words 240 are generated by selecting reference prompt words 232 and a set of reference samples 222 from a verified reference prompt word library 230 and a reference sample library 220, thereby more efficiently recognizing content. Furthermore, the machine learning model can better identify unseen data or tasks based on samples, improving the model's generalization capabilities.
[0055] Searching for a set of reference samples 222 from the reference sample library 220 will be described below with reference to FIG. FIG. 4 illustrates a schematic diagram 400 of searching for a set of reference samples 222 according to some implementations of the present disclosure. In some exemplary implementations, as shown in FIG. 4 , a feature representation 410 of the user request 210 may be obtained and, using this feature representation 410, the reference sample library 220 may be searched for a set of reference samples 222 that match the feature representation 410. For example, the feature representation 410 of the user request 210 may be an embedding of the user request 210. After obtaining the feature representation 410, the distances between the feature representation 410 and the feature representations of multiple reference samples in the reference sample library 220 may be determined. These distances are then sorted to determine the k reference samples with the smallest distances (i.e., the top k most similar samples). Alternatively and / or additionally, an index of the reference sample library 220 may be utilized to accelerate the search for the set of reference samples 222. By using the exemplary implementation of the present disclosure, reference samples similar to the feature representation of the user request 210 can be determined from the reference sample library 220 , thereby generating more accurate prompt words 240 .
[0056] In some exemplary implementations, the index 420 of the reference sample library 220 can be used to determine a group of reference samples 222, and the index 420 is created based on multiple reference samples in the reference sample library 220. An index system can be created based on the content in the library and the corresponding embedding, so that the corresponding small number of samples (i.e., a group of reference samples 222) can be quickly found during the retrieval process. The retrieval method here can be based on any retrieval method, not limited to retrieval by converting to an embedding method, and the retrieval method is not limited to Faiss, Annoy, NMSLIB, Scikit-learn's nearest neighbor algorithm, and BallTree and KDTree in SciPy, etc. By using the exemplary implementation of the present disclosure, the retrieval process can be accelerated and a group of reference samples 222 can be quickly determined in the reference sample library 220. By making full use of the reference sample library 220 and allowing the model to perform contextual learning based on retrieval, the content recognition rate can be improved.
[0057] In some exemplary implementations, a set of reference samples can be used to update the reference prompt words. The prompt word configuration may also include prompt word configuration and assembly steps. In this step, after the prompt words for each type of task are designed, they need to be assembled with the retrieved few samples. The following will describe the use of a set of reference samples to update the reference prompt words with reference to Figure 5. Figure 5 shows an example data structure 500 of the reference prompt word library 230 according to some implementations of the present disclosure. In the example data structure 500, a set of reference samples can be used to update the reference prompt words 520. In one example, a set of reference samples for the text classification task in task type 510 is to classify different texts into different school types, then the reference prompt words 520 can be updated as follows:
[0058] “We want to classify some texts into these categories (Education-Elementary, Education-Middle, Education-College). Here are some examples:…Please classify the following texts into these categories based on given examples”.
[0059] In another example, for a set of reference samples of an image classification task in task type 510 , different images are classified as different animals. Then, the reference prompt words 520 may be updated as follows:
[0060] “We want to classify some images into these categories (cat,dog,...). Here are some examples:…Please classify the following images into these categories based on given examples”.
[0061] In some exemplary implementations, updated reference prompts and user requests can be combined to generate prompts. The following describes combining updated reference prompts and user requests to generate prompts, with reference to FIG6 . FIG6 shows a schematic diagram 600 of generating responses based on prompts according to some implementations of the present disclosure. As shown in FIG6 , user request 210, reference prompt 232, and a set of reference samples 222 can be combined to generate prompt 240. In this case, user request 210 corresponds to portion 616 of prompt 240, reference prompt 232 corresponds to portions 610 and 614 of prompt 240, and a set of reference samples 222 corresponds to portion 612 of prompt 240. By combining updated reference prompts and user requests into prompts using exemplary implementations of the present disclosure, new environments and tasks can be dynamically adapted, thereby improving the ability to handle dynamic changes and uncertainties in real-world problems.
[0062] In some exemplary implementations, a language processing model can be used to generate a response to a user request based on a prompt word. The language processing model here is a model with contextual learning capabilities and is not limited to various language models. For example, it can include multiple language models known in the past and / or to be developed in the future. Continuing with reference to FIG6 , the combined prompt word 240 is input into the language processing model, and the language processing model can output a response 620. In the example of FIG6 , the user request 210 is “some children are learning basic words and pronunciation, they are so cute”, and the response 620 to the user request 210 is “Education-Elementary”.
[0063] FIG7 shows another schematic diagram 700 of generating a response based on a prompt word according to some implementations of the present disclosure. As shown in FIG7 , in prompt word 710, the portion 716 corresponding to the user request is "students are conducting experiments in lab to test some methods." The response 720 to this user request is "Education-Middle." Using exemplary implementations of the present disclosure, the language processing model generates more accurate responses, better satisfying user requests.
[0064] In some exemplary implementations, a sample is created based on a user request and response, including both the user request and response, and added to a reference sample library. By using the exemplary implementations of the present disclosure, adding the created sample to the reference sample library can increase the task diversity of the reference sample library, keeping its content up to date and continuously capturing and covering key online tasks and patterns.
[0065] In some exemplary implementations, the sample can be used to update the index of the reference sample library. Using the exemplary implementation of the present disclosure, by updating the index of the reference sample library, it can be ensured that the index matches the updated reference sample library, thereby performing searches at a faster speed.
[0066] In some exemplary implementations, the language processing model supports multimodal processing and contextual learning. Multimodal processing includes processing at least one of the following: text, images, audio, and video. For example, the language processing model can support text-to-text (Text2Text, i.e., input text, generate text), image-to-text (Image2Text, i.e., input text and image, generate text), and text-to-image (Text2Image, i.e., input text, generate image). Using the exemplary implementations of the present disclosure, a language processing model with contextual learning capabilities can support the processing of multimodal data, improving the generalization of the language processing model.
[0067] Example Process
[0068] FIG8 illustrates a flow chart of a method 800 for processing a user request according to some implementations of the present disclosure. At block 810, in response to receiving a user request, a reference prompt word matching the user request is determined. At block 820, a set of reference samples matching the user request is determined, wherein the reference samples in the set of reference samples include a reference user request and a reference response to the reference user request, wherein the task type specified by the user request is the same as the reference task type specified by the reference user request. At block 830, a prompt word for executing the user request is generated based on the user request, the reference prompt word, and the set of reference samples.
[0069] In some exemplary implementations, the method 800 further includes: executing the method in response to determining that the number of samples associated with the task type is less than a predetermined threshold.
[0070] In some exemplary implementations, determining a reference prompt word includes: selecting a reference prompt word from a reference prompt word library including multiple sample prompt words, the reference prompt word library being generated based on prompt words associated with the task type; and determining a set of reference samples includes: determining a set of reference samples from a reference sample library including multiple reference samples, the reference sample library being generated based on samples associated with the task type.
[0071] In some exemplary implementations, determining a set of reference samples includes: obtaining a feature representation of a user request; and searching a reference sample library for a set of reference samples that matches the feature representation using the feature representation.
[0072] In some exemplary implementations, searching for a set of reference samples in the reference sample library includes determining a set of reference samples using an index of the reference sample library, where the index is created based on a plurality of reference samples in the reference sample library.
[0073] In some exemplary implementations, generating the prompt word includes: updating a reference prompt word using a set of reference samples; and combining the updated reference prompt word with the user request to generate the prompt word.
[0074] In some exemplary implementations, method 800 further includes: utilizing a language processing model to generate a response to the user request based on the prompt word.
[0075] In some exemplary implementations, the method 800 further includes: creating a sample based on the user request and the response, the sample including the user request and the response; and adding the sample to a reference sample library.
[0076] In some exemplary implementations, the method 800 further includes: updating an index of a reference sample library using the sample.
[0077] In some exemplary implementations, the language processing model supports multimodal processing and contextual learning, and the multimodal processing includes processing of at least any of the following: text, image, audio, and video.
[0078] Example devices and equipment
[0079] FIG9 shows a block diagram of an apparatus 900 for processing a user request according to some implementations of the present disclosure. The apparatus 900 includes: a reference prompt word determination module 910 configured to, in response to receiving a user request, determine a reference prompt word that matches the user request; a reference sample determination module 920 configured to determine a set of reference samples that match the user request, wherein the reference samples in the set of reference samples include a reference user request and a reference response to the reference user request, wherein the task type specified by the user request is the same as the reference task type specified by the reference user request; and a prompt word generation module 930 configured to generate a prompt word for executing the user request based on the user request, the reference prompt word, and the set of reference samples.
[0080] In some exemplary implementations, apparatus 900 is invoked in response to determining that a number of samples associated with a task type is less than a predetermined threshold.
[0081] In some exemplary implementations, the reference prompt word determination module 910 further includes a reference prompt word selection module, which is configured to select reference prompt words from a reference prompt word library including multiple sample prompt words, and the reference prompt word library is generated based on prompt words associated with the task type; and the reference sample determination module 920 further includes a reference sample selection module, which is configured to determine a set of reference samples from a reference sample library including multiple reference samples, and the reference sample library is generated based on samples associated with the task type.
[0082] In some exemplary implementations, the reference sample determination module 920 further includes an index utilization module configured to utilize an index of the reference sample library to determine a set of reference samples, where the index is created based on multiple reference samples in the reference sample library.
[0083] In some exemplary implementations, the prompt word generation module 930 further includes a combining module configured to update the reference prompt words using a set of reference samples; and combine the updated reference prompt words with the user request to generate the prompt words.
[0084] In some exemplary implementations, the apparatus 900 further includes a response generation module configured to generate a response to the user request based on the prompt word using a language processing model.
[0085] In some exemplary implementations, the apparatus 900 further includes a sample adding module configured to create a sample based on a user request and a response, the sample including the user request and the response; and add the sample to a reference sample library.
[0086] In some exemplary implementations, the apparatus 900 further includes an index updating module configured to update an index of the reference sample library using the sample.
[0087] In some exemplary implementations, the language processing model supports multimodal processing and contextual learning, and the multimodal processing includes processing of at least any of the following: text, image, audio, and video.
[0088] FIG10 shows a block diagram of a device 1000 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1000 shown in FIG10 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. The computing device 1000 shown in FIG10 can be used to implement the methods described above.
[0089] As shown in FIG10 , computing device 1000 is in the form of a general-purpose computing device. Components of computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processing unit 1010 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1000.
[0090] The computing device 1000 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1020 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1030 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1000.
[0091] The computing device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0092] The communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1000 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0093] Input device 1050 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1060 may be one or more output devices, such as a display, a speaker, or a printer. Computing device 1000 may also communicate with one or more external devices (not shown) via communication unit 1040 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with computing device 1000, or with any device that allows computing device 1000 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0094] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0095] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0096] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0097] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0098] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0099] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for processing a user request, comprising: In response to receiving the user request, determining a reference prompt word that matches the user request; Determining a set of reference samples that match the user request, where the reference samples in the set of reference samples include a reference user request and a reference response to the reference user request, and the task type specified by the user request is the same as the reference task type specified by the reference user request; as well as Based on the user request, the reference prompt word, and the set of reference samples, a prompt word for executing the user request is generated.
2. The method according to claim 1, further comprising: In response to determining that the number of samples associated with the task type is less than a predetermined threshold, the method is performed.
3. The method according to claim 1, wherein determining the reference prompt word comprises: Selecting the reference prompt word from a reference prompt word library comprising a plurality of sample prompt words, wherein the reference prompt word library is generated based on prompt words associated with the task type; as well as Determining the set of reference samples includes: determining the set of reference samples from a reference sample library comprising a plurality of reference samples, wherein the reference sample library is generated based on samples associated with the task type.
4. The method of claim 3, wherein determining the set of reference samples comprises: Obtaining a feature representation of the user request; as well as The feature representation is used to search the reference sample library for the set of reference samples that match the feature representation.
5. The method according to claim 4, wherein searching the reference sample library for the set of reference samples comprises: The group of reference samples is determined by using an index of the reference sample library, where the index is created based on the plurality of reference samples in the reference sample library.
6. The method according to claim 1, wherein generating the prompt word comprises: Updating the reference prompt words using the set of reference samples; as well as The updated reference prompt word and the user request are combined to generate the prompt word.
7. The method according to claim 5, further comprising: A language processing model is used to generate a response to the user request based on the prompt word.
8. The method according to claim 7, further comprising: creating a sample based on the user request and the response, the sample including the user request and the response; as well as Add the sample to the reference sample library.
9. The method according to claim 8, further comprising: The index of the reference sample library is updated using the sample.
10. The method according to claim 7, wherein the language processing model supports multimodal processing and contextual learning, and the multimodal processing includes processing for at least any one of the following: text, image, audio, video.
11. A device for processing a user request, comprising: a reference prompt word determination module, configured to determine a reference prompt word matching the user request in response to receiving the user request; a reference sample determination module configured to determine a set of reference samples matching the user request, wherein the reference samples in the set of reference samples include a reference user request and a reference response to the reference user request, and the task type specified by the user request is the same as the reference task type specified by the reference user request; as well as The prompt word generation module is configured to generate a prompt word for executing the user request based on the user request, the reference prompt word, and the set of reference samples.
12. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processor The electronic device comprises a processing unit and stores instructions for execution by the at least one processing unit, wherein when the instructions are executed by the at least one processing unit, the electronic device executes the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Question answering method and device, electronic equipment and computer readable storage medium
CN116244418A
Intelligent question and answer method and device, equipment and storage medium
CN117493505A
Method and device for training cue word generation model, equipment and medium
CN117557885A
Label value determination method and device, equipment and storage medium
CN117574286A