Reply method, apparatus and system based on generative language models
Through the probability weighted combination of multiple independent generative language models, the "illusion" problem of generative language models is solved, efficient and accurate reply generation is achieved, and training costs are reduced.
Patent Information
- Application Number
- PCT/CN2024/117409
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2024-09-06
- Publication Date
- 2025-07-24
AI Technical Summary
Existing generative language models are prone to "illusion" problems during the generation process. Existing solutions are either unable to effectively alleviate them, or they are costly and have uncertain training effects.
Multiple independent generative language models are used to multiply the generation probability of the candidate output token of each model with the preset weight to obtain the target generation probability, and the token with the highest target generation probability is used as the final output token, and the reply text is generated based on the output of multiple models.
There is no need to further train each generative language model, reducing costs, and even if one model outputs an error, other models can correct it, significantly improving the accuracy of the response.
Smart Images

Figure CN2024117409_24072025_PF_FP_ABST
Abstract
Description
Reply method, device and system based on generative language model
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 15, 2024, with application number 202410053208.3. The entire contents of the above application are incorporated by reference into this application. Technical Field
[0002] The present application relates to an automatic reply system, and in particular, to a reply method, device and system based on a generative language model. Background Art
[0003] Currently, generative large language models, represented by GPT, have become mainstream in the field of artificial intelligence, particularly natural language processing. However, due to the probabilistic nature of their generation process, generative large language models inevitably suffer from "hallucinations." For example, for the question "Where will the 2020 Olympics be held?", the model might output "The 2020 Olympics will be held in Doha." The probability of "Doha" in this case is higher than the correct answer, "Tokyo," when the large model generated the token.
[0004] This phenomenon is one manifestation of the "hallucination" problem of large models. There are generally two solutions to this problem. One is to reduce the randomness of generation by adjusting the model generation parameters. However, this only addresses a small number of generation errors caused by random sampling. If the token with the highest probability predicted by the model at the current position is an incorrect answer, this solution will fail. The other is to increase the number of model parameters and the amount of pre-training data. Current researchers generally believe that the knowledge and capabilities of large language models come from pre-training and are proportional to the number of parameters and pre-training data. Therefore, by increasing the number of model parameters and the amount of pre-training data during the pre-training phase, the model's knowledge and capabilities can be improved, thereby fundamentally alleviating the "hallucination" problem. However, this solution requires extremely high upfront training costs, and the effectiveness cannot be predicted in advance.
[0005] Summary of the Invention
[0006] In order to overcome the shortcomings of the existing technology, the present application provides a reply method, device and system based on a generative language model to solve the problems that the existing methods for alleviating the "hallucination" problem of the generative language model are either unable to effectively alleviate it or have high training costs and uncertain training effects.
[0007] The technical solution adopted by this application to solve its technical problems is:
[0008] In a first aspect, a reply method based on a generative language model is provided, wherein the generative language model includes multiple generative language models, and any two generative language models are independent of each other. The method includes:
[0009] Determining, based on the instruction text input by the user, a candidate output token predicted by each generative language model and a generation probability of the candidate output token;
[0010] Multiplying the generation probability of the candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the candidate output token;
[0011] The candidate output token with the highest target generation probability is used as the target output token;
[0012] The reply text is obtained based on the target output token.
[0013] Furthermore, the determining of the candidate output tokens of each generative language model based on the instruction text input by the user includes:
[0014] Get the command text entered by the user;
[0015] Encode the instruction text into an input sequence of tokens using the tokenizer corresponding to each generative language model;
[0016] Perform a forward propagation in the corresponding generative language model based on the input sequence to obtain multiple predicted tokens;
[0017] The token with the highest probability among the multiple predicted token probability distributions is used as the candidate output token of the generative language model.
[0018] Furthermore, obtaining a reply text based on the target output token includes:
[0019] Decoding the target output token using a tokenizer in a generative language model corresponding to the target output token to obtain a target string;
[0020] respectively using each generative language model to encode the target character string, and adding the target output token to the input sequence of the corresponding generative language model;
[0021] Determining, based on the input sequence, a next candidate output token predicted by each generative language model and a generation probability of the next candidate output token;
[0022] Multiplying the generation probability of the next candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the next candidate output token;
[0023] The next candidate output token with the highest target generation probability is used as the next target output token;
[0024] If the currently generated target output token meets the preset stop condition, the text corresponding to the target output token sequence when the stop condition is met will be used as the reply text.
[0025] Furthermore, if the currently generated target output token does not meet the preset stop condition, the step of obtaining the next target output token is repeated.
[0026] Furthermore, the preset stopping condition is: the generated target output token is the stop mark of the corresponding generative language model.
[0027] Furthermore, the preset stopping condition is: the target output token sequence length reaches the maximum generation length.
[0028] Furthermore, it also includes:
[0029] Determining a classification of the instruction text based on the instruction text input by the user;
[0030] A preset weight for each generative language model is determined based on the classification.
[0031] Furthermore, each generative language model is loaded on a GPU, and any two generative language models are loaded on different GPUs.
[0032] In a second aspect, a reply device based on a generative language model is provided, wherein the generative language model includes a plurality of generative language models, and any two generative language models are independent of each other, and the device includes:
[0033] A candidate output token determination module is used to determine the candidate output token predicted by each generative language model and the generation probability of the candidate output token based on the instruction text input by the user;
[0034] A target generation probability calculation module is used to multiply the generation probability of the candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the candidate output token;
[0035] A target output token determination module is used to select the candidate output token with the highest target generation probability as the target output token;
[0036] The reply text acquisition module is used to obtain the reply text based on the target output token.
[0037] In a third aspect, a response system based on a generative language model is provided, comprising:
[0038] at least one processor and at least one memory;
[0039] The memory stores executable instructions of the processor;
[0040] The processor is configured to execute the above method. Beneficial effects:
[0041] The technical solution of the present application provides a reply method, device and system based on a generative language model. After the user inputs the instruction text, multiple generative language models are used to predict the next token respectively, and the token with the highest generation probability is used as the candidate output token. Then, the product of the generation probability of each candidate output token and the preset weight of the corresponding generative language model is used as the final target generation probability, and the candidate input token with the highest target generation probability is used as the target output token. That is, a target output token is obtained by combining the advantages of multiple generative language models, and the text corresponding to the target output token is used as the output of all generative language models. In this way, each token in the final reply text is obtained by combining the outputs of all generative language models. The present application solution does not require further training of each generative language model, which greatly reduces the cost. Even if the candidate output token of a generative language model is incorrect, it can be eliminated in combination with the candidate output tokens of other generative language models, which greatly increases the accuracy of the final reply. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0043] FIG1 is a flow chart of a reply method based on a generative language model provided in an embodiment of the present application;
[0044] FIG2 is a flowchart of a specific reply method based on a generative language model provided in an embodiment of the present application;
[0045] FIG3 is a schematic diagram of the structure of a reply device based on a generative language model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application are described in detail below with reference to the accompanying drawings and examples. Obviously, the described embodiments are only some of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other implementation methods obtained by ordinary technicians in this field without making any creative work are within the scope of protection of this application.
[0047] First, it's important to note that in large language models, tokens are the fundamental building blocks of text processing. They represent a discrete element in a text, such as a word, character, subword, or character. Tokens break text into its smallest processable units for subsequent text analysis and application.
[0048] In a large language model, tokens have the following characteristics and functions:
[0049] Model input and output: The input and output of large language models are processed in token units. Both input text and generated text are represented and analyzed using tokens.
[0050] In addition, the generative language model generates responses by tokenizing the user input text based on a tokenizer (called a word segmenter in Chinese, which breaks sentences into small chunks (tokens) based on a preset vocabulary). Based on the input token sequence, it predicts one or more output tokens and uses the one with the highest probability as the final input token. The next token is then predicted based on the input token sequence and the generated output tokens, and the process continues until it stops. The output token sequence at the tokenizer's stop point is then converted into text.
[0051] However, single-source generative language models currently suffer from inaccurate output. This is because either the model has too few parameters, in which case the output is incorrect regardless of training or parameter adjustment. Alternatively, the model has been trained sparingly, but the training cost is too high to achieve accurate output.
[0052] Based on the above problems, the present application provides an output method based on multiple generative language models, wherein the generative language models include multiple, and any two generative language models are independent of each other, that is, each generative language model is different. Referring to Figure 1, the method includes:
[0053] S11: Based on the instruction text input by the user, determine the candidate output token predicted by each generative language model and the generation probability of the candidate output token;
[0054] Specifically, the method involves obtaining a command text input by a user; encoding the command text into a token input sequence using a tokenizer corresponding to each generative language model; performing a forward propagation in the corresponding generative language model based on the input sequence to obtain multiple predicted tokens; and selecting the one with the highest probability among the multiple predicted tokens as a candidate output token of the generative language model.
[0055] S12: Multiply the generation probability of the candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the candidate output token.
[0056] S13: Taking the candidate output token with the highest target generation probability as the target output token.
[0057] S14: Obtain a reply text based on the target output token.
[0058] Specifically, the target output token is decoded by the tokenizer in the generative language model corresponding to the target output token to obtain a target string; the target output token of each generative language model that encodes the target string is respectively obtained by using each generative language model, and the target output token is added to the input sequence of the corresponding generative language model; because the target output token obtained at this time is only a recognizable token of a generative language model, and due to different vocabularies, other generative language models may not be able to recognize the target output token, so the target output token needs to be decoded into text and then encoded into a recognizable target output token according to the vocabulary of each generative language model.
[0059] Based on the input sequence, the next candidate output token predicted by each generative language model and the generation probability of the next candidate output token are determined; the generation probability of the next candidate output token is multiplied by the preset weight of the corresponding generative language model to obtain a target generation probability of the next candidate output token; the next candidate output token with the highest target generation probability is used as the next target output token; if the currently generated target output token meets a preset stopping condition, the text corresponding to the target output token sequence when the stopping condition is met is used as the reply text. If the currently generated target output token does not meet the preset stopping condition, the step of obtaining the next target output token is repeated.
[0060] The preset stopping condition is: the generated target output token is the stop mark of the corresponding generative language model. Alternatively, the stopping condition is: the length of the target output token sequence reaches the maximum generated length. It should be noted that after obtaining the next target output token, the next target output token needs to be decoded into text and then encoded using the vocabulary of other generative language models. The decision to stop is then made based on the encoded token.
[0061] Repeat the steps to obtain the next target output token. Specifically, add the newly obtained target output token to the current input sequence, then obtain the candidate output tokens and corresponding generation probabilities of each generative language model based on the current input sequence, and obtain the next target output token based on the candidate output tokens and corresponding generation probabilities.
[0062] As an optional implementation of the present invention, each generative language model is loaded on a GPU, and any two generative language models are loaded on different GPUs. That is, each generative language model is loaded on a separate GPU. This allows for parallel processing when processing data, significantly improving processing speed.
[0063] In one embodiment, the preset weight of each generative language model in this application is set to a fixed value based on the user's needs. However, because each generative language model has different focuses, its accuracy in responding to different types of questions varies. For example, a generative language model may have an accuracy of 90% when responding to questions of type A, but only 70% when responding to questions of type B. Therefore, if a fixed value is used, it cannot adapt to all situations.
[0064] Therefore, in another embodiment of the present application, the category of the instruction text input by the user is determined, and the preset weight of each generative language model is determined based on the category. That is, each generative language model uses different preset weights when faced with instruction texts of different categories. This can improve the final accuracy when faced with texts of different categories.
[0065] To more clearly illustrate the above solution, the present application provides a specific reply method based on a generative language model, including the following steps:
[0066] 1. Model Parallel Loading: Multiple generative language models involved in the combination are independently loaded on different GPUs to fully utilize the high performance of multi-GPU parallel computing.
[0067] 2. Candidate output token generation: For the same instruction text input by the user, all generative language models in step 1 encode it into a token sequence in parallel using the corresponding tokenizer. A forward propagation is performed based on this sequence, and the token with the highest probability in the predicted token probability distribution is selected as the candidate output token.
[0068] 3. Determine the target output token: For each candidate output token of the generative language model in step 2, multiply its generation probability by the preset weight coefficient of each model. Compare the weighted generation probabilities and take the one with the highest probability as the target output token for this round.
[0069] 4. Token alignment: Because different generative language models have different vocabularies and cannot share tokens, the target output token in step 3 needs to be decoded into the actual representative string by its corresponding generative language model tokenizer, and then encoded using all the generative language models in step 1 and added to the token sequence in step 2.
[0070] 5. Iterative generation: Repeat steps 2 to 4 until the generated target output token is the stop mark of its corresponding generative language model (such as EOS, etc.), or the maximum generation length is reached.
[0071] 6. Output result: At this time, the token sequence generated by decoding any generative language model will correspond to the same text, and the output of this text is the final output result of the user input instruction in step S22.
[0072] The specific implementation steps are shown in Figure 2:
[0073] 1. Initialize each of the multiple GLMs involved in the combination. To fully utilize parallel processing, load each GLM onto a different GPU. Also, set a weight coefficient w for each GLM a priori, with the default value being 1.0.
[0074] 2. Input the user's input instruction text I into each model. According to the general processing flow of the generative language model, for each generative language model m, the input instruction text I is first encoded into a sequence of tokens S through its own tokenizer. m .
[0075] 3. Starting from the round of generating the first token, for each generative language model m, based on the current token sequence S m , the candidate token is generated by taking the greedy method with the maximum probability and recorded as T m , and its corresponding probability is recorded as P m ;
[0076] 4. For each generative language model m, the probability P m Multiply by the pre-defined weight coefficient w of the generative language model m m get
[0077] 5. Compare all models after weight adjustment The token corresponding to the largest value is selected as the preferred token for this round, denoted as T.
[0078] 6. Decode the preferred token T into the string s it actually represents through the tokenizer of model m;
[0079] 7. All models re-encode the string s and add the encoded tokens to their respective sequences S m middle;
[0080] 8. Repeat steps 3 to 7 until the token T generated in this round is the stop mark of its source model m (such as EOS, etc.), or the set maximum generation length is reached.
[0081] 9. Take the token sequence S of any model and decode it as the final output text O of the user instruction text I.
[0082] The generative language model-based reply method provided by this application fully utilizes the different knowledge reserves and capabilities of multiple generative language models, combining multiple models in a probabilistically weighted manner during the token generation phase, thereby compensating for the limitations and weaknesses of individual generative language models and alleviating the "hallucination" problem commonly seen when using only a single generative language model with smaller parameters. Compared with the prior art, the advantages of this invention are:
[0083] 1. Supports any number of generative language models to participate in the combination, so the combination scale can be flexibly adjusted according to resources and needs;
[0084] 2. The generative language models involved in the combination can be Transformer-based causal language models of any parameter scale and any base, such as LLaMA, Bloom, ChatGLM, etc., and the greater the difference between the models, the better;
[0085] 3. The combined generative language model does not require additional training steps, thus avoiding huge pre-training resources and time investment;
[0086] 4. This method adopts a probability-weighted combination method, so the priority of different models can be flexibly adjusted through customized weight allocation.
[0087] 5. The combined generative language models can be calculated in parallel, and the combined system generates faster than a single generative language model with the same number of parameters.
[0088] Based on the same inventive concept, the present application provides a reply device based on a generative language model, wherein the generative language model includes multiple, any two generative language models are independent of each other, wherein each generative language model is loaded on a GPU, and any two generative language models are loaded on different GPUs. As shown in Figure 3, the device includes:
[0089] The candidate output token determination module 31 is used to determine the predicted candidate output tokens and the generation probability of the candidate output tokens for each generative language model based on the instruction text input by the user. Specifically, the instruction text input by the user is obtained; the instruction text is encoded into an input sequence of tokens using the tokenizer corresponding to each generative language model; a forward propagation is performed in the corresponding generative language model based on the input sequence to obtain multiple predicted tokens; and the one with the highest probability among the multiple predicted tokens is used as the candidate output token of the generative language model.
[0090] The target generation probability calculation module 32 is used to multiply the generation probability of the candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the candidate output token.
[0091] In one embodiment, the target generation probability calculation module 32 determines the classification of the instruction text based on the instruction text input by the user; and determines the preset weight of each generative language model based on the classification.
[0092] The target output token determination module 33 is configured to select the candidate output token with the highest target generation probability as the target output token.
[0093] The reply text acquisition module 34 is configured to obtain the reply text based on the target output token.
[0094] Specifically, the reply text acquisition module 34 uses the tokenizer in the generative language model corresponding to the target output token to decode the target output token to obtain a target string; uses the target output token of each generative language model that encodes the target string to add the target output token to the input sequence of the corresponding generative language model; based on the input sequence, determines the next candidate output token predicted by each generative language model and the generation probability of the next candidate output token; multiplies the generation probability of the next candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the next candidate output token; uses the next candidate output token with the highest target generation probability as the next target output token; if the currently generated target output token meets the preset stop condition, uses the text corresponding to the target output token sequence when the stop condition is met as the reply text. If the currently generated target output token does not meet the preset stop condition, repeats the step of obtaining the next target output token.
[0095] The preset stopping condition is: the generated target output token is the stop mark of the corresponding generative language model, or the target output token sequence length reaches the maximum generation length.
[0096] The reply device based on the generative language model provided in the embodiment of the present application uses multiple generative language models to predict the next token after the user inputs the instruction text, and uses the token with the highest generation probability as the candidate output token, and then uses the product of the generation probability of each candidate output token and the preset weight of the corresponding generative language model as the final target generation probability, and uses the candidate input token with the highest target generation probability as the target output token. That is, a target output token is obtained by combining the advantages of multiple generative language models, and the text corresponding to the target output token is used as the output of all generative language models. In this way, each token in the final reply text is obtained by combining the outputs of all generative language models. The present application solution does not require further training of each generative language model, which greatly reduces the cost, and even if the candidate output token of a generative language model is incorrect, it can be eliminated in combination with the candidate output tokens of other generative language models, which greatly increases the accuracy of the final reply.
[0097] Based on the same inventive concept, the present application provides a response system based on a generative language model, comprising:
[0098] at least one processor and at least one memory;
[0099] The memory stores executable instructions of the processor;
[0100] The processor is configured to execute the generative language model-based reply method provided in the above embodiment.
[0101] The reply system based on the generative language model provided by the embodiment of the present application stores the executable instructions of the processor in the memory. When the executable instructions are executed, the processor can use multiple generative language models to predict the next token after the user inputs the instruction text, and use the token with the highest generation probability as the candidate output token, and then use the product of the generation probability of each candidate output token and the preset weight of the corresponding generative language model as the final target generation probability, and use the candidate input token with the highest target generation probability as the target output token. That is, a target output token is obtained by combining the advantages of multiple generative language models, and the text corresponding to the target output token is used as the output of all generative language models. In this way, each token in the final reply text is obtained by combining the outputs of all generative language models. The present application solution does not require further training of each generative language model, which greatly reduces the cost, and even if the candidate output token of a generative language model is incorrect, it can be eliminated in combination with the candidate output tokens of other generative language models, which greatly increases the accuracy of the final reply.
[0102] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0103] It should be noted that, in the description of this application, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" refers to at least two.
[0104] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0105] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0106] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0107] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0108] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0109] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0110] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A response method based on a generative language model, wherein, The generative language models include multiple ones, and any two generative language models are independent of each other. The method includes: Based on the instruction text input by the user, determining the predicted candidate output tokens of each generative language model and the generation probabilities of the candidate output tokens; Multiplying the generation probability of the candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the candidate output token; Taking the candidate output token with the highest target generation probability as the target output token; Obtaining a response text based on the target output token.
2. The method according to claim 1, wherein, The determining the candidate output tokens of each generative language model based on the instruction text input by the user includes: Obtaining the instruction text input by the user; Encoding the instruction text into an input sequence of tokens by using the tokenizer corresponding to each generative language model; Performing a forward pass in the corresponding generative language model based on the input sequence to obtain multiple predicted tokens; Taking the token with the highest probability in the probability distribution of the multiple predicted tokens as the candidate output token of the generative language model.
3. The method according to claim 2, wherein, The obtaining a response text based on the target output token includes: Decoding the target output token by using the tokenizer in the generative language model corresponding to the target output token to obtain a target string; Encoding the target string by using each generative language model to obtain the target output tokens of each generative language model, and adding the target output tokens to the input sequence of the corresponding generative language model; Based on the input sequence, determining the predicted next candidate output tokens of each generative language model and the generation probabilities of the next candidate output tokens; Multiplying the generation probability of the next candidate output token by the preset weight of the corresponding generative language model to obtain the target generation probability of the next candidate output token; Taking the next candidate output token with the highest target generation probability as the next target output token; If the currently generated target output token meets the preset stop condition, then taking the text corresponding to the target output token sequence when the stop condition is met as the response text.
4. The method according to claim 3, wherein, If the currently generated target output token does not meet the preset stop condition, then repeating the steps to obtain the next target output token.
5. The method according to claim 3, wherein, The preset stop condition is: the generated target output token is the stop token of the corresponding generative language model.
6. The method according to claim 3, wherein, The preset stop condition is: the length of the target output token sequence reaches the maximum generation length.
7. The method according to claim 1, further comprising: Based on the instruction text input by the user, determining the classification of the instruction text; Determining the preset weight of each generative language model based on the classification.
8. The method according to claim 1, wherein, Each generative language model is loaded on a GPU, and the GPUs loaded by any two generative language models are different.
9. A response device based on a generative language model, wherein, The generative language models include multiple ones, and any two generative language models are independent of each other. The apparatus includes: a candidate output token determination module, configured to determine, based on the instruction text input by the user, the predicted candidate output tokens of each generative language model and the generation probabilities of the candidate output tokens; a target generation probability calculation module, configured to multiply the generation probabilities of the candidate output tokens by the preset weights of the corresponding generative language models to obtain the target generation probabilities of the candidate output tokens; a target output token determination module, configured to use the candidate output token with the highest target generation probability as the target output token; a response text acquisition module, configured to obtain a response text based on the target output token.
10. A response system based on a generative language model, including: at least one processor and at least one memory; the memory stores executable instructions of the processor; the processor is configured to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Named entity identification method and device, electronic equipment and storage medium
CN116842951A
Dialogue generation method and device, electronic equipment and computer readable storage medium
CN116955554A
Replay method, device and system based on generative language model
CN118012999A
Multi-model approach to natural language processing and recommendation generation
US20230046851A1
Using large language model(s) in generating automated assistant response(s)
WO2023038654A1