Model-based reasoning method, electronic equipment and intelligent agent
By adjusting the probability values of candidate codes using a word segmenter and a pre-stored intent graph, the problem of inconsistent or meaningless output information from large language models is solved, thus improving the accuracy of the output information and its matching degree with the pre-stored intent.
Patent Information
- Application Number
- CN202511073768.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-04
AI Technical Summary
Large language models may output information that does not match the user input or has no real meaning in reasoning tasks; this is called "illusion".
The input information is converted into input words by a word segmenter, the input encoding sequence is obtained based on the vocabulary, and the output encoding is generated through multiple rounds of reasoning. The probability values of candidate encodings are adjusted using a pre-stored meaning map to determine a more accurate output encoding.
This reduces the probability of large language models outputting inconsistent or meaningless information, and improves the accuracy of the output information and its matching degree with pre-stored intents.
Smart Images

Figure CN120893579A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of model inference, in particular to a model-based inference method, an electronic device and an agent. BACKGROUND
[0002] Large language models have been widely used to process inference tasks in various scenarios. When used for inference, large language models can obtain input information, perform inference calculation based on the input information, and obtain corresponding output information.
[0003] In some cases, the output information of the large language model may be some information inconsistent with the user input or objective facts, or some meaningless random codes, which is generally referred to as the "hallucination" of the large language model. SUMMARY
[0004] To this end, the present application discloses the following technical solutions:
[0005] The first aspect of the present application provides a model-based inference method, comprising:
[0006] obtaining first input information;
[0007] converting the first input information into a plurality of input tokens through a tokenizer, the input token being the smallest unit for text processing by a large language model;
[0008] obtaining an input encoding of each input token based on a vocabulary to form an input encoding sequence;
[0009] generating output information composed of a plurality of output encodings through multiple rounds of inference based on the input encoding sequence and a large language model, one output encoding being generated in each round of inference;
[0010] wherein, generating one output encoding in each round of inference comprises:
[0011] calculating a first probability value of a candidate encoding in the Nth round based on the encoding in the (N-1)th round, N being an integer greater than or equal to 1;
[0012] querying a pre-stored intent table to obtain a plurality of different pre-stored encodings included in the Nth column, the pre-stored intent table including pre-stored encodings corresponding to a plurality of intent items;
[0013] adjusting the first probability value of the target encoding to a second probability value, the target encoding being an encoding consistent with the plurality of different pre-stored encodings included in the Nth column among the candidate encodings;
[0014] determining a candidate encoding as one output encoding generated in the Nth round of inference according to the probability value of the candidate encoding.
[0015] Optionally, further comprising:
[0016] Executing an operation instruction corresponding to the output information;
[0017] Or, outputting a word corresponding to each output encoding in the output information.
[0018] Optionally, the encoding of the N-1th round and an output encoding generated by the Nth round of inference are used as the encoding of the Nth round;
[0019] The encoding of the Nth round is used for the N+1th round of inference.
[0020] Optionally, the first probability value of the target encoding is adjusted to the second probability value, comprising:
[0021] Multiplying the first probability value of the target encoding by a coefficient to obtain the second probability value.
[0022] Optionally, if the output encoding generated by the Nth round of inference belongs to the target encoding, the second probability value of the target encoding in the candidate encoding representing the process of generating the output encoding by the Nth round of inference is greater than the probability value of each candidate encoding except the target encoding.
[0023] Optionally, if the output encoding generated by the Nth round of inference does not belong to the target encoding, the second probability value of the target encoding in the candidate encoding representing the process of generating the output encoding by the Nth round of inference is not greater than the probability value of each candidate encoding except the target encoding.
[0024] Optionally, the coefficient is determined according to the output information when the large language model processes the sample input information.
[0025] Optionally, determining a candidate encoding as an output encoding generated by the Nth round of inference according to the probability value of the candidate encoding, comprising:
[0026] Determining the candidate encoding with the maximum probability value as the output encoding generated by the Nth round of inference based on the probability value of the candidate encoding;
[0027] Or, determining multiple candidate encodings in descending order of probability values and determining any one of the candidate encodings as the output encoding generated by the Nth round of inference.
[0028] Optionally, N is greater than 1;
[0029] The query pre-stored intention table obtains multiple different pre-stored encodings included in the Nth column, comprising:
[0030] Querying the candidate pre-stored encoding sequence in the pre-stored intention table obtains multiple different pre-stored encodings included in the Nth column of the candidate pre-stored encoding sequence;
[0031] wherein the pre-stored encoding of the N-1th column of the candidate pre-stored encoding sequence is identical to one output encoding generated by the N-1th round of inference.
[0032] The second aspect of the present application provides an electronic device, comprising:
[0033] a memory configured to store a large language model;
[0034] a processor configured to:
[0035] obtain first input information;
[0036] convert the first input information into a plurality of input tokens by a tokenizer, the input token being a minimum unit for text processing by the large language model;
[0037] obtain an input encoding of each of the input tokens based on a vocabulary to form an input encoding sequence;
[0038] generate output information composed of a plurality of output encodings through a plurality of rounds of inference based on the input encoding sequence and the large language model, one output encoding being generated by each round of inference in the plurality of rounds of inference;
[0039] wherein one output encoding is generated by each round of inference comprises:
[0040] calculating to obtain a first probability value of a candidate encoding of the Nth round based on the encoding of the N-1th round, N being an integer greater than or equal to 1;
[0041] querying a pre-stored intent table to obtain a plurality of different pre-stored encodings included in the Nth column, the pre-stored intent table including pre-stored encodings corresponding to a plurality of intent items;
[0042] adjusting the first probability value of the target encoding to a second probability value, the target encoding being an encoding consistent with the plurality of different pre-stored encodings included in the Nth column among the candidate encodings;
[0043] determining one candidate encoding as one output encoding generated by the Nth round of inference based on the probability value of the candidate encoding.
[0044] Optionally, the precision of the model parameters of the large language model stored in the memory is less than the precision of the model parameters of the large language model when the large language model is trained.
[0045] The third aspect of the present application provides an agent comprising executable computer instructions;
[0046] the agent, when invoked, is configured to implement:
[0047] obtain first input information;
[0048] convert the first input information into a plurality of input tokens by a tokenizer, the input token being a minimum unit for text processing by a large language model;
[0049] obtain an input encoding of each of the input tokens based on a vocabulary to form an input encoding sequence;
[0050] generate output information composed of a plurality of output encodings through a plurality of rounds of reasoning based on the input encoding sequence and the large language model, one output encoding being generated in each round of reasoning;
[0051] wherein one output encoding is generated in each round of reasoning includes:
[0052] based on the encoding of the N-1th round, calculate to obtain a first probability value of the candidate encoding of the Nth round, N being an integer greater than or equal to 1;
[0053] query a pre-stored intent table to obtain a plurality of different pre-stored encodings included in the Nth column, the pre-stored intent table including pre-stored encodings corresponding to a plurality of intent items;
[0054] adjust the first probability value of the target encoding to a second probability value, the target encoding being an encoding consistent with the plurality of different pre-stored encodings included in the Nth column among the candidate encodings;
[0055] determine a candidate encoding as one output encoding generated in the Nth round of reasoning according to the probability value of the candidate encoding. BRIEF DESCRIPTION OF DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.
[0057] Figure 1 is a flowchart of a model-based reasoning method provided by an embodiment of the present application;
[0058] Figure 2 is an example diagram of a plurality of rounds of reasoning to obtain output information provided by an embodiment of the present application;
[0059] Figure 3 is a diagram of a probability value corresponding to a candidate encoding provided by an embodiment of the present application;
[0060] Figure 4 is another diagram of a probability value corresponding to a candidate encoding provided by an embodiment of the present application;
[0061] Figure 5Fig. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0062] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0063] The present embodiment provides a model-based reasoning method, please refer to Figure 1 , which is a flowchart of the method. The method can include the following steps.
[0064] S101, obtaining first input information.
[0065] The reasoning method provided by the present embodiment can be executed by any electronic device capable of deploying a large language model. The type of the electronic device is not limited, which can be a terminal type device such as a smart phone, a notebook computer, etc., or a server type device.
[0066] The obtaining method of the first input information is not limited, for example, the text information input by the user can be obtained through a keyboard or a virtual key as the first input information, the text information can be recognized from an image through a character recognition technology as the first input information, the text information can be captured from any document as the first input information, and the text information can be obtained by voice recognition of audio as the first input information.
[0067] S102, converting the first input information into a plurality of input tokens through a tokenizer, the input token being the smallest unit for text processing by the large language model.
[0068] The tokenizer is a software module pre-configured in the electronic device for splitting text information into a plurality of tokens, and the token refers to the smallest unit for text processing by the large language model. The plurality of tokens obtained by splitting the first input information are the plurality of input tokens in S102. The structure and working principle of the tokenizer can be referred to related technologies, and the present embodiment is not limited.
[0069] In the Chinese context, a token can be a single character, a word composed of two or more characters, an Arabic numeral, or a specific English abbreviation. As an example, the first input information "I want to watch TV" can obtain four input tokens after splitting, which are "I", "want", "watch", and "TV" in turn.
[0070] In the English context, a word piece can be a word, a specific phrase composed of two or more words, an Arabic numeral, or a specific English abbreviation. As an example, after the first input information "I have a cat" is split, four input word pieces can be obtained, in order: "I", "have", "a", and "cat".
[0071] A large language model is an artificial intelligence model (also referred to as a deep learning model) based on deep learning technology and relying on a self-attention mechanism, and has a large parameter size compared to other artificial intelligence models. For example, a large language model commonly used in the related technical field generally contains a parameter amount of about 100 billion, and other artificial intelligence models other than large language models generally contain a parameter amount of about 1 million.
[0072] Based on the large parameter size, the large language model has the ability to understand natural language, such as understanding the complex semantics, context logic, and implied intent of natural language. Based on this ability, the large language model can reason on input information in the form of natural language and perform various tasks based on the reasoning results.
[0073] S103, obtaining an input encoding of each input word piece based on the vocabulary to form an input encoding sequence.
[0074] The vocabulary is a data table pre-configured in an electronic device deploying a large language model, which records a plurality of word pieces and the encoding corresponding to each word piece. The correspondence between the word piece and the encoding is unique, that is, the encoding corresponding to different word pieces is different, and the word piece corresponding to different encodings is also different.
[0075] In S103, for each input word piece, the input word piece can be queried in the vocabulary to obtain the encoding corresponding to the input word piece as the input encoding. In the first input information, a plurality of input word pieces are arranged in a certain order, and the input encodings corresponding to these input word pieces are arranged in the same order to form an input encoding sequence.
[0076] Table 1
[0077]
[0078] Continuing the preceding example, a part of the word pieces in the vocabulary and the encodings corresponding thereto can be shown in Table 1. According to the vocabulary, the input encoding 16394 corresponding to "I", the input encoding 16248 corresponding to "think", the input encoding 20398 corresponding to "watch", and the input encoding 4624 corresponding to "television" can be obtained. These input encodings are arranged in the order of the input word pieces in the first input information to form an input encoding sequence: (16394, 16248, 20398, 4624).
[0079] In some embodiments, each token corresponds to an encoding, which can also be referred to as an identifier of the token.
[0080] In step S104, the electronic device can call the large language model to read the input encoding sequence, and make the large language model perform multiple rounds of reasoning based on the input encoding sequence. In each round of reasoning, the large language model can generate an output encoding. From the first round of reasoning to the termination of the reasoning (i.e., the end of the last round of reasoning), the multiple output encodings generated by the large language model constitute the output information of the large language model.
[0081] In step S104, the electronic device can call the large language model to read the input encoding sequence, and make the large language model perform multiple rounds of reasoning based on the input encoding sequence. In each round of reasoning, the large language model can generate an output encoding. From the first round of reasoning to the termination of the reasoning (i.e., the end of the last round of reasoning), the multiple output encodings generated by the large language model constitute the output information of the large language model.
[0082] In each round of reasoning, the large language model can terminate the reasoning or continue the next round of reasoning based on the output encoding generated in the current round of reasoning. For example, the vocabulary can include one or more encodings indicating termination of the reasoning. If the output encoding generated in the current round of reasoning is an encoding indicating termination of the reasoning, the large language model terminates the reasoning. If the output encoding generated in the current round of reasoning is not an encoding indicating termination of the reasoning, the large language model continues the next round of reasoning.
[0083] For example, the vocabulary can include an encoding corresponding to a token “EOF”. If the output encoding generated in the current round of reasoning is the encoding corresponding to the token “EOF”, the large language model terminates the reasoning. If the output encoding generated in the current round of reasoning is not the encoding corresponding to the token “EOF”, the large language model continues the next round of reasoning.
[0084] In each round of reasoning, the large language model can terminate the reasoning or continue the next round of reasoning based on the number of output encodings generated from the first round of reasoning. For example, if the number of output encodings generated from the first round of reasoning is equal to a preset number threshold, the large language model terminates the reasoning. If the number of output encodings generated from the first round of reasoning is less than the preset number threshold, the large language model continues the next round of reasoning.
[0085] In step S104, the process of generating an output encoding in each round of reasoning can include the following steps:
[0086] In step S1041, the first probability value of the candidate encoding of the Nth round is calculated based on the encoding of the (N-1)th round, where N is an integer greater than or equal to 1.
[0087] In the case that the first round of inference is performed by the large language model, i.e., N is equal to 1, the encoding of the 0th round can only include all encodings constituting the input encoding sequence, see the example of S103, the encoding of the 0th round can include (16394, 16248, 20398, 31198).
[0088] In the case that N is greater than 1, the encoding of the (N-1)th round can include all encodings constituting the input encoding sequence, and N-1 output encodings generated in the first round to the (N-1)th round of inference.
[0089] The manner of calculating the first probability value of the candidate encoding of the Nth round can be:
[0090] The large language model calculates, based on the weight parameters contained in the model, the key vector Key, the value vector Value corresponding to each encoding in the encoding of the (N-1)th round, and the query vector Query of the current encoding of the (N-1)th round;
[0091] In the case that the encoding of the (N-1)th round only includes all encodings constituting the input encoding sequence, the current encoding of the (N-1)th round can be the last input encoding in the input encoding sequence; in the case that the encoding of the (N-1)th round includes N-1 output encodings generated in the first round to the (N-1)th round of inference, the current encoding of the (N-1)th round can be the output encoding generated by the large language model in the (N-1)th round of inference;
[0092] The correlation between the key vector of each encoding in the encoding of the (N-1)th round and the query vector of the current encoding is calculated, and the value vectors of each encoding in the encoding of the (N-1)th round are fused by taking the correlation as a weight, to obtain an output feature vector of the Nth round;
[0093] Finally, based on the output feature vector of the Nth round and the weight parameters contained in the large language model, the first probability value of the candidate encoding of the Nth round is calculated.
[0094] The weight parameters contained in the large language model can be determined when the large language model is trained, and the manner of determining the weight parameters and performing corresponding calculations based on the weight parameters can refer to related technologies, which will not be described herein.
[0095] The first probability value corresponding to each candidate encoding can be any real number, such as -0.5, 10, 0.6, etc., and the specific value depends on the calculation result of the large language model.
[0096] The number of candidate encodings in the Nth round is not limited. In some embodiments, all encodings recorded in the aforementioned vocabulary can be used as candidate encodings. For example, if there are 1000 encodings recorded in the vocabulary, the candidate encodings in the Nth round include the 1000 encodings recorded in the vocabulary. In some embodiments, the large language model can calculate the probability value of each encoding recorded in the vocabulary, and the encodings with a probability value greater than a certain threshold are used as candidate encodings.
[0097] In step S1041, the obtained can be a first probability value sequence composed of a plurality of first probability values, wherein each first probability value corresponds to a candidate encoding at a corresponding position in the vocabulary in order. For example, the plurality of first probability values obtained in the Nth round of reasoning can include (0.234, -3.9, 4.083, 1.45, 0.346, 0.467, 0.942), wherein 0.234 corresponds to the candidate encoding in the first column of the vocabulary, -3.9 corresponds to the candidate encoding in the second column of the vocabulary, and so on.
[0098] In S1042, the pre-stored intent table is queried to obtain a plurality of different pre-stored encodings included in the Nth column. The pre-stored intent table includes pre-stored encodings corresponding to a plurality of intent items.
[0099] The pre-stored intent table can include a plurality of pre-stored encoding sequences composed of pre-stored encodings. Each pre-stored encoding sequence is stored in a row of the pre-stored intent table, and each pre-stored encoding sequence corresponds to an intent item. The Nth column of the pre-stored intent table includes the Nth pre-stored encoding of each pre-stored encoding sequence in the table.
[0100] The pre-stored intent table can be configured into the memory of the electronic device when the large language model is deployed to the electronic device. It can be pre-written into the memory of the electronic device by the relevant manufacturer during the production stage of the electronic device, or it can be downloaded by the electronic device through the network. The electronic device can update the pre-stored intent table in various ways, or it can not update the pre-stored intent table, which is not limited.
[0101] The pre-stored intent table specifically stores which pre-stored encoding sequences correspond to which intent items, which can be configured as needed and is not limited. The relevant manufacturer can write pre-stored encoding sequences corresponding to different intent items in the pre-stored intent table according to the main use scenarios of the large language model.
[0102] As some examples, if the large language model configured in the electronic device is mainly used to control various appliances, the pre-stored encoding sequences in the pre-stored intent table can correspond to the intent items of turning on Bluetooth search, turning on air conditioner, turning off air conditioner, increasing air conditioner temperature, turning on TV, turning on sound, etc.
[0103] If the large language model configured in the electronic device is mainly used in daily life scenarios, the pre-stored encoding sequences in the pre-stored intent table can correspond to the intent items of booking a train ticket, searching for nearby food, taking a taxi, etc.
[0104] If the large language model configured in the electronic device is mainly used for managing and optimizing the system of the electronic device, each pre-stored coding sequence in the pre-stored intent table can correspond to: virus killing, memory optimization, garbage file cleaning, movie downloading, and other intent items.
[0105] The number of pre-stored codes contained in different pre-stored coding sequences can be the same or different.
[0106] For each intent item, the intent item can be decomposed into multiple word units by a word segmenter, and then the codes corresponding to each word unit contained in the intent item are obtained from a word table as pre-stored codes. These pre-stored codes are arranged in the order of the arrangement of the word units in the intent item, which constitutes the pre-stored coding sequence corresponding to the intent item.
[0107] As an example, the pre-stored intent table can include several pre-stored coding sequences as shown in Table 2.
[0108] Table 2
[0109]
[0110] The pre-stored coding sequence (55, 643, 4624, 35) of the first row of Table 2 corresponds to the intent item “turn on the TV”, and the four pre-stored codes correspond to the word units “turn on”, “TV”, and “machine” in turn;
[0111] The pre-stored coding sequence (55, 643, 430, 800) of the second row corresponds to the intent item “turn on Bluetooth search”, and the four pre-stored codes correspond to the word units “turn on”, “Bluetooth”, and “search” in turn;
[0112] The pre-stored coding sequence (100, 318, 4624, 35) of the third row corresponds to the intent item “turn off the TV”, and the four pre-stored codes correspond to the word units “turn off”, “TV”, and “machine” in turn;
[0113] The pre-stored coding sequence (100, 318, 592, 7301, 561) of the fourth row corresponds to the intent item “turn off the air conditioner cooling”, and the five pre-stored codes correspond to the word units “turn off”, “air conditioner”, “cooling”, and “cool” in turn.
[0114] The pre-stored coding sequence (55, 6071, 592, 4423, 8402) of the fifth row corresponds to the intent item “call a taxi home”, and the five pre-stored codes correspond to the word units “call a taxi”, “home”, and “home” in turn.
[0115] In S1042, the Nth column of the pre-stored intent table can be queried to obtain all pre-stored codes in the Nth column. After filtering out the remaining pre-stored codes, the pre-stored codes are the multiple different pre-stored codes included in the Nth column.
[0116] In the example of Table 2, when N equals 1, the first column of Table 2 is queried to filter out the duplicate pre-stored codes, and two different pre-stored codes 55 and 100 included in the first column are obtained; when N equals 3, the third column of Table 2 is queried to filter out the duplicate pre-stored codes, and three different pre-stored codes 4624, 430 and 592 included in the third column are obtained.
[0117] In S1043, the first probability value of the target code is adjusted to the second probability value, and the target code is a code in the candidate codes that is consistent with the plurality of different pre-stored codes included in the Nth column.
[0118] The target code refers to a candidate code that is consistent with any pre-stored code in the Nth column queried in S1042.
[0119] In S1043, whether there is a target code in the Nth round of candidate codes that is consistent with the plurality of different pre-stored codes included in the Nth column can be queried based on the plurality of different pre-stored codes included in the Nth column, and the first probability value corresponding to each target code is adjusted to the second probability value.
[0120] In the example of Table 2, when N equals 1, two different pre-stored codes 55 and 100 included in the first column are obtained by querying, and the first probability value corresponding to the first round of candidate codes can include: the candidate code 89, the corresponding probability value is 0.1; the candidate code 62, the corresponding probability value is 0.2; the candidate code 55, the corresponding probability value is 0.7; the candidate code 784, the corresponding probability value is 0.9; the candidate code 100, the corresponding probability value is 0.3; and the candidate code 7203, the corresponding probability value is 1.0.
[0121] After querying, it is found that the candidate code 55 is the same as one of the two different pre-stored codes included in the first column, and the candidate code 100 is the same as one of the two different pre-stored codes included in the first column, so the first probability value 0.7 corresponding to the candidate code 55 is adjusted to the second probability value 1.4, and the first probability value 0.3 corresponding to the candidate code 100 is adjusted to the second probability value 0.6.
[0122] In the Nth round of candidate codes, there can be a target code or no target code, and the number of target codes can be one or more when there is a target code.
[0123] If there is no target code in the Nth round of candidate codes, S1044 can be directly executed, and if there is at least one target code in the Nth round of candidate codes, the first probability value of each target code can be adjusted to the second probability value.
[0124] The first probability values corresponding to different candidate encodings can be the same or different, and the second probability values obtained after adjustment of different target encodings can be the same or different.
[0125] In S1044, a candidate encoding is determined as an output encoding generated by the Nth round of reasoning based on the probability values of the candidate encodings.
[0126] In S1044, a candidate encoding is determined as an output encoding generated by the Nth round of reasoning based on the probability values of the candidate encodings.
[0127] In combination with the example of S1043, a candidate encoding can be determined as an output encoding generated by the Nth round of reasoning based on the probability values corresponding to the following candidate encodings: candidate encoding 89, corresponding probability value 0.1; candidate encoding 62, corresponding probability value 0.2; candidate encoding 55, corresponding probability value 1.4; 784, corresponding probability value 0.9; candidate encoding 100, corresponding probability value 0.6; and candidate encoding 7203, corresponding probability value 1.0.
[0128] Generally, the greater the probability value corresponding to a candidate encoding, the more likely the candidate encoding is to be determined as an output encoding generated by the Nth round of reasoning, and the smaller the probability value corresponding to a candidate encoding, the less likely the candidate encoding is to be determined as an output encoding generated by the Nth round of reasoning. The determination of an output encoding generated by the Nth round of reasoning based on the probability values of the candidate encodings is not limited.
[0129] The beneficial effects of the embodiment are as follows:
[0130] Before generating an output encoding based on one round of reasoning of a large language model, the first probability values of target encodings consistent with the pre-stored encodings in the candidate encodings are adjusted according to the pre-stored intention table. Thus, the probability of the target encodings consistent with the pre-stored encodings in the candidate encodings being determined as output encodings can be fine-tuned, the large language model is more likely to output output information matching the pre-stored intention table, and the probability of the large language model outputting information inconsistent with user input or objective facts and outputting random codes without actual meaning is reduced to a certain extent, which is beneficial to avoid the "hallucination" of the large language model.
[0131] The beneficial effects of the method of the embodiment will be described below in combination with an example.
[0132] Please refer to Figure 2 An example diagram for performing multiple rounds of reasoning to obtain output information is provided for the embodiment.
[0133] The electronic device obtains the first input information "I want to watch TV", and obtains the input coding sequence (134, 5246, 9254, 374) corresponding to the first input information based on the method of the foregoing embodiment;
[0134] The input coding sequence is input into the large language model, and the large language model performs the first round of reasoning based on the input coding sequence, generating the output coding 55 of the first round of reasoning. The word element corresponding to this coding is "hit";
[0135] Based on the coding of the first round, that is, based on the input coding sequence and the output coding of the first round of reasoning, the large language model continues to perform the second round of reasoning, generating the output coding 643 of the second round of reasoning. The word element corresponding to this coding is "turn on";
[0136] Based on the coding of the second round, which includes the input coding sequence and the output codings of the first two rounds of reasoning, the large language model performs the third round of reasoning, generating the output coding 4624 of the third round of reasoning. The word element corresponding to this coding is "TV";
[0137] Based on the coding of the third round, which includes the input coding sequence and the output codings of the first three rounds of reasoning, the large language model performs the fourth round of reasoning, generating the output coding 35 of the fourth round of reasoning. The word element corresponding to this coding is "set-top box";
[0138] Based on the coding of the fourth round, which includes the input coding sequence and the output codings of the first four rounds of reasoning, the large language model performs the fifth round of reasoning, generating the output coding 9000 of the fifth round of reasoning. This coding corresponds to the identifier EOF indicating the termination of reasoning in the word list. Thus, when the fifth round of reasoning ends, the large language model terminates reasoning. The output codings other than the coding indicating the termination of reasoning output in the last round of reasoning constitute the output information of the large language model. That is, the output information of the large language model can be: 55, 643, 4624, 35, and the corresponding text is "Turn on the TV set-top box".
[0139] In the above reasoning process, taking the first round of reasoning as an example, the first probability value of the first-round candidate coding calculated by the large language model based on the input coding sequence can include the first probability values shown in Table 3.
[0140] Table 3
[0141]
[0142] The first row of Table 3 shows five candidate codings, and the second row shows the first probability values corresponding to the above five candidate codings obtained during the first round of reasoning. Among them, the first probability value 8 corresponding to the candidate coding 800 is the largest. If the output coding of the first round of reasoning is directly determined based on the first probability values of each candidate coding, there is a relatively high probability that the output coding of the first round is determined to be 800, and the corresponding word element is "search".
[0143] In the case of applying the method of this embodiment, after obtaining the first probability value of the first-round candidate encoding, query the first column of the pre-stored intention table described in Table 2, and obtain two different pre-stored encodings 55 and 100 included in the first column. Then, adjust the first probability value 7 of the candidate encoding 55 in Table 3 to the second probability value 14, and adjust the first probability value 0.5 of the candidate encoding 100 in Table 3 to the second probability value 1.
[0144] After adjustment, the second probability value 14 corresponding to the candidate encoding 55 is the largest probability value corresponding to the candidate encodings in Table 3. Therefore, there is a high probability that the output encoding of the first round is determined to be 55, and the corresponding token is "hit".
[0145] It can be seen from this that after applying the method of this embodiment, the output encoding 800 that is highly likely to be generated in the first-round inference can be corrected to the output encoding 55 that is more in line with the first input information "I want to watch TV", which can reduce the probability of the large language model outputting information that does not match the user input or objective facts, and the probability of outputting meaningless garbled characters to a certain extent.
[0146] After obtaining the output information, the electronic device can also perform corresponding operations based on the output information. For example, it can execute the operation instruction corresponding to the output information;
[0147] Or it can output the tokens corresponding to each output encoding in the output information.
[0148] What specific operation the electronic device performs can be related to the current application scenario.
[0149] In a scenario where the electronic device is communicatively connected to other devices and can thus control other devices (such as air conditioners, TVs), the electronic device can execute the operation instruction corresponding to the output information. For example, if the output information is: 55, 643, 4624, 35 (corresponding to "turn on the TV"), then the electronic device controls the TV to turn on.
[0150] In a scenario where the electronic device is not communicatively connected to other devices, the electronic device can convert each output encoding in the output information into a token based on the token table and output these tokens.
[0151] What specific operation the electronic device performs can also be related to the output information of the large language model. If the output information of the large language model can represent an operation instruction, the electronic device executes the operation instruction corresponding to the output information. If the output information of the large language model cannot represent an operation instruction, the electronic device outputs the tokens corresponding to each output encoding in the output information.
[0152] For example, if the text corresponding to the output information is "virus scan", the electronic device will execute the virus scan command corresponding to the output information. If the text corresponding to the output information is "sunny weather", the electronic device will output the tokens corresponding to each output code in the output information.
[0153] When outputting word units, they can be output one by one according to the order in which the large language model generates each output code. That is, first output the word units corresponding to the output codes generated in the first round of inference, then output the word units corresponding to the output codes generated in the second round of inference, then output the word units corresponding to the output codes generated in the third round of inference, and so on. Alternatively, these word units can be combined into an output text according to the order in which the output codes are generated, and then the output text can be output directly.
[0154] The output format is not limited, including but not limited to display output, voice output, and other possible output formats.
[0155] Optionally, the encoding of round N-1 and an output encoding generated by the reasoning of round N are used as the encoding of round N;
[0156] The encoding in round N is used for reasoning in round N+1.
[0157] In other words, when a large language model finishes a round of reasoning, it can combine the output code generated in the final round with the code used at the beginning of the previous round to calculate the first probability value of the candidate code for the next round of reasoning.
[0158] Combination Figure 2 For example, in the first round of reasoning, the large language model calculates the first probability value of the candidate codes for the first round based on the codes from the 0th round, that is, based on the input code sequence. In the second round of reasoning, the input code sequence and the output codes from the 1st round form the codes for the 1st round, and the large language model calculates the first probability value of the candidate codes for the 2nd round based on the codes from the 1st round. In the third round of reasoning, the codes from the 1st round and the output codes from the 2nd round form the codes for the 2nd round, and the large language model calculates the first probability value of the candidate codes for the 3rd round based on the codes from the 2nd round, and so on, until the large language model terminates the reasoning.
[0159] Using the output encoding generated in each round of inference to calculate the first probability value in the next round of inference helps improve the semantic coherence between the output encoding generated in each round of inference and the semantics of the previously generated output encoding, making the final output information more in line with the user's needs.
[0160] And, after adjusting the first probability value obtained in the Nth round of reasoning based on the method of the embodiment, the output encoding generated in the Nth round of reasoning is used for the N+1th round of reasoning, which can make the adjustment of the probability value in the Nth round of reasoning affect the N+1th round of reasoning, so as to avoid generating encoding that does not meet the requirements and is meaningless in the Nth round of reasoning, and also help to avoid generating encoding that does not meet the requirements and is meaningless in the N+1th round of reasoning.
[0161] Optionally, adjusting the first probability value of the target encoding to the second probability value comprises:
[0162] Multiplying the first probability value of the target encoding by a coefficient to obtain the second probability value.
[0163] The coefficient can be equal to any pre-set real number greater than 1, for example, it can be equal to 1.5, it can be equal to 2, it can be equal to 2.1, and its actual value is not limited.
[0164] Taking the coefficient equal to 2 as an example, when performing the Nth round of reasoning, after determining a target encoding, the first probability value corresponding to the target encoding in the first probability value generated in the Nth round of reasoning can be multiplied by 2, and the result is taken as the second probability value.
[0165] The way of adjusting the first probability value to the second probability value is not limited, and in some embodiments, the way of adjusting the first probability value to the second probability value can also be adding a pre-set probability value increment to the first probability value to obtain the second probability value.
[0166] The effect of adjusting the first probability value to the second probability value by multiplying the coefficient is that when adjusting the first probability value obtained in the Nth round of reasoning, the first probability value of the target encoding is only fine-tuned on the original basis by multiplying the coefficient, and the first probability value calculated by the large language model in the Nth round of reasoning will not be greatly changed.
[0167] Specifically, if a target encoding originally corresponds to a larger first probability value, the larger first probability value can be further amplified by multiplying the coefficient, so as to sufficiently improve the possibility of the target encoding being determined as the output encoding; if a target encoding originally corresponds to a smaller first probability value, even if it is adjusted to a second probability value by multiplying the coefficient, the adjusted second probability value is still smaller and will not exceed the larger first probability value originally output by the large language model, so these target encodings will not be determined as the output encoding.
[0168] It can be seen that after the large language model calculates the first probability value of the candidate encoding of the Nth round, the method of the embodiment can amplify the first probability value of the target encoding that is originally likely to be determined as the output encoding after the large language model calculates to the second probability value, so that these target encodings are more likely to be output in the Nth round of reasoning, and the first probability value of the target encoding that is basically not likely to be determined as the output encoding after the large language model calculates is not excessively improved, and these target encodings that are basically not likely to be determined as the output encoding are not output.
[0169] Therefore, if the output encoding generated in the Nth round of reasoning belongs to the target encoding, among the candidate encodings representing the process of generating an output encoding in the Nth round of reasoning, the second probability value of the target encoding is greater than the probability value of each candidate encoding except the target encoding, and the first probability value originally corresponding to the target encoding is the larger first probability value among the multiple first probability values generated in the Nth round of reasoning.
[0170] For example, if the first probability value corresponding to a target encoding is ranked in the top 5 among the multiple first probability values obtained in the Nth round of reasoning from large to small, or the difference between the first probability value and the largest first probability value obtained in the Nth round of reasoning is less than a certain threshold, then the second probability value obtained by multiplying the first probability value corresponding to the target encoding by the coefficient will be greater than the probability value corresponding to each other candidate encoding, so when finally determining the output encoding, the target encoding will be determined as the output encoding generated in the Nth round of reasoning.
[0171] In an example, please refer to Figure 3 , it is assumed that among the first probability values of the candidate encodings generated in the Nth round of reasoning, the first probability values of four candidate encodings are as shown in Figure 3 (1), wherein taking candidate encoding 1 as the output encoding of the Nth round meets the user's demand, the first probability value of candidate encoding 1 is less than the first probability value of the largest candidate encoding 2, but is obviously greater than the first probability values of the other two candidate encodings. Based on the first probability values at this time, candidate encoding 2 will be determined as the output encoding, resulting in that the output of the large language model does not meet the user's demand.
[0172] After determining that candidate encoding 1 belongs to the target encoding, the first probability value of candidate encoding 1 is multiplied by the coefficient 2 to obtain the second probability value as shown in Figure 3 (2), at this time, the larger first probability value of candidate encoding 1 is amplified to the largest second probability value among the probability values of the four candidate encodings, so that candidate encoding 1 can be determined as the target encoding based on the probability values at this time, so that the output of the large language model meets the user's demand.
[0173] In some scenarios, it is also possible Figure 3The candidate encoding 2 is more in line with the user's needs as the output encoding of the Nth round. If the output encoding is directly determined based on the first probability values of the candidate encodings without adjustment, the candidate encoding 2 will be determined as the output encoding. However, after adjusting the first probability value of the candidate encoding 1 that belongs to the target encoding to the second probability value, the second probability value of the candidate encoding 1 is the largest probability value among all candidate encodings, so that the large language model determines the candidate encoding 1 as the output encoding generated by the Nth round of reasoning.
[0174] In this case, although the generated output encoding is not the candidate encoding 2, because the large language model calculates that the first probability value of the candidate encoding 1 is large, it can be considered that the candidate encoding 1 is also related to the first input information and to some extent also meets the user's needs.
[0175] If one of the output encodings generated by the Nth round of reasoning does not belong to the target encoding, the second probability value of the target encoding in the candidate encoding representing the process of generating one of the output encodings by the Nth round of reasoning is not greater than the probability value of each candidate encoding except the target encoding, and among the first probability values obtained by the Nth round of reasoning, the first probability values corresponding to these target encodings are small.
[0176] For example, the first probability value corresponding to one of the target encodings is ranked last 20th from large to small among the multiple first probability values obtained by the Nth round of reasoning, or the difference between the first probability value and the largest first probability value obtained by the Nth round of reasoning is less than a certain threshold. Then, the second probability value obtained by multiplying the first probability value by the coefficient may still be ranked last 20th from large to small among the probability values of the multiple candidate encodings, or the difference from the largest probability value may still be less than a certain threshold. Therefore, the target encodings with the second probability values that are still small after adjustment will not be determined as the output encoding generated by the Nth round of reasoning, and the output encoding generated by the Nth round of reasoning will still be the candidate encodings with larger first probability values.
[0177] Please refer to Figure 4 , it is assumed that among the first probability values of the candidate encodings generated by the Nth round of reasoning, the first probability values of four candidate encodings are as shown in (1) of Figure 4 , wherein the candidate encoding 1 as the output encoding of the Nth round meets the user's needs, the first probability value of the candidate encoding 1 is the smallest probability value among the probability values of the four candidate encodings, and the first probability value of the candidate encoding 2 is the largest first probability value among the first probability values of the four candidate encodings. Based on the first probability values at this time, the candidate encoding 2 can be determined as the output encoding, and the output of the large language model meets the user's needs.
[0178] After determining that the candidate encoding 1 belongs to the target encoding, the first probability value of the candidate encoding 1 is multiplied by the coefficient 2 to obtain Figure 4The second probability value of the candidate encoding 1 is the third in the four candidate encoding probability values, and is still much smaller than the probability value of the maximum candidate encoding 2. Based on the probability value at this time, the candidate encoding 2 is still determined as the output encoding, and the output of the large language model still meets the user's demand.
[0179] From the above examples, it can be seen that:
[0180] The target encoding that has a certain relevance to the first input information and basically meets the user's demand is more likely to be the output encoding of the Nth round of reasoning after adjustment by multiplying the coefficient due to the originally large first probability value. The target encoding that has little relevance to the first input information and does not meet the user's demand will not be the output encoding of the Nth round of reasoning even after adjustment by multiplying the coefficient due to the originally small first probability value. Therefore, adjusting the first probability value by multiplying the coefficient can make the target encoding with a larger first probability value and consistent with the pre-stored encoding more likely to be output, thereby obtaining output information that meets the user's intention, and can also avoid outputting the target encoding with a smaller first probability value.
[0181] Optionally, the coefficient is determined according to output information of the large language model when processing sample input information.
[0182] The large language model used for reasoning in this embodiment can be a large language model obtained by performing model quantization processing on a trained large language model, denoted as a quantized large language model.
[0183] The model quantization processing of the large language model refers to converting the parameters contained in the large language model from a data format with higher precision and larger storage space occupation to a data format with lower precision and smaller storage space occupation, so as to reduce the storage space required for storing the parameters contained in the large language model.
[0184] For example, the data format of the model parameters contained in the trained large language model can be a 32-bit floating point number (float), and the data format of the model parameters contained in the quantized large language model obtained by performing model quantization processing can be a 4-bit integer (int), or even a 1-bit integer.
[0185] The large language model after quantization processing will cause the accuracy of the processing result of the model to decrease, so the quantized large language model is more likely to output an encoding that does not meet the intention or output a "hallucination" phenomenon such as random code. The method of this embodiment can at least alleviate the "hallucination" problem of the quantized large language model, and improve the accuracy of the processing result of the model under the premise of reducing the storage space occupation.
[0186] In the above background, in order to make the quantized large language model obtain the most accurate result, the sample input information can be tested to determine the coefficient suitable for the quantized large language model.
[0187] In this embodiment, the coefficient can be determined based on the sample input information in the following manner.
[0188] For a quantized large language model, a large language model corresponding to the quantized large language model, which is trained and not subjected to model quantization processing, is denoted as a pre-quantized large language model. For example, a trained large language model A is subjected to model quantization processing to obtain a large language model B, and the large language model B is the quantized large language model, and the large language model A is the pre-quantized large language model corresponding to the quantized large language model.
[0189] The value of the coefficient is initialized, for example, the coefficient is set to 1.2, 2 or other values.
[0190] The pre-quantized large language model is used to process the input coding sequence corresponding to the sample input information to obtain first output information generated by the pre-quantized large language model after multiple rounds of reasoning; the quantized large language model is used to process the input coding sequence corresponding to the sample input information to obtain second output information generated by the quantized large language model after multiple rounds of reasoning; wherein the quantized large language model adjusts the first probability value calculated in each round based on the coefficient by using the reasoning method of the foregoing embodiments in each round of reasoning, and the sample input information processed by the pre-quantized large language model and the quantized large language model is the same information.
[0191] On this basis, the first output information and the second output information can be compared, and the probability values corresponding to the two models can be compared, and based on the comparison results, the value of the currently used coefficient is adjusted to make the first output information and the second output information as similar as possible, and the corresponding probability values as identical as possible, and the above process of processing the sample input information is repeated after the adjustment until the similarity of the first output information and the second output information is greater than a certain threshold value, and / or the difference between the corresponding probability values is less than a certain threshold value, and then the currently used coefficient can be determined as the coefficient applied in the foregoing reasoning method.
[0192] For example, after comparison, it is found that the first output information and the second output information differ too much, and the value of the currently used coefficient can be appropriately increased to increase the adjustment range; or after comparison, it is found that the second probability value is much greater than the first probability value output by the pre-quantized large language model, and the value of the currently used coefficient can be appropriately decreased to reduce the difference between the two.
[0193] The corresponding probability values refer to the second probability value of the target encoding output by the quantized large language model and the first probability value of the same candidate encoding output by the large language model before quantization in the same round of reasoning. For example, the second probability value of the target encoding 55 in the first round of reasoning output by the quantized large language model is 12, and the first probability value of the candidate encoding 55 in the first round of reasoning output by the large language model before quantization is 5. These two probability values are a pair of corresponding probability values.
[0194] In the above process of determining the coefficient, the sample input information used can be different from the first input information of the aforementioned reasoning method. The sample input information can be input information set by the relevant manufacturer based on experience in the test link.
[0195] Optionally, determining a candidate encoding as an output encoding generated in the Nth round of reasoning based on the probability value of the candidate encoding includes:
[0196] Determining a candidate encoding with the maximum probability value as an output encoding generated in the Nth round of reasoning based on the probability value of the candidate encoding.
[0197] Alternatively, a plurality of candidate encodings are determined in descending order of probability values, and any one of the candidate encodings is determined as an output encoding generated in the Nth round of reasoning.
[0198] In the first way of determining the output encoding, the probability values corresponding to all candidate encodings in the Nth round of reasoning can be compared. For the target encoding, the probability value used for comparison is the adjusted second probability value, and for the candidate encoding other than the target encoding, the probability value used for comparison is the first probability value.
[0199] After comparison, the maximum probability value among the probability values corresponding to all candidate encodings is determined, and the candidate encoding corresponding to the maximum probability value is determined as the output encoding.
[0200] As an example, after adjustment, the probability values corresponding to the five candidate encodings 55, 800, 592, 7301, and 100 obtained in the Nth round of reasoning are 14, 8, 0.2, 1, and 0.5, respectively. Among them, the probability values corresponding to the candidate encoding 55 and the candidate encoding 7301 are the adjusted second probability values, and the probability values corresponding to the other candidate encodings are the first probability values calculated by the large language model. Based on the above probability values, the candidate encoding 55 can be determined as the output encoding generated in the Nth round of reasoning from the above 5 candidate encodings.
[0201] As another example, after adjustment, the probability values corresponding to the five candidate encodings 55, 800, 592, 7301, 100 obtained by the Nth round of reasoning are 14, 16, 0.2, 1, and 0.5 in turn, where the probability values corresponding to the candidate encoding 55 and the candidate encoding 7301 are the second probability values after adjustment, and the probability values corresponding to the other candidate encodings are the first probability values calculated by the large language model. Based on the above probability values, the candidate encoding 800 can be determined as an output encoding generated by the Nth round of reasoning from the above five candidate encodings.
[0202] In the second way of determining the output encoding, all candidate encodings at the Nth round of reasoning can be sorted in descending order of the corresponding probability values. For the target encoding, the probability value used for sorting is the second probability value after adjustment, and for the candidate encodings other than the target encoding, the probability value used for sorting is the first probability value.
[0203] Then, any one of the top K candidate encodings in the sorted order is selected as an output encoding generated by the Nth round of reasoning. K can be any preset integer greater than 1, for example, it can be equal to 3, 5, or other values.
[0204] When selecting the output encoding from the top K candidate encodings, the probability of each candidate encoding in the top K candidate encodings being selected can be equal, or it can be positively correlated with the respective corresponding probability value, that is, the greater the corresponding probability value, the greater the probability of being selected as the output encoding.
[0205] Optionally, in the case where N is greater than 1, the way of obtaining the plurality of different pre-stored encodings included in the Nth column of the pre-stored intention table can also be:
[0206] Querying the candidate pre-stored encoding sequence in the pre-stored intention table to obtain a plurality of different pre-stored encodings included in the Nth column of the candidate pre-stored encoding sequence.
[0207] Among them, the pre-stored encoding of the N-1th column of the candidate pre-stored encoding sequence is the same as the output encoding generated by the N-1th round of reasoning.
[0208] That is, from the 2nd round of reasoning, when querying the Nth column of the pre-stored intention table each time, pre-stored encoding sequences whose pre-stored encodings of the N-1th column are different from the output encoding of the N-1th round of reasoning can be excluded, and only pre-stored encoding sequences whose pre-stored encodings of the N-1th column are different from the output encoding of the N-1th round of reasoning are queried.
[0209] Further, at each round of reasoning, the pre-stored encoding sequence can be further excluded based on the output encoding generated by the previous round of reasoning on the basis of the pre-stored encoding sequence remaining after the exclusion of the previous round of reasoning. In this way, the larger the number of rounds N of reasoning, the fewer the pre-stored encoding sequences that need to be queried, and the faster the query speed.
[0210] For example, assuming that there are 20 pre-stored encoding sequences in the pre-stored intention table, the output encoding generated at the end of the first round of reasoning is encoding 1, then in the second round of reasoning, 10 pre-stored encoding sequences in which the encoding in the first column is not encoding 1 are excluded from the pre-stored encoding sequences, and only the second column of the remaining 10 pre-stored encoding sequences in which the encoding in the first column is encoding 1 is queried, and the target encoding is determined based on the result of the query in the second round of reasoning.
[0211] The output encoding generated in the second round of reasoning is encoding 2, and in the third round of reasoning, 6 pre-stored encoding sequences in which the encoding in the second column is not encoding 2 are excluded from the 10 pre-stored encoding sequences in which the encoding in the first column is encoding 1, and only the third column of the remaining 4 pre-stored encoding sequences in which the encoding in the second column is encoding 2 is queried, and the target encoding is determined based on the result of the query in the third round of reasoning. By analogy.
[0212] In combination with the example in Table 2, assuming that the output encoding generated in the first round of reasoning is 55, then in the second round of reasoning, only the second column of the pre-stored encoding sequences in the first row, the second row and the fifth row of Table 2 is queried, and two different pre-stored encodings 643 and 6071 are obtained, so that the first probability value corresponding to 643 and the first probability value corresponding to 6071 in the first probability values of the candidate encodings obtained in the second round of reasoning can be adjusted.
[0213] Assuming that the output encoding generated in the second round of reasoning is 643, then in the third round of reasoning, the fifth row is excluded from the pre-stored encoding sequences in the first row, the second row and the fifth row, and only the third column of the first row and the second row of Table 2 is queried, and two different pre-stored encodings 4624 and 430 are obtained, so that the first probability value corresponding to 4624 and the first probability value corresponding to 430 in the first probability values of the candidate encodings obtained in the third round of reasoning can be adjusted.
[0214] The advantage of querying the pre-stored encoding in the above manner is that:
[0215] In the case where N is greater than 1, the number of pre-stored encoding sequences that need to be queried each time and the number of different pre-stored encodings in the Nth column obtained after the query can be reduced, and because the number of different pre-stored encodings in the Nth column obtained is reduced, the adjustment of the first probability values of the candidate encodings obtained in the Nth round of reasoning can also be completed faster, so that querying the pre-stored encoding in the above manner can improve the execution efficiency of the method of the present embodiment in the case where N is greater than 1.
[0216] The present application also provides an electronic device, please see Figure 5 , comprising:
[0217] The memory 501 is configured to store a large language model.
[0218] The at least one processor 502 is configured to:
[0219] obtaining first input information;
[0220] converting the first input information into a plurality of input tokens by a tokenizer, the input token being a minimum unit for text processing by a large language model;
[0221] obtaining an input encoding of each input token based on a vocabulary to form an input encoding sequence;
[0222] generating output information composed of a plurality of output encodings through a plurality of rounds of reasoning based on the input encoding sequence and the large language model, one output encoding being generated in each round of reasoning;
[0223] wherein one output encoding is generated in each round of reasoning includes:
[0224] calculating a first probability value of a candidate encoding in the Nth round based on the encoding in the (N-1)th round, N being an integer greater than or equal to 1;
[0225] querying a pre-stored intent table to obtain a plurality of different pre-stored encodings included in the Nth column, the pre-stored intent table including pre-stored encodings corresponding to a plurality of intent items;
[0226] adjusting the first probability value of the target encoding to a second probability value, the target encoding being an encoding consistent with the plurality of different pre-stored encodings included in the Nth column among the candidate encodings;
[0227] determining one candidate encoding as one output encoding generated in the Nth round of reasoning based on the probability value of the candidate encoding.
[0228] Optionally, the large language model used in the method of the present embodiment can be a large language model stored locally in the electronic device, that is, the weight parameters contained in the large language model are stored in the memory of the electronic device in the form of data. Therefore, in order to save storage space, the model parameters of the large language model stored in the memory can be in a lower precision data format, for example, an integer occupying 4 bits or 1 bit, and correspondingly, the training of the large language model is usually performed on a server device with sufficient computing power and storage resources, so that the model parameters of the large language model during training can be in a higher precision data format, such as a 32-bit floating point number.
[0229] That is, the precision of the model parameters of the large language model stored in the memory of the electronic device of the present embodiment can be lower than the precision of the model parameters of the large language model during training when the electronic device applies the large language model stored in the memory for reasoning.
[0230] These low-precision model parameters applied during reasoning can be obtained by model quantization processing of high-precision model parameters in the trained large language model.
[0231] Optionally, the processor 502 is further configured to:
[0232] execute the operation instruction corresponding to the output information;
[0233] or, output the word units corresponding to each output encoding in the output information.
[0234] Optionally, one output encoding generated in the Nth round of inference is used as the encoding in the (N+1)th round of inference.
[0235] The encoding in the Nth round is used for the (N+1)th round of inference.
[0236] Optionally, the processor 502 adjusts the first probability value of the target encoding to a second probability value, including:
[0237] multiply the first probability value of the target encoding by a coefficient to obtain the second probability value.
[0238] Optionally, if the one output encoding generated in the Nth round of inference belongs to the target encoding, the second probability value of the target encoding in the candidate encoding representing the process of generating the one output encoding in the Nth round of inference is greater than the probability value of each candidate encoding except the target encoding.
[0239] Optionally, if the one output encoding generated in the Nth round of inference does not belong to the target encoding, the second probability value of the target encoding in the candidate encoding representing the process of generating the one output encoding in the Nth round of inference is not greater than the probability value of each candidate encoding except the target encoding.
[0240] Optionally, the coefficient is determined according to the output information when the large language model processes the sample input information.
[0241] Optionally, the processor 502 determines one candidate encoding as the one output encoding generated in the Nth round of inference according to the probability values of the candidate encodings, including:
[0242] determining the candidate encoding with the maximum probability value as the one output encoding generated in the Nth round of inference based on the probability values of the candidate encodings;
[0243] or, determining a plurality of candidate encodings in descending order of probability values and determining any one of the candidate encodings as the one output encoding generated in the Nth round of inference.
[0244] Optionally, N is greater than 1.
[0245] The processor 502 queries the pre-stored intent table to obtain a plurality of different pre-stored encodings included in the Nth column, including:
[0246] querying the candidate pre-stored encoding sequence in the pre-stored intent table to obtain a plurality of different pre-stored encodings included in the Nth column of the candidate pre-stored encoding sequence.
[0247] wherein the pre-stored encoding of the N-1th column of the candidate pre-stored encoding sequence is identical to one output encoding generated by the N-1th round of inference.
[0248] The working principle of the electronic device of the embodiment can refer to the related steps of the model-based inference method in the foregoing embodiments, and will not be described herein.
[0249] The embodiment provides an agent, including executable computer instructions.
[0250] The agent is called to implement:
[0251] obtain first input information;
[0252] convert the first input information into a plurality of input tokens through a tokenizer, the input token being a minimum unit for text processing by a large language model;
[0253] obtain an input encoding of each input token based on a vocabulary to constitute an input encoding sequence;
[0254] generate output information composed of a plurality of output encodings through a plurality of rounds of inference based on the input encoding sequence and the large language model, one output encoding being generated by each round of inference in the plurality of rounds of inference;
[0255] wherein generating one output encoding by each round of inference includes:
[0256] calculate a first probability value of a candidate encoding of the Nth round based on the encoding of the N-1th round, N being an integer greater than or equal to 1;
[0257] query a pre-stored intent table to obtain a plurality of different pre-stored encodings included in the Nth column, the pre-stored intent table including pre-stored encodings corresponding to a plurality of intent items;
[0258] adjust the first probability value of the target encoding to a second probability value, the target encoding being an encoding consistent with the plurality of different pre-stored encodings included in the Nth column among the candidate encodings;
[0259] determine one candidate encoding as one output encoding generated by the Nth round of inference based on the probability value of the candidate encoding.
[0260] In some embodiments, the agent can be installed in a terminal-type electronic device, such as a smartphone, a desktop or notebook computer used by an individual.
[0261] An agent is an application of artificial intelligence technology, which can be implemented based on a target model (such as a large language model), and the behavior of the agent is determined by the target model in the agent according to the current state and external input, which can be used to solve a certain type of problem. Among them, the target model provides the basis for decision-making through the learning and training of a large amount of data; the agent can also call tools, plug-ins and knowledge bases to provide reasoning, decision-making and execution capabilities.
[0262] The way the agent is installed on the terminal type electronic device can be:
[0263] The agent is embedded in the application program such as the browser and the application store as part of the application program, thereby providing intelligent functions for the application program when the application program is running;
[0264] The agent is embedded in the operating system of the electronic device as part of the operating system, thereby providing intelligent functions for the operating system during the operation of the operating system.
[0265] In the case of being embedded in the application program or the operating system, the application program or the operating system can call the agent to provide intelligent functions in various ways. Some examples of calling the agent include: the agent can be called to provide intelligent functions when a specific trigger operation is recognized, such as when a specific key is triggered; or the agent can be called to provide intelligent functions when any input information is obtained, such as when the user's voice is obtained; or the agent can be called to provide intelligent functions in a specific application scenario, such as when browsing a webpage with a browser, browsing goods with a shopping application, etc.
[0266] The specific working principle of the above agent can be referred to the related steps of the model-based reasoning method in the foregoing embodiments, which will not be described here.
[0267] It should be noted that each embodiment in the present specification adopts a progressive manner for description, and each embodiment focuses on the difference from other embodiments. The same and similar parts between the embodiments can be referred to each other.
[0268] For the convenience of description, the above system or device is described as various modules or units respectively described in terms of functions. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware in the implementation of the present application.
[0269] Those skilled in the art can clearly understand the application by the description of the above embodiments. The technical solutions of the application can be implemented by means of software and necessary universal hardware platforms. Based on such an understanding, the technical solutions of the application can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the application.
[0270] Finally, it should be noted that the terms such as first, second, third, and fourth, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0271] The above description is only the preferred embodiments of the application, and it should be pointed out that those skilled in the art can make some improvements and refinements without departing from the principles of the application, and these improvements and refinements should also be regarded as the protection scope of the application.
Claims
1. A model-based reasoning method, comprising: Obtain the first input information; The first input information is converted into multiple input tokens by a word segmenter. The input token is the smallest unit for text processing in a large language model. The input code for each input word is obtained based on the vocabulary to form an input code sequence; Based on the input encoding sequence and the large language model, output information consisting of multiple output codes is generated through multiple rounds of reasoning, with each round of reasoning generating one output code. Each round of reasoning generates an output code, including: Based on the encoding of the (N-1)th round, calculate the first probability value to obtain the candidate encoding of the Nth round, where N is an integer greater than or equal to 1; Query the pre-stored intent chart to obtain multiple different pre-stored codes included in the Nth column. The pre-stored intent chart includes pre-stored codes corresponding to multiple intent items. The first probability value of the target code is adjusted to the second probability value, wherein the target code is the candidate code that is consistent with the multiple different pre-stored codes included in the Nth column; A candidate code is determined based on the probability value of the candidate code and used as an output code generated in the Nth round of inference.
2. The method according to claim 1, further comprising: Execute the operation instructions corresponding to the output information; Alternatively, output the tokens corresponding to each output code in the output information.
3. The method according to claim 1, wherein the encoding of the (N-1)th round and an output encoding generated by the reasoning of the Nth round are used as the encoding of the Nth round; The encoding of the Nth round is used for the reasoning of the N+1th round.
4. The method according to claim 1, wherein adjusting the first probability value of the target encoding to the second probability value comprises: The second probability value is obtained by multiplying the first probability value of the target encoding by a coefficient.
5. The method according to claim 4, wherein if an output code generated by the Nth round of inference belongs to the target code, the second probability value of the target code in the candidate codes of the output code generated by the Nth round of inference is greater than the probability value of each candidate code other than the target code.
6. The method according to claim 4, wherein if an output code generated by the Nth round of inference does not belong to the target code, the second probability value of the target code in the candidate codes representing the output code generated by the Nth round of inference is not greater than the probability value of each candidate code other than the target code.
7. The method according to claim 1, wherein determining a candidate code as an output code generated in the Nth round of inference based on the probability value of the candidate code comprises: Based on the probability values of the candidate codes, the candidate code with the highest probability value is used as an output code generated in the Nth round of inference; Alternatively, multiple candidate codes can be determined in descending order of probability value, and any one of these candidate codes can be used as an output code generated in the Nth round of reasoning.
8. The method according to claim 1, wherein N is greater than 1; The query pre-stored meaning chart obtains multiple different pre-stored codes included in the Nth column, including: Query the candidate pre-stored encoding sequence in the pre-stored intent table to obtain multiple different pre-stored codes included in the Nth column of the candidate pre-stored encoding sequence; The pre-stored code of the N-1th column of the candidate pre-stored code sequence is the same as an output code generated in the N-1th round of reasoning.
9. An electronic device, comprising: Memory, used to store large language models; Processor, used for: Obtain the first input information; The first input information is converted into multiple input tokens by a word segmenter. The input token is the smallest unit for text processing in a large language model. The input code for each input word is obtained based on the vocabulary to form an input code sequence; Based on the input encoding sequence and the large language model, output information consisting of multiple output codes is generated through multiple rounds of reasoning, with each round of reasoning generating one output code. Each round of reasoning generates an output code, including: Based on the encoding of the (N-1)th round, calculate the first probability value to obtain the candidate encoding of the Nth round, where N is an integer greater than or equal to 1; Query the pre-stored intent chart to obtain multiple different pre-stored codes included in the Nth column. The pre-stored intent chart includes pre-stored codes corresponding to multiple intent items. The first probability value of the target code is adjusted to the second probability value, wherein the target code is the candidate code that is consistent with the multiple different pre-stored codes included in the Nth column; A candidate code is determined based on the probability value of the candidate code and used as an output code generated in the Nth round of inference.
10. An intelligent agent, comprising executable computer instructions; The intelligent agent is invoked to implement: Obtain the first input information; The first input information is converted into multiple input tokens by a word segmenter. The input token is the smallest unit for text processing in a large language model. The input code for each input word is obtained based on the vocabulary to form an input code sequence; Based on the input encoding sequence and the large language model, output information consisting of multiple output codes is generated through multiple rounds of reasoning, with each round of reasoning generating one output code. Each round of reasoning generates an output code, including: Based on the encoding of the (N-1)th round, calculate the first probability value to obtain the candidate encoding of the Nth round, where N is an integer greater than or equal to 1; Query the pre-stored intent chart to obtain multiple different pre-stored codes included in the Nth column. The pre-stored intent chart includes pre-stored codes corresponding to multiple intent items. The first probability value of the target code is adjusted to the second probability value, wherein the target code is the candidate code that is consistent with the multiple different pre-stored codes included in the Nth column; A candidate code is determined based on the probability value of the candidate code and used as an output code generated in the Nth round of inference.