Large model decoding constraint method, device and medium

By building a joint constraint model based on FST and DFA, the output of large language model is constrained, and the problem of generating invalid output is solved, achieving higher controllability and dynamic optimization capabilities.

CN119830959BActive Publication Date: 2025-05-13INSPUR GENERSOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510300772.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-13
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Large language models lack controllability in the content generation process, and may generate invalid or unexecutable plans, affecting the quality of task completion.

Method used

A joint constraint model is adopted based on finite state converter (FST) and finite state automaton (DFA), and the minimum text unit output from the large language model is restored to a string through the FST model, and the output is verified by the DFA model to form a joint constraint model to exclude illegal vocabulary.

Benefits of technology

It realizes controllability of output of large language models, avoids invalid output, and supports dynamic structure optimization, balances generation freedom and controllability, and improves dynamic controllability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830959B_ABST
    Figure CN119830959B_ABST
Patent Text Reader

Abstract

The present application discloses a large model decoding constraint method, device and medium, which relates to the field of model processing. The method includes: forming a joint constraint model based on an FST model and a DFA model; decoding based on a large language model to generate a probability score for a corresponding vocabulary; excluding the minimum text unit in the vocabulary based on the joint constraint model, and outputting a specified minimum text unit; dynamically adjusting the structure of the specified minimum text unit based on factors to be determined, and updating the output of the large language model. The output format is not forced to be fixed, and the model is allowed to freely generate intermediate reasoning steps. Compared with traditional structured generation, it avoids conflicts with natural generation logic and retains generation flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of model processing, and in particular to a large model decoding constraint method, device and medium. Background Art

[0002] With the development of technology, Large Language Model (LLM) has gradually come into people's view.

[0003] Although large language models have made remarkable achievements in creativity and flexibility, controllability issues in the content generation process have gradually emerged. For example, due to the lack of sufficient supervision in the generation process of large language models, they sometimes generate invalid or unexecutable plans, affecting the quality of task completion.

[0004] In traditional solutions, a common solution is structured generation, which uses format restrictions to allow large language models to provide output in standardized formats such as JSON or XML. However, most structured generation methods limit the model's ability to generate necessary intermediate reasoning steps, and the mandatory format requirements may be incompatible with the way the model naturally generates answers. Summary of the invention

[0005] In order to solve the above problems, this application proposes a large model decoding constraint method, including:

[0006] Based on the finite state converter, an FST model is generated to restore the smallest text unit output by the large language model to a string; and based on the deterministic finite state automaton, a DFA model is generated by compiling the regular expression to verify the constraint conditions of the string output by the FST model;

[0007] Based on the FST model and the DFA model, a joint constraint model is formed;

[0008] Decode based on the large language model and generate probability scores for the corresponding vocabulary;

[0009] Excluding the minimum text unit in the vocabulary based on the joint constraint model, and outputting the specified minimum text unit;

[0010] Based on the factors to be determined, the structure of the specified minimum text unit is dynamically adjusted to update the output of the large language model.

[0011] In one example, based on a finite state converter, an FST model for restoring a minimum text unit output by a large language model to a string is generated, specifically including:

[0012] Based on the finite state converter, generate the FST model, initialize it, and store the root state into the FST structure;

[0013] For the vocabulary, traverse each word in it and perform vocabulary conversion processing on each word until all words are traversed;

[0014] The vocabulary conversion process includes:

[0015] For each word, starting from the first character of the word, set the input label of each character to empty, set the output label to the current character, and update the edge set and state set;

[0016] Loop through all characters of the vocabulary until the last character is output, use the complete vocabulary as the input label, and return the state to the root node.

[0017] In one example, based on determining a finite state automaton, by compiling a regular expression, a DFA model for verifying the constraint conditions of a string output by the FST model is generated, specifically including:

[0018] Get a predefined regular expression;

[0019] Based on determining a finite state automaton, compiling the regular expression to generate a DFA model;

[0020] Initializing the DFA model;

[0021] The character string output by the FST model is obtained through the DFA model, and each character in the character string is verified in turn according to the constraint conditions obtained by compiling the regular expression until the character string is completely verified.

[0022] In one example, the minimum text unit in the vocabulary is excluded based on the joint constraint model, and the output specified minimum text unit specifically includes:

[0023] Based on the joint constraint model, the smallest text units in the vocabulary are verified in turn through the regular expression constraints corresponding to the regular expression, and illegal words are excluded, and the remaining legal words form a candidate set;

[0024] Based on a preset sampling algorithm, a specified minimum text unit is selected from the candidate set as output.

[0025] In one example, based on the factors to be determined, dynamically adjusting the structure of the specified minimum text unit and updating the output of the large language model specifically includes:

[0026] Mark the first specified minimum text unit outputted last as a factor to be determined;

[0027] Obtaining a second specified minimum text unit generated in the next round, combining the factor to be determined with the second specified minimum text unit, and determining whether the obtained combination meets the constraint condition corresponding to the regular expression;

[0028] If it is in accordance with the combination, the factor to be determined is updated according to the combination;

[0029] If not, the factor to be determined is directed to a new specified minimum text unit.

[0030] In one example, determining whether the obtained combination meets the constraint condition corresponding to the regular expression specifically includes:

[0031] Determine the current length corresponding to the combination;

[0032] If the current length exceeds the preset length threshold, a fallback strategy is triggered. In the combination, the latest obtained number of specified minimum text units are retained, the remaining specified minimum text units are deleted, and the number of deletions is recorded.

[0033] Based on the regular expression and the number of deletions, a corresponding starting judgment position is selected;

[0034] According to the starting judgment position, it is judged whether the combination meets the constraint condition corresponding to the regular expression.

[0035] In one example, the method further includes:

[0036] Whenever the factors to be determined are updated, a cumulative probability score is obtained based on the probability scores corresponding to each specified minimum text unit in the vocabulary in the factors to be determined;

[0037] If the cumulative probability score is lower than the corresponding preset dynamic score, starting from the last specified minimum text unit in the factor to be determined, it is gradually replaced by other specified minimum text units until the cumulative probability score is higher than the corresponding preset dynamic score.

[0038] In one example, the method further includes:

[0039] Obtaining a preset dynamic score corresponding to the previous length of the factor to be determined;

[0040] Based on the standard deviation and mean of the probability score of each word in the vocabulary, estimating the interval in which the probability score lies;

[0041] Based on the maximum value of the interval and the preset dynamic score corresponding to the previous length, the preset dynamic score corresponding to the current length is obtained.

[0042] On the other hand, the present application also proposes a large model decoding constraint device, comprising:

[0043] at least one processor; and,

[0044] a memory communicatively connected to the at least one processor; wherein,

[0045] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the large model decoding constraint method as described in any of the above examples.

[0046] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to be: the large model decoding constraint method described in any of the above examples.

[0047] The large model decoding constraint method proposed in this application can bring the following beneficial effects:

[0048] 1. It does not force a fixed output format, allowing the model to freely generate intermediate reasoning steps. Compared with traditional structured generation (JSON / XML), it avoids conflicts with natural generation logic and retains generation flexibility.

[0049] 2. Real-time screening and adjustment of the minimum text unit through the joint constraint model (FST+DFA) not only constrains invalid output, but also supports dynamic structural optimization, balances the generation freedom and controllability, and improves dynamic controllability. In addition, the DFA model quickly verifies string constraints by compiling regular expressions, which is lighter than traditional post-processing verification and reduces the probability of generating invalid content.

[0050] 3. Based on the vocabulary probability score and minimum unit exclusion, the constraint process adapts the model's native decoding logic, without the need for forced format rewriting, reducing generation performance loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0052] Figure 1 This is a flow chart of a large model decoding constraint method in an embodiment of the present application;

[0053] Figure 2 A schematic diagram of a large model decoding constraint method in one scenario in an embodiment of the present application;

[0054] Figure 3It is a schematic diagram of a large model decoding constraint device in an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0056] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0057] like Figure 1 and Figure 2 As shown, the embodiment of the present application provides a large model decoding constraint method, including:

[0058] S101: Generate an FST model based on a finite state converter to restore the smallest text unit output by a large language model to a string; and generate a DFA model based on a deterministic finite state automaton to compile an input regular expression and verify the string output by the FST model.

[0059] Finite State Automaton (FSA) is a computational model used to process and analyze strings or symbol sequences. The FSA model can be formally represented as A collection of Represents a finite set of input symbols that the FSA model can accept. These symbols can be characters, numbers, or other types of data. represents a finite set of all possible states in the FSA model, each state representing a certain stage or condition when processing the input symbol; and They represent the initial and final states of the FSA model respectively; Represents the set of edges that transfer from one state to the next state, where , is an empty symbol, each edge It can be expressed as , For the original state, For input tags, is the target state.

[0060] The finite automaton (DFA) model is a type of FSA model. For each state and input symbol, the state transition function uniquely determines the next state. In other words, in the DFA model, given a state and an input symbol, there will not be multiple possible state transitions.

[0061] If the DFA model is in one of the terminal states after processing the input string, then the input string is considered to be accepted. The DFA model can be used in the lexical analysis stage to decompose the input source code string into a series of minimal text units (tokens). Each token represents an element in the source code, for example, it can be a keyword, identifier, operator or numeric constant.

[0062] By constructing an appropriate FSA model (DFA model is used in this application), efficient recognition and matching of specific languages ​​or patterns can be achieved. Regular languages ​​can be defined by regular expressions, and regular expressions can be compiled to obtain DFA models. Here, a regular expression search tool can be used to compile regular expressions into The DFA model is then used as the judgment formula of the text to verify the string.

[0063] The finite state transducer (FST) is an extended computational model of the FSA model. The FST model can be formally expressed as A collection of Consistent with the definition of the FSA model, is the set of output symbols, Represents a set of edges, each edge It can be expressed as , For the original state, For input tags, is the output label, is the target state.

[0064] Unlike the FSA model, the FST model not only receives input strings but also produces output strings, so it is used to model the relationship between input and output, such as string conversion and mapping.

[0065] Based on this, Figure 2As shown, a series of minimum text unit tokens of the preliminary output (not the final output) of the large language model LLM are restored to string expressions based on the vocabulary by the FST model, and the FSA model (the DFA model is used in the embodiment of the present application) can verify the string output by the FST model through regular expressions.

[0066] Here is a pseudo code of the FST model for vocabulary conversion, which is explained as an example:

[0067] ← , ←{ } / / Input is tokens, output is string

[0068] ← { }, ← , ← { }, ← {} / / Initialization, root state

[0069] for do:

[0070] ← , ←

[0071] for to do:

[0072] ← { }, ← {( , , , )}

[0073] ← {( , , , )} / / Last character operation, return to the root node

[0074] Pseudo code for vocabulary conversion based on the FST model. In the process of vocabulary conversion of the FST model, the FST model is generated based on the finite state converter and initialized, and the root state Store into FST structure;

[0075] For vocabulary , iterate through each word in it , and perform vocabulary conversion on each word until all words are traversed.

[0076] The vocabulary conversion process includes:

[0077] For each word , starting from the first character of the word, and continuing until the n-1th character (assuming there are n characters in total), set the input label to the empty symbol for each character , and set the output label to the current character , and update the edge set and state collection .

[0078] Loop through all characters of the word until the last character is output , the complete word As input label, return the state to the root node.

[0079] For the DFA model, a predefined regular expression is obtained. A regular expression is a syntax for describing a string pattern (for example, \d{4}-\d{2} represents "4 digits + hyphen + 2 digits"), which can be used to represent the year and month in time.

[0080] Based on the deterministic finite state automaton, the regular expression is compiled to generate a DFA model. The regular expression can be first converted into an NFA (non-deterministic finite state automaton) through recursive decomposition, and multiple transfer paths are allowed. Then, the NFA is converted into an equivalent DFA to eliminate non-determinism.

[0081] Initialize the DFA model, obtain the string output by the FST model through the DFA model, and verify each character in the string in turn according to the constraints obtained by compiling the regular expression until the string is verified. In the verification process of each character, the current state of the character and the input character can be determined, and the response of the state transition function to the character can be obtained by looking up the table, and the state is updated until the last character is verified, the termination state is determined and verified, and the conclusion is drawn as to whether the string is accepted.

[0082] S102: forming a joint constraint model based on the FST model and the DFA model.

[0083] like Figure 2 As shown, the output of the FST model is associated with the input of the DFA model, and the string output by the FST model is verified using the constraints in the DFA model, thereby forming a joint constraint model.

[0084] S103: Decoding is performed based on the large language model to generate a probability score of the corresponding vocabulary.

[0085] When a user enters text as a query or prompt, the large language model first encodes the input and converts it into a vector representation that the model can process. Then, based on its trained knowledge and parameters, the model begins to generate the corresponding output. During this generation process, the model decodes position by position, predicts the next possible minimum text unit, and outputs it.

[0086] During the decoding process, the large language model generates a corresponding vocabulary, which contains multiple words and the probability score corresponding to each word. This vocabulary can be used as the smallest text unit in the decoding generation process.

[0087] The vocabulary table can be pre-built after data collection, cleaning, word segmentation, and word screening. The probability score of each word is calculated based on the generated words and the context sequence, which will not be described here.

[0088] S104: Excluding the minimum text units in the vocabulary based on the joint constraint model, and outputting a specified minimum text unit.

[0089] The joint constraint model includes corresponding constraint conditions. Based on the joint constraint model, the smallest text units in the vocabulary can be verified in turn through the regular expression constraints corresponding to the regular expression, and illegal words can be excluded, and the remaining legal words can form a candidate set.

[0090] For example, taking the regular expression "\d{4}-\d{2}" (YYYY-MM format) and the vocabulary ["2023", "1999","-", "12", "ab"] as an example, in the initial generation stage, the candidate minimum text units include "2023", "1999", "-", "12", "ab".

[0091] When "2023" or "1999" is selected, the string is "2023" or "1999", which is legal but incomplete. When other candidate minimum text units are selected, they do not start with a 4-digit number and are illegal words.

[0092] Next is the intermediate generation stage. Assuming that the current string is "2023-", among the remaining candidate minimum text units "12", "31", and "ab", through verification, it is finally determined that "12" and "31" are legal words, and "ab" is an illegal word.

[0093] When selecting a candidate minimum text unit as the designated minimum text unit, the designated minimum text unit may be selected from the candidate set based on a preset sampling algorithm as an output.

[0094] The sampling strategy can be greedy sampling, which directly selects the word with the highest probability from all candidate minimum text units as the specified minimum text unit. Alternatively, random sampling can be selected, which randomly selects from the top k candidate minimum text units with the highest probability.

[0095] Still taking the example above as an example, after determining that the current string is "2023-", the subsequent "12" and "31" are both legal words. At this time, the sampling strategy can be used to select the corresponding candidate minimum text unit as the designated minimum text unit.

[0096] S105: Based on the factors to be determined, dynamically adjust the structure of the specified minimum text unit to update the output of the large language model.

[0097] Since the generation of the large language model (LLM) is based on probability distribution and may be affected by factors such as hallucination and non-robustness, the generated intermediate results may not fully conform to the target grammar.

[0098] Based on this, in order to effectively manage the uncertainty of the LLM generation process, the to-be-determined factor (Rest) is designed. By retaining the minimum text unit token that may change in the future, the structure of the minimum text unit can be dynamically adjusted during the generation process. The last minimum text unit token in the current iteration is regarded as the Rest that cannot be finalized, and the Rest can be re-determined in the subsequent minimum text unit generation process.

[0099] Here is a pseudo code based on dynamic adjustment of the factors to be determined, which is explained as an example:

[0100] ← DETOKENZATION_FST( ) / / Use FST to convert tokens into string original expression

[0101] ← REGEX_DFA( ) / / Build FSA based on regular expression

[0102] ← / / Combined FSA processing tokens

[0103] ← / / Composite FSA start state

[0104] Rest ← / / Initialize the factors to be determined

[0105] for step=1 to T do: / / decoding step

[0106] Score ← COMPUTELOGITS(LLM) / / Large model generates vocabulary probability score

[0107] ← { }

[0108] for to do: / / Check vocabulary

[0109] if thenScore[ ]← -

[0110] ← SAMPLE_NEXT_TOKEN(Score) / / Sampling selection token

[0111] if Rest. then = Rest. / / Connect the factor to be determined and the token, and update the output

[0112] Rest ← / / Update the factors to be determined

[0113] ← , / / Update status

[0114] For the pseudo code based on the dynamic adjustment of the factor to be determined, represents the vocabulary, Represents a regular expression constraint. DETOKENZATION_FST ( ) function converts the vocabulary into an FST model, which can be pre-calculated; REGEX_DFA( ) function converts the regular expression constraint into a DFA model; COMPUTELOGITS(LLM) function represents the probability score of the corresponding vocabulary generated by the large language model LLM; SAMPLE_NEXT_TOKEN(Score) function represents the determination of the next specified minimum text unit token according to the Score value.

[0115] Specifically, in the LLM generation process, the FST and FDA composite computational model is first used to constrain the large model to decode tokens that conform to the specific grammar (for the convenience of description, the token currently output last is called the first specified minimum text unit), and then the current token is marked as a factor to be determined.

[0116] In the next generation phase, a new token generated in the next round is obtained (for the convenience of description, it is called the second specified minimum text unit), and the factor to be determined is combined with the second specified minimum text unit, and it is determined whether the obtained combination meets the constraints corresponding to the regular expression. When the constraints are met, the combination is considered to be accepted.

[0117] If the combination conforms, that is, the combination conforms to the regular expression constraint corresponding to the regular expression, then the factor to be determined is updated according to the combination, that is, the factor to be determined is updated to the combination, and the subsequent new round of generation of the specified minimum text unit is continued.

[0118] If the combination does not meet the requirements, the factor to be determined is directed to a new specified minimum text unit, that is, the second specified minimum text unit is reselected and the regular constraint condition is determined again.

[0119] 1. It does not force a fixed output format, allowing the model to freely generate intermediate reasoning steps. Compared with traditional structured generation (JSON / XML), it avoids conflicts with natural generation logic and retains generation flexibility.

[0120] 2. Real-time screening and adjustment of the minimum text unit through the joint constraint model (FST+DFA) not only constrains invalid output, but also supports dynamic structural optimization, balances the generation freedom and controllability, and improves dynamic controllability. In addition, the DFA model quickly verifies string constraints by compiling regular expressions, which is lighter than traditional post-processing verification and reduces the probability of generating invalid content.

[0121] 3. Based on the vocabulary probability score and minimum unit exclusion, the constraint process adapts the model's native decoding logic, without the need for forced format rewriting, reducing generation performance loss.

[0122] In one embodiment, when determining whether the obtained combination meets the constraint conditions corresponding to the regular expression, considering that some regular expressions are long, repeated calculations are likely to increase the splicing and verification overhead, so the current length corresponding to the combination is determined.

[0123] If the current length exceeds the preset length threshold (the length refers to the number of minimum text units contained therein), it is considered that the length is too long and the fallback strategy needs to be triggered.

[0124] The fallback strategy means that in the combination, the latest specified minimum text units are retained, the remaining specified minimum text units are deleted, and the number of deletions is recorded. For example, the preset length threshold is 8, the number of retained specified minimum text units is 5, and when the current length corresponding to the combination reaches 8, the latest 5 specified minimum text units are retained, the first 3 specified minimum text units are deleted, and the number of deletions is recorded, that is, 3 are deleted.

[0125] Based on the regular expression and the number of deletions, the corresponding starting position is selected. For example, when the number of deletions is 3, the 4th specified minimum text unit is used as the starting position. Of course, if the specified minimum text unit has been deleted before, the number of deletions can be accumulated.

[0126] According to the starting judgment position, it is judged whether the combination meets the constraint conditions corresponding to the regular expression. That is, when judging whether the constraint conditions are met, it is enough to start judging from the starting judgment position. Through the fallback strategy, the verification overhead in the judgment process is reduced, and the calculation pressure is reduced.

[0127] In one embodiment, in order to consider the reliability of the generated content from a global perspective, whenever the factor to be determined is updated, the cumulative probability score is obtained based on the probability score corresponding to each specified minimum text unit in the vocabulary in the factor to be determined. When the factor to be determined is a single specified minimum text unit, the cumulative probability score is its probability score in the vocabulary. When the factor to be determined is a combination of two specified minimum text units, the cumulative probability score is the product of the probability scores corresponding to each of the two specified minimum text units, and so on.

[0128] If the cumulative probability score is lower than the corresponding preset dynamic score (the corresponding preset dynamic score can be set for the length of each factor to be determined), then starting from the last specified minimum text unit in the factor to be determined, it is gradually replaced by other specified minimum text units until the cumulative probability score is higher than the corresponding preset dynamic score. In other words, as the factors to be determined are updated, the actual output content is also gradually increasing, and the cumulative probability score will also decrease accordingly. If it is lower than a certain level (that is, the preset dynamic score), it is considered that the content of this output is unqualified as a whole. At this time, adjustments are made starting from the last specified minimum text unit, and it is still changed to other specified minimum text units through the sampling algorithm. If it still does not meet the requirements, continue to change. If all the last specified minimum text units do not meet the preset dynamic score corresponding to their length, the penultimate specified minimum text unit is changed to other specified minimum text units (of course, it must meet the preset dynamic score corresponding to the length of the second specified minimum text unit), and then the last specified minimum text unit is changed, and so on, until the content corresponding to the current regular expression is generated. Of course, there may be extreme cases where the cumulative probability scores of all generated content fail to meet all requirements of the preset dynamic score. In this case, the content with the highest cumulative probability score can be selected for generation.

[0129] Among them, as the more content is generated, the cumulative probability score will decrease accordingly, so the preset dynamic score is set to be dynamic and not a constant value.

[0130] Specifically, the preset dynamic score corresponding to the previous length of the factor to be determined is obtained. If the current length is 1, that is, there is no preset dynamic score corresponding to the previous length, a default value (eg, 0.6) can be selected as the preset dynamic score corresponding to the previous length.

[0131] Based on the standard deviation and mean of the probability score of each word in the vocabulary, the probability score range is estimated. Considering that the distribution of probability scores may be different in different vocabularies, the preset dynamic score is modified based on the standard deviation and mean of the probability score of each word in the vocabulary (here refers to the legal words).

[0132] Based on the standard deviation and mean of the probability scores of each word in the vocabulary, the probability score range is estimated. The standard deviation reflects the degree of dispersion between words, and the mean reflects the concentration trend, which can roughly estimate the range where most of the data is located. For example, when the mean is 0.6 and the standard deviation is 0.2, it can be reflected that most of the probability scores are concentrated in the range of 0.4~0.8.

[0133] Based on the maximum value of the interval and the preset dynamic score corresponding to the previous length, the preset dynamic score corresponding to the current length is obtained. For example, assuming the current length is 2, the preset dynamic score corresponding to the previous length is 0.6. Since the content output by the large language model itself needs to find words with relatively high probability scores, considering that the maximum value of the interval is both relatively high and can radiate to a relatively large number of words, the maximum value of 0.8 is selected instead of the default value of 0.6, which is more in line with the overall probability score of the words in the current vocabulary. Calculate the product of the maximum value 0.8 and the exponent 0.6, 0.48, which can be used as the preset dynamic score corresponding to the current length.

[0134] like Figure 3 As shown, the embodiment of the present application also proposes a large model decoding constraint device, including:

[0135] at least one processor; and,

[0136] a memory communicatively connected to the at least one processor; wherein,

[0137] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the large model decoding constraint method as described in any of the above embodiments.

[0138] The present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to be: the large model decoding constraint method described in any of the above embodiments.

[0139] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0140] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0141] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A large model decoding constraint method, characterized in that: include: Based on the finite state converter, an FST model is generated to restore the smallest text unit output by the large language model to a string; Based on the deterministic finite state automaton, a DFA model for verifying the constraint conditions of the string output by the FST model is generated by compiling the regular expression; Based on the FST model and the DFA model, a joint constraint model is formed; Decode based on the large language model and generate probability scores for the corresponding vocabulary; Excluding the minimum text unit in the vocabulary based on the joint constraint model, and outputting the specified minimum text unit; Based on the factors to be determined, the structure of the specified minimum text unit is dynamically adjusted to update the output of the large language model.

2. The large model decoding constraint method according to claim 1, characterized in that: Based on the finite state converter, an FST model is generated to restore the smallest text unit output by the large language model to a string, including: Based on the finite state converter, generate the FST model, initialize it, and store the root state into the FST structure; For the vocabulary, traverse each word in it and perform vocabulary conversion processing on each word until all words are traversed; The vocabulary conversion process includes: For each word, starting from the first character of the word, set the input label of each character to empty, set the output label to the current character, and update the edge set and state set; Loop through all characters of the vocabulary until the last character is output, use the complete vocabulary as the input label, and return the state to the root node.

3. The large model decoding constraint method according to claim 1, characterized in that: Based on the deterministic finite state automaton, by compiling the regular expression, a DFA model is generated for verifying the constraint conditions of the string output by the FST model, including: Get a predefined regular expression; Based on determining a finite state automaton, compiling the regular expression to generate a DFA model; Initializing the DFA model; The character string output by the FST model is obtained through the DFA model, and each character in the character string is verified in turn according to the constraint conditions obtained by compiling the regular expression until the character string is completely verified.

4. The large model decoding constraint method according to claim 3, characterized in that: The minimum text unit in the vocabulary is excluded based on the joint constraint model, and the output specified minimum text unit specifically includes: Based on the joint constraint model, the smallest text units in the vocabulary are verified in turn through the regular expression constraints corresponding to the regular expression, and illegal words are excluded, and the remaining legal words form a candidate set; Based on a preset sampling algorithm, a specified minimum text unit is selected from the candidate set as output.

5. The large model decoding constraint method according to claim 1, characterized in that: Based on the factors to be determined, dynamically adjusting the structure of the specified minimum text unit and updating the output of the large language model specifically includes: Mark the first specified minimum text unit outputted last as a factor to be determined; Obtaining a second specified minimum text unit generated in the next round, combining the factor to be determined with the second specified minimum text unit, and determining whether the obtained combination meets the constraint condition corresponding to the regular expression; If it is in compliance, then the factor to be determined is updated according to the combination formula; If not, the factor to be determined is directed to a new specified minimum text unit.

6. The large model decoding constraint method according to claim 5, characterized in that: Determining whether the obtained combination meets the constraint conditions corresponding to the regular expression specifically includes: Determine the current length corresponding to the combination; If the current length exceeds the preset length threshold, a fallback strategy is triggered. In the combination, the latest obtained number of specified minimum text units are retained, the remaining specified minimum text units are deleted, and the number of deletions is recorded. Based on the regular expression and the number of deletions, a corresponding starting judgment position is selected; According to the starting judgment position, it is judged whether the combination meets the constraint condition corresponding to the regular expression.

7. The large model decoding constraint method according to claim 5, characterized in that: The method further comprises: Whenever the factors to be determined are updated, a cumulative probability score is obtained based on the probability scores corresponding to each specified minimum text unit in the vocabulary in the factors to be determined; If the cumulative probability score is lower than the corresponding preset dynamic score, starting from the last specified minimum text unit in the factor to be determined, it is gradually replaced by other specified minimum text units until the cumulative probability score is higher than the corresponding preset dynamic score.

8. The large model decoding constraint method according to claim 7, characterized in that: The method further comprises: Obtaining a preset dynamic score corresponding to the previous length of the factor to be determined; Based on the standard deviation and mean of the probability score of each word in the vocabulary, estimating the interval in which the probability score lies; Based on the maximum value of the interval and the preset dynamic score corresponding to the previous length, the preset dynamic score corresponding to the current length is obtained.

9. A large model decoding constraint device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the large model decoding constraint method as described in any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are set to: the large model decoding constraint method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Transformation of modular finite state transducers

    CN101517533A

  • Unknown protocol behavior reverse inference method based on optimized stochastic converter model

    CN114172972A