Method and system for improving json format output precision of large language model and readable storage medium
By building a policy set dictionary and applying corresponding control strategies, the problems of incomplete output, high token consumption and long processing time when generating JSON structured results of large language models are solved, and efficient and accurate JSON generation effect is achieved.
Patent Information
- Application Number
- CN202510094974.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
AI Technical Summary
When generating JSON structured results, existing large language models have problems such as incomplete output, high token consumption and long processing time, resulting in low accuracy and efficiency of generating results.
By constructing a strategy set dictionary, the input processing, output score and output probability steps of the large language model are controlled, and strategies such as modifying input text, type control, length control, probability graph modification, subvocabulary control and pause word control are adopted to ensure that the model generates 100% JSON structured results.
It realizes 100% JSON structured result generation, reduces token consumption and processing time, improves the accuracy and efficiency of the generated result, and adapts to more decoding methods.
Smart Images

Figure CN120068932A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of generative large language models, and particularly to a method, system and readable storage medium for improving the output accuracy of large language models in JSON format, which are applicable to probability generation models that use text as output, such as current mainstream large language models and multimodal large language models. Background Art
[0002] A large language model (LLM) refers to a natural language processing (NLP) model based on deep learning technology and trained with large-scale data. It usually has hundreds of millions to hundreds of billions of parameters and can understand and generate natural language text. It involves multiple fields, including machine learning, deep learning, neural networks, natural language processing, and computing resources. The underlying architecture of the large language model is Transformer-decoder. Transformer was proposed by Vaswani et al. in 2017. The self-attention mechanism is used inside the model to replace traditional RNNs and CNNs, aiming to solve the long-sequence dependence problem in sequence data processing. This model was mainly used for sequence transcription in the past, so it has great advantages in multilingual translation tasks. The large language model originally came from GPT (Generative Pre-trained Transformer), an autoregressive model based on the Transformer-Decoder architecture. In GPT, the generation of each word needs to consider the previous sequence, so its inference process has unidirectionality and causality.
[0003] The training methods of current mainstream GPT models are divided into three stages: pre-training, fine-tuning, and alignment. The pre-training of the GPT model is carried out on a large scale of unsupervised data, so that the model can learn the real data distribution in this way; in the fine-tuning stage, by fine-tuning on carefully constructed supervised data, the model can quickly master specific reply formats and conversation methods; finally, in the alignment stage, through technologies such as RLHF, the alignment of the model output corpus and the real corpus is realized, further improving the reply accuracy of the model and the security of the reply, so that the model can truly answer in a safe, harmless and accurate manner. The fine-tuned model can be used for question answering, conversation, code generation, text parsing, etc. Its essence is to generate text by sampling the probability distribution.
[0004] Here, a simple demonstration is given of how the large language model makes predictions to achieve tasks such as question answering. As follows Figure 1As shown in , we first need to prepare a vocabulary T, which is constructed by sampling a large amount of text through PBE and other methods. The model learns a probability distribution by learning the relationship between the sequence numbers after word mapping. In this way, when a word in the vocabulary is input, the model can give the distribution of the next word of the word, and then the next word can be sampled from it, such as Figure 2 shown.
[0005] By getting the next word and adding it to the input, we repeat this process to get the third word, the fourth word, and so on. The model will continue to output, such as Figure 3 This process is called "autoregression", which is why GPT-like models are called autoregressive language models.
[0006] Explanation of the autoregressive process: Input "you", converted to 1 through the vocabulary, the model outputs 2, which is concatenated after 1, and converted to "good" through the vocabulary and printed out. Then the model receives [1,2], outputs 3, which is concatenated after 2, and converted to "ah" through the vocabulary. Repeating this process, the model will continue to output. However, there are some problems with such output, such as the inability to determine when the model needs to stop outputting, and the generation of content is completely dependent on the model and cannot be controlled manually. Summary of the invention
[0007] The present invention proposes a method, system and readable storage medium for improving the output accuracy of a large language model in json format, which belongs to a strategy set control method and aims to control the execution of different strategies within the strategy set by constructing a dictionary to ensure that 100% JSON structured results can be generated on a large language model (LLM). The core of the present invention is to design a flexible and extensible strategy set dictionary. The strategy set dictionary is used to store the strategies required for a specific format, which can be probability modifications before model sampling, additional splicing of output results, etc. In the model generation process, the input at a specific time point is customized and modified through the strategy to adjust the output at the next time point, and the model output at a specific time point is specifically modified, so as to achieve control of the output content and improve the correctness of the results. This method is significantly different from the current mainstream practice in terms of efficiency, time and resource consumption, and provides a more optimized solution.
[0008] According to a first aspect of the technical solution of the present invention, a method for improving the output accuracy of a large language model in JSON format is provided, wherein the large language model includes an input processing step, a model calculation step, an output score step, an output probability step, and a decoding step.
[0009] It is characterized in that the large language model is provided with a policy set, and the policy set includes a policy for modifying the input text, a type control policy, a length control policy, a probability graph modification policy, a sub-word table control policy, and multiple stop word control policies;
[0010] In the input processing step of the large language model, a policy for modifying the input text is set;
[0011] In the output score step of the large language model, a type control policy and a sub-word table control policy are set. Among them, the type control policy includes a first stop word control policy; the sub-word table control policy includes a probability graph modification policy;
[0012] In the output probability step of the large language model, a sub-word table control policy and a length control policy are set. Among them, the sub-word table control policy includes a probability graph modification policy; the length control policy includes a second stop word control policy;
[0013] Among them, the output result of the large language model is in the complete json format.
[0014] Furthermore, the policy set is bound to the large language model in the form of a hook.
[0015] Furthermore, both the first and second stop word control policies include multiple different types of tokens, which are executed when the score is converted into a probability and the next token is sampled. When the next token output is the token in this policy, the model inference ends.
[0016] Furthermore, the type control policy is specifically:
[0017] By controlling the stop symbol as the closing symbol of the local element and adding partial splicing of the next content to be inferred.
[0018] Furthermore, the type control policy is implemented based on a stack-type data structure.
[0019] Furthermore, the stop symbols include the following types:
[0020] Integer (int): If the current key element being processed is not the last item in the given set, then insert "," in the stop word policy; on the contrary, if it is the last item, then insert "}";
[0021] Floating point number (float): According to the position of the current processed element in the overall sequence, change the insertion type of the stop word;
[0022] String (str): Depending on whether the string elements in the dictionary use double quotes ("") or single quotes ('), add " or' to the stop word strategy respectively;
[0023] List (list) and dictionary (dict): For lists, insert "]"; for dictionaries, insert "}".
[0024] Furthermore, the length control strategy is implemented by setting stop words or modifying the probability graph.
[0025] Furthermore, when the length control strategy is combined with the type control strategy, the probability graph is modified according to the stop words of the type control strategy to be paused by the type control strategy. That is
[0026] When it is necessary to pause due to exceeding the length, different token probability graphs need to be set according to the current pause symbol of the type control strategy.
[0027] The example is as follows (assuming the length strategy is 1):
[0028] For example, when inferring to {"A": "B", at this time, the length strategy first obtains the current stop word of the type strategy (closing double quotes), and then uses the probability graph modification method to increase the probability corresponding to "(closing double quotes) to the maximum to trigger the pause of the type strategy.
[0029] Furthermore, the sub-word list control strategy is specifically:
[0030] During the sampling process, a sub-word list is set for a specific type, thereby reducing the sampling range.
[0031] Furthermore, the sub-word list control strategy includes an integer sub-word list control strategy and a Chinese sub-word list control strategy.
[0032] Furthermore, the probability graph modification strategy achieves targeted sampling by modifying the probabilities output by the model.
[0033] Furthermore, the probability graph modification strategy is specifically:
[0034] Suppose the structured information to be extracted includes first type information and second type information. The first type information only needs the words between the first index item groups in the word list, and the second type information only needs the numbers between the second index item groups in the word list;
[0035] The probability graph output by the model is a score of size 1*100, indicating the probability distribution of the model for the next word among these 100 words;
[0036] By marking the probabilities in other index items except the first index item group and the second index item group as "-inf" to modify the output probabilities of the model, the only remaining word list is recorded as the sub-word list.
[0037] According to the second aspect of the technical solution of the present invention, a system for improving the output accuracy of the large language model in JSON format is provided. The system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to perform the method for improving the output accuracy of the large language model in JSON format as described in any of the above aspects.
[0038] According to the third aspect of the technical solution of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method for improving the output accuracy of the large language model in JSON format as described in any of the above aspects is implemented.
[0039] Advantages of the present invention:
[0040] 1. 100% generate a complete structured dictionary: In the past, the decoding method of sampling decoding was generally used in the industry to generate a structured dictionary, and there was a certain probability that a complete structure could not be generated, which led to the inability to parse the Key-Value key-value pairs using interpreters in formats such as JSON. Therefore, it was necessary to repeatedly ask questions. However, using the policy set control program in the present invention can generate a complete dictionary structure 100%. By inputting in advance for different Keys, the hallucination problem of the model for Keys can be avoided. At the same time, by limiting the length of Values and setting the vocabulary, etc., the model will not cause infinite output due to getting stuck in repetition. At the same time, the model finally returns a complete JSON format dictionary, and the Key-Value key-value pairs can surely be obtained through JSON format parsing.
[0041] 2. Lower Token consumption and shorter processing time: Currently, in the inference process of the large language model for structured information output, since the model output may not be parsed into a complete dictionary format, the common practice is to use decoding methods such as increasing the temperature, increasing top_p, and increasing top_n for sampling decoding, and generate a complete structured dictionary through a large number of repeated questions and parsing. In this process, the decoding method is fixed, and there is a large amount of token waste. All the previous unparsable answers need to be discarded. However, using this solution only requires one question and setting, and then after one generation by the model, the generated result can be directly parsed into a JSON dictionary, with 0 token waste, and there is no need to spend too much time repeating questions and generating. On the other hand, in some specific business scenarios, some tokens in the policy dictionary can also be pre-computed for KVCache to further accelerate the inference speed. Therefore, compared with the current common practice, this solution has lower Token consumption and shorter processing time.
[0042] 3. Adapt to more decoding methods: In the past, a complete structured dictionary was mainly generated by writing prompts and other forms. To prevent the model from being unable to generate completely, the common practice in the industry is to output through decoding methods such as sampling decoding with temperature, increasing top_p, and top_n. As a result, problems such as unstable output results and poor quality of the output content will occur. Essentially, this is due to the overly large range of tokens that can be selected by sampling decoding. The sub-word table in the solution of the present invention essentially improves the quality of the reply effect by pruning the sampling range. At the same time, it naturally adapts to the diverse decoding methods of large language models. Using different decoding methods for different scenarios can enable the model to have higher output quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0044] Figure 1 Show a schematic diagram of word representation.
[0045] Figure 2 Show a schematic diagram of the operation of predicting the next word.
[0046] Figure 3 Show a schematic diagram of the autoregressive process.
[0047] Figure 4 Show a schematic diagram of the problems existing in the prior art.
[0048] Figure 5 Show a schematic diagram of the inference process of a large language model.
[0049] Figure 6 Show a schematic diagram of the stop word control strategy in the embodiment of the technical solution of the present invention.
[0050] Figure 7 Show a schematic diagram of word representation in the embodiment of the technical solution of the present invention.
[0051] Figure 8 Show a schematic diagram of setting detailed descriptions for each key to control the generation method of the model in the embodiment of the technical solution of the present invention.
[0052] Figure 9 Show a schematic diagram of the execution process in the prompt format or dialogue in the embodiment of the technical solution of the present invention.
[0053] Figure 10Schematic diagram showing the combination of the embodiment of the technical solution of the present invention with the current mainstream decoding method.
[0054] Figure 11 Schematic diagram showing the model of the program with a policy set and an LLM in the embodiment of the technical solution of the present invention.
[0055] Figure 12 Schematic diagram showing the binding manner of the program with a policy set in the embodiment of the technical solution of the present invention to each function in the LLM.
[0056] Figure 13 Schematic diagram showing the inference on the user side in the embodiment of the technical solution of the present invention.
[0057] Figure 14 Schematic diagram showing the internal inference in the embodiment of the technical solution of the present invention.
[0058] Figure 15 Schematic diagram showing the input processing flow after adding a hook in the embodiment of the technical solution of the present invention.
[0059] Figure 16 Schematic diagram showing the inference of the first Key-Value output in the embodiment of the technical solution of the present invention.
[0060] Figure 17 Schematic diagram showing the inference of the second Key-Value output in the embodiment of the technical solution of the present invention.
[0061] Figure 18 Schematic diagram showing the inference of the third Key-Value output in the embodiment of the technical solution of the present invention.
[0062] Figure 19 Schematic diagram showing the inference of the fourth Key-Value output in the embodiment of the technical solution of the present invention.
[0063] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0064] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0065] The terms "first", "second", etc. in the description and claims of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented, for example, in an order other than those illustrated or described herein.
[0066] In addition, the terms "comprising", "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0067] Multiple, including two or more.
[0068] And / or, it should be understood that for the term "and / or" used in the present disclosure, it is merely an association relationship describing associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, these three situations.
[0069] Nowadays, large language models led by Transformer and GPT have been used in many fields such as natural language understanding, natural language generation, cross-modal learning (vision, speech, etc.) due to their very powerful versatility. As Figure 4 shown, for the parsing of a specific format such as json, there are two problems with the common practice in the prior art of repeatedly asking questions by modifying prompt words and adjusting parameters. This is because when using specific sampling parameters, if the output of the model cannot be parsed as json, it can only be solved by modifying the predefined prompt words or increasing the divergence degree of the model's answer, but these two methods rely on the capabilities of the model itself. Such repeated questioning wastes time and Tokens in multiples.
[0070] Compared with the current common practice of repeatedly asking questions by modifying prompt words and adjusting parameters, the technical solution of the present invention proposes a method, system and readable storage medium for improving the json format output accuracy of large language models. The policy set control method of the present invention realizes a 100% JSON structured result, and at the same time, by optimizing the policy execution process, reducing resource consumption and increasing processing speed, it realizes a significant reduction in token consumption and time. This innovation not only improves the quality and direct usability of the generated results, but also significantly improves the overall production efficiency, providing a more efficient and economical solution for fixed format text generation tasks.
[0071] Specifically, the technical solution of the present invention first provides a method for improving the accuracy of the JSON format output of a large language model. Among them, the large language model includes an input processing step, a model calculation step, an output score step, an output probability step, and a decoding step, where:
[0072] The large language model is provided with a policy set, and the policy set includes a policy for modifying the input text, a type control policy, a length control policy, a probability graph modification policy, a sub-word table control policy, and multiple stop word control policies;
[0073] A policy for modifying the input text is set in the input processing step of the large language model;
[0074] A type control policy and a sub-word table control policy are set in the output score step of the large language model. Among them, the type control policy includes a first stop word control policy; the sub-word table control policy includes a probability graph modification policy;
[0075] A sub-word table control policy and a length control policy are set in the output probability step of the large language model. Among them, the sub-word table control policy includes a probability graph modification policy; the length control policy includes a second stop word control policy;
[0076] Among them, the output result of the large language model is in the complete JSON format.
[0077] In a preferred embodiment, the policy set is bound to the large language model in the form of a hook.
[0078] In a preferred embodiment, the first / second stop word control policy includes multiple different types of tokens and is executed when the score is converted into a probability and the next token is sampled. When the next token output is the token in this policy, the model inference ends.
[0079] In a preferred embodiment, the type control policy is specifically:
[0080] By controlling the stop symbol as the closing symbol of the local element and adding partial splicing of the next content to be inferred.
[0081] In a preferred embodiment, the type control policy is implemented based on a stack-type data structure.
[0082] In a preferred embodiment, the stop symbols include the following types:
[0083] Integer: If the current key element being processed is not the last item in the given set, then insert "," in the stop word policy; on the contrary, if it is the last item, then insert "}";
[0084] Floating point numbers: Change the insertion type of stop words according to the position of the current processing element in the overall sequence;
[0085] Strings: Add " or ' to the stop word strategy according to whether the string elements in the dictionary use double quotes "" or single quotes '' respectively;
[0086] Lists and dictionaries: Insert ] for lists and} for dictionaries.
[0087] In a preferred embodiment, the length control strategy is implemented by setting stop words or modifying the probability graph.
[0088] In a preferred embodiment, when the length control strategy is combined with the type control strategy, the probability graph is modified according to the stop words of the type control strategy to achieve suspension by the type control strategy.
[0089] In a preferred embodiment, the sub-word list control strategy is specifically as follows:
[0090] During the sampling process, set a sub-word list for a specific type to reduce the sampling range.
[0091] In a preferred embodiment, the sub-word list control strategy includes an integer sub-word list control strategy and a Chinese sub-word list control strategy.
[0092] In a preferred embodiment, the probability graph modification strategy achieves targeted sampling by modifying the probabilities output by the model.
[0093] In a preferred embodiment, the probability graph modification strategy is specifically as follows:
[0094] Assume that the structured information to be extracted contains first type information and second type information. The first type information only requires the words between the first index item groups in the word list, and the second type information only requires the numbers between the second index item groups in the word list;
[0095] The probability graph output by the model is a score of size 1*100, representing the probability distribution of the model for the next word among these 100 words;
[0096] Modify the output probabilities of the model by marking the probabilities in other index items except the first index item group and the second index item group as "-inf", so as to record the only remaining word list as the sub-word list.
[0097] The technical solution of the present invention also provides a system for improving the accuracy of JSON format output of large language models. The system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to perform the method for improving the accuracy of JSON format output of large language models as described above.
[0098] The technical solution of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for improving the accuracy of JSON format output of large language models as described above.
[0099] The following elaborates on each technical term in the technical solution of the present invention:
[0100] Policy set
[0101] Definition: Policy set - a collection of multiple policies. By constructing a dictionary to control the execution of different policies within the policy set, it is ensured that 100% JSON-structured results can be generated on large language models (LLMs).
[0102] Currently, some inference processes of large language models can be divided into stages as shown in Figure 5 The policy set can be inserted at the input, output score, and output probability stages. At the input stage, the input text is monitored and modified, and different policies are initialized accordingly; while at the output score and output probability stages, the corresponding policies are executed.
[0103] The following are the policy definitions during the design of the present invention:
[0104] Stop word control policy: A policy for controlling the end of output tokens, which is a commonly used policy in the industry. It is executed when the score is converted to probability and the next token is sampled. When the next token output is the token in this policy, the model inference ends.
[0105] Type control policy: The type policy mainly controls the pause symbol as the closing symbol of local elements and adds partial splicing of the next content to be inferred. For a specific format, such as json, there are always closing symbols for specific elements. For example, the closing symbol of the dictionary format must be the right curly brace, and the closing symbol of the string must be double quotes.
[0106] Length control policy: Setting this policy can control the length range of the next generation. If it is insufficient, continue to generate; if it exceeds, stop immediately.
[0107] Probability graph modification policy: A policy for controlling the output probability. Modify the probability graph at the output probability stage of inference
[0108] Sub-vocabulary list control strategy: A strategy for controlling the output probability. Set up a sub-vocabulary list, modify the strategy through a probability graph, and modify the sampling vocabulary in the model output to the sub-vocabulary list, so that the next token is the content in the sub-vocabulary list. It includes but is not limited to integer sub-vocabulary lists, Chinese sub-vocabulary lists, etc.
[0109] Stop word strategy
[0110] Since the large language model is an autoregressive model and cannot stop generating spontaneously, it is necessary to accurately judge the stopping time for the model. By adding special characters during training, the model can learn to spit out these characters at the end of certain texts, and judge the output of special characters during inference. That is to say, the technical solution of the present invention controls the model to pause by relying on the model to generate a pause symbol, making this seemingly completely unachievable function possible mainly relying on the large language model. After pre-training with a large amount of data, the large language model can obviously control the position information of the pause very well. Taking the fine-tuning of dialogue format data as an example, the technical solution of the present invention adds a special symbol "stop" at the end of each round of dialogue in the supervised data as a sign for the model to stop. After fine-tuning the model, the model will output the stop symbol and end the inference when the reply to a sentence ends, as Figure 6 shown.
[0111] Type control strategy
[0112] In the generation of specific formats, instead of relying on model inference to generate by pre-concatenating the Key content, it is usually necessary to "predict" the closing position of the previous element in advance and terminate the model inference to ensure the efficiency and pertinence of the inference process. To achieve this goal, the technical solution of the present invention adds a one-time pause symbol to the existing stop word strategy, that is, add this stop word at a certain moment in the inference process, and when the pause is triggered during the inference, perform specific operations and delete this stop word to prevent the impact of this stop word on subsequent inferences. The implementation of this strategy relies on a data structure called a stack to ensure that the later added stop words can be evaluated and applied first.
[0113] The configuration methods of different types of pause symbols are as follows:
[0114] 1. Integer (int): If the current key element being processed (such as a dictionary key) is not the last item in the given set, then insert "," in the stop word strategy; on the contrary, if it is the last item, then insert "}". Such a design aims to adapt to the needs of different logical branches through the dynamic change of stop words and achieve a more flexible control process.
[0115] 2. Floating point number (float): Consistent with integer logic, whether it is int or float, the same pause word insertion rules are followed in logical processing. That is, based on the position of the current processing element in the overall sequence, the insertion type of pause words is dynamically changed to meet the requirements of different scenarios.
[0116] 3. String (str): According to whether the string elements in the dictionary are constructed with double quotes ("") or single quotes ('), add " or' in the pause word strategy respectively. This step aims to handle string boundaries to ensure that in subsequent logical judgments, string elements can be accurately identified and processed, avoiding misjudgment or missed judgment.
[0117] 4. List and dictionary: For a list, insert "]"; for a dictionary, insert "}". Such a design principle also aims to meet the specific requirements of different data structures in the logical process, enabling the control logic of the system to more precisely adapt to the processing of different data types, thereby enhancing the flexibility and efficiency of the overall strategy.
[0118] By implementing the above strategies, not only can the accurate control of the output score be achieved, but also the process flow, automation, and efficiency of the reasoning process are ensured, providing stronger support for subsequent data analysis and decision-making.
[0119] Length control strategy
[0120] This strategy can be implemented by setting pause words or modifying the probability graph. When combined with the type strategy, the specific settings of pause words need to be considered. This strategy is mainly used to control the length of the next generated content. For fields that may have repeated generation, this strategy can immediately stop repeated generation.
[0121] Sub-word list control strategy
[0122] Generally, the model samples from the entire dictionary. However, for specific types of generation, there is no need to sample from the entire dictionary. By setting a sub-word list and reducing the sampling range of the model, the next output of the model will surely be the output result of a specific type. For example, for the age in a json dictionary, the technical solution of the present invention only needs to set a digital sub-word list.
[0123] Probability graph modification strategy
[0124] This method mainly aims to achieve targeted sampling by modifying the output probability of the model.
[0125] Suppose there is Figure 7The shown vocabulary list (with a vocabulary size of 100), where the words from "4" to "10" in the vocabulary list are in Chinese, and the words from "11" to "20" are numbers. Assume that the structured information to be extracted includes name and age. The name only requires the words between 4 and 10 in the vocabulary list, while the age only requires the numbers between 11 and 20 in the vocabulary list. The probability graph output by the model is a score of size 1*100, which represents the probability distribution of the model for the next word among these 100 words. The technical solution of the present invention can modify the output probability of the model by marking the probabilities of other index items (1 to 3 and 21 to 100) as "-inf", so as to record only the retained vocabulary list as a sub-vocabulary list.
[0126] Generally, denote the input as
[0127] x = {w 1 , w 2 , w 3 …, w m}, where m = the length of the input characters, the output is y, the model is M, and the vocabulary list is
[0128] V = {v 1 , v 2 , v 3 …, v n}, where n = the vocabulary size, the output vector is s = [s 1 , s 2 , s 3 …, s n , and the output probability is P(y|x). Then the calculation process of the model is:
[0129] s = M(x)
[0130]
[0131] y = argmax([P(y 1 │x), P(y 2 │x), …, P(y n │x)])
[0132] Assume that using the sub-vocabulary list for calculation is only equivalent to selecting the output vector s, and keeping the other calculation parts unchanged:
[0133]
[0134] s = s * V *
[0135]
[0136] y = argmax([P(y 1 │x), P(y2 │x),…,P(y n │x)])
[0137] At this time, the probability output by the model is only for a certain sub-word table and the corresponding content is output, and there is not much modification to the inference process of the model.
[0138] Policy set dictionary
[0139] So how to control the selection of the policy set? The technical solution of the present invention proposes a policy set dictionary. During the inference process of the large language model, by deconstructing this dictionary, the selection of the policy set can be controlled. The specific approach is to design a special format description. Here, taking the extraction of two structured information, namely name and age, as an example, first, a detailed description needs to be set for each key to control the generation method of the model. The example is as Figure 8 shown.
[0140] From Figure 8 It can be seen that the technical solution of the present invention constructs a dictionary with detailed information about name and age.
[0141] Since the name is in Chinese, the type is str, and the length is controlled within 2 to 4 characters. The sub-word table selected for it is 4 to 10;
[0142] Age is a number, with the type of integer int, and the length is between 1 and 4 (considering that some models will add a + sign, resulting in an output of +100, which occupies 4 characters). Therefore, the sub-word table selected should all be integers, so the sub-word table is set to 11 to 20.
[0143] Deconstruct the policy set dictionary to control policy execution
[0144] Through the above steps, the technical solution of the present invention converts a data structure that needs to be structured into a dictionary that deeply describes it. This dictionary is called a policy set dictionary. After that, the technical solution of the present invention will deconstruct the policy set dictionary during the inference process of the large language model and map it to the corresponding policy, so that the large language model can generate a complete and parsable dictionary data according to the format regulations.
[0145] In the policy set dictionary of the above example, the technical solution of the present invention requires the model to reply with the name and age. Therefore, the technical solution of the present invention can construct multiple description information for the name field to correspond to different policies, and then process them during inference. For example, the processing of the type policy is as follows: when structured is required, the type policy needs to be set to "dict" (meaning the data type is a dictionary). This policy will add a one-time pause symbol: a closing brace (}). If the type of the name must be str, then the type policy is set to "str", and this policy will add a one-time pause symbol: a double quote ("). If the type of the age is int, then the type policy is set to "int". This policy will first determine whether the age is the last key in the structured dictionary. If it is, no pause symbol will be added because a closing brace pause symbol has already been added in the initial operation. If not, a comma will be added. The example is as follows:
[0146] First, "{"name":"" can be constructed from the Json and sent as the above text to the model. Since the type of the name is a string, a double quote (") can be used as the control symbol for the pause policy, and the length control is 2 - 3. Similarly, "age":"" is also sent as the above text, and a closing brace (}) is used as the control symbol for the pause policy. This method can be used in both the dialogue format and the prompt format (such as Figure 9 )
[0147] Through the above method, the result inferred by the model at one time is in the complete json format, and there are no problems of repeated questions and token consumption.
[0148] The definition of the policy dictionary is very open, and different dictionary contents should be analyzed and implemented according to the specific constructed data to control different output results.
[0149] Combined with the existing decoding method
[0150] The above-mentioned structured parsing can be extended in the form of a plugin, so it can be better combined with the currently commonly used decoding methods (sampling, greedy decoding, Beam Search, etc.). As Figure 10 shown.
[0151] As Figure 10 can be seen, after using the current structured parsing, various decoding methods will decode on the selected probability subgraph and select the appropriate output from it.
[0152] Embodiment
[0153] The complete processing flow includes initialization and inference. For inference, it can be divided into two aspects: the user side and the internal details.
[0154] Initialization process
[0155] First, initialize the model with the policy set program, as Figure 11 shown. Figure 11 In the figure, it is a model with a policy set program and an LLM. For the user side, it is still a complete model. The policy set program will be bound before and after each function in the LLM in the form of a hook, as Figure 12 shown. During initialization, each control module in the policy set program hooks to the input and output layers of the LLM. The inference process of the LLM can be simply divided into input processing, model calculation, output score, output probability, and decoding layer. Here, different modules of the policy and program will hook to different intermediate layers in advance and perform related operations under specific conditions.
[0156] User-side inference
[0157] Assume that under a specific question and a pre-constructed policy dictionary in a specific format, when using the Model with the policy set program for inference, the final output result is a json result as Figure 13 shown. On the user side, when the user uses the model for inference, it will return a json result in a specific set format.
[0158] Internal inference details
[0159] Taking the example in Figure 14 as an example, if you want the model to return a specific json, first send the input data and the pre-constructed policy set dictionary into the model.
[0160] Prepare the question and the pre-constructed policy dictionary in the json format that needs to be generated. As Figure 15 Before performing the input processing of the LLM, due to the addition of the hook, the input will first execute the input modification control module of the policy and program:
[0161] Then the model infers the first Key-Value ( Figure 16 at the current inference position in the figure). Usually, when the model generates a double-quotation mark closed result, it will stop. In special cases, when the generated result is too long, it will also be stopped by the length policy control. At this time, the inference of the first Key-Value ends, and the result is concatenated to prepare for the inference of the second Key-Value. In the subsequent figures, the decoding layer of the LLM will be ignored, so no operation needs to be performed on it. The solid lines in the figure are the actual data flows, and the dotted lines are the set conditions. Only when the conditions are triggered will it jump to the corresponding control model after the hook for additional processing. The solid lines already assume that the model has output to the trigger condition and will not be further explained later.
[0162] Figure 16Show the inference of the first Key-Value output. The solid lines in the figure are the actual data flows, and the dashed lines are the set conditions. The inference process of the final model only infers the specific time, and the other parts are pieced together by preprocessing or postprocessing. Since the generation of all the specified contents in the policy dictionary has not been completed, the inference needs to continue.
[0163] Next, perform the inference of the second Key-Value ( Figure 17 at the current inference position in). Here, the length is set to 1, so the control of the length policy will be triggered directly after outputting a character, and then "Company" will be concatenated as the output, and then the inference of the third Key-Value will be performed.
[0164] Figure 17 Show the inference of the second Key-Value output. The solid lines in the figure are the actual data flows, and the dashed lines are the set conditions. The type of the highlighted part of this policy dictionary is format_str. Compared with str, it only concatenates the inferred str result into "Content" using Python's f-string, and there is no difference from type being str in other aspects. Since the generation of all the specified contents in the policy dictionary has not been completed, the inference needs to continue.
[0165] In the inference of the third Key-Value ( Figure 18 at the current inference position in), the model will freely answer, only restricted by the type policy.
[0166] Figure 18 Show the inference of the third Key-Value output. There is no special specification here, so the model will freely output until the model outputs an unescaped double quote and stops. Since the generation of all the specified contents in the policy dictionary has not been completed, the inference needs to continue.
[0167] In the last Key-Value inference ( Figure 19 at the current inference position in), the model will generate a number and complete the entire inference process.
[0168] Therefore, at present, the technical solution of the present invention can achieve the generation of the structured dictionary with a probability of 100%, solving the problems of high token consumption, long time consumption, and unsatisfactory generation effect when the existing large language models answer structured information.
[0169] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.
[0170] The serial numbers of the embodiments of the present invention described above are for description only and do not represent the superiority or inferiority of the embodiments.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the above implementation methods can be realized by means of software plus a necessary general hardware platform. Of course, they can also be realized by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0172] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A method for improving the accuracy of JSON format output of a large language model, wherein: The large language model includes an input processing step, a model calculation step, a score output step, a probability output step, and a decoding step, and is characterized in that: The large language model is provided with a strategy set, wherein the strategy set includes a strategy for modifying input text, a type control strategy, a length control strategy, a probability graph modification strategy, a subword table control strategy, and a plurality of pause word control strategies; The input processing step of the large language model is provided with a strategy for modifying the input text; The output score step of the large language model is provided with a type control strategy and a sub-word table control strategy, wherein the type control strategy includes a first pause word control strategy; the sub-word table control strategy includes a probability map modification strategy; The output probability step of the large language model is provided with a subword table control strategy and a length control strategy, wherein the subword table control strategy includes a probability graph modification strategy; the length control strategy includes a second pause word control strategy; The output result of the large language model is in a complete JSON format.
2. The method according to claim 1, characterized in that The policy set is bound to the large language model in the form of a hook.
3. The method according to claim 1, characterized in that: The first and second pause word control strategies both include a plurality of different types of tokens, which are executed when the score is converted into a probability and the next token is sampled. When the next token output is a token in the strategy, the model reasoning ends.
4. The method according to claim 1, characterized in that The type control strategy is specifically: By controlling the pause symbol as the closing symbol of the local element, a partial splicing of the next content to be inferred is added.
5. The method according to claim 4, characterized in that The type control strategy is implemented based on a stack-type data structure.
6. The method according to claim 4, characterized in that The pause symbols include the following types: Integer: If the key element currently being processed is not the last item in a given set, insert "," in the pause word strategy; otherwise, if it is the last item, insert "}"; Floating point number: Change the insertion type of pause words according to the position of the currently processed element in the overall sequence; String: Depending on whether the string element in the dictionary uses double quotes """" or single quotes """, add """ or "'" to the pause word strategy; Lists and dictionaries: For lists, insert "]"; for dictionaries, insert "}".
7. The method according to claim 1, characterized in that The length control strategy is implemented by setting pause words or modifying the probability map.
8. The method according to claim 7, characterized in that When the length control strategy is combined with the type control strategy, the probability map is modified according to the pause words of the type control strategy to achieve the pause by the type control strategy.
9. The method according to claim 1, characterized in that: The sub-vocabulary control strategy is specifically as follows: During the sampling process, a sub-word list is set for a specific type to reduce the sampling scope.
10. The method according to claim 9, characterized in that The sub-word table control strategy includes an integer sub-word table control strategy and a Chinese sub-word table control strategy.
11. The method according to claim 1, characterized in that: The probability map modification strategy achieves targeted sampling by modifying the model output probability.
12. The method according to claim 11, characterized in that The probability map modification strategy is specifically as follows: Assume that the structured information to be extracted includes first type information and second type information, the first type information only requires words between the first index item group in the vocabulary, and the second type information only requires numbers between the second index item group in the vocabulary; The probability graph output by the model is a 1*100 score, which represents the probability distribution of the model for the next word among these 100 words. The output probability of the model is modified by marking the probabilities in the index items other than the first index item group and the second index item group as "-inf", so that only the retained vocabulary is recorded as a sub-vocabulary.
13. A system for improving the accuracy of JSON format output of a large language model, the system comprising: A processor and a memory for storing executable instructions; characterized in that the processor is configured to execute the executable instructions to execute the method for improving the output accuracy of a large language model in json format according to any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method for improving the output accuracy of a large language model in JSON format according to any one of claims 1 to 12 is implemented.
Citation Information
Cited By
Generative large model-oriented dynamic adaptive interaction system and method
CN121541952A