Large language model efficient reasoning method and system based on preset Schema structure
By introducing efficient inference methods based on preset Schema structures in large language models, using serial or parallel inference filling technology and KV Cache to accelerate the generation process, the problem of accuracy and low efficiency in generating structured data in large language models is solved, and more stable and efficient structured data generation is achieved.
Patent Information
- Application Number
- CN202411854091.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing large language models have problems such as inaccurate output, low efficiency, and easy eviction of requests when generating structured data, resulting in incomplete generation.
An efficient inference method of large language model based on preset Schema structure is proposed. The preset Schema structure is requested at one time to perform fill inference, and structured data is generated using serial or parallel inference filling technology, and the generation process is accelerated through KV Cache.
It improves the accuracy and efficiency of the large language model to output structured data, enhances the robustness and response speed of the system, and is suitable for multi-instance load deployment scenarios, ensuring the stability and completeness of the generated structured data.
Smart Images

Figure CN120012908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and in particular to an efficient reasoning method and system for a large language model based on a preset Schema structure. Background Art
[0002] In recent years, large language model technology has developed rapidly. Direct input and output through large language models can no longer meet some of the needs of modern technology. Technical personnel often need to embed large language models into workflows and let large language models interact with other work components to implement complex systems. Large language models need to efficiently output stable and reliable structured data.
[0003] The current mainstream approach is to design prompt words through prompt word engineering to initiate requests to the large language model to obtain the expected structured data. Through prompt word engineering, you only need to initiate a request to the large language model once to obtain the final structured data. Although the accuracy of the output of the large language model can be improved by designing prompt words, there is still the possibility of non-standard output format. Once unexpected errors occur in the actual workflow, the entire system may be abnormal; and there is output deviation directly through prompt word engineering, which is easy to cause parsing abnormalities.
[0004] Another approach is to preset the output Schema structure, and make multiple requests to infer the field values corresponding to the Schema structure to obtain the final structured data. By making multiple requests to the large language model to generate the final content, it can be ensured that the generated content is standardized structured data. However, multiple requests involve a lot of repetitive and redundant calculations, and the overall efficiency is low. Each request requires establishing a connection with the server and transmitting the request content, which results in a lot of network overhead.
[0005] Moreover, in order to reduce repeated calculations during the inference process, KV Cache is usually introduced to speed up the generation speed. However, since the requests are processed in segments, the requests cannot effectively utilize KV Cache.
[0006] Furthermore, although Prefix Caching can be enabled to cache the common parts of different requests during the process of inferring the schema structure in multiple requests, this method will occupy additional video memory. Even for block-level reuse, a small part needs to be recalculated. Prefix Caching often fails in multi-instance load deployment scenarios.
[0007] In addition, when the concurrency is large, requests are easily evicted, resulting in incomplete generation of the preset Schema structure.
[0008] In order to solve the above problem of generating structured data with a large language model, the present application proposes an efficient reasoning method and system for a large language model based on a preset Schema structure. Summary of the invention
[0009] The present invention proposes the following technical solutions to address one or more technical deficiencies in the above-mentioned prior art.
[0010] Based on the first aspect of the present application, an efficient reasoning method for a large language model based on a preset Schema structure is proposed, comprising:
[0011] S1: Prompt and the preset Schema structure are sent to the large language model server, and the large language model in the large language model server parses the reasoning part and the stop condition of the Schema structure;
[0012] S2: selectively set the association between the fields of the Schema structure;
[0013] S3: If there is a set association between the fields of the Schema structure, the large language model performs serial reasoning and filling, splices the Prompt prompt with a field of the preset Schema structure and uses it as the input of the large language model, generates field values corresponding to the fields of the Schema structure token by token, and fills the field values of the reasoning part, performs a stop generation check on the fields of the Schema structure according to the stop condition, and if there are still fields of the Schema structure that have not generated field values, continues to splice the Prompt prompt with the next field of the Schema structure and repeats this step;
[0014] S4: Output complete structured data when all fields of the Schema structure have generated field values.
[0015] This application can guide the model to continue generating in the desired direction, and the generated content has a higher accuracy.
[0016] Furthermore, if there is no set association between the fields of the Schema structure in S3, the large language model performs parallel reasoning and filling, uniformly encodes the common parts of the prompt prompts in parallel, and splices each of the Prompt prompts with a field of the Schema structure to perform parallel field value reasoning. Each token word simultaneously generates a field value corresponding to the field of the Schema structure, and assembles the inferred field value structure into complete structured data.
[0017] Furthermore, it is verified whether each of the token words contains the stop condition or the stop symbol. If the token word contains the stop condition or the stop symbol, the field generation of the Schema structure is completed.
[0018] Furthermore, if the length of the generated field value exceeds a preset length and the stop condition or the stop symbol does not appear, the field value corresponding to the field of the Schema structure is regenerated.
[0019] Furthermore, if the number of times the field value corresponding to the field of the Schema structure is regenerated reaches a preset maximum number of regeneration times, the currently generated field value content is used as the field value corresponding to the field of the Schema structure.
[0020] Furthermore, the Prompt prompt is concatenated with a field of the preset Schema structure through parallel encoding.
[0021] Furthermore, the field value content corresponding to the generated field of the Schema structure, as well as the key value and value value of the prompt for concatenating the content of the field of the Schema structure are added to the KV cache.
[0022] The Prompt word can help the model better understand the input intent and make corresponding responses when training a supervised learning or unsupervised learning model.
[0023] The fields of the preset Schema structure are used to describe the structure and relationship of the data model, which can standardize and automate the data structure, ensure correctness and consistency, and improve development efficiency.
[0024] Furthermore, preferably, the Schema structure is:
[0025]
[0026] Among them, the "[Model Generation]" part is the part that needs to be filled by reasoning for the large language model. Serial reasoning filling or parallel reasoning filling is performed according to the association between the fields of the Schema structure. Each "Model Generation" part has a "[stop_str]" stop mark.
[0027] Based on the second aspect of the present application, an efficient reasoning system for a large language model based on a preset Schema structure is also proposed, including:
[0028] Parsing module: Send the Prompt prompt and the preset Schema structure to the large language model server, and the large language model in the large language model server parses the reasoning part and the stop condition of the Schema structure;
[0029] Setting module: selectively setting the association between the fields of the Schema structure;
[0030] Reasoning module: if there is a set association between the fields of the Schema structure, the large language model performs serial reasoning and filling, splices the Prompt prompt with a field of the preset Schema structure and uses it as the input of the large language model, generates field values corresponding to the fields of the Schema structure token by token, and fills the field values of the reasoning part, performs a stop generation check on the fields of the Schema structure according to the stop condition, and if there are still fields of the Schema structure that have not generated field values, continues to splice the Prompt prompt with the next field of the Schema structure and repeats this step;
[0031] Output module: Output complete structured data when all fields of the Schema structure generate field values.
[0032] Based on the third aspect of the present application, a computer program product is also proposed, which has one or more computer programs thereon, and when the computer program is executed by a computer processor, implements any of the methods described above.
[0033] The technical effect of the present invention is that: the present application proposes an efficient reasoning method and system for a large language model with a preset Schema structure, which uses a fill-in reasoning technology that outputs a preset Schema structure with a single request, solves the problem of inaccurate and low efficiency in the output of structured data by a large language model, improves the robustness and response speed of the large language model output, and can be applied to all scenarios where a large language model is expected to generate stable structured data, especially to scenarios where a large language model is embedded in a workflow to realize business functions, and the large language model can interact stably with other components in the workflow. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Other features, objects and advantages of the present application will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings.
[0035] Figure 1 It is a flowchart of an efficient reasoning method for a large language model based on a preset Schema structure provided according to an embodiment of the present application.
[0036] Figure 2It is an architectural diagram of an efficient reasoning method for a large language model based on a preset Schema structure provided according to an embodiment of the present application.
[0037] Figure 3 It is a framework diagram of an efficient reasoning system for a large language model based on a preset Schema structure provided according to an embodiment of the present application.
[0038] Figure 4 It is a schematic diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION
[0039] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0040] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0041] Figure 1 The present invention shows an efficient reasoning method for a large language model based on a preset Schema structure, including:
[0042] S1: Send the Prompt prompt and the preset Schema structure to the large language model server, and the large language model in the large language model server parses the reasoning part and the stop condition of the Schema structure;
[0043] S2: selectively set the association between the fields of the Schema structure;
[0044] S3: If there is a set association between the fields of the Schema structure, the large language model performs serial reasoning and filling, splices the Prompt prompt with a field of the preset Schema structure and uses it as the input of the large language model, generates field values corresponding to the fields of the Schema structure token by token, and fills the field values of the reasoning part, performs a stop generation check on the fields of the Schema structure according to the stop condition, and if there are still fields of the Schema structure that have not generated field values, continues to splice the Prompt prompt with the next field of the Schema structure and repeats this step;
[0045] S4: Output complete structured data when all fields of the Schema structure have generated field values.
[0046] It should be noted that if there is no set association between the fields of the Schema structure in S3, the large language model performs parallel reasoning and filling, uniformly encodes the common parts of the prompt prompts in parallel, and splices each of the Prompt prompts with a field of the Schema structure to perform parallel field value reasoning. Each token word simultaneously generates a field value corresponding to the field of the Schema structure, and assembles the inferred field value structure into complete structured data.
[0047] It should be noted that it is verified whether each of the token words contains the stop condition or the stop symbol. If the token word contains the stop condition or the stop symbol, the field generation of the Schema structure is completed.
[0048] It should be noted that if the length of the generated field value exceeds the preset length and the stop condition or the stop symbol does not appear, the field value corresponding to the field of the Schema structure is regenerated.
[0049] It should be noted that if the number of times the field value corresponding to the field of the Schema structure is regenerated reaches a preset maximum number of regeneration times, the currently generated field value content is used as the field value corresponding to the field of the Schema structure.
[0050] It should be noted that the Prompt prompt is concatenated with a field of the preset Schema structure through parallel coding.
[0051] It should be noted that the field value content corresponding to the generated field of the Schema structure, as well as the key value and value value of the prompt for splicing the content of the field of the Schema structure are added to the KV cache.
[0052] In a specific embodiment, the Prompt prompt and the preset Schema structure are sent to the large language model server at one time, and the request carries the parameters of the preset Schema structure and performs fill-in reasoning. Compared with other conventional practices of multiple requests for segmented reasoning, the present application can complete the structured data output with only one request, which can improve the request efficiency and response speed of the large language model.
[0053] In a specific embodiment, the Schema structure of the preset output is as follows:
[0054]
[0055] The "[Model Generation]" part is the part that needs to be filled in for reasoning in the large language model. Each "Model Generation" part has a "[stop_str]", which indicates the stop mark for the generation of this part.
[0056] It should be noted that the large language model reasoning filling method includes serial reasoning filling and parallel reasoning filling. The user can set the association between the fields of the Schema structure as needed. Serial reasoning filling is used when there is an association between the fields of the Schema structure, and parallel reasoning filling is used when there is no association between the fields of the Schema structure.
[0057] It should be noted that the Prompt word can help the model better understand the input intent and make corresponding responses when training a supervised learning or unsupervised learning model.
[0058] It should be noted that the fields of the preset Schema structure are used to describe the structure and relationship of the data model, which can achieve standardization and automation of the data structure, ensure correctness and consistency, and improve development efficiency.
[0059] Table 1
[0060]
[0061]
[0062] In a specific embodiment, as shown in Table 1, the Prompt prompt and a field of the preset Schema structure, such as "Prompt+{"time"," are used as inputs of the large language model to generate a field value "2024-10-16" corresponding to the Schema structure in the token word;
[0063] When performing the second inference, the Prompt prompt of the first inference is concatenated with a new field "location" of the Schema structure, and "Prompt+{"time":"2024-10-16","location":" is used as the new input of the large language model to generate the field value "12 Guanri Road, Xiamen City" corresponding to the Schema result in the token word;
[0064] When performing the third inference, the Prompt prompt of the second inference is concatenated with a new field "name" of the preset Schema structure, and "Prompt+{"Time":"2024-10-16","Location":"12 Guanri Road, Xiamen","Person":{"Name":" is used as the new input of the large language model to generate the field value "Zhang San" corresponding to the Schema result in the token word;
[0065] When performing the fourth inference, the Prompt prompt of the third inference is concatenated with a new field "gender" of the preset Schema structure, and "Prompt+{"Time":"2024-10-16","Location":"12 Guanri Road, Xiamen","Person":{"Name":"Zhang San","Gender":" is used as the new input of the large language model to generate the field value "Male" corresponding to the Schema result in the token word;
[0066] Finally, each field in the Schema structure generates a corresponding field value, and the output of the serial reasoning filling result is:
[0067]
[0068] It should be noted that, compared with directly generating the fields of the Schema structure from a large language model, the present application can guide the model to continue generating in the desired direction, and the generated content has higher accuracy.
[0069] It should be noted that the fields of the preset Schema structure are spliced, and no model generation is required, which speeds up the output of results. The part of the fields of the preset Schema structure in the Prompt splicing can improve the encoding efficiency by adopting parallel encoding calculation. The Key and Value values calculated for the part of the fields of the preset Schema structure in the Prompt splicing can be cached in the KV cache to accelerate the generation of subsequent content.
[0070] Table 2
[0071]
[0072] In a specific embodiment, as shown in Table 2, the Prompt prompt is uniformly encoded as a public part, and is spliced with the fields of the preset Schema structure respectively to obtain "Prompt+{"time"", "Prompt+{"place"", "Prompt+{"person":{"name":", "Prompt+{"person":{"gender":", as a one-time input of the large language model, the large language model performs parallel reasoning and filling, and obtains the corresponding values of the fields of the preset Schema structure "2024-10-16", "12 Guanri Road, Xiamen City", "Zhang San" and "Male" respectively;
[0073] Finally, each field in the Schema structure generates a corresponding field value, and the output of the serial reasoning filling result is:
[0074]
[0075] It should be noted that the overall architecture of an efficient reasoning method for a large language model based on a preset Schema structure is as follows Figure 2 As shown, the Prompt prompt and the preset Schema structure are sent to the large language model server, and the large language model of the large language model server performs the Schema reasoning part and the stop condition analysis. The user selectively sets the association between the fields of the Schema structure as needed, including serial relationship and parallel relationship;
[0076] If it is a serial relationship, the Prompt prompt is concatenated with a field of the preset Schema structure and used as the input of the large language model, field values corresponding to the fields of the Schema structure are generated token by token, and the field values are filled in the inference part, and the fields of the Schema structure are checked for stop generation according to the stop condition until all the fields of the preset Schema structure generate corresponding field values, and the complete structured data is output;
[0077] If it is a parallel relationship, the large language model performs parallel reasoning and filling, uniformly encodes the common parts of the prompts in parallel, splices each of the prompts with a field of the preset Schema structure, and each token word generates a field value corresponding to the field of the Schema structure at the same time, and assembles the inferred field value structure into a complete structured data;
[0078] The field encoding of the concatenated Schema structure and the generated field values are saved in the KV cache.
[0079] It should be noted that KV cache is a caching mechanism for storing key-value pairs. During the reasoning process of a large language model, the model needs to access the same data multiple times. KV cache avoids repeated calculations by caching this data in memory, thereby significantly improving the reasoning speed and reducing resource consumption. Not only that, it can also maintain the performance of the model while reducing the amount of calculation.
[0080] It should be noted that token is the smallest unit processed by the model. After splitting the original text into tokens, the model can calculate the relationship between each token and understand the meaning of the sentence. In the generation task, the model gradually predicts the next token. Each time a new token is generated, the model will re-evaluate and output the most likely next token until the complete text is generated.
[0081] Figure 3A framework diagram of an efficient reasoning system for a large language model based on a preset Schema structure of the present application is shown, including a parsing module a, a setting module b, a reasoning module c and an output module d.
[0082] In a specific embodiment, the parsing module a is configured to: send a Prompt prompt and a preset Schema structure to a large model server, and the large language model in the large model server parses the reasoning part and the stop condition of the Schema structure.
[0083] In a specific embodiment, the setting module b is configured to: selectively set the association between the fields of the Schema structure.
[0084] In a specific embodiment, the reasoning module c is configured as follows: if there is a set association between the fields of the Schema structure, the large language model performs serial reasoning and filling, the Prompt prompt is spliced with a field of the preset Schema structure and used as the input of the large language model, the field value corresponding to the field of the Schema structure is generated token by token and the field value is filled in the reasoning part, the field of the Schema structure is checked for stop generation according to the stop condition, if there are still fields of the Schema structure that have not generated field values, the Prompt prompt is continued to be spliced with the next field of the Schema structure and this step is repeated.
[0085] In a specific embodiment, the output module d is configured to output complete structured data when all fields of the Schema structure generate field values.
[0086] Reference below Figure 4 , which shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. Figure 4 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0087] like Figure 4 As shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage part 408 into a random access memory (RAM) 403. Various programs and data required for system operation are also stored in the RAM 403. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0088] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage section 408 as needed.
[0089] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.
[0090] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0091] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0092] The modules involved in the embodiments of the present application may be implemented by software or hardware.
[0093] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist alone and not be assembled into the electronic device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device: sends the Prompt prompt and the preset Schema structure to the large language model server, and the large language model of the large language model server parses the reasoning part and the stop condition of the Schema structure; selectively sets the association between the fields of the Schema structure; if there is a set association between the fields of the Schema structure, the large language model performs serial reasoning filling, splices the Prompt prompt with a field of the preset Schema structure and uses it as the input of the large language model, generates the field value corresponding to the field of the Schema structure token by token word, and fills the field value of the reasoning part, performs stop generation verification on the field of the Schema structure according to the stop condition, and if there are still fields of the preset Schema structure that have not generated field values, then continue to splice the Prompt prompt with the next field of the preset Schema structure and repeat this step; when all the fields of the preset Schema structure generate field values, output complete structured data.
[0094] Finally, it should be noted that the above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. An efficient reasoning method for a large language model based on a preset Schema structure, characterized in that: include: S1: Send the Prompt prompt and the preset Schema structure to the large language model server, and the large language model in the large language model server parses the reasoning part and the stop condition of the Schema structure; S2: selectively set the association between the fields of the Schema structure; S3: If there is a set association between the fields of the Schema structure, the large language model performs serial reasoning and filling, splices the Prompt prompt with a field of the preset Schema structure and uses it as the input of the large language model, generates field values corresponding to the fields of the Schema structure token by token, and fills the field values of the reasoning part, performs a stop generation check on the fields of the Schema structure according to the stop condition, and if there are still fields of the Schema structure that have not generated field values, continues to splice the Prompt prompt with the next field of the Schema structure and repeats this step; S4: Output complete structured data when all fields of the Schema structure have generated field values.
2. The method according to claim 1, characterized in that If there is no set association between the fields of the Schema structure in S3, the large language model performs parallel reasoning and filling, uniformly encodes the common parts of the prompt prompts in parallel, and concatenates each of the Prompt prompts with a field of the Schema structure to perform parallel field value reasoning. Each token word simultaneously generates a field value corresponding to the field of the Schema structure, and assembles the inferred field value structure into complete structured data.
3. The method according to claim 1, characterized in that Verify whether each of the token words contains the stop condition or the stop symbol. If the token word contains the stop condition or the stop symbol, the field generation of the Schema structure is completed.
4. The method according to claim 1, characterized in that If the length of the generated field value exceeds the preset length and the stop condition or the stop symbol does not appear, the field value corresponding to the field of the Schema structure is regenerated.
5. The method according to claim 4, characterized in that If the number of times the field value corresponding to the field of the Schema structure is regenerated reaches a preset maximum number of regeneration times, the currently generated field value content is used as the field value corresponding to the field of the Schema structure.
6. The method according to claim 1, characterized in that The Prompt prompt is concatenated with a field of the preset Schema structure through parallel encoding.
7. The method according to claim 1, characterized in that The field value content corresponding to the generated field of the Schema structure, as well as the key value and value value of the prompt for concatenating the content of the field of the Schema structure are added to the KV cache.
8. The method according to claim 1, characterized in that Preferably, the Schema structure is: "{"time":"[model generation][stop_str]", "Location":"[Model Generation][stop_str]", "Character":{"Name":"[Model Generation][stop_str]", "gender":"[model generated][stop_str]",},}"; Among them, the "[Model Generation]" part is the part that needs to be filled in by reasoning for the large language model. Serial reasoning filling or parallel reasoning filling is performed according to the association between the fields of the Schema structure. Each "Model Generation" part has a "[stop_str]" stop mark.
9. An efficient reasoning system for a large language model based on a preset Schema structure, characterized in that: include: Parsing module: Send the Prompt prompt and the preset Schema structure to the large language model server, and the large language model in the large language model server parses the reasoning part and the stop condition of the Schema structure; Setting module: selectively setting the association between the fields of the Schema structure; Reasoning module: if there is a set association between the fields of the Schema structure, the large language model performs serial reasoning and filling, splices the Prompt prompt with a field of the preset Schema structure and uses it as the input of the large language model, generates field values corresponding to the fields of the Schema structure token by token, and fills the field values of the reasoning part, performs a stop generation check on the fields of the Schema structure according to the stop condition, and if there are still fields of the Schema structure that have not generated field values, continues to splice the Prompt prompt with the next field of the Schema structure and repeats this step; Output module: Output complete structured data when all fields of the Schema structure generate field values.
10. A computer program product having one or more computer programs thereon, characterized in that: When the computer program is executed by a computer processor, the method according to any one of claims 1 to 8 is implemented.