A method and system for extreme optimization of the output unit of a large model agent
Through multi-level compression and optimization of the output format of the big model, the problem of excessive tokens in the low-latency system of the big model is solved, and the running speed and user experience are improved.
Patent Information
- Application Number
- CN202411301633.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-18
AI Technical Summary
In the low-latency response system, the number of tokens caused by the redundant information of JSON format and the direct output of the original function name in the existing large model has increased response time and reduced system performance.
Through multi-level compression and optimization, including compression algorithms, vocabulary layer, structural layer, symbol layer and model layer, combining coding and training, reduce the number of tokens, optimize the output format, and use custom tokenization tools and fine-tuning models to generate concise encoding.
Significantly reduce the number of tokens, improve the running speed and overall efficiency of large models, improve user experience, and is especially suitable for low-latency systems.
Smart Images

Figure CN119248280B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model agents; specifically, it relates to a method and system for extreme optimization of the output unit of large model agents. Background Art
[0002] With the rapid evolution of artificial intelligence technology, large model applications in the fields of natural language processing, machine translation, and intelligent question answering have received extensive attention and use.
[0003] However, in these large model applications, the large model tool invocation method faces significant efficiency challenges, especially in cases where low-latency responses are required, and this problem is particularly obvious.
[0004] The progress of natural language processing (NLP) technology enables large models to perform various language tasks with higher accuracy and speed. Through large-scale pre-training and fine-tuning, these models demonstrate powerful language understanding and generation capabilities.
[0005] However, in the traditional method for the agent to use tools, the large model outputs a tool list in a specific form (jsonified tool generation), waits for the tool to be generated, and then processes these tools, which greatly increases the response latency. When applied to a system requiring low latency, this method directly leads to a significant reduction in the user experience.
[0006] In the prior art, large models usually output the tool list in JSON format ({"name": "device_control", "parameters": {"device": "air conditioner", "command": "turn off"}}). Although the JSON format has the advantages of clear structure and easy parsing and performs well in simple tasks, it will cause parsing delay and a decrease in processing efficiency in complex tasks. The redundant information in the JSON format and the direct output of the original function name further increase the number of Tokens, thus increasing the response time. The redundant Tokens brought by using the JSON format and the direct output of the original function name significantly increase the parsing and processing time and reduce the system performance.
[0007] Through in-depth analysis of the technical problems of the prior art, it is found that the performance bottlenecks of current large models at high Token output rates are mainly concentrated in two aspects: one is the processing of redundant information in the JSON format, and the other is the excessive number of Tokens caused by the direct output of the original function name.
[0008] Therefore, how to reduce the number of Tokens and optimize the output format to improve the system operation efficiency and user experience has become a difficult problem to be solved urgently. Summary of the Invention
[0009] In view of this, the purpose of the present invention is to design a new large model tool call method and system. By improving the output format of large model tool calls, the number of Token outputs is reduced. By encoding the original function name and removing redundant JSON formats, the number of Tokens is effectively reduced, the output Tokens are streamlined, the output format is directly processed, the tool information output format is optimized, and the output time of each tool is reduced to achieve the purpose of fast response, solving the problem of low efficiency caused by excessive Token outputs in the current large model tool call system, and being suitable for applications in systems requiring low latency; by reducing the number of output units and optimizing the output format, the running speed and overall running efficiency of the large model are significantly improved, and a better user experience can also be brought.
[0010] The present invention provides a method for extreme optimization of the output unit of a large model intelligent agent, including the following steps:
[0011] S1. Use a compression algorithm to compress the content generated by the large model and directly generate the compressed encoding.
[0012] S2. Optimize the number of tokens output by the large model and perform customized encoding on the tokens, including: compression of tool names and parameter names, compression of JSON structures, and compression of special symbol conjunctions.
[0013] The present invention defines specific tokens within a specific domain through a custom tokenization tool (such as Byte-Pair Encoding, BPE method), making the word segmentation and generation processes more efficient.
[0014] S3. Focus on data in a specific domain, fine-tune the large model (such as GPT), train the large model to recognize the data in this specific domain, and generate a more concise encoding method to adapt to the compression requirements of steps S1 and S2. Improve the processing efficiency.
[0015] The present invention further reduces the number of tokens and response time through multi-level compression and optimization, combined with encoding and training.
[0016] Specifically, the multi-level compression includes the following levels: compression algorithm layer, vocabulary layer, structure layer, symbol layer, model layer.
[0017] The compression algorithm layer applies compression algorithms such as gzip and Brotli to compress the entire content generated by the language model.
[0018] The vocabulary layer is used for customized field name compression to reduce the overhead caused by long field names.
[0019] The structure layer is used to optimize structures such as JSON to reduce format redundancy.
[0020] The symbol layer reduces the number of tokens by compressing and representing special symbols;
[0021] The model layer optimizes the generation method to meet the compression requirements of the above layers by fine-tuning the model.
[0022] Further, the method for fine-tuning the large model in step S3 includes the following steps:
[0023] S31. Collect and prepare a large amount of data in this specific field, including specific abbreviations, coding methods, and sentence structures in this specific field;
[0024] S32. Use the collected data to fine-tune and train the large model so that the large model understands and uses the concise coding method in this specific field.
[0025] Further, the method for compressing the tool name and parameter name in step S2 includes: establishing short codes for the tool name and parameters to reduce the overhead caused by long field names; for example, `getWeather` can be abbreviated as `GW`.
[0026] Example: `getWeather` -> `GW`
[0027] Example: `getDeviceStatus` -> `GDS`
[0028] Further, the method for compressing the JSON structure in step S2 includes: encoding field names and common values to reduce redundant information in the JSON structure and make the JSON structure more concise;
[0029] Example: `{'name': 'device_control', 'parameters': {'device': 'air conditioner', 'command': 'turn off'}}` -> `DC|d=air conditioner,c=turn off`
[0030] Example: `{'status': 'active', 'data': [123, 456, 789]}` -> `A|D[123,456,789]`.
[0031] Further, the compression of special symbol conjunctions in step S2 includes: specifically optimizing special symbols (such as equal signs, delimiters) so that special symbols become a single token in the field, reducing the number of tokens; preferably, the effect can be further enhanced by fine-tuning the training model;
[0032] An example is: `DC|d=air conditioner, c=off` -> `DC| is a token`.
[0033] Furthermore, the compression algorithm in the S1 step includes: gzip compression algorithm, Brotli compression algorithm, or Lempel-Ziv-Welch (LZW) compression algorithm;
[0034] For example, use compression algorithms such as gzip, Lempel-Ziv-Welch (LZW), or Brotli to compress the generated content.
[0035] The gzip compression algorithm is a popular compression algorithm widely used to reduce file size.
[0036] The Brotli compression algorithm is a compression algorithm developed by Google, which has a higher compression ratio and faster decompression speed.
[0037] The present invention also provides a system for extremely optimizing the output unit of a large model agent, which executes the method for extremely optimizing the output unit of a large model agent as described above, including:
[0038] Compression encoding module: used to compress the content generated by the large model using a compression algorithm to directly generate the compressed encoding;
[0039] Vertical domain encoding module: used to optimize the number of tokens output by the large model and perform customized encoding on the tokens, including: compression of tool names and parameter names, compression of JSON structures, and compression of special symbol conjunctions;
[0040] Fine-tuning training module: used to focus on data in a specific domain, fine-tune the large model, and train the large model to be able to recognize the data in that specific domain and generate a more concise encoding method to meet the compression requirements of steps S1 and S2.
[0041] Furthermore, the fine-tuning training module includes:
[0042] Fine-tuning training data preparation unit: used to collect and prepare a large amount of data in that specific domain, including specific abbreviations, encoding methods, and sentence structures in that specific domain;
[0043] Fine-tuning model training unit: used to fine-tune and train the large model using the collected data so that the large model understands and uses the concise encoding method in that specific domain.
[0044] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method for extremely optimizing the output unit of a large model agent as described above.
[0045] The present invention also provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for extremely optimizing the output unit of the large model agent as described above are implemented.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] The method and system for extremely optimizing the output unit of the large model agent provided by the present invention significantly improve the running speed of large model tool calls by reducing the number of Tokens and optimizing string processing; the streamlined Token output reduces the processing time, improves the processing efficiency, and significantly improves the overall efficiency of the system, and is particularly suitable for applications in systems requiring low latency; the optimized string processing method makes tool calls faster and more accurate, enhances the accuracy of tool calls, and reduces various problems caused by excessive parsing time; the improvement of the running efficiency also improves the operation efficiency, and users have a faster response speed and a smoother interaction experience in actual operations, significantly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.
[0049] In the drawings:
[0050] Figure 1 is a schematic flow chart of a user controlling a smart home system device through a natural language command in an application example of the present invention;
[0051] Figure 2 is a flow chart of a method for extremely optimizing the output unit of a large model agent of the present invention;
[0052] Figure 3 is a flow chart of a method for fine-tuning a large model of the present invention;
[0053] Figure 4 is a schematic diagram of the composition of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of systems and products consistent with some aspects of the present disclosure as detailed in the appended claims.
[0055] The terms used in the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0056] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0057] The embodiments of the present invention will be further described in detail below.
[0058] The embodiments of the present invention provide a method for extremely optimizing the output unit of a large model agent. Refer to Figure 2 as shown, including the following steps:
[0059] S1. Use a compression algorithm to compress the content generated by the large model and directly generate the compressed encoding;
[0060] Specifically, use the gzip compression algorithm to compress the generated content.
[0061] S2. Optimize the number of tokens output by the large model and perform customized encoding on the tokens, including: compression of tool names and parameter names, compression of JSON structures, and compression of special symbol conjunctions;
[0062] Specifically, define the tokens unique to the specific domain through the custom tokenization tool BPE method, and the word segmentation and generation process are more efficient;
[0063] The method for compressing tool names and parameter names includes: establishing short codes for tool names and parameters to reduce the overhead caused by long field names;
[0064] Examples in this embodiment are as follows:
[0065] `getWeather` -> `GW`
[0066] `getDeviceStatus` -> `GDS`
[0067] The method for compressing JSON structures includes: encoding field names and common values to reduce redundant information in the JSON structure and make the JSON structure more concise;
[0068] Examples in this embodiment are as follows:
[0069] `{'name': 'device_control', 'parameters': {'device': 'air conditioner', 'command': 'turn off'}}` -> `DC|d=air conditioner,c=turn off`
[0070] `{'status': 'active', 'data': [123, 456, 789]}` -> `A|D[123,456,789]`.
[0071] Compression of special symbol connectors includes: specifically optimizing special symbols (such as equal signs, delimiters) so that special symbols become a single token in the field and reduce the number of tokens; preferably, the effect can be further enhanced by fine-tuning the training model;
[0072] Examples in this embodiment are as follows: `DC|d=air conditioner,c=turn off` -> `DC|is a token`.
[0073] S3. Focus on data in a specific field, fine-tune the large model, and train the large model to be able to recognize the data in this specific field and generate a more concise encoding method to meet the compression requirements of steps S1 and S2.
[0074] The method for fine-tuning the large model includes the following steps (see Figure 3 shown):
[0075] S31. Collect and prepare a large amount of data in this specific field, including specific abbreviations, encoding methods, and sentence structures in this specific field;
[0076] S32. Use the collected data to fine-tune and train the large model so that the large model understands and uses the concise encoding method in this specific field.
[0077] This embodiment further reduces the number of tokens and response time through multi-level compression and optimization, combined with encoding and training.
[0078] Specifically, the multi-level compression includes the following levels: compression algorithm layer, vocabulary layer, structure layer, symbol layer, and model layer; the compression algorithm layer applies compression algorithms such as gzip and Brotli to compress the overall content generated by the language model; the vocabulary layer is used for customizing the compression of field names to reduce the overhead caused by long field names; the structure layer is used to optimize structures such as JSON to reduce format redundancy; the symbol layer utilizes the compression and representation of special symbols to reduce the number of tokens; the model layer optimizes the generation method through fine-tuning the model to meet the compression requirements of the above layers.
[0079] The embodiment of the present invention also provides a system for extremely optimizing the output unit of the large model agent, which executes the method for extremely optimizing the output unit of the large model agent as described above, including:
[0080] Compression encoding module: used to compress the content generated by the large model using a compression algorithm and directly generate the compressed encoding;
[0081] Vertical domain encoding module: used to optimize the number of tokens output by the large model and perform customized encoding on the tokens, including: compression of tool names and parameter names, JSON structure compression, and compression of special symbol conjunctions;
[0082] Fine-tuning training module: used to focus on data in a specific domain, fine-tune the large model, and train the large model to recognize the data in this specific domain and generate a more concise encoding method to meet the compression requirements of steps S1 and S2.
[0083] The fine-tuning training module includes:
[0084] Fine-tuning training data preparation unit: used to collect and prepare a large amount of data in this specific domain, including specific abbreviations, encoding methods, and sentence structures in this specific domain;
[0085] Fine-tuning model training unit: used to fine-tune and train the large model using the collected data, so that the large model understands and uses the concise encoding method in this specific domain.
[0086] An application example of the embodiment of the present invention is (see Figure 1 shown):
[0087] Suppose the user has a smart home system, and the user can control the device through natural language commands:
[0088] Input command: `Raise the temperature in the living room to 25 degrees`
[0089] Standardized command: `{'device': 'thermostat', 'location': 'living_room', 'command':'set_temperature', 'value': 25}`
[0090] Compressed output: `T|lr|st|25`
[0091] By the above method, the number of output characters is significantly reduced, the number of tokens is reduced by two-thirds, the latency is increased by 1.5 times, effectively improving the processing speed and response speed.
[0092] By streamlining the number of output tokens, the number of tokens in the generated text is reduced, which can reduce the number of output tokens, accelerate the response time, reduce the model latency, and reduce the computational cost, improving the user experience.
[0093] An embodiment of the present invention also provides a computer device. Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention; see the accompanying drawings. Figure 4 As shown, the computer device includes: an input system 23, an output system 24, a memory 22, and a processor 21; the memory 22 is used to store one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the method for extreme optimization of the output unit of the large model agent as provided in the above embodiment; where the input system 23, the output system 24, the memory 22, and the processor 21 can be connected through a bus or other means. Figure 4 Taking connection through a bus as an example.
[0094] The memory 22, as a readable and writable storage medium of a computing device, can be used to store software programs and computer-executable programs, such as program instructions corresponding to the method for extreme optimization of the output unit of the large model agent as described in the embodiment of the present invention; the memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc.; in addition, the memory 22 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices; in some instances, the memory 22 can further include a memory remotely set relative to the processor 21, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and their combinations.
[0095] The input system 23 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the device; the output system 24 can include display devices such as a display screen.
[0096] The processor 21 executes various functional applications and data processing of the device by running software programs, instructions, and modules stored in the memory 22, that is, implements the method for extremely optimizing the output unit of the large model intelligent agent as described above.
[0097] The computer device provided above can be used to execute the method for extremely optimizing the output unit of the large model intelligent agent provided in the above embodiments, and has corresponding functions and beneficial effects.
[0098] The embodiment of the present invention further provides a storage medium containing computer-executable instructions. The computer-executable instructions are used to execute the method for extremely optimizing the output unit of the large model intelligent agent provided in the above embodiments when executed by a computer processor. The storage medium is any of various types of memory devices or storage devices, including: installation media such as CD-ROMs, floppy disks, or tape systems; computer system memories or random access memories such as DRAM, DDRRAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memories such as flash memories, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc.; the storage medium can also include other types of memories or combinations thereof; additionally, the storage medium can be located in the first computer system in which the program is executed, or can be located in a different second computer system, and the second computer system is connected to the first computer system through a network (such as the Internet); the second computer system can provide program instructions to the first computer for execution. The storage medium includes two or more storage media that can reside in different locations (such as in different computer systems connected through a network). The storage medium can store program instructions executable by one or more processors (such as specifically implemented as a computer program).
[0099] Of course, the computer-executable instructions of the storage medium containing computer-executable instructions provided by the embodiment of the present invention are not limited to the method for extremely optimizing the output unit of the large model intelligent agent as described in the above embodiments, and can also execute related operations in the method for extremely optimizing the output unit of the large model intelligent agent provided by any embodiment of the present invention.
[0100] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, those skilled in the art can easily understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0101] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for extreme optimization of the output unit of a large model intelligent agent, characterized in that, It includes the following steps: S1. Use a compression algorithm to compress the content generated by the large model and directly generate the compressed encoding; S2. Optimize the number of tokens output by the large model and perform customized encoding on the tokens, including: compression of tool names and parameter names, compression of JSON structures, and compression of special symbol conjunctions; S3. Focus on data in a specific field, fine-tune the large model, and train the large model to be able to recognize the data in that specific field and generate a more concise encoding method to meet the compression requirements of steps S1 and S2; The method for fine-tuning the large model in step S3 includes the following steps: S31. Collect and prepare a large amount of data in that specific field, including specific abbreviations, encoding methods, and sentence structures in that specific field; S32. Use the collected data to perform fine-tuning training on the large model so that the large model understands and uses the concise encoding method in that specific field.
2. The method for extreme optimization of the output unit of the large model intelligent agent according to claim 1, wherein The method for compressing tool names and parameter names in step S2 includes: establishing short codes for tool names and parameters to reduce the overhead caused by long field names.
3. The method for extreme optimization of the output unit of the large model intelligent agent according to claim 1, wherein The method for compressing JSON structures in step S2 includes: encoding field names and common values to reduce redundant information in the JSON structure and make the JSON structure more concise.
4. The method for extreme optimization of the output unit of the large model intelligent agent according to claim 1, wherein, The compression of special symbol conjunctions in step S2 includes: specifically optimizing special symbols so that special symbols become a single token in the field and reducing the number of tokens.
5. The method for extreme optimization of the output unit of the large model intelligent agent according to claim 1, wherein The compression algorithm in step S1 includes: gzip compression algorithm, Brotli compression algorithm, or Lempel-Ziv-Welch compression algorithm.
6. A system for extremely optimizing the output unit of a large model agent, which executes the method for extremely optimizing the output unit of a large model agent according to any one of claims 1-5, characterized in that, It includes: Compression Encoding Module: Used to compress the content generated by the large model using a compression algorithm and directly generate the compressed encoding; Vertical Domain Encoding Module: Used to optimize the number of tokens output by the large model and perform customized encoding on the tokens, including: compression of tool names and parameter names, compression of JSON structures, and compression of special symbol conjunctions; Fine-tuning Training Module: Used to focus on data in a specific field, fine-tune the large model, and train the large model to be able to recognize the data in that specific field and generate a more concise encoding method to meet the compression requirements of steps S1 and S2.
7. The system for extreme optimization of the output unit of the large model intelligent agent according to claim 6, characterized in that, The fine-tuning training module includes: Fine-tuning Training Data Preparation Unit: Used to collect and prepare a large amount of data in that specific field, including specific abbreviations, encoding methods, and sentence structures in that specific field; Fine-tuning Model Training Unit: Used to perform fine-tuning training on the large model using the collected data so that the large model understands and uses the concise encoding method in that specific field.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method for extremely optimizing the output unit of the large model agent as described in any one of claims 1-5.
9. A computer device, the computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for extremely optimizing the output unit of the large model agent as described in any one of claims 1-5.
Citation Information
Patent Citations
Incremental multi-mode expert large model architecture based on integrated model
CN118607665A