Large language model dynamic hierarchical prompt compression method, system and device and storage medium

Through a dynamic hierarchical prompt compression method, Markov decision process and reinforcement learning are used to optimize the prompt word compression of large language models, which solves the problem of high computational cost caused by lengthy prompt words and improves the inference efficiency and output performance of the model on resource-constrained devices.

CN120633874APending Publication Date: 2025-09-12SOUTH CHINA UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510795084.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-14
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The prompt words of existing large language models are long and complex, resulting in high computational costs, especially affecting the efficiency of model inference on resource-constrained devices. Black-box compression methods are difficult to effectively compress while maintaining semantic integrity and clarity of task instructions.

Method used

A dynamic hierarchical prompt compression method is adopted. The prompt word compression is modeled through Markov decision process. The compression agent is trained by combining reinforcement learning and curriculum learning. The reward function is designed to optimize the compression strategy and dynamically adjust the prompt length.

Benefits of technology

Without affecting the quality of model output, it effectively compresses the prompt word length, improves inference efficiency, achieves adaptive adjustment and generalization, and is suitable for resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633874A_ABST
    Figure CN120633874A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model dynamic grading prompt compression method, system and device and a storage medium, and belongs to the technical fields of deep learning, reinforcement learning, large language models and the like. The method comprises the following steps: constructing a Markov decision process for prompting compression; the training language model is aligned with output distribution of the target large model; comprehensively designing a compression ratio, and outputting a reward function for alignment and information retention; training a compression agent according to a reinforcement learning algorithm and curriculum learning optimized by a near-end strategy; and dynamically compressing the input prompt by using the compression agent. The invention discloses a dynamic hierarchical prompt compression method based on reinforcement learning, and aims to solve the problems that in the current prompt compression technology, the compression ratio and key information retention are difficult to balance, the method generalization is insufficient, a self-adaptive adjustment mechanism is lacked and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to technical fields such as deep learning, reinforcement learning, and large language models, and in particular to a large language model dynamic hierarchical prompt compression method, system, device, and storage medium. Technical Background

[0002] With the widespread application of large language models in various natural language processing tasks, their powerful language understanding and generation capabilities have played a significant role in scenarios such as dialogue systems, text summarization, machine translation, and code generation. To better guide large language models in performing specific tasks, researchers have proposed the concept of "cue words." This involves using artificially constructed input text templates to provide task instructions or contextual information to the model, thereby stimulating its potential. Cues play a crucial guiding role in the practical application of large language models, becoming a key bridge between user needs and model output.

[0003] However, with the increasing complexity of tasks and the advancement of prompt engineering technology, the constructed prompts have become increasingly lengthy and complex, sometimes even incorporating multiple context snippets, task instructions, example questions, and more. While this significant increase in prompt length has improved model generation quality and task completion to a certain extent, it also incurs significant computational overhead. This severely impacts model inference efficiency, especially when processing large batches of prompt requests or when deployed on resource-constrained devices, becoming a significant bottleneck in the practical deployment of large language models.

[0004] To maintain model performance while reducing computational overhead, prompt compression technology has emerged. Prompt compression aims to prune or restructure the content of prompts, reducing their length and improving inference speed, without significantly impacting model output. Existing prompt compression methods can be broadly categorized into two types, depending on whether they rely on the internal structure of a large language model: white-box and black-box. White-box methods typically rely on accessing and modifying the internal structure of a large language model, for example by analyzing the self-attention mechanism or fusing semantically similar tokens to achieve prompt simplification. These methods typically require access to the source code and parameter information of the large language model, making them suitable for open models or research environments. However, their application to mainstream commercial closed-source models is limited, making them difficult to meet practical implementation requirements. Black-box compression methods, on the other hand, do not rely on the model's internal structure and instead tailor or optimize prompts based solely on the model's input and output behavior. This makes them more versatile and deployable. However, due to the lack of direct access to the model's attention distribution or intermediate representations during black-box compression, effectively identifying redundant information and appropriately compressing it while maintaining prompt semantic integrity and task clarity remains a challenging technical challenge. Summary of the Invention

[0005] In order to at least partially solve one of the technical problems existing in the prior art, the present invention aims to provide a large language model dynamic hierarchical prompt compression method, system, device and storage medium.

[0006] The technical solution adopted in the present invention is:

[0007] A large language model dynamic hierarchical prompt word compression method includes the following steps:

[0008] Constructing a Markov decision process for prompt compression;

[0009] Train the language model to align the output distribution of the target large model;

[0010] Comprehensively design reward functions for compression ratio, output alignment, and information preservation;

[0011] Training compressed agents using reinforcement learning algorithms and curriculum learning based on proximal policy optimization;

[0012] Use compression agents to dynamically compress input prompts.

[0013] Furthermore, the Markov decision process is constructed as follows:

[0014] Modeling cue word compression as a Markov decision process , where the state space is the prompt word text state during the compression process, action space For each word to be retained or deleted, the state transition function Update the prompt state according to the action performed, the reward function Give immediate rewards based on compression ratio, key information retention and output consistency, strategy Represents the behavior strategy of the compressed agent.

[0015] Furthermore, the trained language model is aligned with the target large model output distribution by minimizing the Kullback-Leibler divergence, which is expressed as follows:

[0016]

[0017] in, represents the fine-tuned language model, is the output distribution of the large language model, It is the Kulbeck-Leibler divergence, which is fine-tuned to make the output distribution consistent with the target large language model, thereby enhancing the semantic accuracy of the compressed hint.

[0018] Furthermore, the reward function comprehensively considers the compression rate of the prompt word, the Kullback-Leibler divergence of the large model output corresponding to the original prompt word and the compressed prompt word, the degree of key information retention and the penalty term constraint compression rate within the set upper and lower limits. and The expression is as follows:

[0019]

[0020] in, 、 and is an adjustable weight, Key information preserving metrics for compressed vs. original cues.

[0021] Furthermore, the reinforcement learning algorithm optimized according to the proximal policy and the curriculum learning training compressed agent maximizes the expected cumulative reward, which is expressed as follows:

[0022]

[0023] in, It means taking the expectation of the following formula. Represents the advantage function, which is used to evaluate the gap between the benefits brought by the actions taken by the agent in the current state and the expected benefits. Combined with course learning, the prompt word compression task is divided into multiple stages, and different compression ratio ranges are set for each stage. , to achieve gradual learning from low difficulty to high difficulty.

[0024] Another technical solution adopted in the present invention is:

[0025] a modeling module for constructing a Markov decision process for prompt compression;

[0026] The fine-tuning module is used to train the language model to align with the output distribution of the target large model;

[0027] Reward building module for comprehensively designing reward functions for compression ratio, output alignment, and information preservation;

[0028] Agent training module, used for training compressed agents using reinforcement learning algorithms and curriculum learning based on proximal policy optimization;

[0029] The compression execution module is used to apply the compression agent to dynamically compress the input prompts.

[0030] Another technical solution adopted in the present invention is:

[0031] A large language model dynamic hierarchical prompt compression device, comprising:

[0032] at least one processor;

[0033] at least one memory for storing at least one program;

[0034] When the at least one program is executed by the at least one processor, the at least one processor implements the method described above.

[0035] Another technical solution adopted in the present invention is:

[0036] A computer-readable storage medium stores a program executable by a processor, wherein the program executable by the processor is used to perform the method described above when executed by the processor.

[0037] The beneficial effects of the present invention are: the present invention aims to solve the problems in current prompt compression technology such as the difficulty in balancing the compression ratio and retaining key information, insufficient generalization of the method, and lack of an adaptive adjustment mechanism through a dynamic hierarchical prompt compression method based on reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 This is a flowchart of a large language model dynamic hierarchical prompt compression method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention. The step numbers in the following embodiments are provided for ease of explanation only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0041] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention.

[0042] In the description of the present invention, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0043] Furthermore, in the description of this invention, unless otherwise specified, "plurality" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0044] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, and connecting should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0045] This embodiment provides a method for dynamic hierarchical prompt compression for a large language model. First, prompt compression is modeled as a Markov decision process. Next, a language model is trained to align the output distribution of a target large language model. A reward function is then designed that comprehensively considers compression ratio, output alignment, and information preservation. Finally, a compression agent is trained using a reinforcement learning algorithm optimized for proximal policies and curriculum learning. Finally, the compression agent is used to dynamically compress input prompts. The method specifically includes the following steps:

[0046] S1. Modeling prompt compression as a Markov decision process.

[0047] This embodiment proposes to model the prompt word compression as a Markov decision process , where the state space is the prompt word text state during the compression process, action space For each word to be retained or deleted, the state transition function Update the prompt state according to the action performed, the reward function Give immediate rewards based on compression ratio, key information retention and output consistency, strategy Represents the behavior strategy of the compressed agent.

[0048] S2. Train a language model to align the output distribution of the target large model.

[0049] This embodiment proposes to train the language model to align the output distribution of the target large model by minimizing the Kullback-Leibler divergence, which is expressed as follows:

[0050]

[0051] in, represents the fine-tuned language model, is the output distribution of the large language model, It is the Kulbeck-Leibler divergence, which is fine-tuned to make the output distribution consistent with the target large language model, thereby enhancing the semantic accuracy of the compressed hint.

[0052] S3. Design a reward function that integrates compression ratio, output alignment, and information preservation.

[0053] This embodiment proposes a reward function that comprehensively considers the compression rate of the prompt word, the Kullback-Leibler divergence of the large model output corresponding to the original prompt word and the compressed prompt word, the degree of key information retention, and the penalty term to constrain the compression rate within the set upper and lower limits. and The expression is as follows:

[0054]

[0055] in, 、 and is an adjustable weight, Key information preserving metrics for compressed vs. original cues.

[0056] S4. Compressed agents are trained based on reinforcement learning algorithms optimized with proximal policies and curriculum learning.

[0057] This example proposes a proximal policy optimization reinforcement learning algorithm and curriculum learning to train a compressed agent to maximize the expected cumulative reward, which is expressed as follows:

[0058]

[0059] in, It means taking the expectation of the following formula. Represents the advantage function, which is used to evaluate the gap between the benefits brought by the actions taken by the agent in the current state and the expected benefits. Combined with course learning, the prompt word compression task is divided into multiple stages, and different compression ratio ranges are set for each stage. , to achieve gradual learning from low difficulty to high difficulty.

[0060] S5. Use compression agents to dynamically compress input prompts.

[0061] After training from step S1 to step S5, a large language model dynamic hierarchical prompt compression agent has been successfully developed, which can effectively compress the length of the prompt word.

[0062] In summary, in order to solve the existing technical problems, the present invention proposes a method that aims to solve the problems in the current prompt compression technology, such as the difficulty in balancing the compression ratio and the retention of key information, the lack of generalization of the method, and the lack of an adaptive adjustment mechanism. This method models prompt compression as a Markov decision process, in which the compression agent interacts with the environment and dynamically chooses whether to retain the current prompt segment in each decision step; at the same time, a fine-tuning language model is introduced to simulate the behavior of the target large model, thereby designing a high-quality reward signal and guiding the compression strategy to be optimized in a more accurate and efficient direction. By combining curriculum learning with reinforcement learning algorithms, the present invention can achieve stable training under different compression task difficulties, and ultimately achieve the technical effect of compressing the prompt length, improving the inference efficiency and maintaining stable output performance without accessing the source code of the large language model.

[0063] This embodiment also provides a large language model dynamic hierarchical prompt compression system, including:

[0064] a modeling module for constructing a Markov decision process for prompt compression;

[0065] The fine-tuning module is used to train the language model to align with the output distribution of the target large model;

[0066] Reward building module for comprehensively designing reward functions for compression ratio, output alignment, and information preservation;

[0067] Agent training module, used for training compressed agents using reinforcement learning algorithms and curriculum learning based on proximal policy optimization;

[0068] The compression execution module is used to apply the compression agent to dynamically compress the input prompts.

[0069] A large language model dynamic hierarchical prompt compression system of this embodiment can execute a large language model dynamic hierarchical prompt compression method provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0070] This embodiment further provides a large language model dynamic hierarchical prompt compression device, comprising:

[0071] at least one processor;

[0072] at least one memory for storing at least one program;

[0073] When the at least one program is executed by the at least one processor, the at least one processor implements the following Figure 1 The method shown.

[0074] A large language model dynamic hierarchical prompt compression device of this embodiment can execute a large language model dynamic hierarchical prompt compression method provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0075] The present application also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs Figure 1 The method shown.

[0076] This embodiment also provides a storage medium storing instructions or programs that can execute a low-light image enhancement processing method provided by an embodiment of the method of the present invention. When the instructions or program are run, any combination of implementation steps of the method embodiment can be executed, and the corresponding functions and beneficial effects of the method can be obtained.

[0077] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.

[0078] Furthermore, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise indicated, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It will also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skill of an engineer. Therefore, a person skilled in the art using ordinary skill will be able to implement the present invention set forth in the claims without undue experimentation. It will also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0079] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0080] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0081] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in another suitable manner as necessary, and then storing it in a computer memory.

[0082] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0083] In the above description of this specification, reference to the terms "one embodiment / example," "another embodiment / example," or "certain embodiments / examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0084] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

[0085] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A large language model dynamic hierarchical prompt word compression method, characterized in that: The following steps are involved: Modeling cue word compression as a Markov decision process , where the state space is the prompt word text state during the compression process, action space For each word to be retained or deleted, the state transition function Update the prompt state according to the action performed, the reward function Give immediate rewards based on compression ratio, key information retention and output consistency, strategy Represents the behavior strategy of the compressed agent; Construct a compression agent and train it using reinforcement learning algorithm. Output the retention probability of each word and at each time step Select Action To decide whether to keep or delete the current word; Fine-tune a language model so that its output distribution aligns with the target large language model, and and its target output Perform an optimization to minimize the Kullback–Lebler divergence: in represents the fine-tuned language model, is the output distribution of the large language model, The Kulbeck-Leibler divergence is fine-tuned to align the output distribution with the target large language model, thereby enhancing the semantic accuracy of the compressed hints. Constructing a reward function , taking into account the following factor: (1) Compression rate of prompt words; (2) Kullback-Leibler divergence of the large model output corresponding to the original prompt word and the compressed prompt word; (3) Degree of retention of key information; (4) The penalty term constrains the compression rate to be within the set upper and lower limits. and between; Training the agent policy network based on the Proximal Policy Optimization (PPO) algorithm and value network , so that the expected cumulative reward is maximized: in It means taking the expectation of the following formula. represents the advantage function, which is used to evaluate the gap between the benefits brought by the actions taken by the agent in the current state and the expected benefits; Combined with course learning, the prompt word compression task is divided into multiple different stages, and different compression ratio ranges are set in each stage. , to achieve gradual learning from low difficulty to high difficulty.

2. The method according to claim 1, characterized in that The policy network of the compressed agent It takes the feature vector of the input prompt word as input, outputs the retention probability of each word unit, and samples the compression action accordingly.

3. The method according to claim 1, characterized in that The language model fine-tuning is done by supervising the training data pairs so that the fine-tuned language model simulates the behavior of the target large language model, and constructs a training set (the original prompt word , large language model output ).

4. The method according to claim 1, wherein The reward function for: in 、 and is an adjustable weight, Key information preserving metrics for compressed vs. original cues.

5. The method according to claim 1, characterized in that The course learning strategy gradually tightens the upper and lower limits of the compression ratio and , in order to control the compression difficulty and guide the agent to gradually learn from high retention ratio to high compression rate tasks.

6. A large language model dynamic hierarchical prompt compression system implementing the method according to any one of claims 1 to 5, characterized in that: include: a modeling module for constructing a Markov decision process for prompt compression; The fine-tuning module is used to train the language model to align with the output distribution of the target large model; Reward building module for comprehensively designing reward functions for compression ratio, output alignment, and information preservation; Agent training module, used for training compressed agents using reinforcement learning algorithms and curriculum learning based on proximal policy optimization; The compression execution module is used to apply the compression agent to dynamically compress the input prompts.

7. A large language model dynamic hierarchical prompt compression device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 5 when executed by the processor.