Method and device for compressing cue words of large language model and medium

By grading the importance of prompt words in large language models, grammar tree construction, information entropy alignment and recursive pruning, efficient compressed prompts are generated, which solves the problems of increased computational cost and limited model processing capabilities caused by long prompts, and achieves more efficient generation efficiency and semantic coherence.

CN119940540APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510010671.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When a large language model based on Transformer architecture processes long prompts, the calculation cost increases sharply, and the model has a fixed upper limit on the number of input tokens, so it cannot process longer inputs, affecting performance.

Method used

By calculating the importance score of each vocabulary, removing redundant vocabulary and statements, building a local syntax tree and a global parse tree, calculating local information entropy for node alignment, recursively pruning the global parse tree, generating compression prompts, and ensuring the quality of compression prompts through iterative optimization.

Benefits of technology

It significantly improves the refinement of prompt words, reduces the processing burden of large language models, improves generation efficiency, retains the hierarchy and semantic coherence of prompt words, and ensures accurate communication of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940540A_ABST
    Figure CN119940540A_ABST
Patent Text Reader

Abstract

The invention discloses a cue word compression method and device of a large language model and a medium. The method comprises the steps that redundant vocabularies and redundant statements are recognized and removed according to the size relation between the vocabulary importance score and the importance threshold value of each vocabulary after original cue word segmentation; constructing a local syntax tree of each statement, and constructing a global analysis tree by the local syntax trees according to the dependency relationship among the cue words; calculating a local information entropy so as to align the LLM word segmentation device with nodes on the global analysis tree and adjust node values of the global analysis tree; trimming the global analysis tree through a recursive algorithm according to the adjusted node value, and generating a compression prompt corresponding to the prompt word; according to the method, the compression prompt is evaluated, the compression prompt is iteratively optimized according to an evaluation result, the cue word is reconstructed according to the optimized compression prompt, and compression of the cue word is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular to a prompt word compression method, device and medium for a large language model. Background Art

[0002] In the field of Natural Language Processing (NLP), with the rapid development of artificial intelligence technology, large language models (LLMs) have become an important tool for solving complex language tasks. These models, such as BERT and GPT series, are based on the Transformer architecture and use deep learning technology to learn language rules and patterns from massive data, thereby demonstrating unprecedented performance in multiple aspects such as question answering, text summarization, multimodal generation, and information extraction. They can understand and generate natural language close to human level, which has greatly promoted the advancement of NLP technology and the expansion of its application scope.

[0003] In order to further improve the performance of LLMs, researchers have found that providing the model with contextual information closely related to the task, namely "prompts", is an effective way. Through carefully designed prompts, the model can be guided to better understand the task requirements and produce more accurate and expected outputs. In recent years, a series of advanced prompting technologies have emerged, such as contextual learning, which uses historical conversations or related documents to build rich contexts to help the model better understand the current task; thought chain reasoning, which guides the model to gradually derive the answer by simulating the human reasoning process; retrieval-enhanced generation, which combines external knowledge bases to provide additional information support for the model to enhance its generation capabilities. These technologies have shown significant advantages in dealing with problems that require long-tail knowledge (i.e. rare or domain-specific knowledge) and complex reasoning.

[0004] However, LLMs based on the Transformer architecture face the problem of sharply rising computational costs as the input length increases. This is because the self-attention mechanism in the Transformer model needs to calculate the relationship between each token and all other tokens, resulting in a computational complexity that grows with the square of the input length. In addition, these models usually have a fixed upper limit on the number of input tokens, and once this limit is exceeded, they cannot handle longer inputs. Therefore, long prompts not only increase the computational burden, but may also cause the model to be unable to process the entire input sequence, thus affecting its performance.

[0005] To address this challenge, researchers have proposed prompt compression techniques, which aim to reduce computational costs by shortening prompts and enable the model to handle longer contexts. Existing prompt compression methods are mainly divided into two categories: generative compression and selective compression. Generative compression uses the language model itself to generate short prompts, but this method may introduce information distortion because the generated prompts may not fully retain the key information in the original prompts. Selective compression, on the other hand, selectively retains the key parts by evaluating the importance of each token. Although this method avoids the problem of information distortion, it often ignores the language rules and the overall structure of the prompt. Especially for long prompts, existing methods may not be able to fully retain the logical relationship and overall hierarchical structure between sentences when processing, resulting in the compressed prompts being semantically incoherent or missing important information. Summary of the invention

[0006] The embodiments of the present application provide a method, device and medium for compressing prompt words of a large language model to solve the above-mentioned technical problems.

[0007] On the one hand, an embodiment of the present application provides a method for compressing prompt words of a large language model, including:

[0008] Calculating the vocabulary importance score corresponding to each word after the original prompt word segmentation, and identifying and removing redundant words and redundant sentences according to the relationship between the vocabulary importance score and the importance threshold;

[0009] Constructing a corresponding local syntax tree for each sentence in the prompt words after removing redundancy, and constructing a global parse tree from the local syntax trees according to the dependency relationship between the prompt words;

[0010] Calculating local information entropy to align the LLM word segmenter with the nodes on the global parse tree based on the local information entropy, and adjusting the node values ​​on the global parse tree;

[0011] Pruning the global parse tree according to the adjusted node values ​​through a recursive algorithm, and generating a compressed prompt corresponding to the prompt word;

[0012] The compression prompt is evaluated, the compression prompt is iteratively optimized according to the evaluation result, and the prompt word is reconstructed according to the optimized compression prompt to achieve compression of the prompt word.

[0013] In one implementation of the present application, the vocabulary importance score corresponding to each vocabulary after the original prompt word segmentation is calculated, specifically including:

[0014] The original prompt words are segmented by the LLM segmenter to obtain several independent lexical units;

[0015] Obtaining a high-dimensional embedding vector for each word through a pre-trained language model, and calculating the cosine similarity between the high-dimensional embedding vectors to evaluate the semantic relevance between the words;

[0016] According to the frequency of occurrence and position information of the vocabulary in the original prompt words and the strength of semantic association with other vocabulary, a corresponding semantic contribution score is assigned to each vocabulary;

[0017] The vocabulary importance score corresponding to each vocabulary word is calculated according to the occurrence frequency, the position information, the semantic association strength and the semantic contribution score.

[0018] In one implementation of the present application, redundant words and redundant sentences are identified and removed according to the relationship between the word importance score and the importance threshold, specifically including:

[0019] Compare the vocabulary importance score with a preset importance threshold, and determine that the vocabulary whose vocabulary importance score is less than the importance threshold is a redundant vocabulary;

[0020] Determining sentences containing redundant words, and calculating a comprehensive score of the sentence based on the lexical importance scores of the redundant words, the correlation between the words, and the semantic completeness of the sentence;

[0021] If the comprehensive score is less than a preset sentence score threshold, the sentence is determined to be a redundant sentence, and an overall semantic check is performed on the prompt word based on the marked redundant words and redundant sentences;

[0022] According to the inspection results, redundant words and redundant sentences that have no impact on the overall semantics are removed, and the prompt words after the redundancy is removed are verified for integrity, and the prompt words are adjusted according to the verification results to obtain optimized prompt words.

[0023] In one implementation of the present application, a corresponding local syntax tree is constructed for each sentence in the prompt word after removing redundancy, and a global parsing tree is constructed from the local syntax tree according to the dependency relationship between the prompt words, specifically including:

[0024] Based on the language rules, a corresponding local syntax tree is constructed for each sentence in the prompt word after removing redundancy, and a corresponding virtual node is added to the root node of each local syntax tree to serve as the local root node of the local syntax tree;

[0025] According to the dependency relationship between the sentences, paragraphs and parts in the prompt word, the hierarchy and association between the local root nodes are determined to construct a global virtual root node;

[0026] The local root node is used as a child node of the global virtual root node, and the corresponding local syntax tree is connected through directed edges according to the actual order of the prompt word to construct the corresponding global parse tree.

[0027] In one implementation of the present application, calculating the local information entropy to align the LLM word segmenter with the node on the global parse tree based on the local information entropy specifically includes:

[0028] Divide the prompt words after removing redundancy into several sentences, calculate the conditional probability corresponding to each word in each sentence, and calculate the local information entropy corresponding to the word;

[0029] According to the segmentation of the vocabulary corresponding to each node on the global parse tree in the LLM word segmenter, the corresponding fine-grained token is determined, and the information entropy of the fine-grained token is calculated as the alignment information entropy of the node; the alignment information entropy is used to represent the information amount and uncertainty of the vocabulary corresponding to the node in the context;

[0030] In the descending order of the alignment information entropy, the nodes on the global parse tree are compared with the nodes in the LLM word segmenter to see if they are aligned. If not, the nodes are aligned according to the alignment information entropy of the nodes and the semantic relationship of the surrounding nodes.

[0031] In one implementation of the present application, pruning the global parse tree according to the adjusted node values ​​by a recursive algorithm specifically includes:

[0032] Based on a preset recursive function, calculating the importance value of each node on the global parse tree;

[0033] When the importance value is less than a preset pruning threshold, the node is marked as pruned, and the sum of the pruned subtree values ​​is returned;

[0034] When the importance value is greater than the pruning threshold, continue to recursively traverse the child nodes of the node and calculate the sum of the subtree values ​​corresponding to the node.

[0035] In one implementation of the present application, generating a compressed prompt corresponding to a prompt word specifically includes:

[0036] Traverse the pruned global parse tree and extract the vocabulary corresponding to the unmarked pruned nodes;

[0037] Merge the word into the serialized string and check whether the word is the last word in the serialized string. If not, add a space after the word.

[0038] Until the word is the last word of the serialized character string, the serialized character string is checked, and abnormal adjustments are made according to the checking result to generate a compressed prompt corresponding to the prompt word; the abnormal adjustment includes: removing redundant spaces and adjusting punctuation marks.

[0039] In one implementation of the present application, the compression hint is evaluated, and the compression hint is iteratively optimized according to the evaluation result, specifically including:

[0040] Inputting the compressed prompt into a large language model, and evaluating the generation results of the compressed prompt and the original prompt in the large language model;

[0041] According to the evaluation results, identifying and recording abnormal nodes with problems, and adjusting the weights of the abnormal nodes on the global parse tree through a bidirectional path dependency propagation method;

[0042] According to the adjusted node values, the global parse tree is pruned through a recursive algorithm to generate new compression hints, until there are no abnormalities in the evaluation results corresponding to the new compression hints, thus completing the iteration of the compression hints.

[0043] On the other hand, the embodiment of the present application further provides a prompt word compression device for a large language model, the device comprising:

[0044] at least one processor;

[0045] and, a memory communicatively coupled to the at least one processor;

[0046] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the prompt word compression method for a large language model as described above.

[0047] On the other hand, an embodiment of the present application further provides a non-volatile computer storage medium storing computer executable instructions, which, when executed, implement the above-mentioned method for compressing prompt words of a large language model.

[0048] The present application provides a method, device and medium for compressing prompt words of a large language model, which at least have the following beneficial effects:

[0049] By calculating the importance score of each word and removing redundant words and sentences based on the importance threshold, the refinement of the prompt words can be significantly improved, ensuring that the retained information is the most critical and core part of the generation task, thereby reducing the processing burden of the large language model and improving the generation efficiency; by constructing local syntax trees and global parse trees, not only the grammatical structure within a single sentence is considered, but also the dependency relationship between the prompt words as a whole is comprehensively considered, which helps to retain the hierarchical structure and semantic coherence of the prompt words, and ensure the accurate transmission of information even in the compression process; by calculating the local information entropy and aligning the LLM segmenter with the nodes on the global parse tree, the node value can be adjusted more accurately, so that the compressed prompt words are more in line with the word segmentation habits and semantic understanding mechanism of the large language model, thereby improving the quality of the sentence. The fluency and accuracy of the generated text are improved; the application of the recursive algorithm makes the pruning process of the global parse tree more efficient and automated, and can dynamically decide which information to retain and which to discard based on the node value, which not only improves the compression efficiency, but also ensures that the compressed prompt words still have enough information to support high-quality generation; by evaluating the compressed prompts and iteratively optimizing according to the evaluation results, the compression strategy can be continuously improved to ensure that the compressed prompt words are both concise and informative, which helps to gradually improve the compression effect until it reaches the optimal state; the compressed and optimized prompt words can be processed by the large language model more quickly to generate text content that is more in line with user expectations, which not only improves the user experience, but also enhances the user's satisfaction and trust in the generation results of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0051] Figure 1 A flowchart of a method for compressing prompt words of a large language model provided in an embodiment of the present application;

[0052] Figure 2 A schematic diagram of the internal structure of a prompt word compression device for a large language model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0054] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0055] Figure 1 A flowchart of a prompt word compression method for a large language model provided in an embodiment of the present application.

[0056] The analysis method involved in the embodiments of the present application can be implemented by a terminal device or a server, and the present application does not impose any special restrictions on this. For the convenience of understanding and description, the following embodiments are described in detail by taking a server as an example.

[0057] It should be noted that the server may be a single device or a system consisting of multiple devices, that is, a distributed server, and this application does not make any specific limitation on this.

[0058] like Figure 1 As shown, the embodiment of the present application provides a method for compressing prompt words of a large language model, including:

[0059] 101. Calculate the vocabulary importance score corresponding to each word after the original prompt word segmentation, and identify and remove redundant words and redundant sentences based on the relationship between the vocabulary importance score and the importance threshold.

[0060] Specifically, in one embodiment of the present application, the vocabulary importance score corresponding to each vocabulary after the original prompt word segmentation is calculated, specifically including:

[0061] The original prompt words are segmented by the LLM segmenter to obtain several independent lexical units;

[0062] Through the pre-trained language model, the high-dimensional embedding vector of each word is obtained, and the cosine similarity between the high-dimensional embedding vectors is calculated to evaluate the semantic relevance between words;

[0063] Combined with the frequency of occurrence of the word in the original prompt word, position information and the strength of semantic association with other words, each word is assigned a corresponding semantic contribution score;

[0064] The lexical importance score corresponding to each word is calculated based on the frequency of occurrence, position information, semantic association strength and semantic contribution score.

[0065] In one embodiment, the original prompt words are firstly subjected to detailed word segmentation using natural language processing technology to decompose the text into independent vocabulary units; then the relative importance of each word in the entire prompt word is calculated. Specifically, a high-dimensional embedding vector of each word is obtained through a pre-trained language model, and the cosine similarity between the word embedding vectors is calculated to evaluate the semantic relevance between words. The semantic contribution of each word is comprehensively evaluated by combining the frequency, position and semantic association of the word with other words; finally, a multi-dimensional comprehensive evaluation of the prompt words is performed by combining the importance score and semantic contribution of the word.

[0066] In one embodiment of the present application, redundant words and redundant sentences are identified and removed according to the relationship between the word importance score and the importance threshold, specifically including:

[0067] Compare the word importance score with a preset importance threshold, and determine that words whose word importance score is less than the importance threshold are redundant words;

[0068] Identify sentences containing redundant words and calculate a comprehensive score of the sentences based on the lexical importance scores of the redundant words, the degree of relevance between the words, and the semantic completeness of the sentences;

[0069] If the comprehensive score is less than the preset sentence score threshold, the sentence is determined to be a redundant sentence, and the prompt word is subjected to an overall semantic check based on the marked redundant words and redundant sentences;

[0070] According to the inspection results, redundant words and sentences that have no impact on the overall semantics are removed, and the prompt words after the redundancy is removed are verified for integrity. According to the verification results, the prompt words are adjusted to obtain optimized prompt words.

[0071] In one embodiment, based on the calculated vocabulary importance score, an importance threshold T is set, and the vocabulary with a score lower than the threshold T is marked as redundant information; then, the comprehensive score of the sentences or phrases containing redundant vocabulary is calculated, and if the comprehensive score of a sentence is lower than the predetermined sentence score threshold Ts, the entire sentence is marked as redundant. Then, after marking the redundant vocabulary and sentences, an overall check is performed to identify and mark the parts that contribute less or no contribution to the final result, ensuring that the prompt words can still maintain semantic integrity and logical coherence after removing these redundant information. Finally, the coherence and integrity of the prompt words after removing the redundant information are repeatedly verified, and fine-tuned when necessary to ensure that the final prompt words can maintain both semantic integrity and effectively reduce length.

[0072] 102. Construct a corresponding local syntax tree for each sentence in the prompt words after removing redundancy, and construct a global parse tree from the local syntax trees according to the dependency relationship between the prompt words.

[0073] Specifically, in one embodiment of the present application, a corresponding local syntax tree is constructed for each sentence in the prompt word after removing redundancy, and a global parse tree is constructed from the local syntax tree according to the dependency relationship between the prompt words, specifically including:

[0074] Based on the language rules, a corresponding local syntax tree is constructed for each sentence in the prompt word after removing redundancy, and a corresponding virtual node is added to the root node of each local syntax tree as the local root node of the local syntax tree;

[0075] According to the dependencies among sentences, paragraphs and parts in the prompt word, the hierarchy and association between local root nodes are determined to construct a global virtual root node;

[0076] The local root node is used as the child node of the global virtual root node, and the corresponding global parse tree is constructed by connecting the local syntax trees through directed edges according to the actual order of the corresponding local syntax trees in the prompt word.

[0077] In one embodiment, a grammar analysis tool is used to generate each sentence P based on the language rules of each sentence. j Construct a local syntax tree T j , and obtain the corresponding n local parse trees [T1,T2,…,T n ]. For each local tree T j Add a virtual node to the root node of the global tree T, and make the root nodes of all local trees the child nodes of the virtual root node of the global tree T. Build the global tree T to reflect the structure of the entire prompt, including sentences, paragraphs, chapters and other levels.

[0078] 103. Calculate the local information entropy to align the LLM tokenizer with the nodes on the global parse tree based on the local information entropy, and adjust the node values ​​on the global parse tree.

[0079] Specifically, in one embodiment of the present application, the local information entropy is calculated to align the LLM word segmenter with the nodes on the global parse tree based on the local information entropy, specifically including:

[0080] The prompt words after removing redundancy are divided into several sentences, so as to calculate the conditional probability corresponding to each word in each sentence, and calculate the local information entropy corresponding to the word;

[0081] According to the segmentation of the vocabulary corresponding to each node on the global parse tree in the LLM tokenizer, the corresponding fine-grained token is determined, and the information entropy of the fine-grained token is calculated as the alignment information entropy of the node; the alignment information entropy is used to represent the amount of information and uncertainty of the vocabulary corresponding to the node in the context;

[0082] In the order from the largest to the smallest alignment information entropy, compare whether the nodes on the global parse tree are aligned with the nodes in the LLM tokenizer one by one. If not, align the nodes according to the alignment information entropy of the nodes and the semantic relationships of the surrounding nodes.

[0083] In one embodiment, a parse tree is obtained for each sentence according to language rules. Specifically, a grammar parsing tool is used to analyze the grammatical structure of each sentence in the prompt, and a tree diagram representing the grammatical structure of the sentence is obtained.

[0084] In one embodiment, given a long prompt P containing multiple sentences, it is split into a list of sentences, [P1, P2,..., P n . For each word r j in the sentence P ji calculate the conditional probability: p(r j,i ∣ r j,<i ). Where j represents the sentence index, i represents the word index, and r<i represents all the words before the word r ji .

[0085] Calculate the information entropy E(r ji ) of each word, and ignore the words outside the sentence: E(r j,i ) = -logp(r j,i ∣ r j,<i ) ≈ -logp(r j,i ∣ r j,<i, r<j), where r<j represents the words in all the sentences before the sentence P j .

[0086] To ensure that the minimum tokens maintain semantic integrity and achieve alignment between the parse tree tokenizer and the LLM tokenizer, first obtain the actual nodes of the parse tree tokenizer and the fine-grained tokens of the LLM tokenizer and their information entropy; then, for each actual node v ji in the parse tree, check its segmentation in the LLM tokenizer, and calculate the sum of the information entropy of these fine-grained tokens as the alignment information entropy E(v ji) ; ensure semantic integrity.

[0087] In one embodiment, a two-way path-dependent propagation is used to adjust the node values on the global tree. Starting from the root node of the global tree, adjust the weights of each node from top to bottom. The weight of the root node affects its child nodes to ensure that important information is retained. Leaf-to-root propagation: Starting from the leaf nodes, adjust the weights of each node from bottom to top.

[0088] For each node v, calculate its degree of dependence on the path. The degree of dependence is based on factors such as the distance between nodes (i.e., path length), the importance of nodes (such as the level of information entropy), and the strength of connections between nodes (such as weight distribution). Adjust the weight or information entropy value of the node according to the degree of path dependence. When adjusting, consider both the influence from the parent node and the influence from the child node, so that important information can be better retained while ensuring the accuracy of information transmission.

[0089] The weight or information entropy value of each node is continuously adjusted in multiple iterations until convergence or the predetermined number of iterations is reached. The above steps are repeated in each iteration to ensure that the information is evenly distributed and consistent throughout the tree structure.

[0090] 104. Prune the global parse tree according to the adjusted node values ​​through a recursive algorithm, and generate a compressed prompt corresponding to the prompt word.

[0091] Specifically, in one embodiment of the present application, pruning the global parse tree according to the adjusted node value through a recursive algorithm specifically includes:

[0092] Based on the preset recursive function, calculate the importance value of each node on the global parse tree;

[0093] When the importance value is less than the preset pruning threshold, the node is marked as pruned and the sum of the pruned subtree values ​​is returned;

[0094] When the importance value is greater than the pruning threshold, continue to recursively traverse the child nodes of the node and calculate the sum of the subtree values ​​corresponding to the node.

[0095] In one embodiment, a recursive algorithm is used to prune the global tree according to the adjusted node values. Unimportant nodes in the global tree T are pruned according to the adjusted values ​​of the nodes. Recursive traversal starts from the root node, and for each node Ni, if pruning the node can maximize the compression ratio of the sum of the values ​​of the subtrees, the node is pruned.

[0096] Specifically, we first define a recursive function PruneTree, which starts from the root node of the global tree, recursively traverses each node, and evaluates the adjusted importance value of each node. In this process, each node is judged: if the importance value of the node is lower than a dynamically set threshold, or the length of the subtree containing the node exceeds the predetermined limit, we mark the node for pruning. In order to more accurately control the compression process, several optimization strategies are implemented: dynamically adjust the pruning threshold according to the position of the node in the tree and the number of words currently used, increase the importance value weight of key sentences or structurally important nodes, and perform a comprehensive post-processing check after the recursion is completed to ensure that the compressed prompt words meet the length requirements while retaining key information and ensuring the integrity and accuracy of semantics.

[0097] In one embodiment of the present application, generating a compressed prompt corresponding to a prompt word specifically includes:

[0098] Traverse the pruned global parse tree and extract the vocabulary corresponding to the unmarked pruned nodes;

[0099] Merge the word into the serialized string and check if the word is the last word in the serialized string. If not, add a space after the word.

[0100] Until the word is the last word of the serialized string, the serialized string is checked, and abnormal adjustments are made according to the check results to generate a compressed prompt corresponding to the prompt word; the abnormal adjustment includes: removing redundant spaces and adjusting punctuation marks.

[0101] In one embodiment, the pruned global tree is traversed to generate the compression hint P cp :Starting from the root node, perform a depth-first search, check each node one by one, skip the nodes and their subtrees that have been marked for deletion; for the remaining non-deleted nodes, extract the corresponding words or tokens and merge them into the serialized string SS, and pay attention to adding appropriate spaces after each token (except for the last token or special cases where spaces are not required); then continue to traverse all child nodes of the current node until the entire tree is completely traversed; finally, check and perform necessary post-processing on the generated serialized string S to ensure that it is grammatically correct and semantically complete, thereby forming a compression hint P cp , this prompt is a simplified version of the original prompt P. Although it is shortened in length, it still retains the key information.

[0102] 105. Evaluate the compression prompt, iteratively optimize the compression prompt according to the evaluation result, and reconstruct the prompt word according to the optimized compression prompt to achieve compression of the prompt word.

[0103] Specifically, in one embodiment of the present application, the compression hint is evaluated, and the compression hint is iteratively optimized according to the evaluation result, which specifically includes:

[0104] Feed the compressed prompt into the large language model and evaluate the compressed prompt and the original prompt in the large language model.

[0105] According to the evaluation results, identify and record the abnormal nodes with problems, and adjust the weights of the abnormal nodes on the global parse tree through the bidirectional path dependency propagation method;

[0106] According to the adjusted node values, the global parse tree is pruned through a recursive algorithm to generate new compression hints, until there are no abnormalities in the evaluation results corresponding to the new compression hints, thus completing the iteration of the compression hints.

[0107] In one embodiment, the compression hint P cp Input it into LLM and observe the performance of the model when processing this prompt, including the quality of the generated text, relevance, and consistency of the context. Record the performance indicators of LLM when processing compression prompts, such as generation speed, response time, fluency and logic of generated text, etc. cp The original prompt P is input into the same LLM and the differences in the generated results are compared to evaluate whether the compressed prompt achieves the expected simplification effect while retaining enough information to support high-quality generation.

[0108] Based on the test results, identify the problems in the compression prompts and record them for the next step of loop node adjustment.

[0109] After evaluation, if compression hint P is found cp There are deficiencies, and the node values ​​on the global tree need to be adjusted again to optimize the compression effect.

[0110] Based on the evaluation results of the compression hint, determine which nodes need further adjustment to improve the quality of the compression hint. For the nodes that need to be adjusted, recalculate their weights or information entropy values. Use the bidirectional path dependency propagation method to adjust the node weights from top to bottom starting from the root node, and adjust the node weights from bottom to top starting from the leaf node to ensure that the adjusted node values ​​can better reflect the importance and semantic integrity of the information. Based on the adjusted node values, use the recursive algorithm again to prune the global tree and remove nodes that are no longer important to further optimize the compression hint. Repeat the above steps until the compression hint P is reached. cp Until satisfactory evaluation criteria are met or the predetermined number of iterations is reached.

[0111] Then, the prompt words are reconstructed, and a new prompt version is generated based on the optimized prompt words using the language model. By comparing multiple generated results, the optimal version is selected to ensure output quality. While keeping the core meaning of the prompt words unchanged, redundant words or sentences are removed, and strategies such as vocabulary replacement and syntactic simplification are applied. Ensure that the reconstructed prompt words still have good readability and logical coherence, so that the language model can correctly understand and execute the prompt instructions.

[0112] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a prompt word compression device for a large language model, whose structure is as follows: Figure 2 shown.

[0113] Figure 2 The internal structure diagram of a prompt word compression device for a large language model provided in an embodiment of the present application. Figure 2 As shown, the device includes:

[0114] at least one processor;

[0115] and, a memory communicatively coupled to the at least one processor;

[0116] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to:

[0117] Calculate the vocabulary importance score of each word after the original prompt word segmentation, and identify and remove redundant words and redundant sentences according to the relationship between the vocabulary importance score and the importance threshold;

[0118] Construct a corresponding local syntax tree for each sentence in the prompt words after removing redundancy, and construct a global parse tree from the local syntax trees according to the dependency relationship between the prompt words;

[0119] Calculate the local information entropy to align the LLM tokenizer with the nodes on the global parse tree based on the local information entropy and adjust the node values ​​on the global parse tree;

[0120] Prune the global parse tree according to the adjusted node values ​​through a recursive algorithm, and generate compressed prompts corresponding to the prompt words;

[0121] The compression prompt is evaluated, the compression prompt is iteratively optimized according to the evaluation result, and the prompt word is reconstructed according to the optimized compression prompt to achieve compression of the prompt word.

[0122] The present application also provides a non-volatile computer storage medium storing computer executable instructions. When the computer executable instructions are executed, they can:

[0123] Calculate the vocabulary importance score of each word after the original prompt word segmentation, and identify and remove redundant words and redundant sentences according to the relationship between the vocabulary importance score and the importance threshold;

[0124] Construct a corresponding local syntax tree for each sentence in the prompt words after removing redundancy, and construct a global parse tree from the local syntax trees according to the dependency relationship between the prompt words;

[0125] Calculate the local information entropy to align the LLM tokenizer with the nodes on the global parse tree based on the local information entropy and adjust the node values ​​on the global parse tree;

[0126] Prune the global parse tree according to the adjusted node values ​​through a recursive algorithm, and generate compressed prompts corresponding to the prompt words;

[0127] The compression prompt is evaluated, the compression prompt is iteratively optimized according to the evaluation result, and the prompt word is reconstructed according to the optimized compression prompt to achieve compression of the prompt word.

[0128] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0129] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0130] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0131] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0132] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0133] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0135] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0136] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0137] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0138] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0139] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for compressing prompt words of a large language model, characterized in that: The method comprises: Calculating the vocabulary importance score corresponding to each word after the original prompt word segmentation, and identifying and removing redundant words and redundant sentences according to the relationship between the vocabulary importance score and the importance threshold; Constructing a corresponding local syntax tree for each sentence in the prompt words after removing redundancy, and constructing a global parse tree from the local syntax trees according to the dependency relationship between the prompt words; Calculating local information entropy to align the LLM word segmenter with the nodes on the global parse tree based on the local information entropy, and adjusting the node values ​​on the global parse tree; Pruning the global parse tree according to the adjusted node values ​​through a recursive algorithm, and generating a compressed prompt corresponding to the prompt word; The compression prompt is evaluated, the compression prompt is iteratively optimized according to the evaluation result, and the prompt word is reconstructed according to the optimized compression prompt to achieve compression of the prompt word.

2. The method for compressing prompt words of a large language model according to claim 1, characterized in that: Calculate the vocabulary importance score of each word after the original prompt word segmentation, including: The original prompt words are segmented by the LLM segmenter to obtain several independent lexical units; Obtaining a high-dimensional embedding vector for each word through a pre-trained language model, and calculating the cosine similarity between the high-dimensional embedding vectors to evaluate the semantic relevance between the words; According to the frequency of occurrence and position information of the vocabulary in the original prompt words and the strength of semantic association with other vocabulary, a corresponding semantic contribution score is assigned to each vocabulary; The vocabulary importance score corresponding to each vocabulary word is calculated according to the occurrence frequency, the position information, the semantic association strength and the semantic contribution score.

3. The method for compressing prompt words of a large language model according to claim 1, characterized in that: According to the relationship between the importance score of the vocabulary and the importance threshold, redundant vocabulary and redundant sentences are identified and removed, specifically including: Compare the vocabulary importance score with a preset importance threshold, and determine that the vocabulary whose vocabulary importance score is less than the importance threshold is a redundant vocabulary; Determining sentences containing redundant words, and calculating a comprehensive score of the sentence based on the lexical importance scores of the redundant words, the correlation between the words, and the semantic completeness of the sentence; If the comprehensive score is less than a preset sentence score threshold, the sentence is determined to be a redundant sentence, and an overall semantic check is performed on the prompt word based on the marked redundant words and redundant sentences; According to the inspection results, redundant words and redundant sentences that have no impact on the overall semantics are removed, and the prompt words after the redundancy is removed are verified for integrity, and the prompt words are adjusted according to the verification results to obtain optimized prompt words.

4. The method for compressing prompt words of a large language model according to claim 1, characterized in that: Constructing a corresponding local syntax tree for each sentence in the prompt words after removing redundancy, and constructing a global parse tree from the local syntax tree according to the dependency relationship between the prompt words, specifically including: Based on the language rules, a corresponding local syntax tree is constructed for each sentence in the prompt word after removing redundancy, and a corresponding virtual node is added to the root node of each local syntax tree to serve as the local root node of the local syntax tree; According to the dependency relationship between the sentences, paragraphs and parts in the prompt word, the hierarchy and association between the local root nodes are determined to construct a global virtual root node; The local root node is used as a child node of the global virtual root node, and the corresponding local syntax tree is connected through directed edges according to the actual order of the prompt word to construct the corresponding global parse tree.

5. The method for compressing prompt words of a large language model according to claim 1, characterized in that: Calculating local information entropy to align the LLM word segmenter with the nodes on the global parse tree based on the local information entropy, specifically including: Divide the prompt words after removing redundancy into several sentences, calculate the conditional probability corresponding to each word in each sentence, and calculate the local information entropy corresponding to the word; According to the segmentation of the vocabulary corresponding to each node on the global parse tree in the LLM word segmenter, the corresponding fine-grained token is determined, and the information entropy of the fine-grained token is calculated as the alignment information entropy of the node; the alignment information entropy is used to represent the information amount and uncertainty of the vocabulary corresponding to the node in the context; In the descending order of the alignment information entropy, the nodes on the global parse tree are compared with the nodes in the LLM word segmenter to see if they are aligned. If not, the nodes are aligned according to the alignment information entropy of the nodes and the semantic relationship of the surrounding nodes.

6. The method for compressing prompt words of a large language model according to claim 1, characterized in that: The global parse tree is pruned according to the adjusted node values ​​by a recursive algorithm, specifically comprising: Based on a preset recursive function, calculating the importance value of each node on the global parse tree; When the importance value is less than a preset pruning threshold, the node is marked as pruned, and the sum of the pruned subtree values ​​is returned; When the importance value is greater than the pruning threshold, continue to recursively traverse the child nodes of the node and calculate the sum of the subtree values ​​corresponding to the node.

7. The method for compressing prompt words of a large language model according to claim 1, characterized in that: Generate a compressed prompt corresponding to the prompt word, including: Traverse the pruned global parse tree and extract the vocabulary corresponding to the unmarked pruned nodes; Merge the word into the serialized string and check whether the word is the last word of the serialized string. If not, add a space after the word. Until the word is the last word of the serialized character string, the serialized character string is checked, and abnormal adjustments are made according to the checking result to generate a compressed prompt corresponding to the prompt word; the abnormal adjustment includes: removing redundant spaces and adjusting punctuation marks.

8. The method for compressing prompt words of a large language model according to claim 1, characterized in that: The compression hint is evaluated, and the compression hint is iteratively optimized according to the evaluation result, specifically including: Inputting the compressed prompt into a large language model, and evaluating the generation results of the compressed prompt and the original prompt in the large language model; According to the evaluation results, identifying and recording abnormal nodes with problems, and adjusting the weights of the abnormal nodes on the global parse tree through a bidirectional path dependency propagation method; According to the adjusted node values, the global parse tree is pruned through a recursive algorithm to generate new compression hints, until there are no abnormalities in the evaluation results corresponding to the new compression hints, thus completing the iteration of the compression hints.

9. A prompt word compression device for a large language model, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the prompt word compression method for a large language model as described in any one of claims 1-8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: When the computer executable instructions are executed, the prompt word compression method for a large language model as described in any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Cue word layered optimization method for complex problem decomposition

    CN120234399A

  • TF-IDF and cross entropy-based cue word compression method and system

    CN120449893A

  • Hint word compression method and system based on TF-IDF and cross entropy

    CN120449893B

  • Data query method and device based on large language model and storage medium

    CN120508640A

  • Data query method, device and storage medium based on large language model

    CN120508640B