A Prompt Compression Method and System for Large Models on Edge Devices
By performing performance testing and Prompt division on the edge device large model, combined with the method of semantic correlation calculation, multiple rounds of compression of the input Prompt are solved, and the problems of semantic retention and inference delay in the existing technology are achieved, and efficient Prompt compression is achieved.
Patent Information
- Application Number
- CN202510409012.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The prior art performs poorly in semantic retention and leads to inference delays, or leads to high memory consumption and inference delays when semantics are retained.
By performing performance testing on the edge device large model, a regression model of input length and processing time is established to determine the maximum Prompt length; the input Prompt is divided into key paragraphs and non-critical paragraphs, the semantic correlation degree is calculated, and multiple rounds of Prompt compression are performed.
It achieves excellent performance in semantic retention, while limiting memory consumption and inference delay, reducing the workload of the compression process, and taking into account the inference accuracy and delay requirements.
Smart Images

Figure CN119918678B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of prompt compression, and particularly relates to a prompt compression method and system for large models of edge devices. Background Art
[0002] Technologies for Prompt compression are mainly divided into two categories: traditional natural language processing technologies and language model technologies. Among them, traditional natural language processing technologies usually include: (1) rule-based methods, which use predefined language rules and patterns to identify and remove redundant information, such as data dictionaries and stop word removal techniques; (2) extractive summarization, which generates a compressed version by selecting key paragraphs in the text. For example, sentences are sorted and compressed using term frequency-inverse document frequency (TF-IDF); (3) topic modeling, which uses Latent Dirichlet Allocation (LDA) to identify the main topics in the Prompt, thereby retaining relevant information and removing irrelevant content. Language model technologies rely on language models to capture the key parts in the text and achieve text compression, aiming to remove redundancy while retaining important information.
[0003] Most traditional natural language processing technologies have been phased out due to their poor performance in semantic retention.
[0004] Language model technologies are currently widely used. For example, Microsoft's LLMLingua uses the LLaMa2 7b model for Prompt compression, while some other cutting-edge works use FLAN-T5-XXL or GPT2-XL. However, due to the limitations of end-side computing power and available memory, using advanced language models for compression may not be suitable for real-time Prompt compression during inference. This is because these compression technologies consume additional computing resources during execution, resulting in an increase in inference latency, thereby offsetting the time savings brought by compression. Traditional semantic extraction or segment truncation can quickly shorten the Prompt length, but it will lose a large amount of key information and significantly increase the zero-shot perplexity of the model.
[0005] Therefore, the existing technologies have two main drawbacks: on the one hand, the technologies with faster compression speed perform poorly in semantic retention; on the other hand, although using language models for compression can retain most of the semantics, it will bring high memory consumption and inference latency. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a prompt compression method and system for large models of edge devices, aiming to solve the technical problems existing in the prior art mentioned in the background art.
[0007] The embodiments of the present invention are implemented as follows:
[0008] A method for compressing prompts for large models on edge devices, the method specifically includes the following steps:
[0009] Perform a performance test on the large model of the edge device, establish a regression model of the input length and the processing time, and determine the longest Prompt length based on a preset expected processing time;
[0010] Receive the input Prompt, divide the input Prompt into multiple key paragraphs and multiple non-key paragraphs, and calculate the semantic correlation degree between the multiple non-key paragraphs and the multiple key paragraphs;
[0011] Calculate the total key length of the multiple key paragraphs, compare the total key length with the longest Prompt length, and perform multiple rounds of Prompt compression on the input Prompt.
[0012] As a further limitation of the technical solution of the embodiment of the present invention, the step of performing a performance test on the large model of the edge device, establishing a regression model of the input length and the processing time, and determining the longest Prompt length specifically includes the following steps:
[0013] Perform a performance test on the large model of the edge device to obtain performance test data;
[0014] According to the performance test data, select random forest regression, and establish a regression model of the input length and the processing time through cross-validation and parameter tuning;
[0015] Import the preset expected processing time into the regression model;
[0016] Export the longest Prompt length.
[0017] As a further limitation of the technical solution of the embodiment of the present invention, the step of receiving the input Prompt, dividing the input Prompt into multiple key paragraphs and multiple non-key paragraphs, and calculating the semantic correlation degree between the multiple non-key paragraphs and the multiple key paragraphs specifically includes the following steps:
[0018] Receive the input Prompt;
[0019] Perform text preprocessing and sentence similarity analysis on the input Prompt, and divide the input Prompt into multiple topic paragraphs;
[0020] Calculate the attention scores of the multiple topic paragraphs, and divide the multiple topic paragraphs into multiple key paragraphs and multiple non-key paragraphs;
[0021] Obtain a set of non-critical paragraphs and the average vector of critical paragraphs, and calculate the semantic correlation degree between multiple non-critical paragraphs and multiple critical paragraphs.
[0022] As a further limitation of the technical solution of the embodiment of the present invention, the text preprocessing and sentence similarity analysis of the input Prompt, and dividing the input Prompt into multiple topic paragraphs specifically include the following steps:
[0023] Perform text preprocessing on the input Prompt to eliminate segmentation noise, and divide the input Prompt into multiple sentences;
[0024] Convert multiple sentences into fixed-length vectors;
[0025] Calculate the cosine similarity between multiple sentences according to the multiple fixed-length vectors;
[0026] Compare multiple cosine similarities with a preset similarity threshold, and record the comparison results;
[0027] Aggregate multiple sentences according to the comparison results to obtain multiple topic paragraphs.
[0028] As a further limitation of the technical solution of the embodiment of the present invention, the expression of the fixed-length vector is:
[0029] ;
[0030] Where represents the th sentence, is the total number of sentences;
[0031] The calculation formula for the cosine similarity between multiple sentences is:
[0032] ;
[0033] Where represents the inner product of vectors, and represent the two-norm of vectors.
[0034] As a further limitation of the technical solution of the embodiment of the present invention, the calculation of the attention scores of multiple topic paragraphs, and dividing multiple topic paragraphs into multiple critical paragraphs and multiple non-critical paragraphs specifically include the following steps:
[0035] Obtain the attention weights of each attention head in each layer through a hook function;
[0036] Calculate the global attention scores of multiple topic paragraphs;
[0037] Normalize the multiple global attention scores to obtain multiple attention scores;
[0038] Divide the multiple topic paragraphs into multiple key paragraphs and multiple non - key paragraphs according to the multiple attention scores.
[0039] As a further limitation of the technical solution of the embodiment of the present invention, the calculation formula of the attention weight is:
[0040] ;
[0041] Wherein, represents the layer's attention head, and are the matrices of query and key respectively, is the dimension of the key;
[0042] The calculation formula of the global attention score is:
[0043] ;
[0044] Wherein, represents the th topic paragraph, and the topic paragraph contains the token set , is the attention matrix for the position ;
[0045] The calculation formula of the normalization process is:
[0046] ;
[0047] Wherein, is the attention score of the topic paragraph .
[0048] As a further limitation of the technical solution of the embodiment of the present invention, the set of non - key paragraphs is expressed as:
[0049] ;
[0050] Wherein, is the set of all topic paragraphs;
[0051] The average vector of the key paragraphs is expressed as:
[0052] ;
[0053] Wherein, A set of key paragraphs, is the embedded vector representation of;
[0054] The calculation formula of the semantic correlation degree is:
[0055] ;
[0056] wherein, is the th non - key paragraph the embedded vector representation of.
[0057] As a further limitation of the technical solution of the embodiment of the present invention, calculating the total key length of the multiple key paragraphs, comparing the total key length with the longest Prompt length, and performing multi - round Prompt compression on the input Prompt specifically includes the following steps:
[0058] Calculating the total key length of the multiple key paragraphs;
[0059] Comparing the total key length with the longest Prompt length and calculating the allocated total length;
[0060] If the allocated total length is not greater than 0, retain the multiple key paragraphs and compress the multiple non - key paragraphs;
[0061] If the allocated total length is greater than 0, calculate the total non - key length of the multiple non - key paragraphs;
[0062] Comparing the total non - key length with the allocated total length;
[0063] If the total non - key length is not less than 0, retain the multiple non - key paragraphs;
[0064] If the total non - key length is less than 0, perform selective compression on the multiple non - key paragraphs according to the multiple semantic correlation degrees.
[0065] A prompt compression system for large models of edge devices, the system includes a longest Prompt length determination module, a key paragraph recognition and processing module, and a multi - round Prompt compression module, wherein:
[0066] The longest Prompt length determination module is used to perform performance testing on the large model of the edge device, establish a regression model between the input length and the processing time, and determine the longest Prompt length based on a preset expected processing time;
[0067] The key paragraph recognition and processing module is used to receive the input Prompt, divide the input Prompt into multiple key paragraphs and multiple non-key paragraphs, and calculate the semantic correlation degree between the multiple non-key paragraphs and the multiple key paragraphs;
[0068] The multi-round Prompt compression module is used to calculate the total key length of the multiple key paragraphs, compare the total key length with the longest Prompt length, and perform multi-round Prompt compression on the input Prompt.
[0069] Compared with the prior art, the beneficial effects of the present invention are:
[0070] In the embodiment of the present invention, by performing performance testing on the large model of the edge device, a regression model of the input length and the processing time is established, and based on the preset expected processing time, the longest Prompt length is determined; the input Prompt is divided into multiple key paragraphs and multiple non-key paragraphs, and the semantic correlation degree between the multiple non-key paragraphs and the multiple key paragraphs is calculated; the total key length of the multiple key paragraphs is calculated, the total key length is compared with the longest Prompt length, and multi-round Prompt compression is performed on the input Prompt. It can reduce the workload in the compression process through the key paragraph-oriented compression method, balance the inference accuracy and latency requirements, not only perform excellently in semantic retention, but also limit the memory consumption and inference latency, and effectively support the subsequent inference work. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 The flowchart of the prompt word compression method for the large model of the edge device provided by the embodiment of the present invention is shown;
[0072] Figure 2 The application architecture diagram of the prompt word compression system for the large model of the edge device provided by the embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0074] It can be understood that in the prior art, there are two major categories of Prompt compression techniques: traditional natural language processing techniques, which have mostly been phased out due to poor performance in semantic retention; and language model techniques, which, due to limitations in edge-side computing power and available memory, require additional computing resources during execution, resulting in increased inference latency, thus offsetting the time savings brought about by compression. Traditional semantic extraction or segment truncation, although able to quickly shorten the Prompt length, will lose a large amount of key information and significantly increase the zero-shot perplexity of the model. Therefore, the prior art has two main drawbacks: on the one hand, techniques with faster compression speed perform poorly in semantic retention; on the other hand, using a language model for compression can retain most of the semantics but will incur high memory consumption and inference latency.
[0075] To solve the above problems, a method and system for Prompt compression for large models on edge devices disclosed in an embodiment of the present invention, by performing performance testing on the large model of the edge device, establishing a regression model between the input length and the processing time, and determining the longest Prompt length based on a preset expected processing time; receiving the input Prompt, dividing the input Prompt into multiple key paragraphs and multiple non-key paragraphs, and calculating the semantic correlation degree between the multiple non-key paragraphs and the multiple key paragraphs; calculating the total key length of the multiple key paragraphs, comparing the total key length with the longest Prompt length, and performing multiple rounds of Prompt compression on the input Prompt. It can reduce the workload of the compression process through a key-paragraph-oriented compression method, taking into account both inference accuracy and latency requirements, being able to perform excellently in semantic retention and at the same time limiting memory consumption and inference latency, effectively supporting subsequent inference work.
[0076] Specifically, Figure 1 shows a flowchart of the method for Prompt compression for large models on edge devices provided by an embodiment of the present invention.
[0077] In a preferred embodiment provided by the present invention, a method for Prompt compression for large models on edge devices, the method specifically includes the following steps:
[0078] Step S101, perform performance testing on the large model of the edge device, establish a regression model between the input length and the processing time, and determine the longest Prompt length based on a preset expected processing time.
[0079] In the embodiments of the present invention, by comprehensively testing and evaluating the performance of the large model of the edge device on the target edge device, including but not limited to key performance indicators such as the response time, memory occupancy, and energy consumption of the model under different Prompt lengths, performance test data is obtained. Based on the performance test data, random forest regression is selected, and through cross-validation and parameter tuning, a regression model of the input length and processing time is established. By importing the preset expected processing time into the regression model for processing, the longest Prompt length is derived.
[0080] It can be understood that through systematic performance testing, the operating limit of the large model of the edge device on the target edge device can be clarified, thereby providing data support for the formulation of the Prompt compression strategy.
[0081] It can be understood that in terms of model selection, random forest regression is selected to improve the accuracy and robustness of prediction; through cross-validation and parameter tuning, the generalization ability of the regression model on different data sets is ensured.
[0082] It can be understood that the longest Prompt length is not a constant, and it will vary depending on the computing power of the end-side device and the deployed model, and this value needs to be dynamically determined before inference. Specifically, by designing an adaptive mechanism, according to the current operating state and load conditions of the device, the longest Prompt length is adjusted in real time to adapt to different application requirements and environmental changes.
[0083] Step S102: Receive the input Prompt, divide the input Prompt into multiple key paragraphs and multiple non-key paragraphs, and calculate the semantic association degree between the multiple non-key paragraphs and the multiple key paragraphs.
[0084] In the embodiments of the present invention, by receiving the input Prompt, performing text preprocessing to eliminate the influence of segmentation noise on the input Prompt, dividing the input Prompt into multiple sentences, then converting the multiple sentences into fixed-length vectors, and further calculating the cosine similarity between the multiple sentences according to the multiple fixed-length vectors. By comparing the multiple cosine similarities with a preset similarity threshold, recording the comparison results, and then aggregating the multiple sentences according to the comparison results to obtain multiple topic paragraphs. After that, the attention weights of each attention head in each layer are obtained through a hook function, the global attention scores of the multiple topic paragraphs are calculated, and then the multiple global attention scores are normalized to obtain multiple attention scores. Furthermore, according to a preset score threshold, the multiple attention scores are compared, the multiple topic paragraphs are divided into multiple key paragraphs and multiple non-key paragraphs, and then the set of non-key paragraphs and the average vector of the key paragraphs are obtained, and the semantic association degree between the multiple non-key paragraphs and the multiple key paragraphs is calculated. Specifically, the expression of the fixed-length vector is:
[0085] ;
[0086] Among them, represents the th sentence, is the total number of sentences;
[0087] The calculation formula for the cosine similarity between multiple sentences is:
[0088] ;
[0089] Among them, represents the inner product of vectors, and represent the two-norm of vectors;
[0090] The calculation formula for the attention weight is:
[0091] ;
[0092] Among them, represents the th layer of attention head, and are the matrices of query and key respectively, is the dimension of the key;
[0093] The calculation formula for the global attention score is:
[0094] ;
[0095] Among them, represents the th topic paragraph, and the topic paragraph contains the token set , is the attention matrix for the position ;
[0096] The calculation formula for normalization processing is:
[0097] ;
[0098] Among them, is the attention score of the topic paragraph ;
[0099] The set representation of non-critical paragraphs is:
[0100] ;
[0101] Among them, is the set of all topic paragraphs, is the score threshold;
[0102] The average vector representation of the key paragraphs is:
[0103] ;
[0104] where, is the set of key paragraphs, is the embedding vector representation of;
[0105] The calculation formula for the semantic correlation degree is:
[0106] ;
[0107] where, is the th non-key paragraph 's embedding vector representation.
[0108] It can be understood that the text preprocessing of eliminating the influencing segmentation noise for the input Prompt is to remove the meaningless symbols such as extra spaces in the input Prompt, but punctuation marks, modal particles and emojis need to be retained to reflect the user's emotions and the structural distribution of the Prompt.
[0109] It can be understood that the final value range of the cosine similarity will be [-1, 1]. The closer the value is to 1, the higher the similarity, and it is suitable to be grouped into the same paragraph. By setting the similarity threshold , the standard for controlling sentence aggregation is ensured to make the paragraph division reasonable.
[0110] It can be understood that the difficulty of identifying key paragraphs depends on the quality of the Prompt. High-quality Prompts usually have "problem" paragraphs at the beginning and end, and only need to quickly locate the question sentences to identify the key paragraphs. For Prompts with complex content, the attention weights of the first few layers of the large language model are extracted to determine the key paragraphs, that is, the large language model selects the important paragraphs it deems by itself. The specific method is to obtain the attention matrix of each attention head in each layer to represent the attention weight for the position .
[0111] It is understandable that by calculating the semantic correlation degrees between multiple non-critical paragraphs and multiple critical paragraphs, the semantic relevance between non-critical paragraphs and critical paragraphs can be quantified, thereby providing a basis for subsequent compression or deletion decisions. Non-critical paragraphs with high correlation degrees may contain important information related to critical paragraphs, while non-critical paragraphs with low correlation degrees can be preferentially considered for compression or deletion to minimize the loss of semantic information.
[0112] Step S103: Calculate the total key length of multiple said critical paragraphs, compare the total key length with the longest Prompt length, and perform multi-round Prompt compression on the input Prompt.
[0113] In the embodiments of the present invention, by calculating the total key length of multiple critical paragraphs, then comparing the total key length with the longest Prompt length, calculating the allocated total length, when the allocated total length is not greater than 0, multiple critical paragraphs are retained and multiple non-critical paragraphs are compressed; while when the allocated total length is greater than 0, calculate the total non-critical length of multiple non-critical paragraphs, by comparing the total non-critical length with the allocated total length, when the total non-critical length is not less than 0, multiple non-critical paragraphs are retained; while when the total non-critical length is less than 0, select and compress multiple non-critical paragraphs according to multiple semantic correlation degrees. Specifically, the calculation formula for the total key length is:
[0114] ;
[0115] where is the number of Tokens of paragraph ;
[0116] The calculation formula for the allocated total length is:
[0117] ;
[0118] where is the longest Prompt length;
[0119] The calculation formula for the total non-critical length is:
[0120] ;
[0121] where is the number of Tokens of paragraph ;
[0122] It is understandable that by performing selective compression on multiple non-critical paragraphs, non-critical paragraphs can be effectively compressed while ensuring the information of critical paragraphs, reducing the overall Prompt length. The calculation formula for the corresponding compression ratio is:
[0123] ;
[0124] It can be understood that if the longest Prompt length cannot be satisfied after processing non-critical paragraphs, , and it is difficult to delete non-critical paragraphs with high relevance, then it is necessary to consider compressing critical paragraphs. The calculation formula for the corresponding compression ratio is:
[0125] .
[0126] Furthermore, Figure 2 shows an application architecture diagram of the prompt compression system for the large model of edge devices provided by the embodiments of the present invention.
[0127] Among them, in another preferred embodiment provided by the present invention, a prompt compression system for the large model of edge devices includes:
[0128] The longest Prompt length determination module 101 is used to perform performance testing on the large model of edge devices, establish a regression model between the input length and the processing time, and determine the longest Prompt length based on a preset expected processing time.
[0129] In the embodiments of the present invention, the longest Prompt length determination module 101 comprehensively tests and evaluates the performance of the large model of edge devices on the target edge device, including but not limited to key performance indicators such as the response time, memory occupancy, and energy consumption of the model under different Prompt lengths, obtains performance test data, selects random forest regression based on the performance test data, establishes a regression model between the input length and the processing time through cross-validation and parameter tuning, and exports the longest Prompt length by importing the preset expected processing time into the regression model for processing.
[0130] The critical paragraph identification and processing module 102 is used to receive the input Prompt, divide the input Prompt into multiple critical paragraphs and multiple non-critical paragraphs, and calculate the semantic relevance between the multiple non-critical paragraphs and the multiple critical paragraphs.
[0131] In an embodiment of the present invention, the key paragraph recognition processing module 102 performs text preprocessing on the input Prompt to eliminate influencing segmentation noise by receiving the input Prompt, divides the input Prompt into multiple sentences, then converts the multiple sentences into fixed-length vectors, and further calculates the cosine similarity between the multiple sentences according to the multiple fixed-length vectors. By comparing the multiple cosine similarities with a preset similarity threshold, recording the comparison results, and then aggregating the multiple sentences according to the comparison results, multiple topic paragraphs are obtained. After that, the attention weights of each attention head in each layer are obtained through a hook function, the global attention scores of the multiple topic paragraphs are calculated, and then the multiple global attention scores are normalized to obtain multiple attention scores. Further, according to a preset score threshold, the multiple attention scores are compared, and the multiple topic paragraphs are divided into multiple key paragraphs and multiple non-key paragraphs. Then, the set of non-key paragraphs and the average vector of the key paragraphs are obtained, and the semantic correlation degree between the multiple non-key paragraphs and the multiple key paragraphs is calculated. Specifically, the expression of the fixed-length vector is:
[0132] ;
[0133] Among them, represents the th sentence, is the total number of sentences;
[0134] The calculation formula for the cosine similarity between multiple sentences is:
[0135] ;
[0136] Among them, represents the inner product of vectors, and represent the two-norm of vectors;
[0137] The calculation formula for the attention weight is:
[0138] ;
[0139] Among them, represents layer's attention head, and are the matrices of query and key respectively, is the dimension of the key;
[0140] The calculation formula for the global attention score is:
[0141] ;
[0142] Among them, represents the A topic paragraph, the topic paragraph contains a set of tokens , is the attention matrix for the position ;
[0143] The calculation formula for the normalization process is:
[0144] ;
[0145] where, is the attention score of the topic paragraph ;
[0146] The set of non-critical paragraphs is represented as:
[0147] ;
[0148] where, is the set of all topic paragraphs, is the score threshold;
[0149] The average vector of the critical paragraphs is represented as:
[0150] ;
[0151] where, is the set of critical paragraphs, is the embedding vector representation of;
[0152] The calculation formula for the semantic correlation degree is:
[0153] ;
[0154] where, is the th non-critical paragraph 's embedding vector representation.
[0155] The multi-round Prompt compression module 103 is used to calculate the total critical length of multiple said critical paragraphs, compare the total critical length with the longest Prompt length, and perform multi-round Prompt compression on the input Prompt.
[0156] In the embodiments of the present invention, the multi-round Prompt compression module 103 calculates the total key length of multiple key paragraphs, then compares the total key length with the longest Prompt length to calculate the allocated total length. When the allocated total length is not greater than 0, multiple key paragraphs are retained and multiple non-key paragraphs are compressed; when the allocated total length is greater than 0, the total non-key length of multiple non-key paragraphs is calculated, and by comparing the total non-key length with the allocated total length, when the total non-key length is not less than 0, multiple non-key paragraphs are retained; when the total non-key length is less than 0, multiple non-key paragraphs are selectively compressed according to multiple semantic relevance degrees. Specifically, the calculation formula for the total key length is:
[0157] ;
[0158] where is the number of Tokens of paragraph .
[0159] The calculation formula for the allocated total length is:
[0160] ;
[0161] where is the longest Prompt length;
[0162] The calculation formula for the total non-key length is:
[0163] ;
[0164] where is the number of Tokens of paragraph .
[0165] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0166] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0167] The above embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A prompt word compression method for a large model of an edge device, characterized in that: The method specifically comprises the following steps: Perform performance tests on large models of edge devices, build a regression model of input length and processing time, and determine the maximum prompt length based on the preset expected processing time; Receive an input prompt, divide the input prompt into a plurality of key paragraphs and a plurality of non-key paragraphs, and calculate semantic associations between the plurality of non-key paragraphs and the plurality of key paragraphs; Calculating the total key lengths of the plurality of key paragraphs, comparing the total key length with the longest Prompt length, and performing multiple rounds of Prompt compression on the input Prompt; The calculating of the total key lengths of the plurality of key paragraphs, comparing the total key length with the longest Prompt length, and performing multiple rounds of Prompt compression on the input Prompt specifically comprises the following steps: Calculate the total key length of the plurality of key paragraphs; Compare the key total length with the longest prompt length to calculate the total allocation length; If the total length of the allocation is not greater than 0, retaining multiple key sections and compressing multiple non-key sections; If the total allocated length is greater than 0, then calculating the total non-critical length of the plurality of non-critical sections; comparing the non-critical total length to the allocated total length; If the total length of the non-critical paragraphs is not less than 0, multiple non-critical paragraphs are retained; If the total length of the non-critical paragraphs is less than 0, multiple non-critical paragraphs are selectively compressed according to multiple semantic associations.
2. The prompt word compression method for edge device large model according to claim 1, characterized in that: The performance test of the edge device large model is performed to establish a regression model of input length and processing time, and based on the preset expected processing time, the longest prompt length is determined, specifically including the following steps: Perform performance tests on large models of edge devices and obtain performance test data; According to the performance test data, random forest regression is selected, and a regression model of input length and processing time is established through cross-validation and parameter tuning; Importing the preset expected processing time into the regression model; Export the maximum prompt length.
3. The prompt word compression method for edge device large model according to claim 1, characterized in that: The receiving of the input Prompt, dividing the input Prompt into a plurality of key paragraphs and a plurality of non-key paragraphs, and calculating the semantic association between the plurality of non-key paragraphs and the plurality of key paragraphs specifically comprises the following steps: Receive input prompt; Performing text preprocessing and sentence similarity analysis on the input prompt, and dividing the input prompt into multiple topic paragraphs; Calculating the attention scores of the plurality of topic paragraphs, and dividing the plurality of topic paragraphs into a plurality of key paragraphs and a plurality of non-key paragraphs; A set of non-key paragraphs and an average vector of key paragraphs are obtained, and semantic associations between a plurality of the non-key paragraphs and a plurality of the key paragraphs are calculated.
4. The prompt word compression method for edge device large model according to claim 3 is characterized in that: The performing text preprocessing and sentence similarity analysis on the input prompt and dividing the input prompt into a plurality of topic paragraphs specifically comprises the following steps: Performing text preprocessing on the input prompt to eliminate noise that affects segmentation, and dividing the input prompt into multiple sentences; Converting the plurality of sentences into fixed-length vectors; Calculating the cosine similarity between the plurality of sentences according to the plurality of fixed-length vectors; Compare the multiple cosine similarities with a preset similarity threshold, and record the comparison result; According to the comparison result, multiple sentences are aggregated to obtain multiple topic paragraphs.
5. The prompt word compression method for edge device large model according to claim 4 is characterized in that: The expression of the fixed-length vector is: ; in, Representative Sentences, is the total number of sentences; The calculation formula for calculating the cosine similarity between the multiple sentences is: ; in, represents the inner product of vectors, and Represents the bi-norm of a vector.
6. The prompt word compression method for edge device large model according to claim 3, characterized in that: The step of calculating the attention scores of the plurality of topic paragraphs and dividing the plurality of topic paragraphs into a plurality of key paragraphs and a plurality of non-key paragraphs specifically comprises the following steps: Get the attention weight of each attention head in each layer through the hook function; Calculating global attention scores for a plurality of the topic paragraphs; Normalizing the multiple global attention scores to obtain multiple attention scores; According to the plurality of attention scores, the plurality of topic paragraphs are divided into a plurality of key paragraphs and a plurality of non-key paragraphs.
7. The prompt word compression method for edge device large model according to claim 6, characterized in that: The calculation formula of the attention weight is: ; in, represent Layer Attention head, and are the matrices of query and key respectively, is the dimension of the key; The calculation formula of the global attention score is: ; in, Representative topic paragraphs, topic paragraphs Contains a token collection , for About Location The attention matrix; The calculation formula for the normalization process is: ; in, Topic paragraph Attention score.
8. The prompt word compression method for edge device large model according to claim 7, characterized in that: The set of non-critical paragraphs is expressed as: ; in, is the collection of all thematic paragraphs; is the score threshold; The average vector of the key paragraph is expressed as: ; in, is a collection of key paragraphs, for Embedded vector representation of ; The calculation formula of the semantic relevance is: ; in, For the Non-critical paragraphs The embedded vector representation of .
9. A prompt word compression system for a large model of edge devices, characterized in that: The system includes a longest prompt length determination module, a key paragraph identification processing module and a multi-round prompt compression module, wherein: The longest prompt length determination module is used to perform performance testing on the edge device large model, establish a regression model of input length and processing time, and determine the longest prompt length based on the preset expected processing time; A key paragraph identification processing module, used for receiving an input prompt, dividing the input prompt into a plurality of key paragraphs and a plurality of non-key paragraphs, and calculating the semantic association between the plurality of non-key paragraphs and the plurality of key paragraphs; A multi-round prompt compression module, used for calculating the total key length of a plurality of the key paragraphs, comparing the total key length with the longest prompt length, and performing multi-round prompt compression on the input prompt; The calculating of the total key lengths of the plurality of key paragraphs, comparing the total key length with the longest Prompt length, and performing multiple rounds of Prompt compression on the input Prompt are specifically as follows: Calculate the total key length of the plurality of key paragraphs; Compare the key total length with the longest prompt length to calculate the total allocation length; If the total length of the allocation is not greater than 0, retaining multiple key sections and compressing multiple non-key sections; If the total allocated length is greater than 0, then calculating the total non-critical length of the plurality of non-critical sections; comparing the non-critical total length to the allocated total length; If the total length of the non-critical paragraphs is not less than 0, multiple non-critical paragraphs are retained; If the total length of the non-critical paragraphs is less than 0, multiple non-critical paragraphs are selectively compressed according to multiple semantic associations.
Citation Information
Patent Citations
Chinese cue word compression method and device
CN117725036A
Constructing Prompt Information for Submission to a Language Model by Dynamically Compressing Source Information
US20240394479A1