Text optimization method and system for dynamic de-duplication of cue words in big language model dialogue

Through the dynamic deduplication mechanism of ring buffers and composite weight models, redundant prompt words in large language model dialogue are optimized, redundant occupation and information loss problems are solved, and dialogue processing efficiency and resource utilization are improved.

CN120256578APending Publication Date: 2025-07-04FUJIAN TQ ONLINE INTERACTIVE INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510393362.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing large language model has redundant prompt words in continuous dialogue scenarios that occupy a large amount of context token capacity, resulting in limited effective dialogue rounds. The redundant prompt words interfere with the model's attention allocation and cannot adapt to the semantic evolution needs of multiple rounds of dialogues. A simple truncation strategy will cause the loss of key conversation information.

Method used

The ring buffer is used to dynamically maintain dialogue records, and repetitive long text is compressed through the progressive prefix matching algorithm. The quantitative index system is built with the composite weight model to achieve dynamic deduplication, give priority to retaining important prompt words and eliminating old data to prevent semantic faults.

Benefits of technology

Effectively reduce the occupation of invalid tokens, improve the efficiency of model dialogue processing, improve the retention rate of key information, reduce API call costs, and ensure that the model focuses on efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256578A_ABST
    Figure CN120256578A_ABST
Patent Text Reader

Abstract

The invention relates to the field of large language model dialogue dynamic optimization, in particular to a text optimization method and system for dynamic de-duplication of cue words in large language model dialogues, and the method comprises the following steps: dynamically maintaining dialogue records by adopting an annular buffer area, preferentially retaining cue words and eliminating old data; compressing a repeated long text of the dialogue record through a progressive prefix matching algorithm to generate a mark format capable of being reversely analyzed, and meanwhile, establishing a compressed content self-checking mechanism to prevent semantic fault; according to the method, content value evaluation and deduplication are realized based on a composite weight model, a quantitative index system of three dimensions of time, frequency and semantics is constructed, low-weight redundant information in dialogue record content is automatically filtered, context length reduction in large language model dialogues can be realized, the retention rate of key information is increased, and the method is suitable for large language model dialogues. The big language model dialogue processing efficiency and the resource utilization rate can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of dynamic optimization of large language model conversations, in particular to a method and system for dynamically removing duplicate text in prompts during large language model conversations. Background Art

[0002] In existing large language model application systems in continuous conversation scenarios, the context splicing method of full historical conversation + fixed prompts is usually adopted. Its defects are as follows: 1) Repeated fixed prompts occupy 30%-50% of the context token capacity, resulting in limited effective conversation turns; 2) Redundant prompts interfere with the model's attention allocation and reduce the generation quality; 3) Fixed-position prompts lack a dynamic adjustment mechanism and cannot adapt to the semantic evolution requirements of multi-turn conversations. Current solutions mostly adopt simple truncation strategies, but this will cause the loss of key conversation information. Summary of the Invention

[0003] To overcome the problems of information redundancy, low conversation processing efficiency, and easy loss of key conversation information in large language models in the prior art, the purpose of the present invention is to provide a method and system for dynamically removing duplicate text in prompts during large language model conversations, realizing the reduction of context length, improving the retention rate of key information, and being able to improve the conversation processing efficiency and resource utilization rate of large language models.

[0004] The present invention is implemented by the following scheme:

[0005] A method for dynamically removing duplicate text in prompts during large language model conversations, the method steps are as follows:

[0006] Step 1: Dynamically maintain the conversation record using a circular buffer, giving priority to retaining prompts and eliminating stale data;

[0007] Step 2: Compress the repeated long text in the conversation record through a progressive prefix matching algorithm to generate a reversibly parsable token format, and at the same time establish a self-check mechanism for the compressed content to prevent semantic breaks;

[0008] Step 3: Based on a composite weight model, realize content value evaluation and duplicate removal, construct a quantitative index system in three dimensions of time, frequency, and semantics, and automatically filter out low-weight redundant information in the conversation record content.

[0009] Further, step 1 is further specified as follows: Maintain the conversation history using a dynamic queue structure, and balance information retention and resource consumption through a capacity control strategy; Define a structure to encapsulate metadata, where the metadata includes the conversation role, content, timestamp, and occurrence frequency, and initialize a circular buffer with a capacity of x + 2; When adding a new entry, if the current context quantity exceeds x, the first x prompt entries are preferentially retained, and subsequent old data is eliminated according to the first-in, first-out principle; In terms of thread safety, concurrent control is achieved through mutex locks, and at the same time, a global frequency mapping table freqMap is maintained. The global frequency mapping table freqMap is used to statistically count the occurrence times of each compressed text in real time, providing data support for subsequent filtering.

[0010] Further, step 2 is further specified as follows: Scan the context content item by item, and execute a three-level compression strategy for non-fixed prompt entries: First, convert the text into a rune array, and then start from the longest substring in the rune array to initiate a matching query to the prefix cache; If a historical repeated prefix is found, a differential compression mark is generated according to the length. If the text length ≤ 4 characters, the complete prefix is retained; If the text length ≥ 4 characters, the first and last 2 characters are extracted and merged to form a compression mark, and the unmatched text is added to the cache library as a whole. After compression, reverse verification is performed. By traversing the context, it is detected whether the associated original text of the compression mark exists. When it is found that the original text corresponding to the mark is no longer in the current context, content recovery is automatically triggered to prevent the information chain from breaking.

[0011] Further, step 3 is further specified as follows: The weight calculation method CalculateWeight of the composite weight model has the following formula:

[0012] Weight = Time decay factor × 0.4 + Frequency factor × 0.3 + Keyword bonus × 0.3

[0013] Among them, the time decay factor is in seconds; The frequency factor Freq is taken from the global statistical table. The more times the content is repeated, the lower the weight of this item; The frequency factor Freq is taken from the global statistical table. The global statistical table is used to record the occurrence times, first / most recent occurrence timestamps, and cumulative occurrence times of all content in real time. The keyword bonus item gives a fixed gain of 0.3 to the content marked as "important", calculates the weight values of each entry in real time, retains the high-quality content with a weight > 0.65, and automatically eliminates the remaining low-value information.

[0014] A system for dynamically removing duplicates and optimizing text in prompts during large language model conversations, the system includes: a dynamic storage management module, a text compression and fusion module, and a duplicate removal module;

[0015] The dynamic storage management module is used to dynamically maintain conversation records using a circular buffer, preferentially retain prompt words, and eliminate stale data;

[0016] The text compression and fusion module is used to compress the repeated long texts in the conversation record through a progressive prefix matching algorithm, generate a reversely resolvable token format, and establish a self-verification mechanism for the compressed content to prevent semantic breaks.

[0017] The deduplication module is used to implement content value evaluation and deduplication based on a composite weight model, construct a quantitative index system in three dimensions of time, frequency, and semantics, and automatically filter out low-weight redundant information in the conversation record content.

[0018] Further, the dynamic storage management module is specifically as follows: maintaining the conversation history using a dynamic queue structure, balancing information retention and resource consumption through a capacity control strategy; defining a structure to encapsulate metadata, where the metadata includes conversation roles, content, timestamps, and occurrence frequencies, and initializing a circular buffer with a capacity of x + 2; when adding a new entry, if the current number of contexts exceeds x, the first x prompt entries are preferentially retained, and subsequent old data is eliminated according to the first-in, first-out principle; in terms of thread safety, concurrent control is achieved through mutex locks, and at the same time, a global frequency mapping table freqMap is maintained, and the global frequency mapping table freqMap is used to count the occurrence times of each compressed text in real time, providing data support for subsequent filtering.

[0019] Further, the text compression and fusion module is specifically as follows: scanning the context content item by item, and executing a three-level compression strategy for non-fixed prompt entries: first, convert the text into a rune array, and then start from the longest substring in the rune array to initiate a matching query for the prefix cache; if a historical repeated prefix is found, generate a differential compression token according to the length. If the text length ≤ 4 characters, the complete prefix is retained; if the text length ≥ 4 characters, the first and last two characters are extracted and merged to form a compression token, and the unmatched text is added to the cache library as a whole. After compression, reverse verification is performed, and whether the associated original text of the compression token exists is detected by traversing the context. When it is found that the original text corresponding to the token is no longer in the current context, content recovery is automatically triggered to prevent the information chain from breaking.

[0020] Further, the deduplication module is specifically as follows: The weight calculation method CalculateWeight of the composite weight model has the following formula:

[0021] Weight = time decay factor × 0.4 + frequency factor × 0.3 + keyword bonus × 0.3

[0022] Among them, the time decay factor is in seconds; the frequency factor Freq is taken from the global statistical table, and the higher the content repetition times, the lower the weight of this item; the frequency factor Freq is taken from the global statistical table, and the global statistical table is used to record the occurrence times, first / most recent occurrence timestamps, and cumulative occurrence times of all content in real time. The keyword addition item gives a fixed gain of 0.3 to the content marked as "important", calculates the weight values of each item in real time, retains the high-quality content with a weight > 0.65, and automatically eliminates the remaining low-value information.

[0023] The beneficial effects of the present invention are as follows:

[0024] This patent provides a method and system for optimizing prompt dynamic deduplication text in large language model conversations. This method reduces the occupation of invalid tokens through dynamic context queue management, improves the model's attention efficiency based on a weight scoring-based prompt screening mechanism, and reduces the repetition rate while ensuring that the core prompts are not lost through a hierarchical prompt injection architecture, and at the same time reduces the API call cost. Brief Description of the Drawings

[0025] Figure 1 is a flowchart of the method of the present invention;

[0026] Figure 2 is a structural block diagram of the system of the present invention. Detailed Embodiments

[0027] The present invention will be further described below with reference to the accompanying drawings.

[0028] See Figure 1 , a method for optimizing prompt dynamic deduplication text in large language model conversations, and the method steps are as follows:

[0029] Step 1: Use a circular buffer to dynamically maintain the conversation record, preferentially retain the prompts and eliminate the stale data;

[0030] Step 2: Compress the repeated long texts in the conversation record through a progressive prefix matching algorithm, generate a reversibly parsable markup format, and establish a self-verification mechanism for the compressed content to prevent semantic breaks;

[0031] Step 3: Based on a composite weight model, implement content value evaluation and deduplication, construct a quantitative index system in three dimensions of time, frequency, and semantics, and automatically filter out the low-weight redundant information in the conversation record content.

[0032] The present invention will be further described below with reference to a specific embodiment:

[0033] A method for optimizing prompt dynamic deduplication text in large language model conversations, and the method includes the following steps:

[0034] Step 1: Maintain the conversation history using a dynamic queue structure, and balance information retention and resource consumption through a capacity control strategy; define a structure to encapsulate metadata, where the metadata includes conversation roles, content, timestamps, occurrence frequencies, etc., and initialize a circular buffer with a capacity of "MaxContextLength + 2"; when adding a new entry, if the current context quantity exceeds MaxContextLength, prioritize retaining the first MaxContextLength prompt entries, and subsequent old data is eliminated according to the first-in, first-out principle; for thread safety, implement concurrent control through mutex locks to ensure data consistency under multi-coroutine operations; at the same time, maintain a global frequency mapping table freqMap, which is used to statistically count the occurrence times of each compressed text in real time to provide data support for subsequent filtering.

[0035] The core code is as follows:

[0036]

[0037]

[0038]

[0039] Step 2: The text compression engine adopts a progressive prefix matching algorithm to establish a two-way mapping relationship between compression tags and the original content;

[0040] Scan the context content item by item. For non-fixed prompt word entries (fixed prompt words refer to the content that needs to be carried in each conversation, and non-fixed prompt words refer to random questions), execute a three-level compression strategy: First, convert the text into a rune array to avoid Chinese character truncation, and then start from the longest substring in the rune array to initiate a matching query to the prefix cache (prefixCache); if a historical repeated prefix is found, generate a differentiated compression tag according to the length. If the text length ≤ 4 characters, retain the complete prefix; if the text length ≥ 4 characters, extract the first and last 2 characters each and combine them to form a compression tag. The unmatched text is added to the cache library as a whole. After compression, perform a reverse verification by traversing the context to detect whether the associated original text of the compression tag exists. When it is found that the original text corresponding to the tag is no longer in the current context, the content recovery is automatically triggered to prevent the information chain from breaking.

[0041] For example, questions with multiple repeated types are as follows:

[0042] [assistant]

Large Language Model

[0043] [user]Please explain the following poem: "Do you not see the Yellow River come from the sky, rushing to the sea and never returning? Do you not see the grief of white hair in the bright mirror of the high hall, turning from black as silk in the morning to white as snow at dusk? When life is going well, one should enjoy to the fullest; do not let the golden goblet face the moon in vain. Since heaven has given me talent, it must be useful; though a thousand pieces of gold are spent, they will come back again."

[0044] [user]Please re - explain the poem "Do you not see the Yellow River come from the sky... they will come back again."

[0045] [user]What poems are similar to "Do you not see the Yellow River come from the sky... they will come back again."?

[0046] [user]What are the differences between the styles of Li Bai and Du Fu?

[0047] After compression, it becomes:

[0048] [assistant][Large Language Model] You are a professional assistant, please answer rigorously. The content with [xx...xx] in the user's question is the compressed content, which has appeared in the context.

[0049] [user]Please explain the following poem: "Do you not see the Yellow River come from the sky, rushing to the sea and never returning? Do you not see the grief of white hair in the bright mirror of the high hall, turning from black as silk in the morning to white as snow at dusk? When life is going well, one should enjoy to the fullest; do not let the golden goblet face the moon in vain. Since heaven has given me talent, it must be useful; though a thousand pieces of gold are spent, they will come back again"

[0050] [user]Please re - explain the poem [Do you not... they will come back again]

[0051] [user]What poems are similar to [Do you not... they will come back again]?

[0052] [user]What are the differences between the styles of Li Bai and Du Fu?

[0053] Core code:

[0054]

[0055]

[0056]

[0057]

[0058] Step 3: Implement a multi - factor dynamic deduplication mechanism as follows:

[0059] The calculation formula for calculating weights in the composite weight model CalculateWeight is as follows:

[0060] Weight = Time decay factor × 0.4 + Frequency factor × 0.3 + Keyword bonus × 0.3;

[0061] Among them, the time decay factor is in seconds, and the half-life design of 900 seconds (15 minutes) makes the weight of old content decrease linearly; the frequency factor Freq is taken from the global statistical table, and the more times the content is repeated, the lower the weight of this item; the frequency factor Freq is taken from the global statistical table, and the global statistical table is used to record the occurrence times, first / latest occurrence timestamps, and cumulative occurrence times of all content in real time. The keyword bonus item gives a fixed gain of 0.3 to the content marked as "important", calculates the weight values of each entry in real time, retains the high-quality content with a weight > 0.65, and automatically eliminates the remaining low-value information.

[0062] Core code:

[0063]

[0064]

[0065] See Figure 2 , the system for optimizing the dynamic deduplication of prompt words in the large language model dialogue, the system includes: a dynamic storage management module, a text compression and fusion module, and a deduplication module;

[0066] The dynamic storage management module is used to dynamically maintain the dialogue records using a circular buffer, giving priority to retaining the prompt words and eliminating stale data;

[0067] The text compression and fusion module is used to compress the repeated long texts in the dialogue records through a progressive prefix matching algorithm, generate a mark format that can be reversely parsed, and establish a self-check mechanism for the compressed content to prevent semantic breaks;

[0068] The deduplication module is used to implement content value evaluation and deduplication based on the composite weight model, construct a quantitative index system in three dimensions of time, frequency, and semantics, and automatically filter out the low-weight redundant information in the dialogue record content.

[0069] In one embodiment of the present invention, the dynamic storage management module specifically is: maintaining the conversation history using a dynamic queue structure, and balancing information retention and resource consumption through a capacity control strategy; defining a structure to encapsulate metadata such as conversation roles, content, timestamps, and occurrence frequencies, and initializing a circular buffer with a capacity of x + 2; when a new entry is added, if the current context quantity exceeds x, the first x prompt entries are preferentially retained, and subsequent old data is eliminated according to the first-in, first-out principle; in terms of thread safety, concurrent control is achieved through mutex locks, and at the same time, a global frequency mapping table freqMap is maintained. The global frequency mapping table freqMap is used to count the occurrence times of each compressed text in real time, providing data support for subsequent filtering.

[0070] In one embodiment of the present invention, the deduplication module specifically is: The weight calculation formula CalculateWeight of the composite weight model is:

[0071] Weight = time decay factor × 0.4 + frequency factor × 0.3 + keyword bonus × 0.3

[0072] Among them, the time decay factor is in seconds; the frequency factor Freq is taken from the global statistical table, and the more times the content is repeated, the lower the weight of this item; the frequency factor Freq is taken from the global statistical table, and the global statistical table is used to record the occurrence times, first / most recent occurrence timestamps, and cumulative occurrence times of all content in real time. The keyword bonus item gives a fixed gain of 0.3 to the content marked with "important", calculates the weight values of each entry in real time, retains the high-quality content with a weight > 0.65, and automatically eliminates the remaining low-value information.

[0073] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. A method for optimizing text by dynamically removing duplicates in prompts during large language model conversations, characterized in that, The method steps are as follows: Step 1: Dynamically maintain the conversation record using a circular buffer, giving priority to retaining the prompt words and eliminating stale data; Step 2: Compress the repeated long texts in the conversation record through a progressive prefix matching algorithm to generate a reversibly parsable token format, and at the same time establish a self-check mechanism for the compressed content to prevent semantic breaks; Step 3: Implement content value evaluation and deduplication based on a composite weight model, construct a quantitative index system in three dimensions of time, frequency, and semantics, and automatically filter out low-weight redundant information in the conversation record content.

2. The method for optimizing the dynamically deduplicated text of the prompt in the large language model dialogue according to claim 1, wherein, Step 1 is further specifically as follows: Use a dynamic queue structure to maintain the conversation history, and balance information retention and resource consumption through a capacity control strategy; Define a structure to encapsulate the metadata, and the metadata includes the conversation role, content, timestamp, and occurrence frequency. Initialize a circular buffer with a capacity of x + 2; When a new entry is added, if the current context quantity exceeds x, then give priority to retaining the first x prompt word entries, and the subsequent old data is eliminated according to the first-in, first-out principle; In terms of thread safety, concurrent control is achieved through mutex locks, and at the same time a global frequency mapping table freqMap is maintained. The global frequency mapping table freqMap is used to count the occurrence times of each compressed text in real time, providing data support for subsequent filtering.

3. The method for optimizing the dynamically deduplicated text of prompts in the large language model conversation according to claim 1, wherein, Step 2 is further specifically as follows: Scan the context content item by item, and execute a three-level compression strategy for non-fixed prompt word entries: First, convert the text into a rune array, and then start from the longest substring in the rune array to initiate a matching query to the prefix cache; If a historical repeated prefix is found, generate a differential compression token according to the length. If the text length ≤ 4 characters, retain the complete prefix; If the text length ≥ 4 characters, extract the first and last 2 characters each and merge them to form a compression token. The unmatched text is added to the cache library as a whole. After compression, perform a reverse check. Detect whether the associated original text of the compression token exists by traversing the context. When it is found that the original text corresponding to the token is no longer in the current context, content recovery is automatically triggered to prevent the information chain from breaking.

4. The method for optimizing the dynamic deduplication of prompt words in the large language model dialogue according to claim 1, wherein Step 3 is further specifically as follows: The weight calculation method CalculateWeight of the composite weight model has the following formula: Weight = Time decay factor × 0.4 + Frequency factor × 0.3 + Keyword bonus × 0.3 Among them, the time decay factor is in seconds; The frequency factor Freq is taken from the global statistical table. The more times the content is repeated, the lower the weight of this item; The frequency factor Freq is taken from the global statistical table. The global statistical table is used to record the occurrence times, first / latest occurrence timestamps, and cumulative occurrence times of all content in real time. The keyword bonus item gives a fixed gain of 0.3 to the content marked with "important". Calculate the weight value of each item in real time, retain the high-quality content with a weight > 0.65, and automatically eliminate the remaining low-value information.

5. A system for optimizing dynamic deduplication of prompt words in large language model conversations, characterized in that the system includes: a dynamic storage management module, a text compression and fusion module, and a deduplication module; The dynamic storage management module is used to dynamically maintain the conversation record using a circular buffer, giving priority to retaining the prompt words and eliminating stale data; The text compression and fusion module is used to compress the repeated long texts in the conversation records through a progressive prefix matching algorithm, generate a reversibly parsable tag format, and establish a self-check mechanism for the compressed content to prevent semantic breaks; The deduplication module is used to implement content value evaluation and deduplication based on a composite weight model, construct a quantitative index system in three dimensions of time, frequency, and semantics, and automatically filter out low-weight redundant information in the conversation record content.

6. The system for optimizing the dynamic deduplication of prompt words in the large language model conversation according to claim 5, wherein, The dynamic storage management module is specifically as follows: maintaining the conversation history using a dynamic queue structure, balancing information retention and resource consumption through a capacity control strategy; defining a structure to encapsulate metadata, where the metadata includes the conversation role, content, timestamp, and occurrence frequency, and initializing a circular buffer with a capacity of x + 2; when a new entry is added, if the current context quantity exceeds x, the first x prompt entries are preferentially retained, and the subsequent old data is eliminated according to the first-in, first-out principle; In terms of thread safety, concurrent control is achieved through mutex locks, and at the same time, a global frequency mapping table freqMap is maintained. The global frequency mapping table freqMap is used to count the occurrence times of each compressed text in real time, providing data support for subsequent filtering.

7. The system for optimizing the dynamic deduplication of prompt words in the large language model conversation according to claim 5, wherein, The text compression and fusion module is specifically as follows: scanning the context content item by item, and executing a three-level compression strategy for non-fixed prompt entries: first, converting the text into a rune array, and then starting from the longest substring in the rune array to initiate a matching query to the prefix cache; If a historical repeated prefix is found, a differential compression tag is generated according to the length. If the text length ≤ 4 characters, the complete prefix is retained; if the text length ≥ 4 characters, the first and last two characters are extracted and merged to form a compression tag, and the unmatched text is added to the cache library as a whole. After compression, reverse verification is performed, and it is detected whether the associated original text of the compression tag exists by traversing the context. When it is found that the original text corresponding to the tag is no longer in the current context, content recovery is automatically triggered to prevent the information chain from breaking.

8. The prompt dynamic deduplication text optimization system in the large language model conversation according to claim 5, wherein The deduplication module is specifically as follows: The weight calculation method CalculateWeight of the composite weight model has the following formula: Weight = Time decay factor × 0.4 + Frequency factor × 0.3 + Keyword bonus × 0.3 Among them, the time decay factor is in seconds; the frequency factor Freq is taken from the global statistical table, and the more times the content is repeated, the lower the weight of this item; the frequency factor Freq is taken from the global statistical table, and the global statistical table is used to record the occurrence times, first / most recent occurrence timestamp, and cumulative occurrence times of all content in real time. The keyword bonus item gives a fixed gain of 0.3 to the content marked with "important", calculates the weight value of each entry in real time, retains the high-quality content with a weight > 0.65, and automatically eliminates the remaining low-value information.

Citation Information

Cited By

  • Compression method and device for cue words of large language model system, electronic equipment, storage medium and program product

    CN122021922A