Data management method based on sentence semantic perception and related equipment
Through the data management method of sentence semantic perception, the problem of insufficient utilization of syntax and semantic structure in the existing technology is solved, and the inference efficiency and memory utilization of large language models in long text tasks are improved, ensuring the semantic consistency of the generated content.
Patent Information
- Application Number
- CN202510733515.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing KV cache management method is based on token-level granularity and fails to effectively utilize syntax and semantic structures, resulting in fragmentation of context information, insufficient semantic representation ability, and large computing overhead, making it difficult to meet the needs of online reasoning and low latency.
Using a sentence semantic perception data management method, we use segmentation and sorting of long texts, calculate the importance score of sentence tokens, filter key tokens, and build sentence-level semantic vectors, combine context dependence to dynamically load KV caches, and use a multi-scale attention mechanism to generate semantic index caches.
A more efficient and intelligent reasoning process is realized, and the inference efficiency and memory utilization of large language models in long text tasks are improved, ensuring the semantic consistency of the generated content.
Smart Images

Figure CN120277117A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cache management, and in particular, to a data management method and related devices based on sentence semantic perception. Background Art
[0002] With the widespread application of large language models (LLMs) in natural language processing tasks, they have demonstrated remarkable capabilities in scenarios such as long text generation, multi-turn conversations, and document understanding. However, such models rely on key-value (KV) caches during the inference process to maintain the consistency of context semantics.
[0003] Existing KV cache management methods are often based on token-level granularity and fail to effectively utilize syntactic and semantic structures, resulting in fragmented context information and insufficient semantic representation capabilities. At the same time, relying on external clustering or auxiliary models incurs high computational overhead and poor real-time performance, making it difficult to meet the requirements of online inference and low latency.
[0004] In this context, although techniques such as fixed windows and dynamic pruning have been attempted to optimize KV cache usage; due to fixed windows being prone to truncating long-distance dependency information, affecting semantic coherence; while dynamic pruning is more flexible, it requires continuous calculation of attention scores during the generation process, increasing the system burden.
[0005] Therefore, there is an urgent need for a KV cache management method with sentence-level semantic perception capabilities that can adaptively retain important context, eliminate redundant content, and dynamically load in combination with context dependencies to achieve a more efficient and intelligent inference process. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a data management method and related devices based on sentence semantic perception, and a KV cache management method that dynamically loads in combination with context dependencies to achieve a more efficient and intelligent inference process.
[0007] To achieve the above object, embodiments of the present invention provide the following technical solutions:
[0008] A first aspect of embodiments of the present invention shows a data management method based on sentence semantic perception, the method including:
[0009] When receiving a text caching instruction input by a user, splitting the long text carried by the text caching instruction to obtain multiple sentences and their sorting order;
[0010] For each sentence, calculating the importance score of the tokens corresponding to the sentence based on the sentence and its sorting order;
[0011] For each sentence, based on the importance score of the tokens corresponding to the sentence and the dynamic budget corresponding to the sentence, filter out the key tokens in the sentence and store them in the retention set corresponding to the sentence;
[0012] For the retention set of each sentence, calculate the local weighted attention score of each token based on all the key tokens in the retention set;
[0013] Based on the local weighted attention scores of each token in the sentence and the fusion weights of the preset attention heads, perform weighted fusion to obtain the sentence-level semantic vector;
[0014] Based on the sentence-level semantic vector of the sentence and the corresponding retention set, construct a semantic index cache and a key-value pair KV content cache.
[0015] Optionally, split the long text carried by the text cache instruction to obtain multiple sentences and their sorting order, including:
[0016] Preprocess the long text carried by the text cache instruction to obtain the preprocessed long text;
[0017] If the preprocessed long text contains a preset symbol, split the long text into multiple sentences and their sorting order according to the preset end symbol;
[0018] If the preprocessed long text does not contain a preset symbol, call a preset constructed segmentation model to segment the long text to obtain multiple sentences and their sorting order, where the preset constructed segmentation model is trained based on historical long texts and the corresponding sentences.
[0019] Optionally, for each sentence, calculate the importance score of the tokens corresponding to the sentence based on the sentence and its sorting order, including:
[0020] Based on multiple sentences and their sorting order, determine the currently to-be-processed sentence;
[0021] For the currently to-be-processed sentence, generate corresponding tokens based on the features in the sentence, and the number of tokens is at least one;
[0022] Based on the tokens within the preset context range and the semantic vectors of the tokens, calculate the target vector of the currently to-be-processed sentence;
[0023] For each token of the currently to-be-processed sentence, calculate the similarity between the token and the target vector;
[0024] For each token, calculate the importance score of the token based on the token, the target vector of the sentence corresponding to the token, and the similarity between the token and the target vector.
[0025] Optionally, for each sentence, based on the importance scores of the tokens corresponding to the sentence and the dynamic budget corresponding to the sentence, filter out the key tokens in the sentence and store them in the reserved set corresponding to the sentence, including:
[0026] Select tokens based on the importance scores of the tokens and the sentence weights of each sentence to construct an initial reserved set, where the sentence weight is the average of the importance scores of each token corresponding to the sentence;
[0027] Optimize the importance score of each token based on the initial reserved set and the importance score of the token to obtain the target importance score of each token;
[0028] For each sentence, process the sentence based on the target importance scores of each token in the sentence and the tokens of the sentence to obtain the sentence importance weight of the sentence;
[0029] Determine the dynamic budget allocated to each sentence based on the sentence importance weight of each sentence and the preset total dynamic budget;
[0030] For each sentence, based on the sentence and its corresponding dynamic budget, filter out the key tokens from the sentence and write them into the corresponding reserved set.
[0031] Optionally, based on the local weighted attention scores of each token within the sentence and the fusion weights of the preset attention heads, perform weighted fusion to obtain the sentence-level semantic vector, including:
[0032] For each attention head in the model, perform weighted fusion on the key vectors based on the local weighted attention scores of each token within the sentence to obtain the sentence vector under each attention head, where the key vectors refer to the key vectors of each token under the attention head;
[0033] Perform weighted fusion on the sentence vectors under each attention head and the fusion weights of the preset attention heads to obtain the sentence-level semantic vector of the sentence.
[0034] Optionally, it further includes:
[0035] If it is determined that the current is in the decoding stage, obtain the current generated content and process the current generated content to generate a corresponding joint score;
[0036] Retrieve from the semantic index cache and the key-value pair KV content cache using the combined score to load the candidate sentence set corresponding to the current generated content and the key-value pair KV into the corresponding memory.
[0037] Optionally, retrieving from the semantic index cache and the key-value pair KV content cache using the combined score to load the candidate sentence set corresponding to the current generated content and the key-value pair KV into the corresponding memory includes:
[0038] Retrieve the candidate sentence set from the semantic index cache according to the combined score;
[0039] Traverse the key-value pair KV content cache according to the identifier of each sentence in the candidate sentence set to determine the corresponding key-value pair KV;
[0040] Store the candidate sentence set and the key-value pair KV into the corresponding memory.
[0041] A second aspect of the embodiments of the present invention shows a data management device based on sentence semantic perception, and the device includes:
[0042] A segmentation unit, configured to segment the long text carried in the text cache instruction when receiving the text cache instruction input by the user to obtain multiple sentences and their sorting orders;
[0043] A calculation unit, configured to calculate the importance score of the token corresponding to each sentence based on the sentence and its sorting order;
[0044] A screening unit, configured to screen out the key tokens in each sentence according to the importance score of the token corresponding to the sentence and the dynamic budget corresponding to the sentence, and store them in the reserved set corresponding to the sentence;
[0045] A processing unit, configured to calculate the local weighted attention score of each token based on all the key tokens in the reserved set of each sentence; perform weighted fusion based on the local weighted attention score of each token in the sentence and the fusion weight of the preset attention head to obtain a sentence-level semantic vector;
[0046] A caching unit, configured to construct a semantic index cache and a key-value pair KV content cache based on the sentence-level semantic vector of the sentence and the corresponding reserved set.
[0047] In a third aspect of the embodiments of the present invention, an electronic device is shown. The electronic device includes a processor and a memory. The memory is used to store program codes and data for sentence semantics-aware data management. The processor is used to call the program instructions in the memory to execute the sentence semantics-aware data management method shown in the first aspect of the embodiments of the present invention.
[0048] In a fourth aspect of the embodiments of the present invention, a storage medium is shown. The storage medium includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the sentence semantics-aware data management method shown in the first aspect of the embodiments of the present invention.
[0049] Based on the sentence semantics-aware data management method and related devices provided in the above embodiments of the present invention, the method includes: when receiving a text caching instruction input by a user, splitting the long text carried by the text caching instruction to obtain multiple sentences and their sorting order; for each sentence, calculating the importance score of the token corresponding to the sentence based on the sentence and its sorting order; for each sentence, screening out the key tokens in the sentence according to the importance score of the token corresponding to the sentence and the dynamic budget corresponding to the sentence, and storing them in the reserved set corresponding to the sentence; for the reserved set of each sentence, calculating the local weighted attention score of each token based on all the key tokens in the reserved set; performing weighted fusion based on the local weighted attention score of each token in the sentence and the fusion weight of the preset attention head to obtain a sentence-level semantic vector; constructing a semantic index cache and a key-value pair KV content cache based on the sentence-level semantic vector of the sentence and the corresponding reserved set. In the embodiments of the present invention, first, semantic parsing and sentence splitting are performed on the input long text, and the importance score of the token of each sentence is calculated; then, combined with context adaptive weights, semantic similarity, and attention mechanisms, key tokens corresponding to each sentence are dynamically screened and stored in the reserved set corresponding to the sentence; on this basis, a sentence-level semantic vector is constructed, and a sentence-level semantic vector is generated by means of a multi-scale attention mechanism, and then the corresponding semantic index cache and KV content cache are constructed to uniformly store the sentence-level semantic vectors of all sentences in the semantic index structure in the GPU memory as a fast semantic retrieval basis for the subsequent decoding stage. Thus, combined with the context-dependent dynamic loading KV cache management method, a more efficient and intelligent inference process is realized. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0051] Figure 1 It is a schematic flowchart of a data management method based on sentence semantic perception shown in an embodiment of the present invention;
[0052] Figure 2 It is a schematic flowchart of another data management method based on sentence semantic perception shown in an embodiment of the present invention;
[0053] Figure 3 It is a schematic structural diagram of a data management device based on sentence semantic perception shown in an embodiment of the present invention. Detailed implementation manners
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0055] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above accompanying drawings of this application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0056] It should be noted that in the present invention, the descriptions involving "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0057] In this application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0058] An embodiment of the present invention shows a management method based on sentence semantic perception, aiming to improve the inference efficiency and memory utilization rate of the large language model (LLM) in long text tasks, while ensuring the semantic consistency of the generated content. Specifically, in the caching stage, taking sentences as the basic unit, first perform semantic parsing and sentence segmentation on the input long text, and calculate the importance scores of the tokens of each sentence; then combine context adaptive weights, semantic similarity and attention mechanism to dynamically screen key tokens, and introduce a diversity penalty strategy to optimize the retained content; on this basis, construct sentence-level semantic vectors, generate semantic representations by means of multi-scale attention mechanism, and then construct corresponding semantic index caches and KV content caches. In the decoding stage, according to the semantic vector of the currently generated content, load relevant key-value pairs KV from the cache on demand to form a lightweight semantic support set, and through a dynamic eviction strategy and an asynchronous preloading mechanism, realize the efficient management of the GPU cache. Thus, semantic-driven fine-grained KV cache control is achieved, significantly improving the performance and context retention ability in the long text processing process.
[0059] See Figure 1 , which is a schematic flowchart of a sentence semantic perception cache management method shown in an embodiment of the present invention, and the method includes:
[0060] Step S101: When receiving a text caching instruction input by a user, segment the long text carried by the text caching instruction to obtain multiple sentences and their sorting order.
[0061] Optionally, after the user inputs a long text, a corresponding text caching instruction is triggered, so that the text caching instruction carries the long text input by the user.
[0062] It should be noted that the process of specifically implementing step S101 includes the following steps.
[0063] Step S11: Preprocess the long text carried by the text caching instruction to obtain the preprocessed long text.
[0064] In the process of specifically implementing step S11, traverse the long text according to the preset irrelevant symbols and preset special characters. If it is determined that the corresponding preset irrelevant symbols and preset special characters exist in the long text, delete them to obtain the preprocessed long text.
[0065] It should be noted that both the preset irrelevant symbols and the preset special characters are preset in advance. The preset irrelevant symbols include redundant spaces, line breaks, etc.; the preset special characters include emoticons, website addresses, etc.
[0066] Optionally, for the preset special characters, they can be retained or deleted according to actual needs.
[0067] Step S12: Determine whether the preprocessed long text contains a preset symbol. If it does, execute step S13; otherwise, execute step S14.
[0068] In the process of specifically implementing step S12, traverse the long text to determine whether the long text contains a preset symbol. If it does, execute step S13; otherwise, execute step S14.
[0069] It should be noted that the preset symbol is preset according to the actual situation in advance and can generally be punctuation marks.
[0070] Step S13: Split the long text into multiple sentences and their sorting order according to the preset end symbol.
[0071] It should be noted that the preset end symbol is set in advance based on experience or multiple experiments and can generally be a full stop “。”, a question mark “?”, etc.
[0072] For example: for the input long text A “The weather is very good today. I go for a walk in the park.”, split the long text A into two sentences according to the preset end symbol, specifically including sentence 1: The weather is very good today. Sentence 2: I go for a walk in the park, where 1 and 2 represent the sorting order of the sentences.
[0073] Among them, sentence 1 and sentence 2 are sorted according to the text order, that is, the order of sentence 1 is 1, the order of sentence 2 is 2, and sentence 1 is before sentence 2.
[0074] Step S14: Invoke the pre-built segmentation model to segment the long text, obtaining multiple sentences and their sorting order.
[0075] Among them, the pre-built segmentation model is trained based on historical long texts and corresponding sentences.
[0076] It should be noted that the process of training the segmentation model based on historical long texts and corresponding sentences includes:
[0077] Obtain historical long texts and corresponding sentences as samples; divide the samples into a training set and a test set; train the initial model based on the training set to obtain a trained initial model; use the trained initial model to perform context understanding on the long texts in the test set of the samples to infer the start and end positions of the sentences, thereby obtaining the corresponding sentences; determine whether the obtained sentences are consistent with the sentences corresponding to the long texts in the test set. If they are consistent, use the trained initial model as the segmentation model. Otherwise, continue to train the initial model using the training set.
[0078] It should be noted that the initial model is a natural language processing BERT model.
[0079] In the specific implementation process of step S14, use the pre-built segmentation model to perform context understanding on the long text to infer the start and end positions of the sentences, thereby obtaining multiple sentences and the sorting order of the sentences.
[0080] In this application, for each sentence, each sentence is used as an independent unit for subsequent token importance calculation and semantic vector calculation.
[0081] Step S102: For each sentence, calculate the importance score of the tokens corresponding to the sentence based on the sentence and its sorting order.
[0082] It should be noted that the specific implementation process of step S102 includes:
[0083] Step S21: Based on multiple sentences and their sorting order, determine the currently pending sentence.
[0084] In the specific implementation process of step S21, the sentences need to be executed in the sorting order. Mark the executed sentences as executed, and sequentially determine the currently pending sentence in ascending order of the sorting order.
[0085] Step S22: For the currently pending sentence, generate corresponding tokens based on the features in the sentence, and the number of tokens is at least one.
[0086] In the process of specifically implementing step S22, in the order of the sentences in the text, the sentences are sequentially divided into features such as words, phrases, symbols, or others, that is, tokens.
[0087] For example, the sentence "I like cats." might be split into a list of tokens like ["I", "like", "cats", "."]. In the English sentence "I like cats.", tokenization might produce ["I", "like", "cats", "."], where even punctuation marks might be treated as separate tokens.
[0088] Step S23: Based on the tokens within a preset context range and the semantic vectors of the tokens, calculate the target vector of the currently to-be-processed sentence.
[0089] It should be noted that the preset context range refers to the range from the first sentence in the sorting order to the sentence corresponding to the currently to-be-processed sentence.
[0090] In the process of specifically implementing step S23, all the tokens from the first sentence in the sorting order to the currently to-be-processed sentence are obtained within the preset context range, which can be expressed as , ,... , where L represents the number of tokens that have been generated currently; then, for each token within the preset context range , calculate its semantic vector through the last layer of the network hidden states Hidden States of the large speech model Transformer , the semantic vector is Model( ); then, based on the semantic vector of each token , substitute it into formula (1) to calculate the corresponding average vector, and use it as the target vector of the currently to-be-processed sentence .
[0091] Formula (1):
[0092]
[0093] Step S24: For each token of the currently to-be-processed sentence , calculate the similarity between the token and the target vector.
[0094] In the process of specifically implementing step S24, for each token of the currently to-be-processed sentence , take the token and the target vector of this currently to-be-processed sentence , substitute them into formula (2) to calculate the similarity between the token and the target vector .
[0095] Formula (2):
[0096]
[0097] Based on this, through the above steps S21 to S24, the target vector corresponding to each sentence in the long text can be calculated, as well as the similarity between each token under each sentence and the corresponding target vector.
[0098] Step S25: For each token, calculate the importance score of the token based on the token, the target vector of the sentence corresponding to the token, and the similarity between the token and the target vector.
[0099] In the process of specifically implementing step S25, the final importance score of each token not only depends on its similarity to the target semantics, but also combines its relationship with other tokens in the current context. Based on this, for each token, the token , the target vector of the sentence corresponding to the token , and the similarity between the token and the target vector are substituted into formula (3) for calculation to obtain the importance score of the token .
[0100] Formula (3):
[0101]
[0102] Wherein, is the importance score of the token , is the similarity between the token and the target vector, represents the attention score between token and token . This attention score represents the importance of token in the current context, combining semantic relevance and attention information; is a variable window size dynamically selected by calculating the contribution degree of the token pair within the preset context range of the currently processed sentence to generate the target vector.
[0103] Optionally, a variable window size is dynamically selected by calculating the contribution degree of the token pair within the preset context range of the currently processed sentence , and this window size is related to factors such as the complexity of the currently generated target and the richness of the context.
[0104] Wherein, the size of the window is adjusted according to the change of the context range to avoid information loss caused by a fixed window.
[0105] Step S103: For each sentence, based on the importance score of the tokens corresponding to the sentence and the dynamic budget corresponding to the sentence, filter out the key tokens in the sentence and store them in the reserved set corresponding to the sentence;
[0106] It should be noted that in the process of specifically implementing step S103, the following steps are included.
[0107] Step S31: Select tokens based on the importance score of the tokens and the sentence weights of each sentence to construct an initial reserved set.
[0108] Among them, the sentence weight is the average value of the importance scores of each token corresponding to the sentence.
[0109] In the process of specifically implementing step S31, first, for each sentence, calculate the average value of the importance scores of each token corresponding to the sentence and use it as the sentence importance weight of the sentence; then, determine whether the sentence weight of the sentence is greater than the first threshold. If so, reserve the tokens corresponding to the sentence with a weight greater than the first threshold in the initial reserved set; then, judge the importance scores of each token of other sentences and reserve the tokens with an importance score greater than the second threshold in the initial reserved set.
[0110] It should be noted that both the first threshold and the second threshold are set by technicians in advance according to the actual situation.
[0111] Optionally, in the preliminary reserved set, introduce a diversity penalty mechanism based on the semantic similarity between tokens to reduce the importance scores of semantically redundant tokens and further optimize the semantic dispersion of the reserved token set.
[0112] Step S32: Optimize the importance score of each token based on the initial reserved set and the importance score of the token to obtain the target importance score of each token.
[0113] To avoid losing the overall context coverage due to over-concentration in local contexts (such as a conversation or a narrative), design a diversity penalty term to select tokens from different semantic segments; in the process of specifically implementing step S32, for each token i in each sentence, calculate its semantic similarity with the token j in the initial reserved set ; Substitute the semantic similarity and the token into formula (4) for calculation to correct the importance score and obtain the target importance score , that is, the final importance score of the token.
[0114] Formula (4):
[0115]
[0116] Among them, λ is a preset penalty coefficient, and Selected is the set of selected tokens, that is, the initial retention set. is the importance score of the token . is the target importance score of the token .
[0117] It should be noted that the semantic similarity is a similarity matrix , measuring the semantic similarity between tokens i and j.
[0118] Step S33: For each sentence, process based on the target importance score of each token in the sentence and the tokens of the sentence to obtain the sentence importance weight of the sentence.
[0119] In the process of specifically implementing step S33, for each sentence, substitute all the tokens of the sentence and the target importance score of each token in the sentence into formula (5) for calculation to obtain the sentence importance weight of the sentence .
[0120] Formula (5):
[0121]
[0122] Among them, is the set of all tokens included in sentence s, is the final importance score of the i-th corresponding token .
[0123] Step S34: Based on the sentence importance weight of each sentence and the preset total dynamic budget τ, determine the dynamic budget allocated to each sentence.
[0124] In the process of specifically implementing step S34, when reserving tokens, it is allowed to dynamically adjust the number of tokens reserved in different parts according to the actual complexity of the context. Specifically, for each sentence, substitute the preset total budget τ and the sentence importance weight corresponding to the sentence into formula (6) for calculation to obtain the dynamic budget of the sentence , and then obtain the dynamic budget of each sentence .
[0125] Formula (6):
[0126]
[0127] Among them, is the budget allocated to sentence s, is the sentence importance weight of sentence s, and τ is the preset total budget τ.
[0128] Optionally, through formula (6), the retention budget can be appropriately increased for long sentences and important sentences (such as paragraphs with high scores); the budget can be compressed for short sentences and redundant content.
[0129] Step S35: For each sentence, based on the sentence and its corresponding dynamic budget, filter out the key tokens from the sentence and write them into the corresponding retention set.
[0130] It should be noted that the number of tokens filtered out from the sentence can be 1 or multiple.
[0131] In the process of specifically implementing step S35, for each sentence s, under its dynamic budget limit, filter out τ tokens with the highest scores and reasonable distributions in the sentence as the key tokens, and then retain the key tokens in the retention set corresponding to the sentence.
[0132] Optionally, the retention set is the GPU memory, and the remaining tokens of the sentence are transferred to the CPU or discarded.
[0133] Step S104: For the retention set of each sentence, calculate the local weighted attention score of each token based on all the key tokens in the retention set;
[0134] In the process of specifically implementing step S104, for the retention set of each sentence, first, calculate the attention values between the key tokens in the retention set; then substitute the attention values between the key tokens and the retention set into formula (7) for calculation to obtain the local weighted attention score of each token , that is, for all the retained tokens in each sentence s, first calculate the local weighted attention score of each token , used to measure the contribution degree of token x to the overall sentence under attention head h.
[0135] Formula (7):
[0136]
[0137] Among them, is the attention value between token x and token y under attention head h, x is a key token in the retention set , y is another key token in the retention set ; is the number of key tokens included in the retention set .
[0138] Step S105: Perform weighted fusion based on the local weighted attention scores of each token in the sentence and the fusion weights of the preset attention heads to obtain a sentence-level semantic vector;
[0139] It should be noted that in the process of specifically implementing Step S105, the following steps are included.
[0140] Step S41: For each attention head in the model, perform weighted fusion on the key vectors based on the local weighted attention scores of each token in the sentence to obtain a sentence vector under each attention head, where the key vectors refer to the key vectors of each token under the attention head;
[0141] In the process of specifically implementing Step S41, first, for each sentence s, assume that the retention set of the sentence contains tokens, and the key vector of token x is , where h is the attention head; within each attention head h, based on the local weighted attention score of each token and the key vectors under the sentence, substitute them into formula (8) to obtain the sentence vector under this attention head, and then obtain the sentence vectors under each attention head.
[0142] Formula (8):
[0143]
[0144] Step S42: Perform weighted fusion on the sentence vectors under each attention head and the fusion weights of the preset attention heads to obtain the sentence-level semantic vector of the sentence.
[0145] In the process of specifically implementing Step S42, input the sentence vectors under each attention head and the fusion weights of the preset attention heads into formula (9) for calculation to obtain the sentence-level semantic vector of the sentence.
[0146] It should be noted that one attention head corresponds to one fusion weight, and the fusion weight corresponding to the h-th attention head is , where h belongs to H.
[0147] Among them, the number of attention heads is H.
[0148] Formula (9):
[0149]
[0150] Among them, is the fusion weight corresponding to the h-th attention head.
[0151] In an embodiment of the present invention, for sentence s, first, within each attention head h, local normalized weights are calculated based on the attention scores among all tokens within the sentence, that is, local weighted attention scores , and the key vectors of each token are weighted to generate a single-head sentence, that is, a sentence vector . Subsequently, based on the heterogeneous features of different attention heads, learnable fusion weights are used to fuse the sentence vectors of all attention heads , and finally a unified sentence-level semantic vector is formed. Based on the tokens in the retention set, a sentence-level semantic vector is generated through a multi-scale, hierarchical attention weighting mechanism .
[0152] Step S106: Construct a semantic index cache and a key-value pair KV content cache based on the sentence-level semantic vector of the sentence and the corresponding retention set
[0153] During the specific implementation of step S106, the following steps are included
[0154] Step S51: Form a corresponding semantic index based on the sentence-level semantic vector of the sentence and cache it
[0155] Specifically, extract the sentence-level semantic vector of each sentence ; store the sentence-level semantic vector in the GPU high-speed memory to form a semantic index cache for each sentence
[0156] It should be noted that for each sentence, its additional structural feature information is further recorded, including but not limited to: sentence length normalization value, token importance mean, and sentence position information
[0157] The semantic index cache is used for retrieval operations based on multi-factor matching in the decoding stage
[0158] Step S52: Establish a key-value pair KV content cache based on the tokens in the retention set
[0159] During the specific implementation of step S52, for the key tokens filtered and retained in the retention set of each sentence, their corresponding Key-Value pairs are saved; the above Key-Value pairs are stored in the CPU memory or a low-frequency persistent storage medium, that is, in the key-value pair KV content cache; to support subsequent retrieval and loading of Key-Value pairs at the sentence level or clause level
[0160] The present invention implements the prefill stage of the LLM through the above steps S101 to S106. This stage occurs after the LLM receives the complete input long text but before starting to generate the first output token. In this stage, the input long text is mainly processed to determine the sentence-level semantic vectors of the sentences corresponding to the long text, as well as the key-value pairs KV of the tokens in the retention set corresponding to each sentence, and to construct a semantic index cache and a key-value pair KV content cache.
[0161] In an embodiment of the present invention, in the caching stage, taking sentences as the basic unit, first perform semantic parsing and sentence segmentation on the input long text, and calculate the importance scores of the tokens in each sentence; then, in combination with context adaptive weights, semantic similarity, and attention mechanisms, dynamically filter the key tokens corresponding to each sentence and store them in the retention set corresponding to the sentence; on this basis, construct sentence-level semantic vectors, generate sentence-level semantic vectors by means of a multi-scale attention mechanism, and then construct the corresponding semantic index cache and KV content cache to uniformly store the sentence-level semantic vectors of all sentences in the semantic index structure in the GPU memory as the basis for fast semantic retrieval in the subsequent decoding stage. At the same time, record the corresponding Key / Value for each key token in the retention set, and these data structures are uniformly organized into cache units divided by sentences to provide a fine-grained context loading basis for the decoding stage.
[0162] Optionally, after the prefill stage of the above steps S101 to S105 ends, the LLM enters the decoding stage; in this stage, the LLM generates output tokens one by one in an autoregressive manner. Each time a token is generated, the token is added to the generated sequence and used as the input for the next generation.
[0163] Among them, the decoding stage starts from the last token of the long text processed in the prefill stage, and the goal is to generate the subsequent output sequence, and the output sequence includes multiple output tokens.
[0164] It should be noted that if it is the first step of the decoding stage, use the last token of the long text processed in the prefill stage as the input of the current Transformer layer of the LLM; if it is not the first step of the decoding stage, use the token generated in the previous step as the input of the current Transformer layer.
[0165] Based on the management method based on sentence semantic perception shown in the above embodiments of the present invention, correspondingly, the present invention also shows a flow schematic diagram of another management method based on sentence semantic perception, as Figure 2 shown, the method includes:
[0166] Step S201: If it is determined that the current is in the decoding stage, obtain the current generated content and process the current generated content to generate a corresponding joint score.
[0167] In the process of specifically implementing step S201, if it is determined that the current is in the decoding stage, the current generated content of a certain attention mechanism of the LLM is obtained, and the target vector corresponding to the current generated content and the corresponding sentence-level semantic vector are input into formula (10) for processing to generate a corresponding joint score. 。
[0168] Formula (10):
[0169]
[0170] Wherein, includes the sentence length normalization value, position information, token importance density, etc.; and are coefficients, which are preset, is the target vector corresponding to the content, is the corresponding sentence-level semantic vector.
[0171] It should be noted that the target vector corresponding to the current generated content and can be obtained through the processes of the above steps S101 to step 105.
[0172] The current generated content is generally the token output by the current Transformer layer of the LLM.
[0173] Step S202: Use the joint score to retrieve from the semantic index cache and the key-value pair KV content cache to load the candidate sentence set and the key-value pair KV corresponding to the current generated content into the corresponding memory.
[0174] It should be noted that the self-attention mechanism inside the LLM can still access the semantic index cache and the key-value pair KV content cache.
[0175] It should be noted that in the process of specifically implementing step S202, the following steps are included.
[0176] Step S61: Retrieve the candidate sentence set from the semantic index cache according to the joint score;
[0177] In the process of specifically implementing step S61, retrieve and filter the top-K most relevant sentences from the semantic index cache according to the combined score. That is to say, retrieve and filter the top-K, i.e., the preset number of, sentence-level semantic vectors that are the same as or closest to the combined score from the semantic index cache, and use the corresponding sentences as candidate loading targets, i.e., the candidate sentence set.
[0178] Step S62: Traverse the key-value pair KV content cache according to the identifier of each sentence in the candidate sentence set to determine the corresponding key-value pair KV.
[0179] In the process of specifically implementing step S62, each sentence is associated with its sentence number or unique identifier ID; use the identifier of each sentence in the candidate sentence set to quickly locate in the KV content cache (in CPU memory or disk) to obtain the storage offset and index information of the reserved key-value pair.
[0180] Step S63: Store the candidate sentence set and the key-value pair KV into the corresponding memory.
[0181] In the process of specifically implementing step S63, selectively load the KV and the candidate sentence set from the CPU into an independent memory unit in the GPU based on the combined score, so that the next generated content in the LLM decoding stage can be directly called.
[0182] Optionally, in addition to loading the candidate sentence set shown above, the loading range of the cache can also be divided according to the granularity of units such as clauses, phrases, keyword paragraphs, etc.
[0183] Among them, the clause division method includes but is not limited to punctuation marks (commas, semicolons, etc.), and syntactic analysis includes token attention clustering.
[0184] Optionally, to improve the management efficiency of the KV cache in the GPU memory, a multi-dimensional state-aware dynamic cache control mechanism is introduced.
[0185] Specifically, it includes: a soft eviction policy for preferentially releasing cache units with low popularity and long-term inaccessibility;
[0186] Dynamic quota adjustment, flexibly adjusting the cache residency capacity according to the remaining video memory of the GPU; and an asynchronous pre-loading mechanism, which pre-loads the possibly relevant sentence cache into the pre-buffer in advance by predicting the future generated semantic vectors, so as to achieve a low-latency and high-semantic-consistency generation process.
[0187] Optionally, to ensure the decoding efficiency and semantic consistency, after executing step S202, the present application further includes the following steps.
[0188] Step S71: Determine the context semantic information required for the currently generated content.
[0189] In the process of specifically implementing Step S71, during the generation process of each round of content, collect the query vectors of all tokens in the currently generated window. When it is determined that a certain query vector is a sentence boundary (such as a period, a question mark, a carriage return), calculate the sentence-level query semantic vector. As shown in formula (11).
[0190] Formula (11):
[0191]
[0192] Where, represents the semantic target of the currently generated sentence and is used for subsequent matching operations.
[0193] Next, based on the semantic index cache saved in the GPU (including the semantic vectors of each sentence), ), by calculating the cosine similarity between and the sentence-level semantic vectors of each sentence in the semantic index cache, obtain the preset number Top-K sentence cache units with the largest cosine similarity, that is, the most relevant context semantic information required for the currently generated content from the semantic index cache.
[0194] It should be noted that to improve the matching accuracy, the LLM can combine other structured features (such as sentence length, number of reserved tokens, etc.) to perform multi-factor sorting and weighted matching.
[0195] The Top-K sentence cache units include the corresponding sentences.
[0196] In the present invention. To improve the matching accuracy, other structured features (such as sentence length, number of reserved tokens, etc.) can be combined to perform multi-factor sorting and weighted matching.
[0197] Step S72: Based on the context semantic information required for the currently generated content, load relevant Key-Value pairs from the cache to form a lightweight semantic support set for supporting the current generation.
[0198] In the process of specifically implementing Step S72, load some or all of the Key-Value pairs of the Top-K sentence cache units from the CPU memory. This process supports the KV loading at the sub-clause or token subset level. The loaded cache units are added to the active state table and participate in the attention calculation during the current generation process. In addition, the soft eviction, dynamic residency, and asynchronous preloading mechanisms of the GPU cache are also executed to ensure that the optimal sentence set is always maintained in the GPU memory, thereby improving the generation response speed and semantic accuracy.
[0199] Optionally, it further includes:
[0200] During the generation stage, the LLM queries the vector of the current generated content's tokens and performs efficient attention calculations with the Key-Value pairs already loaded in the GPU to generate the next token, which is automatically added to the cache of the current generated content.
[0201] Whenever a complete sentence is generated, recalculate the semantic vector of the sentence and update the matching basis for the subsequent query stage according to the sentence-level query semantic vector, thereby enhancing the context consistency of the generated content.
[0202] Optionally, the LLM continuously records the matching relationship between sentences and cache units during each round of generation, constructing a dynamic semantic matching feedback matrix. This matrix is used to optimize the retrieval accuracy and matching efficiency of subsequent sentences in real time, thereby ensuring that the generated content is highly relevant to the cached information.
[0203] Optionally, during the long-term generation process, i.e., the decoding stage, to improve the generation efficiency and memory utilization, semantic compression is performed on the historical context, converting low-frequency sentences into summary representations. At the same time, the LLM adopts a soft freezing strategy to ensure that some high-value semantic sentences always remain in the GPU memory, thereby maximizing semantic coherence and optimizing the efficient use of memory.
[0204] Optionally, it further includes:
[0205] During each round of generation in the decoding stage, the LLM records the association strength between the candidate sentence cache unit and the tokens actually called by the attention mechanism, forming a triple:
[0206] < , , call frequency>
[0207] These triples provide the basic data for the construction of the subsequent feedback matrix. The LLM constructs a trainable semantic matching matrix (where d is the vector dimension). The matrix parameters are optimized by minimizing the following loss function, as shown in Equation (12).
[0208] Equation (12):
[0209] +
[0210] Where, is the matching error, is the regularization term to prevent overfitting. The optimized matrix M is used to enhance the semantic alignment ability of subsequent retrievals and improve the accuracy and efficiency of retrievals.
[0211] Optionally, the LLM dynamically evaluates the popularity of cached sentences based on access frequency and maintains a popularity score. The popularity update follows time decay, as shown in formula (13).
[0212] Formula (13):
[0213]
[0214] where, is the decay coefficient, and I (number of calls) is an indicator function representing the number of calls to the cached sentence. Through this formula, the popularity score can flexibly reflect the usage frequency of sentence caching and is continuously adjusted according to time decay.
[0215] The LLM classifies the cached sentences into three categories according to the popularity score: hot data are sentences with frequent access and are retained in the GPU memory to ensure efficient access. Warm data are sentences with medium frequency and are stored in the CPU memory. Cold data are sentences with low access frequency and are stored on disk.
[0216] In the embodiment of the present invention, in the decoding stage, according to the semantic vector of the currently generated content, relevant key-value pairs KV are loaded from the cache as needed to form a lightweight semantic support set, and through a dynamic eviction policy and an asynchronous preloading mechanism, efficient management of the GPU cache is achieved. Thus, semantic-driven fine-grained KV cache control is realized, significantly improving the performance and context retention ability during long text processing.
[0217] Corresponding to the data management method based on sentence semantic perception shown in the above embodiment of the present invention, a data management device based on sentence semantic perception is also shown, as Figure 3 shown, the device includes:
[0218] A segmentation unit 301, configured to segment the long text carried in the text caching instruction when receiving the text caching instruction input by the user, to obtain multiple sentences and their sorting order;
[0219] A calculation unit 302, configured to calculate the importance score of the tokens corresponding to each sentence based on the sentence and its sorting order;
[0220] A screening unit 303, configured to screen out the key tokens in each sentence according to the importance score of the tokens corresponding to the sentence and the dynamic budget corresponding to the sentence, and store them in the retention set corresponding to the sentence;
[0221] A processing unit 304 is configured to calculate local weighted attention scores for each token based on all the key tokens in the retention set for each sentence; perform weighted fusion based on the local weighted attention scores of each token in the sentence and the fusion weights of the preset attention heads to obtain a sentence-level semantic vector;
[0222] A cache unit 305 is configured to construct a semantic index cache and a key-value pair KV content cache based on the sentence-level semantic vector of the sentence and the corresponding retention set.
[0223] The specific principles and execution processes of each unit in the data management device based on sentence semantic perception disclosed in the above embodiments of the present invention are the same as the corresponding content in the data management method based on sentence semantic perception provided in the above embodiments of the present invention. Reference can be made to the corresponding parts in the data management method based on sentence semantic perception disclosed in the above embodiments of the present invention, and details will not be described here.
[0224] In the embodiments of the present invention, in the caching stage, taking sentences as basic units, first perform semantic parsing and sentence segmentation on the input long text, and calculate the importance scores of the tokens in each sentence; then, in combination with context adaptive weights, semantic similarity, and attention mechanisms, dynamically screen the key tokens corresponding to each sentence and store them in the retention set corresponding to the sentence; on this basis, construct a sentence-level semantic vector, generate a sentence-level semantic vector by means of a multi-scale attention mechanism, and then construct a corresponding semantic index cache and KV content cache to uniformly store the sentence-level semantic vectors of all sentences in a semantic index structure in the GPU memory as a basis for fast semantic retrieval in the subsequent decoding stage. At the same time, record the corresponding Key Value for each key token in the retention set, and these data structures are uniformly organized into cache units divided by sentences to provide a fine-grained context loading basis for the decoding stage.
[0225] Optionally, based on a data management device based on sentence semantic perception shown in the above embodiments of the present invention, the segmentation unit 301 includes:
[0226] Preprocess the long text carried by the text cache instruction to obtain a preprocessed long text;
[0227] If the preprocessed long text contains a preset symbol, split the long text into multiple sentences and their sorting orders according to the preset end symbol;
[0228] If the preprocessed long text does not contain a preset symbol, call a preset constructed segmentation model to segment the long text to obtain multiple sentences and their sorting orders, where the preset constructed segmentation model is trained based on historical long texts and corresponding sentences.
[0229] Optionally, based on a data management device based on sentence semantic perception shown in the above embodiments of the present invention, the calculation unit 302 includes:
[0230] Determine the currently to-be-processed sentence based on the plurality of sentences and their sorting order;
[0231] For the currently to-be-processed sentence, generate corresponding tokens based on the features in the sentence, and the number of the tokens is at least one;
[0232] Based on the tokens within a preset context range and the semantic vectors of the tokens, calculate the target vector of the currently to-be-processed sentence;
[0233] For each token of the currently to-be-processed sentence, calculate the similarity between the token and the target vector;
[0234] For each token, calculate the importance score of the token based on the token, the target vector of the sentence corresponding to the token, and the similarity between the token and the target vector.
[0235] Optionally, based on a data management device based on sentence semantic perception shown in the above embodiments of the present invention, the screening unit 303 includes:
[0236] Select tokens based on the importance scores of the tokens and the sentence weights of each sentence to construct an initial retention set, where the sentence weight is the average of the importance scores of each token corresponding to the sentence;
[0237] Optimize the importance score of each token based on the initial retention set and the importance score of the token to obtain the target importance score of each token;
[0238] For each sentence, process based on the target importance score of each token in the sentence and the tokens of the sentence to obtain the sentence importance weight of the sentence;
[0239] Based on the sentence importance weights of each sentence and a preset total dynamic budget, determine the dynamic budget allocated to each sentence;
[0240] For each sentence, based on the sentence and its corresponding dynamic budget, screen out key tokens from the sentence and write them into the corresponding retention set.
[0241] Optionally, based on a data management device based on sentence semantic perception shown in the above embodiments of the present invention, the processing unit 304 for weighted fusion of the local weighted attention scores of each token within the sentence and the fusion weights of a preset attention head to obtain a sentence-level semantic vector includes:
[0242] For each attention head in the model, based on the local weighted attention scores of each token in the sentence, the key vectors are weighted and fused to obtain the sentence vector under each attention head, where the key vectors refer to the key vectors of each token under the attention head.
[0243] The sentence vectors under each attention head and the fusion weights of the preset attention heads are weighted and fused to obtain the sentence-level semantic vector of the sentence.
[0244] Optionally, based on the data management device based on sentence semantic perception shown in the embodiments of the present invention above, it further includes:
[0245] A generation unit, configured to obtain the current generated content and process the current generated content to generate a corresponding joint score if it is determined that the current is in the decoding stage.
[0246] A retrieval unit retrieves from the semantic index cache and the key-value pair KV content cache using the joint score to load the candidate sentence set and the key-value pair KV corresponding to the current generated content into the corresponding memory.
[0247] Among them, retrieving from the semantic index cache and the key-value pair KV content cache using the joint score to load the candidate sentence set and the key-value pair KV corresponding to the current generated content into the corresponding memory includes:
[0248] Retrieving the candidate sentence set from the semantic index cache according to the joint score;
[0249] Traversing the key-value pair KV content cache according to the identifier of each sentence in the candidate sentence set to determine the corresponding key-value pair KV;
[0250] Storing the candidate sentence set and the key-value pair KV into the corresponding memory.
[0251] An embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory is used to store the data management program code and data based on sentence semantic perception, and the processor is used to call the program instructions in the memory to execute the data management method based on sentence semantic perception as described in the above embodiments.
[0252] An embodiment of the present invention provides a storage medium, which includes the electronic device provided in the above embodiment of the present application, and this electronic device is used to execute the data management method based on sentence semantic perception disclosed in the embodiments of the present application.
[0253] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant content. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0254] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0255] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data management method based on sentence semantic perception, characterized in that, The method includes: When receiving a text caching instruction input by a user, splitting the long text carried by the text caching instruction to obtain multiple sentences and their sorting order; For each sentence, calculating the importance score of the token corresponding to the sentence based on the sentence and its sorting order; For each sentence, screening out the key tokens in the sentence according to the importance score of the token corresponding to the sentence and the dynamic budget corresponding to the sentence, and storing them in the reserved set corresponding to the sentence; For the reserved set of each sentence, calculating the local weighted attention score of each token based on all the key tokens in the reserved set; Based on the local weighted attention scores of each token in the sentence and the fusion weights of the preset attention heads, performing weighted fusion to obtain a sentence-level semantic vector; Based on the sentence-level semantic vector of the sentence and the corresponding reserved set, constructing a semantic index cache and a key-value pair KV content cache.
2. The method according to claim 1, wherein Splitting the long text carried by the text caching instruction to obtain multiple sentences and their sorting order includes: Performing preprocessing on the long text carried by the text caching instruction to obtain a preprocessed long text; If the preprocessed long text contains a preset symbol, splitting the long text into multiple sentences and their sorting order according to the preset end symbol; If the preprocessed long text does not contain a preset symbol, calling a preset constructed segmentation model to segment the long text to obtain multiple sentences and their sorting order, where the preset constructed segmentation model is trained based on historical long texts and corresponding sentences.
3. The method according to claim 1, characterized in that, For each sentence, calculating the importance score of the token corresponding to the sentence based on the sentence and its sorting order includes: Based on multiple sentences and their sorting order, determining the currently to-be-processed sentence; For the currently to-be-processed sentence, generating corresponding tokens based on the features in the sentence, and the number of tokens is at least one; Based on the tokens within a preset context range and the semantic vectors of the tokens, calculating the target vector of the currently to-be-processed sentence; For each token in the currently to-be-processed sentence, calculating the similarity between the token and the target vector; For each token, calculating the importance score of the token based on the token, the target vector of the sentence corresponding to the token, and the similarity between the token and the target vector.
4. The method according to claim 1, characterized in that, For each sentence, screening out the key tokens in the sentence according to the importance score of the token corresponding to the sentence and the dynamic budget corresponding to the sentence, and storing them in the reserved set corresponding to the sentence includes: Selecting tokens based on the importance score of the token and the sentence weights of each sentence to construct an initial reserved set, where the sentence weight is the average of the importance scores of each token corresponding to the sentence; Based on the initial reserved set and the importance score of the token, optimizing the importance score of each token to obtain the target importance score of each token; For each sentence, processing based on the target importance score of each token in the sentence and the tokens of the sentence to obtain the sentence importance weight of the sentence; Determine the dynamic budget allocated to each of the sentences based on the sentence importance weight of each sentence and a preset total dynamic budget; For each sentence, based on the sentence and its corresponding dynamic budget, filter out the key tokens from the sentence and write them into the corresponding retention set.
5. The method according to claim 1, wherein Perform weighted fusion based on the local weighted attention scores of each token within the sentence and the fusion weights of preset attention heads to obtain a sentence-level semantic vector, including: For each attention head in the model, perform weighted fusion on the key vectors based on the local weighted attention scores of each token within the sentence to obtain a sentence vector under each attention head, where the key vectors refer to the key vectors of each token under the attention head; Perform weighted fusion on the sentence vectors under each attention head and the fusion weights of the preset attention heads to obtain the sentence-level semantic vector of the sentence.
6. The method according to claim 1, characterized in that Further include: If it is determined that the current is in the decoding stage, obtain the current generated content and process the current generated content to generate a corresponding joint score; Use the joint score to retrieve from the semantic index cache and the key-value pair KV content cache to load the candidate sentence set and the key-value pair KV corresponding to the current generated content into the corresponding memory.
7. The method according to claim 6, characterized in that, Use the joint score to retrieve from the semantic index cache and the key-value pair KV content cache to load the candidate sentence set and the key-value pair KV corresponding to the current generated content into the corresponding memory, including: Retrieve the candidate sentence set from the semantic index cache according to the joint score; Traverse the key-value pair KV content cache according to the identifier of each sentence in the candidate sentence set to determine the corresponding key-value pair KV; Store the candidate sentence set and the key-value pair KV into the corresponding memory.
8. A data management device based on sentence semantic perception, characterized in that The device includes: A segmentation unit, configured to segment the long text carried in the text cache instruction into multiple sentences and their sorting order if a text cache instruction input by a user is received; A calculation unit, configured to calculate the importance score of the tokens corresponding to each sentence based on the sentence and its sorting order; A screening unit, configured to screen out the key tokens in each sentence according to the importance score of the tokens corresponding to the sentence and the dynamic budget corresponding to the sentence, and store them in the retention set corresponding to the sentence; A processing unit, configured to calculate the local weighted attention score of each token based on all the key tokens in the retention set of each sentence; perform weighted fusion based on the local weighted attention scores of each token within the sentence and the fusion weights of preset attention heads to obtain a sentence-level semantic vector; A caching unit, configured to construct a semantic index cache and a key-value pair KV content cache based on the sentence-level semantic vector of the sentence and the corresponding retention set.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory is used to store program codes and data for data management based on sentence semantic perception, and the processor is used to call the program instructions in the memory to execute the data management method based on sentence semantic perception according to any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the data management method based on sentence semantic perception according to any one of claims 1-7.
Citation Information
Patent Citations
Text summarization method and system based on deep learning
CN114385806A
Document level relationship extraction based on span negative sample and enhanced context representation
CN117195075A
Threat entity extraction method based on autoregression tag subsequences
CN118536508A
Method and device for reasoning cache optimization of generative language model
CN119761500A
Transform calculation optimization method based on mixed depth
CN119883647A
Cited By
Key value cache allocation method, electronic equipment and computer program
CN120973836A
Intelligent tour guide method and system based on intelligent token and semantic fusion
CN121301533A
An intelligent guide method and system based on intelligent token and semantic fusion
CN121301533B
Algorithm method for large-model long-context reasoning
CN121365738A
Large language model-oriented data management method
CN122020132A