Fusion caching method, system and equipment based on LLM cue word and medium

By converting prompt words into embedded vectors and calculating semantic similarity using vector database, storing them in Redis cache, and organizing cross-requested KV cache in video memory, the problems of repeated calculations and resource waste in large language models are solved, and efficient inference performance and response time are achieved.

CN120277069APending Publication Date: 2025-07-08上海和今信息科技有限公司
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510367407.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art fails to effectively utilize the similarity of prompt words when dealing with long-sequence tasks of large language models, resulting in increased latency and resource consumption of repeated calculations, especially the demand for GPU and video memory.

Method used

By converting the prompt words entered by the user into an embedded vector, the vector database is used to calculate semantic similarity, filtering similar prompt words, and storing the inference results in the Redis cache, combining PagedAttention technology to organize cross-requested KV cache in the video memory to reduce duplicate calculations.

Benefits of technology

It significantly reduces GPU computing requirements, improves inference performance, shortens response time, reduces redundant storage and computing burden, and improves the overall efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277069A_ABST
    Figure CN120277069A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of large language models, and discloses a fusion caching method, system and device based on LLM cue words and a medium. Cue words input by a user are converted into embedded vectors and stored in a vector database; on the basis of the vector database, the semantic similarity between the new cue word and the historical cue word is calculated through the vector retrieval technology, and similar cue words are screened; a reasoning result generated by the LLM large language model according to the cue word is stored in a Redis cache; performing retrieval based on the similar cue word, and quickly returning a corresponding reasoning result in the Redis cache; and splitting the cue word and carrying out fragmentation storage, and independently storing the corresponding KV Cache for each fragmentation. The technical problem that computing resources and video memory resources generated in the large language model reasoning process are wasted can be at least solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of large language models, and in particular, to a fusion caching method, system, device, and medium based on LLM prompts. Background Art

[0002] With the wide application of large language models (LLMs) in the field of natural language processing, their powerful generation capabilities have enabled such models to achieve remarkable success in a variety of tasks. However, LLMs, especially models based on the Transformer architecture, have very high requirements for computing and storage resources when dealing with long sequence tasks, especially for GPUs and video memory. During the inference process, each input prompt may be similar to previous prompts. However, existing technologies usually ignore this similarity, resulting in unnecessary repeated calculations, increased latency, and consumption of a large amount of resources.

[0003] In the prior art, when dealing with similar prompts, there is usually a lack of an effective caching mechanism, resulting in a significant increase in repeated calculations, latency, and resource consumption. At the same time, in large models based on the Transformer architecture, the KV Cache (key-value cache), as a key component for accelerating inference, can cache the intermediate results of model calculations to avoid repeated calculations. However, existing implementations are often limited to the reuse of KV Cache within a single request and fail to fully exploit the similarity characteristics between requests, resulting in video memory waste and performance bottlenecks. Summary of the Invention

[0004] An object of this application is to provide a fusion caching method, system, device, and medium based on LLM prompts, at least to solve the technical problem of waste of computing resources and video memory resources generated during the inference process of large language models.

[0005] To achieve the above object, some embodiments of this application provide the following aspects:

[0006] In a first aspect, some embodiments of the present application further provide a fusion caching method based on LLM prompts, including converting the prompts input by the user into embedding vectors and storing them in a vector database; calculating the semantic similarity between the new prompt and the historical prompts using vector retrieval technology based on the vector database, and screening similar prompts; storing the inference results generated by the LLM large language model according to the prompts in the Redis cache; retrieving based on the similar prompts to quickly return the corresponding inference results in the Redis cache; splitting the prompts and storing them in slices, and independently storing the corresponding KV Cache for each slice; when the LLM large language model performs inference, using the PagedAttention technology to organize and reuse the cross-request KV Cache in the video memory.

[0007] In a second aspect, some embodiments of the present application further provide a fusion caching system based on LLM prompts, including a prompt embedding and similarity calculation module for converting the prompts input by the user into embedding vectors and storing them in a vector database; calculating the semantic similarity between the new prompt and the historical prompts using vector retrieval technology based on the vector database, and screening similar prompts; a result caching and quick retrieval module for storing the inference results generated by the LLM large language model according to the prompts in the Redis cache; retrieving based on the similar prompts to quickly return the corresponding inference results in the Redis cache; an inference module based on structured prompts for splitting the prompts and storing them in slices, and independently storing the corresponding KV Cache for each slice; when the LLM large language model performs inference, using the PagedAttention technology to organize and reuse the cross-request KV Cache in the video memory.

[0008] In a third aspect, some embodiments of the present application further provide an electronic device, the electronic device includes: one or more processors; and a memory storing computer program instructions, the computer program instructions when executed cause the processor to execute the steps of the method described above.

[0009] In a fourth aspect, some embodiments of the present application further provide a computer-readable medium, on which computer program instructions are stored, the computer program instructions can be executed by a processor to implement the method described above.

[0010] Compared with the related technologies, in the solution provided by the embodiments of the present application, by combining a vector database and a Redis cache system, duplicate calculations are reduced, and the computing requirements of the GPU are significantly reduced; by leveraging the semantic similarity and structural characteristics of the prompt words and through cross-request KV Cache reuse, the inference performance is significantly improved, and the response time is shortened. By efficiently storing and retrieving the prompt words and inference results, redundant storage and computing burdens are reduced, and the overall efficiency of the system is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0012] Figure 1 FIG. [FIGURE NUMBER] is a schematic flowchart of a fusion caching method based on LLM prompt words provided according to an embodiment of the present application;

[0013] Figure 2 FIG. [FIGURE NUMBER] is a schematic architecture diagram of a fusion caching method based on LLM prompt words provided according to an embodiment of the present application;

[0014] Figure 3 FIG. [FIGURE NUMBER] is a schematic structural diagram of a fusion caching system based on LLM prompt words provided according to an embodiment of the present application;

[0015] Figure 4 FIG. [FIGURE NUMBER] is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0017] The following terms are used in this document:

[0018] Redis is an open-source in-memory data structure storage system that is commonly used as a database, cache, message broker, etc.;

[0019] Transformer is a deep learning model architecture based on the attention mechanism;

[0020] KV Cache key-value cache;

[0021] PagedAttention is a technique for optimizing the attention mechanism of large Transformer models.

[0022] First Embodiment

[0023] The first embodiment of this application relates to a fusion caching method based on LLM prompt words. As Figure 1 shown, the method may include the following steps:

[0024] S101, Convert the prompt words input by the user into embedding vectors and store them in the vector database. Receive the prompt words input by the user, which may be queries, commands, or conversation contents in natural language form. For example, the user inputs the prompt word "How to improve the performance of machine learning models?" Convert this prompt word into an embedding vector.

[0025] The process of converting the prompt words into embedding vectors can use existing natural language processing models (such as BERT, GPT, etc.) to encode the input prompt words and convert them into vector representations of a fixed dimension. This vector contains the semantic information of the prompt words. Store the generated vector in the vector database. The vector database such as FAISS (Facebook AI Similarity Search) or other efficient vector storage systems is used to store and quickly retrieve the embedding vectors of the prompt words.

[0026] S102, Calculate the semantic similarity between the new prompt words and the historical prompt words based on the vector database using vector retrieval technology, and filter out the similar prompt words. The newly received prompt words (such as "How to optimize deep learning algorithms?") will be converted into embedding vectors and retrieved in the vector database. The retrieval process calculates the similarity between the embedding vectors of the new prompt words and the historical prompt words.

[0027] Adopt common similarity calculation methods to calculate the semantic similarity between the new prompt words and the historical prompt words stored in the database. Set a threshold according to the similarity value, and select the historical prompt words that are semantically similar to the current prompt words. If the similarity exceeds the threshold, it is considered a similar prompt word. The similar prompt words will be filtered out for use in subsequent steps.

[0028] S103, Store the inference results generated by the LLM large language model according to the prompt words in the Redis cache. According to the user's prompt words, call the LLM large language model (such as GPT, T5, etc.) for inference to generate inference results. The inference results include the generated text, answers, or other response contents. Use the LLM model to perform inference on the input prompt words to generate results, and then store the inference results in the Redis cache. Use the embedding vector ID of the prompt words or the prompt words themselves as the key of the cache, and the inference results as the value.

[0029] Use the Redis database because its in-memory data storage and fast reading characteristics are very suitable for cache management. Redis stores inference results through key-value pairs and can quickly respond to subsequent query requests.

[0030] S104, retrieve based on the similar prompt words and quickly return the corresponding inference results in the Redis cache;

[0031] When a new prompt word request is received, it will first query from the Redis cache according to the similar prompt words filtered out in step S102. Use the similar prompt words as keys to query the Redis cache. If there is a matching inference result in the Redis cache, directly return the inference result in the cache without having to call the LLM model for inference again, thus reducing the computational burden and accelerating the response time.

[0032] S105, split the prompt words and store them in slices, and independently store the corresponding KV Cache for each slice; when the LLM large language model performs inference, use the PagedAttention technology to organize and reuse the KV Cache across requests in the video memory.

[0033] To further optimize the use of video memory and reduce the consumption of computing resources, split the prompt words into multiple parts for slice storage, and at the same time optimize the reuse of KV Cache across requests. Split the prompt words into a prefix, a suffix, and a middle variable part according to their structure. For example, the prompt word "How to improve the performance of machine learning models?" can be split into: prefix part "How to improve"; suffix part "the performance of the model"; middle part "machine learning". Each part (slice) will be stored and managed independently. For each slice of the prompt word, the corresponding key-value pair (KV) cache will be stored independently. The purpose of this move is to reduce the part that needs to be recalculated each time inference is performed.

[0034] PagedAttention technology: In the video memory, the PagedAttention technology can efficiently organize these cache data. By calculating the hash value for each Token ID, mark the KV Cache data blocks. When the hash values of the Token IDs at the same position in different requests are the same, the existing KV Cache blocks can be directly reused without having to recalculate. The PagedAttention technology enables these cache data across requests to be efficiently reused in the video memory, significantly reducing the GPU computational burden.

[0035] It is not difficult to find that, compared with the related technologies, in the solution provided by the embodiments of the present application, the cache system based on the vector library and Redis reduces the dependence on repeated calculations and lowers the GPU computing power requirements; by utilizing the semantic similarity and structural characteristics of the prompt words, as well as the reuse mechanism of the cross-request KV Cache, the inference efficiency is significantly improved; by optimizing the storage and retrieval of the prompt words and inference results, the burden of repeated calculations is reduced.

[0036] Second Embodiment

[0037] The second embodiment of the present application relates to a fusion caching method based on LLM prompt words. The second implementation is an improvement based on the first embodiment, and the specific improvement lies in:

[0038] Furthermore, the calculating the semantic similarity between the new prompt word and the historical prompt words by using the vector retrieval technology includes: calculating the semantic similarity between the new prompt word and the historical prompt words based on the cosine similarity or the Euclidean distance.

[0039] When calculating the similarity between the new prompt word and the historical prompt words, the cosine similarity or the Euclidean distance is used to compare the embedding vectors of the prompt words. Through these two distance measurement methods, the system can more accurately judge the semantic similarity of the prompt words, thereby optimizing the subsequent cache hit rate. The cosine similarity calculates the cosine value of the included angle between two vectors, and the value closer to 1 indicates more similarity. The Euclidean distance calculates the straight-line distance between two vectors, and the smaller the distance, the more similar.

[0040] Furthermore, the inference results in the Redis cache are indexed by the prompt word or the embedding vector ID corresponding to the prompt word.

[0041] The inference results not only use the prompt word itself as the cache key, but also can use the embedding vector ID corresponding to the prompt word as the index, thereby further improving the efficiency of cache query. By using the embedding vector ID as the cache key, the cache invalidation problem caused by the change of the prompt word can be effectively reduced. For example, for different prompt words with similar semantics, they can be associated with the same cache result through the vector ID, avoiding repeated calculations. Through this method, even if the user inputs prompt words that are semantically similar but have different texts, the cache can be hit and fast inference results can be obtained.

[0042] Furthermore, the splitting and sharding storage of the prompt words includes: splitting the prompt word into a prefix, a suffix, and an intermediate variable part.

[0043] When splitting prompt words, the granularity and quantity of splitting will be adjusted according to the length and structural complexity of the prompt words. For example, for a short and simple prompt word, it may only be split into two parts, the front and the back; while for a long and complex prompt word, it may be split into multiple finer-grained segments (for example, a prefix, an infix, and multiple suffix parts). Through this dynamic sharding method, the computational efficiency and storage burden can be more efficiently balanced.

[0044] Furthermore, the inference of the LLM large language model includes: marking the KV Cache by calculating the hash value of the Token ID. When the hash values of the Token IDs at the same position in different requests are the same, the corresponding KV Cache in the Block Table is directly reused.

[0045] By hashing each Token ID, it is possible to mark the Token in different requests and its corresponding KV Cache block. If the hash values of the Token IDs at the same position in different requests are the same, the KV Cache in the Block Table can be directly reused without having to perform calculations again. This method can significantly reduce the video memory occupancy and speed up the inference. For example, if the same Token ID of "deep learning model" is involved in the prompt words of multiple user requests, the KV Cache of this Token can be directly reused through the hash marking system without having to perform new calculations each time.

[0046] Furthermore, the method further includes: balancing the computational efficiency and storage burden by dynamically adjusting the quantity and granularity of the prompt word shards according to the length and complexity of the prompt words.

[0047] According to the length and complexity of the input prompt word, the quantity and granularity of the prompt word shards will be automatically adjusted. For example, for a simple short sentence, it may only be split into a few coarse-grained segments; while for a complex long sentence, a finer-grained sharding strategy will be adopted to balance the consumption of computational resources and the requirements of storage space.

[0048] The strategy of dynamically adjusting the sharding granularity ensures that under different prompt word inputs, the cache and inference resources can be efficiently used, redundant calculations are reduced, and at the same time, the storage overhead caused by excessive sharding is avoided.

[0049] It is not difficult to find that in the embodiments of the present application, the large language model based on the LLM can significantly reduce the consumption of computational resources while maintaining efficient inference, especially when dealing with complex long prompt words, and shows better performance.

[0050] Third Embodiment

[0051] The third embodiment of this application relates to a fusion caching method based on LLM prompts. The third embodiment is an improvement based on the first embodiment. The specific improvements are as follows:

[0052] As Figure 2 shown, receive the prompt word input by the user; then retrieve similar prompt words through the vector database, and judge whether the inference result of the LLM large model corresponding to the prompt word already exists in the Redis cache; if the cache exists, directly return the inference result from Redis; if the cache does not exist, then perform inference through the LLM large language model, generate the result and store it in the Redis cache, and finally return the result. Specifically:

[0053] The user inputs a prompt word (such as "How to optimize a machine learning model?") through the front-end interface, receive the prompt word and process it. Use the vector database to perform similarity retrieval of prompt words, convert the prompt word input by the user into an embedding vector, and store it in the vector database. Then, use vector retrieval techniques (such as cosine similarity or Euclidean distance) to find the prompt words of historical inputs and calculate the similarity between the input prompt word and the historical prompt words.

[0054] If the retrieval result shows that the similarity between the current prompt word and the historical prompt words exceeds the preset threshold, it is considered that the prompt word is similar to the historical prompt words. According to the similarity retrieval result, judge whether the inference result corresponding to the current prompt word exists in the Redis cache. At this time, if the similar prompt word already exists in the Redis cache, directly obtain the result from the cache; if the cache is not hit, perform a new inference through the LLM large language model.

[0055] Obtain the result from the Redis cache. If the prompt word is already in the cache, directly obtain the corresponding inference result from the Redis cache. At this time, the Redis cache stores the inference results based on historical prompt words and can be directly returned to the user, avoiding the need to perform inference operations again and greatly reducing the consumption of computing resources.

[0056] If the prompt word does not hit the cache, call the LLM large language model to perform inference. The LLM generates an inference result based on the input prompt word. The inference result can be text, answer, or other forms of response content.

[0057] Store the inference result generated by the LLM in the Redis cache, using the prompt word or its embedding vector ID as the cache key and the inference result as the value. In this way, when a similar prompt word is requested again, the result can be directly obtained from the cache, improving the response speed and system efficiency.

[0058] Finally, return the inference result (whether obtained from the cache or generated by the LLM model inference) to the user to complete the response.

[0059] It should be noted that the third embodiment of this application can also be an improvement based on the second embodiment.

[0060] It is not difficult to find that in the embodiments of this application, by caching the inference results in Redis, the repeated calculation of multiple identical prompt words is avoided, significantly improving the inference efficiency and the system response speed. In the similarity retrieval of prompt words, cache hits can be accurately judged, reducing unnecessary consumption of computing resources. Especially when dealing with similar prompt words, results can be quickly returned, reducing the burden on the GPU and video memory. Through vector retrieval and cache mechanisms, diverse prompt word inputs can be handled, and at the same time, with the accumulation of the system cache, the overall response performance will continue to improve.

[0061] The step divisions of the above various methods are only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are all within the protection scope of this application.

[0062] Fourth Embodiment

[0063] The fourth embodiment of this application relates to a fusion cache system / apparatus based on LLM prompt words, as Figure 3 shown. The system includes:

[0064] A prompt word embedding and similarity calculation module, which is used to convert the prompt word input by the user into an embedding vector and store it in the vector database; calculate the semantic similarity between the new prompt word and the historical prompt words using vector retrieval technology based on the vector database, and screen for similar prompt words;

[0065] A result cache and fast retrieval module, which is used to store the inference results generated by the LLM large language model according to the prompt words in the Redis cache; retrieve based on the similar prompt words and quickly return the corresponding inference results in the Redis cache;

[0066] An inference module based on structured prompt words, which is used to split the prompt words and store them in slices, and independently store the corresponding KV Cache for each slice; when the LLM large language model performs inference, use the PagedAttention technology to organize and reuse the cross-request KV Cache in the video memory.

[0067] It is not difficult to find that this embodiment is a system embodiment corresponding to the first embodiment, and this embodiment can be implemented in cooperation with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.

[0068] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0069] In addition, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0070] The electronic device includes: one or more processors; and a memory storing computer program instructions, which when executed cause the processor to perform the steps of the method provided in any one or more of the above embodiments. Figure 4 An exemplary structural diagram of the electronic device is disclosed. As Figure 4 shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component is interconnected using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Among them, the components, their connections and relationships, and their functions shown herein are only examples and are not intended to limit the implementation of this application described and / or claimed herein.

[0071] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 may be connected through a bus or other means. Figure 4 Taking the connection through the bus as an example.

[0072] The input device 1103 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 1104 may include a display device, an auxiliary lighting device (e.g., LED), and a haptic feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touchscreen.

[0073] To provide interaction with the user, the electronic device may be a computer. The computer has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and the input from the user may be received in any form (including acoustic input, voice input, or haptic input).

[0074] In the embodiments of the present application, a computer program / instructions is stored on a computer-readable medium. When the computer program / instructions are executed by the processor, the steps of the method provided in any one or more of the above embodiments are implemented. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0075] The memory 1102 may be used as a non-transitory computer-readable storage medium, and may be used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.

[0076] The memory 1102 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device and the like. In addition, the memory 1102 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely disposed relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0077] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0078] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0079] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0080] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of this application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.

[0081] The computer program product provided by the embodiments of this application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they wholly or partly generate the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available media can be magnetic media (e.g., floppy disk, hard disk, magnetic tape), optical media (e.g., DVD), or semiconductor media (e.g., solid state disk (SSD)), etc.

[0082] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted in the blocks may occur in a different order than that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0083] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present application. Any reference numerals in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices recited in the apparatus claims may also be implemented by one element or device through software or hardware. The words "first", "second", etc. are only used for distinguishing descriptions and do not indicate any particular order, nor can they be construed as indicating or implying relative importance.

[0084] As described above, these are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily make changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A fusion caching method based on LLM prompts, characterized in that The method includes: Converting the prompt word input by the user into an embedding vector and storing it in a vector database; calculating the semantic similarity between the new prompt word and the historical prompt words using vector retrieval technology based on the vector database, and screening for similar prompt words; Storing the inference results generated by the LLM large language model according to the prompt words in a Redis cache; retrieving based on the similar prompt words and quickly returning the corresponding inference results in the Redis cache; Splitting the prompt word and storing it in fragments, and independently storing the corresponding KV Cache for each fragment; when the LLM large language model performs inference, using the PagedAttention technology to organize and reuse the cross-request KV Cache in the video memory.

2. The method according to claim 1, wherein The calculating the semantic similarity between the new prompt word and the historical prompt words using vector retrieval technology includes: calculating the semantic similarity between the new prompt word and the historical prompt words based on cosine similarity or Euclidean distance.

3. The method according to claim 1, wherein The inference results in the Redis cache use the prompt word or the embedding vector ID corresponding to the prompt word as the key index.

4. The method according to claim 3, wherein The splitting the prompt word and storing it in fragments includes: splitting the prompt word into a prefix, a suffix, and a variable middle part.

5. The method according to claim 4, characterized in that, The LLM large language model performing inference includes: Marking the KV Cache by calculating the hash value of the Token ID. When the hash values of the Token IDs at the same position in different requests are the same, directly reuse the corresponding KV Cache in the Block Table.

6. The method according to claim 1, characterized in that, The method further includes: By dynamically adjusting the number and granularity of the prompt word fragments, balancing the calculation efficiency and storage burden according to the length and complexity of the prompt word.

7. A fusion caching system based on LLM prompts, characterized in that, The system includes: A prompt word embedding and similarity calculation module, configured to convert the prompt word input by the user into an embedding vector and store it in a vector database; calculate the semantic similarity between the new prompt word and the historical prompt words using vector retrieval technology based on the vector database, and screen for similar prompt words; A result cache and fast retrieval module, configured to store the inference results generated by the LLM large language model according to the prompt words in a Redis cache; retrieve based on the similar prompt words and quickly return the corresponding inference results in the Redis cache; An inference module based on structured prompt words, configured to split the prompt word and store it in fragments, and independently store the corresponding KV Cache for each fragment; when the LLM large language model performs inference, use the PagedAttention technology to organize and reuse the cross-request KV Cache in the video memory.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, where the computer program instructions, when executed, cause the processor to execute the steps of the method according to any one of claims 1 to 6.

9. A computer-readable medium having computer programs / instructions stored thereon, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Method and device for generating reasoning result, storage medium and electronic equipment

    CN120930810A

  • Method and device for generating reasoning result, storage medium and electronic device

    CN120930810B

  • Multi-level cache access method, system and equipment and storage medium

    CN121387772A

  • Large model reasoning control method and device, equipment and medium

    CN121581247A