Key value pair caching method, system and equipment of large model and storage medium
By detecting dialogue modes in the big model and establishing an adaptive cardinal tree to store key-value pair data, and dynamically adjusting the tree nodes, the problem of the large model requirements in the existing technology that are difficult to compatible with different dialogue modes is solved, and caching efficiency and inference performance are improved.
Patent Information
- Application Number
- CN202510345128.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
The key-value caching technology of existing large models is difficult to compatible with the requirements of large models of different dialogue modes, resulting in performance impacts and low resource utilization.
By detecting the dialogue mode of the large model, an adaptive cardinality tree is established to store key-value pair data, and dynamically adjust the tree nodes according to the dialogue mode and data volume to improve cache efficiency.
It effectively improves the key-value pair cache efficiency of large models, is suitable for large models of various dialogue modes, and improves inference performance and storage resource utilization.
Smart Images

Figure CN120197676A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large models, and in particular, to a key-value pair caching method, system, device, and storage medium for large models. Background Art
[0002] In the field of medical and health, the application of large model technology is gradually becoming popular, mainly due to its powerful data processing and analysis capabilities. A large language model (LLM) is a language model that is trained with a large amount of text data and has a very high number of parameters. These large models can perform well in various natural language processing tasks, such as text generation, translation, question answering, and summary generation. For example, in the field of medical and health, online medical consultations can be realized through large model technology, or corresponding diagnosis and treatment plans can be generated based on the personalized information of patients, thereby improving the efficiency of medical consultations.
[0003] In related technologies, key-value (KV) caching is a commonly used optimization technology in the inference process of large models. By caching the previously calculated results, repeated calculations can be reduced, thereby accelerating the inference speed and reducing the calculation cost. Key-value pair caching needs to store a large amount of intermediate results. Especially for large-scale language models, these intermediate results may be very large. As the scale of the large model increases, the memory required for caching will also increase sharply, and the requirements for caching performance are relatively high. Moreover, the conversation modes of different large models are different, and the requirements for key-value pair caching also vary. In current applications, corresponding caching strategies are often preset, and a tree-like data structure (such as B-tree, Trie, etc.) is established to implement the caching of key-value pair data. This application method is difficult to be compatible with the requirements of various different large models, resulting in the performance of the large model being affected and the resource utilization rate being low.
[0004] In summary, the problems existing in the related technologies need to be solved urgently. Summary of the Invention
[0005] The purpose of this application is to solve at least one of the technical problems existing in the related technologies to a certain extent.
[0006] To this end, an object of an embodiment of this application is to provide a key-value pair caching method for large models. This method can effectively improve the caching efficiency of key-value pairs of large models, is applicable to large models with various conversation modes, is beneficial to improving the inference performance of large models, and improves the utilization rate of storage resources.
[0007] To achieve the above technical purpose, the technical solutions adopted in the embodiments of this application include:
[0008] On the one hand, an embodiment of this application provides a key-value pair caching method for large models, including:
[0009] Obtain the inference service data of multiple large models, and convert the inference service data into key-value pair data;
[0010] Detect the dialogue modes of each of the large models, establish a corresponding adaptive radix tree for each of the large models, and store the key-value pair data corresponding to the large model through the adaptive radix tree; wherein, the adaptive radix tree includes multiple nodes;
[0011] Detect the data volume of the inference service data of each of the large models;
[0012] Dynamically adjust the nodes of the adaptive radix tree corresponding to each of the large models according to the dialogue mode and the data volume.
[0013] In addition, according to the key-value pair caching method of the large model in the above embodiments of the present application, the following additional technical features may also be included:
[0014] Further, in an embodiment of the present application, the detecting the dialogue modes of each of the large models, establishing a corresponding adaptive radix tree for each of the large models, and storing the key-value pair data corresponding to the large model through the adaptive radix tree includes:
[0015] If the dialogue mode of the large model is the few-shot learning mode, build a first adaptive radix tree including multiple compact nodes, and store the key-value pair data corresponding to the large model in the few-shot learning mode through the first adaptive radix tree;
[0016] If the dialogue mode of the large model is the multi-turn dialogue mode, build a second adaptive radix tree including multiple leaf nodes, and store the key-value pair data corresponding to the large model in the multi-turn dialogue mode through the second adaptive radix tree;
[0017] If the dialogue mode of the large model is the decision tree mode, build a third adaptive radix tree including leaf nodes, internal nodes and compact nodes, and store the key-value pair data corresponding to the large model in the decision tree mode through the third adaptive radix tree.
[0018] Further, in an embodiment of the present application, the dynamically adjusting the nodes of the adaptive radix tree corresponding to each of the large models according to the dialogue mode and the data volume includes:
[0019] If the dialogue mode of the large model is the few-shot learning mode, detect whether the size of the data volume exceeds a first preset threshold;
[0020] If the size of the data volume exceeds the first preset threshold, at least one of the compact nodes in the first adaptive radix tree is converted into an internal node.
[0021] Further, in an embodiment of the present application, the dynamically adjusting the nodes of the adaptive radix tree corresponding to each of the large models according to the conversation mode and the data volume includes:
[0022] If the conversation mode of the large model is a multi-round conversation mode, detect whether the size of the data volume exceeds a second preset threshold;
[0023] If the size of the data volume exceeds the second preset threshold, create at least one leaf node in the second adaptive radix tree.
[0024] Further, in an embodiment of the present application, the method further includes:
[0025] If the conversation mode of the large model is a few-shot learning mode, detect the similarity of the key-value pair data existing in the adaptive radix tree;
[0026] Compare the size of the similarity and a preset similarity threshold;
[0027] If the similarity is greater than the similarity threshold, perform path compression processing on the corresponding key-value pair data.
[0028] Further, in an embodiment of the present application, the detecting the conversation mode of each of the large models and establishing a corresponding adaptive radix tree for each of the large models includes:
[0029] If the conversation mode of the large model is a few-shot learning mode, build an adaptive radix tree with a survival period of a first duration;
[0030] If the conversation mode of the large model is a multi-round conversation mode, build an adaptive radix tree with a survival period of a second duration;
[0031] Wherein, the first duration is greater than the second duration.
[0032] Further, in an embodiment of the present application, the method further includes:
[0033] Detect the historical call times of each of the adaptive radix trees and the historical time node of the last call;
[0034] If the historical call times of the adaptive radix tree are lower than a preset number threshold, and the duration from the historical time node to the current time node is greater than a preset duration threshold, perform deletion processing on the adaptive radix tree.
[0035] On the other hand, an embodiment of the present application also provides a key-value pair caching system for a large model, including:
[0036] An acquisition unit, configured to acquire inference service data of multiple large models, and convert the inference service data into key-value pair data;
[0037] A building unit, configured to detect the dialogue modes of the respective large models, establish a corresponding adaptive radix tree for each of the large models, and store the key-value pair data corresponding to the large model through the adaptive radix tree; wherein, the adaptive radix tree includes multiple nodes;
[0038] A detection unit, configured to detect the data volume of the inference service data of each of the large models;
[0039] An adjustment unit, configured to dynamically adjust the nodes of the adaptive radix tree corresponding to each of the large models according to the dialogue mode and the data volume.
[0040] On the other hand, an embodiment of the present application provides a computer device, including:
[0041] At least one processor;
[0042] At least one memory, configured to store at least one program;
[0043] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned key-value pair caching method for the large model.
[0044] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above-mentioned key-value pair caching method for the large model when executed by the processor.
[0045] The advantages and beneficial effects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present application:
[0046] A key-value pair caching method for a large model disclosed in an embodiment of the present application obtains inference service data of multiple large models, converts the inference service data into key-value pair data; detects the conversation modes of each large model, establishes a corresponding adaptive radix tree for each large model, and stores the key-value pair data corresponding to the large model through the adaptive radix tree; wherein, the adaptive radix tree includes multiple nodes; detects the data volume of the inference service data of each large model; dynamically adjusts the nodes of the adaptive radix tree corresponding to each large model according to the conversation mode and the data volume. The method in the embodiment of the present application can effectively improve the caching efficiency of the key-value pairs of the large model, is applicable to large models of various conversation modes, is beneficial to improving the inference performance of the large model, and improves the utilization rate of storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present application or the prior art. It should be understood that the accompanying drawings below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 Schematic diagram of the implementation environment of a key-value pair caching method for a large model provided in an embodiment of the present application;
[0049] Figure 2 Schematic diagram of the flow of a key-value pair caching method for a large model provided in an embodiment of the present application;
[0050] Figure 3 Schematic diagram of the flow of establishing an adaptive radix tree provided in an embodiment of the present application;
[0051] Figure 4 Schematic diagram of the flow of dynamically adjusting the nodes of an adaptive radix tree provided in an embodiment of the present application;
[0052] Figure 5 Schematic diagram of the flow of another dynamic adjustment of the nodes of an adaptive radix tree provided in an embodiment of the present application;
[0053] Figure 6 Schematic diagram of the flow of deleting an adaptive radix tree provided in an embodiment of the present application;
[0054] Figure 7 Schematic diagram of the structure of a key-value pair caching system for a large model provided in an embodiment of the present application;
[0055] Figure 8This is a schematic structural diagram of a computer device provided in an embodiment of the present application. Specific embodiments
[0056] The present application will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0057] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0059] First, several nouns involved in the present application are analyzed:
[0060] 1) Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0061] 2) Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning (deep learning) usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0062] 3) LLM (Large Language Model), a large language model, refers to those language models that are trained with a large amount of text data and have an extremely high number of parameters, also known as large models. These large models can perform well in various natural language processing tasks, such as text generation, translation, question answering, and summary generation.
[0063] 4) Key-Value Pair, a common data structure used to store and retrieve data. In this structure, each data item consists of a unique key and its associated value.
[0064] 5) Adaptive Radix Tree (ART), an efficient data structure used to store and retrieve key-value pair data, especially suitable for in-memory data indexing.
[0065] In the field of healthcare, the application of large model technology is gradually becoming popular, mainly due to its powerful data processing and analysis capabilities. The large model (Large Language Model, LLM) is a large language model that refers to those language models that are trained with a large amount of text data and have an extremely high number of parameters, also known as large models. These large models can perform well in various natural language processing tasks, such as text generation, translation, question answering, and summary generation. For example, in the field of healthcare, online consultations can be achieved through large model technology, or corresponding diagnosis and treatment plans can be generated based on the personalized information of patients using large model technology, thereby improving the efficiency of medical consultations.
[0066] In the related art, the key-value (KV) cache is an optimization technique commonly used in the inference process of large models. By caching the previously computed results, it reduces repeated computations, thereby accelerating the inference speed and lowering the computational cost. The key-value cache needs to store a large amount of intermediate results. Especially for large-scale language models, these intermediate results can be extremely large. As the scale of the large model increases, the memory required for the cache also increases sharply, posing high requirements for cache performance. Moreover, the conversation patterns of different large models are different, and there are also differences in the requirements for the key-value cache. In current applications, corresponding cache strategies are often preset, and a tree-shaped data structure (such as B-tree, Trie, etc.) is established to implement the caching of key-value data. This application method is difficult to be compatible with the requirements of various different large models, resulting in the performance of the large model being affected and the resource utilization rate being low.
[0067] To solve the problems existing in the related art, the embodiments of the present application provide a method, a system, a device, and a storage medium for key-value caching of large models. The method includes obtaining the inference service data of multiple large models and converting the inference service data into key-value data; detecting the conversation patterns of each of the large models and establishing a corresponding adaptive radix tree for each of the large models, and storing the key-value data corresponding to the large model through the adaptive radix tree; wherein the adaptive radix tree includes multiple nodes; detecting the data volume of the inference service data of each of the large models; and dynamically adjusting the nodes of the adaptive radix tree corresponding to each of the large models according to the conversation pattern and the data volume. The method in the embodiments of the present application can effectively improve the caching efficiency of the key-value pairs of large models, is applicable to large models with various conversation patterns, is beneficial to improving the inference performance of large models, and improves the utilization rate of storage resources. When applying the above large model in the field of medical and health, it can better assist doctors in disease diagnosis and prediction, help patients obtain accurate and reliable medical advice, and reduce the implementation cost of applying large models.
[0068] The method for key-value caching of large models provided in the embodiments of the present application can be executed in various application scenarios in the field of medical and health:
[0069] Exemplarily, in some embodiments, the method for key-value caching of large models in the embodiments of the present application can be applied to the scenario of medical consultation. For example, through a large model, medical consultation services can be provided, such as building a chatbot to provide basic medical consultation and advice. In this scenario, the method for key-value caching of large models provided in the embodiments of the present application can be applied to implement the caching of inference service data related to the large model, thereby improving the application performance of the large model.
[0070] Exemplarily, in some embodiments, the key-value pair caching method of the large model in the embodiments of the present application can be applied to the scenario of medical record query. For example, in the medical record storage databases of each medical institution, a large model can be built to achieve more accurate search and provide the corresponding information to relevant users (such as doctors or patients, etc.). In this scenario, the key-value pair caching method of the large model provided in the embodiments of the present application can be applied to implement the caching of inference service data related to the large model, thereby improving the application performance of the large model.
[0071] Of course, it should be noted that the above application scenarios only serve as examples and do not mean to limit the actual application of the method in the embodiments of the present application. Those skilled in the art can understand that in different application scenarios, the method provided in the embodiments of the present application can be used to perform specified tasks.
[0072] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the implementation environment of the key-value pair caching method of the large model provided in the embodiments of the present application. The main software and hardware entities of this implementation environment mainly include a terminal device 110 and a server 120, and the terminal device 110 is communicatively connected to the server 120. Among them, the key-value pair caching method of the large model can be configured to be executed on the server 120 side and implemented based on the data interaction between the terminal device 110 and the server 120. For example, relevant large model applications can be deployed on the terminal device 110, and the server 120 can be the background server of this application.
[0073] Specifically, the terminal device 110 in the present application can include, but is not limited to, any one or more of a smart watch, a smart phone, a computer, a personal digital assistant (PDA), a smart voice interaction device, a smart home appliance, or a vehicle-mounted terminal. The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0074] A communication connection can be established between the terminal device 110 and the server 120 through a wireless network or a wired network. The wireless network or the wired network uses standard communication technologies and / or protocols. The network can be set as the Internet or any other network, such as any combination including but not limited to a Local Area Network (LAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a mobile, wired or wireless network, a private network or a virtual private network.
[0075] Of course, it can be understood that Figure 1 the implementation environment in Figure 1 is only an optional application scenario of the key-value pair caching method for the large model provided in the embodiments of the present application. The actual application is not fixed to the
[0076] Next, in combination with Figure 1 the implementation environment shown, the key-value pair caching method for the large model provided in the embodiments of the present application will be described in detail.
[0077] First, please refer to Figure 2 , Figure 2 which is a schematic flowchart of the key-value pair caching method for the large model provided in the embodiments of the present application. Figure 2 The key-value pair caching method shown can be applied to relevant computer devices in the server 120, but is not limited to the above form. Figure 2 The method in
[0078] Step 210: Obtain inference service data of multiple large models, and convert the inference service data into key-value pair data;
[0079] Step 220: Detect the conversation modes of the respective large models, establish a corresponding adaptive radix tree for each of the large models, and store the key-value pair data corresponding to the large model through the adaptive radix tree; wherein, the adaptive radix tree includes multiple nodes;
[0080] Step 230: Detect the data volume of the inference service data of each of the large models;
[0081] Step 240: Dynamically adjust the nodes of the adaptive radix tree corresponding to each of the large models according to the conversation mode and the data volume.
[0082] In the embodiments of the present application, a key-value pair caching method for large models is provided. This method can effectively improve the caching efficiency of key-value pairs in large models, is applicable to large models of various dialogue modes, is beneficial to improving the inference performance of large models, and improves the utilization rate of storage resources. When applying the above large model in the field of medical and health, it can better assist doctors in disease diagnosis and prediction, help patients obtain accurate and reliable medical advice, and reduce the implementation cost of applying large models.
[0083] Specifically, when executing the key-value pair caching method for large models provided in the embodiments of the present application, first, in step 210, inference service data can be collected from each large model. Here, the inference service data refers to the service data involved when the large model executes relevant inference tasks. For example, in some embodiments, the inference service data may include question information proposed by users and reply information given by the large model. For example, the question information may be the symptoms of certain diseases described by users, and the reply information given by the large model may be the type of disease and related diagnosis and treatment suggestions, etc. Specifically, when collecting inference service data, the original data in the logs corresponding to each large model can be obtained. After collecting the original data, it needs to be parsed to extract useful fields. For example, regular expressions or other text parsing tools can be used to identify and extract specific data.
[0084] In step 210, after obtaining the inference service data of the large model, it can be converted into key-value pair format data, which is denoted as key-value pair data in the embodiments of the present application. Each key-value pair data consists of a key and a value. The key is usually of string type and is used to identify the specific meaning of the data; the value can be of any data type, depending on the business requirements. Exemplarily, the value in the inference service data of the large model can be the dialogue information generated during a medical consultation, and the key can be the data identifying this dialogue information.
[0085] It should be noted that in the embodiments of the present application, the number of each large model involved can be any number, and the present application does not limit the dialogue modes of these large models.
[0086] In step 220, the dialogue modes of each large model can be detected, and corresponding adaptive radix trees can be established for each large model according to their corresponding dialogue modes. In the embodiments of the present application, the adaptive radix tree used combines the advantages of multiple tree structures, such as B-trees, Trie trees, etc., and can provide high-performance insertion, search, and deletion operations on dynamically changing data sets. It can be understood that for different large models, the specific structures of the corresponding adaptive radix trees may be different, and the present application does not limit this.
[0087] In the embodiments of the present application, for each adaptive radix tree, it generally includes multiple nodes. There are various types of nodes in the adaptive radix tree. Exemplarily, the common node types and corresponding functions are as follows: 1. Leaf Node: It can be used to directly store key-value pair data and is suitable for storing specific cache data. 2. Internal Node: It is used to fork paths. Each child node corresponds to a prefix. An internal node usually contains multiple child nodes, and each child node points to a node at the next level. 3. Compact Node: When the prefix is long and unique, the prefix can be compressed into one node to reduce the height of the tree. The compact node is suitable for the case of long keys and can significantly reduce the storage overhead. 4. Hybrid Node: It combines the characteristics of leaf nodes and internal nodes and can be dynamically converted when needed. For example, when the number of key-value pair data under a certain node is small, it can be processed as a compact node; when the number increases, it can be converted into an internal node.
[0088] Of course, it can be understood that the node types that can be included in the actual adaptive radix tree are not limited to the above cases, and the present application does not make any restrictions on this.
[0089] In the embodiments of the present application, the key-value pair data corresponding to the large model can be stored through the adaptive radix tree. Specifically, a database that supports the adaptive radix tree can be selected, such as libart, JavaART, etc. An ART tree instance is initialized in the database, that is, the root node of the tree is created. Then, the key-value pair data corresponding to the large model can be inserted into the adaptive radix tree. The key of each key-value pair data is unique to ensure the correct storage and retrieval of the data.
[0090] In step 230, the data volume of the inference service data of each large model can be detected. Specifically, for each large model, it can be monitored while generating the service data, and the corresponding data volume can be counted. Exemplarily, in some embodiments, a logging function can be added to the inference service of each large model to record information such as the input and output data sizes of each inference. Then, the input and output data of each inference are accumulated to obtain the total data volume. For example, taking the large model used to generate medical consultation data as an example, the data volume of the medical consultation data output by the large model each time can be counted, and these data volumes are accumulated to obtain the data volume of the inference service data of the large model.
[0091] In step 240, the nodes of the adaptive radix tree can be dynamically adjusted according to the dialogue mode and data volume of the large model. Specifically, it can be understood that the adaptive radix tree itself can optimize performance by dynamically adjusting the node type. In the embodiments of the present application, the node adjustment of the adaptive radix tree can be achieved based on the dialogue mode and data volume of the large model. For example, in some embodiments, if the dialogue mode and data volume of the large model indicate that there are many high-frequency short conversations, that is, the data volume is small each time, but the requests are frequent. In this case, the branching of the internal nodes in the adaptive radix tree can be optimized to reduce the search path length and improve the search efficiency. In some embodiments, if the dialogue mode and data volume of the large model indicate that it has a large data volume but few requests, the storage of the leaf nodes in the adaptive radix tree can be optimized to reduce memory occupancy and improve data storage efficiency.
[0092] Of course, in the embodiments of the present application, the strategy for dynamically adjusting the nodes of the adaptive radix tree is not limited, and it can be flexibly set according to actual needs. In the embodiments of the present application, according to the dialogue mode and data volume of the large model, real-time monitoring and adjustment can be achieved to adapt to changing needs, thereby improving the overall performance of the large model.
[0093] It can be understood that a key-value pair caching method for a large model provided in the embodiments of the present application includes obtaining inference service data of multiple large models, converting the inference service data into key-value pair data; detecting the dialogue mode of each large model, establishing a corresponding adaptive radix tree for each large model, and storing the key-value pair data corresponding to the large model through the adaptive radix tree; wherein, the adaptive radix tree includes multiple nodes; detecting the data volume of the inference service data of each large model; and dynamically adjusting the nodes of the adaptive radix tree corresponding to each large model according to the dialogue mode and the data volume. The method in the embodiments of the present application can effectively improve the caching efficiency of the key-value pairs of the large model, is applicable to large models with various dialogue modes, is beneficial to improving the inference performance of the large model, and improves the utilization rate of storage resources.
[0094] Specifically, in some embodiments, referring to Figure 3 , the detecting the dialogue mode of each large model, establishing a corresponding adaptive radix tree for each large model, and storing the key-value pair data corresponding to the large model through the adaptive radix tree includes:
[0095] If the dialogue mode of the large model is the few-shot learning mode, build a first adaptive radix tree including multiple compact nodes, and store the key-value pair data corresponding to the large model in the few-shot learning mode through the first adaptive radix tree;
[0096] If the dialogue mode of the large model is a multi-round dialogue mode, construct a second adaptive radix tree including multiple leaf nodes, and store the key-value pair data corresponding to the large model in the multi-round dialogue mode through the second adaptive radix tree;
[0097] If the dialogue mode of the large model is a decision tree mode, construct a third adaptive radix tree including leaf nodes, internal nodes, and compact nodes, and store the key-value pair data corresponding to the large model in the decision tree mode through the third adaptive radix tree.
[0098] In the embodiments of the present application, when constructing the adaptive radix tree corresponding to the large model, the adaptive radix tree structure suitable for different dialogue models can be constructed in combination with the dialogue mode of the large model. Exemplarily, in some embodiments, the dialogue mode of the large model may be the few-shot learning mode. In the few-shot learning mode, the prompt and few-shot examples are usually fixed, and relatively speaking, compact nodes can be used to store its business data. Therefore, in the embodiments of the present application, for the large model in the few-shot learning mode, an adaptive radix tree including multiple compact nodes can be constructed, which is denoted as the first adaptive radix tree in the embodiments of the present application, and then the key-value pair data corresponding to the large model in the few-shot learning mode can be stored through the first adaptive radix tree.
[0099] In some embodiments, the dialogue mode of the large model may be a multi-round dialogue mode. In the multi-round dialogue mode, the relative data volume is much larger. It is very likely that except for the initial chat-prompt, the subsequent content will be different. Therefore, in the embodiments of the present application, an adaptive radix tree including multiple leaf nodes can be established, which is denoted as the second adaptive radix tree. Through the leaf nodes in the second adaptive radix tree, the initial chat-prompt data of the large model in the multi-round dialogue mode can be stored. Subsequently, according to the situation of the business data of the large model, the second adaptive radix tree can be adjusted accordingly to support the business requirements of multi-round dialogues.
[0100] In some embodiments, the dialogue mode of the large model may be the decision tree mode (Tree-of-Thought mode), and the decision tree mode simulates the human thinking process by constructing a decision tree. In this mode, the model gradually decomposes the problem, generates multiple branches, and performs reasoning on each branch, and finally selects the optimal path. Its characteristic is that the model generates multiple branches, each branch represents a possible thinking path, and recursive reasoning is performed on each branch to gradually approach the solution to the problem. Therefore, when building an adaptive radix tree corresponding to the large model in the decision tree mode, an adaptive radix tree including leaf nodes, internal nodes, and compact nodes can be built, denoted as the third adaptive radix tree. The leaf nodes therein can be used to store the initial dialogue context and the first candidate response generated. If the initial context is long and unique, compact nodes can be used to store it. The internal nodes can coordinate the path relationships between the nodes to facilitate the implementation of reasoning.
[0101] Of course, it can be understood that the dialogue mode of the large model involved in the embodiments of the present application is not limited to the above-given examples. For the dialogue modes of various large models, corresponding adaptive radix trees can be built, and in the embodiments of the present application, no limitation is imposed on the specific structure of the adaptive radix tree.
[0102] Specifically, in some embodiments, referring to Figure 4 , the dynamically adjusting the nodes of the adaptive radix tree corresponding to each large model according to the dialogue mode and the data volume includes:
[0103] If the dialogue mode of the large model is the few-shot learning mode, detecting whether the size of the data volume exceeds a first preset threshold;
[0104] If the size of the data volume exceeds the first preset threshold, converting at least one of the compact nodes in the first adaptive radix tree into an internal node.
[0105] In the embodiments of the present application, for a large model with a few-shot learning mode, when dynamically adjusting the nodes of its adaptive radix tree, it can be detected whether the size of its data volume exceeds a first preset threshold. The size of this first preset threshold can be flexibly set, and the present application does not limit this.
[0106] If it is found that the size of the data volume of a large model in the few-shot learning mode exceeds the first preset threshold, the type of the nodes of its adaptive radix tree can be dynamically adjusted. For example, some compact nodes can be converted into internal nodes, which can better support the business of the large model. The specific number of the converted nodes is not limited in the present application.
[0107] Specifically, in some embodiments, referring to Figure 5, dynamically adjusting the nodes of the adaptive radix tree corresponding to each large model according to the dialogue mode and the data volume, including:
[0108] If the dialogue mode of the large model is a multi-turn dialogue mode, detect whether the size of the data volume exceeds a second preset threshold;
[0109] If the size of the data volume exceeds the second preset threshold, create at least one new leaf node in the second adaptive radix tree.
[0110] In the embodiments of the present application, for a large model with a multi-turn dialogue mode, when dynamically adjusting the nodes of its adaptive radix tree, it is possible to detect whether the size of its data volume exceeds a second preset threshold. The size of this second preset threshold can be flexibly set, and the present application does not limit this.
[0111] If it is found that the size of the data volume of a large model with a multi-turn dialogue mode exceeds the second preset threshold, the number of nodes of its adaptive radix tree can be dynamically adjusted. For example, at least one new leaf node can be created.
[0112] Of course, it should be noted that in the embodiments of the present application, for a large model with a multi-turn dialogue mode, during adjustment, it is also possible to perform type conversion of its nodes. For example, if the prefix of a certain node appears frequently, it can be converted into a compact node to save space; if the prefix is short and has high diversity, it can be converted into an internal node for branching, and the present application does not limit this.
[0113] Specifically, in some embodiments, the method further includes:
[0114] If the dialogue mode of the large model is a few-shot learning mode, detect the similarity of the key-value pair data existing in the adaptive radix tree;
[0115] Compare the size of the similarity and a preset similarity threshold;
[0116] If the similarity is greater than the similarity threshold, perform path compression processing on the corresponding key-value pair data.
[0117] In the embodiments of the present application, for a large model with a few-shot learning mode as the dialogue mode, path compression and lazy expansion can also be performed on it. Specifically, the similarity of the existing key-value pair data in its corresponding adaptive radix tree can be detected, and the similarity is compared with a preset similarity threshold. If it is found that the similarity between some key-value pair data is greater than the similarity threshold, path compression processing can be performed on them. Path compression can perform redundancy processing on similar key-value pair data (i.e., few-shot examples). In the embodiments of the present application, when processing key-value pair data, the idea of dynamic programming can be adopted to save intermediate states and avoid repeated calculations. For example, in dialogue generation, the state information of the dialogue history can be saved. When the subsequent dialogue content overlaps with the previous content, the existing state information can be directly used to continue generation without starting from scratch.
[0118] It can be understood that through path compression, the response speed of the model can be improved well, and the resource consumption can also be significantly reduced.
[0119] Specifically, in some embodiments, the detecting the dialogue mode of each of the large models and establishing a corresponding adaptive radix tree for each of the large models includes:
[0120] If the dialogue mode of the large model is the few-shot learning mode, an adaptive radix tree with a survival period of a first duration is built;
[0121] If the dialogue mode of the large model is the multi-turn dialogue mode, an adaptive radix tree with a survival period of a second duration is built;
[0122] Wherein, the first duration is greater than the second duration.
[0123] In the embodiments of the present application, for each large model, when establishing a corresponding adaptive radix tree, the survival period when configuring the corresponding adaptive radix tree can also be set. In the embodiments of the present application, the survival period of the adaptive radix tree refers to the continuous duration after the adaptive radix tree is established.
[0124] It can be understood that if the dialogue mode of the large model is the few-shot learning mode, such models will have a demand for calls for a long time after being established. Therefore, in the embodiments of the present application, for a large model with a few-shot learning mode, an adaptive radix tree with a longer survival period, that is, a long-term tree, can be established to avoid the cold start problem and improve the inference speed at the same time.
[0125] If the dialogue mode of the large model is the multi-turn dialogue mode, since the content of each inference is likely to be different, it is difficult to reuse the data. In the embodiments of the present application, for a large model with a multi-turn dialogue mode, an adaptive radix tree with a shorter survival period, that is, a short-term tree, can be established.
[0126] In the embodiments of the present application, the survival duration of the long-term tree is recorded as the first duration, and the survival duration of the short-term tree is recorded as the second duration, where the first duration is greater than the second duration. The specific lengths of the two durations are not limited in the present application.
[0127] Specifically, in some embodiments, with reference to Figure 6 , the method further includes:
[0128] Detecting the historical call times and the historical time nodes of the last call of each of the adaptive radix trees;
[0129] If the historical call times of the adaptive radix tree are lower than a preset times threshold, and the duration from the historical time node to the current time node is greater than a preset duration threshold, performing a deletion process on the adaptive radix tree.
[0130] In the embodiments of the present application, for the maintenance and management of each adaptive radix tree, a strategy of the most recent call can be combined. For example, the historical call times and the historical time nodes of the last call of each adaptive radix tree can be detected. If the historical call times of the adaptive radix tree are lower than a preset times threshold, it indicates that it is overall less used. It can be further determined whether the duration from the historical time node of its last call to the current time node is greater than a preset duration threshold. Here, the preset duration threshold can be used to determine which adaptive radix trees have not been called recently. If the duration from the historical time node to the current time node is greater than the preset duration threshold, it indicates that the corresponding adaptive radix tree has not been called recently. If the duration from the historical time node to the current time node is less than or equal to the preset duration threshold, it indicates that the corresponding adaptive radix tree has been called recently.
[0131] In the embodiments of the present application, for an adaptive radix tree with historical call times lower than a preset times threshold and the duration from the historical time node of the call to the current time node greater than a preset duration threshold, a deletion process can be performed on it, thereby saving relevant storage space. For an adaptive radix tree with the duration from the historical time node of the call to the current time node greater than a preset duration threshold but the historical call times greater than the preset times threshold, it can be retained.
[0132] It can be understood that the method in the embodiments of the present application has at least the following advantages:
[0133] By adopting an Adaptive Radix Tree (ART) as the storage structure for key-value pair (KV) caching, this application achieves efficient storage and dynamic expansion and update, bringing multiple advantages. The Adaptive Radix Tree is a tree-shaped data structure optimized for modern processors. Through core technologies such as dynamic node types, path compression, and lazy expansion, it effectively overcomes the performance bottlenecks and space efficiency problems encountered by traditional data structures when dealing with large-scale data.
[0134] Specifically, the design of dynamic node types enables this application to flexibly adjust the storage structure according to the characteristics of different dialogue modes, reducing unnecessary storage overhead. This means that for highly repetitive dialogue content, such as the same prompts and examples in few-shot learning scenarios, through dynamic adjustment of node types, waste of storage space can be reduced while maintaining high data access speed.
[0135] Secondly, the path compression and lazy expansion strategies further improve storage efficiency and data processing speed. Path compression reduces the number of nodes by merging identical path nodes, thereby reducing memory occupancy and increasing access speed. The lazy expansion strategy ensures real-time data updates while avoiding frequent tree structure adjustments, guaranteeing the system's response speed.
[0136] In addition, the design of long-term trees and short-term trees meets the data storage requirements in different scenarios. Long-term trees are optimized for storing data that is frequently accessed over a long period, avoiding cold start problems and accelerating the inference speed. Short-term trees focus on the data during the current dialogue process and release resources after the dialogue ends, ensuring efficient utilization of system resources.
[0137] Finally, this application adopts a combination of less frequent calls and most recent calls in the tree update strategy. Compared with the traditional LRU (Least Recently Used) strategy, it can more reasonably retain valuable data, avoid frequent and unnecessary data deletion, and improve the accuracy and speed of data access.
[0138] In summary, the key-value pair caching method for the large model of this application, through its efficient storage structure and intelligent data processing strategy, provides a solution that saves space and improves performance for key-value pair caching management. It is particularly suitable for processing large-scale and dynamically changing data scenarios, with important practical value and broad application prospects.
[0139] Refer to Figure 7 , in the embodiments of this application, a key-value pair caching system for a large model is also proposed, including:
[0140] An acquisition unit 710, configured to acquire inference service data of multiple large models and convert the inference service data into key-value pair data;
[0141] A building unit 720 is configured to detect the dialogue patterns of each of the large models, establish a corresponding adaptive radix tree for each of the large models, and store the key-value pair data corresponding to the large model through the adaptive radix tree; wherein, the adaptive radix tree includes a plurality of nodes.
[0142] A detection unit 730 is configured to detect the data volume of the inference service data of each of the large models.
[0143] An adjustment unit 740 is configured to dynamically adjust the nodes of the adaptive radix tree corresponding to each of the large models according to the dialogue pattern and the data volume.
[0144] It can be understood that the content in the embodiments of the key-value pair caching method for the large model described above is applicable to the embodiments of this processing system. The functions specifically implemented by the embodiments of this processing system are the same as those of the embodiments of the key-value pair caching method for the large model, and the beneficial effects achieved are also the same as those of the embodiments of the key-value pair caching method for the large model.
[0145] Referring to Figure 8 , an embodiment of the present application also discloses a computer device, including:
[0146] At least one processor 810;
[0147] At least one memory 820, configured to store at least one program;
[0148] When at least one program is executed by at least one processor 820, at least one processor 820 implements the embodiments of the key-value pair caching method for the large model described above.
[0149] It can be understood that the content in the embodiments of the key-value pair caching method for the large model described above is applicable to the embodiments of this computer device. The functions specifically implemented by the embodiments of this computer device are the same as those of the embodiments of the key-value pair caching method for the large model, and the beneficial effects achieved are also the same as those of the embodiments of the key-value pair caching method for the large model.
[0150] An embodiment of the present application also discloses a computer-readable storage medium, in which a processor-executable program is stored, and the processor-executable program is used to implement the embodiments of the key-value pair caching method for the large model when executed by a processor.
[0151] It can be understood that the content in the embodiments of the key-value pair caching method for the large model described above is applicable to the embodiments of this computer-readable storage medium. The functions specifically implemented by the embodiments of this computer-readable storage medium are the same as those of the embodiments of the key-value pair caching method for the large model, and the beneficial effects achieved are also the same as those of the embodiments of the key-value pair caching method for the large model.
[0152] A key-value pair caching method for a large model disclosed in an embodiment of this application obtains inference service data of multiple large models, and converts the inference service data into key-value pair data; detects the conversation modes of each of the large models, and establishes a corresponding adaptive radix tree for each of the large models, and stores the key-value pair data corresponding to the large model through the adaptive radix tree; wherein, the adaptive radix tree includes multiple nodes; detects the data volume of the inference service data of each of the large models; dynamically adjusts the nodes of the adaptive radix tree corresponding to each of the large models according to the conversation mode and the data volume. The method in the embodiments of this application can effectively improve the caching efficiency of the key-value pairs of the large model, is applicable to large models of various conversation modes, is beneficial to improving the inference performance of the large model, and improves the utilization rate of storage resources.
[0153] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of this application are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.
[0154] In addition, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present application as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0155] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0156] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in connection with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0157] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other suitable processing, and then stored in a computer memory.
[0158] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0159] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0160] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents.
[0161] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
[0162] In the description of this specification, the descriptions referring to terms such as "one embodiment", "another embodiment", or "certain embodiments" mean that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0163] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.
Claims
1. A key-value pair caching method for a large model, characterized in that: include: Acquire inference business data of multiple large models, and convert the inference business data into key-value pair data; Detecting the dialogue mode of each of the large models, establishing a corresponding adaptive radix tree for each of the large models, and storing the key-value pair data corresponding to the large model through the adaptive radix tree; wherein the adaptive radix tree includes a plurality of nodes; Detecting the data volume of the inference business data of each of the large models; The nodes of the adaptive radix tree corresponding to each of the large models are dynamically adjusted according to the dialogue mode and the data volume.
2. A key-value pair caching method for a large model according to claim 1, characterized in that: The detecting the dialogue mode of each of the large models, establishing a corresponding adaptive radix tree for each of the large models, and storing the key-value pair data corresponding to the large model through the adaptive radix tree includes: If the conversation mode of the large model is a few-sample learning mode, construct a first adaptive radix tree including a plurality of compact nodes, and store the key-value pair data corresponding to the large model of the few-sample learning mode through the first adaptive radix tree; If the conversation mode of the large model is a multi-round conversation mode, construct a second adaptive radix tree including a plurality of leaf nodes, and store the key-value pair data corresponding to the large model of the multi-round conversation mode through the second adaptive radix tree; If the dialogue mode of the large model is a decision tree mode, a third adaptive radix tree including leaf nodes, internal nodes and compact nodes is constructed, and the key-value pair data corresponding to the large model of the decision tree mode is stored through the third adaptive radix tree.
3. A key-value pair caching method for a large model according to claim 2, characterized in that: The dynamically adjusting the nodes of the adaptive radix tree corresponding to each of the large models according to the dialogue mode and the data volume includes: If the conversation mode of the large model is a small sample learning mode, detecting whether the amount of data exceeds a first preset threshold; If the size of the data volume exceeds the first preset threshold, at least one of the compact nodes in the first adaptive radix tree is converted into an internal node.
4. A key-value pair caching method for a large model according to claim 2, characterized in that: The dynamically adjusting the nodes of the adaptive radix tree corresponding to each of the large models according to the dialogue mode and the data volume includes: If the dialogue mode of the large model is a multi-round dialogue mode, detecting whether the amount of data exceeds a second preset threshold; If the size of the data volume exceeds the second preset threshold, at least one leaf node is newly created in the second adaptive radix tree.
5. A key-value pair caching method for a large model according to any one of claims 1 to 4, characterized in that: The method further comprises: If the conversation mode of the large model is a few-sample learning mode, detecting the similarity of the existing key-value pair data in the adaptive radix tree; Comparing the similarity with a preset similarity threshold; If the similarity is greater than the similarity threshold, path compression processing is performed on the corresponding key-value pair data.
6. A key-value pair caching method for a large model according to claim 1, characterized in that: The detecting the dialogue mode of each of the large models and establishing a corresponding adaptive radix tree for each of the large models includes: If the conversation mode of the large model is a few-sample learning mode, an adaptive radix tree with a first duration is constructed; If the dialogue mode of the large model is a multi-round dialogue mode, construct an adaptive radix tree with a duration of the second time period; The first duration is greater than the second duration.
7. A key-value pair caching method for a large model according to claim 1, characterized in that: The method further comprises: Detect the historical call times and the historical time node of the last call of each of the adaptive radix trees; If the number of historical calls of the adaptive radix tree is lower than a preset number threshold, and the duration from the historical time node to the current time node is greater than a preset duration threshold, the adaptive radix tree is deleted.
8. A key-value pair cache system for a large model, characterized in that: include: An acquisition unit, used to acquire the inference business data of multiple large models and convert the inference business data into key-value pair data; An establishing unit, used for detecting the dialogue mode of each of the large models, establishing a corresponding adaptive radix tree for each of the large models, and storing the key-value pair data corresponding to the large model through the adaptive radix tree; wherein the adaptive radix tree includes a plurality of nodes; A detection unit, used for detecting the data volume of the inference business data of each of the large models; An adjustment unit is used to dynamically adjust the nodes of the adaptive radix tree corresponding to each of the large models according to the dialogue mode and the data volume.
9. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the key-value pair caching method for a large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to implement the key-value pair caching method of the large model as described in any one of claims 1 to 7 when executed by the processor.