Method, medium, device and program product for calculating attention score based on CPU
By calculating attention scores on the CPU and sending them to the GPU, the time-consuming data transmission problem caused by GPU video memory limitations is solved, and the computing efficiency and response speed of the large language model are improved.
Patent Information
- Application Number
- CN202510775011.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In the attention calculation process of large language models, due to the limited GPU video memory capacity, it is necessary to obtain key values from the CPU storage medium, resulting in a long time to transmit data, affecting the computing efficiency.
By calculating attention scores on the CPU and sending them to the GPU, it is avoided to transmit key values between different processors, and generate response information on the GPU directly after the attention calculation is performed on the CPU.
It improves the efficiency of the attention calculation process, reduces the transmission time between key values and processors, and improves the response speed of large language models.
Smart Images

Figure CN120297328A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large language models, and in particular, to a method, medium, device, and program product for calculating attention scores based on a CPU. Background Art
[0002] Retrieval-Augmented Generation (RAG) is a framework that combines retrieval and generation technologies. It enhances the generation ability of a large language model (LLM) by retrieving relevant information from an external knowledge base and combining it with a user query. During the RAG process, the user input query and the retrieved information together serve as the prompt information for the large language model. The large language model can process the obtained prompt information based on the attention mechanism and output a response message. To optimize this process, the TurboRAG solution proposes to pre-compute and store the key-value pairs corresponding to each text block. In this way, after retrieving the information relevant to the user query, the corresponding key-value pairs can be directly obtained, and attention calculation can be performed based on the obtained key-value pairs, avoiding the step of online key-value calculation, significantly reducing the computational overhead, and improving the response efficiency of the large language model.
[0003] Since the Graphics Processing Unit (GPU) has powerful computing performance, large language models are usually deployed on the GPU. However, RAG needs to retrieve information relevant to the user query from a large number of text blocks. Due to the limited video memory capacity of the GPU, the key-value pairs corresponding to these text blocks are usually stored in the memory of the Central Processing Unit (CPU) or in a storage medium such as a disk. During attention calculation, the CPU needs to first obtain the key-value pairs from the storage medium and then send the obtained key-value pairs to the GPU. In cases where the data transfer bandwidth is low or the data volume is large, the data transfer takes a long time, resulting in low efficiency in the attention calculation process. Summary of the Invention
[0004] In a first aspect, an embodiment of this application provides a method for calculating attention scores based on a CPU, and the method includes: Obtain a user query for input to a large language model; Retrieve information relevant to the user query and obtain the key-value pairs corresponding to the information from a storage medium; the storage medium is different from the GPU video memory, and the key-value pairs are pre-computed key-value pairs based on the large language model; The CPU calculates the attention based on the key value to obtain an attention score, and sends the attention score to the GPU, so that the GPU uses the attention score to generate response information for the user query.
[0005] In a second aspect, an embodiment of the present application provides a method for calculating an attention score based on a CPU. The method includes: The GPU receives the attention score sent by the CPU. The attention score is obtained by the CPU calculating the attention based on the key value. The key value is obtained by the CPU from the storage medium, and the key value corresponds to the information retrieved by the CPU. The information is related to the user query for inputting into the large language model. The storage medium is different from the GPU video memory, and the key value is a key value pre-calculated based on the large language model. The GPU generates response information for the user query based on the attention score.
[0006] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.
[0007] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any embodiment of the present application is implemented.
[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.
[0009] In the embodiments of the present application, different processors are used to execute the process of calculating the attention score in the large language model and other processing processes of the large language model. Specifically, the CPU executes the process of calculating the attention score, and the GPU executes other processing processes other than calculating the attention score to obtain response information. In this way, the CPU does not need to send the retrieved key value to the GPU, but directly sends the attention score to the GPU after calculating the attention score. Since the data volume of the attention score is much smaller than the data volume of the key value, the method of the embodiments of the present application can save the transmission time of the key value between different processors, thereby improving the efficiency of the attention calculation process.
[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings
[0011] The accompanying drawings here are incorporated into the specification and form a part of this application. These drawings illustrate embodiments consistent with this application and, together with the specification, are used to explain the technical solutions of this application.
[0012] Figure 1 is a schematic diagram of the overall processes of RAG and TurboRAG in the related art.
[0013] Figure 2 is a flowchart of the method for calculating attention scores based on a CPU in an embodiment of this application.
[0014] Figure 3 is a flowchart of the method for calculating attention scores based on a CPU in another embodiment of this application.
[0015] Figure 4 is a block diagram of the apparatus for calculating attention scores based on a CPU in an embodiment of this application.
[0016] Figure 5 is a block diagram of the apparatus for calculating attention scores based on a CPU in another embodiment of this application.
[0017] Figure 6 is a schematic diagram of a computer device in an embodiment of this application. Detailed Embodiments
[0018] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0019] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality.
[0020] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0021] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application and to make the above-mentioned objects, features, and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0022] Figure 1 Shows the overall processes of standard RAG and TurboRAG in the related art.
[0023] In standard RAG, a knowledge base can be established in advance, vectorize each text block in the knowledge base (embedding), and store the vectors of each text block. During the inference process, first retrieve K (K is a positive integer) text blocks (i.e., context information) related to the user query, and then, use the K text blocks and the user query together to construct a prompt and pass it to the large language model. The large language model can generate response information based on the attention mechanism. Specifically, it can calculate the query vector (Q), key vector (K), and value vector (V) based on the vectors of the user query and the K text blocks, obtain the attention scores based on the calculated Q, K, and V, and generate response information based on the attention scores.
[0024] To accelerate the above process, the TurboRAG solution proposes to pre-compute and store the key values corresponding to each text block. The stored key values are called the key-value cache (KV cache). In this way, after retrieving the information, the corresponding key values can be directly obtained, and the attention calculation can be performed based on the obtained key values, avoiding the step of online key value calculation, significantly reducing the computational overhead, and improving the response efficiency of the large language model.
[0025] Due to the powerful computing performance of the GPU, large language models are usually deployed on the GPU. However, RAG needs to retrieve information from a large number of text chunks. Since the video memory capacity of the GPU is limited, the key values corresponding to these text chunks are usually stored in memory or storage media such as disks that can be accessed by the CPU. When performing attention calculations, the CPU needs to first obtain the key values from the storage media and then send the obtained key values to the GPU. In cases where the data transfer bandwidth is low or the data volume is large, the data transfer takes a long time, resulting in low efficiency in the attention calculation process.
[0026] Based on this, an embodiment of the present application proposes a method for calculating attention scores based on the CPU. The CPU is used to replace the GPU to execute the calculation process of attention scores. In this way, the CPU does not need to send the retrieved key values to the GPU, but directly sends the attention scores to the GPU after calculating the attention scores. Since the data volume of the attention scores is much smaller than that of the key values, the method of the embodiment of the present application can save the transmission time of the key values between different processors, thereby improving the efficiency of the attention calculation process. See Figure 2 , the method includes: Step S12: Obtain a user query for input to the large language model; Step S14: Retrieve information related to the user query and obtain the key values corresponding to the information from the storage medium; the storage medium is different from the GPU video memory, and the key values are pre-calculated key values based on the large language model; Step S16: The CPU performs attention calculation based on the key values to obtain attention scores, and sends the attention scores to the GPU, so that the GPU uses the attention scores to generate response information for the user query.
[0027] In the embodiment of the present application, the step of performing attention calculation is executed by the CPU, and other steps can be executed by the CPU or other processing units. For the sake of simplicity, the solution of the embodiment of the present application will be described below taking the case where each step is executed by the CPU as an example.
[0028] In step S12, the user can input a user query (query) on a client (such as a browser, a mobile application, etc.) and send it to the server of the large model service via an HTTP request. The HTTP request is transmitted over the network to the server. After receiving the HTTP request, the server parses the HTTP request through the CPU and extracts the query input by the user. At this time, the query data is stored in a storage medium. Here, the storage medium can be a medium different from the GPU video memory. For example, it can be a medium directly or indirectly accessible by the CPU of the server, such as memory or disk, etc. Due to hardware architecture limitations, GPUs usually cannot directly access storage media such as memory and disks.
[0029] In step S14, the CPU can retrieve information related to the user query. Specifically, the CPU can vectorize the user query to obtain a vector corresponding to the user query, and based on the similarity between the vector corresponding to the user query and the vectors of each pre-stored text block, retrieve several text blocks related to the user query from each pre-stored text block. These retrieved text blocks are the information related to the user query. Among them, the information related to the user query can be stored in the above-mentioned storage medium, or in the cloud, or in other storage spaces.
[0030] After retrieving the information related to the user query, the CPU can obtain the key value corresponding to the above information. Specifically, after generating the key value, the source information of the key value (such as the text block corresponding to the key value and the document to which the text block belongs) and the storage address of the key value can be recorded. After the CPU retrieves the information related to the user query, it can determine the storage address of the key value bound to the retrieved information based on the recorded above source information and read the key value from the corresponding storage address.
[0031] In step S16, the CPU can obtain a query vector (Q), a key vector (K), and a value vector (V) calculated based on the user query. In some embodiments, the above Q, K, and V can be calculated by the GPU. Specifically, after the CPU parses the query input by the user, it can send the query to the GPU. The GPU can calculate the above Q, K, and V based on the received query and return them to the CPU. The data transmission between the CPU and the GPU can be implemented based on the PCIe bus. Since the data volume of the query and Q, K, and V is usually relatively small, the time consumption for transmitting the query and Q, K, and V is usually short; and the computing power of the GPU is relatively strong, so the time consumption for calculating Q, K, and V based on the query is also short. Therefore, the overall time consumption of the above process is usually short.
[0032] In other examples, after the CPU parses the query, it can directly calculate Q, K, and V based on the query. Although the computing power of the CPU is weaker than that of the GPU, doing so can save the time for the CPU to transmit the query to the GPU and for the GPU to return Q, K, and V to the CPU.
[0033] After obtaining Q, K, and V, the CPU can perform attention calculation based on Q, K, V, and the key-value obtained in step S14 to obtain the attention score. Specifically, the CPU can fuse the key vector in the key-value with the key vector calculated based on the user query to obtain a fused key vector (denoted as K'), and fuse the value vector in the key-value with the value vector calculated based on the user query to obtain a fused value vector (denoted as V'). Then, perform attention calculation based on Q, K', and V' to obtain the attention score.
[0034] After obtaining the attention score, the CPU can send the attention score to the GPU through the PCIe bus. After receiving the above attention score, the GPU can use the attention score to generate response information for the user query.
[0035] In some embodiments, the GPU and the CPU can obtain response information through multiple rounds of iterative interaction. In each iteration, the CPU can perform attention calculation based on the query vector, key vector, and value vector sent by the GPU in this iteration to obtain the attention score in this iteration. Among them, the query vector, key vector, and value vector sent by the GPU in the first iteration are generated based on the user query or based on the historical output tokens (such as the previous output token) of the large language model. The query vector, key vector, and value vector sent by the GPU in the i-th iteration are generated based on the attention score obtained in the (i - 1)-th iteration; 1 < i < N, and N is the preset maximum number of iterations. The attention score obtained in the N-th iteration is used to generate response information for the user query. The iterative process will be illustrated by a specific example below. In this embodiment, for each output token generated in the response information, the CPU and the GPU perform N rounds of iterative interaction. The process of generating the first output token is as follows: In the first iteration, the GPU generates a query vector, a key vector, and a value vector based on the user query, sends the generated query vector, key vector, and value vector to the CPU. The CPU fuses the key vector among them with the key vector in the key-value to obtain a fused key vector, fuses the value vector among them with the value vector in the key-value to obtain a fused value vector, generates an attention score based on the query vector, the fused key vector, and the fused value vector and sends it to the GPU. The GPU calculates a hidden state vector based on the attention score.
[0036] In the second iteration, the GPU obtains the hidden state vector that the CPU obtained in the first iteration, calculates query vectors, key vectors, and value vectors based on the hidden state vector, and sends the calculated query vectors, key vectors, and value vectors to the CPU so that the CPU can recalculate the attention scores. Then, the CPU sends the recalculated attention scores to the GPU so that the GPU can recalculate the hidden state vector.
[0037] The subsequent iteration process is similar to that in the foregoing embodiments and will not be elaborated here. In the Nth iteration, the GPU calculates the probabilities corresponding to each token based on the hidden state vector obtained in the Nth iteration, and determines the token with the highest probability as the first output token.
[0038] The process of generating other output tokens is similar to the above process. The difference is that in the process of generating the second output token, in the first iteration, the GPU generates query vectors, key vectors, and value vectors based on the first output token. In the process of generating the third output token, in the first iteration, the GPU generates query vectors, key vectors, and value vectors based on the second output token. And so on.
[0039] In some embodiments, the large language model may include multiple processing modules, and each processing module includes a sub-module deployed on the CPU and a sub-module deployed on the GPU. The ith iteration process may be executed by the ith processing module, and N is the total number of processing modules. The operations executed by the CPU in the ith iteration are executed by the sub-module deployed on the CPU in the ith processing module, and the operations executed by the GPU in the ith iteration are executed by the sub-module deployed on the GPU in the ith processing module. The specific processing process can be found in the foregoing embodiments and will not be elaborated here.
[0040] In some embodiments, using the CPU instead of the GPU to calculate the attention scores does not necessarily improve the calculation efficiency in all cases. Since the computing power of the CPU is relatively weak, the time taken to calculate the attention scores using the GPU is mainly affected by the computing power of the CPU. When calculating the attention scores using the GPU, key values need to be transferred between the CPU and the GPU. Therefore, the time taken to calculate the attention scores using the GPU is mainly affected by the transmission time of the key values. To improve the calculation efficiency of the attention scores, a preset condition can be determined based on the estimated calculation duration of the CPU for attention calculation and the estimated transmission duration of the CPU for transferring the key values to the GPU. Only when the preset condition is met, step S16 is executed.
[0041] In some embodiments, the preset condition includes: the estimated calculation duration for the CPU to perform attention calculation is less than the estimated transmission duration for the CPU to transmit key values to the GPU. When this preset condition is met, calculating the attention score by the CPU instead of the GPU can reduce the time consumption caused by the transmission of key values and improve the calculation efficiency of the attention score.
[0042] It is possible to determine whether the estimated calculation duration (hereinafter referred to as the calculation duration) for the CPU to perform attention calculation is less than the estimated transmission duration (hereinafter referred to as the transmission duration) for the CPU to transmit key values to the GPU based on at least one of the following conditions: (1) The length of the information related to the user query. The length of the above information can be expressed by the number of characters included in the information. The length of the key values is usually positively correlated with the length of the information, and the length of the information will affect the amount of data to be processed during attention calculation and the amount of data transmission between the CPU and the GPU. Therefore, it is possible to determine whether the above calculation duration is less than the transmission duration based on the length of the information. In some embodiments, if the length of the information is within a preset length range (such as between a preset length lower limit and a preset length upper limit), it can be determined that the calculation duration is less than the transmission duration. Conversely, if the length of the information is not within the preset length range, it can be determined that the calculation duration is greater than or equal to the transmission duration.
[0043] (2) The resource utilization rate of the CPU. The resource utilization rate of the CPU is the ratio between the non-idle resources in the CPU and the total resources of the CPU. For example, the resource utilization rate of the CPU can be characterized by the ratio between the number of non-idle threads in the CPU and the total number of threads of the CPU. The resource utilization rate of the CPU can reflect the load status of the CPU. If the resource utilization rate of the CPU is too high, it means that the CPU is busy processing various tasks and may not have enough resources to calculate the attention score. Therefore, if the resource utilization rate of the CPU is less than a preset resource utilization rate threshold, it can be determined that the calculation duration is less than the transmission duration. Conversely, if the resource utilization rate of the CPU is greater than or equal to the preset resource utilization rate threshold, it can be determined that the calculation duration is greater than or equal to the transmission duration.
[0044] (3) The data transmission bandwidth between the CPU and the GPU. The larger the data transmission bandwidth, the shorter the duration for the CPU to transmit the key value cache to the GPU. Therefore, if the data transmission bandwidth between the CPU and the GPU is less than a preset bandwidth threshold, it can be determined that the calculation duration is less than the transmission duration. Conversely, if the data transmission bandwidth between the CPU and the GPU is greater than or equal to the preset bandwidth threshold, it can be determined that the calculation duration is greater than or equal to the transmission duration.
[0045] In practical applications, the magnitude relationship between the computing duration and the transmission duration can be determined by combining one or more of the above parameters, or can be determined based on other factors.
[0046] If it is determined that the above preset conditions are met, step S16 can be executed, that is, the attention score is calculated by the CPU. Further, if the preset conditions are not met, the key value can be sent to the GPU so that the GPU performs attention calculation based on the query vector, key vector, value vector, and key value to obtain the attention score. That is to say, in the case where the preset conditions are not met, the GPU can execute the calculation process of the attention score. Based on this, the embodiment of the present application realizes the flexible selection of the processor for calculating the attention score between the CPU and the GPU according to the actual situation, and can meet the requirements of various application scenarios.
[0047] In some embodiments, when the CPU calculates the attention score, the CPU can call multiple threads to perform attention calculation based on the key value to obtain the attention score. Among them, the number of threads called can be preset or determined according to the actual situation. Optionally, it can be determined according to the current available thread number of the CPU. For example, a certain proportion of threads can be selected from the current available threads to calculate the attention score.
[0048] When performing natural language processing, the order of tokens is crucial for semantic understanding. Therefore, in the related art, position encoding information for each token is generated based on the position of the token in the original document, such as Rotary Position Embedding (RoPE). After fusing the position encoding information into the query vector and the key vector, the key value is generated and stored based on the key vector fused with the position encoding information. However, when retrieving, several text blocks are retrieved from the original document, and the position of the token in the text block may be different from its position in the original document. For example, the original document includes 100 tokens, and the retrieved text block includes the 51st token to the 100th token in the original document. Then, the 51st token in the original document is the 1st token in the retrieved text block, which is different from its position information in the original document, resulting in a change in the position encoding information of the token.
[0049] To solve the above problem, the key value obtained in the embodiment of the present application is the key value without fused position encoding information. After retrieving the information, the position encoding information of the token in the information is obtained, and the above position encoding information is fused into the key value. The above method separates the process of obtaining the key value from the process of obtaining the position encoding information, avoiding the situation where the attention calculation result is incorrect due to the change in the position encoding information of the token before and after retrieval.
[0050] See Figure 3 , an embodiment of the present application further provides a method for calculating attention scores based on a CPU, and the method includes: Step S22: The GPU receives the attention scores sent by the CPU. The attention scores are obtained by the CPU based on key values for attention calculation. The key values are obtained by the CPU from a storage medium, and the key values correspond to the information retrieved by the CPU. The information is related to the user query for inputting into the large language model. The storage medium is different from the GPU video memory, and the key values are pre-calculated key values based on the large language model; Step S24: The GPU generates response information for the user query based on the attention scores.
[0051] In some embodiments, the GPU receiving the attention scores sent by the CPU includes: the GPU receives the attention scores sent by the CPU when a preset condition is met. The preset condition includes: the estimated calculation duration for the CPU to perform attention calculation is less than the estimated transmission duration for the CPU to transmit the key values to the GPU.
[0052] In some embodiments, the method further includes: the GPU receives the key values sent by the CPU when the preset condition is not met; the GPU performs attention calculation based on the key values to obtain the attention scores.
[0053] For the specific implementation details of the embodiments of the present application, refer to the foregoing method embodiments executed by the CPU, which will not be elaborated here.
[0054] See Figure 4 , an embodiment of the present application further provides a device for calculating attention scores based on a CPU, and the device includes: An acquisition module 102, configured to acquire a user query for inputting into the large language model; A retrieval module 104, configured to retrieve information related to the user query and obtain the key values corresponding to the information from a storage medium. The storage medium is different from the GPU video memory, and the key values are pre-calculated key values based on the large language model; A calculation module 106, configured to perform attention calculation based on the key values to obtain attention scores, and send the attention scores to the GPU, so that the GPU uses the attention scores to generate response information for the user query. The calculation module 106 is a module in the CPU.
[0055] For the specific implementation details of the embodiments of the present application, refer to the foregoing method embodiments executed by the CPU, which will not be elaborated here.
[0056] See Figure 5, an embodiment of the present application further provides a device for calculating attention scores based on a CPU. The method includes: A receiving module 202, configured to receive the attention scores sent by the CPU. The attention scores are obtained by the CPU through attention calculation based on key values. The key values are obtained by the CPU from a storage medium, and the key values correspond to the information retrieved by the CPU. The information is related to the user query for inputting into the large language model. The storage medium is different from the GPU video memory, and the key values are pre-calculated key values based on the large language model. A response generation module 204, configured to generate response information for the user query based on the attention scores. Both the receiving module 202 and the response generation module 204 are modules in the GPU.
[0057] For the specific implementation details of the embodiments of the present application, please refer to the foregoing method embodiments executed by the GPU, which will not be elaborated here.
[0058] An embodiment of the present application further provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in any of the foregoing embodiments.
[0059] Figure 6 FIG. shows a more specific schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. The device may include: a processor 302, a memory 304, an input / output interface 306, a communication interface 308, and a bus 310. Among them, the processor 302, the memory 304, the input / output interface 306, and the communication interface 308 are communicatively connected to each other inside the device through the bus 310.
[0060] The processor 302 may include a CPU and a GPU, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application. The processor 302 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0061] The memory 304 may be implemented in the form of a Read Only Memory (ROM), a Random Access Memory (RAM), a static storage device, a dynamic storage device, etc. When the processor 302 includes a CPU and a GPU, the memory 304 may include the memory corresponding to the CPU (such as memory and a disk) and the memory corresponding to the GPU (such as video memory). Among them, the memory corresponding to the CPU can be accessed by the CPU, and the CPU can obtain the data therein and send it to the GPU. Similarly, the memory corresponding to the GPU can be accessed by the GPU, and the GPU can obtain the data therein and send it to the CPU. The memory 304 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present application through software or firmware, the relevant program codes are stored in the memory 304 and are called and executed by the processor 302.
[0062] The input / output interface 306 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0063] The communication interface 308 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can implement communication in a wired manner (such as USB, network cable, etc.) or can implement communication in a wireless manner (such as a mobile network, Wi-Fi, Bluetooth, etc.).
[0064] The bus 310 includes a path for transmitting information between various components of the device (such as the processor 302, the memory 304, the input / output interface 306, and the communication interface 308).
[0065] It should be noted that although the above device only shows the processor 302, the memory 304, the input / output interface 306, the communication interface 308, and the bus 310, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary for implementing the solutions of the embodiments of the present application and does not necessarily include all the components shown in the figure.
[0066] The embodiments of the present application provide a computer program product, including a computer program, which when executed by a processor implements the method described in any embodiment of the present application.
[0067] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the methods described in any of the foregoing embodiments are implemented.
[0068] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computer device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0069] Each embodiment in the present application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. The relevant parts can be referred to the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of the present application, the functions of the modules can be implemented in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0070] The above is only the specific implementation manner of the embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the embodiments of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present application.
Claims
1. A method for calculating attention scores based on CPU, the method comprising: Obtaining a user query for inputting into a large language model; Retrieving information related to the user query and obtaining the key-value corresponding to the information from a storage medium; The storage medium is different from the GPU video memory, and the key-value is a key-value pre-calculated based on the large language model; The CPU performs attention calculation based on the key-value to obtain an attention score, and sends the attention score to the GPU, so that the GPU uses the attention score to generate response information for the user query.
2. The method according to claim 1, wherein the CPU performs attention calculation based on the key-value to obtain an attention score, comprising: If the estimated calculation duration for the CPU to perform attention calculation is less than the estimated transmission duration for the CPU to transmit the key-value to the GPU, the CPU performs attention calculation based on the key-value to obtain an attention score.
3. The method according to claim 2, the method further comprising: Judging whether the estimated calculation duration for the CPU to perform attention calculation is less than the estimated transmission duration for the CPU to transmit the key-value to the GPU based on at least one of the following conditions: The length of the information; The resource utilization rate of the CPU; The data transmission bandwidth between the CPU and the GPU.
4. The method according to claim 1, the method further comprising: If the estimated calculation duration for the CPU to perform attention calculation is greater than or equal to the estimated transmission duration for the CPU to transmit the key-value to the GPU, the CPU sends the key-value to the GPU, so that the GPU performs attention calculation based on the key-value to obtain the attention score.
5. The method according to claim 1, wherein the CPU performs attention calculation based on the key-value to obtain an attention score, comprising: The CPU calls multiple threads to perform attention calculation based on the key-value to obtain an attention score.
6. The method according to claim 1, wherein the CPU performs attention calculation based on the key-value to obtain an attention score, comprising: The CPU obtains a query vector, a key vector and a value vector sent by the GPU; The CPU fuses the key vector sent by the GPU with the key vector in the key-value to obtain a fused key vector, and fuses the value vector sent by the GPU with the value vector in the key-value to obtain a fused value vector; The CPU performs attention calculation based on the query vector, the fused key vector and the fused value vector to obtain an attention score.
7. The method according to claim 6, wherein the GPU and the CPU obtain the response information through multiple rounds of iterative interaction. In each iteration, the CPU performs attention calculation based on the query vector, the key vector and the value vector sent by the GPU in the current iteration to obtain the attention score in the current iteration; Among them, The query vector, the key vector and the value vector sent by the GPU in the first iteration are generated based on the user query or based on the historical output tokens of the large language model; The query vector, key vector, and value vector sent by the GPU during the i-th iteration are generated based on the attention scores obtained in the (i - 1)-th iteration; 1 < i < N, where N is a preset maximum number of iterations; The attention scores obtained in the N-th iteration are used to generate response information for the user query.
8. The method according to claim 1, before the CPU calculates the attention scores based on the key values, the method further includes: Obtaining the position encoding information of each token in the information; Fusing the position encoding information into the key values.
9. A method for calculating attention scores based on a CPU, the method includes: The GPU receives the attention scores sent by the CPU, the attention scores are obtained by the CPU through attention calculation based on key values, the key values are obtained by the CPU from a storage medium, and the key values correspond to the information retrieved by the CPU; the information is related to the user query for inputting into a large language model, the storage medium is different from the GPU video memory, and the key values are pre-calculated key values based on the large language model; The GPU generates response information for the user query based on the attention scores.
10. The method according to claim 9, where the GPU receives the attention scores sent by the CPU includes: The GPU receives the attention scores sent by the CPU when a preset condition is satisfied; The preset condition includes: the estimated calculation duration for the CPU to perform attention calculation is less than the estimated transmission duration for the CPU to transmit the key values to the GPU.
11. The method according to claim 10, the method further includes: The GPU receives the key values sent by the CPU when the preset condition is not satisfied; The GPU calculates the attention scores based on the key values.
12. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of claims 1 to 11 is implemented.
13. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method described in any one of claims 1 to 11 is implemented.
14. A computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Extraction type machine intelligent reading understanding question-answering system
CN111611361A
Hypergraph partition-based computing power network task unloading method
CN118113367A
Large-scale language model KV Cache optimization method based on recent query attention information
CN119396995A
Energy Efficient Computations of Attention-based Inferences
US20240281428A1
A system and method for a dynamic large language model (LLM) agent
WO2025046584A1
Cited By
Large language model reasoning optimization method, system and equipment and storage medium
CN120996208A
Request processing method and device, equipment, medium and program product
CN121397087A