Data processing method and device, electronic equipment, storage medium and computer program product

CN122047520BActive Publication Date: 2026-09-04SOPHGO TECH LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610510365.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-17
Publication Date
2026-09-04
Estimated Expiration
2046-04-17

AI Technical Summary

Technical Problem

而在大语言模型的推理过程中,算力和带宽消耗较大,不利于节省功耗

Benefits of technology

[0046] In this application, the dynamic range of the numerical values ​​remains unchanged (the exponent remains constant). Only the length of the mantissa of the floating-point numbers in the query vector and/or the key vectors of each historical term is reduced, sacrificing a slight level of precision. This allows for the effective representation of both extremely large and small values ​​in the widely distributed query vector and/or key vectors, minimizing the significant precision loss caused by the uniform scaling factor during quantization. This saves computational power and improves efficiency while minimizing prediction accuracy loss. Furthermore, this application eliminates the need for quantization-dequantization operations, resulting in superior model stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047520B_ABST
    Figure CN122047520B_ABST
Patent Text Reader

Abstract

The application provides a data processing method and device, electronic equipment, storage medium and computer program product; the method comprises the following steps: obtaining a query vector to be queried and a key-value vector pair of each historical word in a text sequence; reducing the length of the mantissa of the floating point number in the query vector and / or the key vector of each historical word, and determining the relevance degree of the query vector and each historical word by using the attention mechanism in the large language model; based on the relevance degree of the query vector and each historical word and the value vector of the historical word, the target word of the query vector is generated by using the large language model. Through the application, the computing power can be saved, the power consumption can be reduced, and the efficiency can be improved while minimizing the loss of prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to data processing technology, and more particularly to a data processing method, apparatus, electronic device, storage medium, and computer program product. Background Technology

[0002] Large Language Models (LLMs) have demonstrated tremendous application potential in various fields, such as text generation and intelligent question-answering systems. However, the inference process of large language models consumes significant computing power and bandwidth, which is not conducive to saving power consumption. Summary of the Invention

[0003] This application provides a data processing method, apparatus, electronic device, storage medium, and computer program product that can save computing power and reduce power consumption while minimizing the loss of prediction accuracy.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a data processing method, including:

[0006] Obtain the query vector to be queried and the key-value vector pairs of each historical word in the text sequence; wherein, the key vector in each key-value vector pair of historical words represents the features of the historical word, and the value vector represents the context features of the historical word in the text sequence; the numbers in the query vector and the key vector of each historical word are all floating-point numbers, and the encoding format of floating-point numbers includes exponent bits and mantissa bits;

[0007] Reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, and use the attention mechanism in the large language model to determine the relevance between the query vector and each historical word.

[0008] Based on the relevance of the query vector to each historical word and the value vector of the historical words, the target word of the query vector is predicted and generated using the large language model.

[0009] In some embodiments, reducing the length of the mantissas of the floating-point numbers in the query vector and / or the key vectors of each historical term includes:

[0010] Based on the mapping relationship between the large language model and the length of the mantissa, the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word is reduced; wherein, the mapping relationship includes the reduced mantissa, and the reduced mantissa is determined based on experimental samples. The difference between the prediction accuracy of the large language model before the reduction of the mantissa and the prediction accuracy of the large language model after the reduction of the mantissa is less than a preset difference threshold.

[0011] In some embodiments, the step of predicting and generating the target word of the query vector using the large language model based on the relevance between the query vector and each historical word and the value vector of the historical words includes:

[0012] Based on the relevance of the query vector to each historical word, target historical words that meet the preset relevance conditions are selected from the text sequence.

[0013] Based on the relevance between the query vector and the target historical word, as well as the value vector of the target historical word, the target word of the query vector is predicted and generated using the large language model.

[0014] In some embodiments, selecting target historical words that meet preset relevance criteria from the text sequence based on the relevance degree between the query vector and each historical word includes:

[0015] The relevance of the query vector to each historical word is sorted to obtain the ranking result of the relevance.

[0016] Based on the ranking results of relevance, target historical words with a sum of relevance values ​​greater than a preset threshold are selected from the text sequence in descending order of relevance.

[0017] In some embodiments, the relevance level is a floating-point number with an exponent and a mantissa in the encoding format; the process of sorting the relevance levels of the query vector with each historical word to obtain a ranking result of the relevance levels includes:

[0018] The index values ​​in each correlation level are sorted to obtain the initial ranking of correlation levels.

[0019] In response to the initial sorting result including multiple target relevance degrees with the same exponent, the last digit of each target relevance degree is sorted to obtain a sorting result after reordering the multiple target relevance degrees in the initial sorting result; wherein, the direction of sorting the last digit is consistent with the direction of sorting the exponent.

[0020] In some embodiments, sorting the relevance between the query vector and each historical word to obtain a relevance ranking result includes:

[0021] The relevance of the query vector to each historical term is amplified by a predetermined factor and then sorted to obtain the ranking result of the relevance.

[0022] In some embodiments, the attention mechanism is a multi-head attention mechanism; reducing the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical word, and using the attention mechanism in the large language model to determine the relevance between the query vector and each historical word, includes:

[0023] Reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, and use the attention mechanism in the large language model to determine the relevance score between the query vector and each historical word corresponding to each attention head;

[0024] For each historical word, the average relevance score between the query vector and the historical word is determined based on the relevance score between the query vector and the historical word corresponding to different attention heads and the number of attention heads.

[0025] The relevance scores of each historical term are averaged using an activation function to obtain the relevance degree between the query vector and each historical term.

[0026] In some embodiments, the numbers in the value vectors of historical words are floating-point numbers, and the encoding format of floating-point numbers includes exponent bits and mantissa bits; the step of predicting and generating the target word of the query vector using the large language model based on the relevance between the query vector and each historical word and the value vectors of the historical words includes:

[0027] The length of the mantissa of the floating-point numbers in the value vector of historical words is reduced, and the target word of the query vector is predicted and generated using the large language model, taking into account the relevance between the query vector and the historical words.

[0028] This application provides a data processing apparatus, including:

[0029] The vector acquisition module is configured to acquire the query vector to be queried and the key-value vector pairs of each historical word in the text sequence; wherein, the key vector in each key-value vector pair of the historical word represents the feature of the historical word, and the value vector represents the contextual feature of the historical word in the text sequence; the numbers in the query vector and the key vector of each historical word are all floating-point numbers, and the encoding format of the floating-point numbers includes exponent bits and mantissa bits;

[0030] The attention module is configured to reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical word, and to use the attention mechanism in the large language model to determine the relevance between the query vector and each historical word.

[0031] The prediction module is configured to predict and generate the target word of the query vector based on the relevance between the query vector and each historical word and the value vector of the historical words, using the large language model.

[0032] In some embodiments, the attention module is further configured to reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word based on the mapping relationship between the large language model and the mantissa length; wherein the mapping relationship includes the reduced mantissa, and the reduced mantissa is determined based on experimental samples. The difference between the prediction accuracy of the large language model before the mantissa reduction and the prediction accuracy of the large language model after the mantissa reduction is less than a preset difference threshold.

[0033] In some embodiments, the prediction module is further configured to select target historical words that meet preset relevance conditions in the text sequence based on the relevance degree between the query vector and each historical word; and to predict and generate the target word of the query vector using the large language model based on the relevance degree between the query vector and the target historical word and the value vector of the target historical word.

[0034] In some embodiments, the prediction module is further configured to sort the relevance between the query vector and each historical word to obtain a relevance ranking result; based on the relevance ranking result, select target historical words from the text sequence whose sum of relevance is greater than a preset sum threshold in descending order of relevance.

[0035] In some embodiments, the correlation degree is a floating-point number with an encoding format including an exponent and a mantissa; the prediction module is further configured to sort the exponents of each correlation degree to obtain an initial sorting result of the correlation degree; in response to multiple target correlation degrees including the same exponent in the initial sorting result, the mantissas of each target correlation degree are sorted to obtain a sorting result after reordering the multiple target correlation degrees in the initial sorting result; wherein, the direction of sorting the mantissas is consistent with the direction of sorting the exponents.

[0036] In some embodiments, the prediction module is further configured to amplify the relevance of the query vector to each historical word by a predetermined factor before sorting them to obtain a ranking result of the relevance.

[0037] In some embodiments, the attention mechanism is a multi-head attention mechanism; the attention module is further configured to reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical word, and to use the attention mechanism in the large language model to determine the relevance score between the query vector and each historical word corresponding to each attention head; for each historical word, based on the relevance score between the query vector and the historical word corresponding to different attention heads and the number of attention heads, to determine the average relevance score between the query vector and the historical word; and to use an activation function to process the average relevance score corresponding to each historical word to obtain the degree of relevance between the query vector and each historical word.

[0038] In some embodiments, the numbers in the value vector of historical words are floating-point numbers, and the encoding format of floating-point numbers includes exponent bits and mantissa bits; the prediction module is further configured to reduce the length of the mantissa bits of the floating-point numbers in the value vector of historical words, and combine the relevance between the query vector and historical words to predict and generate the target word of the query vector using the large language model.

[0039] This application provides an electronic device, including:

[0040] processor;

[0041] Memory used to store computer programs or instructions;

[0042] The processor executes the computer program or instructions to implement the method provided in the embodiments of this application.

[0043] This application provides a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements the method provided in this application.

[0044] This application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the method provided in this application.

[0045] The technical solutions provided by the embodiments of this application may include the following beneficial effects:

[0046] In this application, the dynamic range of the numerical values ​​remains unchanged (the exponent remains constant). Only the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical term is reduced, sacrificing a slight level of precision. This allows for the effective representation of both extremely large and small values ​​in the widely distributed query vector and / or key vectors, minimizing the significant precision loss caused by the uniform scaling factor during quantization. This saves computational power and improves efficiency while minimizing prediction accuracy loss. Furthermore, this application eliminates the need for quantization-dequantization operations, resulting in superior model stability.

[0047] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0048] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0049] Figure 1 This is a schematic diagram of the architecture of a data processing system provided in an embodiment of this application;

[0050] Figure 2 This is a schematic diagram of the terminal structure provided in the embodiments of this application;

[0051] Figure 3 This is a flowchart illustrating the data processing method provided in an embodiment of this application;

[0052] Figure 4 This is an example diagram illustrating the encoding format of floating-point numbers in this application;

[0053] Figure 5 This is an example diagram of encoding with reduced mantissa length in this application;

[0054] Figure 6 The performance of selecting target historical words in this application Figure 1 ;

[0055] Figure 7 The performance of selecting target historical words in this application Figure 2 ;

[0056] Figure 8 An example diagram of attention calculation mechanisms in related technologies;

[0057] Figure 9 This is an example diagram of the attention calculation mechanism in this application;

[0058] Figure 10 This is a flowchart illustrating another data processing method provided in an embodiment of this application. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0060] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0061] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0062] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0063] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0064] KV Cache is an optimization technique used to accelerate the autoregressive generation process of large language models. It avoids redundant calculations of historical words when generating new words by caching the keys K and values ​​V of all previously computed historical words. K determines which historical word is prioritized, and V determines what information the prioritized historical word provides.

[0065] In related technologies, large language models, when predicting new words, compare the attention scores of the currently queried vector with all cached key vectors based on the key-value cache. Then, they rank candidate tokens (historical words) by importance based on these attention scores and select the K tokens with the highest scores for subsequent new word prediction. This process heavily relies on high-precision floating-point multiplication (such as dot product) to generate an accurate ranking, resulting in high computational consumption. Furthermore, the floating-point computing units supporting these complex operations occupy a significant amount of chip area, hindering chip miniaturization.

[0066] This application provides a data processing method, apparatus, electronic device, and computer-readable storage medium that can save computing power, reduce power consumption, and improve data processing efficiency.

[0067] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of user terminals, such as laptops, tablets, desktop computers, set-top boxes, and mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), or as servers. The following will describe exemplary applications when the electronic device is implemented as a server.

[0068] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the data processing system provided in the embodiments of this application. In order to implement the data processing method of this application, the terminal (terminal 200-1 and terminal 200-2 are shown as examples) connects to the server 400 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0069] In some possible implementations, user A can trigger prediction request A through terminal 200-1, and user B can trigger prediction request B through terminal 200-2. Prediction requests A and B are uploaded to server 400 via network 300. Prediction request A may be based on the user inputting text in the text creation application of 210-1, thereby triggering server 400 to automatically complete sentences based on the data processing method of this disclosure embodiment. Prediction request B may be based on the user inputting a voice message in the social application of 210-2 and instructing it to be converted into text, thereby triggering server 400 to predict possible words based on the data processing method of this disclosure embodiment, thus improving recognition accuracy.

[0070] In some embodiments, server 400 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, and big data and artificial intelligence platforms. Terminals (such as terminal 200-1 and terminal 200-2) may be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these. Terminals and servers can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the invention.

[0071] See Figure 2 , Figure 2 This is a schematic diagram of the terminal structure provided in the embodiments of this application. It should be noted that... Figure 2 The terminal 200 shown can be Figure 1 Either terminal 200-1 or terminal 200-2 shown in the figure can be other terminals, and this application embodiment does not limit this. Figure 2 The terminal 200 shown includes at least one processor 210, a memory 250, at least one network interface 220, and a user interface 230. The various components in the terminal 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to implement communication between these components. In addition to a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 240.

[0072] Processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0073] User interface 230 includes one or more output devices 231 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0074] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices physically located away from the processor 210.

[0075] The memory 250 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 250 described in this application embodiment is intended to include any suitable type of memory.

[0076] In some embodiments, memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0077] Operating system 251 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0078] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including: Bluetooth, Wi-Fi, and Universal Serial Bus (USB), etc.

[0079] Presentation module 253 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 231 associated with user interface 230 (e.g., a display screen, a speaker, etc.).

[0080] The input processing module 254 is used to detect and translate one or more user inputs or interactions from one or more input devices 232.

[0081] In some embodiments, the data processing method of this application can also be executed by a terminal, and the apparatus provided in the embodiments of this application can be implemented in software. Figure 2 A data processing device 255 stored in memory 250 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: a vector acquisition module 2551, an attention module 2552, and a prediction module 2553. These modules are logically linked and can therefore be arbitrarily combined or further divided according to their implemented functions. The functions of each module will be described below.

[0082] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0083] In some embodiments, the terminal or server can implement the methods provided in the embodiments of this application by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a local application (APP), that is, a program that needs to be installed in the operating system to run, such as a social APP; it can also be a mini-program, that is, a program that only needs to be downloaded to a browser environment to run; or it can be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.

[0084] The data processing method provided in this application will be described below with reference to exemplary applications and implementations of the electronic devices provided in the embodiments of this application.

[0085] See Figure 3 , Figure 3 This is a flowchart illustrating the data processing method provided in the embodiments of this application, which will be combined with... Figure 3 The steps shown are explained.

[0086] S301. Obtain the query vector to be queried and the key-value vector pairs of each historical word in the text sequence; wherein, the key vector in each key-value vector pair of historical words represents the features of the historical word, and the value vector represents the context features of the historical word in the text sequence; the numbers in the query vector and the key vector of each historical word are all floating-point numbers, and the encoding format of the floating-point numbers includes exponent bits and mantissa bits;

[0087] S302. Reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, and use the attention mechanism in the large language model to determine the relevance between the query vector and each historical word.

[0088] S303. Based on the relevance between the query vector and each historical word, and the value vector of the historical words, the target word of the query vector is predicted and generated using the large language model.

[0089] The data processing method in this application can be applied to the aforementioned server or terminal, collectively referred to herein as an electronic device.

[0090] In step S301, the electronic device acquires the query vector to be queried. When the target word of the query vector is generated for the first time, the query vector can be the embedding vector of the last word in the user's prompt, and the KV Cache is initialized by calculating the K vector and V vector of each word through forward propagation of the prompt and caching them. When the target word of the query vector is not generated for the first time, the query vector can be obtained by transforming the embedding vector of the target word generated in the previous step. When the large language model makes predictions based on the query vector, it combines the key-value vectors of each historical word in the text sequence and uses an attention mechanism to predict the next most likely embedding vector, thereby generating the target word.

[0091] In this application, the key-value vector pairs of each historical word in the text sequence constitute the aforementioned KV Cache. In each key-value vector pair, the key vector represents the feature of the historical word, and the value vector represents the contextual feature of the historical word in the text sequence. The key vector is the aforementioned K, and the value vector is the aforementioned V. In this application, the newly generated key and value vectors of the target word can be used to update the key-value vector pairs of historical words, i.e., to update the KV Cache, so that the large language model can continue to make predictions.

[0092] In this application, the numbers in the query vector and the key vector of each historical word are all floating-point numbers, such as floating-point numbers of data types FP64, FP32, and FP16 in the IEEE 754 standard format, or floating-point numbers of other data types such as BF16 and FP8.

[0093] In this application, the encoding format of floating-point numbers includes exponent bits and mantissa bits. Figure 4 This is an example diagram of the encoding format for floating-point numbers in this application, such as... Figure 4 As shown, it includes identifier 41, exponent 42, and mantissa 43.

[0094] In this system, the flag bit typically occupies 1 bit and determines whether the value is positive or negative; the exponent bit represents the size or order of magnitude of the value through an offset coding strategy; and the mantissa bit determines the precision of the value within the range determined by the exponent bit. Taking FP32 or BF16 as examples, the exponent bit typically occupies 8 bits, and the mantissa bit typically occupies 7 bits.

[0095] Figure 5 Here is an example diagram of the encoding method for reducing the length of the mantissa in this application, as shown below. Figure 5 As shown, Figure 5 The length of the mantissa indicated by 53 is only Figure 4 Part of the last digit 43.

[0096] In step S302, the electronic device reduces the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, that is, changes the precision of each floating-point number in the vector, so as to reduce the computing power in subsequent calculations.

[0097] In some embodiments of this application, the length of the last digits can be reduced based on the task type of the current prediction task. For example, when using a large language model to predict words in the target language during machine translation of important documents, the accuracy requirement is usually relatively high, so the length of the last digits can be reduced slightly, such as reducing the original 7 digits to 4 or 5 digits; while in some scenarios where the accuracy requirement is relatively low, such as when using a large language model in a smart input method to predict the words or phrases that the user wants to input next, the last digits can be discarded directly.

[0098] In other embodiments of this application, the setting permission for the mantissa length can also be enabled, allowing the user of the large language model prediction task to specify the mantissa length in order to reduce the mantissa length.

[0099] In other embodiments, the appropriate mantissa length for the current large language model can be verified experimentally to reduce the mantissa length of floating-point numbers declared in vector floating-point types.

[0100] In this application, after reducing the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical word, the attention mechanism in the large language model is used to determine the relevance between the query vector and each historical word. The attention mechanism can be a single-head attention mechanism or a multi-head attention mechanism; this application does not impose any restrictions on this. The attention mechanism of this application can be expressed by the following formula (1):

[0101] (1)

[0102] Where Q is the query vector, K is the key vector, and V is the value vector. It is a scalar constant (64 or 128) used to prevent the dot product result from being too large.

[0103] In the above formula (1), The result represents the degree of relevance between the query vector and the key vector of the current historical word, which is also known as the attention weight.

[0104] In step S303 of this application, the electronic device predicts the target word of the generated query vector based on the relevance between the query vector and each historical word and the value vector of the historical words using a large language model.

[0105] In some embodiments of this application, the value vectors of all historical words can be used for prediction. For example, after calculating the relevance of all historical words in the text sequence, the value vector V of the corresponding historical words is weighted based on the relevance, such as multiplying the result of the softmax function in the above formula (1) with V, and then summing the weighted V corresponding to different historical words to generate a new vector representation that integrates global context information. Based on this new vector, the target word can be further processed using a large language model.

[0106] In some other embodiments of this application, the value vectors of some historical words may be selected to participate in subsequent predictions.

[0107] It should be noted that if this application employs a multi-head attention mechanism, in some embodiments, the new vectors generated by different attention heads and incorporating global context information can be combined and then further processed by a large language model to obtain the predicted target word. In other embodiments, the intermediate outputs in the attention mechanism can be processed before generating the new vector, so that even with multiple attention heads, only one new vector incorporating global context information is generated to predict the target word. In related technologies, there is also a scheme to reduce the computational power in the KV Cache processing, but this usually involves quantization, converting the KV Cache data type from high precision (such as FP16, FP32) to low precision (such as INT8, INT4, or even INT2), thereby directly reducing the number of bits occupied by each parameter. However, the above quantization method relies on a uniform scaling factor to map floating-point values ​​to the low-bit range, a process that results in a loss of dynamic range and precision. When the numerical distribution in the KV cache is wide, the scaling factor is easily dominated by extreme values. This causes most small but important values ​​to be submerged in low bit precision or reduced to zero after quantization, resulting in the loss of detailed information and affecting the accuracy of predictions. Furthermore, after quantization, dequantization may be required for further processing, and quantization and dequantization operations introduce additional computational overhead and may lead to distribution distortion, ultimately affecting the accuracy and stability of the model output. These problems are particularly pronounced in long texts and complex inference tasks.

[0108] In contrast, this application does not change the dynamic range of the numerical values ​​(the exponent remains unchanged). It only sacrifices some precision by reducing the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical term. This allows for the effective representation of both extremely large and small values ​​in the widely distributed query vector and / or key vectors, reducing the significant precision loss caused by the uniform scaling factor during quantization. This saves computational power while minimizing prediction accuracy loss. Furthermore, this application does not require quantization-dequantization operations, resulting in better performance in maintaining model stability.

[0109] In some embodiments, reducing the length of the mantissas of the floating-point numbers in the query vector and / or the key vectors of each historical term includes:

[0110] Based on the mapping relationship between the large language model and the length of the mantissa, the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word is reduced; wherein, the mapping relationship includes the reduced mantissa, and the reduced mantissa is determined based on experimental samples. The difference between the prediction accuracy of the large language model before the reduction of the mantissa and the prediction accuracy of the large language model after the reduction of the mantissa is less than a preset difference threshold.

[0111] In this embodiment, the appropriate mantissa length for the large language model is verified experimentally. For example, if the original mantissa has 7 digits, the number of mantissa digits can be adjusted multiple times. The prediction accuracy of the large language model after each adjustment is tested, and then compared with the prediction accuracy before adjustment. The mantissa with the smallest accuracy difference, and the difference being less than a preset difference threshold, is selected as the mantissa length corresponding to the large language model. For example, the preset difference threshold is, for example, 5%.

[0112] It should be noted that, in the embodiments of this application, the appropriate mantissa length can be determined separately for different large language models. Furthermore, when testing the prediction accuracy before and after adjustment for a large language model, the same experimental samples can be used to reduce errors caused by sample differences.

[0113] It is understood that in the embodiments of this application, the mapping relationship between the large language model and the mantissa length is determined by experimentation. This approach is simple and makes the mantissa length more compatible with the large language model, thus balancing the accuracy and efficiency of prediction using the large language model.

[0114] In some embodiments, the step of predicting and generating the target word of the query vector using the large language model based on the relevance between the query vector and each historical word and the value vector of the historical words includes:

[0115] Based on the relevance of the query vector to each historical word, target historical words that meet the preset relevance conditions are selected from the text sequence.

[0116] Based on the relevance between the query vector and the target historical word, as well as the value vector of the target historical word, the target word of the query vector is predicted and generated using the large language model.

[0117] In this embodiment, satisfying the preset relevance condition can mean that the relevance is greater than a preset relevance threshold. For example, the preset relevance threshold can be 0.9. In this embodiment, target historical words are selected based on the preset relevance condition, that is, the "unimportant" parts in the KV Cache are discarded, and only the key information (KV information of the target historical words) is retained. This sparsity method can reduce the amount of data that needs to be processed, thereby helping to save computing power.

[0118] In related technologies, when sparsifying a key-value cache, the values ​​in the K and V vectors are filtered one by one. For example, an "if-else" check is needed for each value to set values ​​that do not meet the conditions to 0. When calling the KV cache, memory addresses occupied by zero values ​​must be skipped to retrieve valid data. This irregular memory skipping disrupts pipeline efficiency and significantly reduces the actual utilization of memory access bandwidth, potentially offsetting or even exceeding the performance gains from reduced data. Furthermore, it may cause memory access bottlenecks. This application, however, filters key-value vector pairs as a whole, rather than focusing on each individual value in the vector, thus reducing the aforementioned memory access bottlenecks.

[0119] In some embodiments, selecting target historical words that meet preset relevance conditions from the text sequence based on the relevance degree between the query vector and each historical word includes:

[0120] The relevance of the query vector to each historical word is sorted to obtain the ranking result of the relevance.

[0121] Based on the ranking results of relevance, target historical words with a sum of relevance values ​​greater than a preset threshold are selected from the text sequence in descending order of relevance.

[0122] In this application, the preset sum threshold is set to, for example, 0.9. The electronic device sorts the relevance levels, for example, in descending order, and finds the top K values ​​whose sum is greater than 0.9. The index of the top K values ​​and the historical words corresponding to the index of the top K values ​​are the target historical words.

[0123] As in the aforementioned formula (1), the relevance is the value after being processed by the softmax function. The softmax function transforms a set of numbers into a probability distribution with a total sum of 1, that is, the sum of the relevance between the key vector of each historical word and the query vector is 1.

[0124] In this application, target historical words are dynamically selected based on the sum of their relevance values. For example, if the sum of the relevance values ​​of a small number of top K historical words is greater than 0.9, it means that these few historical words are most relevant to the current query vector, so only a small number of target historical words need to be selected. However, if the sum of the relevance values ​​of more top K historical words is required to satisfy the condition of being greater than 0.9, then more target historical words need to be selected. By dynamically determining the number of target historical words in this way, computational power can be saved while minimizing the loss of prediction accuracy.

[0125] Figure 6 The performance of selecting target historical words in this application Figure 1 ,like Figure 6 As shown, the horizontal axis 61 represents the number of mantissas retained after reducing the length of the floating-point mantissas in the query vector and / or the key vectors of each historical term; 62 represents the overlap between the target historical terms selected based on the number of retained mantissas and the target historical terms selected without reducing the mantissa length (i.e., topK). Taking an original length of 7 mantissas as an example... Figure 6 As shown, when retaining 3 last digits, the overlap rate reaches 90%, and when retaining 5 or 6 last digits, the overlap rate of the target historical words is close to 100%.

[0126] Understandably, even reducing the length of the last digit by one reduces the computational load in each relevance calculation, especially when there are many historical words (i.e., in long text prediction tasks), and the effect is more significant while maintaining prediction accuracy.

[0127] Figure 7 The performance of selecting target historical words in this application Figure 2 ,like Figure 7 As shown, this figure illustrates the sparsity rates of different network layers in a large language model when maintaining a sum of relevance values ​​greater than 90%. The horizontal axis (Figure 71) indicates the network layer, and the vertical axis (Figure 72) indicates the sparsity rate, i.e., the proportion of the selected target historical word to all historical words. The figure shows that the sparsity rates can vary across different network layers, with an overall average sparsity of approximately 80%. It is understandable that sparsifying each network layer in the attention mechanism effectively saves computational resources.

[0128] Figure 8 Example diagram of attention calculation mechanism in related technologies, such as Figure 8 As shown, taking the single-head attention mechanism as an example, for the floating-point number 81 in the query vector and the floating-point number 82 in the key vector of historical words, during the dot product operation, the mantissa is multiplied, the exponent is added, and then the result of the exponent calculation is shifted based on the result of the mantissa calculation to obtain the result of 83. The result of 83 is then divided by... Then, the correlation level of 84 is obtained by using the softmax function.

[0129] Figure 9 Here is an example diagram of the attention calculation mechanism in this application, such as Figure 9 As shown, taking the single-head attention mechanism as an example, in this example, the floating-point number 91 in the query vector and the floating-point number 92 in the key vector of the history word dot product operation do not require the mantissa; only the addition result of the exponent digits is used to obtain the dot product result 93. The calculation result based on 93 is divided by... Then, the softmax function yields a relevance score of 94. By selecting the top K, the target historical words corresponding to the two selected indices 0 and 2, as shown in image 95, are obtained. It's understandable that in this example, removing the mantissa and performing a floating-point dot product, combined with KV cache sparsity, effectively saves computational power.

[0130] In some embodiments, the relevance level is a floating-point number with an exponent and a mantissa in the encoding format; the process of sorting the relevance levels of the query vector with each historical word to obtain a ranking result of the relevance levels includes:

[0131] The index values ​​in each correlation level are sorted to obtain the initial ranking of correlation levels.

[0132] In response to the initial sorting result including multiple target relevance degrees with the same exponent, the last digit of each target relevance degree is sorted to obtain a sorting result after reordering the multiple target relevance degrees in the initial sorting result; wherein, the direction of sorting the last digit is consistent with the direction of sorting the exponent.

[0133] In this application, as mentioned above, the exponent indicates the size or order of magnitude of the value, that is, approximately how large the number is, while the mantissa indicates the precision of the value within the range determined by the exponent. Therefore, it can be understood that the larger the exponent of a floating-point number, the greater the correlation, regardless of the mantissa. Furthermore, with the same number of exponents, the larger the mantissa, the greater the correlation.

[0134] Based on this, this application prioritizes sorting the exponent of the relevance degree. If there are target relevance degrees with the same exponent, then the last digit of the target relevance degree is sorted to obtain the final sorting result.

[0135] For example, the BF16 code for a decimal bit correlation degree of 0.8 is: 0-01111110-1001100, where 0 is the identifier bit, "01111110" is the exponent bit, and "1001100" is the mantissa bit; the BF16 code for a decimal bit correlation degree of 0.2 is: 0-01111100-1001100. By comparing the binary code bits, it can be seen that the value represented by the exponent bit of 0.8 (126) is greater than the value of the exponent bit of 0.2 (124), and the mantissa bits are the same.

[0136] For example, the BF16 code of 0.86 is: 0-01111110-1011100; the BF16 code of 0.24 is: 0-01111101-1110000. By comparing the binary code bits, it can be seen that the value represented by the exponent bit of 0.86 (126) is greater than the value of the exponent bit of 0.24 (125), and the mantissas are different.

[0137] For example, the BF16 encoding of 0.75 is: 0-01111110-000000; the BF16 encoding of 0.625 is: 0-01111110-0100000. By comparing the binary encoding bits, it can be seen that the exponent bits of the two are the same, but the mantissa bits are different. The value represented by the mantissa bits of 0.75 is greater than the value represented by the mantissa bits of 0.625.

[0138] As can be seen from the above, the sorting of the exponent and mantissa is consistent with the sorting of relevance. Therefore, in this application, by adopting a floating-point sorting method that compares the exponent first and then the mantissa, it is not necessary to convert it to a decimal value. Instead, the original bit pattern is directly compared, which essentially transforms the comparison operation into an efficient integer bit-level comparison, thereby significantly improving sorting performance.

[0139] In some embodiments, sorting the relevance between the query vector and each historical word to obtain a relevance ranking result includes:

[0140] The relevance of the query vector to each historical term is amplified by a predetermined factor and then sorted to obtain the ranking result of the relevance.

[0141] In this application, since the correlation degree is less than 1, there is a high probability that the exponents are the same. If the exponents are the same, the last digits need to be compared again. In this regard, the embodiments of this application can first increase the correlation degree by a predetermined factor before sorting, such as increasing it by 10 times. By increasing the factor, the difference in the exponents can be widened, thereby reducing the number of eliminations and also helping to improve the sorting performance.

[0142] In some embodiments, the attention mechanism is a multi-head attention mechanism; reducing the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical word, and using the attention mechanism in the large language model to determine the relevance between the query vector and each historical word, includes:

[0143] Reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, and use the attention mechanism in the large language model to determine the relevance score between the query vector and each historical word corresponding to each attention head;

[0144] For each historical word, the average relevance score between the query vector and the historical word is determined based on the relevance score between the query vector and the historical word corresponding to different attention heads and the number of attention heads.

[0145] The relevance scores of each historical term are averaged using an activation function to obtain the relevance degree between the query vector and each historical term.

[0146] In this application, the electronic device utilizes the attention mechanism in a large language model to determine the relevance score between the query vector corresponding to each attention head and each historical word. The relevance score is... The corresponding results.

[0147] For example, suppose the number of attention heads `head_num` = 2, the number of historical words in the text sequence `seqlen` = 4, the relevance score corresponding to one attention head is [2,3,5,7], and the relevance score corresponding to another attention head is [1,2,4,5], where a value in the matrix represents the relevance score corresponding to a historical word. Then the summation along the `head_num` dimension is [3,5,9,12], and the average relevance score based on the number of attention heads is [1.5,2.5,4.5,6], where a value in the matrix represents an average relevance score corresponding to a historical word.

[0148] In this application, the activation function is the aforementioned softmax function, and the relevance between the query vector processed by the activation function and each historical word is [0.0088, 0.0239, 0.1765, 0.7908].

[0149] It is understandable that in this application, a set of relevance scores from different attention heads is statistically obtained, which enables the generation of a new vector that integrates global context information based on a set of relevance scores for prediction. This saves computational power compared to generating multiple new vectors that integrate global context information based on multiple sets of relevance scores from multiple attention heads and combining them for prediction.

[0150] In some embodiments, the numbers in the value vectors of historical words are floating-point numbers, and the encoding format of floating-point numbers includes exponent bits and mantissa bits; the step of predicting and generating the target word of the query vector using the large language model based on the relevance between the query vector and each historical word and the value vectors of the historical words includes:

[0151] The length of the mantissa of the floating-point numbers in the value vector of historical words is reduced, and the target word of the query vector is predicted and generated using the large language model, taking into account the relevance between the query vector and the historical words.

[0152] In this application, the length of the mantissa in the value vector can also be reduced. Specifically, the reduction in the length of the mantissa of the floating-point numbers in the value vector can be determined by referring to the aforementioned method for determining the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical term.

[0153] Understandably, this application can reduce computational power with minimal loss of prediction accuracy by reducing the length of the mantissa of the value vector.

[0154] It should be noted that the multiple embodiments in the above embodiments of this application can be arbitrarily combined without conflict. For example, reducing the length of the last few bits of the floating-point numbers in the query vector and / or the key vector of each historical word, selecting the sorting method of the target historical words, and reducing the length of the last few bits of the value vector can be combined to reduce computing power as much as possible while taking into account the prediction accuracy.

[0155] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario. See also... Figure 10 , Figure 10 This is a flowchart illustrating another data processing method provided in this application embodiment. For example, in the process of automatically supplementing sentences using target words generated by large language models in text creation, during the attention processing stage of the large language model, the overall computational power and power consumption can be reduced by combining the truncation of the mantissa of floating-point numbers (reducing the number of mantissas) and the sparsity of key-value vector pairs. Figure 10 The steps described herein illustrate the data processing method provided in the embodiments of this application.

[0156] S1001. Obtain the query vector Q, key vector K, and value vector V.

[0157] Here, the key vector K and the value vector V are the key vector and value vector in the key-value vector pair in this application.

[0158] S1002, the attention score operator is used for calculation.

[0159] In this application, the attention scoring operator is the one mentioned above. Before calculating the attention score, the length of the mantissa of floating-point numbers in the query vector and / or key-value vector can be reduced.

[0160] S1003, Obtain the attention score Result_1.

[0161] In this application, the attention score is the relevance score. For example, the shape of Result_1 can be [1, head_num, seqlen], where 1 indicates that the current task request is 1, such as only a text creation task; head_num indicates the number of attention heads, for example, one head may focus on grammatical structure, and another head may focus on referential relations; seqlen indicates the number of historical words in the text sequence.

[0162] S1004, average after summation.

[0163] In this application, the average score after summation is determined by first summing result_1 along the head_num dimension, and then dividing the summed value by head_num. That is, in this application, for each historical word, the average relevance score between the query vector and the historical word is determined based on the relevance score between the query vector and the historical word corresponding to different attention heads and the number of attention heads.

[0164] S1005, Obtain the attention score Result_2.

[0165] In this application, the attention score Result_2 is the average relevance score between the query vector and historical words. For example, the shape of Result_2 can be [1, seqlen].

[0166] S1006, Activation function is used for calculation.

[0167] In this application, the activation function calculation involves processing the average relevance score of each historical term using the activation function to obtain the relevance degree between the query vector and each historical term. The activation function is the softmax function.

[0168] S1007, Obtain attention weight Result_3.

[0169] In this application, the attention weight Result_3 represents the relevance between the query vector and each historical word. For example, the shape of result_3 is also [1, seqlen].

[0170] S1008. Sort, calculate the top K values ​​when the sum of the top K values ​​is greater than 0.9, and the indices corresponding to the top K values.

[0171] In this application, the sorting and calculation of the topK values ​​when the sum of the topK values ​​is greater than 0.9 is the sorting result based on the degree of relevance. The relevance is selected from the text sequence in descending order of relevance, and the relevance values ​​between the relevance values ​​are selected when the sum of the relevance values ​​is greater than a preset sum threshold. The indices of these relevance values ​​are obtained to filter key-value vector pairs.

[0172] S1009, Obtain index Result_4.

[0173] In this application, the length of Result_4 is topK, and the index is the number of the KV Cache that needs to be retained.

[0174] S1010, KV extraction based on index.

[0175] S1011, obtain the sparsed KV cache.

[0176] In this application, KV is extracted based on the index to obtain a sparse KV cache, that is, the key-value vector pairs of the target historical word are extracted from the KV cache, that is, a key vector and a value vector of the target historical word are obtained.

[0177] In this application, after obtaining the key-value vector pairs of the target historical words, the target words of the query vector can be predicted and generated using a large language model based on the relevance between the query vector and the target historical words and the value vector of the target historical words.

[0178] In this embodiment, by reducing the number of mantissas and obtaining a set of relevance scores based on multi-head attention to sparsify the key-value vector pairs, overall computational power and power consumption can be reduced while maintaining prediction accuracy. The following continues to describe the exemplary structure of the data processing device 255 provided in this embodiment as a software module. In some embodiments, such as... Figure 2 As shown, the software modules stored in the data processing device 255 of the memory 250 may include:

[0179] The vector acquisition module 2551 is configured to acquire the query vector to be queried and the key-value vector pairs of each historical word in the text sequence; wherein, the key vector in each key-value vector pair of the historical word represents the feature of the historical word, and the value vector represents the contextual feature of the historical word in the text sequence; the numbers in the query vector and the key vector of each historical word are all floating-point numbers, and the encoding format of the floating-point numbers includes exponent bits and mantissa bits;

[0180] Attention module 2552 is configured to reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, and to use the attention mechanism in the large language model to determine the relevance between the query vector and each historical word.

[0181] The prediction module 2553 is configured to predict and generate the target word of the query vector based on the relevance between the query vector and each historical word and the value vector of the historical words using the large language model.

[0182] In some embodiments, the attention module 2552 is further configured to reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word based on the mapping relationship between the large language model and the mantissa length; wherein the mapping relationship includes the reduced mantissa, and the reduced mantissa is determined based on experimental samples. The difference between the prediction accuracy of the large language model before the mantissa reduction and the prediction accuracy of the large language model after the mantissa reduction is less than a preset difference threshold.

[0183] In some embodiments, the prediction module 2553 is further configured to select target historical words that meet preset relevance conditions in the text sequence based on the relevance degree between the query vector and each historical word; and to predict and generate the target word of the query vector using the large language model based on the relevance degree between the query vector and the target historical word and the value vector of the target historical word.

[0184] In some embodiments, the prediction module 2553 is further configured to sort the relevance between the query vector and each historical word to obtain a relevance ranking result; based on the relevance ranking result, select target historical words from the text sequence whose sum of relevance is greater than a preset sum threshold in descending order of relevance.

[0185] In some embodiments, the correlation degree is a floating-point number with an encoding format including an exponent and a mantissa; the prediction module 2553 is further configured to sort the exponents of each correlation degree to obtain an initial sorting result of the correlation degree; in response to multiple target correlation degrees including the same exponent in the initial sorting result, the mantissas of each target correlation degree are sorted to obtain a sorting result after reordering the multiple target correlation degrees in the initial sorting result; wherein, the direction of sorting the mantissas is consistent with the direction of sorting the exponents.

[0186] In some embodiments, the prediction module 2553 is further configured to amplify the relevance of the query vector to each historical word by a predetermined factor and then sort them to obtain a ranking result of the relevance.

[0187] In some embodiments, the attention mechanism is a multi-head attention mechanism; the attention module 2552 is further configured to reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, and to use the attention mechanism in the large language model to determine the relevance score between the query vector and each historical word corresponding to each attention head; for each historical word, based on the relevance score between the query vector and the historical word corresponding to different attention heads and the number of attention heads, to determine the average relevance score between the query vector and the historical word; and to use an activation function to process the average relevance score corresponding to each historical word to obtain the degree of relevance between the query vector and each historical word.

[0188] In some embodiments, the numbers in the value vector of historical words are floating-point numbers, and the encoding format of floating-point numbers includes exponent bits and mantissa bits; the prediction module 2553 is further configured to reduce the length of the mantissa bits of the floating-point numbers in the value vector of historical words, and combine the relevance between the query vector and historical words to predict and generate the target word of the query vector using the large language model.

[0189] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or executable instructions from the computer-readable storage medium and executes the computer program or executable instructions, causing the electronic device to perform the method described in this application.

[0190] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions. When the computer program or executable instructions are executed by a processor, the processor will execute the data processing method provided in this application. For example, ... Figure 3 The data processing method is shown.

[0191] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0192] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0193] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0194] As an example, computer-executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0195] In summary, the embodiments of this application can save computing power and reduce power consumption while minimizing the loss of prediction accuracy.

[0196] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the query vector to be queried and the key-value vector pairs of each historical word in the text sequence; wherein, the key vector in each key-value vector pair of historical words represents the features of the historical word, and the value vector represents the context features of the historical word in the text sequence; the numbers in the query vector and the key vector of each historical word are all floating-point numbers, and the encoding format of floating-point numbers includes exponent bits and mantissa bits; Based on the mapping relationship between the large language model and the length of the mantissa, the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical word is reduced, and the attention mechanism in the large language model is used to determine the relevance between the query vector and each historical word; wherein, the mapping relationship includes the reduced mantissa, and the reduced mantissa is determined based on experimental samples, which is the prediction accuracy of the large language model before the mantissa reduction, and the difference between the prediction accuracy of the large language model before the mantissa reduction and the prediction accuracy of the large language model after the mantissa reduction is less than a preset difference threshold; Based on the relevance of the query vector to each historical word, target historical words that meet the preset relevance conditions are selected from the text sequence. Based on the relevance between the query vector and the target historical word and the value vector of the target historical word, the target word of the query vector is predicted and generated using the large language model; The attention mechanism is a multi-head attention mechanism; reducing the length of the mantissa of the floating-point numbers in the query vector and / or the key vectors of each historical word, and using the attention mechanism in the large language model to determine the relevance between the query vector and each historical word, includes: Reduce the length of the mantissa of the floating-point numbers in the query vector and / or the key vector of each historical word, and use the attention mechanism in the large language model to determine the relevance score between the query vector and each historical word corresponding to each attention head; For each historical word, the average relevance score between the query vector and the historical word is determined based on the relevance score between the query vector and the historical word corresponding to different attention heads and the number of attention heads. The relevance scores of each historical term are averaged using an activation function to obtain the relevance degree between the query vector and each historical term.

2. The method according to claim 1, characterized in that, The step of selecting target historical words that meet preset relevance conditions from the text sequence based on the relevance degree between the query vector and each historical word includes: The relevance of the query vector to each historical word is sorted to obtain the ranking result of the relevance. Based on the ranking results of relevance, target historical words with a sum of relevance values ​​greater than a preset threshold are selected from the text sequence in descending order of relevance.

3. The method according to claim 2, characterized in that, The relevance level is a floating-point number in the encoding format, including the exponent and mantissa; the process of sorting the relevance levels of the query vector with each historical word to obtain the relevance ranking result includes: The index values ​​in each correlation level are sorted to obtain the initial ranking of correlation levels. In response to the initial sorting result including multiple target relevance degrees with the same exponent, the last digit of each target relevance degree is sorted to obtain a sorting result after reordering the multiple target relevance degrees in the initial sorting result; wherein, the direction of sorting the last digit is consistent with the direction of sorting the exponent.

4. The method according to claim 2 or 3, characterized in that, The process of sorting the relevance between the query vector and each historical word to obtain the relevance ranking result includes: The relevance of the query vector to each historical term is amplified by a predetermined factor and then sorted to obtain the ranking result of the relevance.

5. The method according to claim 1, characterized in that, The values ​​in the historical word value vectors are floating-point numbers, and the encoding format of floating-point numbers includes exponent bits and mantissa bits; the step of predicting and generating the target word of the query vector using the large language model based on the relevance between the query vector and each historical word and the value vectors of the historical words includes: The length of the mantissa of the floating-point numbers in the value vector of historical words is reduced, and the target word of the query vector is predicted and generated using the large language model, taking into account the relevance between the query vector and the historical words.

6. A data processing apparatus, characterized in that, The device includes: The vector acquisition module is configured to acquire the query vector to be queried and the key-value vector pairs of each historical word in the text sequence; wherein, the key vector in each key-value vector pair of the historical word represents the feature of the historical word, and the value vector represents the contextual feature of the historical word in the text sequence; the numbers in the query vector and the key vector of each historical word are all floating-point numbers, and the encoding format of the floating-point numbers includes exponent bits and mantissa bits; An attention module is configured to reduce the length of the last digit in the query vector and / or the key vectors of each historical word based on the mapping relationship between the large language model and the last digit length. It then uses the attention mechanism in the large language model to determine the relevance score between the query vector and each historical word corresponding to each attention head. This attention mechanism is a multi-head attention mechanism. For each historical word, based on the relevance score between the query vector and the historical word corresponding to different attention heads and the number of attention heads, the average relevance score between the query vector and the historical word is determined. An activation function is used to process the average relevance score corresponding to each historical word to obtain the degree of relevance between the query vector and each historical word. The mapping relationship includes the reduced last digit length, which is determined based on experimental samples. The difference between the prediction accuracy of the large language model before the last digit reduction and the prediction accuracy of the large language model after the last digit reduction is less than a preset difference threshold. The prediction module is configured to select target historical words that meet preset relevance conditions in the text sequence based on the relevance degree between the query vector and each historical word; and to predict and generate the target word of the query vector using the large language model based on the relevance degree between the query vector and the target historical word and the value vector of the target historical word.

7. An electronic device, characterized in that, include: processor; Memory used to store computer programs or instructions; The processor executes the computer program or instructions to implement the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Bidirectional block floating point-based large language model reasoning acceleration method

    CN120851185A