Data processing method and equipment
By compressing and filtering the input key-value data, combined with attention weight, the computing resource consumption and duration of the data processing model are reduced, the problem of low processing efficiency under large data volume is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202510480398.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
As the amount of key-value data increases, the time and computing resources consumed by the data processing model when processing key-value data increases significantly, and the prior art is difficult to effectively reduce.
By compressing the input key-value data, the target cache capacity is determined, and the target key-value data is filtered based on the attention weight, and the target key-value data is finally fused with some of the pending key-value data to obtain the output data.
It significantly reduces the amount of data and computing resources required to process each time the output data is obtained, shortens the processing time and improves the data processing efficiency.
Smart Images

Figure CN120336273A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly relates to a data processing method and device. Background Art
[0002] When a data processing model processes data, it will convert the data to be processed into corresponding key-value data, and then process these key-value data to obtain corresponding outputs. For example, when a large language model processes text, it converts the text to be processed into key-value data, processes the key-value data to obtain output data, and then converts the output data into corresponding text and outputs the text.
[0003] In the above processing method, as the data volume of the key-value data increases, the time and computing resources consumed for processing the key-value data will also increase significantly. Therefore, how to reduce the time and computing resources consumed by the data processing model has become an urgent problem to be solved. Summary of the Invention
[0004] For this reason, this application discloses the following technical solutions:
[0005] The first aspect of this application provides a data processing method, including:
[0006] Compress the input key-value data corresponding to the input text unit to obtain the key-value data to be processed;
[0007] Determine the target cache capacity according to the key-value data to be processed and a preset initial cache capacity, where the target cache capacity is the capacity of the target cache area for caching the target key-value data;
[0008] Based on the target cache capacity and the attention weights of the key-value data to be processed, screen out the target key-value data from the multiple key-value data to be processed, where the attention weights are calculated from the key-value data to be processed;
[0009] Fuse the target key-value data and a part of the key-value data to be processed to obtain output data.
[0010] Optionally, the determining the target cache capacity according to the key-value data to be processed and a preset initial cache capacity includes:
[0011] Obtain the target key-value data obtained by the most recent two screenings;
[0012] Adjust the initial cache capacity according to the key-value data to be processed and the target key-value data obtained by the most recent two screenings to obtain the target cache capacity.
[0013] Optionally, the adjusting the initial cache capacity according to the key-value data to be processed and the target key-value data obtained by the most recent two screenings includes:
[0014] Adjust the initial cache capacity according to the data volume of the key-value data to be processed and the difference degree of the target key-value data obtained from the last two screenings.
[0015] Optionally, screening out target key-value data from multiple pieces of the key-value data to be processed based on the target cache capacity and the attention weights of the key-value data to be processed includes:
[0016] Screen the key-value data to be processed according to the attention weights of the key-value data to be processed to obtain a screening result;
[0017] Determine target key-value data from the key-value data to be processed according to the screening result and the target cache capacity, where the quantity of the target key-value data matches the target cache capacity.
[0018] Optionally, screening the key-value data to be processed according to the attention weights of the key-value data to be processed to obtain a screening result includes:
[0019] Screen in the key-value data to be processed with a relative distance greater than or equal to a preset segmentation parameter according to the attention weights of the key-value data to be processed to obtain a screening result, where the relative distance is the distance between the corresponding key-value data to be processed and the last piece of key-value data to be processed.
[0020] Optionally, the screening result includes multiple screening results obtained through multiple screenings;
[0021] Determining target key-value data from the key-value data to be processed according to the screening result and the target cache capacity includes:
[0022] Determine the target key-value data according to the target cache capacity and the number of times the key-value data to be processed appears repeatedly in the multiple screening results.
[0023] Optionally, fusing the target key-value data and a part of the key-value data to be processed to obtain output data includes:
[0024] Fuse the target key-value data and a part of the key-value data to be processed according to the number of times the target key-value data appears repeatedly to obtain output data.
[0025] Optionally, fusing the target key-value data and a part of the key-value data to be processed to obtain output data includes:
[0026] Fuse the target key-value data and the to-be-processed key-value data with a relative distance less than a preset threshold to obtain output data, where the relative distance is the distance between the corresponding to-be-processed key-value data and the last to-be-processed key-value data.
[0027] Optionally, the compressing the input key-value data corresponding to the input text unit to obtain to-be-processed key-value data includes:
[0028] Determine the number K of ranks to be retained according to the input key-value data corresponding to the input text unit;
[0029] Perform dimensionality reduction processing on the input key-value data according to the first K ranks of the input key-value data to obtain to-be-processed key-value data.
[0030] A second aspect of the present application provides a data processing device, including a processor and a memory module;
[0031] The memory module is used to cache key-value data;
[0032] The processor is used to execute a preset computer program to perform the following data processing method:
[0033] Compress the input key-value data corresponding to the input text unit to obtain to-be-processed key-value data;
[0034] Determine a target cache capacity according to the to-be-processed key-value data and a preset initial cache capacity, where the target cache capacity is the capacity of a target cache area for caching target key-value data;
[0035] Based on the target cache capacity and the attention weight of the to-be-processed key-value data, screen out target key-value data from multiple to-be-processed key-value data, where the attention weight is calculated from the to-be-processed key-value data;
[0036] Fuse the target key-value data and a part of the to-be-processed key-value data to obtain output data. Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0038] Figure 1 It is a flowchart of a data processing method provided by an embodiment of the present application;
[0039] Figure 2It is a flowchart of a method for determining the number K of ranks to be retained provided by an embodiment of the present application;
[0040] Figure 3 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0042] The key-value cache (KV-Cache) technology is an important technology for improving the processing efficiency of data processing models (such as large language models). The principle of the key-value cache technology is that the data processing model calculates multiple input text units, obtains the key data and value data corresponding to each text unit, and caches these key data and value data in the memory module. In the subsequent processing process, the data processing model can directly read the cached key data and value data from the memory module when needed and perform calculations, without having to recalculate this part of the data, thereby achieving the effect of improving the processing efficiency. These key data and value data cached in the memory module can be collectively referred to as the key-value cache (KV-Cache) data of the text unit.
[0043] The text units generated during the processing of the data processing model will continue to be input into the data processing model to generate new text units. Therefore, the data volume of the key-value cache data of the text unit will continuously increase, which has an adverse impact on the processing efficiency of the model.
[0044] In view of the above problems, this embodiment provides a data processing method. Please refer to Figure 1 , which is a flowchart of this method. This method may include the following steps.
[0045] S101, compress the input key-value data corresponding to the input text unit to obtain the key-value data to be processed.
[0046] The data processing method of this embodiment can be applied to any data processing model capable of processing text data. For example, this embodiment's data processing method can be applied when using a large language model (LLM) to process text data to reduce the time and resources consumed by the LLM in processing text data.
[0047] The input text units can be obtained by splitting the text to be processed that can be acquired by the data processing model. After obtaining the text to be processed, the text to be processed can be split into multiple text units through various word segmentation methods in related technologies. These text units are the input text units. The text to be processed can be a document uploaded by the user or loaded from a specified path, or it can be a sentence input by the user. The text units can be characters, words, phrases, clauses, words, punctuation marks, etc.
[0048] As an example, the text to be processed can be "Recommend some cost-effective mobile phones", and the input text units obtained by splitting can include: "Recommend", "some", "high", "cost-effective", "of", "mobile phones".
[0049] The input key-value data corresponding to the input text unit includes the key data corresponding to the input text unit and the value data corresponding to the input text unit.
[0050] Each text unit can be converted into a corresponding feature vector through the Word Embeddings technology, and the feature vectors corresponding to different text units are different. For the feature vector corresponding to a text unit, the key matrix contained in the data processing model can be used to calculate the feature vector to obtain the key data corresponding to the text unit, denoted as Key’, and the value matrix contained in the data processing model can be used to calculate the feature vector to obtain the value data corresponding to the text unit, denoted as Value’. Among them, the key matrix and the value matrix belong to the model parameters of the data processing model and can be determined when training the data processing model. The method of calculating the key data and the value data according to the key matrix and the value matrix can refer to related technologies and will not be elaborated.
[0051] In step S101, the key data Key’ of the input key-value data can be compressed to obtain the key data to be processed Key, and the value data Value’ can be compressed to obtain the value data to be processed Value. Each piece of key-value data to be processed includes a piece of key data to be processed Key and a piece of value data to be processed Value. Each piece of key-value data to be processed corresponds to a text unit.
[0052] The data volume of the key-value data to be processed is smaller than that of the input key-value data. Exemplarily, both the key data Key’ and the value data Value’ in the input key-value data can be 256-dimensional vector data, and the key data to be processed Key and the value data to be processed Value after compression can both be 100-dimensional vector data.
[0053] In this embodiment, a text unit is equivalent to a token. The key-value data before compression processing is equivalent to the key-value cache (KV-Cache) data obtained by directly calculating the text unit in the related art. The key-value data to be processed is equivalent to the compressed key-value cache data obtained by compressing the key-value cache data obtained in the related art.
[0054] S102. Determine a target cache capacity according to the key-value data to be processed and a preset initial cache capacity. The target cache capacity is the capacity of a target cache area for caching target key-value data.
[0055] The target key-value data is a part of the key-value data screened from the key-value data to be processed. The target cache area is a storage space in the memory of a processor (such as a graphics processing unit GPU) for caching the target key-value data. After determining the target cache capacity through S102, a target cache area of the corresponding size can be allocated in the memory according to the target cache capacity.
[0056] The initial cache capacity can be configured according to the actual situation without limitation. For example, the initial cache capacity can be the same as the data volume of all the key-value data to be processed obtained for the first time, or it can be half of the data volume.
[0057] S103. Screen out the target key-value data from multiple key-value data to be processed based on the target cache capacity and the attention weights of the key-value data to be processed. The attention weights are calculated from the key-value data to be processed.
[0058] In step S103, the key-value data to be processed can be calculated to obtain the attention weights corresponding to the key-value data to be processed, and then, according to the attention weights, multiple target key-value data with a quantity matching the target cache capacity are screened out from multiple key-value data to be processed. For example, if the target cache capacity is the data volume of 20 target key-value data, 20 key-value data to be processed are screened out as the target key-value data according to the attention weights.
[0059] S104. Fuse the target key-value data and a part of the key-value data to be processed to obtain output data.
[0060] In step S104, the query data Query corresponding to the last text unit among multiple input text units can be obtained, the weight parameter of the query data relative to the key data Key contained in the target key-value data and a part of the key-value data to be processed is calculated, and the target key-value data and the value data to be processed contained in a part of the key-value data to be processed are fused based on the weight parameter to obtain the output data.
[0061] After obtaining an output data, the output data can be converted into an output text unit through the word vector technology. The conversion method can refer to the related art and will not be elaborated here.
[0062] The query data to be processed corresponding to the text unit can be obtained as follows:
[0063] Use the query matrix contained in the data processing model to calculate the feature vector corresponding to the text unit, obtain the query data Query' corresponding to the text unit, and compress the query data Query' in the same compression manner as in S101 to obtain the query data to be processed corresponding to the text unit. Among them, the query matrix belongs to the model parameters of the data processing model and can be determined when training the data processing model. The method for calculating the query data based on the query matrix can refer to the related technology and will not be elaborated.
[0064] When step S104 is executed for the first time, the last text unit can be the text unit at the end of the text to be processed. Combining the foregoing example, for the text to be processed "Recommend some cost-effective mobile phones", the text unit at its end can be "mobile phones".
[0065] The method provided in this embodiment can be executed multiple times until the output stops, so as to obtain a sentence composed of multiple consecutive output text units.
[0066] Combining the foregoing example, after obtaining the text to be processed "Recommend some cost-effective mobile phones", the large language model can be used to obtain the following sentence composed of multiple output text units by executing the method of this embodiment multiple times:
[0067] "Mobile phones of model A1 of brand A, mobile phones of models B1 and B2 of brand B, mobile phones of model C1 matching brand C, all of the above mobile phones have high cost performance and can be purchased from the XX e-commerce platform".
[0068] In the case of repeated execution, it can be determined by the data processing model (such as a large language model) whether to stop the output, and the implementation method can refer to the related technology and will not be elaborated.
[0069] Among them, after each execution of S104, the obtained output data can be converted into an output text unit, use the key matrix and value matrix of the data processing model to calculate the feature vector corresponding to the output text unit, obtain the key data Key' and value data Value' corresponding to the output text unit, compress the key data Key' and value data Value' according to the method of S101, add the obtained key-value data as new query key-value data to the existing query key-value data, and then continue to process according to the methods of steps S102 to S104 to obtain the next output text unit.
[0070] Alternatively, after each execution of S104, the obtained output data can be used as the feature vector corresponding to the output text unit. The key matrix of the data processing model is used to calculate the output data to obtain the corresponding key data to be processed, Key. The value matrix of the data processing model is used to calculate the output data to obtain the corresponding value data to be processed, Value. The key data to be processed, Key, and the value data to be processed, Value, are added as a new key-value data to be processed to the existing key-value data to be processed, and then continue to be processed according to the methods of steps S102 to S104 to obtain the next output data and the corresponding output text unit.
[0071] In the process of repeatedly executing the above method, steps S102 and S103 can be selectively executed according to a certain rule. That is to say, after adding the key-value data of an output text unit to the key-value data to be processed each time, the target cache capacity can be re-determined and the target key-value data can be re-screened, and then S104 is executed to obtain the next output text unit; or the target cache capacity can not be re-determined and the target key-value data can not be re-screened, and S104 is directly executed to obtain the next output text unit.
[0072] After each execution of S102, the determined target cache capacity can be the same as or different from the target cache capacity before the execution of S102. After each execution of S103, the selected target key-value data can be the same as or different from the target key-value data before the execution of S103.
[0073] When repeatedly executing the method of this embodiment multiple times, each obtained output text unit can be added as a new input text unit after the existing input text units. Correspondingly, starting from the second execution of S104, the last text unit in each execution of S104 can be the output text unit obtained in the previous time.
[0074] Combined with the foregoing example, when first executing S104, the query data to be processed, Query, corresponding to "mobile phone" can be used for fusion to obtain the output text unit "Brand A";
[0075] When the second execution of S104, the query data to be processed, Query, corresponding to "Brand A" can be used for fusion to obtain the output text unit "of";
[0076] When the third execution of S104, the query data to be processed, Query, corresponding to "of" can be used for fusion to obtain the output text unit "Model A1";
[0077] And so on, finally obtaining the sentence composed of the above multiple output text units.
[0078] Optionally, for each output text unit, the output data corresponding to the output text unit can also be directly calculated using the query matrix, and the obtained result is used as the query data to be processed corresponding to the output text unit.
[0079] The beneficial effects of this embodiment are as follows:
[0080] On the one hand, the above method compresses the input key-value data to obtain the key-value data to be processed, and processes the key-value data to be processed, which can reduce the data volume of each key-value data used for processing;
[0081] On the other hand, after obtaining the key-value data to be processed, the above method screens out the target key-value data from it, and then fuses the target key-value data and a part of the key-value data to be processed to obtain the output data, without the need to process all the data to be processed. In this way, the number of key-value data to be processed can be reduced;
[0082] Therefore, the method of this embodiment significantly reduces the data volume to be processed each time the output data is obtained by reducing the data volume of each key-value data and the number of key-value data to be processed, achieving the effects of shortening the time and reducing the consumed computing resources.
[0083] Optionally, the method of compressing the input key-value data in S101 may include:
[0084] Determine the number K of ranks to be retained according to the input key-value data corresponding to the input text unit;
[0085] Perform dimensionality reduction processing on the input key-value data according to the first K ranks of the input key-value data to obtain the key-value data to be processed.
[0086] When determining the quantity K, on the one hand, preset initial parameters can be obtained, including increment i, difference m, initial rank k, and information quantity threshold v. Both k and i are integers, and m and v can be percentages greater than 0 and less than 1, such as 99%. The initial rank k can be set to 100, or can be set to other integers, without limitation.
[0087] On the other hand, the input key-value data can be combined into a matrix A to be decomposed.
[0088] The combination method is not limited. As an example, the key data Key’ corresponding to each input text unit can be used as a row of the matrix A, and the value data Value’ corresponding to each input text unit can also be used as a row of the matrix A, thereby obtaining the matrix A to be decomposed.
[0089] After obtaining the initial parameters and the matrix A, the quantity K can be determined according to the Figure 2 method shown.
[0090] S201, perform singular value decomposition on the matrix A.
[0091] S202, Calculate the current information amount based on the singular values of the top k ranks.
[0092] S203, Calculate the difference value between the current information amount and the information amount threshold.
[0093] S204, Determine whether the difference value is less than the difference quantity.
[0094] If the difference value is not less than the difference quantity, execute S205; if the difference value is less than the difference quantity, execute S206.
[0095] S205, Update the value of k in increments.
[0096] After executing step S205, execute step S202 again based on the updated value of k.
[0097] S206, Determine the current value of k as the number K of ranks to be retained.
[0098] In step S201, the matrix A can be decomposed into singular values based on the following formula (1) to obtain the left eigenmatrix U, the singular value matrix Σ, and the right eigenmatrix V corresponding to the matrix A.
[0099]
[0100] Among them, the elements on the diagonal of the singular value matrix Σ are the singular values of the matrix A, and the other elements are equal to 0. Each singular value corresponds to a left singular vector of the matrix A, and all the left singular vectors corresponding to the left singular values form the left eigenmatrix U. Each singular value corresponds to a right singular vector of the matrix A, and all the right singular vectors corresponding to the right singular values form the right eigenmatrix V. T represents the transpose operation on the matrix V. The method for decomposing the matrix A into singular values can refer to the related technology and will not be elaborated here.
[0101] In the singular value matrix Σ, starting from the upper left corner of the diagonal, the first singular value corresponds to the 1st rank of the matrix A, the second singular value corresponds to the 2nd rank of the matrix A, and so on. In step S202, the current information amount q can be calculated according to the following formula (2).
[0102]
[0103] Where p i represents the singular value of the matrix A corresponding to the i-th rank. For example, p1 represents the first singular value starting from the upper left corner of the diagonal in the singular value matrix Σ, p2 represents the second singular value on the diagonal, and n is the total number of singular values contained in the singular value matrix Σ.
[0104] Taking k equal to 100 as an example, formula (2) is equivalent to calculating the sum of the squares of the first 100 singular values corresponding to the first 100 ranks of matrix A, calculating the sum of the squares of all singular values of matrix A, and taking the percentage obtained by dividing the former by the latter as the current information amount q.
[0105] In step S203, the difference value between the current information amount and the information amount threshold can be defined as the absolute value of the difference obtained by subtracting the information amount threshold from the current information amount, that is, the difference value = |q - v|, where v is the aforementioned information amount threshold.
[0106] After obtaining the difference value, if the difference value is not less than the preset difference amount m, it means that the information amount retained by the current k ranks is less, and more effective information will be lost when compressing according to the current k ranks. Therefore, the value of k can be appropriately increased to retain more information after compression. The way to increase the value of k can be to increment the value of k by the increment i, that is, k = k + i. The k on the right side of the equation is the k value before the increase, and the k on the left side of the equation represents the k value after the increase. That is to say, the k value after each execution of S205 is larger than the previous k value by the increment i. The value of the preset increment i is not limited. As an example, i can be equal to 2 or equal to 1.
[0107] After updating the k value, the corresponding current information amount can be calculated again based on the updated k value, and it can be judged whether the corresponding difference value is less than the difference amount m, that is, repeat steps S202 to S204 until it is judged that the difference value is less than the difference amount m.
[0108] If the difference value is less than the preset difference amount m, it means that the information amount retained by the current k ranks is sufficient, and compression can be performed based on the current k ranks. Therefore, the current k value is determined as the number K of ranks to be retained. Exemplarily, assuming that when k is equal to 150, it is judged that the difference value is less than the difference amount m, then the number K of ranks to be retained can be determined to be 150.
[0109] Dimensionality reduction processing is performed on the input key-value data according to the first K ranks of the input key-value data. The way to obtain the key-value data to be processed can be:
[0110] Obtain the first K singular values from the singular value matrix Σ corresponding to matrix A, and use these first K singular values as the diagonal of the matrix to form a sub-singular value matrix Σ K Obtain the K left singular vectors corresponding to the first K singular values from the left eigenmatrix U, and form a sub-left eigenmatrix U with the K left singular vectors K Obtain the K right singular vectors corresponding to the first K singular values from the right eigenmatrix V, and form a sub-right eigenmatrix V with the K right singular vectors K Calculate the above sub-singular value matrix, sub-left eigenmatrix and sub-right eigenmatrix according to formula (3), and the obtained result A KIt is the compressed matrix obtained by compressing matrix A based on the first K ranks. The data contained in the compressed matrix is the above-mentioned key-value data to be processed.
[0111]
[0112] Combined with the foregoing example, if a row of matrix A is the key data corresponding to an input text unit, then this row in the compressed matrix is the compressed key data Key to be processed corresponding to this input text unit. If a row of matrix A is the value data corresponding to an input text unit, then this row in the compressed matrix is the compressed value data Value to be processed corresponding to this input text unit.
[0113] Optionally, only when executing S101 for the first time, the number K of ranks to be retained can be determined according to the above method. Starting from the second execution of S101, a new K value does not need to be determined, and the data to be compressed can be directly dimension-reduced based on the first K ranks to obtain the compressed data.
[0114] When the above method of dimension reduction processing is used to compress query data, only A in the above formula (3) needs to be replaced with the query data, which will not be elaborated.
[0115] Optionally, during the process of repeatedly executing Figure 1 the method, the following strategy can be used to decide whether to execute step S102 to determine the target cache capacity.
[0116] One execution strategy can be that after the first execution of S101, the preset initial cache capacity can be directly determined as the target cache capacity. After that, every t decoding steps, the target cache capacity is determined according to the method of step S102.
[0117] One execution strategy can be that after the first execution of S101, the preset initial cache capacity can be directly determined as the target cache capacity. After that, if the target key-value data is screened according to the method of S103 in the previous decoding step, then in the current decoding step, the target cache capacity is determined according to the method of step S102.
[0118] In the above strategies, the value of t can be set as needed. If higher accuracy of the output data is required, a smaller t value can be set, such as setting t equal to 5. If it is necessary to shorten the processing time and resources used as much as possible, a larger t value can be set, such as setting t equal to 10.
[0119] In this embodiment, a decoding step can be understood as the process through which an output data is obtained. Each time an output data is generated, it is equivalent to the end of a decoding step and the start of the next decoding step. For example, after the first execution of S101, the processing is carried out according to the method of S102 to S104 to obtain the first output data. At this time, the first decoding step ends. The process from obtaining the input text unit to generating the first output data is equivalent to the first decoding step. Then, continue to process according to the method of the foregoing embodiment to obtain the second output data. At this time, the second decoding step ends. The process from after generating the first output data to generating the second output data is equivalent to the second decoding step, and so on.
[0120] Optionally, determining the target cache capacity according to the key-value data to be processed and the preset initial cache capacity includes:
[0121] Obtain the target key-value data obtained by the last two screenings;
[0122] Adjust the initial cache capacity according to the key-value data to be processed and the target key-value data obtained by the last two screenings to obtain the target cache capacity.
[0123] The target key-value data obtained by the last two screenings refers to the target key-value data obtained by screening after the last two executions of step S103 as of the current moment. It may include the target key-value data obtained by screening after the last execution of S103, that is, the target key-value data stored in the target cache area at the current moment, denoted as M2, and the target key-value data obtained by screening the previous execution of S103 included in M2, denoted as M1.
[0124] Exemplarily, assuming that step S103 has been executed 5 times as of the current moment, then M1 may be the target key-value data obtained by the second-to-last screening, that is, the target key-value data obtained when S103 was executed for the 4th time, and M2 may be the target key-value data obtained by the last screening, that is, the target key-value data obtained when S103 was executed for the 5th time.
[0125] Optionally, if the cumulative execution times of step S103 are less than 2 when S103 is executed, the initial cache capacity can be directly determined as the target cache capacity.
[0126] When executing S102, the initial cache capacity can be adjusted according to the data volume of the key-value data to be processed and the difference degree of the target key-value data obtained by the last two screenings.
[0127] After obtaining the text to be processed and before stopping the output, all the key-value data to be processed obtained during this period can be stored in the pre-allocated compression buffer area. Therefore, the total amount of data currently cached in the compression buffer area can be detected, and the detection result is the data volume of the key-value data to be processed, denoted as S kv .
[0128] The degree of difference between the target key-value data obtained from the last two screenings, that is, the degree of difference between M1 and M2, can be denoted as f1(M1, M2), where f1 represents the function for calculating the degree of difference. The degree of difference between the two can be determined by various methods, which are not limited. For example, the cosine similarity between M1 and M2 can be calculated, and then 1 minus the cosine similarity can be used as the degree of difference between the two, or the difference between M1 and M2 can be calculated, and the absolute value of the difference can be used as the degree of difference between the two.
[0129] Based on the degree of difference and the data volume S kv , the adjustment index ΔS of the target cache capacity can be calculated according to formula (4).
[0130]
[0131] f2(S kv ) represents the influence factor of the data volume on the adjustment index, and f2 is the function for calculating this influence factor. In this embodiment, the influence factor can be negatively correlated with the data volume, that is, the larger the data volume, the smaller the influence factor, and the smaller the data volume, the larger the influence factor. The calculation method of the influence factor is not limited. As an example, the data volume S kv can be divided by the total cache space that the electronic device can allocate for the data processing model, and the opposite of the obtained result can be used as the influence factor.
[0132] Each time the initial cache capacity is adjusted, the adjustment can be continued on the basis of the previous adjustment, that is, on the basis of the current target cache capacity of the target cache area, rather than directly adjusting the initial cache capacity.
[0133] Among them, if the adjustment index is greater than the preset first threshold T1, that is, ΔS > T1, the target cache capacity can be increased on the basis of the target cache capacity W0 after the previous adjustment according to formula (5).
[0134]
[0135] If the adjustment index is less than or equal to the preset second threshold T2, that is, ΔS ≤ T2, the target cache capacity can be decreased on the basis of the target cache capacity W0 after the previous adjustment according to formula (6).
[0136]
[0137] If the adjustment index is less than or equal to T1 and greater than T2, the target cache capacity can be kept unchanged, that is, the target cache capacity W0 after the previous adjustment can be used as the target cache capacity W1 after this adjustment, that is, W1 = W0.
[0138] In Formulas (5) and (6), Δw is the preset adjustment range of the cache capacity, and W max is the preset upper limit of the target cache capacity, and W min is the preset lower limit of the target cache capacity. W0 represents the target cache capacity after the previous adjustment, W1 represents the target cache capacity after this adjustment, T1 and T2 are two preset different thresholds, and T1 is greater than T2. The initial cache capacity, the upper limit of the target cache capacity, the lower limit of the target cache capacity, and the adjustment range of the cache capacity can all be integer multiples of the data volume of a piece of key-value data to be processed.
[0139] Through the above method, in this embodiment, when the difference between the target key-value data obtained by the last two screenings is large, the target cache capacity can be appropriately increased to store more target key-value data. When the difference between the target key-value data obtained by the last two screenings is small, the target cache capacity can be appropriately decreased to save cache space. At the same time, when the data volume of the key-value data to be processed is large and a large amount of cache space has been occupied, the target cache capacity can be appropriately decreased to further save cache space.
[0140] Optionally, during the process of repeatedly executing Figure 1 the method, the following strategy can be used to determine whether to execute step S103 to re-screen the target cache capacity.
[0141] After the target cache capacity is first determined, step S103 is executed once to determine the target key-value data. Thereafter, every t decoding steps, the target key-value data is screened according to the method of step S103; and if the target cache capacity increases continuously for t0 times, step S103 is immediately executed once.
[0142] t0 is a preset parameter, and its value can be set as needed, such as set to 4, 6 or other values, without limitation.
[0143] Optionally, based on the target cache capacity and the attention weights of the key-value data to be processed, screening the target key-value data from multiple pieces of key-value data to be processed includes:
[0144] Screening the key-value data to be processed according to the attention weights of the key-value data to be processed to obtain a screening result;
[0145] Determining the target key-value data from the key-value data to be processed according to the screening result and the target cache capacity, where the number of target key-value data matches the target cache capacity.
[0146] For any key-value data to be processed, its attention weight S i can be calculated according to the following formula (7).
[0147]
[0148] Among them, softmax represents the normalized exponential function, b i represents the i-th key-value data to be processed, b k represents each key-value data to be processed except b i and T represents transpose, d a represents the dimension of the key-value data to be processed, which can be equal to the sum of the dimensions of the key data to be processed and the value data to be processed. Exemplarily, assuming that the key data to be processed is a 100-dimensional vector and the value data to be processed is a 100-dimensional vector, then d a can be equal to 200.
[0149] Formula (7) is equivalent to substituting each key-value data b i to be processed except b k into the normalized exponential function for calculation to obtain the corresponding normalized value, and then summing all the normalized values to obtain b i corresponding attention weight S i .
[0150] After obtaining the attention weight of each key-value data to be processed, the key-value data to be processed with an attention weight greater than a preset threshold can be selected as the screening result, or a preset number of key-value data to be processed can be selected from high to low according to the attention weight as the screening result.
[0151] The preset number can match the target cache capacity, that is, the preset number is equal to the total number of key-value data to be processed that the current target cache capacity can store, or the preset number can be slightly larger than the total number of key-value data to be processed that the current target cache capacity can store.
[0152] When determining the target key-value data according to the screening result and the target cache capacity, the corresponding number of target key-value data can be determined from the screening result according to the total number of key-value data to be processed that the target cache capacity can store. Exemplarily, if the current target cache capacity can store at most 150 key-value data to be processed, then 150 key-value data to be processed are determined as the target key-value data based on the screening result.
[0153] In some alternative embodiments, the method of screening the key-value data to be processed according to the attention weight of the key-value data to be processed can also be:
[0154] According to the attention weight of the key-value data to be processed, screening is performed among the key-value data to be processed with a relative distance greater than or equal to a preset segmentation parameter to obtain a screening result, and the relative distance is the distance between the corresponding key-value data to be processed and the last key-value data to be processed.
[0155] The serial number of the key-value data to be processed is equivalent to the serial number of the corresponding input text unit. That is to say, for the i-th input text unit, the corresponding key-value data to be processed is also the i-th key-value data to be processed, and the last key-value data to be processed, that is, the key-value data with the largest corresponding serial number.
[0156] Combined with the foregoing example, for "Recommend some cost-effective mobile phones", a total of 6 input text units are divided. Among them, "mobile phone" is the 6th text unit, and the key-value data corresponding to "mobile phone" is the 6th key-value data to be processed. After that, during the process of repeatedly executing Figure 1 the method shown, text units such as "Brand A", "of", "Model A1", "of", "mobile phone" are obtained successively. Among them, "Brand A" is the 7th text unit, "of" is the 8th text unit, "Model A1" is the 9th text unit, and the corresponding key-value data to be processed are the 7th key-value data to be processed, the 8th key-value data to be processed, and the 9th key-value data to be processed in sequence.
[0157] For the i-th key-value data to be processed, the serial number of the last key-value data to be processed, n, can be subtracted by the serial number of the i-th key-value data to be processed, and the obtained difference is divided by the serial number of the last key-value data to be processed. The result obtained is used as the relative distance of the i-th key-value data to be processed. That is to say, the relative distance d i of the i-th key-value data to be processed can be calculated using the following formula (8).
[0158]
[0159] The segmentation parameter is a preset real number greater than 0 and less than 1, and its value can be set as needed. For example, it is set to 0.5.
[0160] Based on the segmentation parameter, the current multiple key-value data to be processed can be divided into a first part with a relative distance less than the segmentation parameter, denoted as S cur , and a second part with a relative distance greater than or equal to the segmentation parameter, denoted as S pre .
[0161] After segmentation, the attention weight of each key-value data in the second part relative to the first part can be calculated using the foregoing formula (7). When calculating the attention weight relative to the first part, b k in formula (7) represents each key-value data to be processed in the first part, and b i represents the i-th key-value data to be processed in the second part.
[0162] Then, based on the attention weights of each key-value data to be processed in the second part, screening can be performed on the key-value data to be processed in the second part to obtain a screening result. The screening method can refer to the foregoing embodiments and will not be elaborated here.
[0163] The advantage of obtaining the screening result in the above manner is that the screening range can be narrowed, from screening all key-value data to be processed to only screening the key-value data to be processed in the second part, achieving the effect of reducing the calculation amount and improving the efficiency.
[0164] In some alternative embodiments, the foregoing screening result may include multiple screening results obtained through multiple screenings.
[0165] In this embodiment, multiple different segmentation parameters can be preset. After obtaining a screening result based on one segmentation parameter according to the foregoing screening method, the screening result is retained, and then another segmentation parameter is selected, and a new screening result is continuously obtained based on the other segmentation parameter according to the foregoing screening method. This process is repeated until corresponding screening results are obtained for each preset segmentation parameter.
[0166] Exemplarily, 5 segmentation parameters can be set, which are 0.3, 0.4, 0.5, 0.6, and 0.7 in sequence. During screening, the first part and the second part can be divided according to 0.3, and then screening is performed to obtain the screening result corresponding to 0.3; the first part and the second part are divided according to 0.4, and then screening is performed to obtain the screening result corresponding to 0.4; and so on. Finally, 5 screening results corresponding to 5 segmentation parameters are obtained.
[0167] On the basis of obtaining multiple screening results, the method for determining the target key-value data in the key-value data to be processed according to the screening results and the target cache capacity can be:
[0168] Determine the target key-value data according to the target cache capacity and the number of times the key-value data to be processed appears repeatedly in multiple screening results.
[0169] In the above screening method, multiple screening results can be combined to determine the number of times the key-value data to be processed contained in these screening results appears repeatedly, and then n key-value data to be processed are selected as the target key-value data in descending order of the number of repeated occurrences, where n is the total amount of key-value data to be processed that the target cache capacity can cache. Exemplarily, if the current target cache capacity can cache 20 key-value data to be processed, then 20 key-value data to be processed are selected as the target key-value data based on the number of repeated occurrences.
[0170] For any key-value data to be processed, every time a screening result contains the key-value data to be processed, it is equivalent to the key-value data to be processed appearing once. The number of repeated occurrences of the key-value data to be processed is equivalent to the number of screening results that contain the key-value data to be processed among multiple screening results.
[0171] Combined with the foregoing example, assuming that 4 out of 5 screening results all contain the key-value data corresponding to "recommended", then the number of repeated occurrences of this key-value data to be processed is equal to 4.
[0172] The advantage of screening the target key-value data in the above manner is that
[0173] Determining the target key-value data by combining the screening results of multiple screenings is conducive to selecting the key-value data that is more important for the output data, thereby reducing the key-value data to be processed while improving the accuracy of the output data.
[0174] Optionally, the way to obtain the output data in step S104 can be:
[0175] According to the number of repeated occurrences corresponding to the target key-value data, fuse a part of the target key-value data and the key-value data to be processed to obtain the output data.
[0176] Among them, for any target key-value data, for example, for the i-th target key-value data, the weight weight corresponding to this target key-value data can be calculated according to formula (9) based on the number of repeated occurrences of this target key-value data i .
[0177]
[0178] In formula (9), overlap i represents the number of repeated occurrences corresponding to the i-th target key-value data, and max(overlap) represents the maximum value of the number of repeated occurrences among these currently screened target key-value data.
[0179] After obtaining the weight, the target key-value data and a part of the key-value data to be processed can be fused according to the following formula (10) to obtain the output data output i .
[0180]
[0181] In formula (10), q i represents the query data Query to be processed corresponding to the last text unit among multiple input text units, T represents transpose, k c represents the key data Key in the c-th target key-value data and key-value data to be processed for fusion, d kRepresents the dimension of the key data Key, v c Represents the value data Value in the c-th target key-value data and the key-value data to be processed for fusion.
[0182] weight c Represents the weight corresponding to the c-th target key-value data and the key-value data to be processed for fusion. If the c-th key-value data is the above target key-value data, then weight c Is the weight calculated based on formula (9) and the number of occurrences of this target key-value data. If the c-th key-value data is not the target key-value data but belongs to a part of the key-value data to be processed for fusion in S104, then weight c Can be equal to a preset weight coefficient, for example, equal to 1.
[0183] For each target key-value data, the greater the number of occurrences, the higher the importance of the output data corresponding to this target key-value data. Therefore, fusing according to the number of occurrences of the target key-value data is beneficial to improving the accuracy of the obtained output data.
[0184] Optionally, the way to fuse the target key-value data and a part of the key-value data to be processed to obtain the output data can be:
[0185] Fuse the target key-value data and the key-value data to be processed with a relative distance less than a preset threshold to obtain the output data. The relative distance is the distance between the corresponding key-value data to be processed and the last key-value data to be processed.
[0186] That is to say, a part of the key-value data to be processed for fusion in step S104 can be the part of the current key-value data to be processed with a relative distance less than the preset threshold.
[0187] The preset threshold is a real number greater than 0 and less than 1. The specific value can be set as needed and is not limited. In some embodiments, the preset threshold can be set to the minimum value of the foregoing multiple segmentation parameters. Exemplarily, when the segmentation parameters include 0.3, 0.4, 0.5, 0.6, and 0.7, the preset threshold can be 0.3.
[0188] In some embodiments, a buffer can be pre-allocated to cache the key-value data to be processed with a relative distance less than the preset threshold, denoted as the buffer closest to the output by distance. After obtaining multiple key-value data to be processed through S101 for the first time, the key-value data to be processed with a relative distance less than the preset threshold is screened out according to the preset threshold and stored in this buffer closest to the output by distance. Thereafter, when repeatedly executing Figure 1During the process of the method shown, each time a new key-value data to be processed is obtained, the new key-value data to be processed can be written into the buffer closest to the output, and the key-value data to be processed with the largest relative distance in the buffer closest to the output can be removed, so as to ensure that the cached data in this buffer always satisfies the condition that the relative distance is less than the preset threshold.
[0189] On this basis, each time S104 is executed, the key-value data to be processed cached in the buffer closest to the output can be directly read, and these key-value data to be processed and the target key-value data are fused to obtain the output data.
[0190] This embodiment provides a data processing device. Please refer to Figure 3 , and this device may include a processor 301 and a memory module 302;
[0191] The memory module 301 is used to cache key-value data;
[0192] The processor 302 is used to execute a preset computer program to perform the following data processing method:
[0193] Compress the input key-value data corresponding to the input text unit to obtain the key-value data to be processed;
[0194] Determine the target cache capacity according to the key-value data to be processed and the preset initial cache capacity. The target cache capacity is the capacity of the target buffer for caching the target key-value data;
[0195] Based on the target cache capacity and the attention weight of the key-value data to be processed, screen out the target key-value data from multiple key-value data to be processed. The attention weight is calculated from the key-value data to be processed;
[0196] Fuse the target key-value data and a part of the key-value data to be processed to obtain the output data.
[0197] The type of the data processing device in this embodiment is not limited. It can be a server device or a terminal device; it can be the host of a split device (such as a split computer) or an integrated all-in-one device; it can be a general computer device or an intelligent device designed for a specific usage scenario, for example, a meeting machine designed specifically for the meeting scenario.
[0198] The processor 302 can include any type of processor. For example, it can include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU).
[0199] Optionally, when the processor 302 determines the target cache capacity according to the key-value data to be processed and the preset initial cache capacity, it is used for:
[0200] Obtain the target key-value data obtained from the most recent two screenings;
[0201] Adjust the initial cache capacity according to the key-value data to be processed and the target key-value data obtained from the most recent two screenings to obtain the target cache capacity.
[0202] Optionally, when the processor 302 adjusts the initial cache capacity according to the key-value data to be processed and the target key-value data obtained from the most recent two screenings, it is used for:
[0203] Adjust the initial cache capacity according to the data volume of the key-value data to be processed and the degree of difference of the target key-value data obtained from the most recent two screenings.
[0204] Optionally, when the processor 302 filters out the target key-value data from multiple key-value data to be processed based on the target cache capacity and the attention weights of the key-value data to be processed, it is used for:
[0205] Filter the key-value data to be processed according to the attention weights of the key-value data to be processed to obtain the filtering result;
[0206] Determine the target key-value data from the key-value data to be processed according to the filtering result and the target cache capacity, where the number of target key-value data matches the target cache capacity.
[0207] Optionally, when the processor 302 filters the key-value data to be processed according to the attention weights of the key-value data to be processed to obtain the filtering result, it is used for:
[0208] Filter in the key-value data to be processed where the relative distance is greater than or equal to the preset segmentation parameter according to the attention weights of the key-value data to be processed to obtain the filtering result, and the relative distance is the distance between the corresponding key-value data to be processed and the last key-value data to be processed.
[0209] Optionally, the filtering result includes multiple filtering results obtained from multiple screenings;
[0210] When the processor 302 determines the target key-value data from the key-value data to be processed according to the filtering result and the target cache capacity, it is used for:
[0211] Determine the target key-value data according to the target cache capacity and the number of times the key-value data to be processed appears repeatedly in multiple filtering results.
[0212] Optionally, when the processor 302 fuses a part of the target key-value data and the key-value data to be processed to obtain the output data, it is used for:
[0213] Fuse the target key-value data and a part of the key-value data to be processed according to the number of times the target key-value data appears repeatedly to obtain the output data.
[0214] Optionally, when the processor 302 fuses the target key-value data and a part of the key-value data to be processed to obtain output data, it is used for:
[0215] Fuse the target key-value data and the key-value data to be processed with a relative distance less than a preset threshold to obtain output data, where the relative distance is the distance between the corresponding key-value data to be processed and the last key-value data to be processed.
[0216] Optionally, when the processor 302 compresses the input key-value data corresponding to the input text unit to obtain the key-value data to be processed, it is used for:
[0217] Determine the number K of ranks to be retained according to the input key-value data corresponding to the input text unit;
[0218] Perform dimensionality reduction processing on the input key-value data according to the first K ranks of the input key-value data to obtain the key-value data to be processed.
[0219] For the data processing device provided in this embodiment, the working principle can refer to the relevant steps in the data processing method of the foregoing embodiment, which will not be elaborated.
[0220] It should be noted that the various embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0221] For the convenience of description, when describing the above system or device, various modules or units are described separately according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0222] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0223] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising said element.
[0224] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A data processing method, comprising: Compressing the input key - value data corresponding to the input text unit to obtain the key - value data to be processed; Determining a target cache capacity according to the key - value data to be processed and a preset initial cache capacity, where the target cache capacity is the capacity of a target cache area for caching target key - value data; Based on the target cache capacity and the attention weights of the key - value data to be processed, screening out target key - value data from multiple key - value data to be processed, where the attention weights are calculated from the key - value data to be processed; Fusing the target key - value data and a part of the key - value data to be processed to obtain output data.
2. The method according to claim 1, wherein the determining the target cache capacity according to the key - value data to be processed and a preset initial cache capacity comprises: Obtaining the target key - value data obtained from the last two screenings; Adjusting the initial cache capacity according to the key - value data to be processed and the target key - value data obtained from the last two screenings to obtain the target cache capacity.
3. The method according to claim 2, wherein the adjusting the initial cache capacity according to the key - value data to be processed and the target key - value data obtained from the last two screenings comprises: Adjusting the initial cache capacity according to the data volume of the key - value data to be processed and the degree of difference between the target key - value data obtained from the last two screenings.
4. The method according to claim 1, wherein the screening out target key - value data from multiple key - value data to be processed based on the target cache capacity and the attention weights of the key - value data to be processed comprises: Screening the key - value data to be processed according to the attention weights of the key - value data to be processed to obtain a screening result; Determining target key - value data from the key - value data to be processed according to the screening result and the target cache capacity, wherein the number of target key - value data matches the target cache capacity.
5. The method according to claim 4, wherein the screening the key - value data to be processed according to the attention weights of the key - value data to be processed to obtain a screening result comprises: Screening among the key - value data to be processed with a relative distance greater than or equal to a preset segmentation parameter according to the attention weights of the key - value data to be processed to obtain a screening result, where the relative distance is the distance between the corresponding key - value data to be processed and the last key - value data to be processed.
6. The screening result according to claim 4 includes multiple screening results obtained from multiple screenings; The determining target key - value data from the key - value data to be processed according to the screening result and the target cache capacity comprises: Determining target key - value data according to the target cache capacity and the number of times the key - value data to be processed appears repeatedly in the multiple screening results.
7. The method according to claim 6, wherein the fusing the target key - value data and a part of the key - value data to be processed to obtain output data comprises: Fusing the target key - value data and a part of the key - value data to be processed according to the number of times the target key - value data appears repeatedly to obtain output data.
8. The method according to claim 1, wherein the fusing the target key-value data and a part of the to-be-processed key-value data to obtain output data comprises: fusing the target key-value data and the to-be-processed key-value data with a relative distance less than a preset threshold to obtain output data, where the relative distance is the distance between the corresponding to-be-processed key-value data and the last to-be-processed key-value data.
9. The method according to claim 1, wherein the compressing the input key-value data corresponding to the input text unit to obtain the to-be-processed key-value data comprises: determining the number K of ranks to be retained according to the input key-value data corresponding to the input text unit; performing dimensionality reduction processing on the input key-value data according to the first K ranks of the input key-value data to obtain the to-be-processed key-value data.
10. A data processing device, comprising a processor and a memory module; the memory module is used for caching key-value data; the processor is used for executing a preset computer program to perform the following data processing method: compressing the input key-value data corresponding to the input text unit to obtain the to-be-processed key-value data; determining a target cache capacity according to the to-be-processed key-value data and a preset initial cache capacity, where the target cache capacity is the capacity of a target cache area for caching target key-value data; screening out target key-value data from multiple pieces of the to-be-processed key-value data based on the target cache capacity and the attention weight of the to-be-processed key-value data, where the attention weight is calculated from the to-be-processed key-value data; fusing the target key-value data and a part of the to-be-processed key-value data to obtain output data.