Data processing method and electronic equipment
By selecting the first X characters in the target document group to recalculate the kv data, the problem of inaccurate generation results caused by independent calculation of vector data in the cache pool is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202510645935.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-15
AI Technical Summary
In a large language model that generates RAG based on retrieval enhancement, the vector data of each document block in the cache pool is calculated independently, resulting in the contextual relationship between document blocks that cannot be characterized, resulting in inaccurate generation results.
By selecting the first X consecutive characters in the target document group, recalculate the kv data corresponding to these characters, and update them using the intelligent model to ensure the accuracy of the association relationship between the document groups and reduce the calculation amount.
It improves the accuracy of the generated results, while reducing the calculation amount and improving the efficiency of obtaining accurate kv data.
Smart Images

Figure CN120493866A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data processing method and electronic equipment. Background Art
[0002] In the process of using the large language model based on Retrieval-augmented Generation (RAG), RAG retrieves the relevant document blocks based on the user's question, and the large language model calculates vector data in the form of key-value pairs for these document blocks and user questions and generates content.
[0003] In order to improve the efficiency of the large language model, a cache pool is set up for the document block. In this way, the large language model does not need to calculate the vector data of the document block in real time. The vector data can be directly read from the cache pool and provided to the large language model for direct content generation.
[0004] However, the vector data of each document block in the cache pool is calculated independently for the document block, so that the vector data provided to the large language model cannot represent the contextual relationship between the document blocks, resulting in inaccurate results generated by the large language model. Summary of the Invention
[0005] In view of this, the present application provides a data processing method and electronic device as follows:
[0006] A data processing method, comprising:
[0007] Obtaining kv data corresponding to each of a plurality of target document groups; the kv data is kv data obtained by performing independent calculations on the target document groups; the target document groups are sequentially connected;
[0008] In at least one of the target document groups, the first X consecutive characters are selected as target characters; X is a positive integer greater than or equal to 1; X is less than the number of characters included in the corresponding target document group;
[0009] Update the kv data corresponding to each target character in turn.
[0010] The above method preferably updates the kv data corresponding to the target character, including:
[0011] Determine a prefix document group corresponding to the target character; the prefix document group includes all target document groups sorted before the target document group where the target character is located;
[0012] Based on the kv data corresponding to the prefix document group, the kv data corresponding to the target character is recalculated.
[0013] In the above method, preferably, the target document group where the target character is located includes: any target data group among the multiple target document groups except the first target data group sorted;
[0014] Wherein, in the case that the target character is included in the prefix document group, the kv data corresponding to the target character in the prefix document group has been recalculated.
[0015] In the above method, preferably, different target document groups contain different numbers of target characters;
[0016] The number of the target characters in the target document group that is sorted first is greater than the number of the target characters in the target document group that is sorted last.
[0017] In the above method, preferably, the number of target characters contained in the target document groups that are sorted adjacently is linearly varied based on the maximum value and the minimum value;
[0018] The maximum value is an upper limit value of the number of target characters included in one target document group; and the minimum value is a lower limit value of the number of target characters included in one target document group.
[0019] The above method preferably obtains the kv data corresponding to each of the multiple target document groups, including:
[0020] According to the input data, a plurality of neighboring document blocks associated with the input data are retrieved from a database;
[0021] Obtaining the target document group according to the neighboring document blocks, wherein the target document group consists of at least one neighboring document block;
[0022] Read the kv data corresponding to each target document group from the cache pool;
[0023] The maximum number of neighboring document blocks included in the target document group is determined based on the size of the kv data corresponding to the characters and the maximum storage space of the buffer pool.
[0024] In the above method, preferably, when the kv data corresponding to the target document group is not read in the cache pool, the method further comprises:
[0025] Utilize the intelligent model to calculate the kv data corresponding to the target document group.
[0026] Preferably, the above method further comprises, after calculating the kv data corresponding to the target document group using the intelligent model:
[0027] The calculated kv data corresponding to the target document group is stored in the cache pool.
[0028] In the above method, preferably, the kv data read from the cache pool is loaded into the target storage area;
[0029] Wherein, updating the kv data corresponding to each target character in sequence includes:
[0030] For the target characters loaded into the target storage area, updating the kv data corresponding to the target characters in the target storage area in sequence;
[0031] The updating of the kv data corresponding to the target character in the target storage area and the loading of the kv data corresponding to other characters sorted after the updated target character into the target storage area are performed in parallel.
[0032] A data processing device, comprising:
[0033] A data acquisition unit is used to obtain kv data corresponding to each of a plurality of target document groups; the kv data is kv data obtained by performing independent calculations on the target document groups; the target document groups are sequentially connected;
[0034] a character selection unit, configured to select, from at least one of the target document groups, the first X consecutive characters as target characters; X being a positive integer greater than or equal to 1; and X being less than the number of characters contained in the corresponding target document group;
[0035] The data updating unit is used to update the kv data corresponding to each target character in sequence.
[0036] An electronic device, comprising:
[0037] A memory for storing computer programs and data generated by the execution of the computer programs;
[0038] A processor is configured to execute the computer program to achieve the following: obtaining kv data corresponding to each of a plurality of target document groups; the kv data being kv data obtained by independently calculating the target document groups; the target document groups being sequentially connected; in at least one of the target document groups, selecting the first X consecutive characters as target characters; X being a positive integer greater than or equal to 1; X being less than the number of characters contained in the corresponding target document group; and sequentially updating the kv data corresponding to each of the target characters.
[0039] It can be seen from the above technical solution that in a data processing method and electronic device disclosed in the present application, for the kv data obtained by independent calculation of a document group, some characters sorted in the front can be selected from it to recalculate the kv data, and there is no need to recalculate the kv data for other characters in the document group. In this way, while ensuring the accuracy of the kv data, the computational complexity of the kv calculation can be reduced, thereby improving the efficiency of obtaining accurate kv data. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 A flowchart of a data processing method provided in an embodiment of the present application;
[0042] Figure 2 This is an example diagram of the end-to-end connection between target document groups in an embodiment of the present application;
[0043] Figure 3 This is an example diagram of the kv data corresponding to the target document group in the embodiment of the present application;
[0044] Figure 4 A diagram showing the Euclidean distance between the way kv data is calculated for a combination of document blocks and the way kv data is calculated independently and then combined;
[0045] Figure 5 This is an example diagram of target characters in the target document group in the embodiment of the present application;
[0046] Figure 6 Another example diagram of target characters in the target document group in the embodiment of the present application;
[0047] Figure 7 A partial flow chart of a data processing method provided in an embodiment of the present application;
[0048] Figure 8 This is an example diagram of neighboring document blocks forming a target document group in an embodiment of the present application;
[0049] Figure 9 Another partial flow chart of a data processing method provided in an embodiment of the present application;
[0050] Figure 10 This is an example diagram of serial loading and recalculation in an embodiment of the present application;
[0051] Figure 11Another example diagram of serial loading and recalculation in an embodiment of the present application;
[0052] Figure 12 This is an example diagram of parallel loading and recalculation in an embodiment of the present application;
[0053] Figure 13 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0054] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0055] Figure 15 This is an example diagram of implementing content generation in a usage scenario where the application is applicable to a large language model based on RAG;
[0056] Figure 16 This is a schematic diagram of recalculating the key-value data of some tokens in the usage scenario of the RAG-based large language model applicable to this application;
[0057] Figure 17 This is a schematic diagram of iteratively calculating the kv data of some tokens in each document in the usage scenario of the RAG-based large language model applicable to this application;
[0058] Figure 18 This is an example diagram of parallel loading and recalculation in the usage scenario of a large RAG-based language model applicable to this application. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0060] refer to Figure 1 The figure shows a flowchart of a data processing method according to an embodiment of the present application. The method can be applied to electronic devices capable of data processing, such as mobile phones, tablet devices, notebooks, or servers. The technical solution in this embodiment is mainly used to reduce the amount of calculation required to obtain KV data, thereby improving calculation efficiency.
[0061] Specifically, the method in this embodiment may include the following steps:
[0062] Step 101: Obtain kv data corresponding to each of a plurality of target document groups.
[0063] Among them, the kv data is the kv data obtained by independent calculation of the target document group. Kv data is vector data stored in the form of key-value pairs. The target document groups are connected in sequence. For example, Figure 2 As shown in the figure, the four target document groups are document 1, document 2, document 3, and document 4. Each target document group is connected by the characters they contain. Except for the target document group ranked first, all other target document groups have prefix document groups. For example, in the case of document group 3, document groups 1 and 2 are prefix document groups.
[0064] In one implementation, in step 101 , the kv data corresponding to each target document group may be read from the cache pool.
[0065] The cache pool includes the kv data obtained by independently calculating the documents through the intelligent model. Based on this, in step 101, the kv data corresponding to the target document group is searched in the cache pool according to the target document group, and the kv data corresponding to the target document group is read out after the search is completed.
[0066] In another implementation, in step 101 , the kv data of each target document group may be calculated using an intelligent model to obtain the kv data corresponding to each target document group.
[0067] The intelligent model is a machine learning model that can calculate the kv data of characters. A document group can be converted into multiple characters, which can also be called tokens, such as words, single characters, and phrases. The kv data corresponding to the document group consists of the kv data corresponding to each character token converted from the document group, such as Figure 3 As shown in .
[0068] Step 102: In at least one target document group, select the top X consecutive characters as target characters.
[0069] Wherein, X is a positive integer greater than or equal to 1; X is less than the number of characters contained in the corresponding target document group. That is, in this embodiment, some characters in the target document group are selected as target characters, and these target characters are ranked in the first X in the target document group.
[0070] It should be noted that there may be differences between the KV data calculated after combining multiple target document groups and the KV data obtained by performing KV calculations on multiple target document groups independently and then combining them.
[0071] For example, Figure 4As shown in , taking 6 document blocks as an example, full-precision prefix kv data and independent kv data are generated based on the same 6 document blocks. The prefix kv data is the kv data obtained by combining the 6 document blocks end to end and then performing kv calculation on the whole. The independent kv data is obtained by performing independent kv calculation on the 6 document blocks and then combining the obtained kv data. Based on this, the Euclidean distance between the two types of kv data is calculated, including: the distance between the keys in the kv data (such as Figure 4 The part on the top of the middle line), and the distance between the values in the kv data (such as Figure 4 The horizontal axis represents the index of the character token in the document block, and the vertical axis represents the Euclidean distance between the key and value under the two calculation methods. The differences between the method of calculating the key-value data by combining document blocks and the method of calculating the key-value data independently and then combining them can be seen as follows:
[0072] Within each document block, a small number of character tokens at the beginning will have their kv data significantly different between the two calculation methods, while the differences between the majority of character tokens in the middle and end of the document block will be relatively small. Figure 4 As shown in the figure, the Euclidean distance of the key data of the character token at the starting position in document block 2 under the two calculation methods is significantly greater than the key data of the character token at the back middle and end positions in document block 2; the Euclidean distance of the value data of the character token at the starting position in document block 2 under the two calculation methods is significantly greater than the value data of the character token at the back middle and end positions in document block 2.
[0073] Based on this, in this embodiment, in the target document group, the characters ranked in the top X are selected as target characters, such as Figure 5 As shown in , the kv data corresponding to these target characters are significantly different from the kv data obtained by combining all target document groups and performing kv calculations.
[0074] Step 103: Update the kv data corresponding to each target character in sequence.
[0075] Specifically, in this embodiment, the intelligent model can be used to perform kv calculations on these target characters in sequence to obtain kv data corresponding to these target characters.
[0076] Furthermore, in this embodiment, after step 103, all target document groups (wherein at least one target document group has X target characters updated) can be provided to an intelligent model such as a large language model (LLM), and the intelligent model can perform content generation, such as text generation or image generation, based on at least these target document groups to obtain target content.
[0077] It can be seen from the above technical solution that in a data processing method provided by an embodiment of the present application, for the kv data obtained by independent calculation of a document group, some characters sorted in the front can be selected from it to recalculate the kv data, and there is no need to recalculate the kv data for other characters in the document group. This can ensure the accuracy of the kv data while reducing the computational complexity of the kv calculation, thereby improving the efficiency of obtaining accurate kv data.
[0078] In one implementation, when updating the kv data corresponding to each target character in step 103, the prefix document group corresponding to the target character can be determined first. The prefix document group corresponding to the target character includes: all target document groups sorted before the target document group where the target character is located. Then, based on the kv data corresponding to the prefix document group, the kv data corresponding to the target character is recalculated.
[0079] For example, Figure 6 As shown in , taking four document groups as an example, Document 1, Document 2, Document 3, and Document 4 each consist of multiple tokens. For Document 2, Document 3, and Document 4, the target characters l2, l3, and l4 within them correspond to prefix document groups, respectively. The target character l2 in Document 2 corresponds to the prefix document group of Document 1. The target character l3 in Document 3 corresponds to the prefix document groups of Document 1 and Document 2. The target character l4 in Document 4 corresponds to the prefix document groups of Document 1, Document 2, and Document 3. Based on this, in this embodiment, for the target character l2 in document 2, the prefix document group, namely document 1, is determined, and according to the kv data corresponding to document 1, the kv calculation is performed on the target character l2 in document 2 to obtain the kv data corresponding to the target character l2; for the target character l3 in document 3, the prefix document group, namely document 1 and document 2, is determined, and according to the kv data corresponding to document 1 and the kv data corresponding to document 2, the kv calculation is performed on the target character l3 in document 3 to obtain the kv data corresponding to the target character l3; for the target character l4 in document 4, the prefix document group, namely document 1, document 2 and document 3, is determined, and according to the kv data corresponding to document 1, the kv data corresponding to document 2 and the kv data corresponding to document 3, the kv calculation is performed on the target character l4 in document 4 to obtain the kv data corresponding to the target character l4.
[0080] Based on the above implementation, the target document group containing the target character includes any target document group among the multiple target document groups except the first target document group. In other words, the first target document group does not need to have the target character selected, and the target document groups containing the target character do not include the first target document group.
[0081] For example, Figure 6 As shown in , target characters are selected from Document 2, Document 3, and Document 4 respectively. Document 1 does not have a prefix document group and does not need to recalculate kv, so it is not selected as a target character.
[0082] Furthermore, in the case where the prefix document group includes the target character, the kv data corresponding to the target character in the prefix document group has been recalculated.
[0083] It should be noted that in this embodiment, when recalculating the KV data corresponding to the target character, the KV data corresponding to the target character is recalculated sequentially according to the connection order between the target document groups and the front-to-back order between the characters in the target document groups. Based on this, as the KV data corresponding to the target character is calculated, if there are target characters in the prefix document group corresponding to the target character whose KV data is recalculated later, these target characters will be recalculated first.
[0084] For example, Figure 6As shown in , for the target character l2 in document 2, the prefix document group, i.e., document 1, is determined. There is no selected target character in document 1. At this time, kv calculation is performed on each target character l2 in document 2 in turn according to the kv data corresponding to document 1 to obtain the kv data corresponding to the target character l2; for the target character l3 in document 3, the prefix document group, i.e., document 1 and document 2, is determined. Document 2 contains the selected target character l2, but the kv data corresponding to these target characters l2 have been recalculated. At this time, kv calculation is performed on each target character l3 in document 3 in turn according to the kv data corresponding to document 1 and the kv data corresponding to document 2 (the selected target character l2 in document 2 has been recalculated) to obtain the target character l3. The kv data corresponding to the character l3 is obtained; for the target character l4 in document 4, the prefix document group is determined, namely document 1, document 2 and document 3. Document 2 contains the selected target character l2, but the kv data corresponding to these target characters l2 have been recalculated. Document 3 contains the selected target character l3, but the kv data corresponding to these target characters l3 have been recalculated. At this time, according to the kv data corresponding to document 1, the kv data corresponding to document 2 (the selected target character l2 in document 2 has been recalculated) and the kv data corresponding to document 3 (the selected target character l3 in document 3 has been recalculated), the kv calculation is performed on each target character l4 in document 4 in turn to obtain the kv data corresponding to the target character l4.
[0085] Based on the above implementation scheme, Figure 4 As shown in the Euclidean distance, the difference between the way document blocks are combined to calculate KV data and the way they are calculated independently and then combined is as follows:
[0086] The deviation of the kv data calculated independently and then combined increases with the number of previously superimposed document blocks.
[0087] For example, Figure 4 As shown in , the Euclidean distance of the key data of the character token at a certain position in document block 6 under the two calculation methods is greater than the Euclidean distance of the key data of the character token at the same position in document block 4 under the two calculation methods.
[0088] Based on this, in this embodiment, different target document groups may include different numbers of target characters. That is, different numbers of target characters are selected for target document groups at different sorting positions. In a specific implementation, the number of target characters in the target document group ranked earlier is greater than the number of target characters in the target document group ranked later.
[0089] For example, Figure 6As shown in , the number of target characters l2 in document 2 is greater than the number of target characters l3 in document 3, and the number of target characters l3 in document 3 is greater than the number of target characters l4 in document 4.
[0090] In a specific implementation, in the target documents that are adjacent in sorting, the number of target characters included varies linearly based on the maximum value and the minimum value.
[0091] The maximum value is the upper limit of the number of target characters contained in a target document group; the minimum value is the lower limit of the number of target characters contained in a target document group. The maximum and minimum values can be pre-set based on business needs. max Indicates the minimum value with l min In this embodiment, the difference between the maximum value and the minimum value can be used as the slope of the linear change relationship. Specifically, as shown in formula (1):
[0092]
[0093] Wherein, ln is the number of target characters selected in the nth target document group, and N is the total number of target document groups.
[0094] As can be seen, in this embodiment, the difference between the maximum and minimum values is used as the slope to construct an objective function representing a linear change relationship. The objective function uses the ranking position of the target document group among all target document groups as the independent variable and the number of characters selected as target characters as the dependent variable. Therefore, using the objective function, the number X of characters selected as target characters in each target document group can be determined. Furthermore, the characters ranked in the top X in the target document group are selected as target characters, and the KV data for these target characters is recalculated.
[0095] Based on the above implementation scheme, in one implementation, when obtaining the kv data corresponding to each of the multiple target document groups in step 101, it can be implemented in the following way, such as Figure 7 As shown in:
[0096] Step 701: According to input data, a plurality of neighbor document blocks associated with the input data are retrieved from a database.
[0097] The database contains multiple document blocks to be retrieved. Input data is data, such as text or images, that a user enters when using an intelligent model to generate text or images. In this embodiment, based on the input data, the database is searched for document blocks that meet association conditions with the input data, i.e., neighboring document blocks.
[0098] The association condition may be: the Euclidean distance between the input data and the neighboring document block is less than or equal to a corresponding threshold, or the similarity between the input data and the neighboring document block is greater than or equal to a corresponding threshold.
[0099] Step 702: Obtain a target document group based on the neighboring document blocks.
[0100] The target document group consists of at least one neighboring document block. In this embodiment, neighboring document blocks can be combined to obtain multiple target document groups. For example, after retrieving 10 neighboring document blocks, the first and second neighboring document blocks are concatenated into Document 1, the third neighboring document block is used as Document 2, the fourth, fifth, and sixth neighboring document blocks are concatenated into Document 3, and the seventh, eighth, ninth, and tenth neighboring document blocks are concatenated into Document 4. Document 1, Document 2, Document 3, and Document 4 are each a target document group.
[0101] Step 703: Read the kv data corresponding to each target document group from the cache pool.
[0102] The maximum number of neighboring document blocks included in the target document group is determined based on the size of the key-value data corresponding to the character and the maximum storage space of the cache pool. The cache pool includes key-value data obtained through independent calculations of documents using the intelligent model. The documents corresponding to the key-value data included in the cache pool can be document blocks in the database or document groups composed of document blocks.
[0103] It can be seen that in this embodiment, the kv data corresponding to the document group is pre-saved in the cache pool. After obtaining the target document group, the corresponding kv data can be directly searched from the cache pool without using the intelligent model to calculate the kv data of the target document group in real time, thereby reducing the calculation time, thereby speeding up the rate of kv data provided to the intelligent model and improving the efficiency of the intelligent model in content generation.
[0104] It should be noted that the number of document blocks that can be spliced together in a document group corresponding to the KV data included in the cache pool is limited. This is to prevent the KV data corresponding to the resulting document group from being too large due to splicing too many document blocks, thereby occupying too much space in the cache pool. Based on this, in this embodiment, a maximum number of neighboring document blocks is set, and the number of neighboring document blocks included in any target document group does not exceed this maximum number.
[0105] In one implementation, a document group consisting of multiple document blocks is stored in the form of a tree, such as Figure 8As shown in , the document block spliced at the first position is the root node of the tree, the document block spliced at the second position is the child node of the root node, the document block spliced at the third position is the child node of the child node of the root node, and so on. The document block spliced at the last position is the leaf node of the tree. The number of document blocks in a document group, that is, the depth D of the tree, is limited and does not exceed the preset maximum number. The maximum value of D is as shown in formula (2):
[0106]
[0107] Among them, D max is the maximum depth of the tree, α is the preset coefficient, B max is the maximum storage space of the cache pool; B is the maximum size of the kv data corresponding to a single character; M is the number of document blocks in the database.
[0108] It can be seen that in this embodiment, the number of document blocks to be spliced in a document group can be limited to avoid excessive occupation of the buffer pool space, thereby allowing the buffer pool to reserve more kv data of the document group.
[0109] Based on the above implementation scheme, in this embodiment, if the kv data corresponding to the target document group is not read in the cache pool in step 703, the following processing may also be included, such as Figure 9 As shown in:
[0110] Step 704: Calculate the kv data corresponding to the target document group using the intelligent model.
[0111] That is to say, if the kv data corresponding to the target document group is not pre-stored in the cache pool, the intelligent model will be used for real-time calculation.
[0112] Furthermore, in this embodiment, after step 704 , the calculated kv data corresponding to the target document group may be stored in a cache pool for next search.
[0113] It can be seen that in this embodiment, the kv data in the cache pool can be continuously updated, thereby accelerating the efficiency of subsequent kv data acquisition and improving the efficiency of content generation by the intelligent model.
[0114] In a specific implementation, Figure 10As shown in , after this embodiment reads the kv data corresponding to the characters in the target document block from the cache pool, it loads it into the target storage area, and then continues to read the kv data corresponding to the next character from the cache pool and loads it into the target storage area, until the kv data corresponding to the target character is read from the cache pool, the kv data corresponding to the target character is recalculated based on the kv data already loaded into the target storage area, and then the recalculated kv data is loaded into the target storage area, and then the kv data corresponding to other subsequent characters are continued to be read from the cache pool and loaded into the target storage area. In this process, the kv data corresponding to each character (including the target character) is loaded into the target storage area and the kv data is recalculated serially.
[0115] In a specific implementation, Figure 11 As shown in , this embodiment reads the kv data corresponding to each character (including the target character) in each target document block from the cache pool in sequence and loads it into the target storage area. After the loading is completed, the kv data corresponding to the target character is recalculated based on the kv data that has been loaded into the target storage area. In this process, the kv data corresponding to each character (including the target character) is loaded into the target storage area and the kv data is recalculated in series.
[0116] In another implementation, in this embodiment, when reading the KV data corresponding to the target document from the buffer pool, the read KV data can be loaded into a target storage area, such as a video memory area. In other words, in this embodiment, the KV data read from the buffer pool is loaded into the target storage area. Based on this, in step 103, when updating the KV data corresponding to the target character, the KV data corresponding to the target character in the target storage area is updated sequentially for the target character loaded into the target storage area.
[0117] Updating the kv data corresponding to the target character in the target storage area and loading the kv data corresponding to other characters sorted after the updated target character into the target storage area are performed in parallel.
[0118] In specific implementation, such as Figure 12As shown in , after this embodiment reads the kv data corresponding to a character in a target document block from the cache pool, it loads it into the target storage area, and then continues to read the kv data corresponding to the next character from the cache pool and loads it into the target storage area, until the kv data corresponding to the target character is read from the cache pool and loaded into the target storage area, the kv data corresponding to the target character in the target storage area is recalculated based on the kv data already loaded into the target storage area, and at the same time, the kv data corresponding to subsequent other characters are continued to be read from the cache pool in parallel and loaded into the target storage area. In this process, the kv data corresponding to each character (including the target character) is loaded into the target storage area and the kv data loaded into the target storage area is recalculated in parallel.
[0119] It can be seen that in this embodiment, the loading of the kv data corresponding to the current character and the recalculation of the kv data corresponding to the loaded target character are performed in parallel, which can speed up the rate of loading accurate kv data into the target storage area, thereby improving the efficiency of the intelligent model in content generation.
[0120] refer to Figure 13 , is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. This device can be used in electronic devices capable of data processing, such as mobile phones, tablet devices, laptops, or servers. The technical solution in this embodiment is primarily used to reduce the computational complexity of obtaining KV data, thereby improving computational efficiency.
[0121] Specifically, the device in this embodiment may include the following units:
[0122] The data obtaining unit 1301 is used to obtain kv data corresponding to each of a plurality of target document groups; the kv data is the kv data obtained by performing independent calculations on the target document groups; the target document groups are sequentially connected;
[0123] The character selection unit 1302 is configured to select the first X consecutive characters in at least one target document group as target characters; X is a positive integer greater than or equal to 1; and X is less than the number of characters included in the corresponding target document group.
[0124] The data updating unit 1303 is configured to update the kv data corresponding to each target character in sequence.
[0125] It can be seen from the above technical solution that in a data processing device provided by an embodiment of the present application, for the kv data obtained by independent calculation of a document group, some characters sorted in the front can be selected from it to recalculate the kv data, and there is no need to recalculate the kv data for other characters in the document group. In this way, while ensuring the accuracy of the kv data, the computational complexity of the kv calculation can be reduced, thereby improving the efficiency of obtaining accurate kv data.
[0126] In one implementation, when updating the kv data corresponding to the target character, the data update unit 1303 is specifically used to: determine the prefix document group corresponding to the target character; the prefix document group includes all target document groups sorted before the target document group where the target character is located; based on the kv data corresponding to the prefix document group, recalculate the kv data corresponding to the target character.
[0127] Among them, the target document group where the target character is located includes: any target data group among the multiple target document groups except the first target data group sorted; when the target character is included in the prefix document group, the kv data corresponding to the target character in the prefix document group has been recalculated.
[0128] In one implementation, different target document groups contain different numbers of target characters; wherein the number of target characters in the target document group that is sorted first is greater than the number of target characters in the target document group that is sorted later.
[0129] In one implementation, the number of target characters included in the adjacently sorted target document groups varies linearly based on a maximum value and a minimum value; wherein the maximum value is an upper limit value of the number of target characters included in one target document group; and the minimum value is a lower limit value of the number of target characters included in one target document group.
[0130] In one implementation, when obtaining the kv data corresponding to each of a plurality of target document groups, the data acquisition unit 1301 is specifically used to: retrieve a plurality of neighboring document blocks associated with the input data from a database according to the input data; obtain the target document group based on the neighboring document blocks, wherein the target document group is composed of at least one neighboring document block; read the kv data corresponding to each of the target document groups from a cache pool; wherein the maximum number of neighboring document blocks contained in the target document group is determined based on the size of the kv data corresponding to the character and the maximum storage space of the cache pool.
[0131] Wherein, when the kv data corresponding to the target document group is not read in the cache pool, the data acquisition unit 1301 is further configured to calculate the kv data corresponding to the target document group using an intelligent model.
[0132] Furthermore, after calculating the kv data corresponding to the target document group using the intelligent model, the data acquisition unit 1301 is further configured to: store the calculated kv data corresponding to the target document group into the cache pool.
[0133] Specifically, the kv data read from the cache pool is loaded into the target storage area;
[0134] Based on this, the data update unit 1303 is specifically used to: update the kv data corresponding to the target characters in the target storage area in sequence for the target characters loaded into the target storage area; wherein, the updating of the kv data corresponding to the target characters in the target storage area and the loading of the kv data corresponding to other characters sorted after the updated target characters into the target storage area are performed in parallel.
[0135] It should be noted that the specific implementation of each unit in this embodiment can refer to the corresponding content in the previous text and will not be described in detail here.
[0136] refer to Figure 14 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, and the electronic device may include the following structure:
[0137] Memory 1401, used to store computer programs and data generated by the execution of the computer programs;
[0138] Processor 1402 is configured to execute a computer program to: obtain kv data corresponding to each of a plurality of target document groups; the kv data being kv data obtained by independently calculating the target document groups; the target document groups being sequentially connected; in at least one of the target document groups, selecting the first X consecutive characters as target characters; X being a positive integer greater than or equal to 1; and X being less than the number of characters contained in the corresponding target document group; and sequentially updating the kv data corresponding to each of the target characters.
[0139] It can be seen from the above technical solution that in an electronic device provided by an embodiment of the present application, for the kv data obtained by independent calculation of a document group, some characters sorted in the front can be selected from it to recalculate the kv data, and there is no need to recalculate the kv data for other characters in the document group. In this way, while ensuring the accuracy of the kv data, the computational complexity of the kv calculation can be reduced, thereby improving the efficiency of obtaining accurate kv data.
[0140] Taking the use scenario of a large language model based on RAG as an example, the technical solution of this application is illustrated below:
[0141] like Figure 15 As shown in , it is the overall process of the technical solution of this application. For the user's input question (i.e., the input data in the above text), first retrieve the most relevant N document blocks (i.e., the neighboring document blocks in the above text) through the vector database (i.e., the database in the above text), and generate K document groups (i.e., the target document groups in the above text, when each target document group only contains one neighboring document block, K is the same as N) through kv cache grouping. Afterwards, based on these document groups, search for the corresponding kv cache (i.e., kv data) in the kv cache pool. If the kv cache of a certain document group exists, it is read directly. If it does not exist, the kv cache of the document group is generated by the large language model and stored in the cache pool. After the kv cache is read, the kv cache recalculation selection mechanism guided by prior knowledge obtains some tokens that need to perform the recalculation task, and inputs them into the large language model. By re-prefilling these tokens, the kv cache of the restored accuracy (i.e., accuracy) is obtained. Finally, the kv cache of the restored accuracy is input into the large language model together with the user's input question to obtain the final generation result.
[0142] based on Figure 15 The technical solution of this application has the following technical points:
[0143] 1. KV cache recalculation token selection guided by prior knowledge:
[0144] In the technical solution of this application, by summarizing the distribution law of kv cache deviation in independent documents (documents that independently calculate kv data), such as Figure 4 As shown in , a kv cache recalculation selection mechanism guided by prior knowledge is proposed to improve the generation quality of the model and reduce additional computational overhead. The specific solutions are as follows:
[0145] 1) Document block kv cache recalculation token selection:
[0146] based on Figure 4 It is found that the recalculation token selection strategy of the technical solution of this application mainly includes two key points:
[0147] (1) A small number of tokens at the beginning of each document block (i.e., the target document group mentioned above) will have a significant impact on the calculation results, so these tokens are directly selected for recalculation.
[0148] (2) The precision loss caused by the front independent document kv cache will be accumulated to the following document blocks. Therefore, the closer the document block is to the front, the more tokens need to be recalculated.
[0149] Based on the above observations, for the nth document block, let the total number of tokens that need to be recalculated for the kv data be l n , select the l at the beginning of the document block n tokens to recalculate. When n is larger, l n The smaller the value is, the less the deviation accumulation caused by the first n-1 independent document blocks can be. Specifically, Figure 16 As shown, for the first document block, no recalculation is performed because there is no loss of precision. For the subsequent N-1 document blocks, the proportion of recalculated tokens gradually decreases. This application allocates the number of recalculated tokens for each document block through a linear function as shown in formula (1). These tokens form queries and are provided to the large language model for recalculation of kv data.
[0150] Among them, l max and l min Indicates the number of recalculated tokens in the second and last documents, while the number of recalculated tokens in the middle documents decreases as n increases. According to experimental analysis, in actual scenarios, l max With l min The selection range is about 10% to 20% of the total number of document tokens. That is, if the total number of tokens in the background document is 2000, the number of tokens that need to be recalculated is about 200 to 400, which can achieve a large model generation quality close to that of the full-precision prefix cache.
[0151] 2) Optimization of recalculation of token quantity:
[0152] In the technical solution of the present application, since the selection rule for recalculating tokens is to select only a small number of tokens at the beginning of the discontinuous document block, the fewer the number of groups, the fewer the number of discontinuous documents, and the fewer the number of tokens that need to be recalculated. At the same time, based on analysis, document grouping will further improve the accuracy of RAG reasoning. Therefore, the technical solution of the present application uses document grouping to further reduce the number of recalculated tokens. According to existing methods, the document kv cache of the prefix can be stored using a tree structure. Therefore, the present application controls the number of document groups by adaptively setting the depth D of the cache tree.
[0153] Assuming that there are M documents in the database, for a cache tree with a depth of D, the maximum number of cache nodes stored is Therefore, the greater the depth of the tree, the greater the storage space required. In the technical solution of this application, the depth of the tree is adaptively set according to the total size of the space allocated to the kv cache. Related research shows that in the RAG scenario, a portion of the kv cache will be accessed frequently. Therefore, this application sets a coefficient α to measure the proportion of high-frequency cache. Assuming that the average kv cache size of each document is Bbyte, the estimated maximum tree depth that can be set is as shown in formula (2). Where, B max The maximum size of space that can be allocated to the kv cache.
[0154] 2. Optimization of iterative KV cache recalculation process:
[0155] Based on the recalculation token selection strategy proposed in the technical solution of this application, the entire recalculation process is optimized to further improve the RAG reasoning efficiency based on document kv cache.
[0156] 1) Iterative KV cache recalculation:
[0157] In the technical solution of this application, for each document block, only a small number of tokens at the beginning are selected for recalculation. Therefore, an iterative recalculation strategy can be adopted to further reduce the computational overhead of the recalculation process. The main process is as follows: Figure 17 shown.
[0158] For a scenario with N background document blocks (i.e., neighboring document blocks), an iterative method is used to recalculate the initial tokens of the 2nd to Nth document blocks in turn. Figure 17 As shown in the figure, the first l2 tokens of document 2 are used as the query, and the key-value cache of document 1 is used as the key and value. The mutual attention is recalculated to obtain an updated key-value cache for these l2 tokens. The independent key-value cache of the remaining L2-l2 tokens of document 2 is then directly concatenated to the recalculated tokens. Next, the updated key-value caches of documents 1 and 2 are used as the key and value, and the l3 tokens at the beginning of document 3 are used as the query to recalculate the updated key-value cache for these l3 tokens. This process continues in this manner until all document blocks are recalculated.
[0159] It can be seen that the iterative recalculation strategy adopted in this application can further reduce the recalculation overhead. Specifically, in a certain algorithm, a specified proportion of tokens is directly selected as the query, and all document block kv caches are used as keys and values. The complexity of the attention calculation process is By adopting the iterative method of this solution, the l nEach token only needs to calculate attention with the n-1 document blocks before it, and the computational complexity is further reduced to
[0160] 2) KV cache recalculation process optimization:
[0161] Since the technical solution of this application adopts an iterative recalculation strategy, the entire recalculation process can be further optimized. Figure 18 As shown in .
[0162] It can be seen that for the existing recalculation method ( Figure 18 In the upper part of the figure), all kv caches such as kv cache 1 (kv data of document 1), kv cache 2 (kv data of document 2), kv cache 3 (kv data of document 3) and kv cache 4 (kv data of document 4) need to be loaded into the video memory before the subsequent recalculation process can be started. Therefore, it adds a part of the time overhead for recalculation. In the technical solution of this application, due to the use of an iterative recalculation strategy ( Figure 18 ), the kv cache loading process and part of the recalculation process can be overlapped in time. Since the number of recalculated tokens in each iteration is small, the recalculation time is less than the time to load the complete document kv cache from other storage media (such as hard disk) to the video memory. Therefore, each time the application loads, after loading the starting token that needs to be recalculated, it starts to recalculate, and at the same time loads the kv cache of the remaining tokens that do not need to be recalculated in the background. Based on this optimization process, the time consumption caused by recalculation can be minimized to achieve a time efficiency close to that of not recalculating.
[0163] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0164] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0165] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0166] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data processing method, comprising: Obtain the kv data corresponding to each of the multiple target document groups; The kv data is the kv data obtained by independently calculating the target document group; The target document groups are connected in sequence; In at least one of the target document groups, the first X consecutive characters are selected as target characters; X is a positive integer greater than or equal to 1; X is less than the number of characters included in the corresponding target document group; Update the kv data corresponding to each target character in turn.
2. The method according to claim 1, updating the kv data corresponding to the target character, comprising: Determine a prefix document group corresponding to the target character; the prefix document group includes all target document groups sorted before the target document group where the target character is located; Based on the kv data corresponding to the prefix document group, the kv data corresponding to the target character is recalculated.
3. The method according to claim 2, wherein the target document group in which the target character is located comprises: any target data group among the plurality of target document groups except the first target data group sorted; Wherein, in the case that the target character is included in the prefix document group, the kv data corresponding to the target character in the prefix document group has been recalculated.
4. The method according to claim 1 or 2, wherein different target document groups contain different numbers of target characters; in, The number of the target characters in the target document group that is sorted first is greater than the number of the target characters in the target document group that is sorted last.
5. The method according to claim 1 or 2, wherein the number of the target characters contained in the adjacently sorted target document groups varies linearly based on a maximum value and a minimum value; in, The maximum value is an upper limit value of the number of target characters contained in one target document group; The minimum value is a lower limit value of the number of the target characters contained in one target document group.
6. The method according to claim 1 or 2, wherein obtaining the kv data corresponding to each of the plurality of target document groups comprises: According to the input data, a plurality of neighboring document blocks associated with the input data are retrieved from a database; Obtaining the target document group according to the neighboring document blocks, wherein the target document group consists of at least one neighboring document block; Read the kv data corresponding to each target document group from the cache pool; The maximum number of neighboring document blocks included in the target document group is determined based on the size of the kv data corresponding to the characters and the maximum storage space of the buffer pool.
7. The method according to claim 6, wherein when the kv data corresponding to the target document group is not read in the cache pool, the method further comprises: Utilize the intelligent model to calculate the kv data corresponding to the target document group.
8. The method according to claim 7, after calculating the key-value data corresponding to the target document group using the intelligent model, the method further comprises: The calculated kv data corresponding to the target document group is stored in the cache pool.
9. The method according to claim 6, wherein the kv data read from the cache pool is loaded into a target storage area; in, Update the kv data corresponding to each target character in sequence, including: For the target characters loaded into the target storage area, updating the kv data corresponding to the target characters in the target storage area in sequence; The updating of the kv data corresponding to the target character in the target storage area and the loading of the kv data corresponding to other characters sorted after the updated target character into the target storage area are performed in parallel.
10. An electronic device comprising: A memory for storing computer programs and data generated by the execution of the computer programs; A processor is configured to execute the computer program to achieve the following: obtaining kv data corresponding to each of a plurality of target document groups; the kv data being kv data obtained by independently calculating the target document groups; the target document groups being sequentially connected; in at least one of the target document groups, selecting the first X consecutive characters as target characters; X being a positive integer greater than or equal to 1; X being less than the number of characters contained in the corresponding target document group; and sequentially updating the kv data corresponding to each of the target characters.