Text processing method, apparatus, device, and storage medium

CN122779014APending Publication Date: 2026-09-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510318795.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]但将长文本内容压缩为短文本内容的过程中,会存在语义变化甚至丢失的问题,以短文本内容替代长文本内容作为模型输入,可能会影响模型输出的准确性

Benefits of technology

[0029] In this embodiment, the input text is divided into N text windows. A large language model is then used to chain-compress these N text windows, resulting in N compression tokens corresponding to each window. This avoids the semantic loss caused by directly compressing the input text. Furthermore, by mapping the N compression tokens to leaf nodes in a tree structure, the N compression tokens are aggregated to obtain aggregated compression tokens corresponding to each non-leaf node in the tree structure. Based on the N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text corresponding to the question text, the large language model generates the target answer text corresponding to the question text. This enables the generation of the target answer text based on the global text features of the input text, thereby improving the accuracy of answering the question text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779014A_ABST
    Figure CN122779014A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a text processing method, device and equipment, and a storage medium, and belong to the technical field of text processing. The method comprises: obtaining N compressed tokens corresponding to N text window through chain compression of the N text window by a large language model, the N text window being obtained by dividing input text, the input text comprising a question text and a reference text; aggregating the N compressed tokens based on a tree structure to obtain an aggregated compressed token corresponding to each non-leaf node in the tree structure, the N compressed tokens corresponding to leaf nodes of the tree structure one by one; and generating a target answer text corresponding to the question text by the large language model based on the N compressed tokens, the aggregated compressed token corresponding to each non-leaf node in the tree structure and an initial answer text of the question text. The embodiments provided in the present application can improve the accuracy of the answer to the question text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text processing technology, and in particular to a text processing method, apparatus, device, and storage medium. Background Technology

[0002] With the explosion of knowledge and the rapid spread of information, a large amount of question-and-answer text data is being generated and stored, which places higher demands on storage devices and network bandwidth.

[0003] In related technologies, in order to process long text content through a model, the long text content is compressed according to the maximum input length of the model before being input into the model. The compressed short text content is then input into the model to enable the model to process the text content.

[0004] However, compressing long text content into short text content may result in semantic changes or even loss. Using short text content instead of long text content as model input may affect the accuracy of the model output. Summary of the Invention

[0005] This application provides a text processing method, apparatus, device, and storage medium, the technical solutions of which are as follows:

[0006] On the one hand, embodiments of this application provide a text processing method, the method comprising:

[0007] By chaining and compressing N text windows using a large language model, N compression tokens are obtained corresponding to the N text windows. The N text windows are obtained by dividing the input text. The input text includes question text and reference text. The reference text is the source of the answer text corresponding to the question text. The i-th compression token in the N compression tokens is used to characterize the text features of the i-th text window. N is a positive integer, and i is less than or equal to N.

[0008] The N compressed tokens are aggregated based on a tree structure to obtain the aggregated compressed tokens corresponding to each non-leaf node in the tree structure. The N compressed tokens correspond one-to-one with the leaf nodes of the tree structure.

[0009] Based on the N compression tokens, the aggregate compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text, the target answer text corresponding to the question text is generated by the large language model. The initial answer text is the answer text output by the large language model after performing chained compression on the N text windows, and the target answer text is the answer text output by the large language model after performing text processing on the initial answer text.

[0010] On the other hand, embodiments of this application provide a text processing method, the method comprising:

[0011] By chaining and compressing N sample text windows using a large language model, N sample compression tokens corresponding to the N sample text windows are obtained. The N sample text windows are obtained by dividing the sample input text, which includes sample question text and sample reference text.

[0012] Based on the sample tree structure, the N sample compression tokens are aggregated to obtain the sample aggregate compression token corresponding to each non-leaf node of the sample in the sample tree structure. The N sample compression tokens correspond one-to-one with the sample leaf nodes of the sample tree structure.

[0013] Based on the N sample compression tokens, the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure, and the initial sample answer text of the sample question text, the sample answer text corresponding to the sample question text is generated through the large language model.

[0014] Based on the sample answer text corresponding to the sample question text and the ground truth value of the answer text, the model prediction loss of the large language model is determined;

[0015] Based on the model prediction loss, the model parameters of the large language model are updated.

[0016] On the other hand, embodiments of this application provide a text processing apparatus, the apparatus comprising:

[0017] The text compression module is used to chain-compress N text windows using a large language model to obtain N compression tokens corresponding to the N text windows. The N text windows are obtained by dividing the input text. The input text includes question text and reference text. The reference text is the source of the answer text corresponding to the question text. The i-th compression token among the N compression tokens is used to characterize the text features of the i-th text window. N is a positive integer, and i is less than or equal to N.

[0018] The token aggregation module is used to aggregate the N compressed tokens based on a tree structure to obtain the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, wherein the N compressed tokens correspond one-to-one with the leaf nodes of the tree structure.

[0019] The model output module is used to generate the target answer text corresponding to the question text based on the N compression tokens, the aggregate compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text through the large language model. The initial answer text is the answer text output by the large language model after performing chain compression on the N text windows, and the target answer text is the answer text output by the large language model after performing text processing on the initial answer text.

[0020] On the other hand, embodiments of this application provide a text processing apparatus, the apparatus comprising:

[0021] The text compression module is used to chain-compress N sample text windows using a large language model to obtain N sample compression tokens corresponding to the N sample text windows. The N sample text windows are obtained by dividing the sample input text, and the sample input text includes sample question text and sample reference text.

[0022] The token aggregation module is used to aggregate the N sample compressed tokens based on the sample tree structure to obtain the sample aggregated compressed tokens corresponding to each sample non-leaf node in the sample tree structure. The N sample compressed tokens correspond one-to-one with the sample leaf nodes of the sample tree structure.

[0023] The model output module is used to generate sample answer text corresponding to the sample question text through the large language model based on the N sample compression tokens, the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure, and the sample initial answer text of the sample question text.

[0024] The loss determination module is used to determine the model prediction loss of the large language model based on the sample answer text corresponding to the sample question text and the truth value of the answer text.

[0025] The parameter update module is used to update the model parameters of the large language model based on the model prediction loss.

[0026] On the other hand, embodiments of this application provide a computer device including a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the text processing method as described above.

[0027] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the text processing method as described above.

[0028] On the other hand, embodiments of this application provide a computer program product including at least one instruction stored in a computer-readable storage medium. A processor of a computer device reads the at least one instruction from the computer-readable storage medium and executes the at least one instruction, causing the computer device to perform the text processing method described above.

[0029] In this embodiment, the input text is divided into N text windows. A large language model is then used to chain-compress these N text windows, resulting in N compression tokens corresponding to each window. This avoids the semantic loss caused by directly compressing the input text. Furthermore, by mapping the N compression tokens to leaf nodes in a tree structure, the N compression tokens are aggregated to obtain aggregated compression tokens corresponding to each non-leaf node in the tree structure. Based on the N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text corresponding to the question text, the large language model generates the target answer text corresponding to the question text. This enables the generation of the target answer text based on the global text features of the input text, thereby improving the accuracy of answering the question text. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This application shows a structural block diagram of a computer system provided in an exemplary embodiment;

[0032] Figure 2 A flowchart illustrating a text processing method provided in an exemplary embodiment of this application is shown;

[0033] Figure 3 A flowchart illustrating a chain compression process provided in an exemplary embodiment of this application is shown;

[0034] Figure 4 This illustration shows a schematic diagram of a binary tree-structured aggregated token provided in an exemplary embodiment of this application;

[0035] Figure 5 A schematic diagram of a binary tree-structured aggregated token provided in another exemplary embodiment of this application is shown;

[0036] Figure 6 A flowchart illustrating a process for generating target response text provided in an exemplary embodiment of this application is shown;

[0037] Figure 7 This illustration shows a process of performing text processing through a large language model, provided by an exemplary embodiment of this application.

[0038] Figure 8 A flowchart of a text processing method provided by another exemplary embodiment of this application is shown;

[0039] Figure 9 This illustration shows a question-and-answer interface diagram in a question-and-answer scenario provided by an exemplary embodiment of this application;

[0040] Figure 10 This invention provides a structural block diagram of a text processing apparatus according to an exemplary embodiment of the present application.

[0041] Figure 11 A structural block diagram of a text processing apparatus provided in another exemplary embodiment of this application is shown;

[0042] Figure 12 A schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0044] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0045] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0046] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0047] First, a brief introduction to the terms used in the embodiments of this application:

[0048] Large Language Models (LLMs) are neural network models trained on large-scale text data that learn language patterns, structure, and semantics, thereby enabling the understanding and generation of natural language. LLMs can be used for a variety of natural language processing tasks, including but not limited to text generation, translation, question answering, text classification, and sentiment analysis.

[0049] Chained compression is a data compression method that uses specific encoding techniques to compress data, thereby reducing data storage space or improving data transmission efficiency. In this embodiment, by applying chained compression to the processing of input text by a large language model, the large language model can improve its inference efficiency by generating rich, sequentially compressed tokens.

[0050] Attention mechanism: A computational model that mimics human visual attention, designed to dynamically focus on the parts of the input data most relevant to the current task, rather than processing all input information equally. In the attention mechanism, Q (Query), K (Key), and V (Value) are the core components. They perform weighted processing of the input data through specific computational procedures, allowing the model to dynamically focus on the information most relevant to the current task.

[0051] Please refer to Figure 1 This illustration shows a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system may include a terminal 110 and a server 120. The terminal 110 and the server 120 communicate via a communication network. Optionally, the communication network may be a wired network or a wireless network, and the communication network may be at least one of a local area network (LAN), a metropolitan area network (MAN), and a wide area network (WAN).

[0052] Terminal 110 is an electronic device with an application that has a question-and-answer function installed. The question-and-answer function can be a function of a native application in terminal 110, or a function of a third-party application; terminal 110 can be a smartphone, tablet, laptop, desktop computer, smart TV, wearable device, or vehicle terminal, etc. Figure 1 The example of terminal 110 being a desktop computer is used for illustration only, but it is not a limitation.

[0053] Server 120 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. In this embodiment, server 120 can be a backend server for an application with question-and-answer functionality.

[0054] In some embodiments, data interaction exists between server 120 and terminal 110. (Illustrative example, such as...) Figure 1 As shown, in response to a question-and-answer input operation, terminal 110 receives user input text and sends it to server 120. Server 120 then divides the input text into N text windows and performs chained compression on the N text windows using a large language model to obtain N compression tokens corresponding to the N text windows. These N compression tokens are then mapped one-to-one to the leaf nodes of a tree structure. Based on the tree structure, the N compression tokens are aggregated to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure. Thus, based on the N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text to the question text, the large language model generates the target answer text corresponding to the question text and returns the target answer text to terminal 110 so that terminal 110 can display the target answer text to the user.

[0055] Based on the above introduction, the text processing method provided in this application will be described. This method can be executed by a server or a terminal, or by both a server and a terminal.

[0056] Please refer to Figure 2 This document illustrates a flowchart of a text processing method provided in an exemplary embodiment of this application. This embodiment uses the method applied to a computer device (including a terminal and / or a server) as an example for illustration. The method includes the following steps:

[0057] Step 210: Chain compression of N text windows using a large language model to obtain N compression tokens corresponding to the N text windows. The N text windows are obtained by dividing the input text. The input text includes the question text and the reference text. The reference text is the source of the answer text corresponding to the question text. The i-th compression token in the N compression tokens is used to characterize the text features of the i-th text window. N is a positive integer, and i is less than or equal to N.

[0058] Optionally, a large language model is a neural network model trained on large-scale text data that can learn language patterns, structure, and semantics, thereby achieving the understanding and generation of natural language. Large language models can be used for a variety of natural language processing tasks, including but not limited to text generation, translation, question answering, text classification, and sentiment analysis.

[0059] Optionally, the input to the large language model is natural language text, which can be words, sentences, paragraphs, documents, etc. In this embodiment, the input text to the large language model includes question text and reference text, wherein the reference text is the source of the answer text corresponding to the question text. That is, the large language model needs to understand the question text and the reference text, and generate the answer text corresponding to the question text based on the reference text.

[0060] Optionally, the reference text can be in the form of paragraph text, document files, web page links, etc. Optionally, the question text is a statement used to obtain information or answers, and the question text has a clear intent. For example, the question text could be "Based on the following, explain what XX is."

[0061] In some embodiments, considering that large language models usually have a limit on the length of the input, and the input text includes both the question text and the reference text, that is, the length of the input text may exceed the maximum input length of the large language model, in order to ensure that the large language model can understand all the input text, the computer device can divide the input text into N text windows, and then call the large language model N times to perform semantic understanding and generation on the N text windows respectively.

[0062] Optionally, the window segment length of each text window can be the same or different. Optionally, the window segment length of each text window is w, and the i-th text window can be represented as X. i For example, the window segment length of each text window can be 1024.

[0063] Optionally, to avoid the large language model being unable to understand the context information after text segmentation, the computer device can use chain compression to chain compression of N text windows through the large language model, thereby obtaining N compression tokens corresponding to the N text windows.

[0064] Chain compression refers to sequentially performing text processing on N text windows to obtain a compression token corresponding to each text window, and using the first i-1 compression tokens corresponding to the first i-1 text windows in the compression process of the ith text window, so that the large language model can pay attention to the features of the preceding text.

[0065] In this model, the i-th compressed token among the N compressed tokens represents the text features of the i-th text window, and the token length of the i-th compressed token is significantly shorter than the text length of the i-th text window. For example, if the text length of the i-th text window is 1024, the token length of the i-th compressed token is 20. By compressing the i-th text window into the i-th compressed token, the length of the input text can be reduced while ensuring the model's understanding of the contextual semantics.

[0066] In other words, during the semantic understanding and generation of the i-th text window, it is not necessary to input the first i-1 text windows and the i-th text window together into the large language model. Instead, the large language model performs semantic understanding and generation of the i-th text window based on the i-1 compression tokens corresponding to the first i-1 text windows, thereby avoiding the text input length from exceeding the maximum input length of the large language model.

[0067] In one possible implementation, the computer device performs semantic understanding on the i-th text window using a large language model based on the i-1 compression tokens corresponding to the first i-1 text windows, thereby outputting the i-th answer text corresponding to the i-th text window. Simultaneously, during the generation of the i-th answer text, the i-th compression token corresponding to the i-th text window is obtained.

[0068] Step 220: Aggregate N compressed tokens based on the tree structure to obtain the aggregated compressed tokens corresponding to each non-leaf node in the tree structure. The N compressed tokens correspond one-to-one with the leaf nodes of the tree structure.

[0069] In the chained compression process, when understanding the Nth text window, the large language model has already understood the text features of the previous N-1 text windows. That is, after the large language model completes the semantic understanding of the Nth text window, the Nth answer text corresponding to the Nth text window can be used as the answer text to the question text. However, considering that the length of the input text may be long, during the chained compression process, as the large language model understands subsequent text windows, it may cause semantic loss of the previous text windows, thereby reducing the accuracy of the answer text. Therefore, in order to generate the answer text corresponding to the question text based on global semantic features, in this embodiment, the Nth answer text corresponding to the Nth text window is not directly used as the answer text to the question text.

[0070] Optionally, to enable the large language model to focus on the global text features of the input text, the computer device can aggregate the text features corresponding to each text window using a tree structure. In some embodiments, after obtaining N compression tokens corresponding to N text windows, the computer device can map the N compression tokens one-to-one with the leaf nodes in the tree structure, thereby aggregating the N compression tokens based on the tree structure to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure. Finally, the aggregated compression token corresponding to the parent node represents the local text features corresponding to the child node, and the aggregated compression token corresponding to the root node represents the global text features of the input text.

[0071] Optionally, the tree structure can be a binary tree structure, a multi-branch tree structure, or other structures that can achieve token aggregation. This application embodiment does not limit this.

[0072] Optionally, if the tree structure is a binary tree structure, the computer device can start from the leaf node and aggregate the compressed tokens corresponding to the adjacent leaf nodes, thereby aggregating upwards layer by layer to obtain the aggregated compressed tokens corresponding to each non-leaf node in the binary tree structure.

[0073] Optionally, when the tree structure is a multi-branch tree structure, the tree structure can be a two-level multi-branch tree structure, where the computer device directly aggregates the compressed tokens corresponding to each leaf node to obtain the aggregated compressed token corresponding to the root node; or it can be a multi-level multi-branch tree structure, such as a ternary tree structure, where the computer device aggregates the compressed tokens corresponding to three adjacent leaf nodes, and then aggregates the aggregated compressed tokens corresponding to three adjacent non-leaf nodes layer by layer upwards until the aggregated compressed token corresponding to the root node is obtained.

[0074] Optionally, the computer device can aggregate tokens based on a tree structure after obtaining N compressed tokens, or it can aggregate tokens synchronously based on a tree structure during the chain generation of compressed tokens. This application embodiment does not limit this.

[0075] Step 230: Based on N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text, the target answer text corresponding to the question text is generated through the large language model. The initial answer text is the answer text output by the large language model after performing chained compression on N text windows, and the target answer text is the answer text output by the large language model after performing text processing on the initial answer text.

[0076] In some embodiments, after obtaining the aggregated compression tokens corresponding to each non-leaf node in the tree structure, in order to output the answer text based on the global text features of the input text, the computer device can re-input the initial answer text into the large language model, and use the large language model to combine the N compression tokens and the aggregated compression tokens corresponding to each non-leaf node in the tree structure to generate the target answer text corresponding to the question text.

[0077] The initial response text is the response text output by the large language model after chaining compression across N text windows; specifically, it is the Nth response text corresponding to the Nth text window. The target response text is the response text output by the large language model after processing the initial response text. Compared to the initial response text, the target response text, generated based on N compression tokens and the aggregated compression tokens corresponding to each non-leaf node in the tree structure, takes into account global text features more fully, resulting in higher accuracy.

[0078] In summary, in this embodiment, the input text is divided into N text windows, and then the N text windows are chain-compressed using a large language model to obtain N compression tokens corresponding to the N text windows, thus avoiding the text semantic loss problem caused by directly compressing the input text. Furthermore, by mapping the N compression tokens one-to-one with the leaf nodes in the tree structure, the N compression tokens are aggregated to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure. Based on the N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text corresponding to the question text, the target answer text corresponding to the question text is generated using the large language model. This enables the generation of the target answer text based on the global text features of the input text, thereby improving the accuracy of answering the question text.

[0079] Chain compression process

[0080] To meet the maximum input length limit of a large language model while enabling it to generate response text based on its understanding of the entire input text, this embodiment divides the input text into N text windows. The large language model then performs chained compression on these N text windows, resulting in N compression tokens corresponding to each window. (Illustrative example follows.) Figure 3 As shown, the process includes the following steps:

[0081] Step 211: Perform text preprocessing on the i-th text window and the i-th initial compression token corresponding to the i-th text window through the large language model to obtain the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compression token.

[0082] Optionally, the i-th initial compression token is pre-set initial compression data. By inputting the i-th initial compression token into the large language model, the i-th initial compression token can be updated during the text processing of the large language model to obtain the i-th compression token corresponding to the i-th text window.

[0083] Optionally, the initial compression tokens corresponding to different text windows can have the same value, that is, the token embedding is shared between different text windows.

[0084] In one possible implementation, the computer device inputs the i-th text window and the i-th initial compression token corresponding to the i-th text window into the large language model, thereby performing text preprocessing on the i-th text window and the i-th initial compression token through the embedding layer of the large language model to obtain the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compression token.

[0085] Optionally, the i-th text window can represent X. i The i-th initial compressed token can be represented as <ct> i The i-th hidden layer representation obtained after text preprocessing can be represented as: The i-th compressed hidden layer can be represented as follows: Optional, Where w is the sequence length of the i-th text window, i.e., the window segment length; k i Let be the sequence length of the i-th initial compressed token, i.e., the number of tokens, and D be the feature dimension (the size of the hidden layer).

[0086] Step 212: Based on the QKV projection matrix corresponding to the attention mechanism in the large language model, project the i-th hidden layer representation and the i-th compressed hidden layer representation to obtain the i-th initial QKV vector.

[0087] In some embodiments, after obtaining the i-th hidden layer representation and the i-th compressed hidden layer representation, in order to perform attention learning on the i-th hidden layer representation and the i-th compressed hidden layer representation and output the i-th answer text corresponding to the i-th text window, the computer device needs to first project the i-th hidden layer representation and the i-th compressed hidden layer representation onto the QKV space to obtain the i-th initial QKV vector corresponding to the i-th text window.

[0088] Optionally, the computer device projects the i-th hidden layer representation and the i-th compressed hidden layer representation according to the QKV projection matrix corresponding to the attention mechanism in the large language model, thereby converting the hidden layer representation into three different vectors (Q, K, V vectors) to obtain the i-th initial QKV vector corresponding to the i-th text window.

[0089] Here, the Q vector represents the content that needs to be focused on, the K vector represents the features of the reference information, and the V vector provides content related to K.

[0090] Optionally, in order to process the uncompressed hidden layer representation and the compressed hidden layer representation separately, the present application embodiment sets a first QKV projection matrix and a second QKV projection matrix in the attention mechanism of the large language model. The first QKV projection matrix is ​​used to project the uncompressed hidden layer representation to the QKV space, and the second QKV projection matrix is ​​used to project the compressed hidden layer representation to the QKV space.

[0091] Optionally, the first QKV projection matrix can be represented as W nt The second QKV projection matrix can be represented as W ct .

[0092] In one possible implementation, the computer device projects the i-th hidden layer representation based on the first QKV projection matrix corresponding to the attention mechanism in the large language model to obtain the i-th initial text QKV vector, and projects the i-th compressed hidden layer representation based on the second QKV projection matrix corresponding to the attention mechanism in the large language model to obtain the i-th initial compressed QKV vector.

[0093] Optionally, the i-th hidden layer is represented Projection into Q space can be represented as Represent the i-th hidden layer Projection onto K-space can be represented as Represent the i-th hidden layer Projecting onto V space can be represented as in, This is the first Q-projection matrix. Let K be the first projection matrix. This is the first V projection matrix.

[0094] Optionally, the i-th compressed hidden layer is represented Projection into Q space can be represented as Represent the i-th compressed hidden layer Projection onto K-space can be represented as Represent the i-th compressed hidden layer Projecting onto V space can be represented as in, This is the second Q projection matrix. This is the second K-projection matrix. This is the second V projection matrix.

[0095] Furthermore, the computer device can fuse the i-th initial text QKV vector and the i-th initial compressed QKV vector to obtain the i-th initial QKV vector for attention calculation.

[0096] Optionally, regarding the vector fusion process, in one possible implementation, the computer device can directly concatenate the i-th initial text QKV vector and the i-th initial compressed QKV vector to obtain the i-th initial QKV vector for attention calculation, thus preserving the original information represented by each vector. In another possible implementation, the computer device can fuse the i-th initial text QKV vector and the i-th initial compressed QKV vector based on an attention layer to obtain the i-th initial QKV vector.

[0097] To combine the text features of the first i-1 text windows for attention calculation and achieve chained compression, thereby improving the accuracy of the large language model's understanding of the i-th text window, the computer device also needs to perform vector fusion by combining the first i-1 compression tokens corresponding to the first i-1 text windows. That is, the computer device obtains the i-th initial QKV vector based on the i-th initial text QKV vector, the i-th initial compressed QKV vector, and the first i-1 compression tokens corresponding to the first i-1 text windows.

[0098] Optionally, the i-th compressed token is used to represent the text features of the i-th text window. In the large language model, after the input text is transformed into vectors through the embedding layer, these vectors are transformed into QKV vectors through linear transformation. Then, through the attention mechanism, attention weights are generated by calculating the similarity between Q and K, and the attention weights are multiplied by the V vector to obtain the context vector. This process enables the KV vector to capture the semantic relationship between different tokens in the text. Therefore, the i-th compressed KV vector corresponding to the i-th text window can be used as the i-th compressed token corresponding to the i-th text window.

[0099] That is, the i-th compression token includes the i-th compressed KV vector corresponding to the i-th text window, and the i-1 compression tokens corresponding to the first i-1 text windows are the i-1 compressed KV vectors corresponding to the first i-1 text windows.

[0100] Optionally, in order to minimize the data processing volume of a large language model while combining global text features to output the target answer text, the computer device may also take only the last token of the i-th compressed K vector and the last token of the i-th compressed V vector as the i-th compressed token corresponding to the i-th text window.

[0101] Optionally, the process of obtaining the i-th initial QKV vector based on the i-th initial text QKV vector, the i-th initial compressed QKV vector, and the i-1 compressed tokens corresponding to the i-1 text windows is the process of obtaining the i-th initial QKV vector based on the i-th initial text QKV vector, the i-th initial compressed QKV vector, and the i-1 compressed KV vectors corresponding to the i-1 text windows.

[0102] In one possible implementation, the computer device fuses the i-th initial text Q vector and the i-th initial compressed Q vector to obtain the i-th initial Q vector; fuses the i-th initial text K vector, the i-th initial compressed K vector, and the first i-1 compressed K vectors corresponding to the first i-1 text windows to obtain the i-th initial K vector; and fuses the i-th initial text V vector, the i-th initial compressed V vector, and the first i-1 compressed V vectors corresponding to the first i-1 text windows to obtain the i-th initial V vector.

[0103] Optionally, when vector fusion is achieved through vector concatenation, the i-th initial Q-vector can be represented as... The initial K vector of the i-th order can be represented as... The initial V vector of the i-th order can be represented as... Among them, K cA V is used to represent the first i-1 compressed K vectors corresponding to the first i-1 text windows. ca Used to represent the first i-1 compressed V vectors corresponding to the first i-1 text windows.

[0104] The process of extracting key-value vectors from problem text

[0105] Optionally, the input text includes the question text and the reference text. The question text is usually much shorter than the reference text, meaning that text partitioning is essentially a partitioning of the reference text. However, to ensure that the large language model can combine the question text with the output when processing each text window, the computer can first divide the reference text into N text windows, and then add the question text to each text window. Obviously, with this text partitioning method, the large language model needs to process the question text N times, resulting in a waste of computing resources.

[0106] Based on the above problems, in one possible implementation, the computer device can perform position encoding and division of the problem text and the reference text according to the text sequence order of problem text first and reference text last. The data length of the problem text is usually small, which is equivalent to the problem text being divided into the first text window.

[0107] Optionally, the first text window can be represented as Among them, X que This indicates the problem text. This represents the reference text assigned to the first text window. Optionally, the i-th text window (i≥2) can be represented as... This refers to the reference text that has been assigned to the i-th text window.

[0108] Optionally, if the question text is only in the first text window, in order to ensure that the large language model can still combine the text features of the question text to output results when processing other text windows, the computer device can extract the question text KV vector corresponding to the question text after compressing the first text window, and cache it together with the i-1 compressed KV vectors corresponding to the first i-1 text windows.

[0109] In one possible implementation, the computer device compresses the first text window using a large language model to obtain the first text KV vector and the first compressed KV vector corresponding to the first text window, and then extracts the question text KV vector corresponding to the question text from the first text KV vector based on the text position encoding corresponding to the question text.

[0110] Regarding the compression process of the first text window, in one possible implementation, the computer device inputs the first text window and the first initial compression token corresponding to the first text window into a large language model. The first text window and the first initial compression token are encoded through the embedding layer of the large language model to obtain the first hidden layer representation corresponding to the first text window and the first compressed hidden layer representation corresponding to the first initial compression token. The first hidden layer representation and the first compressed hidden layer representation are then projected onto the QKV space, respectively, and vector fusion is used to obtain the first initial Q vector, the first initial K vector, and the first initial V vector. This is then processed... After cross-attention calculation, feedforward neural network processing, residual connections, and normalization, during the decoding stage, for each token output by the large language model, the first initial Q-vector needs to be updated to increase the contextual relevance of the response text. The updated first initial Q-vector and the first initial KV vector then undergo attention calculation again to obtain the updated first initial KV vector. Another token is then output. Finally, after outputting the first response text, the computer obtains the fully updated first initial KV vector, which includes the first text KV vector. and the first compressed KV vector

[0111] Furthermore, the computer device will compress the first KV vector. As the first compressed token cache, it is used for the attention calculation process of the second text window, and is encoded according to the text position corresponding to the question text, from the first text KV vector. Extract the KV vector corresponding to the question text. que .

[0112] For example, if the text position encoding corresponding to the question text is 0-20, then the computer device will assign the first text vector K. The 0th to 20th tokens in the vector K are determined as the question text K. que , the first text vector V The 0th to 20th tokens in the vector V are determined as the question text V. que .

[0113] Optionally, after caching the question text KV vector and the i-1 compressed KV vectors corresponding to the first i-1 text windows together, during the QKV vector fusion process for the i-th text window, the computer device can fuse the question text K vector, the i-th initial text K vector, the i-th initial compressed K vector, and the first i-1 compressed K vectors corresponding to the first i-1 text windows to obtain the i-th initial K vector; and fuse the question text V vector, the i-th initial text V vector, the i-th initial compressed V vector, and the first i-1 compressed V vectors corresponding to the first i-1 text windows to obtain the i-th initial V vector. This allows the large language model to output answers based on the text features of the question text during the processing of each text window.

[0114] Optionally, when vector fusion is achieved through vector concatenation, K ca That is, it can be expressed as V is used to represent the question text K vector and the compressed K vector corresponding to the first i-1 text windows; ca That is, it can be expressed as Used to represent the question text V vector and the first i-1 compressed V vectors corresponding to the first i-1 text windows.

[0115] Combining the above problem text KV vectors and the first i-1 compressed KV vectors corresponding to the first i-1 text windows, the i-th initial Q vector after vector concatenation can be represented as: The initial K vector of the i-th order can be represented as... The initial V vector of the i-th order can be represented as...

[0116] Step 213: Compress the i-th initial QKV vector using the large language model to obtain the i-th compressed token corresponding to the i-th text window.

[0117] Optionally, after obtaining the i-th initial QKV vector, the computer device can compress the i-th initial QKV vector through a large language model to obtain the i-th compressed token corresponding to the i-th text window.

[0118] In one possible implementation, the computer device performs cross-attention calculation on the i-th initial QKV vector using the large language model to obtain the i-th attention output result corresponding to the i-th text window. Then, the large language model performs feedforward neural network processing, residual connections, and normalization on the i-th attention output result, thereby decoding and outputting the i-th answer text corresponding to the i-th text window. During the output of the i-th answer text, the large language model updates the i-th initial QKV vector each time it outputs a token. After obtaining the updated i-th initial QKV vector, the computer device can determine the i-th compressed KV vector corresponding to the i-th text window based on the updated i-th initial QKV vector.

[0119] Optionally, the cross-attention calculation process may include V i =A i V i , Among them, A i Represents the attention score of the i-th digit. This represents the attention output result for the i-th text. This represents the output result of the i-th compressed attention. This represents the first output projection matrix. V[:w] represents the second output projection matrix, V[:w] represents the i-th text V vector, and V[w:] represents the i-th compressed V vector.

[0120] Regarding the process of updating the i-th initial QKV vector, optionally, during the output of the i-th answer text, the large language model first outputs a token and uses this token to update the i-th initial Q vector. The updated i-th initial Q vector and the i-th initial KV vector are then subjected to attention calculation again to obtain the attention output result. After the attention calculation, the i-th initial KV vector is updated. Subsequently, after processing by a feedforward neural network, residual connections, and normalization, the large language model continues to decode and output a token, and continues to update the i-th initial Q vector. Thus, after multiple attention calculations, the large language model outputs the complete i-th answer text and obtains the updated i-th initial QKV vector.

[0121] Optionally, the updated i-th initial QKV vector includes the i-th text QKV vector and the i-th compressed QKV vector. Thus, the computer device extracts the i-th compressed QKV vector from the updated i-th initial QKV vector according to the relative position encoding of each token in the vector, and uses the i-th compressed QKV vector as the i-th compressed token corresponding to the i-th text window.

[0122] Furthermore, computer equipment upgrades K ca and V ca And continue to compress the (i+1)th text window until the chained compression of N text windows is completed.

[0123] Among them, K ca The update process can be represented as K ca ←{K ca ;K que ;K ct After obtaining the i-th compression token, K ca The update process can be specifically represented as follows: V ca The update process can be represented as V ca ←{V ca V quE V ct After obtaining the i-th compression token, V Ca The update process can be specifically represented as follows:

[0124] In the above embodiments, during the chained compression of N text windows, the compressed KV vector corresponding to each text window is cached as a compression token. Thus, during the compression of the i-th text window, the i-1 compressed KV vector corresponding to the first i-1 text windows can be combined to generate the i-th compressed KV vector corresponding to the i-th text window, thereby ensuring the compression accuracy of each text window and improving the accuracy of the model's answer to the question text.

[0125] Furthermore, by dividing the question text into the first text window and extracting the question text KV vector from the first text KV vector after compressing the first text window, the question text KV vector is cached together with the compressed KV vectors corresponding to each text window. This allows the model to combine the question text features for text compression during the compression of each text window, thereby reducing the amount of model data processing while ensuring the accuracy of the model output.

[0126] Token aggregation process

[0127] In some embodiments, in order to improve the aggregation efficiency of compressed tokens and make the text features represented by the aggregated compressed tokens relatively balanced, the tree structure can adopt a binary tree structure.

[0128] Optionally, the N compressed tokens correspond one-to-one with the leaf nodes in the binary tree structure. Optionally, the binary tree structure requires aggregation of the compressed tokens corresponding to each pair of adjacent nodes.

[0129] In one possible implementation, when N is even, the computer device can aggregate the compressed tokens corresponding to adjacent leaf nodes pairwise. Optionally, the computer device aggregates the i-th compressed token with the (i+1)-th compressed token to obtain the aggregated compressed token corresponding to the (i+1) / 2-th parent node, where i is odd.

[0130] In another possible implementation, when N is odd, the compressed tokens corresponding to the first N-1 leaf nodes can be aggregated pairwise, while the Nth leaf node has no adjacent nodes that can be aggregated. In this case, the computer device can move the Nth leaf node up one level and determine the Nth compressed token as the aggregated compressed token corresponding to the (N+1) / 2th parent node. Here, the (N+1) / 2th parent node is the last node of the second-to-last level in the binary tree structure.

[0131] Furthermore, if the number of nodes in the penultimate level of the binary tree structure is even, the computer device can aggregate the aggregated compressed token corresponding to the (N-1) / 2th parent node with the aggregated compressed token corresponding to the (N+1) / 2th parent node; if the number of nodes in the penultimate level of the binary tree structure is odd, the computer device will continue to move the Nth leaf node up one level and determine the Nth compressed token as the aggregated compressed token corresponding to the (N+3) / 4th node in the third-to-last level.

[0132] Based on the two aggregation methods mentioned above, the computer device can aggregate the aggregated compressed tokens corresponding to each parent node from bottom to top, until the aggregated compressed tokens corresponding to each parent node and the root node in the binary tree structure are obtained.

[0133] Optionally, the set of nodes contained in a binary tree structure can be represented as C. v The compressed token corresponding to a node can be represented as h, and the aggregated compressed token corresponding to the root node can be represented as h. v The process of obtaining the aggregated compressed token corresponding to the root node through token aggregation can be represented as h. v =Agg({h u |u∈C v }).

[0134] Indicative, such as Figure 4 As shown, taking the division into 4 text windows as an example, 4 compression tokens are obtained through chain compression. Thus, the 4 compression tokens correspond one-to-one with the leaf nodes of the binary tree structure. The first compression token h1 and the second compression token h2, the third compression token h3 and the fourth compression token h4 can be aggregated in pairs to obtain aggregated compression token h5 and aggregated compression token h6. Then, further aggregating aggregated compression token h5 and aggregated compression token h6 will yield the aggregated compression token h7 corresponding to the root node.

[0135] Indicative, such as Figure 5 As shown, taking the division into 5 text windows as an example, 5 compression tokens are obtained through chained compression. Thus, the 5 compression tokens correspond one-to-one with the leaf nodes of the binary tree structure. The 1st compression token h1 and the 2nd compression token h2, the 3rd compression token h3 and the 4th compression token h4 can be aggregated in pairs to obtain aggregated compression token h6 and aggregated compression token h7. Since the 5th compression token h5 cannot be paired in pairs, it is moved up to the second to last level of the binary tree structure. However, aggregated compression token h6, aggregated compression token h7 and the 5th compression token h5 still cannot be paired in pairs. Therefore, according to the node order, aggregated compression token h6 and aggregated compression token h7 are aggregated first to obtain aggregated compression token h8. Then, aggregated compression token h8 and the 5th compression token h5 are aggregated to finally obtain aggregated compression token h9 corresponding to the root node.

[0136] Optionally, the i-th compressed token is used to represent the text features of the i-th text window. In the large language model, after the input text is transformed into vectors through the embedding layer, these vectors are transformed into QKV vectors through linear transformation. Then, through the attention mechanism, attention weights are generated by calculating the similarity between Q and K, and the attention weights are multiplied by the V vector to obtain the context vector. This process enables the KV vector to capture the semantic relationship between different tokens in the text. Therefore, the i-th compressed KV vector corresponding to the i-th text window can be used as the i-th compressed token corresponding to the i-th text window.

[0137] That is, the i-th compressed token includes the i-th compressed KV vector corresponding to the i-th text window. Optionally, N compressed tokens can be aggregated based on a tree structure, which is equivalent to aggregating N compressed KV vectors corresponding to N text windows.

[0138] Optionally, in order to minimize the data processing volume of a large language model while combining global text features to output the target answer text, the computer device can also take only the last token of the i-th compressed K vector and the last token of the i-th compressed V vector as the i-th compressed token corresponding to the i-th text window, so that the data volume of the compressed token corresponding to each node in the tree structure is two tokens.

[0139] Regarding the aggregation process between the i-th compressed token and the (i+1)-th compressed token, the computer device can aggregate the i-th compressed K vector and the (i+1)-th compressed K vector, and the i-th compressed V vector and the (i+1)-th compressed V vector, respectively.

[0140] In one possible implementation, the computer device performs a weighted summation of the i-th compressed K-vector and the (i+1)-th compressed K-vector based on key aggregation weights to obtain the aggregated compressed K-vector corresponding to the (i+1) / 2-th parent node. Simultaneously, it performs a weighted summation of the i-th compressed V-vector and the (i+1)-th compressed V-vector based on value aggregation weights to obtain the aggregated compressed V-vector corresponding to the (i+1) / 2-th parent node.

[0141] Among them, the bond aggregation weight is calculated by cross-attention based on the i-th compressed K vector and the (i+1)-th compressed K vector, and the value aggregation weight is calculated by cross-attention based on the i-th compressed V vector and the (i+1)-th compressed V vector.

[0142] Taking the determination of key aggregation weights as an example, the computer device can use the i-th compressed K-vector as Q in the cross-attention calculation, and the (i+1)-th compressed K-vector as K and V in the cross-attention calculation, thereby obtaining the key aggregation weight α1 of the i-th compressed K-vector relative to the (i+1)-th compressed K-vector through cross-attention calculation. Similarly, the computer device can use the (i+1)-th compressed K-vector as Q in the cross-attention calculation, and the i-th compressed K-vector as K and V in the cross-attention calculation, thereby obtaining the key aggregation weight α2 of the (i+1)-th compressed K-vector relative to the i-th compressed K-vector through cross-attention calculation. Furthermore, the computer device, based on the key aggregation weights α1 and α2, performs a weighted summation of the i-th compressed K-vector and the (i+1)-th compressed K-vector to obtain the aggregated compressed K-vector corresponding to the (i+1) / 2-th parent node.

[0143] By using the vector aggregation method described above, the compressed key-value vectors corresponding to each child node in the binary tree structure can be aggregated, thereby obtaining the aggregated compressed key-value vector corresponding to the root node in the binary tree structure. This allows the compressed key-value vectors corresponding to each node in the binary tree structure to represent the global text features of the input text.

[0144] In the above embodiments, during the aggregation of compressed tokens, a binary tree structure is used to aggregate the compressed tokens corresponding to each node to obtain the aggregated compressed tokens corresponding to each non-leaf node. This enables the large language model to understand the text features of the input text from a global perspective, thereby improving the accuracy of the model's response to the question text.

[0145] Output process of target response text

[0146] In some embodiments, after completing the chained compression of N text windows and obtaining an aggregated compression token representing the global text features of the input text through token aggregation, the computer device can call the large language model again to optimize the initial answer text, thereby generating the target answer text corresponding to the question text.

[0147] Indicative, such as Figure 6 As shown, the process may include the following steps:

[0148] Step 231: Input the initial answer text of the question text into the large language model, and perform text preprocessing on the initial answer text through the large language model to obtain the hidden layer representation of the text corresponding to the initial answer text.

[0149] In one possible implementation, the computer device re-inputs the initial answer text of the question text into the large language model, thereby first performing text preprocessing on the initial answer text through the embedding layer of the large language model, encoding the initial answer text into the corresponding text hidden layer representation.

[0150] Step 232: Based on the QKV projection matrix corresponding to the attention mechanism in the large language model, project the hidden layer representation of the text to obtain the text QKV vector corresponding to the hidden layer representation of the text.

[0151] Furthermore, the computer device projects the text hidden layer representation onto the QKV space based on the QKV projection matrix corresponding to the attention mechanism in the large language model, thereby obtaining the text QKV vector corresponding to the text hidden layer representation.

[0152] In this embodiment, the input attention mechanism of the large language model corresponds to a first QKV projection matrix and a second QKV projection matrix. The first QKV projection matrix is ​​used to project the uncompressed token, and the second QKV projection matrix is ​​used to project the compressed token. The initial response text is the Nth response text corresponding to the Nth text window, and the Nth response text is uncompressed text. Therefore, the computer device only needs to project the hidden layer representation of the text onto the QKV space through the first QKV projection matrix to obtain the text QKV vector.

[0153] Step 233: Based on N compressed tokens, the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, and the text QKV vector, the target answer text corresponding to the question text is generated through the large language model.

[0154] Optionally, after obtaining the text QKV vector corresponding to the initial answer text, the computer device can combine the N compression tokens and the aggregate compression tokens corresponding to each non-leaf node in the tree structure, call the large language model to understand the text QKV vector, and output the target answer text corresponding to the question text.

[0155] Optionally, the i-th compression token includes the i-th compression KV vector corresponding to the i-th text window, so that each aggregated compression token obtained after token aggregation processing includes the aggregated compression K vector and the aggregated compression V vector.

[0156] In one possible implementation, the computer device fuses the N compressed KV vectors corresponding to N text windows, the aggregated compressed KV vectors corresponding to each non-leaf node in the tree structure, and the text QKV vector to obtain the text Q vector, text K vector, and text V vector for cross-attention calculation. Then, based on the text Q vector, text K vector, and text V vector, the computer device obtains the attention output result corresponding to the text QKV vector through cross-attention calculation. Finally, the attention output result is processed by a large language model using a feedforward neural network, residual connections, and normalization to obtain the target answer text corresponding to the question text.

[0157] Optionally, the computer device performs an attention calculation on the text Q vector, text K vector, and text V vector using a large language model. After processing by a feedforward neural network, residual connections, and normalization, a response token can be output during the decoding stage. The text Q vector is then updated based on this response token. The updated text Q vector, text K vector, and text V vector are then processed again with attention calculation, feedforward neural network processing, residual connections, and normalization. Another response token is then decoded and output. This process is repeated multiple times until the response prediction is completed and the complete target response text is output.

[0158] In the above embodiments, after completing the chained compression of N text windows and obtaining the aggregated compression token through token aggregation, the initial answer text is input into the large language model again. By combining the large language model with the global text features represented by the aggregated compression token, the target answer text is regenerated, which can obtain a more accurate answer text and improve the model output quality.

[0159] Please refer to the above embodiments for details. Figure 7 This illustrates a schematic diagram of a process for performing text processing through a large language model, provided in an exemplary embodiment of this application.

[0160] In response to a text input operation, the computer device acquires input text 701, which includes question text and reference text. To ensure that the large language model 702 can understand all text while satisfying its maximum input length, the computer device divides the input text 701 into N text windows. Then, by calling the large language model 702 N times in a chained compression manner, N compression tokens corresponding to the N text windows are obtained.

[0161] Taking the compression of the i-th text window through the large language model 702 as an example, in one possible implementation, the computer device performs text preprocessing on the i-th text window and the i-th initial compression token corresponding to the i-th text window through the large language model 702 to obtain the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compression token. Based on the QKV projection matrix corresponding to the attention mechanism in the large language model 702, the i-th hidden layer representation and the i-th compressed hidden layer representation are projected to obtain the i-th initial QKV vector. Thus, the i-th initial QKV vector is compressed through the large language model 702 to obtain the i-th compression token corresponding to the i-th text window.

[0162] Furthermore, after compressing the Nth text window, the computer device can obtain the Nth answer text corresponding to the Nth text window output by the large language model 702. However, considering that the input text 701 may be quite long, even with chained compression, there may be semantic loss of previous text windows during the processing of subsequent text windows. Therefore, the Nth answer text corresponding to the Nth text window is not directly used as the answer text to the question text, but only the Nth answer text corresponding to the Nth text window is used as the initial answer text 703.

[0163] Furthermore, to improve the accuracy of the response text, during the process of generating N compression tokens corresponding to N text windows, the computer device also maps the N compression tokens one-to-one with the leaf nodes in the tree structure, thereby aggregating the N compression tokens based on the tree structure 704 to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure. The aggregated compression tokens corresponding to each non-leaf node in the tree structure collectively represent the global text features of the input text.

[0164] Finally, by calling the large language model 702 again, the computer device can generate the target answer text 705 corresponding to the question text based on the N compressed tokens, the aggregated compressed tokens corresponding to each non-leaf node in the tree structure 704, and the initial answer text 703 of the question text.

[0165] Model training process

[0166] In some embodiments, considering that during the attention calculation process, the hidden layer representation and the compressed hidden layer representation need to be projected into the QKV space respectively, and the text attention output result and the compressed attention output result are output projected respectively, in order to improve the accuracy of attention calculation, the large language model needs to be trained before generating the answer text using the large language model.

[0167] Please refer to Figure 8 This document illustrates a flowchart of a text processing method provided in another exemplary embodiment of this application. This embodiment uses the method applied to a computer device (including a terminal and / or a server) as an example for illustration. The method includes the following steps:

[0168] Step 801: Chain compression of N sample text windows is performed using a large language model to obtain N sample compression tokens corresponding to the N sample text windows. The N sample text windows are obtained by dividing the sample input text, which includes sample question text and sample reference text.

[0169] Optionally, the sample input text includes sample question text and sample reference text, where the sample reference text is the source of the sample answer text corresponding to the sample question text. That is, the large language model needs to understand the sample question text and the sample reference text, and generate the sample answer text corresponding to the sample question text based on the sample reference text.

[0170] Optionally, the sample reference text can be in the form of paragraph text, document files, web links, etc. Optionally, the sample question text is a statement used to obtain information or provide answers, and the sample question text has a clear intent. For example, the sample question text might be "Based on the following, explain what XX is."

[0171] In some embodiments, considering that large language models usually have a limit on the length of the input, and that the sample input text includes both sample question text and sample reference text, the length of the sample input text may exceed the maximum input length of the large language model. Therefore, in order to ensure that the large language model can understand all the sample input text, the computer device can divide the sample input text into N sample text windows, and then perform semantic understanding and generation on the N sample text windows by calling the large language model N times respectively.

[0172] Optionally, the window segment length of each sample text window can be the same or different. Optionally, the window segment length of each sample text window is w, and the i-th sample text window can be represented as X. i For example, the window segment length of each sample text window can be 1024.

[0173] Optionally, to avoid the large language model being unable to understand the contextual information after text segmentation, the computer device can use chain compression to chain-compress N sample text windows through the large language model, thereby obtaining N sample compression tokens corresponding to the N sample text windows.

[0174] In this model, the i-th sample compression token among the N sample compression tokens represents the text features of the i-th sample text window. The token length of the i-th sample compression token is significantly shorter than the text length of the i-th sample text window. For example, if the text length of the i-th sample text window is 1024, the token length of the i-th sample compression token is 20. By compressing the i-th sample text window into the i-th sample compression token, the length of the sample input text can be reduced while ensuring the model's understanding of the contextual semantics.

[0175] In other words, during the semantic understanding and generation of the i-th sample text window, it is not necessary to input the first i-1 sample text windows and the i-th sample text window together into the large language model. Instead, the large language model performs semantic understanding and generation of the i-th sample text window based on the i-1 sample compression tokens corresponding to the first i-1 sample text windows, thereby avoiding the text input length from exceeding the maximum input length of the large language model.

[0176] In one possible implementation, the computer device performs semantic understanding on the i-th sample text window using a large language model based on the i-1 sample compression tokens corresponding to the first i-1 sample text windows, thereby outputting the i-th sample response text corresponding to the i-th sample text window. Simultaneously, during the generation of the i-th sample response text, the i-th sample compression token corresponding to the i-th sample text window is obtained.

[0177] Step 802: Aggregate N sample compression tokens based on the sample tree structure to obtain the sample aggregate compression token corresponding to each non-leaf node of the sample in the sample tree structure. The N sample compression tokens correspond one-to-one with the sample leaf nodes of the sample tree structure.

[0178] In the chained compression process, when understanding the Nth sample text window, the large language model has already understood the text features of the previous N-1 sample text windows. That is, after the large language model completes the semantic understanding of the Nth sample text window, the Nth sample answer text output by the Nth sample text window can be used as the sample answer text for the sample question text. However, considering that the length of the sample input text may be relatively long, during the chained compression process, as the large language model understands subsequent sample text windows, it may cause semantic loss of the preceding sample text windows, thereby reducing the accuracy of the sample answer text. Therefore, in order to generate the sample answer text corresponding to the sample question text based on global semantic features, in this embodiment, the Nth sample answer text corresponding to the Nth sample text window is not directly used as the sample answer text for the sample question text.

[0179] Optionally, to enable the large language model to focus on the global text features of the sample input text, the computer device can aggregate the text features corresponding to each sample text window using a sample tree structure. In some embodiments, after obtaining N sample compression tokens corresponding to N sample text windows, the computer device can map each of the N sample compression tokens to a leaf node in the sample tree structure, thereby aggregating the N sample compression tokens based on the sample tree structure to obtain the sample aggregated compression token corresponding to each sample non-leaf node in the sample tree structure. Finally, the sample aggregated compression token corresponding to the sample parent node represents the local text features corresponding to the sample child node, and the sample aggregated compression token corresponding to the sample root node represents the global text features of the sample input text.

[0180] Optionally, the sample tree structure can be a binary tree structure, a multi-branch tree structure, or other structures that can achieve token aggregation. This application embodiment does not limit this.

[0181] Optionally, if the sample tree structure is a binary tree structure, the computer device can start from the sample leaf node and aggregate the sample compression tokens corresponding to the adjacent sample leaf nodes, thereby aggregating upwards layer by layer to obtain the sample aggregate compression tokens corresponding to each sample non-leaf node in the binary tree structure.

[0182] Optionally, when the sample tree structure is a multi-branch tree structure, the sample tree structure can be a two-level multi-branch tree structure, that is, the computer device directly aggregates the sample compression tokens corresponding to each sample leaf node to obtain the sample aggregated compression token corresponding to the sample root node; or it can be a multi-level multi-branch tree structure, such as a ternary tree structure, in which the computer device aggregates the sample compression tokens corresponding to the three adjacent sample leaf nodes, and then aggregates the sample aggregated compression tokens corresponding to the three adjacent sample non-leaf nodes layer by layer upwards until the sample aggregated compression token corresponding to the sample root node is obtained.

[0183] Optionally, the computer device can perform token aggregation based on the sample tree structure after obtaining N sample compressed tokens, or it can perform token aggregation synchronously based on the sample tree structure during the chain generation of sample compressed tokens. This application embodiment does not limit this.

[0184] Step 803: Based on N sample compression tokens, the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure, and the initial sample answer text of the sample question text, the sample answer text corresponding to the sample question text is generated through the large language model.

[0185] In some embodiments, after obtaining the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure, in order to output sample answer text based on the global text features of the sample input text, the computer device can re-input the initial sample answer text into the large language model, and use the large language model to combine the N sample compression tokens and the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure to generate the sample answer text corresponding to the sample question text.

[0186] The initial sample response text is the sample response text output by the large language model after chain compression of N sample text windows, that is, the Nth sample response text corresponding to the Nth sample text window.

[0187] Step 804: Based on the sample answer text corresponding to the sample question text and the ground truth value of the answer text, determine the model prediction loss of the large language model.

[0188] In one possible implementation, after obtaining the sample answer text corresponding to the sample question text, the computer device can determine the model prediction loss of the large language model based on the sample answer text and the truth value of the answer text.

[0189] Optionally, the model prediction loss can be expressed as P LM (X i,j |T;θ,θ ct ), where T represents ( <ct> 1,… <ct> i-1 ,X i,1 ,…X i,j-1 In the context of ), θ represents the projection matrix used to process uncompressed text, θ ct This represents the projection matrix used to process compressed text.

[0190] Step 805: Update the model parameters of the large language model based on the model prediction loss.

[0191] Furthermore, the computer equipment updates the model parameters of the large language model based on the model prediction loss, thereby completing the training of the large language model when the training completion conditions are met.

[0192] In this embodiment, the model parameters of the large language model mainly include the QKV projection matrix and the output projection matrix corresponding to the attention mechanism in the large language model. That is, the first QKV projection matrix, the second QKV projection matrix, the first output projection matrix, and the second output projection matrix in the above embodiment.

[0193] Optionally, θ in the model prediction loss includes the first Q projection matrix. First K-projection matrix First V projection matrix and the first output projection matrix θ ct That is, including the second Q projection matrix Second K-projection matrix Second V projection matrix and the second output projection matrix

[0194] In one possible implementation, the computer device can determine the QKV projection matrix and the output projection matrix when the training completion conditions are met as the target QKV projection matrix and the target output projection matrix, and then update the large language model based on the target QKV projection matrix and the target output projection matrix.

[0195] Optionally, the training completion condition can be the loss convergence condition, the number of training rounds, or other conditions. This application embodiment does not limit this.

[0196] Optionally, the training completion condition can be that the model's prediction loss reaches its minimum value, i.e. Where N is the number of sample text windows, and w is the window segment length of each sample text window, i.e., the number of tokens contained in the sample text window. Furthermore, considering that the above information prompt does not exist during the compression of the first text window, the first compressed token is not included when calculating the model's prediction loss.

[0197] In the above embodiments, during model training, the QKV projection matrix and output projection matrix that meet the training completion conditions are determined as the target QKV projection matrix and target output projection matrix, thereby optimizing the text compression process based on attention calculation and improving the text compression capability of the large language model.

[0198] Furthermore, training large language models based on the model prediction loss combined with context can improve model training efficiency and optimize model training results.

[0199] It should be noted that the above embodiments only provide a general description of the training process of the large language model. In the training process of the large language model, the chain compression process of N sample text windows can refer to the chain compression process of N text windows in the above application-side embodiments, the aggregation process of N sample compression tokens can refer to the aggregation process of N compression tokens in the above application-side embodiments, and the output process of sample answer text can refer to the output process of target answer text in the above application-side embodiments. This embodiment will not elaborate on these details.

[0200] As illustrated in Table 1, this table presents test data comparing the text compression method provided in this application with other related technologies on the Longbench dataset. Columns 1-3 represent test data for answering questions on a single reference document; columns 4-6 represent test data for answering questions on multiple reference texts; columns 7-9 represent test data for answering summary-type questions; columns 10-12 represent test data for answering few-sample tasks; columns 13-14 represent test data for answering code generation tasks; and column 15 represents the average value of the test data.

[0201] Table 1

[0202]

[0203] Optionally, the text processing method provided in this application embodiment can be applied to various scenarios, such as question-and-answer generation scenarios, information summarization scenarios, etc.

[0204] For question-and-answer generation scenarios:

[0205] In an optional example, the text processing method provided in this application is applied to a question-answer generation scenario, with the large language model serving as a question-answer model, as an example. This is illustrative, such as... Figure 9 As shown, it illustrates an interface diagram in a question-and-answer scenario provided by an exemplary embodiment of this application.

[0206] In some embodiments, the computer device performs chained compression of N text windows using a question-and-answer model to obtain N compression tokens corresponding to the N text windows. The N text windows are obtained by dividing the input text, which includes the question text and the reference text. The reference text is the source of the answer text corresponding to the question text. The i-th compression token among the N compression tokens is used to characterize the text features of the i-th text window. N is a positive integer, and i is less than or equal to N. The N compression tokens are aggregated based on a tree structure to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure. The N compression tokens correspond one-to-one with the leaf nodes of the tree structure. Based on the N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text, the question-and-answer model generates the target answer text corresponding to the question text. The initial answer text is the answer text output by the question-and-answer model after performing chained compression of the N text windows. The target answer text is the answer text output by the large language model after performing text processing on the initial answer text.

[0207] Optionally, based on the binary tree structure, when N is even, the i-th compressed token and the (i+1)-th compressed token are aggregated to obtain the aggregated compressed token corresponding to the (i+1) / 2-th parent node, where i is odd; when N is odd, the N-th compressed token is determined as the aggregated compressed token corresponding to the (N+1) / 2-th parent node; the aggregated compressed tokens corresponding to each parent node are aggregated layer by layer to obtain the aggregated compressed tokens corresponding to each parent node in the binary tree structure and the aggregated compressed token corresponding to the root node.

[0208] Optionally, the i-th compressed token includes the i-th compressed KV vector corresponding to the i-th text window. In one possible implementation, the computer device performs a weighted summation of the i-th compressed K vector and the (i+1)-th compressed K vector based on the key aggregation weight to obtain the aggregated compressed K vector corresponding to the (i+1) / 2-th parent node. The key aggregation weight is calculated based on the i-th compressed K vector and the (i+1)-th compressed K vector through cross-attention. Similarly, the computer device performs a weighted summation of the i-th compressed V vector and the (i+1)-th compressed V vector based on the value aggregation weight to obtain the aggregated compressed V vector corresponding to the (i+1) / 2-th parent node. The value aggregation weight is calculated based on the i-th compressed V vector and the (i+1)-th compressed V vector through cross-attention.

[0209] Optionally, the computer device inputs the initial answer text of the question text into the question-answering model, performs text preprocessing on the initial answer text through the question-answering model to obtain the text hidden layer representation corresponding to the initial answer text; projects the text hidden layer representation based on the QKV projection matrix corresponding to the attention mechanism in the question-answering model to obtain the text QKV vector corresponding to the text hidden layer representation; and generates the target answer text corresponding to the question text through the question-answering model based on N compression tokens, the aggregate compression tokens corresponding to each non-leaf node in the tree structure, and the text QKV vector.

[0210] Optionally, the aggregated compression token includes an aggregated compressed K vector and an aggregated compressed V vector. In one possible implementation, the computer device fuses the N compressed KV vectors corresponding to N text windows, the aggregated compressed KV vectors corresponding to each non-leaf node in the tree structure, and the text QKV vector to obtain the text Q vector, text K vector, and text V vector for cross-attention calculation; based on the text Q vector, text K vector, and text V vector, the attention output result corresponding to the text QKV vector is obtained through cross-attention calculation; the attention output result is processed by a question-answering model using a feedforward neural network, residual connections, and normalization to obtain the target answer text corresponding to the question text.

[0211] Optionally, the computer device performs text preprocessing on the i-th text window and the i-th initial compressed token corresponding to the i-th text window through a question-answering model to obtain the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compressed token; based on the QKV projection matrix corresponding to the attention mechanism in the question-answering model, the i-th hidden layer representation and the i-th compressed hidden layer representation are projected to obtain the i-th initial QKV vector; the i-th initial QKV vector is compressed through the question-answering model to obtain the i-th compressed token corresponding to the i-th text window.

[0212] Optionally, the computer device projects the i-th hidden layer representation based on the first QKV projection matrix corresponding to the attention mechanism in the question-answering model to obtain the i-th initial text QKV vector; projects the i-th compressed hidden layer representation based on the second QKV projection matrix corresponding to the attention mechanism in the question-answering model to obtain the i-th initial compressed QKV vector; and obtains the i-th initial QKV vector based on the i-th initial text QKV vector, the i-th initial compressed QKV vector, and the first i-1 compressed tokens corresponding to the first i-1 text windows.

[0213] Optionally, the i-th compression token includes the i-th compressed KV vector corresponding to the i-th text window. In one possible implementation, the computer device fuses the i-th initial text Q vector and the i-th initial compressed Q vector to obtain the i-th initial Q vector; fuses the i-th initial text K vector, the i-th initial compressed K vector, and the first i-1 compressed K vectors corresponding to the first i-1 text windows to obtain the i-th initial K vector; and fuses the i-th initial text V vector, the i-th initial compressed V vector, and the first i-1 compressed V vectors corresponding to the first i-1 text windows to obtain the i-th initial V vector.

[0214] Optionally, the question text is assigned to the first text window. In one possible implementation, the computer device compresses the first text window using a question-and-answer model to obtain the first text key-value vector and the first compressed key-value vector corresponding to the first text window; based on the text position encoding corresponding to the question text, the question text key-value vector corresponding to the question text is extracted from the first text key-value vector.

[0215] Furthermore, the computer device merges the problem text K vector, the i-th initial text K vector, the i-th initial compressed K vector, and the i-1 compressed K vectors corresponding to the first i-1 text windows to obtain the i-th initial K vector; and merges the problem text V vector, the i-th initial text V vector, the i-th initial compressed V vector, and the i-1 compressed V vectors corresponding to the first i-1 text windows to obtain the i-th initial V vector.

[0216] Optionally, the computer device performs cross-attention calculation on the i-th initial QKV vector through a question-answering model to obtain the i-th attention output result corresponding to the i-th text window; performs feedforward neural network processing, residual connection, and normalization processing on the i-th attention output result through the question-answering model to obtain the i-th answer text corresponding to the i-th text window; during the output of the i-th answer text, the i-th initial QKV vector is updated to obtain the updated i-th initial QKV vector; based on the updated i-th initial QKV vector, the i-th compressed KV vector corresponding to the i-th text window is determined.

[0217] Optionally, the question-and-answer scenario can be located in a standalone question-and-answer application or in a question-and-answer webpage.

[0218] Applications in AI question-answering applications:

[0219] In one application scenario, the text processing method provided in this application embodiment can be applied to an AI question-answering application. Optionally, in the AI ​​question-answering application, a question-answering interface is displayed, and in response to the user's question input operation and document upload operation, the computer device obtains the input text, which includes the question text and the reference text. Then, the computer device divides the input text into N text windows, and performs chained compression on the N text windows using a large language model to obtain N compression tokens corresponding to the N text windows.

[0220] Taking the compression of the i-th text window through a large language model as an example, in one possible implementation, the computer device performs text preprocessing on the i-th text window and the i-th initial compression token corresponding to the i-th text window through the large language model, to obtain the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compression token. Based on the QKV projection matrix corresponding to the attention mechanism in the large language model, the i-th hidden layer representation and the i-th compressed hidden layer representation are projected to obtain the i-th initial QKV vector. Thus, the i-th initial vector is compressed through the large language model to obtain the i-th compression token corresponding to the i-th text window.

[0221] Furthermore, after compressing the Nth text window, the computer device can obtain the Nth response text corresponding to the Nth text window output by the large language model, and use the Nth response text corresponding to the Nth text window as the initial response text. Simultaneously, during the generation of N compression tokens corresponding to the N text windows, the computer device also maps the N compression tokens one-to-one with the leaf nodes in the tree structure, thereby aggregating the N compression tokens based on the tree structure to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure.

[0222] Regarding the token compression process, in one possible implementation, when N is even, the computer device can aggregate the compressed tokens corresponding to adjacent leaf nodes pairwise. Optionally, the computer device aggregates the i-th compressed token with the (i+1)-th compressed token to obtain the aggregated compressed token corresponding to the (i+1) / 2-th parent node, where i is odd. In another possible implementation, when N is odd, the compressed tokens corresponding to the first N-1 leaf nodes can be aggregated pairwise, while the N-th leaf node has no aggregable adjacent nodes. In this case, the computer device can move the N-th leaf node up one level, determining the N-th compressed token as the aggregated compressed token corresponding to the (N+1) / 2-th parent node. Here, the (N+1) / 2-th parent node is the last node of the second-to-last level in the binary tree structure. Furthermore, when the number of nodes in the penultimate level of the binary tree structure is even, the computer device can aggregate the aggregated compressed token corresponding to the (N-1) / 2th parent node with the aggregated compressed token corresponding to the (N+1) / 2th parent node. When the number of nodes in the penultimate level of the binary tree structure is odd, the computer device continues to move the Nth leaf node up one level, determining the Nth compressed token as the aggregated compressed token corresponding to the (N+3) / 4th node in the third-to-last level. Based on these two aggregation methods, the computer device can aggregate the aggregated compressed tokens corresponding to each parent node layer by layer from bottom to top, until it obtains the aggregated compressed tokens corresponding to the parent nodes at each level of the binary tree structure and the aggregated compressed token corresponding to the root node.

[0223] Finally, the computer device, by calling the large language model again, can generate the target answer text corresponding to the question text based on N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text. This target answer text is then displayed on the question-and-answer interface of the AI ​​question-answering application. Compared to the initial answer text, the target answer text generated based on N compression tokens and the aggregated compression tokens corresponding to each non-leaf node in the tree structure takes into account global text features more fully, resulting in higher accuracy in answering the question text.

[0224] Application in AI question-answering web pages:

[0225] In one application scenario, the text processing method provided in this application embodiment can be applied to an AI question-and-answer webpage. Optionally, in the AI ​​question-and-answer webpage, a question-and-answer interface is displayed, thereby responding to the user's question input operation and document upload operation. The computer device obtains the input text, which includes the question text and the reference text. Then, the computer device divides the input text into N text windows, and performs chained compression on the N text windows using a large language model to obtain N compression tokens corresponding to the N text windows.

[0226] Taking the compression of the i-th text window through a large language model as an example, in one possible implementation, the computer device performs text preprocessing on the i-th text window and the i-th initial compression token corresponding to the i-th text window through the large language model, to obtain the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compression token. Based on the QKV projection matrix corresponding to the attention mechanism in the large language model, the i-th hidden layer representation and the i-th compressed hidden layer representation are projected to obtain the i-th initial QKV vector. Thus, the i-th initial vector is compressed through the large language model to obtain the i-th compression token corresponding to the i-th text window.

[0227] Furthermore, after compressing the Nth text window, the computer device can obtain the Nth response text corresponding to the Nth text window output by the large language model, and use the Nth response text corresponding to the Nth text window as the initial response text. Simultaneously, during the generation of N compression tokens corresponding to the N text windows, the computer device also maps the N compression tokens one-to-one with the leaf nodes in the tree structure, thereby aggregating the N compression tokens based on the tree structure to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure.

[0228] Regarding the token compression process, in one possible implementation, when N is even, the computer device can aggregate the compressed tokens corresponding to adjacent leaf nodes pairwise. Optionally, the computer device aggregates the i-th compressed token with the (i+1)-th compressed token to obtain the aggregated compressed token corresponding to the (i+1) / 2-th parent node, where i is odd. In another possible implementation, when N is odd, the compressed tokens corresponding to the first N-1 leaf nodes can be aggregated pairwise, while the N-th leaf node has no aggregable adjacent nodes. In this case, the computer device can move the N-th leaf node up one level, determining the N-th compressed token as the aggregated compressed token corresponding to the (N+1) / 2-th parent node. Here, the (N+1) / 2-th parent node is the last node of the second-to-last level in the binary tree structure. Furthermore, when the number of nodes in the penultimate level of the binary tree structure is even, the computer device can aggregate the aggregated compressed token corresponding to the (N-1) / 2th parent node with the aggregated compressed token corresponding to the (N+1) / 2th parent node. When the number of nodes in the penultimate level of the binary tree structure is odd, the computer device continues to move the Nth leaf node up one level, determining the Nth compressed token as the aggregated compressed token corresponding to the (N+3) / 4th node in the third-to-last level. Based on these two aggregation methods, the computer device can aggregate the aggregated compressed tokens corresponding to each parent node layer by layer from bottom to top, until it obtains the aggregated compressed tokens corresponding to the parent nodes at each level of the binary tree structure and the aggregated compressed token corresponding to the root node.

[0229] Finally, the computer device, by calling the large language model again, can generate the target answer text corresponding to the question text based on N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text. This target answer text is then displayed on the question-and-answer interface of the AI ​​question-and-answer webpage. Compared to the initial answer text, the target answer text generated based on N compression tokens and the aggregated compression tokens corresponding to each non-leaf node in the tree structure takes into account global text features more fully, resulting in higher accuracy in answering the question text.

[0230] Please refer to Figure 10 This illustrates a structural block diagram of a text processing apparatus provided in an exemplary embodiment of this application, the apparatus comprising:

[0231] The text compression module 1001 is used to chain-compress N text windows using a large language model to obtain N compression tokens corresponding to the N text windows. The N text windows are obtained by dividing the input text. The input text includes question text and reference text. The reference text is the source of the answer text corresponding to the question text. The i-th compression token among the N compression tokens is used to characterize the text features of the i-th text window. N is a positive integer, and i is less than or equal to N.

[0232] The token aggregation module 1002 is used to aggregate the N compressed tokens based on a tree structure to obtain the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, wherein the N compressed tokens correspond one-to-one with the leaf nodes of the tree structure.

[0233] The model output module 1003 is used to generate a target answer text corresponding to the question text based on the N compression tokens, the aggregate compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text through the large language model. The initial answer text is the answer text output by the large language model after performing chain compression on the N text windows, and the target answer text is the answer text output by the large language model after performing text processing on the initial answer text.

[0234] Optionally, the tree structure is a binary tree structure; the token aggregation module 1002 is used for:

[0235] Based on the binary tree structure, when N is even, the i-th compressed token and the (i+1)-th compressed token are aggregated to obtain the aggregated compressed token corresponding to the (i+1) / 2-th parent node, where i is odd; when N is odd, the N-th compressed token is determined as the aggregated compressed token corresponding to the (N+1) / 2-th parent node.

[0236] By aggregating the aggregated compressed tokens corresponding to each parent node layer by layer, the aggregated compressed tokens corresponding to each parent node and the root node in the binary tree structure are obtained.

[0237] Optionally, the i-th compressed token includes the i-th compressed KV vector corresponding to the i-th text window;

[0238] The token aggregation module 1002 is used for:

[0239] Based on the key aggregation weights, the i-th compressed K-vector and the (i+1)-th compressed K-vector are weighted and summed to obtain the aggregated compressed K-vector corresponding to the (i+1) / 2-th parent node. The key aggregation weights are calculated based on the i-th compressed K-vector and the (i+1)-th compressed K-vector through cross-attention.

[0240] Based on the value aggregation weight, the i-th compressed V vector and the (i+1)-th compressed V vector are weighted and summed to obtain the aggregated compressed V vector corresponding to the (i+1) / 2-th parent node. The value aggregation weight is calculated based on the i-th compressed V vector and the (i+1)-th compressed V vector through cross attention.

[0241] Optionally, the model output module 1003 is used for:

[0242] The initial answer text of the question text is input into the large language model, and the large language model performs text preprocessing on the initial answer text to obtain the hidden layer representation of the text corresponding to the initial answer text.

[0243] Based on the QKV projection matrix corresponding to the attention mechanism in the large language model, the hidden layer representation of the text is projected to obtain the text QKV vector corresponding to the hidden layer representation of the text.

[0244] Based on the N compressed tokens, the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, and the text QKV vector, the target answer text corresponding to the question text is generated through the large language model.

[0245] Optionally, the aggregated compression token includes an aggregated compressed K vector and an aggregated compressed V vector;

[0246] The model output module 1003 is used for:

[0247] The N compressed KV vectors corresponding to the N text windows, the aggregated compressed KV vectors corresponding to each non-leaf node in the tree structure, and the text QKV vector are fused to obtain the text Q vector, text K vector, and text V vector used for cross-attention calculation.

[0248] Based on the text Q vector, the text K vector, and the text V vector, the attention output result corresponding to the text QKV vector is obtained through cross-attention calculation;

[0249] The large language model performs feedforward neural network processing, residual connections, and normalization on the attention output to obtain the target answer text corresponding to the question text.

[0250] Optionally, the text compression module 1001 includes:

[0251] The text preprocessing unit is used to perform text preprocessing on the i-th text window and the i-th initial compression token corresponding to the i-th text window through the large language model, so as to obtain the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compression token;

[0252] The vector projection unit is used to project the i-th hidden layer representation and the i-th compressed hidden layer representation based on the QKV projection matrix corresponding to the attention mechanism in the large language model to obtain the i-th initial QKV vector.

[0253] The vector compression unit is used to compress the i-th initial QKV vector through the large language model to obtain the i-th compressed token corresponding to the i-th text window.

[0254] Optionally, the vector projection unit is used for:

[0255] Based on the first QKV projection matrix corresponding to the attention mechanism in the large language model, the i-th hidden layer representation is projected to obtain the i-th initial text QKV vector;

[0256] Based on the second QKV projection matrix corresponding to the attention mechanism in the large language model, the i-th compressed hidden layer representation is projected to obtain the i-th initial compressed QKV vector.

[0257] The i-th initial text QKV vector is obtained based on the i-th initial compressed QKV vector, the i-th initial compressed QKV vector, and the i-1 compressed tokens corresponding to the i-1 text windows.

[0258] Optionally, the i-th compressed token includes the i-th compressed KV vector corresponding to the i-th text window;

[0259] The vector projection unit is also used for:

[0260] By fusing the i-th initial text Q-vector and the i-th initial compressed Q-vector, we obtain the i-th initial Q-vector;

[0261] The i-th initial text K vector, the i-th initial compressed K vector, and the i-1 compressed K vectors corresponding to the first i-1 text windows are fused to obtain the i-th initial K vector;

[0262] The i-th initial text V vector, the i-th initial compressed V vector, and the i-1 compressed V vectors corresponding to the first i-1 text windows are fused to obtain the i-th initial V vector.

[0263] Optionally, the question text is divided into a first text window, and the device further includes:

[0264] The text compression module 1001 is used to compress the first text window through the large language model to obtain the first text KV vector and the first compressed KV vector corresponding to the first text window;

[0265] The vector extraction module is used to extract the question text KV vector corresponding to the question text from the first text KV vector based on the text position encoding corresponding to the question text;

[0266] The vector projection unit is used for:

[0267] The initial K vector is obtained by fusing the question text K vector, the i-th initial text K vector, the i-th initial compressed K vector, and the i-1 compressed K vectors corresponding to the first i-1 text windows;

[0268] The vector projection unit is also used for:

[0269] The i-th initial V vector is obtained by fusing the problem text V vector, the i-th initial text V vector, the i-th initial compressed V vector, and the i-1 compressed V vectors corresponding to the first i-1 text windows.

[0270] Optionally, the vector compression unit is used for:

[0271] The large language model is used to perform cross-attention calculation on the i-th initial QKV vector to obtain the i-th attention output result corresponding to the i-th text window;

[0272] The large language model is used to perform feedforward neural network processing, residual connection, and normalization on the i-th attention output result to obtain the i-th answer text corresponding to the i-th text window;

[0273] During the process of outputting the i-th answer text, the i-th initial QKV vector is updated to obtain the updated i-th initial QKV vector;

[0274] Based on the updated i-th initial QKV vector, the i-th compressed KV vector corresponding to the i-th text window is determined.

[0275] Please refer to Figure 11 This illustrates a structural block diagram of a text processing apparatus provided in an exemplary embodiment of this application, the apparatus comprising:

[0276] The text compression module 1101 is used to perform chained compression of N sample text windows through a large language model to obtain N sample compression tokens corresponding to the N sample text windows. The N sample text windows are obtained by dividing the sample input text. The sample input text includes sample question text and sample reference text.

[0277] The token aggregation module 1102 is used to aggregate the N sample compressed tokens based on the sample tree structure to obtain the sample aggregated compressed tokens corresponding to each sample non-leaf node in the sample tree structure. The N sample compressed tokens correspond one-to-one with the sample leaf nodes of the sample tree structure.

[0278] The model output module 1103 is used to generate the sample answer text corresponding to the sample question text through the large language model based on the N sample compression tokens, the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure, and the sample initial answer text of the sample question text.

[0279] The loss determination module 1104 is used to determine the model prediction loss of the large language model based on the sample answer text corresponding to the sample question text and the truth value of the answer text.

[0280] The parameter update module 1105 is used to update the model parameters of the large language model based on the model prediction loss.

[0281] Optionally, the model parameters include the QKV projection matrix and the output projection matrix corresponding to the attention mechanism in the large language model;

[0282] The parameter update module 1105 is used for:

[0283] The QKV projection matrix and output projection matrix that meet the training completion conditions are determined as the target QKV projection matrix and target output projection matrix;

[0284] The large language model is updated based on the target QKV projection matrix and the target output projection matrix.

[0285] In summary, in this embodiment, the input text is divided into N text windows, and then the N text windows are chain-compressed using a large language model to obtain N compression tokens corresponding to the N text windows, thus avoiding the text semantic loss problem caused by directly compressing the input text. Furthermore, by mapping the N compression tokens one-to-one with the leaf nodes in the tree structure, the N compression tokens are aggregated to obtain the aggregated compression tokens corresponding to each non-leaf node in the tree structure. Based on the N compression tokens, the aggregated compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text corresponding to the question text, the target answer text corresponding to the question text is generated using the large language model. This enables the generation of the target answer text based on the global text features of the input text, thereby improving the accuracy of answering the question text.

[0286] It should be noted that the apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their implementation process can be found in the method embodiments, which will not be repeated here.

[0287] Please refer to Figure 12 This illustration shows a schematic diagram of a computer device provided in an exemplary embodiment of this application. Specifically, the computer device 1200 includes a Central Processing Unit (CPU) 1201, a system memory 1204 including a random access memory 1202 and a read-only memory 1203, and a system bus 1205 connecting the system memory 1204 and the CPU 1201. The computer device 1200 may also include a basic input / output system (I / O system) 1206 to facilitate information transfer between various devices within the computer, and a mass storage device 1207 for storing an operating system 1213, application programs 1214, and other program modules 1215.

[0288] In some embodiments, the basic input / output system 1206 includes a display 1208 for displaying information and an input device 1209 for user input, such as a mouse or keyboard. Both the display 1208 and the input device 1209 are connected to the central processing unit 1201 via an input / output controller 1210 connected to the system bus 1205. The basic input / output system 1206 may also include the input / output controller 1210 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1210 also provides output to a display screen, printer, or other types of output devices.

[0289] The mass storage device 1207 is connected to the central processing unit 1201 via a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer-readable media provide non-volatile storage for the computer device 1200. That is, the mass storage device 1207 may include computer-readable media (not shown) such as a hard disk or drive.

[0290] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage media are not limited to the above-mentioned types. The system memory 1204 and mass storage device 1207 described above can be collectively referred to as memory.

[0291] The memory stores one or more programs, which are configured to be executed by one or more central processing units 1201. The one or more programs contain instructions for implementing the methods described above. The central processing unit 1201 executes the one or more programs to implement the text processing methods provided in the various method embodiments described above.

[0292] According to various embodiments of this application, the computer device 1200 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1200 can be connected to the network 1211 via the network interface unit 1212 connected to the system bus 1205, or the network interface unit 1212 can be used to connect to other types of networks or remote computer systems (not shown).

[0293] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the text processing method described in the above embodiments.

[0294] Optionally, the computer-readable storage medium may include ROM, RAM, solid-state drives (SSDs), or optical discs, etc. The RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0295] This application provides a computer program product including at least one instruction stored in a computer-readable storage medium. A processor of a computer device reads the at least one instruction from the computer-readable storage medium and executes the at least one instruction, causing the computer device to perform the text processing method described in the above embodiments.

[0296] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0297] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.< / ct> < / ct> < / ct>

Claims

1. A text processing method characterized by, The method includes: By chaining and compressing N text windows using a large language model, N compression tokens are obtained corresponding to the N text windows. The N text windows are obtained by dividing the input text. The input text includes question text and reference text. The reference text is the source of the answer text corresponding to the question text. The i-th compression token in the N compression tokens is used to characterize the text features of the i-th text window. N is a positive integer, and i is less than or equal to N. The N compressed tokens are aggregated based on a tree structure to obtain the aggregated compressed tokens corresponding to each non-leaf node in the tree structure. The N compressed tokens correspond one-to-one with the leaf nodes of the tree structure. Based on the N compression tokens, the aggregate compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text, the target answer text corresponding to the question text is generated by the large language model. The initial answer text is the answer text output by the large language model after performing chained compression on the N text windows, and the target answer text is the answer text output by the large language model after performing text processing on the initial answer text.

2. The method according to claim 1, characterized in that, The tree structure is a binary tree structure; the aggregation of the N compressed tokens based on the tree structure to obtain the aggregated compressed token corresponding to each non-leaf node in the tree structure includes: Based on the binary tree structure, when N is even, the i-th compressed token and the (i+1)-th compressed token are aggregated to obtain the aggregated compressed token corresponding to the (i+1) / 2-th parent node, where i is odd; when N is odd, the N-th compressed token is determined as the aggregated compressed token corresponding to the (N+1) / 2-th parent node. By aggregating the aggregated compressed tokens corresponding to each parent node layer by layer, the aggregated compressed tokens corresponding to each parent node and the root node in the binary tree structure are obtained.

3. The method according to claim 2, characterized in that, The i-th compressed token includes the i-th compressed KV vector corresponding to the i-th text window; The aggregation of the i-th compressed token and the (i+1)-th compressed token to obtain the aggregated compressed token corresponding to the (i+1) / 2-th parent node includes: Based on the key aggregation weights, the i-th compressed K-vector and the (i+1)-th compressed K-vector are weighted and summed to obtain the aggregated compressed K-vector corresponding to the (i+1) / 2-th parent node. The key aggregation weights are calculated based on the i-th compressed K-vector and the (i+1)-th compressed K-vector through cross-attention. Based on the value aggregation weight, the i-th compressed V vector and the (i+1)-th compressed V vector are weighted and summed to obtain the aggregated compressed V vector corresponding to the (i+1) / 2-th parent node. The value aggregation weight is calculated based on the i-th compressed V vector and the (i+1)-th compressed V vector through cross attention.

4. The method according to any one of claims 1 to 3, characterized in that, The process of generating the target answer text corresponding to the question text using the large language model, based on the N compressed tokens, the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text, includes: The initial answer text of the question text is input into the large language model, and the large language model performs text preprocessing on the initial answer text to obtain the hidden layer representation of the text corresponding to the initial answer text. Based on the QKV projection matrix corresponding to the attention mechanism in the large language model, the hidden layer representation of the text is projected to obtain the text QKV vector corresponding to the hidden layer representation of the text. Based on the N compressed tokens, the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, and the text QKV vector, the target answer text corresponding to the question text is generated through the large language model.

5. The method according to claim 4, characterized in that, The aggregated compression token includes an aggregated compressed K vector and an aggregated compressed V vector; The process of generating the target answer text corresponding to the question text using the large language model, based on the N compressed tokens, the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, and the text QKV vector, includes: The N compressed KV vectors corresponding to the N text windows, the aggregated compressed KV vectors corresponding to each non-leaf node in the tree structure, and the text QKV vector are fused to obtain the text Q vector, text K vector, and text V vector used for cross-attention calculation. Based on the text Q vector, the text K vector, and the text V vector, the attention output result corresponding to the text QKV vector is obtained through cross-attention calculation; The large language model performs feedforward neural network processing, residual connections, and normalization on the attention output to obtain the target answer text corresponding to the question text.

6. The method according to any one of claims 1 to 5, characterized in that, The method of chaining and compressing N text windows using a large language model to obtain N compression tokens corresponding to the N text windows includes: By performing text preprocessing on the i-th text window and the i-th initial compression token corresponding to the i-th text window through the large language model, the i-th hidden layer representation corresponding to the i-th text window and the i-th compressed hidden layer representation corresponding to the i-th initial compression token are obtained; Based on the QKV projection matrix corresponding to the attention mechanism in the large language model, the i-th hidden layer representation and the i-th compressed hidden layer representation are projected to obtain the i-th initial QKV vector; The i-th initial QKV vector is compressed using the large language model to obtain the i-th compressed token corresponding to the i-th text window.

7. The method according to claim 6, characterized in that, The step of projecting the i-th hidden layer representation and the i-th compressed hidden layer representation onto the QKV projection matrix corresponding to the attention mechanism in the large language model to obtain the i-th initial QKV vector includes: Based on the first QKV projection matrix corresponding to the attention mechanism in the large language model, the i-th hidden layer representation is projected to obtain the i-th initial text QKV vector; Based on the second QKV projection matrix corresponding to the attention mechanism in the large language model, the i-th compressed hidden layer representation is projected to obtain the i-th initial compressed QKV vector. The i-th initial text QKV vector is obtained based on the i-th initial compressed QKV vector, the i-th initial compressed QKV vector, and the i-1 compressed tokens corresponding to the i-1 text windows.

8. The method according to claim 7, characterized in that, The i-th compressed token includes the i-th compressed KV vector corresponding to the i-th text window; The process of obtaining the i-th initial QKV vector based on the i-th initial text QKV vector, the i-th initial compressed QKV vector, and the i-1 compressed tokens corresponding to the first i-1 text windows includes: By fusing the i-th initial text Q-vector and the i-th initial compressed Q-vector, we obtain the i-th initial Q-vector; and, By fusing the i-th initial text K vector, the i-th initial compressed K vector, and the i-1 compressed K vectors corresponding to the first i-1 text windows, the i-th initial K vector is obtained; and, The i-th initial text V vector, the i-th initial compressed V vector, and the i-1 compressed V vectors corresponding to the first i-1 text windows are fused to obtain the i-th initial V vector.

9. The method according to claim 8, characterized in that, The question text is divided into the first text window, and the method further includes: The first text window is compressed using the large language model to obtain the first text key-value vector and the first compressed key-value vector corresponding to the first text window; Based on the text position encoding corresponding to the question text, extract the question text KV vector corresponding to the question text from the first text KV vector; The process of fusing the i-th initial text K vector, the i-th initial compressed K vector, and the i-1 compressed K vectors corresponding to the first i-1 text windows to obtain the i-th initial K vector includes: The initial K vector is obtained by fusing the question text K vector, the i-th initial text K vector, the i-th initial compressed K vector, and the i-1 compressed K vectors corresponding to the first i-1 text windows; The process of fusing the i-th initial text V vector, the i-th initial compressed V vector, and the i-1 compressed V vectors corresponding to the first i-1 text windows to obtain the i-th initial V vector includes: The i-th initial V vector is obtained by fusing the problem text V vector, the i-th initial text V vector, the i-th initial compressed V vector, and the i-1 compressed V vectors corresponding to the first i-1 text windows.

10. The method according to claim 8, characterized in that, The step of compressing the i-th initial QKV vector using the large language model to obtain the i-th compressed token corresponding to the i-th text window includes: The large language model is used to perform cross-attention calculation on the i-th initial QKV vector to obtain the i-th attention output result corresponding to the i-th text window; The large language model is used to perform feedforward neural network processing, residual connection, and normalization on the i-th attention output result to obtain the i-th answer text corresponding to the i-th text window; During the process of outputting the i-th answer text, the i-th initial QKV vector is updated to obtain the updated i-th initial QKV vector; Based on the updated i-th initial QKV vector, the i-th compressed KV vector corresponding to the i-th text window is determined.

11. A text processing method, characterized in that, The method includes: By chaining and compressing N sample text windows using a large language model, N sample compression tokens corresponding to the N sample text windows are obtained. The N sample text windows are obtained by dividing the sample input text, which includes sample question text and sample reference text. Based on the sample tree structure, the N sample compression tokens are aggregated to obtain the sample aggregate compression token corresponding to each non-leaf node of the sample in the sample tree structure. The N sample compression tokens correspond one-to-one with the sample leaf nodes of the sample tree structure. Based on the N sample compression tokens, the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure, and the initial sample answer text of the sample question text, the sample answer text corresponding to the sample question text is generated through the large language model. Based on the sample answer text corresponding to the sample question text and the ground truth value of the answer text, the model prediction loss of the large language model is determined; Based on the model prediction loss, the model parameters of the large language model are updated.

12. The method according to claim 11, characterized in that, The model parameters include the QKV projection matrix and the output projection matrix corresponding to the attention mechanism in the large language model; The step of updating the model parameters of the large language model based on the model prediction loss includes: The QKV projection matrix and output projection matrix that meet the training completion conditions are determined as the target QKV projection matrix and target output projection matrix; The large language model is updated based on the target QKV projection matrix and the target output projection matrix.

13. A text processing device, characterized in that, The device includes: The text compression module is used to chain-compress N text windows using a large language model to obtain N compression tokens corresponding to the N text windows. The N text windows are obtained by dividing the input text. The input text includes question text and reference text. The reference text is the source of the answer text corresponding to the question text. The i-th compression token among the N compression tokens is used to characterize the text features of the i-th text window. N is a positive integer, and i is less than or equal to N. The token aggregation module is used to aggregate the N compressed tokens based on a tree structure to obtain the aggregated compressed tokens corresponding to each non-leaf node in the tree structure, wherein the N compressed tokens correspond one-to-one with the leaf nodes of the tree structure. The model output module is used to generate the target answer text corresponding to the question text based on the N compression tokens, the aggregate compression tokens corresponding to each non-leaf node in the tree structure, and the initial answer text of the question text through the large language model. The initial answer text is the answer text output by the large language model after performing chain compression on the N text windows, and the target answer text is the answer text output by the large language model after performing text processing on the initial answer text.

14. A text processing device, characterized in that, The device includes: The text compression module is used to chain-compress N sample text windows using a large language model to obtain N sample compression tokens corresponding to the N sample text windows. The N sample text windows are obtained by dividing the sample input text, and the sample input text includes sample question text and sample reference text. The token aggregation module is used to aggregate the N sample compressed tokens based on the sample tree structure to obtain the sample aggregated compressed tokens corresponding to each sample non-leaf node in the sample tree structure. The N sample compressed tokens correspond one-to-one with the sample leaf nodes of the sample tree structure. The model output module is used to generate sample answer text corresponding to the sample question text through the large language model based on the N sample compression tokens, the sample aggregation compression tokens corresponding to the non-leaf nodes of each sample in the sample tree structure, and the sample initial answer text of the sample question text. The loss determination module is used to determine the model prediction loss of the large language model based on the sample answer text corresponding to the sample question text and the truth value of the answer text. The parameter update module is used to update the model parameters of the large language model based on the model prediction loss.

15. A computer device, characterized in that, The computer device includes a processor and a memory; the memory stores at least one instruction, which is executed by the processor to implement the text processing method as described in any one of claims 1 to 10, or the text processing method as described in any one of claims 11 to 12.

16. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is executed by a processor to implement the text processing method as described in any one of claims 1 to 10, or the text processing method as described in any one of claims 11 to 12.

17. A computer program product, characterized in that, The computer program product includes at least one instruction stored in a computer-readable storage medium; a processor of a computer device reads the at least one instruction from the computer-readable storage medium and executes the at least one instruction to cause the computer device to implement the text processing method as described in any one of claims 1 to 10, or the text processing method as described in any one of claims 11 to 12.