Text processing method and device, electronic equipment and readable storage medium

By multiplexing the intermediate result matrix of the previous cycle in the attention network, the problems of low text processing efficiency and waste of resources are solved, and more efficient text processing is achieved.

CN120337857APending Publication Date: 2025-07-18STREAM COMPUTING INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410073003.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, attention networks have problems such as inefficiency and waste of computing resources in text processing, because the same section of text is repeatedly performed on each processing cycle.

Method used

Repeated operations are reduced by saving the intermediate result matrix of a processing cycle on the attention network in memory and multiplexing these matrices in the next processing cycle. The specific method includes determining a vector of the incremental portion of the target input text and a matrix of repeated text, splicing and storing it as a new intermediate result matrix for operation in the next cycle.

Benefits of technology

It effectively reduces matrix operations, improves text processing efficiency, reduces waste of computing resources, and achieves more efficient text processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337857A_ABST
    Figure CN120337857A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text processing method and device, electronic equipment and a readable storage medium, and relates to the technical field of computers. In the embodiment of the invention, the attention network can receive the target input text of the processing period, and takes the first matrix stored in the memory as the multiplexing result to participate in the processing of the attention network so as to determine the second matrix of the processing period. Furthermore, according to the embodiment of the invention, the output text of the processing period can be determined according to the second matrix, and the second matrix is stored for reuse of the next processing period. Therefore, through the embodiment of the invention, the attention network can reuse the intermediate result corresponding to the repeated text in each processing period, so that the matrix operation is effectively reduced, the text processing efficiency is improved, and the waste of operation resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a text processing method, apparatus, electronic device, and readable storage medium. Background Art

[0002] In the process of an electronic device performing natural language processing (NLP), it is often necessary to perform text generation and other processing on text through an attention network. Among them, the attention network usually gradually determines the output text in a cyclic iteration manner.

[0003] In the related art, the attention network concatenates the input text and the output text of the previous processing cycle, and performs arithmetic processing on the concatenated text. During this process, as the arithmetic cycle is continuously iterated, the attention network repeatedly performs arithmetic processing on the same piece of text, which not only results in low text processing efficiency but also causes waste of arithmetic resources. Summary of the Invention

[0004] In view of this, embodiments of this application provide a text processing method, apparatus, electronic device, and readable storage medium, which can effectively reduce the matrix operations of the attention network, improve the text processing efficiency, and reduce the waste of arithmetic resources.

[0005] In a first aspect, a text processing method is provided. The method includes:

[0006] Receiving, by an attention network, a target input text of the current processing cycle, where the target input text includes a repeated text and an incremental text, the repeated text is the input text of the attention network in the previous processing cycle, and the incremental text is the output text of the attention network in the previous processing cycle.

[0007] Determining a target vector corresponding to the incremental text.

[0008] Determining a first matrix corresponding to the repeated text in a memory.

[0009] Determining and storing a second matrix of the current processing cycle according to the target vector and the first matrix.

[0010] Determining an output text of the current processing cycle according to the second matrix.

[0011] In some embodiments, the determining and storing a second matrix of the current processing cycle according to the target vector and the first matrix includes:

[0012] Determining intermediate results of the target vector in the current processing cycle according to the target vector and a plurality of pre-set weight matrices.

[0013] Determine the second matrix of the present processing cycle according to the splicing result of each of the intermediate results and the first matrix.

[0014] Store the second matrix in the memory.

[0015] In some embodiments, the first matrix at least includes a first Query matrix, a first transposed Key matrix, a first Value matrix, and a first score matrix determined by the attention network in the previous processing cycle, and the intermediate results include a second Query matrix, a second Key matrix, and a second Value matrix of the target vector in the present processing cycle.

[0016] The determining the second matrix of the present processing cycle according to the splicing result of each of the intermediate results and the first matrix includes:

[0017] Splice the first Query matrix and the second Query matrix to determine a third Query matrix of the target input text in the present processing cycle, and splice the first Value matrix and the second Value matrix to determine a third Value matrix of the target input text in the present processing cycle.

[0018] Splice the second transposed Key matrix corresponding to the second Key matrix and the first transposed Key matrix to determine a third transposed Key matrix of the target input text in the present processing cycle.

[0019] Determine a third score matrix of the target input text in the present processing cycle according to the third Query matrix, the third transposed Key matrix, and the first score matrix.

[0020] Determine the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix of the present processing cycle.

[0021] In some embodiments, the third score matrix is obtained by performing a matrix multiplication operation on the third Query matrix and the third transposed Key matrix.

[0022] The determining the third score matrix of the target input text in the present processing cycle according to the third Query matrix, the third transposed Key matrix, and the first score matrix includes:

[0023] Use the first score matrix as a partial operation result in the matrix multiplication operation of the third Query matrix and the third transposed Key matrix to determine the third score matrix of the target input text in the present processing cycle.

[0024] In some embodiments, storing the second matrix in the memory includes:

[0025] Storing the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix of the current processing cycle in the memory.

[0026] In some embodiments, the method further includes:

[0027] Concatenating the input text and the output text of the attention network in the previous processing cycle to determine the concatenated text.

[0028] Padding the concatenated text according to a preset text length threshold to determine the target input text of the current processing cycle.

[0029] In some embodiments, determining the output text of the current processing cycle according to the second matrix includes:

[0030] Performing an activation function process on the third score matrix, and performing a matrix multiplication operation on the result of the activation function process and the third Value matrix to determine the context matrix.

[0031] Determining the output text of the current processing cycle according to the context matrix.

[0032] In a second aspect, a text processing device is provided, and the device includes:

[0033] A target input text receiving module, configured to receive the target input text of the current processing cycle through an attention network, where the target input text includes duplicate text and incremental text, the duplicate text is the input text of the attention network in the previous processing cycle, and the incremental text is the output text of the attention network in the previous processing cycle.

[0034] A target vector determining module, configured to determine the target vector corresponding to the incremental text.

[0035] A first matrix determining module, configured to determine the first matrix corresponding to the duplicate text in the memory.

[0036] A second matrix determining module, configured to determine and store the second matrix of the current processing cycle according to the target vector and the first matrix.

[0037] An output text determining module, configured to determine the output text of the current processing cycle according to the second matrix.

[0038] In some embodiments, the second matrix determining module is specifically configured to perform:

[0039] Determine intermediate results of the target vector in the present processing cycle according to the target vector and multiple preset weight matrices.

[0040] Determine the second matrix of the present processing cycle according to the splicing result of each intermediate result and the first matrix.

[0041] Store the second matrix in the memory.

[0042] In some embodiments, the first matrix at least includes a first Query matrix, a first transposed Key matrix, a first Value matrix, and a first score matrix determined by the attention network in the previous processing cycle, and the intermediate results include a second Query matrix, a second Key matrix, and a second Value matrix of the target vector in the present processing cycle.

[0043] The second matrix determination module is specifically configured to execute:

[0044] Splice the first Query matrix and the second Query matrix to determine a third Query matrix of the target input text in the present processing cycle, and splice the first Value matrix and the second Value matrix to determine a third Value matrix of the target input text in the present processing cycle.

[0045] Splice the second transposed Key matrix corresponding to the second Key matrix and the first transposed Key matrix to determine a third transposed Key matrix of the target input text in the present processing cycle.

[0046] Determine a third score matrix of the target input text in the present processing cycle according to the third Query matrix, the third transposed Key matrix, and the first score matrix.

[0047] Determine the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix of the present processing cycle.

[0048] In some embodiments, the third score matrix is obtained by performing a matrix multiplication operation on the third Query matrix and the third transposed Key matrix.

[0049] The second matrix determination module is specifically configured to execute:

[0050] Use the first score matrix as a partial operation result in the matrix multiplication operation of the third Query matrix and the third transposed Key matrix to determine a third score matrix of the target input text in the present processing cycle.

[0051] In some embodiments, the second matrix determination module is specifically configured to perform:

[0052] Store the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix of this processing cycle in the memory.

[0053] In some embodiments, the apparatus further includes:

[0054] A splicing module, configured to perform splicing processing on the input text and the output text of the attention network in the previous processing cycle to determine the spliced text.

[0055] A padding module, configured to perform padding processing on the spliced text according to a preset text length threshold to determine the target input text of this processing cycle.

[0056] In some embodiments, the output text determination module is specifically configured to perform:

[0057] Perform an activation function process on the third score matrix, and perform a matrix multiplication operation on the result of the activation function process and the third Value matrix to determine the context matrix.

[0058] Determine the output text of this processing cycle according to the context matrix.

[0059] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect.

[0060] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions implement the method as described in the first aspect when executed by a processor.

[0061] In the embodiment of the present application, the attention network can receive the target input text of this processing cycle, and use the first matrix saved in the memory as a reuse result to participate in the processing of the attention network to determine the second matrix of this processing cycle. Further, the embodiment of the present application can determine the output text of this processing cycle according to the second matrix, and store the second matrix for reuse in the next processing cycle. Therefore, through the embodiment of the present application, the attention network can reuse the intermediate results corresponding to the repeated text in each processing cycle, effectively reducing matrix operations, improving the efficiency of text processing, and reducing the waste of computing resources. Description of the Drawings

[0062] Through the following description of the embodiments of the present application with reference to the accompanying drawings, the above and other objects, features, and advantages of the embodiments of the present application will become clearer. In the drawings:

[0063] Figure 1 is a flowchart of the text processing method according to the embodiment of the present application;

[0064] Figure 2 is a flowchart of determining the target input text according to the embodiment of the present application;

[0065] Figure 3 is a flowchart of determining the second matrix according to the embodiment of the present application;

[0066] Figure 4 is a schematic diagram of the matrix splicing result according to the embodiment of the present application;

[0067] Figure 5 is a schematic diagram of the working principle of the Attention Layer;

[0068] Figure 6 is another flowchart of determining the second matrix according to the embodiment of the present application;

[0069] Figure 7 is a schematic diagram of the third transposed Key matrix according to the embodiment of the present application;

[0070] Figure 8 is a schematic diagram of the third score matrix according to the embodiment of the present application;

[0071] Figure 9 is a flowchart of determining the output text according to the embodiment of the present application;

[0072] Figure 10 is a schematic diagram of the structure of the text processing device according to the embodiment of the present application;

[0073] Figure 11 is a schematic diagram of the structure of the electronic device according to the embodiment of the present application. Detailed Embodiments

[0074] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0075] In addition, those of ordinary skill in the art should understand that the accompanying drawings provided here are all for illustrative purposes and are not necessarily drawn to scale.

[0076] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document shall be construed in an inclusive sense rather than an exclusive or exhaustive sense; that is, it means "including but not limited to".

[0077] In the description of this application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0078] In the process of a natural language processing (NLP) by an electronic device, it is often necessary to perform text generation and other processing on the text through an attention network. Among them, the attention network usually gradually determines the output text in a cyclic iterative manner. For example, in a model with a structure such as a transformer, multiple attention layers are usually set to perform text generation and other processing.

[0079] In the related art, in each processing cycle, the attention network concatenates the input text and the output text of the previous processing cycle and performs arithmetic processing on the concatenated text. For example, if the input text of the attention network in processing cycle 1 is A and the output text in processing cycle 1 is B, then in processing cycle 2, the input text of the attention network is A + B. Further, the attention network performs arithmetic processing on the above A + B to determine the output text C of processing cycle 2.

[0080] Further, in processing cycle 3, the input text of the attention network is the concatenation result of the input text and the output text of processing cycle 2 (that is, A + B + C). At this time, the attention network needs to perform arithmetic processing on A + B + C to determine the output text of processing cycle 3.

[0081] It can be seen that in the above processing cycles 1 - 3, the attention network performs 3 arithmetic processes on part A and 2 arithmetic processes on part B. That is to say, as the arithmetic cycle continues to iterate, the attention network will repeatedly perform arithmetic processing on the same piece of text. This will not only result in low text processing efficiency but also cause waste of computing resources.

[0082] Therefore, how to effectively improve the efficiency of text processing and reduce the waste of computing resources is an urgent problem to be solved at present.

[0083] To solve the above problems, an embodiment of the present application provides a text processing method, which can be applied to an electronic device. The electronic device can be a terminal or a server. Among them, the terminal can be a smart phone, a tablet computer, or a personal computer (PC), etc., and the server can be a single server, a server cluster configured in a distributed manner, or a cloud server.

[0084] As Figure 1 shown, the text processing method of the embodiment of the present application may include the following steps:

[0085] In step S100, receive the target input text of the current processing cycle through the attention network.

[0086] Among them, the target input text of the embodiment of the present application includes repeated text and incremental text. The repeated text is the input text of the attention network in the previous processing cycle, and the incremental text is the output text of the attention network in the previous processing cycle.

[0087] In an optional implementation manner, the embodiment of the present application may first determine the target input text of the current processing cycle by means of splicing and padding, and then receive the target input text through the attention network.

[0088] Specifically, as Figure 2 shown, the process of determining the target input text may include the following steps:

[0089] In step S600, perform splicing processing on the input text and output text of the attention network in the previous processing cycle to determine the spliced text.

[0090] In step S700, according to the preset text length threshold, perform padding processing on the spliced text to determine the target input text of the current processing cycle.

[0091] Among them, the text length threshold is used to represent the maximum length of the text that the attention network can process. Specifically, when the attention network performs text processing, the maximum length of the text that it can process at one time is often fixed, and moreover, the model needs to process the text according to this maximum length when performing text processing. Therefore, the embodiment of the present application can fill the input text to the above maximum length according to the preset text length threshold, so that the attention network can process the target input text.

[0092] In step S200, determine the target vector corresponding to the incremental text.

[0093] Among them, since the incremental text is the output text of the attention network in the previous processing cycle, the incremental text is also the text that has not been processed by the attention network. Furthermore, the embodiments of the present application need to determine the target vector corresponding to the incremental text so that the attention network can process the incremental text.

[0094] In step S300, determine the first matrix corresponding to the repeated text in the memory.

[0095] Among them, the first matrix can be used to represent the intermediate result determined by the attention network for processing the repeated text. After determining the first matrix in the previous processing cycle, the attention network can store the first matrix in the memory for subsequent reuse. In addition, the first matrix can include one or more matrix data.

[0096] That is to say, since the repeated text in the embodiments of the present application is the text content that has been processed by the attention network in the previous processing cycle, the embodiments of the present application can store the intermediate result obtained for the processed text content in the memory so that the first matrix can be reused in this processing cycle, improving the efficiency of text processing and reducing the waste of computing resources.

[0097] In step S400, determine and store the second matrix of this processing cycle according to the target vector and the first matrix.

[0098] Among them, the second matrix is the intermediate result determined by the attention network according to the target input text of this processing cycle. That is to say, after the second matrix determined in this processing cycle is stored in the memory, it can be used as the first matrix in the next processing cycle.

[0099] In an alternative embodiment, as Figure 3 shown, step S400 may include the following steps:

[0100] In step S410, determine the intermediate results of the target vector in this processing cycle according to the target vector and multiple preset weight matrices.

[0101] In step S420, determine the second matrix of this processing cycle according to the splicing result of the intermediate results and the first matrix.

[0102] In step S430, store the second matrix in the memory.

[0103] For example, as Figure 4 shown, Figure 4 is a schematic diagram of the splicing result of the target vector and the first matrix. Among them, the splicing result can include the first matrix 41, the target vector 42, and the padding part 43.

[0104] By concatenating the first matrix 41 and the target vector 42, the embodiments of the present application can avoid partial matrix operations in the attention network, thereby improving the efficiency of text processing and reducing the waste of computing resources.

[0105] Taking the attention network as Attention Layer as an example, as Figure 5 shown, after receiving the input text, Attention Layer can multiply the matrix corresponding to the input text by 3 different weight matrices respectively to determine the Query matrix, the Key matrix, and the Value matrix. Among them, the matrix corresponding to the input text can be represented as [Batch, seq_len, Hidden_size]. Batch is used to represent the number of statements processed by a single processing core at a time (this number of statements is a non-negative value, and this value can be less than 1). seq_len is used to represent the statement length, and Hidden_size is used to represent the hidden size. The weight matrix can be represented as [Hidden_size, Hidden_size]. By performing matrix multiplication on the matrix corresponding to the input text and the weight matrices with different parameters, Attention Layer can determine the Query matrix, the Key matrix, and the Value matrix. The Query matrix, the Key matrix, and the Value matrix can be represented as [Batch, seq_len, Hidden_size]. Further, Attention Layer can perform recombination processing on the Query matrix, the Key matrix, and the Value matrix. After recombination, the Query matrix, the Key matrix, and the Value matrix can be represented as [Batch, seq_len, head, Hidden_size / head].

[0106] Further, Attention Layer can transpose the Query matrix by (0, 2, 1, 3). The transposed Query matrix can be represented as [Batch, head, seq_len, Hidden_size / head]. At the same time, Attention Layer can transpose the Key matrix by (0, 2, 3, 1). The transposed Key matrix can be represented as [Batch, head, Hidden_size / head, seq_len]. Further, Attention Layer can perform matrix multiplication on the above 2 transposed results to determine the score matrix [Batch, head, seq_len, seq_len].

[0107] Further, after determining the score matrix, the Attention Layer can perform an overlay operation on the score matrix and the Mask matrix, so that the valid part in the score matrix remains unchanged, and the invalid part in the score matrix becomes a minimum value. Further, the Attention Layer can perform a Softmax calculation on the result of the above overlay operation in the lowest dimension to determine the Softmax result matrix [Batch, head, seq_len, seq_len]. In the Softmax result matrix, the values corresponding to the above invalid parts are 0.

[0108] Further, the Attention Layer can transpose the Value matrix by (0, 2, 1, 3). The transposed Value matrix can be represented as [Batch, head, seq_len, Hidden_size / head]. Among them, the transposed result of the Value matrix is the right matrix, and the above Softmax result matrix is the left matrix. Furthermore, the Attention Layer can perform a matrix multiplication operation on the transposed result of the Value matrix and the Softmax result matrix to determine the operation result [Batch, head, seq_len, Hidden_size / head]. Further, the Attention Layer can perform a (0, 2, 1, 3) transpose and recombination process on the operation result to determine the Context matrix [Batch, seq_len, hidden_size].

[0109] Further, the Attention Layer can sequentially determine the local summation matrix, the local output matrix, the first feature matrix, and the second feature matrix through the Context matrix and the input text, and finally determine the output feature. Among them, the output feature corresponds to the output text of this processing cycle.

[0110] Combined Figure 4 and Figure 5 As shown in the content, the embodiment of the present application can splice the first matrix and the target vector, replace the partial matrix operation result in the attention network with the splicing result, and store the second matrix determined in this processing cycle for reuse in the next processing cycle.

[0111] Taking Figure 5 the Attention Layer shown as an example, the embodiment of the present application can replace the Query matrix, Key matrix, Value matrix, transposed Key matrix, and score matrix determined by the Attention Layer with the above splicing result, saving the operation process of the Attention Layer to determine the above matrices, improving the efficiency of text processing and reducing the waste of computing resources.

[0112] In an alternative embodiment, the first matrix of the embodiments of the present application at least includes a first Query matrix, a first transposed Key matrix, a first Value matrix, and a first score matrix determined by the attention network in the previous processing cycle. The intermediate result may include a second Query matrix, a second Key matrix, and a second Value matrix of the target vector in the current processing cycle.

[0113] In the current processing cycle of the embodiments of the present application, according to multiple pre-set weight matrices, the second Query matrix, the second Key matrix, and the second Value matrix of the target vector in the current processing cycle can be determined, and these matrices are used as intermediate results. Further, the embodiments of the present application can use the first Query matrix, the first transposed Key matrix, the first Value matrix, and the first score matrix to replace the matrix operation results, thereby improving the efficiency of text processing and reducing the waste of computing resources. That is to say, in the related art, when determining the Query matrix, the Key matrix, and the Value matrix, it is necessary to calculate for the entire input text, while the embodiments of the present application only need to determine the target vector (i.e., the part corresponding to the incremental text) for calculation, which can effectively reduce the computing amount of the electronic device, and further improve the efficiency of text processing and reduce the waste of computing resources.

[0114] Further, as Figure 6 shown, the above step S420 may include the following steps:

[0115] In step S421, the first Query matrix and the second Query matrix are concatenated to determine the third Query matrix of the target input text in the current processing cycle, and the first Value matrix and the second Value matrix are concatenated to determine the third Value matrix of the target input text in the current processing cycle.

[0116] Among them, through the concatenation processing method, the embodiments of the present application can avoid matrix operations on the Query matrix and the Value matrix, effectively reduce the computing amount of the electronic device, and further improve the efficiency of text processing and reduce the waste of computing resources.

[0117] In step S422, the second transposed Key matrix corresponding to the second Key matrix is concatenated with the first transposed Key matrix to determine the third transposed Key matrix of the target input text in the current processing cycle.

[0118] Among them, the embodiments of the present application can first determine the second transposed Key matrix corresponding to the second Key matrix (i.e., the transposed Key matrix corresponding to the incremental text), and then concatenate the second transposed Key matrix corresponding to the second Key matrix with the first transposed Key matrix to determine the third transposed Key matrix of the target input text in the current processing cycle.

[0119] During this process, since the second transposed Key matrix is the transposed Key matrix determined for the incremental text and the first transposed Key matrix is a pre-stored intermediate result, the embodiments of the present application can skip the process of determining the Key matrix of the target input text and directly determine the third transposed Key matrix of the target input text, further reducing the computational amount of the electronic device, improving the efficiency of text processing, and reducing the waste of computing resources.

[0120] In step S423, according to the third Query matrix, the third transposed Key matrix, and the first score matrix, determine the third score matrix of the target input text in the current processing cycle.

[0121] Among them, the embodiments of the present application can directly use the first score matrix as the multiplexing result to participate in the matrix operation to reduce the computational amount of determining the third score matrix.

[0122] In an alternative embodiment, the third score matrix of the embodiments of the present application can be obtained by performing matrix multiplication on the third Query matrix and the third transposed Key matrix. Further, step S414 above can be specifically executed as: using the first score matrix as a partial operation result in the process of matrix multiplication of the third Query matrix and the third transposed Key matrix to determine the third score matrix of the target input text in the current processing cycle.

[0123] Specifically, as Figure 7 shown, Figure 7 is a schematic diagram of the third transposed Key matrix. Among them, the third transposed Key matrix can include the first transposed Key matrix 71 (hereinafter referred to as matrix 71 for short), the second transposed Key matrix 72 (hereinafter referred to as matrix 72 for short), and the padding part 73. At the same time, the above Figure 4 can represent the matrix obtained by transposing the third Query matrix. The first matrix 41 is hereinafter referred to as matrix 41 for short, and the target vector 42 is hereinafter referred to as matrix 42 for short.

[0124] Further, when the embodiments of the present application determine the third score matrix, the matrix multiplication result of matrix 41 and matrix 71 can be replaced by the first score matrix. Specifically, the process of determining the third score matrix can be represented by the following formula:

[0125]

[0126] Among them, score is used to represent the third score matrix. Since matrix 41 and matrix 71 are reused parts (that is, matrix 41 and matrix 71 are the score matrices determined in the previous processing cycle), therefore, in the embodiments of the present application, the first score matrix can be used as part of the operation result to replace the matrix multiplication result of matrix 41 and matrix 71, so as to reduce the amount of operations for determining the third score matrix, improve the efficiency of text processing, and reduce the waste of computing resources.

[0127] As Figure 8 shown, Figure 8 is a schematic diagram of the third score matrix. As Figure 8 can be seen, the third score matrix determined in the embodiments of the present application can be represented in the lowest 2 dimensions (that is, the dimensions corresponding to the text length threshold), where the text length threshold is the maximum length of the text that the attention network can process at one time. That is to say, the embodiments of the present application can control the third score matrix in the lowest 2 dimensions to reduce the amount of operations such as subsequent matrix operations and normalization, and improve the efficiency of text processing.

[0128] In step S424, the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix are determined as the second matrix of this processing cycle.

[0129] In an alternative embodiment, the embodiments of the present application can also store the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix of this processing cycle in the memory.

[0130] Among them, after determining the second matrix, the embodiments of the present application can store the second matrix in the memory, and the second matrix stored in the memory can be used as the first matrix for the next processing. That is to say, the embodiments of the present application can update the matrix stored in the memory in each processing cycle, so that the attention network can reduce matrix operations as much as possible in each processing cycle, improve the efficiency of text processing, and reduce the waste of computing resources.

[0131] In step S500, according to the second matrix, the output text of this processing cycle is determined.

[0132] Taking Figure 4 as an example, after determining the second matrix, the embodiments of the present application can further determine the Context matrix, the local sum matrix, the local output matrix, the first feature matrix, and the second feature matrix according to the second matrix, and finally determine the output feature. Among them, the output feature corresponds to the output text of this processing cycle. Further, the output text of this processing cycle is also the incremental text in the input text of the next processing cycle.

[0133] In an embodiment of the present application, the attention network may receive the target input text of the current processing cycle, and use the first matrix saved in the memory as the reuse result to participate in the processing of the attention network to determine the second matrix of the current processing cycle. Further, the embodiment of the present application may determine the output text of the current processing cycle according to the second matrix, and store the second matrix for reuse in the next processing cycle. Therefore, through the embodiment of the present application, the attention network can reuse the intermediate results corresponding to the repeated text in each processing cycle, effectively reducing matrix operations, improving the efficiency of text processing, and reducing the waste of computing resources.

[0134] In an alternative embodiment, as Figure 9 shown, the above step S500 may include the following steps:

[0135] In step S510, perform an activation function processing on the third score matrix, and perform a matrix multiplication operation on the result of the activation function processing and the third Value matrix to determine the context matrix.

[0136] Wherein, the activation function may be Figure 4 the Softmax function shown in

[0137] or other applicable activation functions.

[0138] In step S520, determine the output text of the current processing cycle according to the context matrix.

[0139] After determining the Context matrix, the embodiment of the present application may sequentially determine the local summation matrix, the local output matrix, the first feature matrix, and the second feature matrix through the Context matrix and the target input text, and finally determine the output feature. Further, the embodiment of the present application may determine the output text of the current processing cycle according to the output feature.

[0139] Further, the embodiment of the present application may splice the output text of the current processing cycle and the target input text of the current processing cycle to determine the target input text of the next processing cycle, thereby realizing the step-by-step determination of the output text in a cyclic iteration manner. Further, when the attention network detects the termination character, the attention network will terminate the cyclic iteration and output the final text according to the output text of each processing cycle.

[0140] In this process, since the attention network can reuse the intermediate results corresponding to the repeated text in each processing cycle, the embodiment of the present application can effectively reduce matrix operations, improve the efficiency of text processing, and reduce the waste of computing resources.

[0141] Based on the same technical concept, the embodiment of the present application also provides a text processing device, as Figure 10As shown, the device includes: a target input text receiving module 1001, a target vector determining module 1002, a first matrix determining module 1003, a second matrix determining module 1004, and an output text determining module 1005.

[0142] The target input text receiving module 1001 is configured to receive the target input text of the current processing cycle through an attention network, where the target input text includes a repeated text and an incremental text, the repeated text is the input text of the attention network in the previous processing cycle, and the incremental text is the output text of the attention network in the previous processing cycle.

[0143] The target vector determining module 1002 is configured to determine the target vector corresponding to the incremental text.

[0144] The first matrix determining module 1003 is configured to determine the first matrix corresponding to the repeated text in the memory.

[0145] The second matrix determining module 1004 is configured to determine and store the second matrix of the current processing cycle according to the target vector and the first matrix.

[0146] The output text determining module 1005 is configured to determine the output text of the current processing cycle according to the second matrix.

[0147] In some embodiments, the second matrix determining module 1004 is specifically configured to perform:

[0148] Determine the intermediate results of the target vector in the current processing cycle according to the target vector and a plurality of pre-set weight matrices.

[0149] Determine the second matrix of the current processing cycle according to the splicing result of each intermediate result and the first matrix.

[0150] Store the second matrix in the memory.

[0151] In some embodiments, the first matrix at least includes a first Query matrix, a first transposed Key matrix, a first Value matrix, and a first score matrix determined by the attention network in the previous processing cycle, and the intermediate results include a second Query matrix, a second Key matrix, and a second Value matrix of the target vector in the current processing cycle.

[0152] The second matrix determining module 1004 is specifically configured to perform:

[0153] Concatenate the first Query matrix and the second Query matrix to determine the third Query matrix of the target input text in the current processing cycle, and concatenate the first Value matrix and the second Value matrix to determine the third Value matrix of the target input text in the current processing cycle.

[0154] Concatenate the second transposed Key matrix corresponding to the second Key matrix with the first transposed Key matrix to determine the third transposed Key matrix of the target input text in the current processing cycle.

[0155] Determine the third score matrix of the target input text in the current processing cycle according to the third Query matrix, the third transposed Key matrix, and the first score matrix.

[0156] Determine the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix in the current processing cycle.

[0157] In some embodiments, the third score matrix is obtained by performing matrix multiplication on the third Query matrix and the third transposed Key matrix.

[0158] The second matrix determination module 1004 is specifically configured to perform:

[0159] Use the first score matrix as a partial operation result in the matrix multiplication process of the third Query matrix and the third transposed Key matrix to determine the third score matrix of the target input text in the current processing cycle.

[0160] In some embodiments, the second matrix determination module 1004 is specifically configured to perform:

[0161] Store the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix in the current processing cycle into the memory.

[0162] In some embodiments, the apparatus further includes:

[0163] A concatenation module, configured to perform concatenation processing on the input text and the output text of the attention network in the previous processing cycle to determine the concatenated text.

[0164] A padding module, configured to perform padding processing on the concatenated text according to a preset text length threshold to determine the target input text in the current processing cycle.

[0165] In some embodiments, the output text determination module 1005 is specifically configured to perform:

[0166] Perform an activation function process on the third fractional matrix, and perform a matrix multiplication operation on the result of the activation function process and the third Value matrix to determine the context matrix.

[0167] Determine the output text of the current processing cycle according to the context matrix.

[0168] In the embodiments of the present application, the attention network can receive the target input text of the current processing cycle, and use the first matrix saved in the memory as the reuse result to participate in the processing of the attention network to determine the second matrix of the current processing cycle. Further, the embodiments of the present application can determine the output text of the current processing cycle according to the second matrix and store the second matrix for reuse in the next processing cycle. Therefore, through the embodiments of the present application, the attention network can reuse the intermediate results corresponding to the repeated text in each processing cycle, effectively reducing matrix operations, improving the efficiency of text processing, and reducing the waste of computing resources.

[0169] Figure 11 is a schematic diagram of the electronic device according to the embodiment of the present application. As Figure 11 shown, Figure 11 The electronic device shown is a general address query device, which includes a general computer hardware structure, and at least includes a processor 1101 and a memory 1102. The processor 1101 and the memory 1102 are connected through a bus 1103. The memory 1102 is adapted to store instructions or programs executable by the processor 1101. The processor 1101 can be an independent microprocessor or a set of one or more microprocessors. Thus, the processor 1101 executes the instructions stored in the memory 1102 to execute the method flow of the embodiment of the present application as described above to implement the processing of data and the control of other devices. The bus 1103 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 1104, a display device, and an input / output (I / O) device 1105. The input / output (I / O) device 1105 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices well known in the art. Typically, the input / output device 1105 is connected to the system through an input / output (I / O) controller 1106.

[0170] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device (equipment), or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0171] This application is described with reference to the flowcharts of methods, apparatuses (devices), and computer program products according to embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.

[0172] These computer program instructions can be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the process Figure 1 specified functions in one or more of these processes.

[0173] These computer program instructions can also be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the Figure 1 specified functions in one or more of these processes.

[0174] Another embodiment of the present application relates to a non-volatile storage medium for storing a computer-readable program, and the computer-readable program is used for a computer to execute some or all of the above method embodiments.

[0175] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by specifying relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0176] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A text processing method, characterized in that, The method includes: Receiving, by an attention network, a target input text of the current processing cycle, where the target input text includes a repeated text and an incremental text, the repeated text being the input text of the attention network in the previous processing cycle, and the incremental text being the output text of the attention network in the previous processing cycle; Determining a target vector corresponding to the incremental text; Determining, in a memory, a first matrix corresponding to the repeated text; Determining and storing a second matrix of the current processing cycle according to the target vector and the first matrix; Determining an output text of the current processing cycle according to the second matrix.

2. The method according to claim 1, wherein The determining and storing a second matrix of the current processing cycle according to the target vector and the first matrix includes: Determining intermediate results of the target vector in the current processing cycle according to the target vector and a plurality of preset weight matrices; Determining the second matrix of the current processing cycle according to a concatenation result of each of the intermediate results and the first matrix; Storing the second matrix into the memory.

3. The method according to claim 2, wherein The first matrix at least includes a first Query matrix, a first transposed Key matrix, a first Value matrix, and a first score matrix determined by the attention network in the previous processing cycle, and the intermediate results include a second Query matrix, a second Key matrix, and a second Value matrix of the target vector in the current processing cycle; The determining the second matrix of the current processing cycle according to a concatenation result of each of the intermediate results and the first matrix includes: Concatenating the first Query matrix and the second Query matrix to determine a third Query matrix of the target input text in the current processing cycle, and concatenating the first Value matrix and the second Value matrix to determine a third Value matrix of the target input text in the current processing cycle; Concatenating a second transposed Key matrix corresponding to the second Key matrix with the first transposed Key matrix to determine a third transposed Key matrix of the target input text in the current processing cycle; Determining a third score matrix of the target input text in the current processing cycle according to the third Query matrix, the third transposed Key matrix, and the first score matrix; Determining the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix of the current processing cycle.

4. The method according to claim 3, wherein The third score matrix is obtained by performing a matrix multiplication operation on the third Query matrix and the third transposed Key matrix; The determining a third score matrix of the target input text in the current processing cycle according to the third Query matrix, the third transposed Key matrix, and the first score matrix includes: Using the first score matrix as a partial operation result in the matrix multiplication operation of the third Query matrix and the third transposed Key matrix to determine the third score matrix of the target input text in the current processing cycle.

5. The method according to claim 3 or 4, characterized in that, The storing the second matrix into the memory includes: Store the third Query matrix, the third transposed Key matrix, the third Value matrix, and the third score matrix as the second matrix of this processing cycle in the memory.

6. The method according to claim 1, characterized in that The method further includes: Concatenate the input text and the output text of the attention network in the previous processing cycle to determine the concatenated text. Pad the concatenated text according to a preset text length threshold to determine the target input text of this processing cycle.

7. The method according to claim 3, wherein The determining the output text of this processing cycle according to the second matrix includes: Perform an activation function process on the third score matrix, and perform a matrix multiplication operation on the result of the activation function process and the third Value matrix to determine the context matrix. Determine the output text of this processing cycle according to the context matrix.

8. A text processing device, characterized in that, The apparatus includes: A target input text receiving module configured to receive the target input text of this processing cycle through the attention network, where the target input text includes repeated text and incremental text, the repeated text is the input text of the attention network in the previous processing cycle, and the incremental text is the output text of the attention network in the previous processing cycle. A target vector determining module configured to determine the target vector corresponding to the incremental text. A first matrix determining module configured to determine the first matrix of the repeated text in the previous processing cycle in the memory. A second matrix determining module configured to determine and store the second matrix of this processing cycle according to the target vector and the first matrix. An output text determining module configured to determine the output text of this processing cycle according to the second matrix.

9. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, where the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-7.