Data processing method, device and equipment and computer readable storage medium

By performing sequence traversal and sliding window updates on the splicing matrix of the splicing text, the accuracy reduction problem caused by the optimization of attention calculation speed is solved, and efficient and accurate attention calculation is achieved.

CN120030139APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311555034.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the prior art optimizes the speed of attention calculation, it will lead to a decrease in calculation accuracy and affect the accuracy of attention calculation.

Method used

The splicing weight template matrix is ​​determined by the splicing matrix based on the splicing text, and it is traversed in sequence using the calculation window to perform attention processing to obtain the sub-attention matrix. At the same time, the splicing weight template matrix is ​​traversed and updated through the sliding window to reduce unnecessary calculations.

Benefits of technology

Without affecting the calculation accuracy, the amount of attention calculation is reduced and the efficiency and speed of attention calculation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030139A_ABST
    Figure CN120030139A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method, device and equipment and a computer readable storage medium. The method comprises the steps of determining a splicing weight template matrix of a spliced text based on a splicing matrix of the spliced text; a calculation window is obtained, sequence traversal is carried out on each sequence of the splicing weight template matrix based on the calculation window, and the sequences are rows or columns of the splicing weight template matrix; in response to the condition that at least one of multiple elements, located in the calculation window, of the splicing weight template matrix during traversal is non-empty, attention processing is carried out based on the multiple elements, and a sub-attention matrix is obtained; in response to the situation that multiple elements, located in the calculation window, of the splicing weight template matrix are empty during traversal, sequence traversal is ended; and determining an attention matrix of the spliced text based on the plurality of sub attention matrixes in response to completion of traversing of the plurality of sequences of the spliced weight template matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular to a data processing method, device, equipment and computer-readable storage medium. Background Art

[0002] With the rapid development of Internet technology and the widespread application of machine learning algorithms, the attention mechanism has received more and more attention. At present, the attention mechanism is used in different types of machine learning tasks such as natural language processing and speech recognition. In the process of large model training, the output effect of the model can be effectively improved by introducing the attention mechanism. In addition, the attention mechanism can also simplify the model and increase the speed of model training. Therefore, the training speed of the training model can be further improved by optimizing the attention mechanism.

[0003] In the related art, the similarity matrix obtained in the attention calculation is sparsely processed in a linear manner to improve the computational efficiency of the attention calculation. However, this speed optimization method will reduce the computational precision of the attention model, thereby affecting the accuracy of the attention calculation. Summary of the invention

[0004] The embodiments of the present application provide a data processing method, apparatus, device, computer-readable storage medium, and computer program product, which can improve the calculation rate of attention calculation without affecting the calculation accuracy.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present application provides a data processing method, the method comprising:

[0007] Based on the splicing matrix of the spliced ​​text, determining the splicing weight template matrix of the spliced ​​text, wherein the spliced ​​text is obtained by splicing a plurality of texts, each of the texts includes at least one word, the splicing weight template matrix includes weight template matrices corresponding to the plurality of texts, and each element in the weight template matrix is ​​non-empty;

[0008] Acquire a calculation window, and perform sequence traversal on each sequence of the splicing weight template matrix based on the calculation window, wherein the sequence is a row or a column of the splicing weight template matrix;

[0009] In response to at least one element of the plurality of elements of the splicing weight template matrix located in the calculation window being non-empty during traversal, performing attention processing based on the plurality of elements to obtain a sub-attention matrix;

[0010] In response to the fact that a plurality of elements of the splicing weight template matrix located in the calculation window are all empty during the traversal, ending the sequence traversal;

[0011] In response to the completion of multiple sequence traversals of the splicing weight template matrix, an attention matrix of the spliced ​​text is determined based on multiple sub-attention matrices.

[0012] An embodiment of the present application provides a data processing device, including:

[0013] An acquisition module is used to determine a splicing weight template matrix of the spliced ​​text based on a splicing matrix of the spliced ​​text, wherein the spliced ​​text is obtained by splicing a plurality of texts, each of the texts includes at least one word, and the splicing weight template matrix includes weight template matrices corresponding to the plurality of texts, each element in the weight template matrix is ​​non-empty and is also used to acquire a calculation window;

[0014] A traversal module, configured to perform a sequence traversal on each sequence of the splicing weight template matrix based on the calculation window, wherein the sequence is a row or a column of the splicing weight template matrix, and to terminate the sequence traversal in response to a plurality of elements of the splicing weight template matrix located in the calculation window being empty during the traversal;

[0015] A splicing module, configured to determine an attention matrix of the spliced ​​text based on a plurality of sub-attention matrices in response to completion of a plurality of sequence traversals of the splicing weight template matrix;

[0016] A calculation module is used to perform attention processing based on the multiple elements of the splicing weight template matrix located in the calculation window to obtain a sub-attention matrix in response to at least one element being non-empty during traversal.

[0017] An embodiment of the present application provides an electronic device for data processing, the electronic device comprising:

[0018] Memory for storing computer programs or computer executable instructions;

[0019] The processor is used to implement the data processing method provided in the embodiment of the present application when executing the computer program or computer executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, can implement the data processing method provided in the embodiment of the present application.

[0021] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the data processing method provided in the embodiment of the present application is implemented.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] The splicing weight template matrix of the spliced ​​text is determined by the splicing matrix, and then each sequence of the splicing weight template matrix is ​​traversed through the calculation window, and attention processing is performed on multiple elements contained in the calculation window to obtain a sub-attention matrix. In addition, since there are empty elements in the splicing weight template matrix, when multiple elements contained in the calculation window are empty, the sequence traversal is stopped, thereby reducing the amount of attention calculation and improving the attention calculation efficiency without affecting the calculation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a structural diagram of a data processing system architecture provided by an embodiment of the present application;

[0025] Figure 2 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application;

[0026] Figure 3A It is a first flow chart of a data processing method provided by an embodiment of the present application;

[0027] Figure 3B is a second flow chart of a data processing method provided in an embodiment of the present application;

[0028] Figure 3C It is a third flow chart of a data processing method provided in an embodiment of the present application;

[0029] Figure 3D is a fourth flow chart of a data processing method provided in an embodiment of the present application;

[0030] Figure 3E is a fifth flow chart of a data processing method provided in an embodiment of the present application;

[0031] Figure 3F It is a sixth flow chart of a data processing method provided in an embodiment of the present application;

[0032] Figure 3G It is a seventh flow chart of a data processing method provided in an embodiment of the present application;

[0033] Figure 4 is a structural diagram of an attention model provided in an embodiment of the present application;

[0034] Figure 5 It is a structural diagram of a splicing weight template matrix provided in an embodiment of the present application;

[0035] Figure 6 It is a schematic diagram of a calculation window traversal provided in an embodiment of the present application;

[0036] Fig. 7A is a first structural diagram of a weight template matrix provided in an embodiment of the present application;

[0037] Figure 7B is a second structural diagram of a weight template matrix provided in an embodiment of the present application;

[0038] Figure 8 It is a schematic diagram of a data processing application process provided by an embodiment of the present application;

[0039] Fig. 9 is a schematic diagram of a calculation window traversal process provided by an embodiment of the present application;

[0040] Fig.10 is a structural diagram of a splicing matrix provided in an embodiment of the present application;

[0041] Fig.11 It is a schematic diagram of a matrix splicing process provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0043] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0044] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0045] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0046] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0047] 1) Sparse: Sparse refers to the method of using partial calculations to approximate the entire calculation. The basic idea of ​​model sparsification is to reduce or remove unnecessary parameters or layers by screening the parameters in the network, so as to obtain a lighter model.

[0048] 2) Attention: Attention is the core part of the transformer structure used by mainstream large models. The algorithm principle of the Attention mechanism is based on the idea of ​​weight distribution. By assigning different weights to each element in the input sequence, the model can focus on the information most relevant to the current task.

[0049] 3) Packing samples: Packing samples means combining multiple short samples into one long sample. This is an operation method used to improve computing efficiency and graphics processing unit (GPU) utilization. Packing samples can reduce idle time in computing and thus increase training speed.

[0050] In related technologies, the attention model is mainly accelerated and optimized through linear sparsification methods. Although this method can achieve the effect of acceleration, there will be certain losses in the calculation process, which reduces the accuracy of attention calculation.

[0051] The embodiments of the present application provide a data processing method, apparatus, device, computer-readable storage medium and computer program product, which can traverse the splicing weight template matrix of the spliced ​​text based on the calculation window, obtain multiple sub-attention matrices through attention calculation, and determine the attention matrix of the spliced ​​text based on the multiple sub-attention matrices, thereby solving the problem of calculation accuracy loss in the related art.

[0052] The data processing method provided in the embodiment of the present application can be implemented by the terminal alone; it can also be implemented by the terminal and the server in collaboration. For example, the terminal alone undertakes the data processing method below, or the terminal sends text data to the server, and the server executes the data processing method according to the received text data. In the process of data processing, based on the splicing matrix of the spliced ​​text, the splicing weight template matrix of the spliced ​​text is determined; a calculation window is obtained, and a sequence traversal is performed on each sequence of the splicing weight template matrix based on the calculation window; in response to at least one element of the multiple elements of the splicing weight template matrix located in the calculation window during the traversal being non-empty, attention processing is performed based on the multiple elements to obtain a sub-attention matrix; in response to the multiple elements of the splicing weight template matrix located in the calculation window during the traversal being empty, the sequence traversal is ended; in response to the completion of the traversal of multiple sequences of the splicing weight template matrix, the attention matrix of the spliced ​​text is determined based on the multiple sub-attention matrices to improve the data processing rate.

[0053] The following describes exemplary applications of the electronic device provided in the embodiments of the present application. The electronic device provided in the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (for example, mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices, vehicle-mounted devices), smart phones, smart speakers, smart watches, smart televisions, and vehicle-mounted terminals.

[0054] In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms. Among them, cloud services can be real-time data processing services for terminals to call.

[0055] Next, an exemplary application when the electronic device is implemented as a server will be described.

[0056] See also Figure 1 , Figure 1 It is a schematic diagram of the architecture of the data processing system 100 provided in an embodiment of the present application. The terminal device 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0057] In some embodiments, taking the electronic device as a terminal as an example, the data processing method provided in the embodiment of the present application can be implemented by the terminal. For example, the terminal 400 locally executes the data processing method provided in the embodiment of the present application. In the process of data processing, based on the splicing matrix of the spliced ​​text, the splicing weight template matrix of the spliced ​​text is determined; the calculation window is obtained, and each sequence of the splicing weight template matrix is ​​sequence traversed based on the calculation window; in response to at least one element of the multiple elements of the splicing weight template matrix located in the calculation window during traversal being non-empty, attention processing is performed based on multiple elements to obtain a sub-attention matrix; in response to the multiple elements of the splicing weight template matrix located in the calculation window during traversal being empty, the sequence traversal is terminated; in response to the completion of the traversal of multiple sequences of the splicing weight template matrix, the attention matrix of the spliced ​​text is determined based on multiple sub-attention matrices to improve the data processing rate.

[0058] As an example, in the application scenario of Chinese-English translation, when the input multiple texts are "Good morning" and "Have you had breakfast?", the spliced ​​text is "Good morning! Have you had breakfast?" The translation model can calculate the attention matrix corresponding to the spliced ​​text through the data processing method provided in the embodiment of the present application, and then perform text translation on the attention matrix of the spliced ​​text to obtain the translated text "Good morning! Did you have breakfast".

[0059] In some embodiments, the real-time data processing method provided by the embodiment of the present application can also be implemented by the server and the terminal in collaboration. For example, the text data generated by the terminal device 400 can be transmitted to the server 200 through the network 300, and the server 200 is used to determine the splicing weight template matrix of the splicing text based on the splicing matrix of the splicing text, obtain the calculation window, and perform sequence traversal on each sequence of the splicing weight template matrix based on the calculation window, in response to at least one element of the multiple elements of the splicing weight template matrix located in the calculation window during the traversal being non-empty, performing attention processing based on multiple elements to obtain a sub-attention matrix, in response to the multiple elements of the splicing weight template matrix located in the calculation window during the traversal being empty, ending the sequence traversal, in response to the completion of the traversal of multiple sequences of the splicing weight template matrix, determining the attention matrix of the splicing text based on multiple sub-attention matrices, and improving the computational efficiency of the model.

[0060] As an example, in an intelligent dialogue scenario, when multiple texts are input as "Hello" and "Excuse me, where is the bathroom?", the terminal device will obtain the spliced ​​text as "Hello, where is the bathroom?", and the terminal device will send the spliced ​​text to the server. The server can calculate the attention matrix corresponding to the spliced ​​text through the data processing method provided in the embodiment of the present application, and then perform text reply processing on the attention matrix of the spliced ​​text to obtain the reply text "The location of the bathroom is XXX."

[0061] It should be noted that the application scenarios of the embodiments of the present application are not limited to the above-mentioned translation scenarios and dialogue generation scenarios, but can also be text recommendation scenarios, missing text supplement scenarios, etc., which are not specifically limited in the embodiments of the present application.

[0062] In some embodiments, the terminal or server can implement the data processing method provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run; it can also be a small program embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0063] In some embodiments, multiple servers may form a blockchain, and server 200 is a node on the blockchain. Information connections may exist between each node in the blockchain, and information may be transmitted between nodes through the above information connections. Among them, data (such as text data) related to the data processing method provided in the embodiment of the present application may be stored on the blockchain.

[0064] The structure of the electronic device provided by the embodiment of the present application is described below. Figure 2 , Figure 2 is a schematic diagram of the structure of the data processing terminal 200 provided in an embodiment of the present application, Figure 2 The terminal 200 shown includes: at least one processor 210, a memory 250, at least one network interface 220 and a user interface 230. The various components in the terminal 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 2 Various buses are labeled as bus system 240 .

[0065] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0066] The user interface 230 includes one or more output devices 231 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0067] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.

[0068] The memory 250 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0069] In some embodiments, memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0070] Operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0071] A network communication module 252, used to reach other electronic devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 include: Bluetooth, Wireless Compatibility Authentication (WiFi), and Universal Serial Bus (USB), etc.;

[0072] a presentation module 253 for enabling presentation of information via one or more output devices 231 (e.g., display screen, speaker, etc.) associated with the user interface 230 (e.g., a user interface for operating peripherals and displaying content and information);

[0073] The input processing module 254 is used to detect one or more user inputs or interactions from one of the one or more input devices 232 and translate the detected inputs or interactions.

[0074] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2 The data processing device 255 stored in the memory 250 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: an acquisition module 2551, a traversal module 2552, a splicing module 2553 and a calculation module 2554. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0075] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.

[0076] The data processing method provided by the embodiment of the present application will be described in conjunction with the exemplary application and implementation of the electronic device provided by the embodiment of the present application. Figure 3A , Figure 3A is a flow chart of a data processing method provided in an embodiment of the present application, combined with Figure 3A The steps shown are explained, Figure 3A The step execution subject is an electronic device.

[0077] In some embodiments, before step 101, the lengths of multiple input texts may be determined, and an input text length threshold may be obtained; the multiple input texts may be sorted in descending order according to the length of each input text; according to the sorting sequence of the input texts, N adjacent input texts and N+1 adjacent input texts starting from the target text may be obtained, wherein the length of the target text is less than the input text length threshold; in response to the sum of the lengths of the N adjacent input texts being less than or equal to the input text length threshold, and the sum of the lengths of the N+1 adjacent input texts being greater than the input text length threshold, the N adjacent input texts may be treated as multiple texts, and the N adjacent input texts may be spliced ​​to obtain a spliced ​​text; based on the spliced ​​text, a splicing matrix of the spliced ​​text may be obtained.

[0078] For example, when the input text length threshold is 15, the sum of the lengths of four adjacent input texts obtained from the same starting point in the sorted sequence is 14, and the sum of the lengths of five adjacent input texts obtained in the sorted sequence is 20, the four input texts obtained are concatenated to obtain a concatenated text, and a concatenation matrix is ​​generated based on the concatenated text.

[0079] In some embodiments, obtaining a splicing matrix of the spliced ​​text based on the spliced ​​text is achieved by: determining text matrices corresponding to multiple texts respectively, and performing splicing processing on the text matrices corresponding to the multiple texts respectively to obtain the splicing matrix of the spliced ​​text.

[0080] As an example, when the spliced ​​text "I love China, China is my hometown." is generated by splicing "I love China" and "China is my hometown.", the spliced ​​matrix of the spliced ​​text "I love China, China is my hometown." can be obtained by obtaining the text matrices of "I love China" and "China is my hometown." respectively, and then splicing multiple text matrices. Fig.11 As shown, when the spliced ​​text consists of two texts, the text matrix 1 and the text matrix 2 respectively obtained are spliced ​​to obtain a spliced ​​matrix.

[0081] In an embodiment of the present application, by performing attention processing on a spliced ​​matrix obtained by splicing multiple texts, compared with performing attention processing on the text matrices of multiple texts separately, the idle time in the calculation can be reduced, thereby improving the training speed of the electronic device.

[0082] In some embodiments, before step 101, each text may be vectorized to obtain text data of each text; the text data of each text may be position-encoded to obtain a text vector of each text; and the text vectors of multiple texts may be spliced ​​to obtain a splicing matrix of the spliced ​​texts.

[0083] As an example, by performing vectorization processing on the texts "morning", "on", and "good" respectively, text data x corresponding to each text can be obtained. 1 , x 2 , x 3 , and then perform positional encoding on the text data x 1 , x 2 , x 3 respectively to obtain text vectors X corresponding to each text 1 , X 2 , X 3 . Finally, splice X 1 , X 2 , X 3 to obtain a spliced matrix [X 1 X 2 X 3 .

[0084] In step 101, based on the spliced matrix of the spliced text, determine the spliced weight template matrix of the spliced text.

[0085] Among them, the spliced text is obtained by splicing multiple texts, each text includes at least one word, the spliced weight template matrix includes weight template matrices corresponding to multiple texts respectively, and the elements in the lower triangle of each weight template matrix are non-empty. Among them, the weight template matrix includes two types of elements: empty and non-empty. The dimension of the weight matrix corresponding to the weight template matrix is the same as the dimension of the weight template matrix. Among them, the non-empty elements in the weight template matrix are used to represent that the elements in the same position as the non-empty elements in the weight matrix need to participate in the attention calculation (that is, multiply by the value matrix). As Fig. 7A shown, the elements in the lower triangle of the weight template matrix are all non-empty, and the elements in the upper triangle are all empty. The elements in the lower triangle of the weight matrix corresponding to the weight template matrix need to participate in the attention calculation, and the elements in the upper triangle of the weight matrix corresponding to the weight template matrix do not need to participate in the attention calculation.

[0086] In some embodiments, referring to Figure 3B , Figure 3A shown, step 101 can be implemented by the following steps 1011 to 1013, which will be specifically described below.

[0087] In step 1011, identify the spliced matrix based on the delimiter to obtain sub-matrices corresponding to multiple texts.

[0088] Among them, the spliced matrix is obtained by splicing sub-matrices corresponding to multiple texts respectively. Since after splicing multiple texts to generate the spliced text, delimiters will be inserted between different texts for text partitioning, therefore, the spliced matrix can be cut based on the position of the delimiter by identifying the delimiter to obtain sub-matrices corresponding to each text. As Fig.10 As shown, Fig.10 The concatenated matrix in is composed of the sub-matrices corresponding to four texts (including text 1, text 2, text 3, and text 4), and there are separators between the sub-matrices corresponding to different texts.

[0089] In step 1012, based on the sub-matrix, a weight template matrix corresponding to the text is determined.

[0090] Specifically, the length of each corresponding text can be obtained by identifying the submatrix, and the weight template matrix corresponding to each text is generated according to the text length. For example, the number of rows n of the submatrix is ​​determined, the number of rows is used as the text length, and the weight template matrix corresponding to each text is generated based on the text length, wherein the number of rows and columns of the weight template matrix are both n, and the elements of the lower triangle in the weight template matrix are non-empty.

[0091] In step 1013, the weight template matrices corresponding to the multiple texts are combined to obtain a concatenated weight template matrix.

[0092] For example, the weight template matrix corresponding to multiple texts is [A 1 ]、[A 2 ], the combination process can be to transform the weight template matrix [A 1 ] and the weight template matrix [A 2 ] to obtain the splicing weight template matrix

[0093] In some embodiments, see Figure 3C ,exist Figure 3A After step 101 shown, steps 201 to 203 may also be performed, which are described in detail below.

[0094] In step 201, the text length of the concatenated text is obtained, and a sliding window is determined based on the text length.

[0095] As an example, determining the sliding window based on the text length is achieved in the following way: after obtaining the number of attention layers, the length range of the sliding window can be calculated based on the text length and the number of layers.

[0096] The data processing method in the embodiment of the present application is implemented by an attention model, wherein the attention model includes multiple cascaded attention layers, such as Figure 4 As shown, Figure 4 It is a schematic diagram of the structure of the attention model containing 4 attention layers. Figure 4 The shaded part in the middle represents the concatenated text input to the first attention layer, and the length of the sliding window is 5.

[0097] In the attention model, based on the characteristics of multiple layers, the output of each layer will be used as the input of the next layer, so the character information contained in the text sequence will be propagated in multiple attention layers. Figure 4 , the feature of the first character in the first layer will propagate to the position of the 13th character in the fourth layer. Since the propagation range exceeds the limit of the sliding window, there is no loss in data calculation accuracy during data processing.

[0098] It can be seen that as long as the attention model parameters and the length of the sliding window satisfy formula (1), the weight template matrix can be updated by adding a sliding window, but using the updated weight template matrix will not affect the calculation accuracy.

[0099] l≤(n layer -1)(size window -1)+1 (1)

[0100] Among them, l represents the text length of the concatenated text, n layer Indicates the number of attention layers and size window Indicates the length of the sliding window.

[0101] In some embodiments, the number of attention layers can be obtained, and the sliding window can be determined based on the text length and the number of layers.

[0102] Specifically, after obtaining the number of attention layers and the text length, the minimum length of the sliding window is calculated according to formula (2). The length of the sliding window is greater than or equal to the minimum length and less than any integer within the text length range.

[0103] l=(n layer -1)(size window -1)+1 (2)

[0104] Among them, l represents the text length of the concatenated text, n layer Indicates the number of attention layers and size window Indicates the length of the sliding window, size window The calculation result is rounded up to get the minimum length of the sliding window.

[0105] In an embodiment of the present application, by adding a sliding window to limit the data range involved in the attention calculation, the speed of data processing can be improved without affecting the calculation accuracy.

[0106] In step 202, in response to the maximum text length among the lengths of the plurality of texts being greater than the length of the sliding window, the splicing weight template matrix is ​​traversed based on the sliding window, and the elements outside the sliding window are marked as empty.

[0107] In an embodiment of the present application, by marking the elements outside the sliding window as empty, the amount of attention calculation can be reduced without affecting the calculation accuracy, thereby improving the attention calculation speed.

[0108] In some embodiments, see Figure 3D , Figure 3C The illustrated step 202 may be implemented by following steps 2021 to 2023, which are described in detail below.

[0109] In step 2021, the position information of the elements in the splicing weight template matrix is ​​obtained.

[0110] The position information includes the row position x and column position y of the element in the splicing weight template matrix.

[0111] In step 2022, the difference between the row position and the column position is determined.

[0112] For example, when the element position information is x=10, y=2, the difference between the row position and the column position of the current element is 8.

[0113] In step 2023, in response to the difference being greater than or equal to the size of the sliding window, the element is marked as empty.

[0114] Specifically, if the current element difference is greater than or equal to the size of the sliding window, it means that the current element is outside the sliding window range, so the current element is marked as empty, indicating that the current element does not participate in the attention calculation.

[0115] For example, when the position of element M is (10, 2) and the length of the sliding window is 5, since the difference between the row position and the column position of the current element is greater than the length of the sliding window, the element M is marked as empty.

[0116] In step 203, the splicing weight template matrix is ​​updated according to the traversal result.

[0117] In an embodiment of the present application, the splicing weight template matrix is ​​updated by traversing the results, that is, the elements outside the sliding window range are marked as empty, which can reduce the range of data involved in the calculation, thereby improving the data processing speed.

[0118] In step 102, a calculation window is obtained, and a sequence traversal is performed on each sequence of the splicing weight template matrix based on the calculation window.

[0119] The sequence is the row or column of the splicing weight template matrix.

[0120] In some embodiments, see Figure 3E , Figure 3AThe illustrated step 102 may be implemented by following the steps 1021 to 1022, which are described in detail below.

[0121] In step 1021, the initial position of the calculation window and the sequence traversal direction are determined.

[0122] Specifically, in the splicing weight template matrix, the sequence is identified along the sequence direction by calculating the window starting from the sequence start point, and in response to identifying that at least one of the multiple elements in the calculation window is non-empty, the current calculation window position is used as the initial position, and the sliding window traverses the sequence starting from the initial position. Figure 6 As shown, Figure 6 The box in the figure represents the computation window, and the arrow represents the traversal direction.

[0123] In step 1022, starting from the initial position of the calculation window, each sequence of the splicing weight template matrix is ​​traversed along the sequence traversal direction through the calculation window.

[0124] In the process of traversing each sequence of the splicing weight template matrix through the calculation window, the calculation windows in two adjacent calculations are located adjacently. Fig. 9 As shown, when traversing the sequence, the first calculation position is a 1 , the second calculation position is a 2 、The third calculation position is a 3 , these three calculation positions are adjacent.

[0125] In step 103, in response to at least one element of the multiple elements of the splicing weight template matrix located in the calculation window being non-empty during traversal, attention processing is performed based on the multiple elements to obtain a sub-attention matrix.

[0126] In an embodiment of the present application, in response to at least one element among multiple elements in the calculation window being non-empty, block attention calculation is performed on the elements in the splicing weight template matrix, which can reduce redundancy in the splicing matrix attention calculation process and improve the calculation speed.

[0127] In some embodiments, see Figure 3F , Figure 3A The illustrated step 103 may be implemented by following steps 1031 to 1034, which are described in detail below.

[0128] In step 1031, a sub-input matrix is ​​generated based on a plurality of elements.

[0129] The sub-input matrix is ​​a sub-matrix of the splicing matrix. Specifically, in response to at least one element of the multiple elements of the splicing weight template matrix located in the calculation window during traversal being non-empty, according to the position of the calculation window in the splicing weight template matrix, the elements contained in the corresponding position in the weight matrix are obtained, and the multiple elements are combined into a sub-input matrix.

[0130] In step 1032, a linear transformation is performed on the sub-input matrix to obtain a query matrix, a key matrix, and a value matrix.

[0131] Among them, the value matrix is ​​the vector representing the input features, the query matrix and the key matrix are the feature vectors for calculating the attention weights.

[0132] In step 1033, weight-based attention processing is performed on the query matrix and the key matrix to obtain a sub-weight matrix.

[0133] In an embodiment of the present application, based on the calculation window, the elements in the splicing weight template matrix are processed in blocks, and attention calculation is performed on each sub-input matrix to obtain a sub-weight matrix corresponding to each sub-input matrix.

[0134] In some embodiments, see Figure 3G , Figure 3F The illustrated step 1033 may be implemented by following the steps 10331 to 10332, which are described in detail below.

[0135] In step 10331, weight-based similarity processing is performed on the query matrix and the key matrix to obtain a similarity matrix.

[0136] Specifically, for the query matrix and the key matrix, the similarity matrix can be obtained by calculating using the following formula (3).

[0137] m=qk T (3)

[0138] Among them, m represents the similarity matrix, q represents the query matrix, and k represents the key matrix.

[0139] In step 10332, the similarity matrix is ​​normalized to obtain a sub-weight matrix.

[0140] Specifically, the similarity matrix can be normalized using the following formula (4) to obtain a sub-weight matrix.

[0141]

[0142] Among them, p represents the similarity matrix and d is the rank of the key matrix.

[0143] In step 1034, the sub-weight matrix and the value matrix are weighted and summed to obtain a sub-attention matrix.

[0144] Specifically, for the sub-weight matrix and the vector in the value matrix, the sub-attention matrix can be obtained by calculating using the following formula (5).

[0145] o=pv (5)

[0146] Among them, o represents the sub-attention matrix and v represents the value matrix.

[0147] In step 104, in response to the fact that a plurality of elements of the splicing weight template matrix located in the calculation window are all empty during the traversal, the sequence traversal is terminated.

[0148] In an embodiment of the present application, when it is identified that multiple elements in the calculation window are empty, the traversal of the current sequence is terminated, which can reduce the calculation redundancy of the current sequence and improve the model calculation speed.

[0149] In some embodiments, after step 104, the adjacent sequence of the current sequence in the traversed splicing weight template matrix can also be obtained; based on the calculation window, the adjacent sequence is subjected to attention processing to obtain a sub-attention matrix of the adjacent sequence.

[0150] Specifically, when the traversal of the first sequence ends based on the calculation window, a second sequence adjacent to the first sequence is obtained in the splicing weight template matrix, wherein the first sequence is a row or column of the splicing weight template matrix. The traversal initial position of the calculation window is determined in the second sequence, and attention calculation is performed on the elements in the second sequence to obtain a sub-weight matrix of the second sequence.

[0151] In step 105, in response to the completion of multiple sequence traversals of the splicing weight template matrix, an attention matrix of the spliced ​​text is determined based on multiple sub-attention matrices.

[0152] For example, the concatenated weight template matrix contains two sequences, namely the first sequence X and the second sequence Y. Based on the calculation window, the first sequence is traversed based on the calculation window to obtain the sub-attention matrix x 1 、x 2 , traverse the second sequence and get the sub-attention matrix y 1 , concatenate multiple sub-attention matrices to obtain the attention matrix

[0153] The embodiment of the present application uses a calculation window to perform a sequence traversal on each sequence of the splicing weight template matrix. When at least one of the multiple elements in the calculation window is non-empty, attention processing is performed on the multiple elements, and in response to the multiple elements being empty during the traversal, the traversal of the current sequence is ended, which can improve the rate of attention calculation without affecting the calculation accuracy of the model. In addition, the embodiment of the present application also traverses and updates the splicing weight template matrix through a sliding window, further reducing the amount of attention calculation data, thereby further improving the model calculation speed.

[0154] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.

[0155] During the data processing process, there is a problem of loss of calculation accuracy when accelerating the optimization of the attention model. The embodiment of the present application proposes a data processing method, which determines the splicing weight template matrix of the spliced ​​text through the splicing matrix of the spliced ​​text, and then traverses the splicing weight template matrix through the calculation window and performs attention calculation to obtain multiple sub-attention matrices. Finally, the multiple sub-attention matrices are spliced ​​to obtain the attention matrix of the spliced ​​text, which can improve the model calculation speed while ensuring the calculation accuracy.

[0156] In the data processing process of related technologies, the Transformer structure is often used to improve the training speed of the model. Attention is the core part of the Transformer structure, which includes the following important parts:

[0157] (1) Calculating attention weights: The core of the attention mechanism is to calculate the weights of each element in the input sequence. These weights are usually calculated by a learnable function, for example, by a neural network. Specifically, the query vector, key vector, and value vector can be obtained by linearly transforming the input vector, and the attention weights are obtained by calculating the similarity between the query vector and each key vector, where the similarity measurement method can be a click or additive attention method.

[0158] (2) Normalized weights: In order to ensure that the sum of the attention weights is 1, the calculated attention weights need to be normalized. This can be achieved through the softmax function, which can map any real number to the (0, 1) interval and ensure that the sum of all weights is 1.

[0159] (3) Weighted summation: Based on the calculated normalized weights, each element in the input sequence is weighted and summed, so that elements that are more relevant to the current task can obtain higher weights and thus occupy a larger proportion in the final weighted summation result. This weighted summation result can be used as the output of the attention mechanism for subsequent calculations or prediction models.

[0160] (4) Optional Multi-head Attention: Using the multi-head attention method, the model can focus on multiple different information sources. In this method, the query, key, and value vectors are projected into multiple different subspaces, and then the attention weights are calculated independently in each subspace and weighted summed. Finally, the weighted sum results of all subspaces are concatenated to form the final output.

[0161] Figure 8 A schematic diagram of a data processing application flow provided by an embodiment of the present application. Taking the input text (such as the above text or concatenated text) with a length of 10 and a dimension of 16 as an example, Figure 8 Specifically, the process includes the following steps:

[0162] (1) Perform vectorization on the input text to obtain text data I(15*16) corresponding to the input text, where the dimension in brackets is the text data.

[0163] (2) Perform position encoding on the text data to obtain the text vector I'(15*16) corresponding to the input text.

[0164] (3) The text vector is calculated through the attention mechanism to obtain the attention matrix.

[0165] Specifically, Figure 8 As shown, by calling the qkv proj instruction to process the text vector, we get the query matrix q(15*16), key matrix k(15*16), and value matrix v(15*16). Then, based on formula (6), we calculate the weights of q(15*16) and k(15*16) to get the weight matrix p(15*16). Finally, based on formula (7), we calculate the attention matrix of p(15*16) and v(15*16).

[0166] (4) The attention matrix is ​​processed by the output layer (implemented by a feedforward neural network (FNN)) to obtain the output sequence.

[0167] The attention mechanism algorithm enables the model to automatically learn how to allocate attention, so as to more effectively extract key information from the input data. In fields such as natural language processing, computer vision, and speech recognition, the attention mechanism has been proven to be a very effective model optimization technique that can significantly improve the performance and generalization ability of the model.

[0168] In the Attention calculation, the weight template matrix can be used to represent the part of the input matrix involved in the calculation, such as Fig. 7A As shown in the figure, the two dimensions of the weight template matrix are the query dimension and the key dimension. The blank part represents the part that does not participate in the calculation. In the Large Language Model (LLM), Token is used to represent the smallest data unit of the input text, such as a word or a punctuation mark. For models with a decoder-only architecture such as the Generative Pre-Training (GPT), only the Attention before the current data unit needs to be calculated, that is, only the lower triangular part of the input matrix needs to be calculated, such as Fig. 7A shown.

[0169] When using LLM tokens for training and deploying models, text data needs to be converted into a sequence of numbers, a process called tokenization. The purpose of tokenization is to convert text data into a form that can be processed by computers for tasks such as natural language processing and text generation. In the LLM model, tokenization is similar to other natural language processing models, for example, rule-based or machine learning-based methods can be used for calculations.

[0170] Flash Attention is a new type of attention mechanism that can improve computational efficiency and reduce computational complexity. Flash Attention improves computational speed by reducing computational effort and memory requirements. The principles of the FlashAttention algorithm include:

[0171] (1) Low-rank Matrix Decomposition: The core idea of ​​Flash Attention is to use low-rank matrix decomposition to calculate attention weights. In traditional attention mechanisms, calculating weights requires matrix multiplication of the query matrix and the key matrix, which results in high computational complexity. In Flash Attention, the key matrix is ​​decomposed into the product of two low-rank matrices, thereby reducing the computational complexity to a linear level.

[0172] (2) Block-wise Computation: In order to further improve computational efficiency, FlashAttention adopts a block-wise computation strategy. Specifically, the input sequence is divided into multiple smaller blocks, and the attention weight is calculated independently in each block. This can reduce the amount of computation and the loss of computational accuracy, and improve the computational speed.

[0173] (3) Dynamic Programming: In Flash Attention, dynamic programming can also be used to optimize the calculation process. Specifically, after calculating the local optimal solution in each block, these local solutions are merged into a global solution. This can reduce the calculation complexity and have higher calculation accuracy.

[0174] (4) Optional Multi-head Attention: Similar to the attention mechanism, FlashAttention can also be combined with the multi-head attention method. In this case, FlashAttention can be applied independently in the subspace, allowing the model to focus on multiple different sources of information.

[0175] In summary, Flash Attention is an efficient attention mechanism that improves the computing speed and reduces the computational complexity through technologies such as low-rank matrix decomposition, block calculation, and dynamic programming. This makes Flash Attention a method suitable for processing large-scale data and real-time applications.

[0176] In large model training, in order to improve computing efficiency and GPU utilization, multiple short samples are usually packed into one long sample. The operation of packing samples is mainly aimed at sequence data of variable length, such as text data in natural language processing tasks. In an embodiment of the present application, by combining multiple short texts into a concatenated text and then performing data processing, the idle time of the electronic device during the calculation process can be reduced, thereby increasing the speed of model training. The following is a detailed explanation of the operation of packing multiple short texts into one long text:

[0177] (1) Sorting the input texts: First, multiple input texts need to be sorted in descending order according to the length of the input texts. This makes it easier and faster to splice multiple input texts in the subsequent data processing process.

[0178] (2) Concatenating short texts: After sorting the input text in descending order, multiple adjacent texts can be concatenated to obtain a concatenated text. Specifically, N adjacent input texts and N+1 adjacent input texts starting from the target text can be obtained. In response to the sum of the lengths of the N adjacent input texts being less than or equal to the input text length threshold, and the sum of the lengths of the N+1 adjacent input texts being greater than the input text length threshold, the N adjacent input texts are concatenated to obtain a concatenated text. In order to distinguish multiple input texts in the concatenated text, a special delimiter is inserted between each input text. For example, the delimiter can be a special token or mark.

[0179] (3) Training model: After multiple texts are concatenated to obtain the concatenated text, the concatenated text can be passed as input to the neural network model for training. Since the concatenated text contains multiple texts, the idle time of the GPU during the calculation process will be reduced, thereby improving the calculation efficiency of the model and the utilization rate of the GPU.

[0180] (4) Separating short samples: After the training is completed, the multiple texts in the concatenated text can be separated by detecting the separators in the output text corresponding to the concatenated text for subsequent analysis and evaluation.

[0181] In summary, in large model training, concatenating multiple input texts to obtain concatenated text is an effective optimization operation. By combining multiple texts and filling them, the computational efficiency of the model and the utilization of the GPU can be improved, thereby increasing the execution speed of the training process. The Pack sample operation is of great significance for model training involving variable-length sequence processing, such as natural language processing.

[0182] In the embodiment of the present application, due to the pack sample operation, the concatenated text weight template matrix is ​​a sparse matrix. This is because there is no dependency between the multiple texts packed together, so there is no need to perform Attention calculations between different texts. It can be seen that there is redundancy in the Attention calculation of the concatenated text, especially when the number of multiple texts that make up the concatenated text is large, there are more redundant calculations in the concatenated matrix attention calculation process. See Figure 5 , Figure 5 A schematic diagram of the structure of a splicing weight template matrix provided in an embodiment of the present application. The shaded part in the figure represents the element part involved in the Attention calculation.

[0183] Based on the above reasons, the embodiment of the present application performs optimization and acceleration based on flash attention. Specifically, Figure 6As shown in the figure, in the GPU kernel-based programming optimization process, when performing Attention calculation, the calculation window traverses each sequence in the dimension of the key vector. In response to the fact that multiple elements of the splicing weight template matrix located in the calculation window are empty during traversal, the sequence traversal is terminated, which can improve the operation efficiency by reducing unnecessary calculations. The specific steps of the optimization algorithm are as follows:

[0184] 1) Detect the initial position of the calculation window.

[0185] Specifically, the calculation window is gradually moved from the start position of the sequence, and in response to at least one element of the multiple elements of the splicing weight template matrix located in the calculation window being non-empty during the movement, this position is determined as the initial position of the calculation window.

[0186] 2) Starting from the initial position of the calculation window, perform Attention calculation on the elements in the calculation window to obtain the sub-attention matrix, and then move the calculation window according to the sequence traversal direction.

[0187] Specifically, based on multiple elements in the calculation window, a sub-input matrix is ​​generated, and the query matrix, key matrix and value matrix are obtained by performing linear transformation on the sub-input matrix; then, attention calculation is performed on the query matrix and the key matrix to obtain a sub-weight matrix; finally, a weighted summation process is performed on the sub-weight matrix and the value matrix to obtain a sub-attention matrix.

[0188] The calculation formula of the sub-weight matrix is ​​shown in formula (6):

[0189]

[0190] Among them, p represents the sub-weight matrix, q represents the query matrix, k represents the key matrix, and d represents the dimension of k.

[0191] The calculation formula of the sub-attention matrix is ​​shown in formula (7):

[0192] o=pv (7)

[0193] Among them, o represents the sub-attention matrix and v represents the value matrix.

[0194] 3) When multiple elements in the calculation window are empty, it means that in the current traversal sequence, the calculations after the current calculation window are redundant calculations, so the calculation of the current sequence can be terminated in advance.

[0195] In some embodiments, a sliding window can be used to further limit the scope of Attention calculation during the above optimization process. Figure 7B As shown, Figure 7BTo increase the structural diagram of the weight template matrix of the sliding window, the shaded part in the figure represents the elements involved in the Attention calculation, and the blank part represents the elements that do not need to participate in the Attention calculation. Figure 7B It can be seen that by adding a sliding window, the Attention calculation between elements can be limited, and by limiting the part outside the sliding window range, the Attention calculation does not need to be performed, thereby achieving the purpose of acceleration.

[0196] Since the Transformer model is a multi-layer stacked model, the Transformer structure is repeatedly stacked to form the overall model structure, such as Figure 4 See Figure 4 ,In the large model data processing process, based on the ,multi-layer characteristics, the token will propagate in multiple layers, where the output of the ,previous layer will serve as the input of the next layer, and so on.

[0197] exist Figure 4 In the structure shown, the input text length threshold is 17 and the sliding window length is 5. During the data processing, as the layers increase, by the fourth layer, the features of the first token will propagate to the position of 13 tokens. It can be seen that through multi-layer propagation, the influence range of the token will exceed the limit of the sliding window length, that is, when the sliding window is increased, the calculation accuracy is not affected. The actual influence range of the sliding window can be calculated by formula (8).

[0198] S=(n layer -1)(size window -1)+1 (8)

[0199] Among them, S represents the impact range, n layer Indicates the number of attention layers and size window Indicates the length of the sliding window.

[0200] In order not to affect the calculation accuracy of the model, the influence range of the sliding window should be greater than or equal to the text length of the concatenated text, that is, it should satisfy formula (9).

[0201] l≤(n layer -1)(size window -1)+1 (9)

[0202] Among them, l represents the text length of the concatenated text, n layer Indicates the number of attention layers and size window Indicates the length of the sliding window.

[0203] In the embodiment of the present application, the specific steps of the sliding window algorithm are as follows:

[0204] (1) Obtain the text length of the concatenated text and determine the sliding window based on the text length.

[0205] (2) In response to the maximum text length among the lengths of the multiple texts being greater than the length of the sliding window, the splicing weight template matrix is ​​traversed based on the sliding window.

[0206] (3) Determine the position information of the element in the splicing weight template matrix, where the position information includes the row position and column position of the element in the splicing weight template matrix.

[0207] (4) Calculate the difference between the row position and the column position, and mark the element as empty in response to the difference being greater than or equal to the size of the sliding window.

[0208] (5) Update the splicing weight template matrix according to the traversal results.

[0209] In some embodiments, a sliding window is applied to a weight matrix, and the specific application process is as follows:

[0210] Taking the input text (such as the above text or concatenated text) with a length of 10 and a dimension of 16 as an example, the input text is vectorized to obtain the text data I (15*16) corresponding to the input text; the text data is positionally encoded to obtain the text vector I' (15*16) corresponding to the input text; the text vector is calculated through the attention mechanism to obtain the attention matrix. Specifically, by calling qkv The proj instruction processes the text vector to obtain the query matrix q(15*16), the key matrix k(15*16), and the value matrix v(15*16). Then, the weights of q(15*16) and k(15*16) are calculated based on formula (6) to obtain the weight matrix p(15*16). The weight matrix is ​​traversed based on the sliding window, and the weight matrix is ​​updated according to the traversal results to obtain the updated weight matrix p'(15*16). Finally, the attention matrix is ​​obtained by calculating the attention of p'(15*16) and v(15*16) using formula (7). The attention matrix is ​​processed by the output layer (implemented by the feedforward neural network (FNN)) to obtain the output sequence.

[0211] By implementing the optimization method in the embodiment of the present application through the Trion framework, the efficiency of the attention calculation part can be improved by 30%, and the performance of the model remains unchanged, that is, the loss remains consistent with the dense attention.

[0212] In an embodiment of the present application, a sequence traversal is performed on each sequence of the splicing weight template matrix through a calculation window, and in response to at least one element among the multiple elements in the calculation window being non-empty, attention processing is performed on the multiple elements, and in response to multiple elements being empty during the traversal, the traversal of the current sequence is ended, so that the rate of attention calculation can be improved without affecting the accuracy of model calculation. In addition, the embodiment of the present application also traverses and updates the splicing weight template matrix through a sliding window, further reducing the amount of attention calculation data, thereby further improving the model calculation speed.

[0213] The following further describes an exemplary structure of the data processing device 255 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules stored in the data processing device 255 of the memory 250 may include:

[0214] Acquisition module 2551 is used to determine the splicing weight template matrix of the spliced ​​text based on the splicing matrix of the spliced ​​text, wherein the spliced ​​text is obtained by splicing multiple texts, each of the texts includes at least one word, and the splicing weight template matrix includes weight template matrices corresponding to the multiple texts respectively, and each element in the weight template matrix is ​​non-empty and is also used to obtain the calculation window.

[0215] The traversal module 2552 is used to perform a sequence traversal on each sequence of the splicing weight template matrix based on the calculation window, wherein the sequence is a row or column of the splicing weight template matrix, and is also used to end the sequence traversal in response to a plurality of elements of the splicing weight template matrix located in the calculation window being empty during the traversal.

[0216] The splicing module 2553 is used to determine the attention matrix of the spliced ​​text based on the multiple sub-attention matrices in response to the completion of multiple sequence traversals of the splicing weight template matrix.

[0217] The calculation module 2554 is used to perform attention processing based on the multiple elements of the splicing weight template matrix located in the calculation window in response to at least one element being non-empty during traversal to obtain a sub-attention matrix.

[0218] In some embodiments, the computing module 2554 is further used to generate a sub-input matrix based on the multiple elements, wherein the sub-input matrix is ​​a sub-matrix of the splicing matrix; perform a linear transformation on the sub-input matrix to obtain a query matrix, a key matrix and a value matrix; perform weight-based attention processing on the query matrix and the key matrix to obtain a sub-weight matrix; perform weighted summation processing on the sub-weight matrix and the value matrix to obtain the sub-attention matrix.

[0219] In some embodiments, the calculation module 2554 is further used to perform weight-based similarity processing on the query matrix and the key matrix to obtain a similarity matrix; and perform normalization processing on the similarity matrix to obtain the sub-weight matrix.

[0220] In some embodiments, the acquisition module 2551 is also used to identify the splicing matrix based on the separator to obtain multiple sub-matrices corresponding to the text; based on the sub-matrices, determine the weight template matrix corresponding to the text; and combine the weight template matrices corresponding to the multiple texts to obtain the splicing weight template matrix.

[0221] In some embodiments, the traversal module 2552 is also used to determine the initial position of the calculation window and the sequence traversal direction; starting from the initial position of the calculation window, each sequence of the splicing weight template matrix is ​​traversed through the calculation window along the sequence traversal direction.

[0222] In some embodiments, the acquisition module 2551 is also used to determine the length of multiple input texts and obtain an input text length threshold; sort the multiple input texts in descending order according to the length of each of the input texts; obtain N adjacent input texts and N+1 adjacent input texts starting from the target text according to the sorting sequence of the input texts, wherein the length of the target text is less than the input text length threshold; in response to the sum of the lengths of the N adjacent input texts being less than or equal to the input text length threshold, and the sum of the lengths of the N+1 adjacent input texts being greater than the input text length threshold, taking the N adjacent input texts as the multiple texts, and splicing the N adjacent input texts to obtain the spliced ​​text; and obtaining a splicing matrix of the spliced ​​text based on the spliced ​​text.

[0223] In some embodiments, the acquisition module 2551 is also used to vectorize each of the texts to obtain text data of each of the texts; position encode the text data of each of the texts to obtain a text vector of each of the texts; and splice the text vectors of the multiple texts to obtain a splicing matrix of the spliced ​​texts.

[0224] In some embodiments, the traversal module 2552 is also used to obtain the text length of the spliced ​​text and determine the sliding window based on the text length; in response to the maximum text length among the lengths of the multiple texts being greater than the length of the sliding window, the splicing weight template matrix is ​​traversed based on the sliding window, and the elements outside the sliding window are marked as empty; and the splicing weight template matrix is ​​updated according to the traversal result.

[0225] In some embodiments, the traversal module 2552 is further used to obtain the number of the attention layer; and determine the sliding window based on the text length and the number of layers.

[0226] In some embodiments, the traversal module 2552 is also used to obtain position information of elements in the stitching weight template matrix, wherein the position information includes the row position and column position of the element in the stitching weight template matrix; determine the difference between the row position and the column position; and in response to the difference being greater than or equal to the size of the sliding window, mark the element as empty.

[0227] In some embodiments, the traversal module 2552 is further used to obtain the adjacent sequence of the current sequence in the traversed splicing weight template matrix; based on the calculation window, the adjacent sequence is subjected to attention processing to obtain a sub-attention matrix of the adjacent sequence.

[0228] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device executes the data processing method described above in the embodiment of the present application.

[0229] The present application embodiment provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the data processing method provided by the present application embodiment, for example, Figure 3A The data processing method is shown.

[0230] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0231] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and the computer executable instructions may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0232] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).

[0233] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.

[0234] It is understandable that in the embodiments of the present application, related data such as user information is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0235] In summary, the embodiment of the present application performs sequence traversal on each sequence of the splicing weight template matrix based on the calculation window, and in response to at least one element among the multiple elements in the calculation window being non-empty, performs attention processing on the multiple elements, and in response to the multiple elements being empty during the traversal, ends the traversal of the current sequence, which can improve the rate of attention calculation without affecting the calculation accuracy of the model. In addition, the embodiment of the present application also traverses and updates the splicing weight template matrix through a sliding window, further reducing the amount of attention calculation data, thereby further improving the model calculation speed.

[0236] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A data processing method, It is characterized in that The method comprises: Based on the splicing matrix of the spliced ​​text, determining the splicing weight template matrix of the spliced ​​text, wherein the spliced ​​text is obtained by splicing a plurality of texts, each of the texts includes at least one word, the splicing weight template matrix includes weight template matrices corresponding to the plurality of texts, and each element in the weight template matrix is ​​non-empty; Acquire a calculation window, and perform sequence traversal on each sequence of the splicing weight template matrix based on the calculation window, wherein the sequence is a row or a column of the splicing weight template matrix; In response to at least one element of the plurality of elements of the splicing weight template matrix located in the calculation window being non-empty during traversal, performing attention processing based on the plurality of elements to obtain a sub-attention matrix; In response to the fact that a plurality of elements of the splicing weight template matrix located in the calculation window are all empty during the traversal, ending the sequence traversal; In response to the completion of multiple sequence traversals of the splicing weight template matrix, an attention matrix of the spliced ​​text is determined based on multiple sub-attention matrices.

2. The method according to claim 1, It is characterized in that The attention processing is performed based on the multiple elements to obtain a sub-attention matrix, including: Based on the multiple elements, generate a sub-input matrix, wherein the sub-input matrix is ​​a sub-matrix of the spliced ​​matrix; Performing a linear transformation on the sub-input matrix to obtain a query matrix, a key matrix, and a value matrix; Performing weight-based attention processing on the query matrix and the key matrix to obtain a sub-weight matrix; The sub-weight matrix and the value matrix are weighted and summed to obtain the sub-attention matrix.

3. The method according to claim 2, It is characterized in that The weight-based attention processing is performed on the query matrix and the key matrix to obtain a sub-weight matrix, including: Performing weight-based similarity processing on the query matrix and the key matrix to obtain a similarity matrix; The similarity matrix is ​​normalized to obtain the sub-weight matrix.

4. The method according to claim 1, It is characterized in that The step of determining a splicing weight template matrix of the spliced ​​text based on the splicing matrix of the spliced ​​text includes: Identify the concatenated matrix based on the separator to obtain a plurality of sub-matrices corresponding to the texts; Based on the submatrix, determining a weight template matrix corresponding to the text; The weight template matrices corresponding to the multiple texts are combined to obtain the concatenated weight template matrix.

5. The method according to claim 1, It is characterized in that The step of performing sequence traversal on each sequence of the splicing weight template matrix based on the calculation window includes: Determining the initial position of the calculation window and the sequence traversal direction; Taking the initial position of the calculation window as the starting point, a sequence traversal is performed on each sequence of the splicing weight template matrix through the calculation window along the sequence traversal direction.

6. The method according to claim 1, It is characterized in that Before determining the splicing weight template matrix of the splicing text based on the splicing matrix of the splicing text, the method further includes: Determine the length of multiple input texts and obtain the input text length threshold; According to the length of each input text, the multiple input texts are sorted in descending order; According to the sorting sequence of the input text, obtain N adjacent input texts and N+1 adjacent input texts starting from the target text, wherein the length of the target text is less than the input text length threshold; In response to the sum of the lengths of the N adjacent input texts being less than or equal to the input text length threshold, and the sum of the lengths of the N+1 adjacent input texts being greater than the input text length threshold, taking the N adjacent input texts as the multiple texts, and performing splicing processing on the N adjacent input texts to obtain the spliced ​​text; Based on the concatenated text, a concatenated matrix of the concatenated text is obtained.

7. The method according to claim 1, It is characterized in that Before determining the splicing weight template matrix of the splicing text based on the splicing matrix of the splicing text, the method further includes: Performing vectorization processing on each of the texts to obtain text data of each of the texts; Performing position encoding on the text data of each of the texts to obtain a text vector for each of the texts; The text vectors of the multiple texts are spliced ​​to obtain a splicing matrix of the spliced ​​text.

8. The method according to claim 1, It is characterized in that After determining the splicing weight template matrix of the spliced ​​text based on the splicing matrix of the spliced ​​text, the method further includes: Acquire the text length of the concatenated text, and determine the sliding window based on the text length; In response to a maximum text length among the lengths of the multiple texts being greater than a length of the sliding window, traversing the splicing weight template matrix based on the sliding window, and marking elements outside the sliding window as empty; The splicing weight template matrix is ​​updated according to the traversal result.

9. The method according to claim 8, It is characterized in that The data processing method is implemented by an attention model, and the attention model includes a plurality of cascaded attention layers; The determining of the sliding window based on the text length comprises: Get the number of the attention layer; The sliding window is determined based on the text length and the number of layers.

10. The method according to claim 8, It is characterized in that The traversing the splicing weight template matrix based on the sliding window and marking the elements outside the sliding window as empty includes: Acquire position information of an element in the splicing weight template matrix, wherein the position information includes a row position and a column position of the element in the splicing weight template matrix; determining a difference between the row position and the column position; In response to the difference being greater than or equal to the size of the sliding window, marking the element as empty.

11. The method according to claim 1, It is characterized in that In response to the fact that a plurality of elements of the splicing weight template matrix located in the calculation window are all empty during the traversal, after the sequence traversal is ended, the method further comprises: Obtaining the adjacent sequence of the current sequence in the traversed splicing weight template matrix; Based on the calculation window, attention processing is performed on the adjacent sequence to obtain a sub-attention matrix of the adjacent sequence.

12. A data processing device, It is characterized in that The device comprises: An acquisition module is used to determine a splicing weight template matrix of the spliced ​​text based on a splicing matrix of the spliced ​​text, wherein the spliced ​​text is obtained by splicing a plurality of texts, each of the texts includes at least one word, and the splicing weight template matrix includes weight template matrices corresponding to the plurality of texts, each element in the weight template matrix is ​​non-empty and is also used to acquire a calculation window; A traversal module, configured to perform a sequence traversal on each sequence of the splicing weight template matrix based on the calculation window, wherein the sequence is a row or a column of the splicing weight template matrix, and to terminate the sequence traversal in response to a plurality of elements of the splicing weight template matrix located in the calculation window being empty during the traversal; A splicing module, configured to determine an attention matrix of the spliced ​​text based on a plurality of sub-attention matrices in response to completion of a plurality of sequence traversals of the splicing weight template matrix; A calculation module is used to perform attention processing based on the multiple elements of the splicing weight template matrix located in the calculation window to obtain a sub-attention matrix in response to at least one element being non-empty during traversal.

13. An electronic device, It is characterized in that The electronic device comprises: A memory for storing computer executable instructions; A processor, configured to implement the data processing method according to any one of claims 1 to 11 when executing the computer program or computer executable instructions stored in the memory.

14. A computer-readable storage medium, It is characterized in that A computer program or a computer executable instruction is stored, and when the computer program or the computer executable instruction is executed by a processor, the data processing method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising a computer program or computer executable instructions, It is characterized in that When the computer program or computer executable instructions are executed by a processor, the data processing method according to any one of claims 1 to 11 is implemented.