Task execution method, device, equipment, and storage medium used for large-scale language models

The task execution method for large-scale language models addresses inefficiencies in Transformer-based models by using sparse representations to determine and execute target attention tasks, enhancing processing efficiency and reducing computational overhead without sacrificing accuracy.

JP7778871B2Active Publication Date: 2025-12-02BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024125931
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-06-18
Filing Date
2024-08-01
Publication Date
2025-12-02
Estimated Expiration
2044-08-01

AI Technical Summary

Technical Problem

Current large-scale language models face inefficiencies due to high memory usage and computational overhead from traditional masking methods in the Transformer architecture, particularly in attention tasks, which impact processing efficiency and accuracy.

Method used

A task execution method for large-scale language models that determines and executes target attention tasks based on sparse representations of mask positions within non-intersecting intervals in the mask matrix, reducing unnecessary calculations by skipping complete mask regions and performing attention tasks only on incomplete regions.

Benefits of technology

This approach reduces memory usage and computational overhead while maintaining accuracy by efficiently representing mask shapes in different scenes, improving processing efficiency and reducing unnecessary calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778871000004
    Figure 0007778871000004
  • Figure 0007778871000005
    Figure 0007778871000005
  • Figure 0007778871000006
    Figure 0007778871000006
Patent Text Reader

Abstract

To provide a task execution method and device for a large language model, equipment and a storage medium.SOLUTION: A method comprises: determining a target attention task from multiple attention tasks to be processed by using a judgment unit on the basis of a sparse representation corresponding to a mask position of features to be processed, the target attention task being a task corresponding to a non-complete mask area of the to-be-processed feature, the sparse representation being used for representing the mask position of the to-be-processed feature, the mask positions representing mask endpoint positions in at least two mutually disjoint intervals in a mask matrix corresponding to the to-be-processed features; and executing the target attention task by using a computing unit to obtain attention features.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of artificial intelligence, particularly to technical fields such as deep learning, large-scale language models, natural language processing, and computer vision, and more particularly to a task execution method, device, apparatus, and storage medium used in large-scale language models. [Background technology]

[0002] A large-scale language model (LLM) is an advanced artificial intelligence algorithm trained on large amounts of data. Natural language processing systems with over 100 billion parameters can be used in content generation, text summarization, chatbots, program code creation, and AI applications such as predicting protein structures and biomolecular attributes.

[0003] Current large-scale language models mainly use the Transformer architecture to realize the feature processing process based on the attention mechanism. Summary of the Invention

[0004] The present disclosure provides a task execution method, apparatus, device, and storage medium for use with large-scale language models.

[0005] According to one aspect of the present disclosure, a task execution method for use in a large-scale language model is provided, including: determining a target attention task, which is a task corresponding to an incomplete mask region of the feature to be processed, from a plurality of attention tasks to be processed by a determination unit based on a sparse representation representing a mask position of the feature to be processed, which corresponds to the feature to be processed, wherein the mask position represents a mask endpoint position within at least two mutually non-intersecting intervals in a mask matrix corresponding to the feature to be processed; and executing the target attention task by a calculation unit to obtain an attention feature.

[0006] According to another aspect of the present disclosure, there is provided a task execution device for use in a large-scale language model, including a determination unit and a calculation unit. The determination unit determines a target attention task, which is a task corresponding to an incomplete mask region of the feature to be processed, from a plurality of attention tasks to be processed based on a sparse representation corresponding to the feature to be processed and representing a mask position of the feature to be processed, where the mask position represents a mask endpoint position within at least two mutually disjoint intervals in a mask matrix corresponding to the feature to be processed. The calculation unit executes the target attention task and obtains the attention feature.

[0007] According to another aspect of the present disclosure, there is provided a task execution device for use with large scale language models, the device including a task execution device for use with large scale language models according to the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively coupled to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method according to the present disclosure.

[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions causing a computer to perform a method according to the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a computer program product which, when executed by a processor, implements a method according to the present disclosure.

[0011] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following description. [Brief explanation of the drawings]

[0012] The drawings are for a better understanding of the invention and are not intended to limit the disclosure.

[0013] [Figure 1] FIG. 1 is a schematic diagram illustrating an exemplary system architecture of a task execution method and apparatus to which a large-scale language model according to an embodiment of the present disclosure can be applied. [Figure 2] FIG. 2 is a flowchart of a task execution method for use with a large-scale language model according to an embodiment of the present disclosure. [Figure 3A] FIG. 3A is a schematic diagram illustrating a mask diagram in a causal scene according to an embodiment of the present disclosure. [Figure 3B] FIG. 3B is a schematic diagram illustrating a mask diagram in a causal scene according to an embodiment of the present disclosure. [Figure 3C] FIG. 3C is a schematic diagram illustrating a mask diagram in a causal scene according to an embodiment of the present disclosure. [Figure 3D] FIG. 3D is a schematic illustration of a mask diagram in a causal scene according to an embodiment of the present disclosure. [Figure 4A] FIG. 4A is a schematic diagram illustrating a mask diagram in a non-causal scene according to an embodiment of the present disclosure. [Figure 4B] FIG. 4B is a schematic diagram illustrating a mask diagram in a non-causal scene according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is a schematic diagram illustrating a mask diagram in a complex scene according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a schematic diagram illustrating a block diagram of tasks to be processed according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a schematic diagram illustrating a method for determining a target task to be processed according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a schematic diagram illustrating a schematic diagram of performing attention calculation for a target processing task according to an embodiment of the present disclosure. [Figure 9] FIG. 9 is a block diagram of a task execution device used in a large-scale language model according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a block diagram of an electronic device suitable for implementing a task execution method for use with large-scale language models according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014]

[0023] The following description of exemplary embodiments of the present disclosure will be made with reference to the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included, but these are merely illustrative. Therefore, it should be understood that those skilled in the art can make various changes and modifications to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, the following description will omit descriptions of known functions and structures.

[0015] Transformer is the basis of large-scale language models and can address natural language problems in complex scenes. The Transformer's computational process relies on an attention mechanism. During the training process of a large-scale language model, it usually refers to different mask matrices depending on the training stage and training task. This allows the large-scale language model to selectively ignore masked features when performing attention tasks, thereby improving the processing performance of the large-scale language model.

[0016] In the relevant examples, the mask shape is generally [B, A, S, S], where B represents the batch size, A represents the number of heads, and S represents the length of the feature sequence. This type of masking method not only occupies video memory and memory access overhead that is quadratic in the hardware resources, but also results in a large amount of unnecessary calculations in the mask region when performing attention tasks, which impacts the processing efficiency of the model.

[0017] In another example, sparse or low-rank attention mechanisms employ coarse-grained masking schemes, which can reduce computational overhead but incur significant loss of model accuracy.

[0018] In view of this, in order to reduce the display memory occupied by a highly efficient mask and the memory access overhead of useless calculations, an embodiment of the present disclosure provides a task execution method used in large-scale language models, which includes: determining a target attention task, which is a task corresponding to an incomplete mask region of the feature to be processed, from a plurality of attention tasks to be processed by a determination unit based on a sparse representation representing a mask position of the feature to be processed, which corresponds to the feature to be processed; and executing the target attention task by a calculation unit to obtain an attention feature, wherein the mask position represents the mask endpoint position within at least two non-intersecting intervals in the mask matrix corresponding to the feature to be processed.

[0019] FIG. 1 schematically illustrates an exemplary system architecture of a task execution method and apparatus for use with a large-scale language model according to an embodiment of the present disclosure.

[0020] 1 is an example of a system architecture to which the embodiments of the present disclosure can be applied, and is intended to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenes. For example, in another embodiment, an exemplary system architecture of a task execution method and apparatus used for a large-scale language model may include a terminal device, but the terminal device can realize the task execution method and apparatus used for a large-scale language model provided by the embodiments of the present disclosure without needing to interact with a server.

[0021] 1, the system architecture according to this embodiment may include a terminal device 101, a network 102, and a server cluster 103. The network 102 is used as a medium for providing a communication link between the terminal device 101 and the server cluster 103. The network 102 may also be used as a medium for providing a communication link within the server cluster 103. The network 102 may include various connection types, such as wired and / or wireless communication links.

[0022] A user can use the terminal device 101 to interact with the server cluster 103 via the network 102 and receive or send messages, etc. For example, the terminal device 101 may send a request to train a deep learning model to the server cluster 103 via the network 102.

[0023] The terminal device 101 may have installed thereon various communication client applications such as (by way of illustrative examples only) a knowledge browsing application, a web browser application, a search application, an instant messaging tool, a mailbox client and / or social platform software.

[0024] The terminal device 101 may be any electronic device that has a display and supports web page browsing, including, but not limited to, a smartphone, a tablet computer, a laptop portable computer, a desktop computer, and the like.

[0025] The server cluster 103 may be a server that provides various services, for example, a background management server that provides support for requests sent by users using the terminal device 101 (by way of example only).

[0026] The server cluster 103 may be a cloud server, also called a cloud computing server or cloud host, which is a host product in a cloud computing service system, and solves the drawbacks of traditional physical hosts and VPS services (abbreviated as "Virtual Private Server" or "VPS"), such as high management difficulty and poor service scalability. The server may be a server in a distributed system or a server connected to a blockchain.

[0027] The server cluster 103 includes multiple server nodes 1031, 1032, 1033, and 1034, each including one or more hardware devices. The server cluster 103 or the server nodes can be used to execute the task execution method for large-scale models provided in the present disclosure, enabling deployment, inference, or training of large-scale models with low computational and storage resources.

[0028] Having described the system architecture of the present disclosure, the method of the present disclosure will now be described.

[0029] FIG. 2 shows a schematic flow chart of a task execution method for use with large-scale language models according to an embodiment of the present disclosure.

[0030] As shown in FIG. 2, the method 200 includes operations S210-S220.

[0031] In operation S210, a target attention task is determined from a plurality of attention tasks to be processed by a determining unit based on the sparse representation corresponding to the features to be processed.

[0032] In operation S220, a target attention task is performed by a computing unit to obtain attention features.

[0033] According to an embodiment of the present disclosure, the determining unit and the computing unit may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence computing unit, which may include at least one of a neural network processing unit (NPU), a tensor processing unit (TPU), and a Kunlun chip, etc.

[0034] According to an embodiment of the present disclosure, the sparse representation represents mask positions of a feature to be processed, and may be a representation of a mask shape, where the mask positions represent mask endpoint positions in at least two non-intersecting intervals in a mask matrix corresponding to the feature to be processed.

[0035] For example, the sparse representation can represent a mask with shape [B, A, S, n], where B represents the batch size, A represents the number of heads, S represents the length of the feature sequence, n = dk, k represents the number of non-intersecting intervals into which the feature sequence is separated in the S dimension, and d represents the number of elements for identifying the mask endpoint positions in each interval.

[0036] When d=1, it may indicate that there is one element for identifying the mask endpoint position within the section, and it may be the start endpoint position or the end endpoint position. When the mask endpoint position is identified by the start endpoint position, it indicates that the mask area is from the start endpoint to the area boundary line. When the mask endpoint position is identified by the end endpoint position, it indicates that the mask area is from the area boundary line to the end endpoint.

[0037] When d=2, it can be shown that there are two elements for identifying the mask endpoint positions within the interval, including both the start endpoint position and the end endpoint position.

[0038] FIG. 3A is a schematic illustration of a mask diagram in a causal scene according to an embodiment of the present disclosure.

[0039] FIG. 3B is a schematic illustration of a mask diagram in a causal scene according to an embodiment of the present disclosure.

[0040] In both Figures 3A and 3B, the mask matrix can be divided into a lower-left region and an upper-right region by a diagonal line. Since the upper-right region is entirely masked, only the masking status of the lower-left region can be represented. In the figures, S1 and E1 indicate the mask start and end points of the lower-left region, respectively, and S2 and E2 indicate the mask start and end points of the lower-left region, respectively. Shaded areas in the figures indicate unmasked regions, and light gray areas indicate masked regions.

[0041] As shown in FIG. 3A, in the mask schematic diagram 300A, in the lower left region, the mask area in the first column ends from the fourth row to the boundary line, so S1=4. Columns 0, 7, 8, and 9 are all non-masked areas, so S1=10, or even infinity. In such a scene, because each mask area ends at the boundary line, a perspective representation of the shape [B, A, S, 2] can be created.

[0042] As shown in Figure 3B, in the mask schematic diagram 300B, in the lower left region, the mask region in the 0th column ends from the 4th row to the 7th row, so S1 = 4 and E1 = 7. In such a scene, since the mask region does not end entirely at the boundary, a sparse representation can be performed in the shape of [B, A, S, 4].

[0043] Since the upper right regions of the mask schematic diagrams 300A and 300B are both masked, the sparse representation of the causal scene may be simplified to [B, A, S, 1] (corresponding to FIG. 3A) and [B, A, S, 2] (corresponding to FIG. 3B) to further reduce memory usage and access overhead due to the length of the sequence. Such simplified sparse representation is only applicable to causal scenes where the mask format is simple.

[0044] According to an embodiment of the present disclosure, the attention task may be a multi-head self-attention task corresponding to a large-scale language model, for example, the multiple attention tasks may correspond to modules that reason based on the multi-head self-attention mechanisms of multiple processing layers of the large-scale language model.

[0045] According to an embodiment of the present disclosure, the calculation formula based on the attention mechanism is as follows:

[0046]

number

[0047] Here, Q, K, and V represent the query matrix, key matrix, and value matrix, respectively, and each has the shape [B, S, A, H], where H represents the head size, and B, S, and A have the same meanings as above, so their explanations are omitted here. M represents the mask matrix.

[0048] When attention calculation is performed based on the above formula (1),

[0049]

number

[0050] When the intermediate features obtained by calculating

[0051] Therefore, this part of the early matrix operation is actually invalid and will instead cause extra computation overhead.The judgment unit determines whether the attention task to be currently processed is a task corresponding to a complete mask region, and skips the task corresponding to the complete mask region before executing the attention task, thereby reducing the computation overhead during the execution of the attention task.

[0052] According to an embodiment of the present disclosure, the target attention tasks may be tasks corresponding to incomplete mask regions of the feature to be processed, for example, tasks corresponding to partial mask regions and / or tasks corresponding to unmasked regions, and the computing unit then performs attention calculations only on the tasks corresponding to the incomplete mask regions to obtain attention features.

[0053] According to an embodiment of the present disclosure, based on a sparse representation corresponding to the feature to be processed, a judgment unit determines a task corresponding to an incomplete mask region of the feature to be processed from multiple attention tasks to be processed, so that before executing the attention task, the task corresponding to the complete mask region is skipped, and the calculation unit performs attention calculation only on the task corresponding to the incomplete mask region to obtain the attention feature, thereby reducing the calculation overhead during the execution of the attention task.

[0054] As the application scenarios of large-scale language models increase, how to accurately represent mask shapes in different scenes with spatiotemporal overhead linear to the length of the feature sequence can reduce memory usage and computational overhead without losing accuracy.

[0055] Therefore, the method according to the embodiment of the present disclosure further includes an operation of performing a sparse representation task using a sparse representation unit to perform sparse representation on the features to be processed based on the scene category corresponding to the task to be processed.

[0056] According to an embodiment of the present disclosure, scene categories may include causal scenes, non-causal scenes, and complex scenes. Causal scenes may be, for example, advertising item recommendation scenes. Non-causal scenes may be, for example, text recognition scenes. Complex scenes may be, for example, multimodal task identification scenes.

[0057] According to an embodiment of the present disclosure, performing sparse representation on features to be processed based on a scene category corresponding to a task to be processed may include, in response to the scene category being a causal scene, an operation of dividing a mask matrix along a diagonal of the mask matrix into two non-intersecting first and second intervals, in which all elements in the first interval are masked, and an operation of performing sparse representation on features to be processed using mask endpoint positions in the second interval.

[0058] According to an embodiment of the present disclosure, the mask endpoint locations may be the start mask row and the end mask row of each column element of the second interval in the mask matrix corresponding to the feature to be processed.

[0059] FIG. 3C is a schematic illustration of a mask diagram in a causal scene according to an embodiment of the present disclosure.

[0060] FIG. 3D is a schematic illustration of a mask diagram in a causal scene according to an embodiment of the present disclosure.

[0061] The meanings of the illustrations in Figures 3C and 3D are the same as those in Figures 3A and 3B described above, and therefore will not be described here. The first section may be the upper right region, and the second section may be the lower left region.

[0062] As shown in FIG. 3C , in the mask schematic diagram 300C, the starting mask rows in columns 0-9 are all row 5, the ending mask rows are all borders, and the sequence length S is 10. If the ending mask row needs to be indicated, the sparse representation may be [B, A, 10, 2]. If the ending mask row does not need to be indicated, the sparse representation may be [B, A, 10, 1], and the corresponding numerical values ​​may be [[5], [5], [5], [5], [5], [5], [5], [5], [5], [5]].

[0063] 3D , in the mask schematic diagram 300D, columns 0-3 and rows 7-9 are all unmasked, the starting mask rows in columns 4-6 are all row 7, and the ending mask rows are all boundary lines. The sequence length S is 10. If the ending mask row needs to be indicated, the sparse representation may be [B, A, 10, 2]. If the ending mask row does not need to be indicated, the sparse representation may be [B, A, 10, 1], and the corresponding numerical values ​​may be [

[10] ,

[10] ,

[10] ,

[10] , [7], [7],

[10] ,

[10] ,

[10] ].

[0064] According to an embodiment of the present disclosure, in response to the scene category being a non-causal scene, the mask matrix is ​​divided along the diagonal of the mask matrix into two non-intersecting intervals, a third interval and a fourth interval, and a sparse representation is performed on the feature to be processed using the mask endpoint positions in the third interval and the mask endpoint positions in the fourth interval.

[0065] According to an embodiment of the present disclosure, the mask endpoint positions may be the start mask row and the end mask row of each column element of the third section in the mask matrix corresponding to the feature to be processed, and the start mask row and the end mask row of each column element of the fourth region.

[0066] 4A-4B are schematic illustrations of mask diagrams in a non-causal scene according to an embodiment of the present disclosure.

[0067] The meanings of the illustrations in Figures 4A and 4B are the same as those in Figures 3A and 3B described above, and therefore will not be described here. The third section may be the upper right region, and the fourth section may be the lower left region.

[0068] As shown in Figure 4A, in the mask diagram 400A, in the lower left region, columns 0-2 are all unmasked areas, and the mask end rows of columns 3-9 are all boundary lines, with the starting mask rows being rows 3-9. In the upper right region, columns 0-2 are all unmasked areas, the mask in column 3 has already been shown in the lower left region, and the starting mask row of column 4 is 3 and the ending mask row is 4, so S2 = 3 and E2 = 4 in the corresponding columns.

[0069] As shown in FIG. 4B , in the mask schematic diagram 400B, in the lower left region, the starting mask row in columns 0-3 is 4, and the ending mask row is the boundary line. The starting mask row in columns 4-6 is 7, and the ending mask row is the boundary line. Columns 7-9 are all non-masked regions. In the upper right region, columns 0-3 are non-masked regions. The starting mask row in columns 4-6 is the boundary line, and the ending mask row is 4. The starting mask row in columns 7-9 is the boundary line, and the ending mask row is 7. Therefore, the sparse representation is [B, A, 10, 2], and the numerical values ​​are [[0, 4], [0, 4], [0, 4], [0, 4], [4, 7], [4, 7], [7, 10], [7, 10], [7, 10]].

[0070] In the two scenes above, the mask shape is relatively simple, and accurate representation of the mask shape can be achieved by dividing the mask matrix into only two regions. However, for the mask schematic diagram of a complex scene shown in Figure 5, accurate representation of the mask shape can be achieved by dividing more non-intersecting sections.

[0071] FIG. 5 is a schematic diagram illustrating a mask diagram in a complex scene according to an embodiment of the present disclosure.

[0072] 5, in the mask schematic diagram 500, every two rows are one section, the hatched areas indicate unmasked areas, and the lined gray areas indicate masked areas. S1-S5 indicate the start mask row of each column element in each section, and E1-E5 indicate the end mask row of each column element in each section.

[0073] For example, in the 0th column, in the k1 section, everything is an unmasked area. In the k2 section, the starting mask row is 2, the ending mask row is 4 (the ending mask row is the boundary of the section, so the display of the ending row may be omitted), and the corresponding columns S2 = 2 and E2 = 4. In the k3 section, everything is an unmasked area. In the k4 section, the starting mask row is 6, the ending mask row is 8, and the corresponding columns S4 = 6 and E4 = 8. In the k5 section, everything is an unmasked area. Therefore, the sparse representation may be [B, A, 10, 10], and the corresponding numerical values ​​are shown in Figure 5.

[0074] According to embodiments of the present disclosure, accurate sparse representation of mask shapes in different scenes can reduce memory usage and computational overhead without loss of accuracy, with a linear spatio-temporal overhead in feature sequence length.

[0075] When a large-scale language model executes a task, the number of parameters involved in processing is in the hundreds of billions. To improve the processing efficiency of the model, the features can be divided into blocks and then processed in parallel on multiple distributed GPUs.

[0076] Therefore, the task execution method provided by the embodiments of the present disclosure for use in large-scale language models further includes: executing blocking tasks by a blocking unit, and blocking the parameter matrices corresponding to the features to be processed based on the length of the parameter matrices corresponding to the features to be processed and the number of registers, to obtain parameter matrices including a query matrix, a key matrix, a value matrix, and a mask matrix corresponding to each attention task to be processed; and storing the query matrix, key matrix, value matrix, and mask matrix corresponding to each attention task to be processed by a target storage unit.

[0077] According to an embodiment of the present disclosure, first, based on the length of the parameter matrix corresponding to the features to be processed and the number of registers, hyperparameters for executing the blocking task are determined, and then, based on the matrix multiplication rule, the parameter matrix corresponding to the features to be processed is blocked to obtain a parameter matrix corresponding to each attention task to be processed.

[0078] For example, the shape of the query matrix may be [2, 8, 1024, 128] and can be divided into parameter matrices of [2, 8, 64, 128] or [2, 8, 128, 128]. The shape of the query matrix may be [2, 8, 1024, 256] and can be divided into parameter matrices of [2, 8, 32, 256] or [2, 8, 64, 256] or [2, 8, 128, 256] or [2, 8, 256, 256].

[0079] FIG. 6 is a schematic diagram illustrating a block diagram of tasks to be processed according to an embodiment of the present disclosure.

[0080] As shown in FIG. 6, the Query matrix, Key matrix, Value matrix, and Mask matrix can all be divided into four blocks, and each block includes three elements.

[0081] According to an embodiment of the present disclosure, the processing efficiency of a large-scale language model can be effectively improved by blocking the parameter matrix corresponding to the feature to be processed based on the length of the parameter matrix corresponding to the feature to be processed and the number of registers.

[0082] According to an embodiment of the present disclosure, based on the attention calculation mechanism, the matrix dimension obtained by multiplying the transpose of the Query matrix and the Key matrix is ​​the same as the dimension of the Mask matrix, so the blocking dimension of the Query matrix and the Key matrix determines the blocking dimension of the Mask matrix.

[0083] Therefore, by determining a mask section corresponding to multiple attention tasks to be processed, and determining whether the attention tasks to be processed are invalid tasks, the computational overhead occupied by invalid tasks in the process of executing the attention tasks can be reduced.

[0084] According to an embodiment of the present disclosure, determining a target attention task from a plurality of attention tasks to be processed by a judgment unit based on a sparse representation corresponding to the features to be processed includes: determining a mask interval corresponding to the plurality of attention tasks to be processed based on the sparse representation corresponding to the features to be processed; and determining a target attention task from the plurality of attention tasks to be processed by a judgment unit based on the mask interval.

[0085] According to an embodiment of the present disclosure, the mask interval can represent the area between the mask start end point and the mask end end point, thereby determining which task intermediate calculation result matrix in the attention task to be processed is within the mask interval, and determining that the attention task to be processed located within the mask interval is an invalid task before executing the attention task.

[0086] Since the blocking is performed based on the feature matrices corresponding to the attention task of the same head in the same batch, we will omit the display of B and A here. For example, since the block of the Query matrix is ​​[3, 128], the block of the Key matrix is ​​[3, 128], and Q*K^T=[3, 3], we can determine that the block of the Mask matrix is ​​[3, 3].

[0087] Therefore, each [3, 3] block is traversed to determine whether it is located in the mask interval, and whether the corresponding attention task to be processed is an invalid task.

[0088] FIG. 7 is a schematic diagram illustrating a method for determining a target task to be processed according to an embodiment of the present disclosure.

[0089] As shown in Figure 7, in example 700, the non-shaded areas are mask areas, and 701 indicates 12 tokens included in all attention tasks to be processed for a certain head of a certain batch. All areas in the [3, 3] block 7012 corresponding to attention task B to be processed are mask areas, so it can be determined that attention task B to be processed is an invalid task. Only some areas in the [3, 3] block 7011 corresponding to attention task A to be processed are mask areas, so it can be determined that attention task A to be processed is the target attention task.

[0090] According to an embodiment of the present disclosure, invalid tasks can be screened based on the mask interval, and the invalid tasks can be skipped before executing the attention task, thereby reducing the computation overhead occupied by the invalid tasks.

[0091] According to an embodiment of the present disclosure, determining mask intervals corresponding to a plurality of attention tasks to be processed based on a sparse representation corresponding to the features to be processed may include, for each attention task to be processed, determining a plurality of mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on the sparse representation corresponding to the features to be processed, and determining a mask interval corresponding to each attention task to be processed based on the plurality of mask endpoint positions.

[0092] According to an embodiment of the present disclosure, each attention task to be processed may be a blocked attention task.

[0093] Because the mask shapes of different application scenarios are different, in order to adapt to the needs of different application scenarios, the mask position may be the start mask row and the end mask row of the elements of each column in at least two mutually non-intersecting sections in the mask matrix corresponding to the features to be processed.

[0094] According to an embodiment of the present disclosure, determining mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed may include determining a start mask row and an end mask row of each column element in the mask matrix corresponding to each attention task to be processed based on the sparse representation corresponding to the feature to be processed.

[0095] According to an embodiment of the present disclosure, determining a mask interval corresponding to each attention task to be processed based on multiple mask endpoint positions includes determining the end mask row of the element in each column as the end position of the mask interval, and determining the start mask row of the element in each column as the start position of the mask interval.

[0096] As shown in FIG. 7, the shape of the sparse representation may be [B, A, 12, 4], and the corresponding numerical values, in order, are as follows: the starting mask row of the bottom left region is [12, 5, 5, 5, 6, 6, 9, 9, 9, 12, 12, 12]; the ending mask row of the bottom left region is [12, 11, 11, 11, 11, 11, 11, 11, 12, 12, 12]; the starting mask row of the top right region is [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]; and the ending mask row of the bottom right region is [0, 1, 2, 2, 3, 3, 6, 6, 6, 12, 12, 12].

[0097] According to an embodiment of the present disclosure, for a block corresponding to task A to be processed, the block is in the lower left section, and the corresponding column in the section is 0-2. The end mask row [12, 11, 11] of the corresponding column can be the end position of the mask section, and the start mask row [12, 5, 5] of the corresponding column can be the start position of the mask section.

[0098] According to an embodiment of the present disclosure, for a block corresponding to task B to be processed, the block is located in the lower left section, and corresponds to columns 3-5 within the section. The end mask row [12, 11, 11] of the corresponding column can be the end position of the mask section, and the start mask row [5, 6, 6] of the corresponding column can be the start position of the mask section.

[0099] According to an embodiment of the present disclosure, the mask position is the start mask column and the end mask column of each row element in at least two non-intersecting intervals in the mask matrix corresponding to the feature to be processed.

[0100] According to an embodiment of the present disclosure, determining mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed may include determining a start mask column and an end mask column of each row element in the mask matrix corresponding to each attention task to be processed based on the sparse representation corresponding to the feature to be processed.

[0101] According to an embodiment of the present disclosure, determining a mask interval corresponding to each attention task to be processed based on multiple mask endpoint positions may include determining the end mask column of each row element as the end position of the mask interval, and determining the start mask column of each row element as the start position of the mask interval.

[0102] The method of determining a mask interval by indicating the mask position using the start mask column and end mask column of each row element is the same as the method of determining a mask interval by indicating the mask position using the start mask row and end mask row of each column element, so its explanation will be omitted here.

[0103] According to an embodiment of the present disclosure, since the sparse representation accurately represents the shape of the mask, the mask interval can be quickly determined based on the maximum (ending mask row / column) and minimum (starting mask row / column) values ​​of the corresponding mask rows / columns within the task block to be processed, thereby reducing the computational overhead when determining invalid tasks.

[0104] According to an embodiment of the present disclosure, determining a target attention task from a plurality of attention tasks to be processed by a judgment unit based on a mask interval includes determining the attention task to be processed as the target attention task by the judgment unit in response to an element endpoint position in an intermediate feature matrix corresponding to the attention task to be processed not being within the mask interval, and the intermediate feature matrix is ​​obtained based on a query matrix and a key matrix corresponding to the attention task to be processed.

[0105] According to an embodiment of the present disclosure, it is possible to determine whether the attention task to be processed is a target attention task by calculating the intersection between the element endpoint position and the mask interval. If there is an intersection between the element endpoint position and the mask interval and if the element endpoint position is entirely within the mask interval, it is possible to determine that the task to be processed is an invalid task. If there is a partial intersection between the element endpoint position and the mask interval or if there is no intersection, it is possible to determine that the task to be processed is a target attention task.

[0106] For example, if the element endpoint position in the intermediate feature matrix corresponding to a task to be processed is [3, 4] and the mask interval is [5, 11], it can be determined that the two intervals have no common part, and the attention task to be processed can be determined to be the target attention task.

[0107] According to an embodiment of the present disclosure, the mask interval includes a mask end position and a mask start position. In response to the element end point position in the intermediate feature matrix corresponding to the attention task to be processed not being within the mask interval, determining the attention task to be processed as the target attention task by the determining unit may include an operation of determining the attention task to be processed as the target attention task by the determining unit in response to the element end point position in the intermediate feature matrix corresponding to the attention task to be processed being greater than the mask end position or smaller than the mask start position.

[0108] According to an embodiment of the present disclosure, the intermediate feature matrix is ​​obtained as shown in Equation (1):

[0109]

number

[0110] It can be.

[0111] As shown in FIG. 7, for a block corresponding to task A to be processed, the end position of the mask section corresponding to that block is [12, 11, 11], and the mask start position is [12, 5, 5].

[0112] According to the embodiment of the present disclosure, the block corresponding to task A to be processed is located in rows 6-8 of columns 0-2, and only rows 6-8 of columns 1-2 are located within the mask interval, where 6 is smaller than 12, 7 is located between 5-11, and 8 is between 5-11. Therefore, the elements in the intermediate feature matrix corresponding to task A to be processed are partially located in the mask area and can be determined as the target attention task.

[0113] According to an embodiment of the present disclosure, for a block corresponding to task B to be processed, the end position of the mask interval corresponding to that block is [12, 11, 11] and the start position is [5, 6, 6].

[0114] According to an embodiment of the present disclosure, the block corresponding to task B to be processed is located in rows 6-8 of columns 3-5, where 6 is between 5-12, 7 is between 6-11, and 8 is between 6-11. Therefore, the elements in the intermediate feature matrix corresponding to task B to be processed are all located within the mask interval, and the blocks corresponding to task B to be processed are all located in the mask region, and can be determined as invalid tasks.

[0115] According to an embodiment of the present disclosure, by comparing the mask interval within a block with the element endpoint task, the judgment logic is simple and efficient, and invalid tasks can be quickly determined within linear time, and the invalid tasks can be skipped, thereby reducing the computational overhead in the execution process of the attention task.

[0116] According to an embodiment of the present disclosure, performing a target attention task by a computing unit and obtaining attention features may include reading at least one query matrix, at least one key matrix, at least one value matrix, and at least one mask matrix corresponding to the target attention task from a target memory unit by the computing unit; and performing the target attention task by the target computing unit and obtaining attention features based on the at least one query matrix, at least one key matrix, at least one value matrix, and at least one mask matrix.

[0117] FIG. 8 is a schematic diagram illustrating a schematic diagram of performing attention calculation for a target processing task according to an embodiment of the present disclosure.

[0118] As shown in FIG. 8, in embodiment 800, when performing a target attention task, the target attention task can be performed by a target calculation unit based on at least one query matrix 803, at least one key matrix 802, at least one value matrix 807 and at least one mask matrix 801, to obtain attention features 806.

[0119] For example, a first intermediate feature 8041 can be obtained based on the transpose of the query matrix 8031 ​​and the key matrix 8021, a second intermediate feature 8051 can be obtained based on the first intermediate feature 8041 and the mask matrix 8011, the second intermediate feature 8051 can be processed using an activation function to obtain an activation feature matrix, and an attention feature 8061 can be obtained based on the activation feature matrix and the key matrix 8071.

[0120] According to an embodiment of the present disclosure, the activation function may be a normalization function.

[0121] According to the embodiment of the present disclosure, the computation unit performs computation only for the target attention task, reducing the attention computation process corresponding to the invalid task, and further reducing the computation overhead of performing the attention task.

[0122] FIG. 9 illustrates a block diagram of a task execution device used for a large-scale language model according to an embodiment of the present disclosure.

[0123] As shown in FIG. 9, the apparatus 900 may include a determining unit 910 and a calculating unit 920 .

[0124] The determination unit 910 determines a target attention task, which is a task corresponding to a non-complete mask region of the feature to be processed, from the plurality of attention tasks to be processed based on a sparse representation representing a mask position of the feature to be processed, where the mask position represents a mask endpoint position within at least two non-intersecting intervals in the mask matrix corresponding to the feature to be processed.

[0125] The computation unit 920 performs a target attention task and obtains attention features.

[0126] According to an embodiment of the present disclosure, the determination unit includes a first determination subunit and a second determination subunit, the first determination subunit determines a mask interval corresponding to a plurality of attention tasks to be processed based on a sparse representation corresponding to a feature to be processed, and the second determination subunit determines a target attention task from the plurality of attention tasks to be processed by the determination unit based on the mask interval.

[0127] According to an embodiment of the present disclosure, for each attention task to be processed, the first determination subunit determines multiple mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the features to be processed, and determines a mask interval corresponding to each attention task to be processed based on the multiple mask endpoint positions.

[0128] According to an embodiment of the present disclosure, the mask position is a start mask row and an end mask row of each column element in at least two non-intersecting intervals in a mask matrix corresponding to the feature to be processed, and the first determination subunit determines a start mask row and an end mask row of each column element in the mask matrix corresponding to each attention task to be processed based on the sparse representation corresponding to the feature to be processed.

[0129] According to an embodiment of the present disclosure, the first determining subunit determines the end mask row of each column element as the end position of the mask interval, and determines the start mask row of each column element as the start position of the mask interval.

[0130] According to an embodiment of the present disclosure, the mask position is a start mask column and an end mask column of each row element in at least two non-intersecting intervals in a mask matrix corresponding to a feature to be processed, and a first determination subunit determines a start mask column and an end mask column of each row element in the mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed.

[0131] According to an embodiment of the present disclosure, the first determining subunit determines the end mask column of each row element as the end position of the mask interval, and determines the start mask column of each row element as the start position of the mask interval.

[0132] According to an embodiment of the present disclosure, the second determination subunit determines the attention task to be processed as a target attention task by the judgment unit in response to the element endpoint position in the intermediate feature matrix corresponding to the attention task to be processed being not within the mask interval, and the intermediate feature matrix is ​​obtained based on the query matrix and key matrix corresponding to the attention task to be processed.

[0133] According to an embodiment of the present disclosure, the mask section includes a mask end position and a mask start position, and the second determination subunit determines the attention task to be processed by the judgment unit as the target attention task in response to the element end point position in the intermediate feature matrix corresponding to the attention task to be processed being greater than the mask end position or smaller than the mask start position.

[0134] According to an embodiment of the present disclosure, the computing unit includes a reading subunit and a computing subunit, wherein the reading subunit reads at least one query matrix, at least one key matrix, at least one value matrix, and at least one mask matrix corresponding to a target attention task from the target storage unit through the computing unit, and the computing subunit executes the target attention task through the target computing unit based on the at least one query matrix, at least one key matrix, at least one value matrix, and at least one mask matrix to obtain attention features.

[0135] According to an embodiment of the present disclosure, the calculation subunit obtains first intermediate features based on the transpose of the query matrix and the key matrix, obtains second intermediate features based on the first intermediate features and the mask matrix, processes the second intermediate features using an activation function to obtain an activation feature matrix, and obtains attention features based on the activation feature matrix and the key matrix.

[0136] According to an embodiment of the present disclosure, the apparatus further includes a blocking unit and a target storage unit.

[0137] The blocking unit executes the blocking task to block the parameter matrix corresponding to the feature to be processed according to the length of the parameter matrix corresponding to the feature to be processed and the number of registers to obtain the parameter matrix corresponding to each attention task to be processed, the parameter matrix including a query matrix, a key matrix, a value matrix and a mask matrix. The target storage unit stores the query matrix, the key matrix, the value matrix and the mask matrix corresponding to each attention task to be processed.

[0138] According to an embodiment of the present disclosure, the apparatus further includes a sparse representation unit for performing a sparse representation task, such as performing sparse representation on features to be processed based on a scene category corresponding to the task to be processed.

[0139] According to an embodiment of the present disclosure, the sparse representation unit includes a first division subunit and a first display subunit.

[0140] The first division subunit divides the mask matrix into two non-intersecting first and second intervals along a diagonal of the mask matrix in response to the scene category being a causal scene, and all elements in the first interval are masked. The first display subunit performs sparse representation for the feature to be processed using the mask endpoint positions in the second interval.

[0141] According to an embodiment of the present disclosure, the sparse representation unit includes a second division subunit and a second display subunit.

[0142] The second division subunit, in response to the scene category being a non-causal scene, divides the mask matrix at the diagonal of the mask matrix into two third and fourth intervals that do not intersect with each other.

[0143] The second representation sub-unit uses the mask end point positions in the third interval and the mask end point positions in the fourth interval to perform a sparse representation for the feature to be processed.

[0144] According to an embodiment of the present disclosure, the present disclosure further provides a task execution device for use with a large-scale language model, including the above-described device.

[0145] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.

[0146] According to an embodiment of the present disclosure, an electronic device includes at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, such that the at least one processor can perform the above-described method.

[0147] According to an embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions causing a computer to perform the above-described method.

[0148] According to an embodiment of the present disclosure, a computer program, when executed by a processor, implements the above method.

[0149] 10 shows an exemplary block diagram for implementing an example electronic device 1000 according to an embodiment of the present disclosure. The electronic device is intended to represent various types of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the implementation of the present disclosure as described and / or claimed herein.

[0150] 10, the device 1000 includes a computing unit 1001, which can perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 1002 or loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 can further store various programs and data necessary for the operation of the device 1000. The computing unit 1001, the ROM 1002, and the RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0151] Multiple components in device 1000 are connected to an I / O interface 1005, including an input unit 1006 such as a keyboard, a mouse, etc., an output unit 1007 such as various types of displays, speakers, etc., a storage unit 1008 such as a magnetic disk, an optical disk, etc., and a communication unit 1009 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 enables device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0152] The computing unit 1001 may be various general-purpose and / or specialized processing modules having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various machine learning model algorithm computing units, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc.

[0153] The computing unit 1001 executes the above-described methods and processes, such as the task execution method used for large-scale language models. For example, in some embodiments, the task execution method used for large-scale language models may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, it may perform one or more steps of the above-described task execution method used for large-scale language models. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the task execution method used for large-scale language models in any other suitable form (e.g., via firmware).

[0154] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be embodied in one or more computer programs that can be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special purpose or general purpose programmable processor, and that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0155] Program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program code may be entirely executed on a machine, partially executed on a machine, partially executed on a machine as a separate software package and partially executed on a remote machine, or entirely executed on a remote machine or server.

[0156] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program for use in or in connection with an instruction execution system, device, or appliance. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or appliance, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include an electrical connection of one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0157] To provide interaction with a user, a computer may implement the systems and techniques described herein and include a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices may also provide interaction with a user; for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and may receive input from the user in any form (including voice input, speech input, or tactile input).

[0158] The systems and techniques described herein can be implemented in a computing system including background components (e.g., a data server), or middleware components (e.g., an application server), or front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such background, middleware, or front-end components. The components of the system can be connected to each other by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include, by way of example, a local area network (LAN), a wide area network (WAN), and the Internet.

[0159] The computer system may include a client and a server. The client and server are generally remote and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the corresponding computers and having a client-server relationship. The server may be a cloud server, a server in a distributed system, or a server in combination with a blockchain.

[0160] It should be understood that various types of flows shown above may be used, and steps may be rearranged, added, or deleted. For example, the steps described in the present invention may be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present disclosure can be achieved, and the present specification is not limited thereto.

[0161] The specific embodiments described above do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. A determining unit determines a target attention task from a plurality of attention tasks to be processed, the target attention task being a task corresponding to an incomplete mask region of the feature to be processed, based on a sparse representation corresponding to the feature to be processed and representing a mask position of the feature to be processed, wherein the mask position represents a mask endpoint position within at least two non-intersecting intervals in a mask matrix corresponding to the feature to be processed; performing the target attention task by a computing unit to obtain attention features; Determining a target attention task from a plurality of attention tasks to be processed by a determination unit based on a sparse representation corresponding to the features to be processed includes: determining mask intervals corresponding to the plurality of attention tasks to be processed based on sparse representations corresponding to the features to be processed; determining a target attention task from a plurality of attention tasks to be processed by the determining unit based on the mask section; determining a target attention task from a plurality of attention tasks to be processed by the determining unit based on the mask section, determining, by the determining unit, the attention task to be processed as the target attention task in response to an element endpoint position in the intermediate feature matrix corresponding to the attention task to be processed not being within the mask interval; The intermediate feature matrix is ​​obtained based on a query matrix and a key matrix corresponding to the attention task to be processed. Task execution methods used for large-scale language models.

2. Determining mask intervals corresponding to the plurality of attention tasks to be processed based on sparse representations corresponding to the features to be processed includes: For each attention task to be processed, determine a plurality of mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed; determining a mask section corresponding to each attention task to be processed based on the plurality of mask end point positions. The method of claim 1.

3. the mask positions are a start mask row and an end mask row of each column element in at least two non-intersecting intervals in a mask matrix corresponding to the feature to be processed; Determining mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed includes: determining a start mask row and an end mask row of each column element in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed; The method of claim 2.

4. determining a mask section corresponding to each attention task to be processed based on the plurality of mask end point positions, determining an end mask row of each of the column elements as an end position of the mask section; determining a starting mask row of each of the columns as a starting position of the mask interval. The method of claim 3.

5. the mask positions are start and end mask columns of each row element in at least two non-intersecting intervals in the mask matrix corresponding to the feature to be processed; Determining mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed includes: determining a start mask column and an end mask column of each row element in the mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed; The method of claim 2.

6. determining a mask section corresponding to each attention task to be processed based on the plurality of mask end point positions, determining an end mask column of each row element as an end position of the mask section; determining a starting mask column of each row element as a starting position of the mask interval. The method of claim 5.

7. The mask section includes a mask end position and a mask start position, determining, by the determining unit, the attention task to be processed as the target attention task in response to an element endpoint position in the intermediate feature matrix corresponding to the attention task to be processed not being within the mask interval; determining the attention task to be processed as the target attention task by the determining unit in response to an element end point position in the intermediate feature matrix corresponding to the attention task to be processed being greater than the mask end position or less than the mask start position. The method of claim 1.

8. Executing the target attention task by a computing unit and obtaining attention features includes: Reading, by a calculation unit, at least one query matrix, at least one key matrix, at least one value matrix, and at least one mask matrix corresponding to the target attention task from a target storage unit; performing the target attention task by a target computing unit based on the at least one query matrix, the at least one key matrix, the at least one value matrix, and the at least one mask matrix to obtain the attention features. The method according to any one of claims 1 to 6.

9. Executing the target attention task by a target computing unit based on the at least one query matrix, the at least one key matrix, the at least one value matrix, and the at least one mask matrix to obtain the attention features includes: obtaining first intermediate features based on a transpose of the query matrix and the key matrix; obtaining second intermediate features based on the first intermediate features and the mask matrix; processing the second intermediate features using an activation function to obtain an activation feature matrix; and obtaining the attention feature based on the activation feature matrix and the key matrix. The method of claim 8.

10. The method comprises: Execute a blocking task by a blocking unit, and according to the length of the parameter matrix corresponding to the feature to be processed and the number of registers, block the parameter matrix corresponding to the feature to be processed to obtain a parameter matrix including a query matrix, a key matrix, a value matrix and a mask matrix corresponding to each attention task to be processed; and storing, by a target storage unit, the query matrix, the key matrix, the value matrix, and the mask matrix corresponding to each attention task to be processed. The method according to any one of claims 1 to 6.

11. The method comprises: and performing a sparse representation task corresponding to the feature to be processed by a sparse representation unit, and performing the sparse representation on the feature to be processed based on a scene category corresponding to the attention task to be processed. The method according to any one of claims 1 to 6.

12. The sparse representation of the feature to be processed based on a scene category corresponding to the attention task to be processed includes: In response to the scene category being a causal scene, dividing the mask matrix at a diagonal of the mask matrix into two non-intersecting intervals, a first interval and a second interval, and all elements in the first interval are masked; and performing the sparse representation on the feature to be processed using mask endpoint locations within the second interval. The method of claim 11.

13. The sparse representation of the feature to be processed based on a scene category corresponding to the attention task to be processed includes: In response to the scene category being a non-causal scene, dividing the mask matrix at a diagonal of the mask matrix into two non-intersecting third and fourth intervals; performing the sparse representation on the feature to be processed using mask endpoint locations within the third interval and mask endpoint locations within the fourth interval. The method of claim 11.

14. A task execution device for use in a large-scale language model, comprising: a determining unit for determining a target attention task from a plurality of attention tasks to be processed based on a sparse representation of a mask position of the feature to be processed, the target attention task being a task corresponding to an incomplete mask region of the feature to be processed, the mask position representing a mask endpoint position within at least two non-intersecting intervals in a mask matrix corresponding to the feature to be processed; a computing unit for performing the target attention task and obtaining attention features; The determination unit a first determining subunit for determining mask intervals corresponding to the plurality of attention tasks to be processed based on sparse representations corresponding to the features to be processed; a second determination subunit for determining a target attention task from a plurality of attention tasks to be processed by the determination unit based on the mask section; The second determination subunit: In response to an element endpoint position in the intermediate feature matrix corresponding to the attention task to be processed not being within the mask interval, the determining unit determines the attention task to be processed as the target attention task; The intermediate feature matrix is ​​obtained based on a query matrix and a key matrix corresponding to the attention task to be processed. Task execution device used for large-scale language models.

15. The first determination subunit: For each attention task to be processed, determine a plurality of mask endpoint positions in a mask matrix corresponding to each attention task to be processed based on a sparse representation corresponding to the feature to be processed; A mask section corresponding to each attention task to be processed is determined based on the positions of the plurality of mask end points.

15. The apparatus of claim 14.

16. The mask position is a start mask row and an end mask row of elements of each column in at least two non-intersecting intervals in a mask matrix corresponding to the feature to be processed, and the first determining subunit: Based on the sparse representation corresponding to the features to be processed, a start mask row and an end mask row of each column element in the mask matrix corresponding to each attention task to be processed are determined.

16. The apparatus of claim 15.

17. The first determination subunit: determining the end mask row of each of the column elements as the end position of the mask section; The starting mask row of each column element is determined as the starting position of the mask section.

17. The apparatus of claim 16.

18. The mask position is a start mask column and an end mask column of each row element in at least two non-intersecting intervals in a mask matrix corresponding to the feature to be processed, and the first determining subunit: Based on the sparse representation corresponding to the features to be processed, a start mask column and an end mask column of each row element in the mask matrix corresponding to each attention task to be processed are determined.

16. The apparatus of claim 15.

19. The first determination subunit: determining an end mask column of each row element as an end position of the mask section; The start mask column of each row element is determined as the start position of the mask section.

20. The apparatus of claim 18.

20. The mask section includes a mask end position and a mask start position, The second determination subunit: In response to an element end point position in the intermediate feature matrix corresponding to the attention task to be processed being greater than the mask end position or less than the mask start position, the determination unit determines the attention task to be processed as the target attention task.

15. The apparatus of claim 14.

21. The computing unit a reading sub-unit for reading at least one query matrix, at least one key matrix, at least one value matrix and at least one mask matrix corresponding to the target attention task from the target storage unit by a calculation unit; a computing subunit for performing the target attention task by a target computing unit based on the at least one query matrix, the at least one key matrix, the at least one value matrix, and the at least one mask matrix to obtain the attention features.

20. Apparatus according to any one of claims 14 to 19.

22. The computation subunit: Obtaining first intermediate features based on a transpose of the query matrix and the key matrix; obtaining second intermediate features based on the first intermediate features and the mask matrix; processing the second intermediate features using an activation function to obtain an activation feature matrix; Obtaining the attention feature based on the activation feature matrix and the key matrix.

22. The apparatus of claim 21.

23. The device comprises: a blocking unit that performs a blocking task, such that, according to the length of a parameter matrix corresponding to a feature to be processed and the number of registers, the parameter matrix corresponding to the feature to be processed is blocked to obtain a parameter matrix including a query matrix, a key matrix, a value matrix, and a mask matrix corresponding to each attention task to be processed; a target storage unit for storing the query matrix, the key matrix, the value matrix, and the mask matrix corresponding to each attention task to be processed.

20. Apparatus according to any one of claims 14 to 19.

24. The device comprises: and a sparse representation unit that performs a sparse representation task corresponding to a feature to be processed and performs the sparse representation for the feature to be processed based on a scene category corresponding to the attention task to be processed.

20. Apparatus according to any one of claims 14 to 19.

25. The sparse representation unit is a first division subunit that, in response to the scene category being a causal scene, divides the mask matrix along a diagonal of the mask matrix into two non-intersecting first and second intervals, and all elements in the first interval are masked; a first representation subunit for performing the sparse representation on the feature to be processed using mask endpoint positions within the second interval.

25. The apparatus of claim 24.

26. The sparse representation unit is a second division subunit that divides the mask matrix into two non-intersecting third and fourth intervals along a diagonal of the mask matrix in response to the scene category being a non-causal scene; a second representation subunit for performing the sparse representation on the feature to be processed using mask endpoint positions within the third interval and mask endpoint positions within the fourth interval.

25. The apparatus of claim 24.

27. 20. A method for manufacturing a device comprising: Task execution equipment used for large-scale language models.

28. at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method of any one of claims 1 to 6. electronic equipment.

29. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to carry out the method of any one of claims 1 to 6. A non-transitory computer-readable storage medium.

30. A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Language interaction method and device, communication equipment and storage medium

    CN116932728A

  • Automatic test equipment

    JP1994161814A

  • Methods and devices for accelerating a transformer with a sparse attention pattern

    US20230133305A1