Parallelization strategy optimization method and system, device, and medium
By replacing the scaling factor of the softmax function with the maximum preset fixed value in the parallelization strategy of the large language model and performing optimization operations, the additional overhead problem caused by synchronous update of the attention mechanism calculation pipeline is solved, and a more efficient parallelization strategy operation is achieved.
Patent Information
- Application Number
- PCT/CN2024/125585
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-10-17
- Publication Date
- 2025-05-08
AI Technical Summary
In the inference process of existing large language models, the attention mechanism calculation pipeline adopts partial softmax operations, resulting in nearly 20% of the additional overhead, because the calculation results of each part need to be synchronized and updated.
By replacing the scaling factor in the softmax function with the maximum preset fixed value, the softmax function is optimized, and exponential operation and sequence and operation processing are performed in the parallelization strategy, the matrix multiplication operation is completed and the results are corrected to realize parallelization strategy optimization.
It improves the parallelization strategy computing efficiency of large language models, reduces computational overhead, and improves inference performance.
Smart Images

Figure CN2024125585_08052025_PF_FP_ABST
Abstract
Description
A parallel strategy optimization method, system, device and medium
[0001] Cross-references
[0002] This application claims priority to Chinese patent application No. 202311456221.5 filed on November 3, 2023, the entire contents of which are incorporated by reference in their entirety into this application. Technical Field
[0003] The present invention belongs to the field of deep learning technology, and specifically relates to a parallelization strategy optimization method, system, device and medium. Background Art
[0004] Large language models are deep learning models trained using large amounts of text data. They can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, question-answering, and conversation, and are an important path to artificial intelligence.
[0005] As large language models become increasingly important in various fields, the performance of large language model inference is crucial for large-scale large language model applications. Numerous works have been conducted to optimize large language model inference. As shown in Figure 2, a large language model composed of a Transformer can be divided into two stages: Prefill and Decode. The main difference between the two stages lies in the size of the input matrix Q. The data flow executed is similar, consisting of multiple Transformer layers. Each layer can be divided into linear operations and attention mechanism operations. The attention mechanism operation includes two general matrix multiplications and one softmax operation.
[0006] In the process of inference of large language models, in order to improve the parallelism of calculations and reduce the overhead of reading and writing data, the existing work FlashAttention changed the original overall calculation method in the process of calculating the attention mechanism, as shown in Figure 3(a). It chose to split the attention matrix and then perform partial softmax calculation on each part, as shown in Figure 3(b). Therefore, the calculation process needs to complete the synchronization of current information and past information, and complete the update operation of existing results.
[0007] Currently, in the large language model inference calculation flow, the attention mechanism calculation pipeline has the following problems: the current common attention mechanism calculation pipeline uses partial softmax operation, which uses partial matrix data to calculate the results. Because the data obtained in each part is different, it is necessary to synchronize information and update the results between the calculation results of each part. This partial softmax synchronous update calculation will result in nearly 20% additional overhead.
[0008] Therefore, a parallelization strategy optimization method that can improve the operational efficiency of parallelization strategies for large language models is desired.
[0009] Summary of the Invention
[0010] In response to the problems existing in the prior art, the present invention provides a parallelization strategy optimization method, system, device and medium, which at least partially solve the problems existing in the prior art.
[0011] In a first aspect, an embodiment of the present disclosure provides a parallelization strategy optimization method, comprising the following steps:
[0012] Replace the scaling factor in the softmax function in the parallelization strategy with the maximum preset fixed value to obtain the optimized softmax function;
[0013] The optimized softmax function is processed by exponential operation and sequence operation in parallel. After the exponential operation is completed, the matrix multiplication operation is performed, and the result of the sequence operation is used to correct the result of the matrix multiplication operation to complete the parallelization strategy optimization.
[0014] According to a specific implementation of the embodiment of the present disclosure, the softmax function in the parallelization strategy is:
[0015] Where x is the input data; is a scaling factor, which is a maximum preset fixed value; R is a real number; i is the number of input data; x i is the i-th input data; e is the Napier constant; x d is the dth input data.
[0016] According to a specific implementation of the embodiment of the present disclosure, the process of obtaining the maximum preset fixed value is:
[0017] Execute the model inference record preprocessing phase multiple times to get the input data of the softmax function;
[0018] Analyze the statistical distribution of the input data to obtain a maximum preset fixed value, which satisfies:
[0019] Most of the input data counted by this model do not meet the following requirements: i >>Maximum preset fixed value Or input data x i <<Maximum preset fixed value situation.
[0020] According to a specific implementation of an embodiment of the present disclosure, the majority of the input data counted by the model is 99.99% of the input data.
[0021] According to a specific implementation of the embodiment of the present disclosure, the value range of the maximum preset fixed value is:
[0022] -100<maximum preset fixed value
[0023] According to a specific implementation method of the embodiment of the present disclosure, after the matrix multiplication operation result is corrected using the sequence and operation processing result, an inner loop operation processing is performed to optimize the softmax function result and the feature matrix.
[0024] According to a specific implementation of the embodiment of the present disclosure, the inner loop operation processing is to perform an optimized softmax function operation processing on the feature vector of each sample in the feature matrix to obtain the probability distribution of the sample.
[0025] According to a specific implementation of the embodiment of the present disclosure, during the inner loop operation processing, the input data of the optimized softmax function and the feature matrix are processed asynchronously separately.
[0026] According to a specific implementation of the embodiment of the present disclosure, there is outer accumulation in the inner loop operation processing, and after all partial vectors are processed, the outer accumulation processing is performed.
[0027] According to a specific implementation of the embodiment of the present disclosure, the characteristic matrix is a V matrix, and the inner loop operation processing process is:
[0028] Where x is the input data; is a scaling factor, which is a maximum preset fixed value; R is a real number; For input data x (j) The i-th dimension of the vector data; x i is the i-th dimension of the input vector; e is the Napier constant; x d is the dth dimension of the input vector; p is the input data x (j) The number of vectors; j is the jth vector of input data; d / p is x (j) The number of dimensions of the vector; is the i-th dimension of the j-th column vector in the V matrix; For input data The result of scaling and exponential operation.
[0029] According to a specific implementation of the embodiment of the present disclosure, in the process of the inner loop operation processing, without loss of generality, it is assumed that each x i of like or When x i The asynchronous softmax calculation of the vector x belongs to it, and then the synchronous softmax method is used to recalculate the value of the optimized softmax function.
[0030] In a second aspect, an embodiment of the present disclosure provides a parallelization strategy optimization system, the system comprising:
[0031] The preprocessing unit is configured as
[0032] Replace the scaling factor in the softmax function in the parallelization strategy with the maximum preset fixed value to obtain the optimized softmax function;
[0033] Output unit, configured as
[0034] The optimized softmax function is processed by exponential operation and sequence operation in parallel. After the exponential operation is completed, the matrix multiplication operation is performed, and the result of the sequence operation is used to correct the result of the matrix multiplication operation to complete the parallelization strategy optimization.
[0035] The present disclosure also provides an electronic device, including:
[0036] at least one processor; and,
[0037] a memory communicatively connected to the at least one processor; wherein,
[0038] The memory stores instructions that can be executed by the at least one processor. When the instructions are executed by the at least one processor, the at least one processor executes the method for parallelization strategy optimization in the aforementioned first aspect or any implementation of the first aspect.
[0039] In a fourth aspect, an embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by at least one processor, the at least one processor executes the method for parallelization strategy optimization in the aforementioned first aspect or any implementation of the first aspect.
[0040] In the fifth aspect, an embodiment of the present disclosure also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the method for parallelization strategy optimization in the aforementioned first aspect or any implementation of the first aspect.
[0041] Other optional features and technical effects of the embodiments of the present invention are partially described below, and partially can be understood by reading this document.
[0042] Compared with the prior art, the present invention has the following beneficial technical effects:
[0043] The present invention provides a parallelization strategy optimization method, system, device and medium, comprising the following steps: replacing the scaling factor in the softmax function in the parallelization strategy with a maximum preset fixed value to obtain an optimized softmax function; performing exponential operation processing and sequence sum operation processing on the optimized softmax function in parallel, wherein matrix multiplication operation processing is performed after the exponential operation processing is completed, and the result of the sequence sum operation processing is used to correct the result of the matrix multiplication operation processing to complete the parallelization strategy optimization; the present application can improve the parallelization strategy operation efficiency of large language models. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. The elements shown are not limited to the scale shown in the drawings. The same or similar reference numerals in the drawings represent the same or similar elements, wherein:
[0045] FIG1 is a flow chart of a parallelization strategy optimization method according to an embodiment of the present disclosure;
[0046] FIG2 is a schematic diagram of a large language model reasoning calculation process in the prior art;
[0047] FIG3 is a schematic diagram showing a comparison of different softmax calculation methods in the prior art;
[0048] FIG4 is a schematic diagram of a method flow diagram of a maximum preset fixed value according to an embodiment of the present disclosure;
[0049] FIG5 is a schematic diagram of a process of recalculating all partial vectors according to an embodiment of the present disclosure;
[0050] FIG6 is a diagram showing the effect of asynchronous softmax improvement in the Prefill stage according to an embodiment of the present disclosure;
[0051] FIG7 is a diagram showing the effect of asynchronous softmax improvement in the Decode stage according to an embodiment of the present disclosure;
[0052] FIG8 is a parallelization strategy optimization system according to an embodiment of the present disclosure; and
[0053] FIG9 is a diagram of a parallelization strategy optimization device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0054] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0055] The following describes the embodiments of the present disclosure through specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0056] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.
[0057] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0058] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.
[0059] In order to solve the problem that existing circuit simulation methods cannot give a good balance between simulation accuracy and simulation speed, the present invention proposes a method for modular circuit behavior simulation.
[0060] The method for modular circuit behavior simulation proposed in the present invention includes two parts: a software interface for modeling hardware behavior and a program for executing the simulation process, wherein the software interface for modeling hardware behavior uses tasks as basic units, and each task includes a start event and an end event.
[0061] Next, the parallelization strategy optimization method, system, device and medium according to the embodiments of the present disclosure will be described with reference to Figures 1 to 9.
[0062] FIG1 shows a parallelization strategy optimization method 100 according to the present embodiment. As shown in FIG1 , the method includes the following steps: at step S101 , the scaling factor in the softmax function in the parallelization strategy is replaced with a maximum preset fixed value to obtain an optimized softmax function.
[0063] In an embodiment of the present invention, the softmax function in the parallelization strategy is:
[0064] Where x is the input data; is a scaling factor, which is a maximum preset fixed value; R is a real number; i is the number of input data; x i is the i-th input data; e is the Napier constant; x d is the dth input data.
[0065] In an embodiment of the present invention, a method 200 for obtaining the maximum preset fixed value is shown in FIG4 , and includes the following steps:
[0066] At step S210, the model inference record preprocessing stage is executed multiple times to obtain input data of the softmax function.
[0067] Next, go to step S220.
[0068] In step S220, the statistical distribution of the input data is analyzed to obtain a maximum preset fixed value, which satisfies:
[0069] Most of the input data counted by this model do not meet the following requirements: i >>Maximum preset fixed value Or input data x i <<Maximum preset fixed value situation.
[0070] In an embodiment of the present invention, the majority of input data counted by the model is 99.99% of the input data.
[0071] In an embodiment of the present invention, the maximum preset fixed value has a value range of:
[0072] -100<maximum preset fixed value
[0073] Next, go to step S120.
[0074] At step S120, the optimized softmax function is subjected to exponential operation processing and sequence sum operation processing in parallel, wherein the exponential operation processing is followed by matrix multiplication operation processing, and the sequence sum operation processing result is used to correct the matrix multiplication operation processing result, thereby completing the parallelization strategy optimization; it should be noted that in matrix multiplication, we usually use the sequence sum operation result to correct the matrix multiplication operation result. This process can be regarded as first performing a conventional matrix multiplication, and then correcting the multiplication result according to the result of the sequence sum operation; specifically, assuming that we have two matrices A and B, and we want to calculate A*B, after performing conventional matrix multiplication, we will obtain a preliminary result C, and then we will correct C with the result of a sequence sum operation to obtain the final multiplication result D; this correction process can effectively improve the accuracy and stability of matrix multiplication, especially when processing large-scale and complex data.
[0075] In an embodiment of the present invention, after the matrix multiplication operation result is corrected using the sequence and operation processing result, an inner loop operation processing is performed to optimize the softmax function result and the feature matrix.
[0076] In an embodiment of the present invention, the inner loop operation processing is to perform an optimized softmax function operation processing on the feature vector of each sample in the feature matrix to obtain the probability distribution of the sample.
[0077] In an embodiment of the present invention, during the inner loop operation processing, the input data of the optimized softmax function and the feature matrix are all asynchronously processed separately; the asynchronous processing is a processing method that does not need to wait for the processing to be completed before performing subsequent operations. The core logic of asynchronous processing is that it does not block the current thread to wait for the processing to be completed, but allows subsequent operations until other threads complete the processing and call back to notify this thread. This processing method is similar to SMS communication, and there is no need to remain in a waiting state after sending a message; for example, in programming, asynchronous processing can be used to process long-running operations, such as network requests or file IO operations, to improve the program's responsiveness and concurrency. In natural language processing, asynchronous processing can be used for computationally intensive tasks such as training language models to fully utilize computing resources and improve training efficiency.
[0078] In an embodiment of the present invention, there is an outer accumulation during the inner loop operation processing, and the outer accumulation is performed after all partial vectors are processed. It should be noted that the outer accumulation usually refers to an accumulation operation on an external variable during the loop or iteration process. This external variable is usually used to calculate a certain accumulation sum, or is used to calculate the sum of all loop iterations after the loop ends. In each iteration of the loop, the external variable will be accumulated with an internal variable in the loop, thereby gradually increasing the value of the external variable. When the loop ends, the value of the external variable is the accumulation sum of all loop iterations. The outer accumulation is usually used to count the number of times an event occurs, or to calculate the sum of a variable in the loop. This accumulation operation can easily calculate the sum of the results of all iterations in the loop, and can avoid repeated accumulation operations within the loop, thereby improving the efficiency and readability of the code.
[0079] In an embodiment of the present invention, the characteristic matrix is a V matrix, and the inner loop operation processing process is:
[0080] Where x is the input data; is a scaling factor, which is a maximum preset fixed value; R is a real number; For input data x (j) The i-th dimension of the vector data; x i is the i-th dimension of the input vector; e is the Napier constant; x d is the dth dimension of the input vector; p is the input data x (j) The number of vectors; j is the jth vector of input data; d / p is x (j) The number of dimensions of the vector; is the i-th dimension of the j-th column vector in the V matrix; For input data The result of scaling and exponential operation.
[0081] In the embodiment of the present invention, during the inner loop operation process, without loss of generality, it is assumed that each x i of like or When x i The asynchronous softmax calculation of the vector x belongs to it, and then the synchronous softmax method is used to recalculate the value of the optimized softmax function.
[0082] It should be noted that the purpose of optimizing the softmax function and the V matrix for inner loop operations is to obtain a set of probability distributions. The softmax function is a commonly used function that can map any real number to a value between [0,1], and these values add up to 1, so it can be interpreted as a probability distribution. The V matrix is usually a feature matrix, and each row represents the feature vector of a sample. Therefore, the inner loop operation of the softmax function and the V matrix can be understood as: performing a softmax function operation on the feature vector of each sample to obtain the probability distribution of the sample. This probability distribution can be used to represent the probability of the sample belonging to each category, thereby providing a basis for subsequent tasks such as classification or clustering.
[0083] In an embodiment of the present invention, FIG5 shows an example of a parallelization strategy optimization method; a=-3, b=3 are preset. The two vectors x and y are given by Q·K T Calculated and divided into two local vectors; at the same time, the T To these local vectors. For each x i ,have Use the first partial vector of x and Process the first part of the vector x. There are two asynchronous threads, each thread performs the corresponding calculations, respectively:
[0084] and
[0085] The two threads synchronize after processing all partial vectors and perform the final division operation. For y, the first partial vector is processed in a similar way, but Then both threads will be terminated and the first thread will recalculate all partial vectors based on the calculation results.
[0086] FIG5 shows the process of recalculating all partial vectors in this embodiment:
[0087] Figure 5(a) shows that each partial softmax result is processed separately without synchronous updates, and Figure 5(b) shows that when overflow occurs, all partial softmax calculations need to be recalculated.
[0088] The experimental results in the embodiments of the present invention are:
[0089] In this embodiment, the optimization of the softmax function can also be called an asynchronous softmax solution, which can be applied to both the Prefill stage and the Decode stage. The proposed solution is tested against the most advanced attention implementation solution. TM The test results on the GPU are shown in Figures 6 and 7. In the Prefill stage, the proposed scheme achieves an average speedup of 1.52 times and 1.19 times compared with xformers[5] and FlashAttetion2, respectively. In the Decode stage, the proposed scheme outperforms the xformers implementation of customized decoding, which is represented as xformers-decoder in Figure 8. In the case of long context, it is 2.02 times faster than the existing technology FlashDecoding.
[0090] FIG8 shows a parallel strategy optimization system 300 provided by the present invention. The system 300 includes: a preprocessing unit 310 and an output unit 320.
[0091] The preprocessing unit 310 is configured to replace the scaling factor in the softmax function in the parallelization strategy with a maximum preset fixed value to obtain an optimized softmax function;
[0092] The output unit 320 is configured to perform exponential operation processing and sequence sum operation processing on the optimized softmax function in parallel, wherein matrix multiplication operation processing is performed after the exponential operation processing is completed, and the result of the sequence sum operation processing is used to correct the result of the matrix multiplication operation processing to complete the parallelization strategy optimization.
[0093] FIG9 shows a schematic diagram of an electronic device 1000 that can implement a method or implement an embodiment of the present invention. In some embodiments, more or fewer electronic devices may be included than shown. In some embodiments, the method can be implemented using a single or multiple electronic devices. In some embodiments, the method can be implemented using cloud-based or distributed electronic devices.
[0094] As shown in Figure 9, the electronic device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to the programs and / or data stored in the read-only memory (ROM) 1002 or the programs and / or data loaded from the storage part 1008 into the random access memory (RAM) 1003. The processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 1001 can include a general-purpose main processor and one or more special coprocessors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are also stored. The processor 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0095] The processor and memory are used together to execute the program stored in the memory. When the program is executed by the computer, the methods, steps or functions described in the above embodiments can be implemented.
[0096] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, a touch screen, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk and the like; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is installed in the drive 1010 as needed, so that a computer program read therefrom can be installed into the storage section 1008 as needed. FIG. 9 schematically illustrates only some of the components, and does not mean that the computer system 1000 includes only the components shown in FIG.
[0097] The systems, devices, modules, or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smartphone, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.
[0098] Although not shown, in an embodiment of the present invention, a storage medium is provided, wherein the storage medium stores a computer program, and the computer program is configured to execute any file difference-based compilation method according to any embodiment of the present invention when executed.
[0099] Storage media in embodiments of the present invention include permanent and non-permanent, removable and non-removable items that can be used to store information using any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0100] The methods, programs, systems, and apparatuses of the embodiments of the present invention may be executed or implemented in a single or multiple networked computers, or may be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks may be performed by remote processing devices connected via a communication network.
[0101] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, those skilled in the art will appreciate that the functional modules / units or controllers and related method steps described in the above embodiments may be implemented using software, hardware, or a combination of software / hardware.
[0102] Unless explicitly stated, the actions or steps of the methods, procedures, and methods described in accordance with the embodiments of the present invention do not have to be performed in a specific order and can still achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0103] In this document, multiple embodiments of the present invention are described, but for the sake of brevity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" are intended to apply to at least one embodiment or example according to the present invention, but not all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. Those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually contradictory.
[0104] While the exemplary systems and methods of the present invention have been specifically shown and described with reference to the foregoing embodiments, these are merely examples of the best modes for implementing the present systems and methods. Those skilled in the art will appreciate that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A parallel strategy optimization method, characterized in that: The following steps are involved: Replace the scaling factor in the softmax function in the parallelization strategy with the maximum preset fixed value to obtain the optimized softmax function; The optimized softmax function is processed by exponential operation and sequence and operation in parallel, wherein matrix multiplication operation is performed after the exponential operation is completed, and the result of the sequence and operation is used to correct the result of the matrix multiplication operation to complete the parallelization strategy optimization.
2. The parallelization strategy optimization method according to claim 1, characterized in that: The softmax function in the parallelization strategy is: Where x is the input data; is a scaling factor, which is a maximum preset fixed value; R is a real number; i is the number of input data; x i is the i-th input data; e is the Napier constant; x d is the dth input data.
3. The parallelization strategy optimization method according to claim 1, characterized in that: The process of obtaining the maximum preset fixed value is: Execute the model inference record preprocessing phase multiple times to get the input data of the softmax function; Analyze the statistical distribution of the input data to obtain a maximum preset fixed value, which satisfies: Most of the input data counted by this model do not satisfy: Input data x i >>Maximum preset fixed value Or input data x i <<Maximum preset fixed value situation.
4. The parallelization strategy optimization method according to claim 3, characterized in that: The majority of the input data counted by the model is 99.99% of the input data.
5. The parallelization strategy optimization method according to claim 1, characterized in that: The maximum preset fixed value has a value range of: -100 < maximum preset fixed value 6. The parallelization strategy optimization method according to claim 1, characterized in that: After the matrix multiplication operation result is corrected by using the sequence and operation processing result, an inner loop operation processing of optimizing the softmax function result and the feature matrix is performed.
7. The parallelization strategy optimization method according to claim 6, characterized in that: The inner loop operation process is to perform an optimized softmax function operation process on the feature vector of each sample in the feature matrix to obtain the probability distribution of the sample.
8. The parallelization strategy optimization method according to claim 6, characterized in that: During the inner loop operation processing, the input data of the optimized softmax function and the feature matrix are processed asynchronously separately.
9. The parallelization strategy optimization method according to claim 6, characterized in that: There is an outer accumulation during the inner loop operation processing, and after all the partial vectors are processed, the outer accumulation processing is performed.
10. The parallelization strategy optimization method according to claim 6, characterized in that: The characteristic matrix is a V matrix, and the inner loop operation processing process is: Where x is the input data; is a scaling factor, which is a maximum preset fixed value; R is a real number; For the input data x (j) The i-th dimension of the vector data; x i is the i-th dimension of the input vector; e is the Napier constant; x d is the dth dimension of the input vector; p is the input data x (j) The number of vectors; j is the jth vector of the input data; d / p is x (j) The number of dimensions of the vector; is the i-th dimension of the j-th column vector in the V matrix; For input data The result of scaling and exponential operation.
11. The parallelization strategy optimization method according to claim 10, characterized in that: In the inner loop operation process, without loss of generality, it is assumed that each x i of like or When x i The asynchronous partial softmax calculation of the vector x belongs to, and then the synchronous softmax method is used to recalculate the value of the optimized softmax function.
12. A parallel strategy optimization system, characterized in that: The parallelization strategy optimization method according to any one of claims 1 to 11 comprises: The preprocessing unit is configured as Replace the scaling factor in the softmax function in the parallelization strategy with the maximum preset fixed value to obtain the optimized softmax function; Output unit, configured as The optimized softmax function is processed by exponential operation and sequence and operation in parallel, wherein matrix multiplication operation is performed after the exponential operation is completed, and the result of the sequence and operation is used to correct the result of the matrix multiplication operation to complete the parallelization strategy optimization.
13. A computer device, the electronic device comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor executes the parallelization strategy optimization method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, which, when executed by at least one processor, cause the at least one processor to perform the parallelization strategy optimization method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the parallelization strategy optimization method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Softmax function approximate calculation method and device
CN115222033A
Parallelization strategy optimization method, system, device and medium
CN117407793A
Attention neural networks with locality-sensitive hashing
US10909461B1
Efficient softmax computation
US20220067513A1
Computer-Implemented Method of Executing SoftMax
US20220383077A1