Data processing method, electronic device, storage medium and computer program product

By obtaining data sequence length information, selectively calling multiple threads to calculate the original input data of the large language model, the problems of high computing resource consumption and low training efficiency caused by inconsistent data sequence length are solved, and the optimization of computing resources and the improvement of training efficiency are achieved.

CN120429099APending Publication Date: 2025-08-05HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410160446.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-04
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

When training a large language model, due to inconsistent data sequence lengths, it requires expansion, resulting in high computing resources and low training efficiency. The existing technology has failed to effectively solve this problem.

Method used

By obtaining data sequence length information, multiple threads are selectively called to calculate the original input data, reducing redundant calculations, and optimizing the attention calculation process.

Benefits of technology

It reduces the computing resource consumption caused by expansion operations, improves model training efficiency, and reduces computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429099A_ABST
    Figure CN120429099A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, electronic equipment, a storage medium and a computer program product. The method comprises the steps that to-be-processed data is obtained, and the to-be-processed data is data obtained after data sequence length expansion is conducted on original input data; based on pre-recorded data sequence length information, selecting original input data from the to-be-processed data, the data sequence length information being used for recording an initial sequence length corresponding to the original input data; a plurality of threads on the processor of the preset type are called to execute target calculation on the original input data, a target calculation result is obtained, and the number of the threads is determined by the data sequence length information. The technical problems of high computing resource consumption and low training efficiency caused by performing redundant attention calculation on the data matrix constructed by the expansion operation in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of large models. Specifically, it relates to a data processing method, an electronic device, a storage medium, and a computer program product. Background Art

[0002] With the development of large language models, the data sequence length in the input data during model training is getting larger and larger. However, the computational complexity during model training is proportional to the quadratic equation of the sequence length. Therefore, the problem of large consumption of computing power and memory resources caused by the increase in data sequence length is becoming increasingly serious.

[0003] On this basis, the lengths of multiple data sequences in the input data during model training may also be different. The commonly used solution in related technologies is to expand (padding) the shorter data sequences to a specific length to ensure that the lengths of each data sequence in the input data are the same, which is convenient for completing the fine-tuning operation in model training. Obviously, the above expansion operation of data sequences not only cannot improve the training efficiency or training effect of model training, but also brings additional computational and memory access overhead to the model training process.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide a data processing method, an electronic device, and a storage medium to at least solve the technical problem of large consumption of computing resources and low training efficiency caused by redundant attention calculations on the data matrix constructed by the expansion operation in related technologies.

[0006] According to one aspect of the embodiments of this application, a data processing method is provided, including: obtaining data to be processed, where the data to be processed is data obtained by expanding the data sequence length of the original input data; selecting the original input data from the data to be processed based on the pre-recorded data sequence length information, where the data sequence length information is used to record the initial sequence length corresponding to the original input data; calling multiple threads on a preset type of processor to perform a target calculation on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0007] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining a data processing request through a first application programming interface, wherein the request data carried in the data processing request includes: data to be processed, and the data to be processed is data obtained by expanding the data sequence length of the original input data; returning a data processing response through a second application programming interface, wherein the response data carried in the data processing response includes: a target calculation result, and the target calculation result is obtained by invoking multiple threads on a preset type of processor to perform a target calculation on the original input data. The number of threads of the multiple threads is determined by the data sequence length information, and the original input data is selected from the data to be processed based on the data sequence length information. The data sequence length information is used to record the initial sequence length corresponding to the original input data.

[0008] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining a current input data processing dialogue request, wherein the request data carried in the data processing dialogue request includes: data to be processed, and the data to be processed is data obtained by expanding the data sequence length of the original input data; in response to the data processing dialogue request, returning a data processing dialogue reply, wherein the information carried in the data processing dialogue reply includes: a target calculation result, and the target calculation result is obtained by invoking multiple threads on a preset type of processor to perform a target calculation on the original input data. The number of threads of the multiple threads is determined by the data sequence length information, and the original input data is selected from the data to be processed based on the data sequence length information. The data sequence length information is used to record the initial sequence length corresponding to the original input data; displaying the target calculation result in a graphical user interface.

[0009] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory storing an executable program; a processor for running the program, wherein when the program runs, it executes the data processing method of any one of the above.

[0010] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, and the computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the data processing method of any one of the above.

[0011] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the data processing method of any one of the above.

[0012] In an embodiment of the present application, obtain data to be processed, where the data to be processed is data obtained by expanding the data sequence length of the original input data; based on the pre-recorded data sequence length information, select the original input data from the data to be processed, where the data sequence length information is used to record the initial sequence length corresponding to the original input data; call multiple threads on a preset type of processor to perform target calculations on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0013] It is easy to note that in an embodiment of the present application, according to the initial sequence length corresponding to the original input data, perform attention calculations on part of the data in the data to be processed (i.e., the original input data), reducing the redundant calculations performed on the expanded data in the data to be processed. Thus, the present application achieves the purpose of reducing the calculation consumption by performing attention calculations on the original input data in the data to be processed specifically through the data sequence length information, thereby realizing the technical effects of reducing the redundant calculations brought by the expansion operation during the fine-tuning process, reducing the consumption of computing resources, and improving the model training efficiency, and further solving the technical problem in the related art that redundant attention calculations are performed on the data matrix constructed by the expansion operation, resulting in large consumption of computing resources and low training efficiency.

[0014] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0016] Figure 1 is a schematic diagram of an application scenario of a data processing method according to Embodiment 1 of the present application;

[0017] Figure 2 is a schematic diagram of an attention calculation process according to the related art;

[0018] Figure 3 is a flowchart of a data processing method according to Embodiment 1 of the present application;

[0019] Figure 4 is a schematic diagram of an optional attention calculation process according to Embodiment 1 of the present application;

[0020] Figure 5 is a schematic diagram of another optional attention calculation process according to Embodiment 1 of the present application;

[0021] Figure 6is the flowchart of a data processing method according to Embodiment 2 of the present application;

[0022] Figure 7 is the flowchart of a data processing method according to Embodiment 3 of the present application;

[0023] Figure 8 is the structural schematic diagram of a data processing device according to Embodiment 4 of the present application;

[0024] Figure 9 is the structural schematic diagram of another data processing device according to Embodiment 4 of the present application;

[0025] Figure 10 is the structural schematic diagram of yet another data processing device according to Embodiment 4 of the present application;

[0026] Figure 11 is the structural block diagram of an electronic device according to Embodiment 5 of the present application. Detailed implementation manners

[0027] To enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] The technical solution provided by this application is mainly implemented using large model technology. Here, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one hundred trillion model parameters. A large model can also be called a foundation model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with over hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-trained models, etc.

[0030] It should be noted that in actual applications, the large model can be fine-tuned with a small amount of samples on the pre-trained model, enabling the large model to be applied to different tasks. For example, the large model can be widely applied in fields such as natural language processing (NLP), computer vision, etc. Specifically, it can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., and can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0031] First, some nouns or terms that appear during the description of the embodiments of this application are subject to the following explanations.

[0032] Graphics Processing Unit (GPU): Also known as a graphics card, it is a dedicated processor for graphics processing and data parallel computing, and can be used to accelerate the training and inference of large models.

[0033] Large Language Model (LLM): Refers to an artificial intelligence language model with a large number of parameters and a complex structure.

[0034] Fine tune: Refers to the process of further training with a small amount of data on the basis of a pre-trained model to make the model more adaptable to specific task or domain requirements.

[0035] Data batch: It refers to the unit of input data in model training. A data batch contains multiple data sequences. For example, a data batch of 4 means it contains 4 data sequences.

[0036] Padding: It refers to the process of expanding data sequences to a specified length. Since a data batch contains multiple data sequences and the lengths of each data sequence may be different, however, the data input into the model needs to be constructed in the data format of a multi-dimensional array matrix (Tensor), and the lengths of each data sequence in the multi-dimensional array matrix need to be the same. Therefore, during the training of large models, multiple data sequences are expanded to the same length through padding.

[0037] Attention mechanism: It refers to an artificial intelligence model or algorithm that simulates the characteristics of human cognitive attention, enabling the model to focus on key information and ignore unimportant parts when processing input data. This mechanism enables the model to process a large amount of information more effectively and perform better in the learning and reasoning processes.

[0038] Softmax function: A function used to convert the original numerical values output by the model into a probability distribution, commonly used in multi-classification tasks.

[0039] Convolution kernel: In a convolutional neural network, the convolution kernel refers to the filter used for convolution operations, which can extract features from input data. In this solution, the convolution kernel is also used to initiate computing tasks to the GPU.

[0040] Thread: In this solution, it refers to the smallest unit that executes computing tasks in GPU parallel computing. The GPU can execute multiple threads simultaneously to improve computing efficiency.

[0041] High Bandwidth Memory (HBM for short): It refers to a high-speed and high-bandwidth memory technology, commonly used to accelerate the training and reasoning of large models. In this solution, HBM also refers to a storage unit on the GPU.

[0042] Embodiment 1

[0043] According to the embodiments of the present application, a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0044] Considering that the number of model parameters of large models is huge and the computing resources of mobile terminals are limited, the above data processing method provided by the embodiments of the present application can be applied to, for example, Figure 1 the application scenarios shown, but not limited thereto. In the application scenarios shown in Figure 1 , the large model is deployed in the server 10. The server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include, but are not limited to: smart phones, tablets, laptops, palmtop computers, personal computers, smart home devices, vehicle-mounted devices, etc. The client device 20 can interact with the user through a graphical user interface to implement the invocation of the large model, and further implement the method provided by the embodiments of the present application.

[0045] Since in the fine-tuning process of large language model training, the input data of each batch contains multiple data sequences, and the lengths of the multiple data sequences may be different, but the data input to the model needs to be constructed into a data format of a multi-dimensional array matrix (Tensor), in which the lengths of the multiple data sequences must be the same. Therefore, the common practice in the related art is to expand the shorter data sequences to a specific length to construct a multi-dimensional array matrix.

[0046] A calculation process of an attention mechanism provided according to the above related art is as shown in Figure 2 . In large model training, a large number of attention mechanism calculations are usually required. As shown in Figure 2 , the main calculation process includes: multiplying the query matrix (Query Matrix, also referred to as the Q matrix) by the transpose of the key matrix (Key Matrix, also referred to as the K matrix) to obtain a first intermediate matrix (denoted as the S matrix); processing the S matrix with a normalized exponential function (softmax function) to obtain a second intermediate matrix (denoted as the P matrix); multiplying the P matrix by the value matrix (Value Matrix, also referred to as the V matrix) to obtain the calculation result matrix of the attention mechanism (i.e., the output matrix, Output Matrix, also referred to as the O matrix). The above main calculation process can be represented by the following formulas (1), (2), and (3):

[0047]

[0048]

[0049]

[0050] where, K T represents the transpose matrix of the K matrix, N represents the length of the data sequence (in this example Figure 2The first length, the second length, and the third length are equal in size), d represents the number of attention heads, that is, the dimension of the attention mechanism in the large language model, represents the real number field.

[0051] The above Figure 2 illustrates the calculation process of one attention mechanism. The data corresponding to the diagonal shading represents the redundant operations caused by the expansion of the data sequence in the attention mechanism. Obviously, the above expansion process provided by the related technology introduces additional computational and memory access overheads, and does not substantially improve the efficiency or effect of model training. Based on this, how to reduce the computational and memory access resources consumed in the fine-tuning process and improve the model training efficiency in the attention mechanism calculation has become one of the important problems in the related technical field. Before this application, no effective solution has been proposed in the related technical field to solve the above problems.

[0052] Based on the above related technology, in the above operating environment, this application provides a data processing method as Figure 3 shown. Figure 3 is a flowchart of a data processing method according to Embodiment 1 of this application. As Figure 3 shown, the data processing method includes:

[0053] Step S31, obtain the data to be processed, where the data to be processed is the data obtained after expanding the data sequence length of the original input data;

[0054] Step S32, select the original input data from the data to be processed based on the pre-recorded data sequence length information, where the data sequence length information is used to record the initial sequence length corresponding to the original input data;

[0055] Step S33, call multiple threads on a preset type of processor to perform target calculations on the original input data, and obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0056] The above method steps can be applied to the training scenario of the large language model in a preset application scenario to optimize the training process. In particular, in some scenarios involving the expansion of the data sequence length, by using the above method steps, more accurate calculations can be performed on the data sequence after length expansion, saving computational costs and improving computational efficiency. The above preset application scenario can be but is not limited to: scenarios involving the use of large language models in fields such as e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation. Correspondingly, the above original input data is the model training data pre-collected in the above preset application scenario. The large language model trained can exhibit high performance on specific tasks in the corresponding preset application scenario.

[0057] The above-mentioned original input data may include at least one data batch, and each data batch may include multiple data sequences, where the sequence lengths of each data sequence in the multiple data sequences may be different. During the above fine-tuning process, a multi-dimensional array matrix needs to be constructed based on the original input data. Therefore, the data sequences of different lengths in the original input data are extended to a specific length (i.e., the extended sequence length) to obtain the above-mentioned data to be processed. That is to say, the above-mentioned data to be processed is the data obtained after extending the data sequences of different lengths in the original input data during the fine-tuning process of large language model training.

[0058] The above-mentioned pre-recorded data sequence length information is used to record the initial sequence lengths of multiple data sequences in the original input data before the extension operation. That is to say, the above-mentioned data sequence length information can represent the true sequence lengths of the original input data. Based on the above data sequence length information, it is possible to determine the part of the data belonging to the original input data and the part of the data obtained by extension from the data to be processed obtained after the extension operation.

[0059] The above-mentioned preset type of processor may be a multi-core processor, a graphics processor, or an acceleration processor specific to artificial intelligence, etc. The target calculations performed on the original input data may include, but are not limited to: attention calculation, matrix calculation, activation function calculation, gradient calculation, convolution calculation, etc. According to the specific calculation requirements in the application scenario, the above processor type and the type of calculation to be performed are flexibly selected.

[0060] In an exemplary application scenario, according to the data sequence length information, the number of threads to be called is determined, and multiple threads (Threads) with the number of threads on the GPU are called to perform attention calculation on the original input data to obtain a target calculation result, which may be an output matrix. In addition, by scheduling multiple threads to perform matrix calculation under the attention mechanism, it is possible to achieve distributed training of the large language model and implement the fine-tuning stage in a distributed manner.

[0061] In the embodiments of the present application, the data to be processed is obtained, where the data to be processed is the data obtained after extending the data sequence length of the original input data; based on the pre-recorded data sequence length information, the original input data is selected from the data to be processed, where the data sequence length information is used to record the initial sequence length corresponding to the original input data; multiple threads on a preset type of processor are called to perform target calculations on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0062] It is easy to notice that in the embodiment of the present application, attention calculation is performed on part of the data to be processed (that is, the original input data) according to the initial sequence length corresponding to the original input data, thereby reducing redundant calculations of the expanded data in the data to be processed. As a result, the present application achieves the purpose of reducing computational consumption by performing targeted attention calculations on the original input data in the data to be processed through data sequence length information, thereby achieving the technical effect of reducing redundant calculations caused by expansion operations during fine-tuning, reducing computing resource consumption and improving model training efficiency, thereby solving the technical problem in the related art of performing redundant attention calculations on the data matrix constructed by the expansion operation, resulting in large computing resource consumption and low training efficiency.

[0063] In an embodiment of the present application, the above-mentioned data processing method can be used in a scenario where a large language model is trained through a system composed of a client device and a server, and the server provides data processing services for the client device. Specifically, the client device sends the data to be processed to the server, wherein the data to be processed is the data obtained after the data sequence length is expanded for the original input data; the server selects the original input data from the data to be processed based on the pre-recorded data sequence length information, calls multiple threads on a preset type processor to perform target calculations on the original input data, and obtains the target calculation results, wherein the data sequence length information is used to record the initial sequence length corresponding to the original input data, and the number of threads of the multiple threads is determined by the data sequence length information; the server returns the target calculation results to the client device.

[0064] It should be noted that the data processing method described above in the embodiments of the present application can be executed on a server, which can be a standalone server, a distributed server, or a cloud server. Furthermore, if the operating resources of the client device can meet the training, deployment, and operation requirements of the large model, the data processing method described above in the embodiments of the present application can also be performed on the client device.

[0065] The data processing method described above in the embodiment of the present application is further described below.

[0066] In an optional embodiment, in step S31, obtaining the data to be processed includes the following method steps:

[0067] Step S311, determining a plurality of initial sequences based on input batch information of the original input data, wherein at least some of the plurality of initial sequences have different sequence lengths;

[0068] Step S312, perform data sequence length expansion on multiple initial sequences to obtain the data to be processed. The data to be processed includes: an input matrix and a mask matrix to be used. The input matrix includes: multiple target sequences, and the sequence lengths of the multiple target sequences are the same. The mask matrix is used to determine the data category of each matrix element in the input matrix.

[0069] In an exemplary application scenario, the above method steps are illustrated by taking one attention calculation as an example. For example, in the current attention calculation, the input batch information of the original input data of the current batch (batch). The original input information of this current batch contains 4 data sequences, and the sequence lengths of these 4 data sequences are 1024, 32, 32, and 64 respectively.

[0070] Furthermore, in the process of training a large language model, the sizes of the Q matrix, K matrix, and V matrix participating in the attention calculation of the current batch in the fine-tuning stage are (B, M, D, K), where B represents the number of initial sequences in the original input data of the current batch (4 in this example), M represents the maximum sequence length corresponding to the multiple initial sequences (1024 in this example), D represents the number of attention heads, and K represents the dimension of each attention head. D and K are determined by the model settings of the large language model. In this example, D is set to 16 and K is set to 64. That is, the sizes of the Q matrix, K matrix, and V matrix are (4, 1024, 16, 64).

[0071] Since it is necessary to construct a matrix data format, perform a length expansion (padding) operation on the multiple initial sequences of the current batch to obtain an input matrix and a mask matrix. The input matrix includes multiple target sequences after the length expansion of the multiple initial sequences. The dimension of the mask matrix is the same as that of the input matrix, and each element in the mask matrix is used to determine the data category of the matrix element at the corresponding position in the input matrix.

[0072] Specifically, the data category is one of the following: original data, expanded data. Original data indicates that the matrix element belongs to the original input data (i.e., the real input data), and expanded data indicates that the matrix element is the data obtained during the expansion operation (i.e., the constructed data).

[0073] In the above exemplary application scenario, the mask matrix (Mask) and the input matrix (input_padding) have the same size, both (4, 1024). The matrix elements in the mask matrix are composed of 1 and 0. 1 indicates that the data category of the matrix element at the corresponding position in the input matrix is original data, and 0 indicates that the data category of the matrix element at the corresponding position in the input matrix is expanded data.

[0074] That is to say, the above-mentioned mask matrix is used to identify which elements in the input matrix are obtained through expansion operations and which elements are constructed based on the original input data.

[0075] In an optional embodiment, in step S32, based on the data sequence length information, the original input data is selected from the data to be processed, including the following method steps:

[0076] Step S321, based on the data sequence length information, perform area determination on each data element in the data to be processed to obtain a determination result, where the determination result is used to indicate the data area where each data element in the data to be processed is currently located, and the data area includes: the original data area and the expanded data area;

[0077] Step S322, according to the determination result, select the original input data from the data to be processed.

[0078] In an exemplary application scenario, based on the sequence lengths of multiple initial sequences in the original input data of the current batch, area determination is performed on each data element in the constructed input matrix. If the currently determined data element is constructed based on the original input data, it is determined that the current data element belongs to the original data area; if the currently determined data element is obtained by expanding the data sequence, it is determined that the current data element belongs to the expanded data area.

[0079] The above determination result corresponds to the above mask matrix, that is, each data element in the original data area is 1 at the corresponding position in the mask matrix, and each data element in the expanded data area is 0 at the corresponding position in the mask matrix. In another exemplary application scenario, the above mask matrix can also be generated according to the above determination result.

[0080] According to the determination result, multiple data elements in the input matrix in the data to be processed that belong to the original data area are selected as the original input data.

[0081] According to the above method steps of the embodiments of the present application, a schematic diagram of an attention calculation process as shown in Figure 4 is provided. In the calculation of the attention mechanism, first add the mask matrix to the S matrix, and then use the softmax function to process the superimposed result matrix to obtain the P matrix. Thus, the expanded part in the multi-dimensional array matrix can be made not to affect the calculation of the softmax function.

[0082] In an optional embodiment, the data processing method further includes the following method steps:

[0083] Step S34, refuse to perform the target calculation on the data in the expanded data area and release the thread resources occupied by the data in the expanded data area.

[0084] Still as Figure 4 shown, during the process of calculating attention based on the superposition result of the mask matrix and the input matrix, only the elements of the input matrix at the positions where the mask matrix element value is 1 are used for attention calculation, while skipping (or rejecting) the attention calculation of the elements of the input matrix at the positions where the mask matrix element value is 0, thus avoiding redundant calculations in attention calculation.

[0085] In each matrix calculation task in the above attention calculation, a thread on the GPU needs to be called. Before each thread executes the matrix calculation task, a pre-calculation judgment is made according to the above method steps of the embodiments of the present application. If it is judged that the object currently prepared for calculation by the thread is the data within the extended data area, the calculation task of the thread is controlled to end, so as to release the thread resources and reduce the occupation of calculation resources and memory access resources for attention calculation.

[0086] It is easy to notice that during model training, by extracting the true information of the data sequence length in the original input data, based on this true information, the additional calculation and memory access resource consumption brought by the extension operation are reduced or eliminated, thereby reducing the calculation load during model training.

[0087] In an optional embodiment, the data processing method further includes the following method steps:

[0088] Step S35, adjusting the size of the input matrix based on the mask matrix to obtain data sequence length information.

[0089] The above-mentioned size adjustment operation can be used to perform size mapping processing on the input matrix, which is achieved through the reorganization operation of the matrix. After reorganizing the input matrix into a matrix with a specific size, it can adapt to specific calculation requirements or model structures.

[0090] Still taking the above exemplary application scenario as an example, when the sizes of the Q matrix, K matrix, and V matrix for the attention calculation of the current batch are (4, 1024, 16, 64), the size of the input matrix is adjusted (such as reshape processing) according to the mask matrix to obtain the above data sequence length information. Thus, it is determined that the sequence lengths of the 4 data sequences in the original input data of the current batch are 1024, 32, 32, and 64 respectively, and the above data sequence length information can be recorded in a one-dimensional array, expressed as [1024, 32, 32, 64].

[0091] In an optional embodiment, the data processing method further includes the following method steps:

[0092] Step S361, constructing an initial hidden layer state based on the input matrix;

[0093] Step S362, update the initial hidden layer state using the data sequence length information to obtain the target hidden layer state.

[0094] The above initial hidden layer state can be a hidden state matrix. The data elements of the hidden state matrix include: the number of initial sequences in the original input data of the current batch, the maximum sequence lengths corresponding to multiple initial sequences, and the number of hidden layer neurons. The number of hidden layer neurons is determined by the model settings of the large language model. The number of hidden layer neurons is obtained by multiplying the number of attention heads by the dimension of each attention head.

[0095] Construct a hidden state matrix according to the input matrix. After each attention calculation in the fine-tuning stage of the large language model training, the hidden state matrix is updated according to the attention calculation result. In addition, the hidden state matrix can also be remapped according to the data sequence length information to remove the influence of invalid data introduced during the expansion operation, and obtain the target hidden layer state.

[0096] In the above exemplary application scenario, the above initial hidden layer state (hidden_states) is constructed according to the input matrix (input_padding). As mentioned before, the dimension of the input matrix is (B, M). Correspondingly, the dimension of the initial hidden layer state is (B, M, H), where H represents the number of hidden layer neurons, and H is obtained by multiplying D and K. D represents the number of attention heads, and K represents the dimension of each attention head. In this example, the dimension of the input matrix is (4, 1024), and the dimension of the initial hidden layer state is (4, 1024, 1024). Further, remap the initial hidden layer state according to the data sequence length information. In this example, the data sequence length information is represented as a one-dimensional array [1024, 32, 32, 64]. Then, the dimension of the remapped target hidden layer state is determined by the sum of the elements in the one-dimensional array and the number of hidden layer neurons. That is, the dimension of the target hidden layer state is (1152, 1024).

[0097] Thus, according to the above method steps, performing attention calculation based on the target hidden layer state can remove the influence of invalid data introduced during the expansion process, reduce redundant calculations, reduce the occupation of computing resources, and improve the model training efficiency.

[0098] In an optional embodiment, the data processing method further includes the following method steps:

[0099] Step S371, determine the initial matrix dimension according to the sequence lengths of multiple initial sequences, where the initial matrix dimension is used to determine the initial dimensions of multiple matrices for parameter target calculation, and the multiple matrices include: the query matrix, key matrix, and value matrix corresponding to the original input data;

[0100] Step S372: Update the initial matrix dimension based on the data sequence length information and the target hidden layer state to obtain the target matrix dimension, where the target matrix dimension is used to determine the target dimensions of multiple matrices.

[0101] The above query matrix (i.e., Q matrix), key matrix (i.e., K matrix), and value matrix (i.e., V matrix) are used to represent the attention weights of each element in the input sequence during attention calculation. The query matrix is used to extract information from the input sequence, the key matrix is used to represent the information to be queried, and the value matrix is used to determine the numerical value of attention weighting. The above target matrix dimension is determined by the sum of the lengths of multiple data sequences in the data sequence length information.

[0102] In attention calculation, the query matrix is used to represent the object for which attention needs to be calculated, and this object can be an element in a certain data sequence of the input. The query matrix will be used to calculate the similarity between this object and the key matrix to determine the corresponding value matrix. The key matrix is used to represent the object to be compared, usually including all elements in the input data sequence. The key matrix will perform a similarity calculation with the query matrix to determine the relevance between each element and the query matrix. The value matrix is used to represent the numerical information corresponding to the key matrix, usually the numerical values corresponding to each element in the input data sequence, and this data is used to characterize the degree of similarity. That is to say, under the attention mechanism, the similarity between the query matrix and the key matrix is calculated, and then the value matrix is weighted according to the similarity. The attention mechanism can help the model better focus on the information related to the task in the input data sequence and can effectively process long sequence data.

[0103] In the above exemplary application scenario, before performing attention calculation on the input data of the current batch, construct the initial matrix dimension according to the sequence lengths of multiple initial sequences in the original input data, that is, determine the initial dimensions of the above query matrix, key matrix, and value matrix. The initial matrix dimensions corresponding to the above query matrix, key matrix, and value matrix are represented as (B, M, D, K). In this example, the initial matrix dimensions corresponding to the query matrix, key matrix, and value matrix for parameter target calculation (attention calculation in this example) are (4, 1024, 16, 64).

[0104] Further, based on the data sequence length information (sequence_list) and the target hidden layer state, remap the above initial matrix dimensions, and adjust the matrix dimensions corresponding to the above query matrix, key matrix, and value matrix to (1, sum(sequence_list), D, K) to obtain the target matrix dimensions. In this example, the data sequence length information (sequence_list) is represented as a one-dimensional array [1024, 32, 32, 64], then sum(sequence_list) is 1152, and the target matrix dimensions are represented as (1, 1152, 16, 64).

[0105] Further, perform attention calculation based on the above data sequence length information [1024, 32, 32, 64] and the above target matrix dimensions represented as (1, 1152, 16, 64). During attention calculation, the convolution kernel (kernel) schedules multiple threads (Thread) in the GPU according to the data sequence length information.

[0106] In an exemplary implementation, based on the solution provided by the related technology, calculate the number of threads to be scheduled according to the number of initial sequences (B) in the original input data, the maximum sequence length (M) corresponding to the multiple initial sequences, the number of attention heads (D), and the dimension (K) of each attention head. In this example, the number of threads is B×M / Q×K, where Q represents the data sequence length (QueriesPerBlock) that each thread needs to process. For a specific thread, Q is a constant, usually set to 64. Therefore, the number of threads here is 4×1024 / 64×16 = 1024.

[0107] In an alternative embodiment, in the above step S33, calling multiple threads on a preset type of processor to perform target calculation on the original input data to obtain the target calculation result further includes the following method steps:

[0108] Step S331, calling multiple threads on a preset type of processor to load the multiple matrices corresponding to the original input data from a preset storage area on the preset type of processor to the registers on the preset type of processor, and performing target calculation on the multiple matrices stored in the registers to obtain the target calculation result.

[0109] In the above exemplary application scenario, before each thread in the GPU executes the matrix calculation task, it determines whether the current calculated data element belongs to the extended data area in the input matrix. If so, it directly ends the task to be executed on the current thread. If not, it continues to execute the matrix calculation task. That is, load the above query matrix, key matrix, and value matrix from the high-bandwidth memory HBM of the GPU to the registers. Further, after the attention calculation is completed, write the attention calculation result to the HBM.

[0110] Therefore, in this example, the number of GPU threads actually performing the matrix calculation task is sum(sequence_list) / Q×D. sum(sequence_list) is the sum of the true lengths of multiple initial sequences in the original data, and B×M represents the total number of elements in the input matrix obtained after expanding the length of the data sequence. Obviously, B×M is usually greater than or even much greater than sum(sequence_list). In this example, according to the solution provided in the embodiments of the present application, the number of threads is 1152 / 64×16 = 228. Compared with the solution provided according to the related technology (the number of threads is 1024), the present application saves the number of threads required for attention calculation. Therefore, the above method provided in the embodiments of the present application can significantly reduce the consumption of computing resources and memory access resources in attention calculation during the fine-tuning stage. Moreover, since some redundant calculations are avoided, the above method can also improve the model training efficiency.

[0111] In an optional embodiment, the data processing method further includes the following method steps:

[0112] Step S381, write the target calculation result into a preset storage area;

[0113] Step S382, in response to the successful writing of the target calculation result, restore the target hidden layer state to the initial hidden layer state based on the data sequence length information.

[0114] The above preset storage area can be a specified HBM area in the GPU.

[0115] In the above exemplary application scenario, for a single attention calculation, after writing the calculation result of the current attention calculation into the HBM area, update the target hidden layer state according to the calculation result.

[0116] Restoring the target hidden layer state to the initial hidden layer state based on the data sequence length information can be to restore the target hidden layer state to the initial hidden layer state corresponding to the current batch based on the original input data of the current batch. Restoring the target hidden layer state to the initial hidden layer state based on the data sequence length information can also be to re-determine the initial hidden layer state corresponding to the next batch based on the data sequence length information of the original input data of the next batch, and then adjust the target hidden layer state to the initial hidden layer state corresponding to the next batch. Thus, it is possible to avoid the influence of the attention calculation for the input data of the current batch on the calculation of the next batch, which is beneficial to the continuity of attention calculation in large language model training.

[0117] According to the above method steps provided in the embodiments of the present application, there is also provided a schematic diagram of an attention calculation process as Figure 5 shown. AsFigure 5 As shown, under the attention mechanism, this solution provides an optimized algorithm for freely expanding (padding free) data sequences, making efficient use of the GPU bandwidth, and solving the technical problem in the related art that the expansion operation requires introducing additional computational and memory access overheads and cannot effectively improve the model training efficiency.

[0118] Specifically, as Figure 5 shown, during the attention calculation, Flash Attention is used to perform block calculation on the matrix. Flash Attention refers to an attention mechanism that can simulate human vision. A mask matrix (with values of 0 or 1) is multiplied by the S matrix, making the elements of the partially expanded matrix in the S matrix take the value of 0 at the corresponding positions in the mask matrix.

[0119] Specifically, as Figure 5 shown, the length of the data sequence in the original input data is expanded to obtain a first intermediate matrix (i.e., the input matrix, denoted as the S matrix) and a mask matrix. After superimposing the first intermediate matrix and the mask matrix, the softmax function is used for processing to obtain a second intermediate matrix (denoted as the P matrix); at the model input, the S matrix is processed according to the mask matrix to obtain the real data sequence corresponding to the original input data (denoted as S_r) and the target hidden layer state with adjusted dimensions (denoted as hidden_states_r); before performing the attention calculation, the target dimensions of the query matrix, key matrix, and value matrix are determined according to the real data sequence S_r and the target hidden layer state hidden_states_r, and the corresponding query matrix, key matrix, and value matrix are constructed according to the target dimensions; on the convolutional kernel side, multiple threads of the GPU are scheduled according to the sequence length of the original input data. In each thread, the attention calculation is performed based on the query matrix, key matrix, value matrix, and mask matrix. Specifically, before each thread calculation, it is judged according to the above mask matrix whether the data element to be calculated currently belongs to the expanded data area of the input matrix. If so, the calculation process of this thread is ended. If not, the above attention calculation is performed to obtain the calculation result.

[0120] Still as Figure 5 shown, in a single attention calculation, the transposed matrix of the key matrix is traversed in the outer loop, and the query matrix is traversed in the inner loop. During the traversal, multiple matrix multiplication calculations are performed. Figure 5 The first matrix multiplication shown in it represents the first matrix multiplication calculation in the inner loop, that is, the process of multiplying the Q matrix by the transposed matrix of the K matrix to obtain the S matrix. The second matrix multiplication represents the second matrix multiplication calculation in the inner loop, that is, the process of multiplying the P matrix and the V matrix to obtain the output matrix (O matrix).

[0121] As Figure 5 shown, the data corresponding to the shaded part (including the black-filled part) in the matrix is the part currently being calculated in the two-layer loop. Each calculation is assigned to a thread of the GPU for processing.

[0122] The solution provided by this application can solve the additional computational and memory access resource consumption caused by the need for padding operations due to unequal data sequence lengths of the input data during the fine-tuning stage in the training process of large language models. By extracting the true information of the data sequence length in the input data during model training and reducing or eliminating the additional computational and memory access resource consumption brought by the padding operation based on this true information, the computational load during model training can be reduced. Through the above solution provided by this application, the computational efficiency of model training can be improved at a relatively low cost, enabling cloud computing heterogeneous models to have higher cost performance when performing distributed training on large language models.

[0123] By applying the above method provided by the embodiments of this application in actual application scenarios, the fine-tuning process in the training of large language models is optimized, effectively reducing the computational amount of model training, and improving the computational efficiency of the GPU used to train the model. In an exemplary embodiment, compared with the related technology, using the above method provided by the embodiments of this application can improve the training performance of large language models by at least 20%.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0125] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0127] Embodiment 2

[0128] Under the operating environment as in Embodiment 1, the present application provides another data processing method as Figure 6 shown. Figure 6 is a flowchart of a data processing method according to Embodiment 2 of the present application. As Figure 6 shown, the data processing method includes:

[0129] Step S61, obtaining a data processing request through a first application programming interface. Among them, the request data carried in the data processing request includes: data to be processed, where the data to be processed is data obtained after expanding the data sequence length of the original input data;

[0130] Step S62, returning a data processing response through a second application programming interface. Among them, the response data carried in the data processing response includes: a target calculation result, which is obtained by calling multiple threads on a preset type of processor to perform a target calculation on the original input data. The number of threads of the multiple threads is determined by the data sequence length information, and the original input data is selected from the data to be processed based on the data sequence length information. The data sequence length information is used to record the initial sequence length corresponding to the original input data.

[0131] According to the above method steps, a method for implementing a data processing cloud service is provided. The method runs on a cloud server. The cloud server obtains a data processing request sent by a service caller through a first application programming interface, and based on the data to be processed carried in the data processing request, executes a data processing process to obtain a target calculation result. Further, the cloud server returns a data processing response to the service caller through a second application programming interface to provide the target calculation result to the service caller.

[0132] In addition, when the operating resources of the client device can meet the conditions for training, deploying, and running the large model, the above data processing method of the embodiments of the present application can also be performed in the client device to provide local data processing services for customers.

[0133] The above method steps can be applied to the scenario of assisting in the training of large language models through data processing cloud services in a preset application scenario to optimize the large language model training process. In particular, in some scenarios involving the expansion of the data sequence length, more accurate calculations can be performed on the data sequence after length expansion by using the above method steps, saving computing costs and improving computing efficiency. The above preset application scenario can be but is not limited to: scenarios involving the use of large language models in fields such as e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation. Correspondingly, the above original input data is the model training data pre-collected in the above preset application scenario. The trained large language model can show high performance on specific tasks in the corresponding preset application scenario.

[0134] The above original input data may include at least one data batch, and each data batch may include multiple data sequences. The sequence lengths of each data sequence in the multiple data sequences may be different. In the above fine-tuning process, a multi-dimensional array matrix needs to be constructed based on the original input data. Therefore, the data sequences with different lengths in the original input data are expanded to a specific length (i.e., the expanded sequence length) to obtain the above-mentioned data to be processed. That is to say, the above-mentioned data to be processed is the data obtained after expanding the data sequences with different lengths in the original input data during the fine-tuning process of large language model training.

[0135] The above-mentioned pre-recorded data sequence length information is used to record the initial sequence lengths of multiple data sequences in the original input data before the expansion operation. That is to say, the above-mentioned data sequence length information can represent the true sequence length of the original input data. Based on the above data sequence length information, it is possible to determine the part of the data belonging to the original input data and the part of the data obtained by expansion from the data to be processed obtained after the expansion operation.

[0136] The above-mentioned preset type of processor can be a multi-core processor, a graphics processor, or an acceleration processor specific to artificial intelligence, etc. The target calculations performed on the original input data may include but are not limited to: attention calculation, matrix calculation, activation function calculation, gradient calculation, convolution calculation, etc. According to the specific calculation requirements in the application scenario, the above processor type and the type of calculation to be performed are flexibly selected.

[0137] In an exemplary application scenario, the number of threads to be called is determined according to the data sequence length information. Further, multiple threads on the GPU are called to perform attention calculation on the original input data, and a target calculation result is obtained. The target calculation result can be an output matrix. In addition, by scheduling multiple threads to perform matrix calculation under the attention mechanism, distributed training of the large language model can be achieved, and the fine-tuning stage can be implemented in a distributed manner.

[0138] In the embodiment of the present application, a data processing request is obtained through a first application programming interface. The request data carried in the data processing request includes: data to be processed, where the data to be processed is data obtained after expanding the data sequence length of the original input data; a data processing response is returned through a second application programming interface. The response data carried in the data processing response includes: a target calculation result, which is obtained by calling multiple threads on a preset type of processor to perform a target calculation on the original input data. The number of threads of the multiple threads is determined by the data sequence length information, and the original input data is selected from the data to be processed based on the data sequence length information. The data sequence length information is used to record the initial sequence length corresponding to the original input data.

[0139] It is easy to notice that in the embodiment of the present application, based on the above method, a data processing service is provided, which can run on a cloud server. The cloud server performs attention calculation on part of the data in the data to be processed (i.e., the original input data) according to the initial sequence length corresponding to the original input data, reducing redundant calculation on the expanded data in the data to be processed. Thus, the present application achieves the purpose of reducing calculation consumption by performing attention calculation on the original input data in the data to be processed specifically according to the data sequence length information, thereby realizing the technical effects of reducing redundant calculation caused by the expansion operation in the fine-tuning process, reducing calculation resource consumption, and improving model training efficiency, and further solving the technical problem in the related art that redundant attention calculation on the data matrix constructed by the expansion operation results in large calculation resource consumption and low training efficiency. In addition, implementing the data processing process on the cloud server improves the flexibility and scalability of the large language model training process.

[0140] It should be noted that the preferred implementation manner of this embodiment can refer to the relevant description in Embodiment 1, which will not be elaborated here.

[0141] Embodiment 3

[0142] In the operating environment as in Embodiment 1, the present application provides another data processing method as Figure 7 shown. Figure 7 is a flowchart of a data processing method according to Embodiment 3 of the present application, as Figure 7As shown, the data processing method includes:

[0143] Step S71: Obtain the current input data processing dialogue request. Among them, the request data carried in the data processing dialogue request includes: the data to be processed, where the data to be processed is the data obtained after expanding the data sequence length of the original input data;

[0144] Step S72: In response to the data processing dialogue request, return a data processing dialogue reply. Among them, the information carried in the data processing dialogue reply includes: the target calculation result, which is obtained by invoking multiple threads on a preset type of processor to perform target calculations on the original input data. The number of threads of the multiple threads is determined by the data sequence length information. The original input data is selected from the data to be processed based on the data sequence length information. The data sequence length information is used to record the initial sequence length corresponding to the original input data;

[0145] Step S73: Display the target calculation result in the graphical user interface.

[0146] According to the above method steps, a visualization solution for data processing functions is provided. The terminal device provides a graphical user interface, and at least one data processing scenario is displayed in the graphical user interface. The display content of the graphical user interface also includes input components (such as text input boxes, voice input controls, etc.) and display components (such as text display windows). The user inputs a data processing dialogue request through the input component to specify the data to be processed in the data processing task. After detecting the user's input behavior, the data processing process is executed based on the data to be processed to obtain the target calculation result. Further, the target calculation result is displayed through the display component in the graphical user interface.

[0147] The above method steps can be applied to the visualization interaction scenario of data processing services in large language model training in a preset application scenario to optimize the large language model training process. In particular, in some scenarios involving data sequence length expansion, the above method steps can be used to perform more accurate calculations on the data sequence after length expansion, saving calculation costs and improving calculation efficiency. The above preset application scenario can be but is not limited to: scenarios involving the use of large language models in fields such as e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation. Correspondingly, the above original input data is the model training data pre-collected in the above preset application scenario. The trained large language model can show high performance on specific tasks in the corresponding preset application scenario.

[0148] The above-mentioned original input data may include at least one data batch, and each data batch may include multiple data sequences, where the sequence lengths of each data sequence in the multiple data sequences may be different. During the above-mentioned fine-tuning process, a multi-dimensional array matrix needs to be constructed based on the original input data. Therefore, the data sequences with different lengths in the original input data are extended to a specific length (i.e., the extended sequence length) to obtain the above-mentioned data to be processed. That is to say, the above-mentioned data to be processed is the data obtained by extending the data sequences with different lengths in the original input data during the fine-tuning process of large language model training.

[0149] The above-mentioned pre-recorded data sequence length information is used to record the initial sequence lengths of multiple data sequences in the original input data before the extension operation. That is to say, the above-mentioned data sequence length information can represent the true sequence lengths of the original input data. Based on the above-mentioned data sequence length information, it is possible to determine the part of the data belonging to the original input data and the part of the data obtained by extension from the data to be processed obtained after the extension operation.

[0150] The above-mentioned preset type of processor may be a multi-core processor, a graphics processor, or an acceleration processor specific to artificial intelligence, etc. The target calculations performed on the original input data may include, but are not limited to: attention calculation, matrix calculation, activation function calculation, gradient calculation, convolution calculation, etc. According to the specific calculation requirements in the application scenario, the above-mentioned processor type and the type of calculation to be performed are flexibly selected.

[0151] In an exemplary application scenario, according to the data sequence length information, the number of threads to be called is determined. Further, multiple threads on the GPU are called to perform attention calculation on the original input data to obtain a target calculation result, which may be an output matrix. In addition, by scheduling multiple threads for matrix calculation under the attention mechanism, distributed training of the large language model can be achieved, and the fine-tuning stage can be implemented in a distributed manner.

[0152] In the embodiment of the present application, a current input data processing dialogue request is obtained, where the request data carried in the data processing dialogue request includes: data to be processed, where the data to be processed is the data obtained by extending the data sequence length of the original input data; in response to the data processing dialogue request, a data processing dialogue reply is returned, where the information carried in the data processing dialogue reply includes: a target calculation result, which is obtained by calling multiple threads on a preset type of processor to perform target calculations on the original input data, the number of threads of the multiple threads is determined by the data sequence length information, the original input data is selected from the data to be processed based on the data sequence length information, and the data sequence length information is used to record the initial sequence length corresponding to the original input data; the target calculation result is displayed in the graphical user interface.

[0153] It is easy to note that in the embodiments of the present application, based on the above method, a visualization interaction solution for data processing services is provided. The input information of the user is obtained through the graphical user interface and the target calculation result is returned to the user. During the data processing, according to the initial sequence length corresponding to the original input data, attention calculation is performed on some of the data to be processed (i.e., the original input data), reducing the redundant calculation for the extended data in the data to be processed. Thus, the present application achieves the purpose of performing attention calculation on the original input data in the data to be processed specifically through the data sequence length information to reduce the calculation consumption, thereby realizing the technical effects of reducing the redundant calculation brought by the extension operation during the fine-tuning process, reducing the calculation resource consumption, and improving the model training efficiency, and further solving the technical problem in the related art that redundant attention calculation is performed on the data matrix constructed by the extension operation, resulting in large calculation resource consumption and low training efficiency. In addition, implementing the data processing process in the cloud server improves the flexibility and scalability of the large language model training process.

[0154] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in Embodiment 1, which will not be elaborated here.

[0155] Embodiment 4

[0156] According to the embodiments of the present application, an apparatus embodiment for implementing the above data processing method is also provided. Figure 8 It is a schematic structural diagram of a data processing apparatus according to Embodiment 4 of the present application, as Figure 8 shown. The apparatus includes:

[0157] An acquisition module 801, configured to acquire data to be processed, where the data to be processed is data obtained by extending the data sequence length of the original input data;

[0158] A selection module 802, configured to select the original input data from the data to be processed based on the pre-recorded data sequence length information, where the data sequence length information is used to record the initial sequence length corresponding to the original input data;

[0159] A calculation module 803, configured to call multiple threads on a preset type of processor to perform target calculation on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0160] Optionally, the above-mentioned acquisition module 801 is further configured to: determine a plurality of initial sequences based on the input batch information of the original input data, where at least some of the plurality of initial sequences have different sequence lengths; perform data sequence length expansion on the plurality of initial sequences to obtain data to be processed, where the data to be processed includes: an input matrix and a mask matrix to be used, the input matrix includes: a plurality of target sequences, the sequence lengths of the plurality of target sequences are the same, and the mask matrix is used to determine the data category of each matrix element in the input matrix.

[0161] Optionally, the above-mentioned selection module 802 is further configured to: perform area determination on each data element in the data to be processed based on the data sequence length information to obtain a determination result, where the determination result is used to indicate the data area where each data element in the data to be processed is currently located, and the data area includes: an original data area and an expanded data area; select the original input data from the data to be processed according to the determination result.

[0162] Optionally, in addition to the above-mentioned all modules, the above-mentioned data processing device further includes: a stop module 804 (not shown in the figure), which is used to reject performing target calculations on the data in the expanded data area and release the thread resources occupied by the data in the expanded data area.

[0163] Optionally, in addition to the above-mentioned all modules, the above-mentioned data processing device further includes: an adjustment module 805 (not shown in the figure), which is used to perform a dimension adjustment operation on the input matrix based on the mask matrix to obtain data sequence length information.

[0164] Optionally, in addition to the above-mentioned all modules, the above-mentioned data processing device further includes: an update module 806 (not shown in the figure), which is used to construct an initial hidden layer state based on the input matrix; update the initial hidden layer state by using the data sequence length information to obtain a target hidden layer state.

[0165] Optionally, in addition to the above-mentioned all modules, the above-mentioned data processing device further includes: a determination module 808 (not shown in the figure), which is used to determine an initial matrix dimension according to the sequence lengths of the plurality of initial sequences, where the initial matrix dimension is used to determine the initial dimensions of the plurality of matrices for parameter target calculation, and the plurality of matrices include: a query matrix, a key matrix, and a value matrix corresponding to the original input data; update the initial matrix dimension based on the data sequence length information and the target hidden layer state to obtain a target matrix dimension, where the target matrix dimension is used to determine the target dimensions of the plurality of matrices.

[0166] Optionally, the above calculation module 803 is further configured to: call multiple threads on a preset type of processor to load multiple matrices corresponding to the original input data from a preset storage area on the preset type of processor to registers on the preset type of processor, and perform target calculations on the multiple matrices stored in the registers to obtain a target calculation result.

[0167] Optionally, in addition to the above all modules, the above data processing device further includes: a restoration module 808 (not shown in the figure), configured to write the target calculation result into a preset storage area; and in response to successful writing of the target calculation result, restore the target hidden layer state to the initial hidden layer state based on the data sequence length information.

[0168] It should be noted here that the above acquisition module 801, selection module 802, and calculation module 803 correspond to steps S31 to S33 in Embodiment 1. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules or units may be hardware components or software components stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n), and the above modules may also be part of a device and can run in the computer terminal 10 provided in Embodiment 1.

[0169] In an embodiment of the present application, an acquisition module is used to acquire data to be processed, where the data to be processed is data obtained by expanding the data sequence length of the original input data; a selection module is used to select the original input data from the data to be processed based on pre-recorded data sequence length information, where the data sequence length information is used to record the initial sequence length corresponding to the original input data; and a calculation module is used to call multiple threads on a preset type of processor to perform target calculations on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0170] It is easy to note that in an embodiment of the present application, according to the initial sequence length corresponding to the original input data, attention calculations are performed on some of the data in the data to be processed (that is, the original input data), reducing redundant calculations on the expanded data in the data to be processed. Thus, the present application achieves the purpose of performing attention calculations on the original input data in the data to be processed specifically through the data sequence length information to reduce calculation consumption, thereby realizing the technical effects of reducing redundant calculations brought by the expansion operation during the fine-tuning process, reducing calculation resource consumption, and improving the model training efficiency, and further solving the technical problem in the related art that redundant attention calculations are performed on the data matrix constructed by the expansion operation, resulting in large calculation resource consumption and low training efficiency.

[0171] According to an embodiment of the present application, there is also provided another apparatus embodiment for implementing the data processing method in Embodiment 2 above. Figure 9 is a schematic structural diagram of another data processing apparatus according to Embodiment 4 of the present application, as Figure 9 shown, the apparatus includes:

[0172] An acquisition module 901, configured to obtain a data processing request through a first application programming interface, where the request data carried in the data processing request includes: data to be processed, where the data to be processed is data obtained after expanding the data sequence length of the original input data;

[0173] A return module 902, configured to return a data processing response through a second application programming interface, where the response data carried in the data processing response includes: a target calculation result, which is obtained by invoking multiple threads on a preset type of processor to perform a target calculation on the original input data, the number of threads of the multiple threads is determined by the data sequence length information, the original input data is selected from the data to be processed based on the data sequence length information, and the data sequence length information is used to record the initial sequence length corresponding to the original input data.

[0174] It should be noted here that the above acquisition module 901 and return module 902 correspond to steps S61 to step S62 in Embodiment 2. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 2 above. It should be noted that the above modules or units may be hardware components or software components stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n), and the above modules may also be part of the apparatus and can run in the computer terminal 10 provided in Embodiment 1.

[0175] According to an embodiment of the present application, there is also provided an apparatus embodiment for implementing the data processing method in Embodiment 3 above. Figure 10 is a schematic structural diagram of another data processing apparatus according to Embodiment 4 of the present application, as Figure 10 shown, the apparatus includes:

[0176] An acquisition module 1001, configured to obtain a current input data processing dialogue request, where the request data carried in the data processing dialogue request includes: data to be processed, where the data to be processed is data obtained after expanding the data sequence length of the original input data;

[0177] Return module 1002, configured to return a data processing dialogue reply in response to a data processing dialogue request, wherein the information carried in the data processing dialogue reply includes: a target calculation result obtained by performing a target calculation on original input data by invoking multiple threads on a preset type of processor, the number of threads of the multiple threads being determined by data sequence length information, the original input data being selected from the data to be processed based on the data sequence length information, and the data sequence length information being used to record the initial sequence length corresponding to the original input data;

[0178] Display module 1003, configured to display the target calculation result within a graphical user interface.

[0179] It should be noted here that the above-mentioned acquisition module 1001, return module 1002, and display module 1003 correspond to steps S71 to S73 in Embodiment 3. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 3. It should be noted that the above-mentioned module or unit may be a hardware component or a software component stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n), and the above-mentioned module may also be a part of the device and can run in the computer terminal 10 provided in Embodiment 1.

[0180] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in Embodiment 1 or Embodiment 2, and will not be elaborated here.

[0181] Embodiment 5

[0182] According to an embodiment of the present application, there is also provided an electronic device, which can be any terminal device in a group of electronic devices. Optionally, in this embodiment, the above-mentioned electronic device can also be replaced with a terminal device such as a mobile terminal.

[0183] Optionally, in this embodiment, the above-mentioned electronic device can be at least one network device among multiple network devices in a computer network.

[0184] In this embodiment, the above-mentioned electronic device can execute program codes of the following steps in the data processing method: acquiring data to be processed, wherein the data to be processed is data obtained by performing data sequence length expansion on original input data; selecting original input data from the data to be processed based on pre-recorded data sequence length information, wherein the data sequence length information is used to record the initial sequence length corresponding to the original input data; invoking multiple threads on a preset type of processor to perform a target calculation on the original input data to obtain a target calculation result, wherein the number of threads of the multiple threads is determined by the data sequence length information.

[0185] Optionally, Figure 11 is a structural block diagram of an electronic device according to Embodiment 5 of the present application, as Figure 11 shown. The electronic device 110 may include: one or more (only one is shown in the figure) processors 1102, a memory 1104, a storage controller 1106, and a peripheral interface 1108. Among them, the peripheral interface 1108 is connected to a radio frequency module, an audio module, and a display.

[0186] Among them, the memory 1104 can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the above data processing method. The memory 1104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1104 may further include a memory remotely provided with respect to the processor, and these remote memories may be connected to the electronic device 110 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0187] The processor 1102 can call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining data to be processed, where the data to be processed is data obtained after expanding the data sequence length of the original input data; based on the pre-recorded data sequence length information, selecting the original input data from the data to be processed, where the data sequence length information is used to record the initial sequence length corresponding to the original input data; calling multiple threads on a preset type of processor to perform a target calculation on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0188] Optionally, the above processor 1102 may also execute the program code of the following steps: determining multiple initial sequences based on the input batch information of the original input data, where at least some of the multiple initial sequences have different sequence lengths; expanding the data sequence lengths of the multiple initial sequences to obtain data to be processed, where the data to be processed includes: an input matrix and a mask matrix to be used, the input matrix includes: multiple target sequences, the sequence lengths of the multiple target sequences are the same, and the mask matrix is used to determine the data categories of each matrix element in the input matrix.

[0189] Optionally, the above-mentioned processor 1102 may also execute program code for the following steps: based on the data sequence length information, perform area determination on each data element in the to-be-processed data to obtain a determination result, where the determination result is used to indicate the data area where each data element in the to-be-processed data is currently located, and the data area includes: the original data area and the extended data area; according to the determination result, select the original input data from the to-be-processed data.

[0190] Optionally, the above-mentioned processor 1102 may also execute program code for the following steps: refuse to perform target calculations on the data in the extended data area and release the thread resources occupied by the data in the extended data area.

[0191] Optionally, the above-mentioned processor 1102 may also execute program code for the following steps: perform a resizing operation on the input matrix based on the mask matrix to obtain data sequence length information.

[0192] Optionally, the above-mentioned processor 1102 may also execute program code for the following steps: construct an initial hidden layer state based on the input matrix; update the initial hidden layer state using the data sequence length information to obtain a target hidden layer state.

[0193] Optionally, the above-mentioned processor 1102 may also execute program code for the following steps: determine an initial matrix dimension based on the sequence lengths of multiple initial sequences, where the initial matrix dimension is used to determine the initial dimensions of multiple matrices for parameter target calculations, and the multiple matrices include: the query matrix, key matrix, and value matrix corresponding to the original input data; update the initial matrix dimension based on the data sequence length information and the target hidden layer state to obtain a target matrix dimension, where the target matrix dimension is used to determine the target dimensions of the multiple matrices.

[0194] Optionally, the above-mentioned processor 1102 may also execute program code for the following steps: call multiple threads on a preset type of processor to load the multiple matrices corresponding to the original input data from a preset storage area on the preset type of processor to registers on the preset type of processor, and perform target calculations on the multiple matrices stored in the registers to obtain a target calculation result.

[0195] Optionally, the above-mentioned processor 1102 may also execute program code for the following steps: write the target calculation result into a preset storage area; in response to the successful writing of the target calculation result, restore the target hidden layer state to the initial hidden layer state based on the data sequence length information.

[0196] The processor 1102 can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtain a data processing request through the first application programming interface, where the request data carried in the data processing request includes: data to be processed, where the data to be processed is data obtained after expanding the data sequence length of the original input data; return a data processing response through the second application programming interface, where the response data carried in the data processing response includes: a target calculation result, which is obtained by calling multiple threads on a preset type of processor to perform a target calculation on the original input data, the number of threads of the multiple threads is determined by the data sequence length information, the original input data is selected from the data to be processed based on the data sequence length information, and the data sequence length information is used to record the initial sequence length corresponding to the original input data.

[0197] The processor 1102 can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtain the current input data processing dialogue request, where the request data carried in the data processing dialogue request includes: data to be processed, where the data to be processed is data obtained after expanding the data sequence length of the original input data; in response to the data processing dialogue request, return a data processing dialogue reply, where the information carried in the data processing dialogue reply includes: a target calculation result, which is obtained by calling multiple threads on a preset type of processor to perform a target calculation on the original input data, the number of threads of the multiple threads is determined by the data sequence length information, the original input data is selected from the data to be processed based on the data sequence length information, and the data sequence length information is used to record the initial sequence length corresponding to the original input data; display the target calculation result in the graphical user interface.

[0198] According to an embodiment of the present application, an electronic device for implementing the above-mentioned data processing method is provided. Data to be processed is obtained, wherein the data to be processed is data obtained after the original input data is expanded by the data sequence length; based on the pre-recorded data sequence length information, the original input data is selected from the data to be processed, wherein the data sequence length information is used to record the initial sequence length corresponding to the original input data; multiple threads on a preset type processor are called to perform target calculation on the original input data to obtain a target calculation result, wherein the number of threads of the multiple threads is determined by the data sequence length information. In an embodiment of the present application, according to the initial sequence length corresponding to the original input data, attention calculation is performed on part of the data in the data to be processed (i.e., the original input data), reducing the redundant calculation of the expanded data in the data to be processed. Thus, the present application achieves the purpose of reducing computational consumption by performing attention calculation on the original input data in the data to be processed through the data sequence length information, thereby achieving the technical effect of reducing the redundant calculation caused by the expansion operation in the fine-tuning process, reducing the consumption of computing resources and improving the efficiency of model training, thereby solving the technical problem of redundant attention calculation on the data matrix constructed by the expansion operation in the related art, resulting in large consumption of computing resources and low training efficiency.

[0199] It can be understood by those skilled in the art that Figure 11 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, and a mobile Internet device (MID). Figure 11 It does not limit the structure of the above electronic device. For example, the electronic device 110 may also include Figure 11 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 11 Different configurations shown.

[0200] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.

[0201] Example 6

[0202] According to an embodiment of the present application, a computer-readable storage medium is further provided. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the data processing method provided in the above embodiment 1, embodiment 2 or embodiment 3.

[0203] Optionally, in this embodiment, the above storage medium may be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0204] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining data to be processed, where the data to be processed is data obtained by expanding the data sequence length of the original input data; based on the pre-recorded data sequence length information, selecting the original input data from the data to be processed, where the data sequence length information is used to record the initial sequence length corresponding to the original input data; invoking multiple threads on a preset type of processor to perform target calculations on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0205] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining multiple initial sequences based on the input batch information of the original input data, where at least some of the multiple initial sequences have different sequence lengths; performing data sequence length expansion on the multiple initial sequences to obtain data to be processed, where the data to be processed includes: an input matrix to be used and a mask matrix, the input matrix includes: multiple target sequences, the multiple target sequences have the same sequence length, and the mask matrix is used to determine the data category of each matrix element in the input matrix.

[0206] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing area determination on each data element in the data to be processed based on the data sequence length information to obtain a determination result, where the determination result is used to indicate the data area where each data element in the data to be processed is currently located, and the data area includes: an original data area and an expanded data area; selecting the original input data from the data to be processed according to the determination result.

[0207] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: rejecting to perform target calculations on the data in the expanded data area and releasing the thread resources occupied by the data in the expanded data area.

[0208] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing a resizing operation on the input matrix based on the mask matrix to obtain the data sequence length information.

[0209] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: constructing an initial hidden layer state based on an input matrix; and updating the initial hidden layer state by using data sequence length information to obtain a target hidden layer state.

[0210] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining an initial matrix dimension according to the sequence lengths of multiple initial sequences, where the initial matrix dimension is used to determine the initial dimensions of multiple matrices for parameter target calculation, and the multiple matrices include: a query matrix, a key matrix, and a value matrix corresponding to the original input data; and updating the initial matrix dimension based on the data sequence length information and the target hidden layer state to obtain a target matrix dimension, where the target matrix dimension is used to determine the target dimensions of the multiple matrices.

[0211] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: calling multiple threads on a preset type of processor to load multiple matrices corresponding to the original input data from a preset storage area on the preset type of processor to registers on the preset type of processor, and performing target calculation on the multiple matrices stored in the registers to obtain a target calculation result.

[0212] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: writing the target calculation result to a preset storage area; and in response to successful writing of the target calculation result, restoring the target hidden layer state to the initial hidden layer state based on the data sequence length information.

[0213] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining a data processing request through a first application programming interface, where the request data carried in the data processing request includes: data to be processed, where the data to be processed is data obtained by performing data sequence length expansion on the original input data; and returning a data processing response through a second application programming interface, where the response data carried in the data processing response includes: a target calculation result, which is obtained by calling multiple threads on a preset type of processor to perform target calculation on the original input data, the number of threads of the multiple threads is determined by the data sequence length information, the original input data is selected from the data to be processed based on the data sequence length information, and the data sequence length information is used to record the initial sequence length corresponding to the original input data.

[0214] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining a currently input data processing dialogue request, wherein the request data carried in the data processing dialogue request includes: data to be processed, wherein the data to be processed is data obtained by expanding the data sequence length of the original input data; in response to the data processing dialogue request, returning a data processing dialogue reply, wherein the information carried in the data processing dialogue reply includes: a target calculation result, wherein the target calculation result is obtained by calling multiple threads on a preset type of processor to perform a target calculation on the original input data, the number of threads of the multiple threads is determined by the data sequence length information, the original input data is selected from the data to be processed based on the data sequence length information, and the data sequence length information is used to record the initial sequence length corresponding to the original input data; and displaying the target calculation result in a graphical user interface.

[0215] Using the embodiment of the present application, a computer-readable storage medium for implementing the above-mentioned data processing method is provided. Obtain data to be processed, wherein the data to be processed is the data obtained after the data sequence length is expanded for the original input data; based on the pre-recorded data sequence length information, select the original input data from the data to be processed, wherein the data sequence length information is used to record the initial sequence length corresponding to the original input data; call multiple threads on the preset type processor to perform target calculation on the original input data to obtain the target calculation result, wherein the number of threads of the multiple threads is determined by the data sequence length information. In the embodiment of the present application, according to the initial sequence length corresponding to the original input data, attention calculation is performed on part of the data in the data to be processed (that is, the original input data), reducing the redundant calculation of the expanded data in the data to be processed. Thus, the present application achieves the purpose of reducing computational consumption by performing attention calculation on the original input data in the data to be processed through the data sequence length information, thereby achieving the technical effect of reducing the redundant calculation caused by the expansion operation in the fine-tuning process, reducing the consumption of computing resources and improving the efficiency of model training, thereby solving the technical problem of redundant attention calculation on the data matrix constructed by the expansion operation in the related art, resulting in large consumption of computing resources and low training efficiency.

[0216] According to an embodiment of the present application, a computer program product is further provided. Optionally, in this embodiment, the computer program product can provide data processing services based on the data processing method provided in the above embodiment 1, embodiment 2 or embodiment 3.

[0217] Optionally, in this embodiment, the computer program product may be a set of instructions and codes pre-written according to the data processing method. The computer program product may run on various computer platforms, including personal computers, servers, mobile devices, etc.

[0218] Optionally, in this embodiment, the instructions and codes corresponding to the computer program product are used to implement the following method steps: obtaining data to be processed, where the data to be processed is data obtained by expanding the data sequence length of the original input data; based on the pre-recorded data sequence length information, selecting the original input data from the data to be processed, where the data sequence length information is used to record the initial sequence length corresponding to the original input data; calling multiple threads on a preset type of processor to perform target calculations on the original input data to obtain a target calculation result, where the number of threads of the multiple threads is determined by the data sequence length information.

[0219] Through the above computer program product, in an application scenario involving data sequence length expansion calculation during model training, data processing services can be provided, achieving the purpose of reducing calculation consumption by performing attention calculations on the original input data in the data to be processed according to the data sequence length information, thereby realizing the technical effects of reducing redundant calculations brought by the expansion operation during the fine-tuning process, reducing computing resource consumption, and improving the model training efficiency, and further solving the technical problem in the related art that redundant attention calculations are performed on the data matrix constructed by the expansion operation, resulting in large computing resource consumption and low training efficiency.

[0220] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0221] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0222] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.

[0223] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0224] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0225] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, ROM, RAM, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0226] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A data processing method, characterized in that: include: Acquiring data to be processed, wherein the data to be processed is data obtained by expanding the data sequence length of original input data; Selecting the original input data from the data to be processed based on pre-recorded data sequence length information, wherein the data sequence length information is used to record an initial sequence length corresponding to the original input data; Calling multiple threads on a preset type processor to perform target calculation on the original input data to obtain a target calculation result, wherein the number of threads of the multiple threads is determined by the data sequence length information.

2. The data processing method according to claim 1, wherein: Acquiring the data to be processed includes: determining a plurality of initial sequences based on input batch information of the original input data, wherein at least some of the plurality of initial sequences have different sequence lengths; The data sequence lengths of the multiple initial sequences are expanded to obtain the data to be processed, wherein the data to be processed includes: an input matrix to be used and a mask matrix, the input matrix includes: multiple target sequences, the multiple target sequences have the same sequence length, and the mask matrix is used to determine the data category of each matrix element in the input matrix.

3. The data processing method according to claim 1, wherein: Selecting the original input data from the data to be processed based on the data sequence length information includes: Based on the data sequence length information, performing region determination on each data element in the data to be processed to obtain a determination result, wherein the determination result is used to indicate the data region where each data element in the data to be processed is currently located, and the data region includes: an original data region and an expanded data region; The original input data is selected from the data to be processed according to the determination result.

4. The data processing method according to claim 3, wherein: The data processing method further includes: Refuse to execute target calculation on the data in the extended data area, and release thread resources occupied by the data in the extended data area.

5. The data processing method according to claim 2, wherein: The data processing method further includes: A resizing operation is performed on the input matrix based on the mask matrix to obtain the data sequence length information.

6. The data processing method according to claim 2, wherein: The data processing method further includes: constructing an initial hidden layer state based on the input matrix; The initial hidden layer state is updated using the data sequence length information to obtain a target hidden layer state.

7. The data processing method according to claim 6, characterized in that: The data processing method further includes: Determining initial matrix dimensions based on sequence lengths of the multiple initial sequences, wherein the initial matrix dimensions are used to determine initial dimensions of multiple matrices participating in target calculation, the multiple matrices including: a query matrix, a key matrix, and a value matrix corresponding to the original input data; The initial matrix dimension is updated based on the data sequence length information and the target hidden layer state to obtain a target matrix dimension, wherein the target matrix dimension is used to determine the target dimensions of the multiple matrices.

8. The data processing method according to claim 7, characterized in that: Calling the multiple threads on the preset type processor to perform target calculation on the original input data to obtain the target calculation result includes: Calling the multiple threads on the preset type processor, loading the multiple matrices corresponding to the original input data from a preset storage area on the preset type processor to registers on the preset type processor, and performing target calculations on the multiple matrices stored in the registers to obtain the target calculation results.

9. The data processing method according to claim 8, characterized in that: The data processing method further includes: Writing the target calculation result into the preset storage area; In response to the target calculation result being written successfully, the target hidden layer state is restored to the initial hidden layer state based on the data sequence length information.

10. A data processing method, characterized in that: include: Obtaining a data processing request through a first application programming interface, wherein the request data carried in the data processing request includes: data to be processed, wherein the data to be processed is data obtained by expanding the data sequence length of original input data; A data processing response is returned through a second application programming interface, wherein the response data carried in the data processing response includes: a target calculation result, the target calculation result is obtained by calling multiple threads on a preset type of processor to perform a target calculation on the original input data, the number of threads of the multiple threads is determined by data sequence length information, the original input data is selected from the data to be processed based on the data sequence length information, and the data sequence length information is used to record an initial sequence length corresponding to the original input data.

11. A data processing method, characterized in that: include: Obtaining a currently input data processing session request, wherein the request data carried in the data processing session request includes: data to be processed, wherein the data to be processed is data obtained by expanding the data sequence length of the original input data; In response to the data processing dialogue request, a data processing dialogue reply is returned, wherein the information carried in the data processing dialogue reply includes: a target calculation result, the target calculation result being obtained by invoking multiple threads on a preset type of processor to perform the target calculation on the original input data, the number of the multiple threads being determined by data sequence length information, the original input data being selected from the data to be processed based on the data sequence length information, and the data sequence length information being used to record an initial sequence length corresponding to the original input data; The target calculation results are displayed in a graphical user interface.

12. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the data processing method according to any one of claims 1 to 11 when running.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the data processing method according to any one of claims 1 to 11.

14. A computer program product, characterized in that The computer program comprises a computer program which, when executed by a processor, implements the data processing method according to any one of claims 1 to 11.