Data processing device, system, data processing method and computer device

By splitting the data processing tasks of the large language model into multiple target processing blocks and using multiple data processing units to process data with different parameters, the problem of parameter reading consumption time and resources is solved, which improves task processing efficiency and reduces hardware complexity.

CN118798362BActive Publication Date: 2025-07-08BEIJING PINGXIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411071145.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2025-07-08
Estimated Expiration
2044-08-06

AI Technical Summary

Technical Problem

When processing data from large language models, the parameter reading process consumes more time and resources, resulting in a decrease in task processing efficiency.

Method used

The data processing tasks of the large language model are split into different target processing blocks, and multiple data processing units are used to perform the same data processing tasks with different processing parameters, and the result data is merged by the merging unit to reduce the repeated reading of parameters during the task processing.

Benefits of technology

It improves the efficiency of data processing tasks, reduces the complexity of hardware development, and enhances the hardware's adaptability to different large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118798362B_ABST
    Figure CN118798362B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing device, system, method, and computer device; the data processing device includes: a data input unit, a plurality of processing units, and a merging unit; the data input unit is configured to obtain data to be processed corresponding to a target processing block in a large language model and transmit the data to be processed to the plurality of data processing units; each of the plurality of processing units is configured to perform a target process corresponding to the data processing function of the target processing block on the data to be processed to obtain result data corresponding to each of the processing units; wherein different processing units use different processing parameters when performing the target process on the data to be processed, and the processing methods are the same; the merging unit is configured to perform a merging process on the result data respectively output by the plurality of processing units to obtain target data obtained by processing the data to be processed using the data processing function corresponding to the target processing block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer hardware technology, and in particular, to a data processing device, a system, a data processing method, and a computer device. Background Art

[0002] As the core driving force of the new round of scientific and technological revolution and industrial transformation, artificial intelligence is rapidly giving birth to new products, new services, and new business forms, reshaping the economic and social operation mode, and changing the human production and lifestyle. As an important achievement of artificial intelligence, large language models are widely used in many fields. And large language models rely on hardware when performing processing tasks; the processing efficiency of the hardware affects the task processing efficiency of large language models. Summary of the Invention

[0003] The embodiments of the present disclosure provide at least a data processing device, a system, a data processing method, and a computer device.

[0004] In a first aspect, an embodiment of the present disclosure provides a data processing device, including: a data input unit, a plurality of processing units, and a merging unit;

[0005] The data input unit is configured to obtain the data to be processed corresponding to the target processing block of the large language model and transmit the data to be processed to a plurality of data processing units; the large language model includes: a plurality of network layers composed of an attention network and a feed-forward neural network; the target processing block includes at least one of the following: the network layer, the attention network, the feed-forward neural network, a first sub-network divided by the attention network, and a second sub-network divided by the feed-forward neural network;

[0006] Each of the plurality of processing units is configured to perform a target processing corresponding to the data processing function of the target processing block on the data to be processed, and obtain the result data corresponding to each processing unit; wherein, different processing units use different processing parameters when performing the target processing on the data to be processed, and the processing methods are the same;

[0007] The merging unit is configured to perform a merging process on the result data respectively output by the plurality of processing units to obtain the target data corresponding to the target processing block.

[0008] In a possible implementation manner, it further includes: a broadcast unit;

[0009] When transmitting the data to be processed to a plurality of data processing units, the data input unit is configured to:

[0010] Send the data to be processed to the broadcast unit;

[0011] The broadcast unit is configured to broadcast the data to be processed to a plurality of processing units in response to receiving the data to be processed transmitted by the data input unit, according to the pre-established connection relationship between the broadcast unit and the plurality of processing units.

[0012] In a possible implementation, before transmitting the data to be processed to a plurality of data processing units, the data input unit is further configured to:

[0013] Perform segmentation processing on the data to be processed to obtain multiple sets of sub-data to be processed; different data to be processed correspond to different processing cycles;

[0014] When transmitting the data to be processed to a plurality of processing units, the data input unit is configured to:

[0015] In each processing cycle of the multiple processing cycles, transmit the sub-data to be processed corresponding to each processing cycle to the plurality of processing units.

[0016] In a possible implementation, the processing unit includes: an arithmetic matrix; the arithmetic matrix includes a plurality of arithmetic units composed of hardware circuits;

[0017] When performing segmentation processing on the data to be processed to obtain multiple sets of sub-data to be processed, the data input unit is configured to:

[0018] Determine the matrix width of the data to be processed;

[0019] Compare the matrix width with the number of arithmetic units in the arithmetic matrix;

[0020] In the case where the matrix width is less than or equal to the number of arithmetic units, perform row-wise segmentation on the matrix formed by the data to be processed to obtain multiple sets of the sub-data to be processed;

[0021] In the case where the matrix width is greater than the number of arithmetic units, perform column-wise segmentation on the matrix formed by the data to be processed to obtain multiple sets of the sub-data to be processed.

[0022] In a possible implementation, the arithmetic unit includes at least one of the following: a multiply-accumulate unit, a comparator, an accumulator, and a divider.

[0023] In a possible implementation, when merging the result data respectively output by the plurality of processing units to obtain the target data corresponding to the data to be processed, the merging unit is configured to:

[0024] In each of the multiple processing cycles, perform a first merging process on the result data corresponding to each processing cycle to obtain the merged result data corresponding to each processing cycle;

[0025] Perform a second merging process on the merged result data respectively corresponding to the multiple processing cycles to obtain the target data.

[0026] In a possible implementation manner, it further includes: a configuration unit;

[0027] The configuration unit is used to configure processing parameters for the multiple processing units, and store the processing parameters corresponding to each processing unit into the storage space associated with the processing unit;

[0028] and / or, used to configure the merging method for the merging unit.

[0029] In a second aspect, an embodiment of the present disclosure further provides a data processing system, including:

[0030] Multiple data processing devices according to the first aspect or any one of the first aspect, and a controller;

[0031] The controller is used to divide the large language model into multiple processing blocks based on the model structure of the large language model and the number of data processing devices in the data processing system, and determine the data processing devices for establishing mapping relationships for the multiple processing blocks; and, for each data processing device, deploy the processing parameters of the processing blocks associated with each data processing device to each data processing device;

[0032] Each data processing device among the multiple data processing devices is used to execute the data processing tasks of the processing blocks for which mapping relationships are established.

[0033] In a third aspect, an embodiment of the present disclosure further provides a data processing method, including:

[0034] Use a data input unit to obtain the data to be processed corresponding to the target processing block of the large language model; the large language model includes: multiple network layers composed of an attention network and a feedforward neural network; the target processing block includes at least one of the following: the network layer, the attention network, the feedforward neural network, the first subnetwork divided by the attention network, and the second subnetwork divided by the feedforward neural network;

[0035] Using each of multiple processing units, perform target processing corresponding to the data processing function of the target processing block on the data to be processed, and obtain result data corresponding to each processing unit; wherein, different processing parameters are used when different processing units perform the target processing on the data to be processed, and the processing methods are the same;

[0036] Use a merging unit to perform merging processing on the result data respectively output by the multiple processing units to obtain target data corresponding to the target processing block.

[0037] In a fourth aspect, an embodiment of the present disclosure further provides a computer device, including: the data processing device as described in the first aspect or any item of the first aspect, or the data processing system as described in the second aspect.

[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solutions of the present disclosure.

[0039] The data processing device provided by the embodiment of the present disclosure utilizes the structural consistency of the large language model and the weak coupling between different modules caused by the small data transmission volume between different structures constituting the large language model. According to the model structure of the large language model, the data processing task of the large language model is split into different target processing blocks. For each target processing block, multiple data processing units use different processing parameters that can be used to perform the same target processing on the data to be processed corresponding to the target processing block, and obtain the result data corresponding to each processing unit. Then, a merging unit is used to merge the result data corresponding to the multiple processing units respectively, so as to realize that the processing task of the same data is assigned to multiple different processing units to complete. When different processing units execute the processing task of the same data to be processed, different processing parameters are called. Furthermore, before using the data processing device to process the data processing task, the specific parameters corresponding to each processing unit can be passed into each processing unit. During the task processing process, the read parameters do not need to be replaced or re-passed, thereby reducing the time and resources consumed in the parameter reading process and improving the task processing efficiency.

[0040] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the accompanying drawings. Description of the Drawings

[0041] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. The accompanying drawings herein are incorporated into the specification and form a part of this specification. These accompanying drawings show embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these accompanying drawings without creative efforts.

[0042] Figure 1 The structural schematic diagram of a data processing device provided by some embodiments of the present disclosure is shown;

[0043] Figure 2 The structural schematic diagram of another data processing device provided by some embodiments of the present disclosure is shown;

[0044] Figure 3 The flowchart of a data processing method provided by some embodiments of the present disclosure is shown;

[0045] Figure 4 The structural schematic diagram of a computer device provided by some embodiments of the present disclosure is shown. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present disclosure. The components of the embodiments of the present disclosure described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure claimed, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0047] Through research, it has been found that when using hardware to process data of large language models, in order to improve the processing efficiency of data processing, a distributed processing method is usually used to divide data processing into multiple different subtasks and deliver different subtasks to different computing units for processing. The computing unit usually includes an arithmetic unit array; during the process of executing a computing task, the computing unit needs to read the data to be processed and the relevant parameters of the large language model from the memory into the arithmetic unit array, use the arithmetic unit array to perform corresponding arithmetic processing on the data to be processed and the relevant parameters, and output the processing results. In this process, there are usually many relevant parameters; different parameters need to be read into the arithmetic unit array in different processing cycles to implement the data processing task. This results in a relatively long time and a large amount of resources being consumed in the parameter reading process, leading to a decrease in the efficiency of task processing.

[0048] Based on the above research, the present disclosure provides a data processing device. The data processing device provided by the embodiments of the present disclosure utilizes the consistency in the structure of the large language model and the weak coupling between different modules caused by the small amount of data transmission between different structures constituting the large language model. According to the model structure of the large language model, the data processing task of the large language model is split into different target processing blocks. For each target processing block, multiple data processing units can use different processing parameters to perform the same target processing on the data to be processed corresponding to the target processing block, and obtain the result data corresponding to each processing unit. Then, a merging unit is used to merge the result data corresponding to multiple processing units respectively, so as to achieve the distribution of the processing task of the same data to multiple different processing units for completion. When different processing units execute the processing task of the same data to be processed, different processing parameters are called. Therefore, before using the data processing device to process the data processing task, the specific parameters corresponding to each processing unit can be passed into each processing unit. During the task processing, the read parameters do not need to be replaced or re-passed, thereby reducing the time and resources consumed in the parameter reading process and improving the efficiency of task processing.

[0049] At the same time, the embodiments of the present disclosure also provide a data processing system, which distributes complex large-scale operations to multiple data processing devices for completion, reduces the complexity of hardware development, increases the ability of the hardware to match different large language models, and enables the data processing device to have a wider range of applications.

[0050] All the defects existing in the above solutions are the results obtained by the inventors through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the present disclosure for the above problems in the following text should be the contributions made by the inventors to the present disclosure during the process of the present disclosure.

[0051] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined or explained in subsequent figures.

[0052] To facilitate the understanding of this embodiment, a data processing device disclosed in the embodiments of the present disclosure will be introduced in detail first.

[0053] See Figure 1 As shown, it is a schematic structural diagram of a data processing device provided by the embodiments of the present disclosure, including: a data input unit 10, a plurality of processing units 20, and a merging unit 30.

[0054] The data input unit 10 is configured to obtain the data to be processed corresponding to the target processing block in the large language model and transmit the data to be processed to a plurality of data processing units; the large language model includes: at least one network layer composed of an attention network and a feed-forward neural network; the target processing block includes at least one of the following: the network layer, the attention network, the feed-forward neural network, a first sub-network divided by the attention network, and a second sub-network divided by the feed-forward neural network;

[0055] Each of the plurality of processing units 20 is configured to perform a target process corresponding to the target processing block on the data to be processed to obtain result data corresponding to each processing unit; wherein, different processing units use different processing parameters when performing the target process on the data to be processed, and the processing methods are the same;

[0056] The merging unit 30 is configured to perform a merging process on the result data respectively output by the plurality of processing units to obtain target data corresponding to the target processing block.

[0057] In a specific implementation, large language models (LLMs) usually take text and / or images as inputs and perform processing tasks corresponding to the large language model on the text and / or images. For example, large language models can be used to generate and match text with images, videos, etc., or use the input text and images to use certain image features in the input image to generate new images or videos, etc.

[0058] Large language models typically include at least one network layer composed of an attention network and feed-forward neural networks (FFNs). In different large language models, the number of attention networks and FFNs included in the network layer may vary. For example, one attention network can be connected to multiple parallel and / or serial FFNs; or multiple attention networks are connected in parallel and then connected in series with at least one FFN, etc. The specific connection relationship is not limited in the embodiments of the present disclosure.

[0059] A large language model can be split into multiple processing blocks. According to the structure of the large language model, any one of the processing blocks can include, for example, any one of the following: a network layer of the large language model, the attention network in the network layer, the FFN in the network layer, any one of the first sub-networks after splitting the attention network into multiple first sub-networks, and any one of the second sub-networks after splitting the FFN into multiple second sub-networks.

[0060] The target processing block can be any one of the multiple processing blocks. The target processing block corresponds to a specific data processing task. Depending on the different target processing blocks, the corresponding data processing tasks can also include any one of the following:

[0061] The data processing task corresponding to the network layer.

[0062] The data processing task corresponding to the attention network.

[0063] The data processing task corresponding to the feed-forward neural network.

[0064] The data processing task corresponding to the first sub-network divided from the attention network.

[0065] The data processing task corresponding to the second sub-network divided from the feed-forward neural network.

[0066] The data processing tasks corresponding to different processing blocks are characterized by a small amount of data to be processed and a large number of processing parameters. For example, in the attention network, each group of attention weights in multiple groups of attention weights needs to be used to perform attention processing on the same data to be processed to obtain the result of attention processing.

[0067] For the feed-forward neural network, multiple groups of weight parameters need to be used to perform weighted summation processing on the data to be processed.

[0068] Based on this, the data to be processed in the embodiments of the present disclosure includes, for example, the original data processed by the large language model or the data output by other processing blocks upstream of the target processing block. This data to be processed is usually in the form of vectors or matrices.

[0069] Exemplarily, the original data processed using a large language model, for example, includes: text vectors obtained after converting text into vectors, and / or images. Here, the image data itself is a matrix composed of pixel values of multiple pixels.

[0070] Since the data to be processed is of vector or matrix type, and in order to simplify the circuit structure and reduce the circuit design difficulty, in the embodiments of the present disclosure, the size of the matrix or vector of the data to be processed transmitted to the data processing device is limited. In the case where a single data processing device cannot process all the data output by the previous network layer or the original data due to complex processing processes or low processing efficiency, etc., before the data is transmitted to the data processing device, the data can usually be first segmented to obtain multiple different data blocks, and in a way of connecting multiple data processing devices in parallel, multiple data processing devices are used to process different data blocks, and each data block is the data to be processed by each data processing device.

[0071] When the data input unit 10 obtains the data to be processed, for example, it can be notified by an upstream data processing node and read from the memory of a computer device where a large language model is deployed; or, it can also passively receive the data to be processed sent by the upstream data processing node. After the data input unit 10 obtains the data to be processed, it converts the read data to be processed into an input data stream and transmits it to each processing unit 20.

[0072] In a possible implementation manner, the data input unit 10 can directly establish a hardware connection relationship with multiple processing units 20 and transmit the data to be processed to the multiple processing units 20 according to this connection relationship.

[0073] In addition, as Figure 2 shown, in another embodiment of the present disclosure, it further includes: a broadcast unit 40.

[0074] When the data input unit transmits the data to be processed to multiple data processing units, it is used for:

[0075] Sending the data to be processed to the broadcast unit;

[0076] The broadcast unit 40 is used for, in response to receiving the data to be processed transmitted by the data input unit, broadcasting the data to be processed to multiple processing units according to the pre-established connection relationship between the broadcast unit and the multiple processing units.

[0077] In a specific implementation, the broadcast unit 40 can be an independent hardware structure, and a hardware connection relationship is pre-established with multiple processing units, and it can transmit the data to be processed transmitted by the data input unit 10 to multiple processing units in a broadcast transmission manner according to this connection relationship.

[0078] Before transmitting the data to be processed to multiple processing units 20, the data input unit 10 is further configured to:

[0079] Perform segmentation processing on the data to be processed to obtain multiple groups of sub-data to be processed; different data to be processed correspond to different processing cycles.

[0080] When the data input unit 10 transmits the data to be processed to multiple processing units, it is configured to:

[0081] In each processing cycle among multiple processing cycles, transmit the sub-data to be processed corresponding to each processing cycle to the multiple processing units 20.

[0082] Here, when the data input unit 10 performs segmentation processing on the data to be processed, for example, it segments the data stream converted from a matrix or a vector according to a certain step size to generate multiple groups of sub-data to be processed.

[0083] The processing unit 20 is usually integrated with an operation matrix. The operation matrix includes multiple arithmetic units composed of hardware circuits. The arithmetic unit may include at least one of the following: a multiplier-accumulator, a comparator, an accumulator, and a divider, etc.

[0084] These arithmetic units usually form an operation matrix composed of multiple arithmetic units according to a certain connection method, and the operation matrix can implement the processing of the data to be processed.

[0085] Furthermore, in the embodiment of the present disclosure, when the data input unit 10 performs segmentation processing on the data to be processed to obtain multiple groups of sub-data to be processed, it is configured to:

[0086] Determine the matrix width of the data to be processed;

[0087] Compare the matrix width with the number of arithmetic units in the operation matrix;

[0088] In the case where the matrix width is less than the number of arithmetic units, perform row-wise segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed;

[0089] In the case where the matrix width is greater than or equal to the number of arithmetic units, perform column-wise segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed. Generally, in order to reduce the complexity of the operation, the number of arithmetic units in each processing unit is greater than the number of columns of the matrix formed by the data to be processed.

[0090] Assume that the data to be processed is of size m A matrix of m×n, where m and n are the number of rows and columns of the matrix. Assume that each processing unit 20 includes s arithmetic units. s and n can be compared. If n≤s, the matrix formed by the data to be processed is divided row by row to obtain multiple groups of the data to be processed sub-data. That is, the data to be processed is unfolded row by row to form a data stream, and the data stream is divided with the number of columns as the segmentation step length to obtain multiple groups of data to be processed sub-data.

[0091] If n>s, the matrix formed by the data to be processed is divided column by column to obtain multiple groups of the data to be processed sub-data. That is, the data to be processed is unfolded column by column to form a data stream, and the data stream is divided with the number of rows as the segmentation step length to obtain multiple groups of data to be processed sub-data.

[0092] For example, for a matrix K of m ×n:

[0093] , when the matrix is unfolded row by row, the obtained data stream can be expressed as: . Set the segmentation step length to n, and m groups of data to be processed sub-data are obtained, which are respectively: , 、……、 . When the matrix is unfolded column by column, the obtained data stream is expressed as: , set the segmentation step length to m, and n groups of data to be processed sub-data are obtained, which are respectively: 、 、……、 .

[0094] Then, in different processing cycles, the obtained groups of data to be processed sub-data are sequentially sent to the processing unit 20, so that the processing unit 20 processes each data to be processed sub-data with the same input data but different parameters in each processing cycle.

[0095] In addition, assume that the data to be processed is a matrix of size m ×n ×k, where m and n are the number of rows and columns of the matrix, and k represents the number of channels of the matrix. For example, each channel can be unfolded row by row or column by column in ascending order of channel number to obtain a data stream, and each channel will form a vector of length m ×n; k vectors of m ×n are spliced together, that is, a data stream is formed. Then, the data stream is divided and processed in the above manner to generate multiple groups of data to be processed sub-data.

[0096] Take the case where the data input unit 10 directly transmits the data to be processed to multiple processing units 20: In the above process, the data input unit 10 can process the data to be processed in a way of acquiring and splitting simultaneously. That is, when acquiring the data to be processed, when the amount of acquired data reaches a certain quantity, the data of this quantity is taken as a group of sub-data to be processed, and this group of processed result data is broadcast to multiple downstream processing units. At the same time, the data input unit 10 can also receive subsequent data to be processed simultaneously; when the amount of received data reaches a certain quantity again, these data are taken as a new group of sub-data to be processed, and after the next processing cycle arrives, they are transmitted to multiple processing units 20.

[0097] Take the case where the data input unit 10 transmits the data to be processed to the broadcast unit 40, and the broadcast unit 40 then transmits the data to be processed to multiple processing units 20. In the above process, the data input unit 10 can also process the data to be processed in a way of acquiring and splitting simultaneously. When acquiring the data to be processed, when the amount of acquired data reaches a certain quantity, the data of this quantity is taken as a group of sub-data to be processed and sent to the broadcast unit 40. In response to receiving this group of sub-data to be processed, the broadcast unit 40 broadcasts the received group of sub-data to multiple processing units 20 according to the physical connection relationship with multiple processing units 20. At the same time, the data input unit 10 can also receive subsequent data to be processed simultaneously; when the amount of received data reaches a certain quantity again, these data are taken as a new group of sub-data to be processed, and after the next processing cycle arrives, they are transmitted to the broadcast unit 40. After receiving the new group of sub-data to be processed, the broadcast unit 40 broadcasts this new group of sub-data to multiple processing units 20.

[0098] After receiving a group of sub-data to be processed sent by the data input unit 10, multiple processing units 20 use the pre-configured parameters to perform corresponding processing on the sub-data to be processed, and send the obtained processing results to the merging unit 30. When the next processing cycle arrives, if new sub-data to be processed is received, continue to use the same parameters to perform corresponding processing on the new group of sub-data to be processed, obtain new processing results, and send them to the merging unit 30.

[0099] Here, the pre-configured parameters in the processing unit 20 include, for example, the internal parameters of the large language model. Assume that the data processing device is used to process the data processing task of a certain convolutional layer of the large language model, then the internal parameters of the model include: convolutional kernels; assume that the data processing device is used to process the data processing task of a certain fully connected layer of the large language model, then the internal parameters of the model include fully connected weights. Specifically, according to the different data tasks processed by the data processing device, the parameters will also be different.

[0100] The merging unit 30, in each of the multiple processing cycles, performs a first merging process on the result data corresponding to each processing cycle to obtain the merged result data corresponding to each processing cycle; and performs a second merging process on the merged result data corresponding to the multiple processing cycles respectively to obtain the target data.

[0101] In each processing cycle, the merging unit will receive the processing results of the same set of sub-data to be processed sent by multiple processing units 20 respectively. Then, the processing results of the same set of sub-data to be processed sent by multiple processing units 20 are merged to obtain the processing result of the sub-data to be processed within this processing cycle.

[0102] Each of the multiple processing units 20 independently implements data processing. This independence is reflected in two aspects. First, there is no direct communication between the processing units. Second, the processing units receive the same input, use different parameters to perform calculations with the input data, and output different calculation results.

[0103] Each processing unit 20 can integrate an operation matrix for implementing the corresponding function according to the specific processing function to be implemented. The operation matrix includes multiple arithmetic units composed of hardware circuits, such as multipliers, adders, comparators, accumulators, and dividers, etc. The specific type of hardware circuit to be integrated can be determined according to needs, and the embodiments of the present disclosure do not make limitations.

[0104] In the embodiments of the present disclosure, different processing units 20 can be deployed on the same chip or on different chips, which can be specifically determined according to the processing requirements of the large language model and the hardware design requirements, and the embodiments of the present disclosure do not make limitations.

[0105] After multiple processing units 20 perform target processing on the data to be processed and obtain the result data corresponding to each processing unit, they respectively transmit the result data to the merging unit 30.

[0106] The merging unit 30 performs a merging process on the data to be processed to obtain the target data.

[0107] Specifically, when the merging unit 30 performs a merging process on the result data respectively output by the multiple processing units to obtain the target data corresponding to the data to be processed, it is used for:

[0108] In each of the multiple processing cycles, performing a first merging process on the result data corresponding to each processing cycle to obtain the merged result data corresponding to each processing cycle;

[0109] Performing a second merging process on the merged result data corresponding to the multiple processing cycles respectively to obtain the target data.

[0110] Here, the first merging process may include, for example: splicing, cumulative summation, multiplication, averaging, finding the maximum value, finding the minimum value, etc. The second merging process may also include any one of splicing, cumulative summation, multiplication, averaging, finding the maximum value, finding the minimum value, etc.

[0111] The specific processing types of the first merging process and the second merging process may be the same or different, and are specifically set according to actual processing needs.

[0112] According to the different specific positions of the data processing device in the target network, after obtaining the target data, the merging unit 30 may use the target data as the result data of the large language model, or may also transmit it to the data processing device for performing the data processing task of the next network layer.

[0113] In the data processing device provided in another embodiment of the present disclosure, a configuration unit 50 is further included.

[0114] The configuration unit is used to configure the processing parameters of the multiple processing units, and store the processing parameters corresponding to each processing unit into the storage space associated with the processing unit;

[0115] And / or, used to configure the merging method of the merging unit.

[0116] The configuration unit 50 usually configures the processing data to each processing unit before using the data processing device to process the data processing task. And / or configure the merging method into the merging unit.

[0117] In another embodiment of the present disclosure, a data processing system is further provided, including: a plurality of data processing devices as described in any embodiment of the present disclosure, and a controller;

[0118] The controller is used to divide the large language model into multiple target processing blocks based on the model structure of the large language model and the number of data processing devices in the data processing system, and determine the data processing devices for establishing the mapping relationship for the multiple target processing blocks; and, for each data processing device, deploy the processing parameters of the target processing block associated with each data processing device to each data processing device;

[0119] Each data processing device among the plurality of data processing devices is used to execute the data processing task of the target processing block for which the mapping relationship is established.

[0120] In specific implementations, given a certain model structure of the large language model, the number of data processing devices in the data processing system affects the data processing efficiency of the data processing system. The larger the number, the finer the granularity of the data processing tasks corresponding to the large language model can be divided, so that the data processing tasks of the large language model can be distributed to more data processing devices for synchronous execution, thereby improving the efficiency; while the smaller the number, in order to meet the data processing requirements of the large language model, more multiplexing of each data processing device is required to achieve the execution of more data processing tasks by each data processing device, thereby reducing the data processing efficiency. At the same time, as the model volume of the large language model varies, the larger the model volume of the large language model, the more complex the data processing tasks. Therefore, in the embodiments of the present disclosure, when deploying the large language model to the data processing system, the controller needs to divide the large language model into multiple processing blocks according to the model structure of the large language model and the number of data processing devices, and determine the data processing devices that establish mapping relationships for the multiple processing blocks. Specifically, when dividing, for example, based on the number of data processing devices, first determine whether the number of data processing devices is greater than the number of network layers in the large language model; if it is greater, then determine the network layer in the large language model that requires more complex data processing tasks, and then further divide the attention network and feedforward neural network in this network layer.

[0121] The cascading between different data processing devices can be executed in a pipeline manner, for example, that is, the output result of the previous-level data processing device is used as the data to be processed by the next-level data processing device and transmitted to the next-level data processing device. Alternatively, a one-to-many cascading method can also be used, that is, the output result of the previous-level data processing device will be split into multiple sub-data parts, and different sub-data parts are respectively transmitted to multiple different next-level data processing devices.

[0122] In addition, in the data processing system provided by the embodiments of the present disclosure, for example, a judgment module can also be included, and this judgment module is used to determine the mutual relationship between the cascaded data processing devices.

[0123] Based on the same inventive concept, embodiments of the present disclosure also provide a data processing method corresponding to the data processing device. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to that of the above-mentioned data processing device in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0124] See Figure 3 As shown, embodiments of the present disclosure also provide a data processing method, including:

[0125] S301: The control data input unit acquires the data to be processed corresponding to the target processing block of the large language model; the large language model includes at least one network layer composed of an attention network and a feed-forward neural network; the target processing block includes at least one of the following: the network layer, the attention network, the feed-forward neural network, a first sub-network divided by the attention network, and a second sub-network divided by the feed-forward neural network;

[0126] S302: Control each processing unit among the multiple processing units to perform a target process corresponding to the data processing function of the target processing block on the data to be processed, and obtain result data corresponding to each processing unit; wherein, different processing units use different processing parameters when performing the target process on the data to be processed, and the processing methods are the same;

[0127] S303: Control the merging unit to perform a merging process on the result data respectively output by the multiple processing units to obtain target data corresponding to the target processing block.

[0128] In a possible implementation manner, sending the result data to the merging unit includes:

[0129] Sending the data to be processed to the broadcast unit;

[0130] In response to receiving the data to be processed transmitted by the data input unit, the broadcast unit broadcasts the data to be processed to the multiple processing units according to the pre-established connection relationship between the broadcast unit and the multiple processing units.

[0131] In a possible implementation manner, the method further includes:

[0132] The data input unit performs a splitting process on the data to be processed to obtain multiple groups of sub-data to be processed; different data to be processed correspond to different processing cycles;

[0133] Transmitting the data to be processed to the multiple processing units includes:

[0134] In each of the multiple processing cycles, transmitting the sub-data to be processed corresponding to each processing cycle to the multiple processing units.

[0135] In a possible implementation manner, performing the splitting process on the data to be processed to obtain multiple groups of sub-data to be processed includes:

[0136] Determining the matrix width formed by the data to be processed;

[0137] Comparing the matrix width with the number of arithmetic units in the arithmetic matrix;

[0138] When the width of the matrix is less than or equal to the number of the arithmetic units, the matrix formed by the data to be processed is split row by row to obtain multiple groups of the sub-data to be processed;

[0139] When the width of the matrix is greater than or equal to the number of the arithmetic units, the matrix formed by the data to be processed is split column by column to obtain multiple groups of the sub-data to be processed.

[0140] In a possible implementation manner, the merging the result data respectively output by the multiple processing units to obtain the target data corresponding to the data to be processed includes:

[0141] In each processing cycle of the multiple processing cycles, performing a first merging process on the result data corresponding to each processing cycle to obtain the merged result data corresponding to each processing cycle;

[0142] Performing a second merging process on the merged result data respectively corresponding to the multiple processing cycles to obtain the target data.

[0143] In a possible implementation manner, the method further includes:

[0144] Configuring processing parameters for the multiple processing units, and storing the processing parameters corresponding to each processing unit into the storage space associated with the processing unit;

[0145] And / or, configuring the merging mode for the merging unit.

[0146] See Figure 4 As shown, an embodiment of the present disclosure further provides a computer device, including: a processor 41, an external memory 42, a memory 43, and a data processing device / data processing system 44 provided in any embodiment of the present disclosure;

[0147] The external memory 42 is used to store network parameters of a large language model;

[0148] The processor 41 is used to read the network parameters of the large language into the memory 43, and input the network parameters on the memory and the image to be processed into the data processing device / data processing system;

[0149] The data processing device / data processing system 44 is used to execute the data processing task of the large language model.

[0150] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0151] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0152] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0153] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0154] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting it. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A data processing device, characterized in that, Including: A data input unit, multiple processing units, and a merging unit; The data input unit is configured to obtain the data to be processed corresponding to the target processing block of the large language model and transmit the data to be processed to multiple data processing units; the large language model includes: at least one network layer composed of an attention network and a feed-forward neural network; the target processing block includes at least one of the following: the network layer, the attention network, the feed-forward neural network, a first sub-network divided by the attention network, and a second sub-network divided by the feed-forward neural network; Each of the multiple processing units is configured to perform a target processing corresponding to the target processing block on the data to be processed to obtain result data corresponding to each processing unit; wherein, different processing units use different processing parameters when performing the target processing on the data to be processed, but the processing methods are the same; The merging unit is configured to perform a merging process on the result data respectively output by the multiple processing units to obtain target data corresponding to the target processing block; Wherein, before the data to be processed is transmitted to the processing unit, it is segmented by the data input unit into multiple groups of sub-data to be processed; the multiple groups of sub-data to be processed are transmitted to the processing unit in different processing cycles; in the same processing cycle, the sub-data to be processed transmitted to the multiple processing units is the same group of sub-data to be processed; The processing unit includes: an operation matrix; the operation matrix includes multiple arithmetic units composed of hardware circuits; When the data input unit performs a segmentation process on the data to be processed to obtain multiple groups of sub-data to be processed, it is configured to: Determine the matrix width of the data to be processed; Compare the matrix width with the number of arithmetic units in the operation matrix; In the case where the matrix width is less than or equal to the number of arithmetic units, perform a row-wise segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed; In the case where the matrix width is greater than the number of arithmetic units, perform a column-wise segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed; When performing a row-wise segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed, expand the data to be processed row-wise to form a data stream, and segment the data stream with the number of columns as the segmentation step length to obtain multiple groups of sub-data to be processed; When performing a column-wise segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed, expand the data to be processed column-wise to form a data stream, and segment the data stream with the number of rows as the segmentation step length to obtain multiple groups of sub-data to be processed.

2. The data processing device according to claim 1, characterized in that It further includes: A broadcast unit; When the data input unit transmits the data to be processed to multiple data processing units, it is configured to: Send the data to be processed to the broadcast unit; The broadcast unit is configured to, in response to receiving the data to be processed transmitted by the data input unit, broadcast the data to be processed to multiple processing units according to the pre-established connection relationship between the broadcast unit and the multiple processing units.

3. The data processing device according to claim 1 or 2, characterized in that Before transmitting the data to be processed to a plurality of data processing units, the data input unit is further configured to: Segment the data to be processed to obtain multiple sets of sub-data to be processed; different data to be processed correspond to different processing cycles; When transmitting the data to be processed to a plurality of processing units, the data input unit is configured to: In each of the multiple processing cycles, transmit the sub-data to be processed corresponding to each processing cycle to the plurality of processing units.

4. The data processing device according to claim 1, wherein The arithmetic unit includes at least one of the following: a multiplier-accumulator, a comparator, an accumulator, and a divider.

5. The data processing device according to claim 1 or 4, characterized in that When merging the result data respectively output by the plurality of processing units to obtain the target data corresponding to the data to be processed, the merging unit is configured to: In each of the multiple processing cycles, perform a first merging process on the result data corresponding to each processing cycle to obtain the merged result data corresponding to each processing cycle; Perform a second merging process on the merged result data corresponding to the multiple processing cycles respectively to obtain the target data.

6. The data processing device according to any one of claims 1, 2, and 4, wherein Further included: A configuration unit; The configuration unit is configured to configure processing parameters for the plurality of processing units and store the processing parameters corresponding to each processing unit in the storage space associated with the processing unit; And / or, configure the merging method for the merging unit.

7. A data processing system, characterized in that, Including: The plurality of data processing devices according to any one of claims 1-6, and a controller; The controller is configured to divide the large language model into multiple target processing blocks based on the model structure of the large language model and the number of data processing devices in the data processing system, and determine the data processing devices for establishing mapping relationships for the multiple target processing blocks; And, for each data processing device, deploy the processing parameters of the target processing blocks associated with the each data processing device to the each data processing device; Each of the plurality of data processing devices is configured to execute the data processing tasks of the target processing blocks for which the mapping relationships are established.

8. A data processing method, characterized in that Including: Controlling the data input unit to obtain the data to be processed corresponding to the target processing block of the large language model; The large language model includes: at least one network layer composed of an attention network and a feed-forward neural network; the target processing block includes at least one of the following: the network layer, the attention network, the feed-forward neural network, the first sub-network divided by the attention network, and the second sub-network divided by the feed-forward neural network; Controlling each of the plurality of processing units to perform target processing corresponding to the data processing function of the target processing block on the data to be processed to obtain the result data corresponding to each processing unit; wherein, different processing units use different processing parameters and the same processing method when performing the target processing on the data to be processed; Controlling the merging unit to merge the result data respectively output by the plurality of processing units to obtain the target data corresponding to the target processing block; Among them, before the data to be processed is transmitted to the processing unit, it is divided into multiple groups of sub-data to be processed; the multiple groups of sub-data to be processed are transmitted to the processing unit in different processing cycles; in the same processing cycle, the sub-data to be processed transmitted to multiple processing units is the same group of sub-data to be processed; The processing unit includes: an operation matrix; the operation matrix includes multiple arithmetic units composed of hardware circuits; Performing a segmentation process on the data to be processed to obtain multiple groups of sub-data to be processed, including: Determining the matrix width formed by the data to be processed; Comparing the matrix width with the number of arithmetic units in the operation matrix; In the case where the matrix width is less than or equal to the number of arithmetic units, performing a row-by-row segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed; In the case where the matrix width is greater than the number of arithmetic units, performing a column-by-column segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed; When performing a row-by-row segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed, expanding the data to be processed row by row to form a data stream, and using the number of columns as the segmentation step length to segment the data stream to obtain multiple groups of sub-data to be processed; When performing a column-by-column segmentation on the matrix formed by the data to be processed to obtain multiple groups of the sub-data to be processed, expanding the data to be processed column by column to form a data stream, and using the number of rows as the segmentation step length to segment the data stream to obtain multiple groups of sub-data to be processed.

9. A computer device, characterized in that, Including: The data processing device according to any one of claims 1-6, or the data processing system according to claim 7.

Citation Information

Patent Citations

  • Model data processing method, readable medium and electronic equipment

    CN116302551A