Data processing method, device and equipment based on large language model and readable medium
By dividing and iterating the long-sequence data, the problems of communication delay and imbalance in the large language model are solved, and efficient data processing is achieved.
Patent Information
- Application Number
- CN202510544819.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
In the parallelization technology of large language models, the communication transmission of sequence parallel algorithms mainly relies on point-to-point transmission of blocks, resulting in high communication delay and imbalance in computing and communication.
By dividing long sequence data based on the graphics processor sequence, generating divided data groups, and iterating steps and data transmission are performed between the graphics processors, using bidirectional bandwidth sequence parallelism and multi-node data processing to reduce communication delay and resource waste.
It effectively reduces communication delay, avoids imbalance between computing and communication, and improves the efficiency and accuracy of data processing.
Smart Images

Figure CN120448118A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a data processing method, apparatus, device, and readable medium based on a large language model. Background Art
[0002] In the field of large language model parallelization, how to quickly and efficiently process long data sequences has become an important research topic. Currently, the common approach to processing long data sequences is to split the self-attention feedforward computation into multiple blocks using a sequence parallel algorithm. These blocks are then distributed across multiple devices connected in a ring topology for simultaneous computation and communication.
[0003] However, when using the above method to process long data sequences, the following technical problems often occur:
[0004] The communication transmission of sequential parallel algorithms mainly relies on point-to-point transmission of blocks, which introduces high communication latency. When the number of graphics processors involved in the calculation increases, the calculation time per step will decrease rapidly at a quadratic rate, while the communication volume only decreases linearly, which leads to an imbalance between calculation and communication.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose a data processing method, apparatus, electronic device, and computer-readable medium based on a large language model to solve one or more of the technical problems mentioned in the above background technology section.
[0008] In a first aspect, some embodiments of the present disclosure provide a data processing method based on a large language model, the method comprising: receiving long sequence data input by a target user; dividing the long sequence data based on the above-mentioned graphics processor sequence to generate a divided data group, wherein the divided data in the above-mentioned divided data group corresponds to the graphics processor in the above-mentioned graphics processor sequence; for each graphics processor in the above-mentioned graphics processor sequence, executing the following processing steps: determining the iteration step corresponding to the above-mentioned graphics processor; sending the divided data corresponding to the above-mentioned iteration step to the next graphics processor; determining the previous step data based on the above-mentioned iteration step, and sending the above-mentioned step data to the corresponding graphics processor; generating block output according to the corresponding divided data; and performing update synchronization operation on the above-mentioned block output to generate updated output data.
[0009] In a second aspect, some embodiments of the present disclosure provide a data processing device based on a large language model, the device comprising: a receiving unit configured to receive long sequence data input by a target user; a dividing unit configured to divide and process the long sequence data based on the above-mentioned graphics processor sequence to generate a divided data group, wherein the divided data in the above-mentioned divided data group corresponds to the graphics processor in the above-mentioned graphics processor sequence; an execution unit configured to perform the following processing steps for each graphics processor in the above-mentioned graphics processor sequence: determine the iterative step corresponding to the above-mentioned graphics processor; send the divided data corresponding to the above-mentioned iterative step to the next graphics processor; determine the previous step data based on the above-mentioned iterative step, and send the above-mentioned step data to the corresponding graphics processor; generate block output according to the corresponding divided data; and perform update synchronization operation on the above-mentioned block output to generate updated output data.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.
[0012] The aforementioned embodiments of the present disclosure have the following beneficial effects: The data processing methods based on a large language model in some embodiments of the present disclosure reduce communication latency and avoid an imbalance between computation and communication. Specifically, the high communication latency and imbalance between computation and communication are caused by the fact that the communication transmission of the sequential parallel algorithm primarily relies on point-to-point block transmission, which introduces high communication latency. As the number of GPUs involved in the calculation increases, the computation time per step decreases rapidly at a quadratic rate, while the communication volume decreases only linearly, leading to an imbalance between computation and communication. Based on this, the data processing methods based on a large language model in some embodiments of the present disclosure first receive long sequence data input by a target user. This allows the long sequence data to be input to the model. Second, based on the GPU sequence, the long sequence data is partitioned to generate partitioned data groups. This allows the long sequence data to be partitioned into multiple data sets. Then, for each GPU in the GPU sequence, the following processing steps are performed: First, the iteration step corresponding to the GPU is determined. This allows the iteration step currently in the GPU to be removed. Second, the divided data corresponding to the above iterative step is sent to the next graphics processor. Thus, the divided data can be processed by the next graphics processor. Third, based on the above iterative step, the data of the previous step is determined, and the data of the previous step is sent to the corresponding graphics processor. Thus, it can be transmitted to the corresponding graphics processor by reverse transmission. Fourth, based on the corresponding divided data, a block output is generated. Thus, the block output can be determined by attention calculation. Fifth, the above block output is updated and synchronized to generate updated output data. Thus, by updating and synchronizing the output values, the consistency and accuracy of the data are ensured, and because of the serial parallelism of the bidirectional bandwidth and the multi-node data processing solution, the communication delay of the data is reduced, thereby avoiding the imbalance between calculation and communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of a data processing method based on a large language model according to the present disclosure;
[0015] Figure 2 is a schematic structural diagram of some embodiments of a data processing device based on a large language model according to the present disclosure;
[0016] Figure 3 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0023] Figure 1 The flow chart 100 of some embodiments of the data processing method based on a large language model according to the present disclosure is shown. The data processing method based on a large language model includes the following steps:
[0024] Step 101: Receive long sequence data input by a target user.
[0025] In some embodiments, an executing entity (e.g., a server) of a data processing method based on a large language model may receive long sequence data input by a target user. The large language model is deployed in a data processing system containing a graphics processor array. The graphics processor array in the data processing system includes at least one graphics processor (GPU). The target user may be a user using the large language model. The long sequence data may be data having a length greater than a preset data length.
[0026] In some optional implementations of some embodiments, the above-mentioned data processing system can be applied to a multi-node distributed system, wherein the above-mentioned multi-node distributed system may include two computing nodes, each computing node includes four graphics processors, and each graphics processor included in the above-mentioned two computing nodes is connected through a high-speed bus, and the above-mentioned two computing nodes are connected through a remote direct memory access network.
[0027] Step 102 : Based on the graphics processor sequence, the long sequence data is divided and processed to generate divided data groups.
[0028] In some embodiments, the execution entity may partition the long sequence data based on the GPU sequence to generate partitioned data groups. The partitioned data in the partitioned data groups correspond to GPUs in the GPU sequence. In practice, the long sequence data tensors Q (Query), K (Key), and V (Value) may be evenly partitioned into multiple parts according to preset partition lengths to generate the partitioned data groups.
[0029] In practice, the long sequence data can be divided and processed by the following steps:
[0030] The first step is to determine the number of graphics processors included in the graphics processor sequence to generate a graphics processor number value.
[0031] The second step is to divide the long sequence data into a preset number of divided data to obtain divided data groups, wherein the preset number is the number of the graphics processors.
[0032] While employing technical solutions to address the aforementioned issues, the following technical problem often arises: Inter-GPU communication typically requires global broadcasts for data transmission, which results in data being sent to GPUs that don't need to perform computations, resulting in a waste of transmission resources. Considering these technical issues and considering the current state of technology, the following solution has been adopted.
[0033] In some optional implementations of some embodiments, the execution entity may perform segmentation processing on the long sequence data by the following steps:
[0034] The first step is to determine the number of image processors included in the graphics processor sequence as the number of processors.
[0035] The second step is to determine the sequence length of the long sequence data and the model parameters of the large language model. The model parameters include the number of attention heads. The sequence length can be the character length of the long sequence data.
[0036] The third step is to determine the head dimension corresponding to the large language model based on the number of attention heads and the number of processors. In practice, the head dimension can be determined as the quotient of the number of attention heads and the number of processors.
[0037] In step 4, in response to the number of attention heads being greater than or equal to the number of processors, the long sequence data is partitioned based on the head dimensions to generate a partitioned data set. In practice, the long sequence data can be partitioned into a number of partitioned data equal to the number of head dimensions.
[0038] In step 5, in response to the number of attention heads being less than the number of processors, determining a sequence dimension corresponding to the number of processors, wherein the sequence dimension may be the number of data obtained by dividing the long sequence data.
[0039] The sixth step is to divide the long sequence data based on the sequence dimension to generate divided data groups.
[0040] The above-mentioned steps 1-6 and their related contents, as an inventive feature of an embodiment of the present disclosure, address the technical problem that "when communicating across GPUs, data transmission is typically required via global broadcast, resulting in data being transmitted to other GPUs that do not need to perform computation, thereby wasting transmission resources." Factors that often lead to significant transmission resource waste are as follows: When communicating across GPUs, data transmission is typically required via global broadcast, resulting in data being transmitted to other GPUs that do not need to perform computation, thereby wasting transmission resources. If these factors are resolved, the waste of transmission resources can be avoided. To achieve this, first, the number of image processors included in the GPU sequence is determined as the number of processors. This determines the number of processors involved in the computation. Second, the sequence length of the long sequence data is determined, as well as the model parameters of the large language model. This determines the length of the input data. Third, based on the number of attention heads and the number of processors, the head dimension corresponding to the large language model is determined. This determines the head dimension of the multi-head attention mechanism. Fourth, in response to the number of attention heads being greater than or equal to the number of processors, the long sequence data is divided and processed based on the head dimension to generate a divided data group. Thus, data division can be performed by the head dimension. Fifth, in response to the number of attention heads being less than the number of processors, the sequence dimension corresponding to the number of processors is determined; based on the sequence dimension, the long sequence data is divided and processed to generate a divided data group. Thus, data division can be performed by the sequence dimension. Also, because the relationship between the number of attention heads and the number of processors is determined through the dynamic segmentation mechanism, when the graphics processor communicates across CPUs, it only needs to interact with specific neighbors, avoiding the waste of transmission resources, and calculating local data and sending or receiving data in parallel, thereby reducing the time for processing data.
[0041] While employing technical solutions to address the aforementioned issues, the following technical problem often arises: when transferring data between different graphics processors, the entire historical K and V values for each data block must be transmitted, while some K and V values to be transferred do not require calculation, resulting in a significant waste of transmission resources. Considering these technical issues and considering the current state of technology, the following solution has been adopted.
[0042] In practice, the long sequence data can be divided based on the sequence dimension to generate divided data groups through the following sub-steps:
[0043] The first sub-step is to generate a partition quantity value based on the above sequence dimension. In practice, the partition quantity value can be determined as the product of the above sequence dimension and 2.
[0044] The second sub-step is to divide the long sequence data according to the above-mentioned division quantity value to generate initial division data and obtain an initial division data group.
[0045] The third sub-step is to perform the following generation steps for each initial division data in the initial division data group:
[0046] The first generation step is to generate a local causal mask corresponding to the initial divided data, wherein the local causal mask is used to mask future information of the data.
[0047] The second generation step is to perform masking on the initial divided data based on the local causal mask to generate masked data.
[0048] The fourth sub-step is to combine the generated masked data to generate a combined masked data group as a divided data group.
[0049] Optionally, after the fourth sub-step, the following steps are further included:
[0050] In the fifth sub-step, processor tagging is performed on each of the partitioned data in the partitioned data group to generate tagged data, thereby obtaining a tagged data group. Here, the first and last partitioned data in the partitioned data group may be tagged with the same processor, and the second and second-to-last partitioned data may be tagged with the same processor, to achieve symmetrical data distribution.
[0051] The sixth sub-step is to perform the following allocation steps for each graphics processor in the above graphics processor sequence:
[0052] The first allocation step is to control the graphics processor to determine the attention score corresponding to the first marked data.
[0053] In the second allocation step, the Q value of the last marked data corresponding to the GPU is sent to the target GPU. The target GPU can be a pre-determined GPU. Here, the last marked data can be sent to GPU (j+N / 2) mod N, where j represents the marked data symmetrical to the current marked data and N represents the number of GPUs. For example, GPU0 sends the Q value of marked data 7 to GPU2, and CPU2 is determined by (0+2) mod 4 = 2.
[0054] In the third allocation step, the graphics processor is controlled to receive the K and V values sent to the graphics processor. Here, the K and V values sent from GPU (jN / 2) mod N can be received. As an example, GPU0 can receive the K and V values of the marked data 2 from GPU2, where the K and V values of GPU2 are (0-2) mod 4 = 2.
[0055] The fourth allocation step is to determine the attention score of the last marked data based on the Q value of the last marked data and the received K value and V value.
[0056] The seventh sub-step is to aggregate the determined attention scores to generate an attention score sequence.
[0057] The first through seventh substeps and their related content, as an inventive feature of an embodiment of the present disclosure, address the technical problem that "when data is transmitted between different graphics processors, the entire historical K and V values for each data block must be transmitted, while some of the K and V values to be transmitted do not require calculation, resulting in a significant waste of transmission resources." This significant waste of transmission resources is often caused by the following factors: when data is transmitted between different graphics processors, the entire historical K and V values for each data block must be transmitted, while some of the K and V values to be transmitted do not require calculation, resulting in a significant waste of transmission resources. Resolving this issue can avoid this waste of transmission resources. To achieve this, first, a partition count value is generated based on the sequence dimension. This determines the number of partitioned data. Second, the long sequence data is partitioned based on the partition count value to generate initial partition data, resulting in an initial partition data group. This completes the partitioning of the long sequence data. Third, for each initial partition data in the initial partition data group, the following generation step is performed: a local causal mask corresponding to the initial partition data is generated. Based on the local causal mask, the initial partitioned data is masked to generate masked data. Thus, the causal mask can preserve the effective attention region of the data. Fourth, the generated masked data are combined to generate a combined masked data group as a partitioned data group; each partitioned data in the partitioned data group is processor-labeled to generate labeled data, thereby obtaining a labeled data group. Thus, different partitioned data can be assigned to different graphics processors through a symmetrical allocation. Fifth, for each graphics processor in the graphics processor sequence, the following allocation step is performed: controlling the graphics processor to determine the attention score corresponding to the first labeled data. Thus, the attention score of the first half of the data block in the graphics processor can be determined. Sixth, the Q value of the last labeled data corresponding to the graphics processor is sent to the target graphics processor; controlling the graphics processor to receive the K value and V value sent to the graphics processor; and determining the attention score of the last labeled data based on the Q value of the last labeled data and the received K value and V value. Thus, only the Q, K, and V values can be transmitted to the graphics processor performing the calculation, thereby avoiding the transmission of all Q, K, and V values, thereby avoiding the waste of transmission resources. Seventh, the determined attention scores are aggregated to generate an attention score sequence. In this way, the calculation of the attention score is completed, avoiding the waste of transmission resources.
[0058] Step 103: For each graphics processor in the graphics processor sequence, perform the following processing steps:
[0059] The first processing step is to determine the iteration step corresponding to the graphics processor.
[0060] In some embodiments, the execution subject may determine the iteration step corresponding to the graphics processor. In practice, the iteration step currently in the graphics processor may be determined by querying the status of the graphics processor.
[0061] The second processing step is to send the divided data corresponding to the above iterative step to the next graphics processor.
[0062] In some embodiments, the execution entity may send the divided data corresponding to the iterative step to the next graphics processor.
[0063] In practice, the divided data corresponding to the above iterative steps can be sent to the next graphics processor through the following steps:
[0064] The first step is to determine the Q, K, and V values corresponding to the divided data. In practice, the Q, K, and V values corresponding to the divided data can be determined by matrix multiplication.
[0065] The second step is to copy the above Q value to generate a Q value copy.
[0066] The third step is to send the Q value copy to the next graphics processor, and store the Q value, K value and V value in the associated database.
[0067] In the fourth step, in response to receiving the K value copy and the V value copy sent by other graphics processors, the K value copy and the V value copy are stored in the associated cache.
[0068] In step 5, in response to detecting that the calculation is completed, initialization processing is performed on the above-mentioned associated cache to clear the cache content in the above-mentioned cache.
[0069] Optionally, after the second processing step, transmission processing is performed on the updated output data based on a two-way communication mechanism to generate transmitted output data.
[0070] In some embodiments, the execution subject may perform transmission processing on the updated output data based on a bidirectional communication mechanism in response to a data processing system applied to a diffusion transformer to generate transmitted output data.
[0071] The third processing step is to determine the data of the previous step based on the above iterative step, and send the data of the previous step to the corresponding graphics processor.
[0072] In some embodiments, the execution entity may determine the data of the previous step based on the iterative step, and send the data of the previous step to the corresponding graphics processor, wherein the data of the previous step may be the data outputted by the previous iterative step.
[0073] The fourth processing step is to generate block output according to the corresponding divided data.
[0074] In some embodiments, the execution entity may generate block outputs based on the corresponding divided data.
[0075] The fifth processing step is to perform an update synchronization operation on the block output to generate updated output data.
[0076] In some embodiments, the execution entity may perform an update synchronization operation on the block output to generate updated output data.
[0077] Optionally, after step 103, the updated output data is stored in a preset cache.
[0078] In some embodiments, the execution entity may store the updated output data in a pre-set cache.
[0079] The aforementioned embodiments of the present disclosure have the following beneficial effects: The data processing methods based on a large language model in some embodiments of the present disclosure reduce communication latency and avoid an imbalance between computation and communication. Specifically, the high communication latency and imbalance between computation and communication are caused by the fact that the communication transmission of the sequential parallel algorithm primarily relies on point-to-point block transmission, which introduces high communication latency. As the number of GPUs involved in the calculation increases, the computation time per step decreases rapidly at a quadratic rate, while the communication volume decreases only linearly, leading to an imbalance between computation and communication. Based on this, the data processing methods based on a large language model in some embodiments of the present disclosure first receive long sequence data input by a target user. This allows the long sequence data to be input to the model. Second, based on the GPU sequence, the long sequence data is partitioned to generate partitioned data groups. This allows the long sequence data to be partitioned into multiple data sets. Then, for each GPU in the GPU sequence, the following processing steps are performed: First, the iteration step corresponding to the GPU is determined. This allows the iteration step currently in the GPU to be removed. Second, the divided data corresponding to the above iterative step is sent to the next graphics processor. Thus, the divided data can be processed by the next graphics processor. Third, based on the above iterative step, the data of the previous step is determined, and the data of the previous step is sent to the corresponding graphics processor. Thus, it can be transmitted to the corresponding graphics processor by reverse transmission. Fourth, based on the corresponding divided data, a block output is generated. Thus, the block output can be determined by attention calculation. Fifth, the above block output is updated and synchronized to generate updated output data. Thus, by updating and synchronizing the output values, the consistency and accuracy of the data are ensured, and because of the serial parallelism of the bidirectional bandwidth and the multi-node data processing solution, the communication delay of the data is reduced, thereby avoiding the imbalance between calculation and communication.
[0080] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a data processing device based on a large language model. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the data processing device based on the large language model can be specifically applied to various electronic devices.
[0081] like Figure 2 As shown, the data processing device 200 based on the large language model of some embodiments includes: a receiving unit 201, a dividing unit 202 and an executing unit 203. The receiving unit 201 is configured to receive the long sequence data input by the target user;
[0082] The division unit 202 is configured to perform division processing on the long sequence data based on the graphics processor sequence to generate a divided data group, wherein the divided data in the divided data group corresponds to the graphics processors in the graphics processor sequence;
[0083] The execution unit 203 is configured to perform the following processing steps for each graphics processor in the above-mentioned graphics processor sequence: determine the iteration step corresponding to the above-mentioned graphics processor; send the divided data corresponding to the above-mentioned iteration step to the next graphics processor; based on the above-mentioned iteration step, determine the previous step data, and send the above-mentioned previous step data to the corresponding graphics processor; generate block output according to the corresponding divided data; and perform update synchronization operation on the above-mentioned block output to generate updated output data.
[0084] It is understandable that the units and references in the data processing device 200 based on the large language model are Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the data processing device 200 based on the large language model and the units contained therein, and will not be repeated here.
[0085] Reference below Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0086] like Figure 3 As shown, the electronic device 300 may include a processing device 301 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0087] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0088] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0089] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0090] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0091] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: receives long sequence data input by a target user; divides the long sequence data based on the graphics processor sequence to generate a divided data group, wherein the divided data in the divided data group corresponds to a graphics processor in the graphics processor sequence; performs the following processing steps for each graphics processor in the graphics processor sequence: determines the iteration step corresponding to the graphics processor; sends the divided data corresponding to the iteration step to the next graphics processor; determines the previous step data based on the iteration step, and sends the previous step data to the corresponding graphics processor; generates block output based on the corresponding divided data; and performs update synchronization operations on the block output to generate updated output data.
[0092] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0094] The units described in some embodiments of the present disclosure may be implemented in software or hardware. The units described may also be provided in a processor. For example, they may be described as: a processor including a receiving unit, a partitioning unit, and an execution unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the receiving unit may also be described as a "unit for receiving long sequence data input by a target user."
[0095] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0096] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A data processing method based on a large language model, wherein: The large language model is deployed in a data processing system including a graphics processor sequence, wherein the graphics processor sequence in the data processing system includes at least one graphics processor, and the method includes: Receive long sequence data input by the target user; Based on the graphics processor sequence, the long sequence data is divided and processed to generate divided data groups, wherein the divided data in the divided data groups correspond to graphics processors in the graphics processor sequence; For each graphics processor in the graphics processor sequence, the following processing steps are performed: Determining an iteration step corresponding to the graphics processor; Sending the divided data corresponding to the iterative step to the next graphics processor; Based on the iterative steps, determining data of a previous step, and sending the data of the previous step to a corresponding graphics processor; Generate block output according to the corresponding divided data; An update synchronization operation is performed on the block output to generate updated output data.
2. The method according to claim 1, wherein The method further comprises: The updated output data is stored in a preset cache.
3. The method according to claim 1, wherein The data processing system is applied to a multi-node distributed system, wherein the multi-node distributed system includes two computing nodes, each computing node includes four graphics processors, and the graphics processors included in each computing node of the two computing nodes are connected through a high-speed bus, and the two computing nodes are connected through a remote direct memory access network.
4. The method according to claim 1, wherein The step of dividing the long sequence data based on the graphics processor sequence to generate divided data groups includes: Determining the number of graphics processors included in the graphics processor sequence to generate a graphics processor number value; The long sequence data is divided into a preset number of divided data to obtain a divided data group, wherein the preset number is the number of the graphics processors.
5. The method according to claim 1, wherein The data processing system is applied to a diffusion converter; as well as The method further comprises: Based on the bidirectional communication mechanism, transmission processing is performed on the updated output data to generate transmitted output data.
6. The method according to claim 1, wherein The sending the divided data corresponding to the iterative step to the next graphics processor includes: Determine the Q value, K value and V value corresponding to the divided data; Performing a copy operation on the Q value to generate a Q value copy; Sending the Q value copy to the next graphics processor, and storing the Q value, K value and V value in an associated database; In response to receiving the K value copy and the V value copy sent by other graphics processors, storing the K value copy and the V value copy in the associated cache; In response to detecting that the calculation is finished, initialization processing is performed on the associated cache to clear cache content in the cache.
7. A data processing device based on a large language model, comprising: A receiving unit configured to receive long sequence data input by a target user; a partitioning unit configured to partition the long sequence data based on the graphics processor sequence to generate partitioned data groups, wherein the partitioned data in the partitioned data groups correspond to graphics processors in the graphics processor sequence; The execution unit is configured to perform the following processing steps for each graphics processor in the graphics processor sequence: determining an iteration step corresponding to the graphics processor; sending the divided data corresponding to the iteration step to the next graphics processor; determining the previous step data based on the iteration step, and sending the previous step data to the corresponding graphics processor; generating a block output based on the corresponding divided data; and performing an update synchronization operation on the block output to generate updated output data.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.