Information processing method and device, equipment and storage medium
By deploying the attention and feedforward processing units of the hybrid expert model on different computing nodes and using heterogeneous hardware and pipeline parallel mechanisms, the problems of resource waste and inefficiency in traditional deployment methods are solved, and more efficient computing performance and resource utilization are achieved.
Patent Information
- Application Number
- CN202510412284.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
When dealing with hybrid expert models, traditional model deployment methods face problems such as low GPU utilization, large communication overhead and high hardware costs, and it is difficult to fully realize the hardware potential, resulting in waste of resources and inferred inference efficiency.
The attention processing unit and feedforward processing unit of the model are deployed in different computing node groups, data transmission is carried out through the M2N and N2M communication mechanisms, and pipeline parallel mechanism and heterogeneous hardware configuration are used to optimize resource utilization and processing efficiency.
Improves the overall processing efficiency and resource utilization of the hybrid expert model, reduces communication latency and resource waste, and achieves higher computing performance and throughput.
Smart Images

Figure CN120258062A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for information processing. Background Art
[0002] With the development of computer technology, generative models have been gradually applied to the processing of various tasks, such as search engines, chatbots, and programming assistance tools. In recent years, the scale and complexity of models have been continuously increasing, and the demand for computing resources has also been growing. However, traditional model deployment methods face many challenges when dealing with generative models. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for information processing is provided. The method includes: receiving an input request for a model, where the attention processing unit of the model is deployed on a first set of computing nodes, and the feed-forward processing unit of the model is deployed on a second set of computing nodes different from the first set of computing nodes; using the attention processing unit deployed on the first set of computing nodes to process first data to determine first attention information, where the first data is determined based on the input request; providing the first attention information to the second set of computing nodes via communication between the first set of computing nodes and the second set of computing nodes; processing the first attention information by the feed-forward processing unit deployed on the second set of computing nodes to generate second data; and generating response content for the input request based on the second data.
[0004] In a second aspect of the present disclosure, an apparatus for information processing is provided. The apparatus includes: a receiving module configured to receive an input request for a model, where the attention processing unit of the model is deployed on a first set of computing nodes, and the feed-forward processing unit of the model is deployed on a second set of computing nodes different from the first set of computing nodes; an attention processing module configured to use the attention processing unit deployed on the first set of computing nodes to process first data to determine first attention information, where the first data is determined based on the input request; a communication module configured to provide the first attention information to the second set of computing nodes via communication between the first set of computing nodes and the second set of computing nodes; a feed-forward processing module configured to process the first attention information by the feed-forward processing unit deployed on the second set of computing nodes to generate second data; and a generating module configured to generate response content for the input request based on the second data.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 A schematic diagram of a model deployment system according to an embodiment of the present disclosure is shown;
[0010] Figure 2 A schematic diagram of a pipeline parallel mechanism according to some embodiments of the present disclosure is shown;
[0011] Figure 3A A schematic block diagram of a sending component according to some embodiments of the present disclosure is shown;
[0012] Figure 3B A schematic block diagram of a receiving component according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A flowchart of an example process of information processing according to some embodiments of the present disclosure is shown;
[0014] Figure 5 A schematic structural block diagram of an example device for information processing according to some embodiments of the present disclosure is shown; and
[0015] Figure 6 A block diagram of an electronic device capable of implementing multiple embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0017] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.
[0018] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may be other explicit and implicit definitions hereinafter. Terms such as "first", "second", etc. may refer to different or the same objects. There may be other explicit and implicit definitions hereinafter.
[0019] The embodiments of the present disclosure may involve the user's data, data acquisition and / or use, etc. All of these aspects comply with the corresponding laws, regulations and related provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware and confirms. Accordingly, when implementing the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained through appropriate means in accordance with the relevant laws and regulations. The specific informing and / or authorization methods may vary according to the actual situation and application scenarios, and the scope of the present disclosure is not limited in this regard.
[0020] For the solutions in this specification and embodiments, if they involve personal information processing, they will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for basic functions, it will not affect the user's use of basic functions.
[0021] With the rapid development of generative models, the scale and complexity of the models are constantly increasing, and the demand for computing resources is also growing day by day. However, traditional model deployment methods face many challenges when dealing with Mixture of Experts (MoE) models, such as low GPU utilization, high communication overhead, and high hardware costs.
[0022] The Mixture of Experts model, through a sparse activation mechanism, dynamically distributes input data to different expert networks for processing, thereby improving model performance without significantly increasing computational complexity. The expert network is also called FFN (Feed-Forward Network). However, this sparse activation characteristic leads to significant differences in the computational requirements and resource utilization patterns of the attention module and the FFN module during the inference process. Specifically, the attention module is usually limited by memory access during the decoding phase, while the FFN module is limited by computational power. This difference makes it difficult for traditional single-node or simple parallelization deployment methods to fully utilize the potential of the hardware, resulting in resource waste and low inference efficiency.
[0023] Embodiments of the present disclosure propose a solution for information processing. The solution includes: receiving an input request for a model, where the attention processing unit of the model is deployed on a first set of computing nodes, and the feed-forward processing unit of the model is deployed on a second set of computing nodes different from the first set of computing nodes; using the attention processing unit deployed on the first set of computing nodes to process first data to determine first attention information, where the first data is determined based on the input request; providing the first attention information to the second set of computing nodes via communication between the first set of computing nodes and the second set of computing nodes; processing the first attention information by the feed-forward processing unit deployed on the second set of computing nodes to generate second data; and generating response content for the input request based on the second data.
[0024] Thus, by separating the attention processing unit and the feed-forward processing unit of the model into different computing node groups, the embodiments of the present disclosure achieve independent deployment and processing of modules in the Mixture of Experts model. This separated architecture enables the system to optimize the deployment according to the characteristics of different modules, such as selecting appropriate node groups based on computational requirements and hardware characteristics, thereby improving the overall processing efficiency and resource utilization rate.
[0025] The following further describes various example implementations of this solution in detail with reference to the accompanying drawings.
[0026] Example Deployment System
[0027] Figure 1 A schematic diagram of a model deployment system 100 according to an embodiment of the present disclosure is shown. As Figure 1As shown, the attention processing unit of the model is deployed at the first set of computing nodes 110, and the feed-forward processing unit (also known as the expert unit) of the model is deployed at the second set of computing nodes 120.
[0028] The first set of computing nodes 110 can also be referred to as attention nodes, which are used to store the parameters of the attention processing unit (also known as the attention module) and the key-value cache (KV Cache), and are configured to process the input data to generate attention information. In some deployments, the first set of computing nodes 110 can include multiple graphics processing units (GPUs), and these GPUs can work together through tensor parallelism.
[0029] The second set of computing nodes 120 can also be referred to as expert nodes, which store the parameters of a single expert (i.e., the corresponding feed-forward network FFN) and are responsible for processing the input data assigned to that expert. In some deployments, the second set of computing nodes 120 can include 1 to 8 GPUs, and these GPUs can work together through tensor parallelism.
[0030] Furthermore, data transmission can be achieved between the first set of computing nodes 110 and the second set of computing nodes 120 through the M2N communication mechanism and the N2M communication mechanism. Specifically, the attention nodes can send the generated attention information to the expert nodes through the M2N communication mechanism, and the expert nodes can return the processing results through the N2M communication mechanism. The M2N communication mechanism will be described in detail below with reference to Figure 3A and Figure 3B will not be elaborated here for the time being.
[0031] Based on the deployment system 100 as Figure 1 shown, after receiving an input request, the model can use the attention processing unit deployed at the first set of computing nodes 110 to process the first data determined based on the input request, thereby determining the corresponding first attention information. As an example, the first data can include a batch of data to be processed, or a part of the data in the micro-batch.
[0032] Furthermore, the model can use the communication between the first set of computing nodes 110 and the second set of computing nodes 120 to provide the calculated first attention information to the second set of computing nodes 120.
[0033] Correspondingly, the feed-forward processing unit (also known as the expert module) deployed at the second set of computing nodes 120 can process the received first attention information to generate the second data.
[0034] Further, the model can generate response content for the input request based on the obtained second data. For example, the second set of computing nodes 120 can provide the generated second data to the first set of computing nodes 110 for subsequent attention calculation.
[0035] In some embodiments, the model can include deploying any suitable generative model of a mixture of experts network, such as, for example, a language model, a diffusion model, etc. In some embodiments, the input request of the model can indicate appropriate generation control parameters, such as, for example, a prompt. Accordingly, depending on the type of the model, the response content generated by the model can include text content and / or media content, and the media content can include, for example, but not limited to: picture content, video content, music content, 3D models, etc.
[0036] In some embodiments, the first set of computing nodes 110 and the second set of computing nodes 120 can also adopt a pipeline parallel mechanism as Figure 2 shown. Figure 2 FIG. 200 shows a schematic diagram of a pipeline parallel mechanism according to some embodiments of the present disclosure.
[0037] Since the feed-forward processing unit and the attention processing unit are deployed on different computing nodes, using a single batch of requests will cause the feed-forward processing unit and the attention processing unit to not be able to work simultaneously. In addition, during node communication, the computing nodes will also be in an idle state.
[0038] In some embodiments, the deployment system 100 can divide a batch of data into m micro-batches, thereby creating a pipeline as Figure 2 shown between the attention nodes and the expert nodes.
[0039] Specifically, taking m = 4 as an example, the first set of computing nodes 110 can sequentially process data of 4 micro-batches. Further, after the data of the first micro-batch (for example, 11, also referred to as the first data) is processed, the first set of computing nodes 110 can send the attention information corresponding to the data of the first micro-batch (also referred to as the first attention information) to the second set of computing nodes 120.
[0040] As Figure 2 shown, the process of the first set of computing nodes 110 generating the attention information (also referred to as the second attention information) corresponding to the data of the next micro-batch (for example, 12, also referred to as the third data) can be at least partially parallel to the process of providing the first attention information to the second set of computing nodes 120.
[0041] Further, the latency of the two communications can also be reduced by Figure 2Hidden by the pipeline mechanism shown. For example, by reasonably designing the size and number of micro-batches, the first set of computing nodes 110 can receive the processing results returned by the second set of computing nodes 120 before completing the processing of all the data in the micro-batch, so that the next batch of data can be processed.
[0042] Specifically, the receiving time of the first set of computing nodes for the processing results (also referred to as the second data) of the data of the first micro-batch (e.g., 11) generated by the second set of computing nodes 120 is earlier than the completion time of the first set of computing nodes for generating the attention information corresponding to the data of the last micro-batch (e.g., 14, also referred to as the fourth data).
[0043] In this way, these nodes perform the forward pass of the micro-batch twice in each mixture-of-experts model layer and exchange intermediate results. This setting enables the forward computation to cover the communication overhead, thus achieving higher resource utilization.
[0044] In some embodiments, the deployment plan of the model is determined based on the hardware information and / or the model information of the model, where the hardware information is associated with the first set of computing nodes and / or the second set of computing nodes. Additionally, the deployment plan can at least indicate the number of the multiple micro-batches described above.
[0045] In some embodiments, the deployment system 100 can search to determine the deployment plan of the model to maximize the throughput per unit cost while meeting the latency constraint by reasonably configuring the parallel strategy of the model and the hardware resources.
[0046] In some implementations, the deployment system 100 can first initialize a variable to store the currently found optimal deployment plan. Then, the deployment system 100 starts to traverse all possible combinations of tensor parallel sizes, corresponding to the attention nodes and the expert nodes respectively. For each pair of tensor parallel sizes, the deployment system 100 can determine whether they meet the limit conditions of the hardware memory capacity. If they meet, the deployment system 100 will further calculate the number of attention nodes to balance the computing time of the attention module and the expert module as much as possible and reduce the idle time between the modules.
[0047] Furthermore, the deployment system 100 will enumerate different numbers of micro-batches. For each number of micro-batches, the deployment system 100 will generate a deployment configuration and call a simulation function to evaluate the performance of this configuration. The simulation function will calculate the maximum global batch size that the system can achieve under the given configuration and ensure that this configuration meets the latency constraint. If the throughput per unit cost of the current configuration is better than the recorded optimal configuration, then the deployment system 100 will update the optimal configuration.
[0048] In this way, the deployment system 100 continuously traverses all possible configurations and returns the optimal deployment plan. Specifically, this plan can achieve the highest throughput per unit cost while meeting the latency requirements. In this way, embodiments of the present disclosure can obtain an efficient deployment solution.
[0049] In some embodiments, the deployment system 100 can also support heterogeneous deployment, that is, using heterogeneous hardware settings for the attention nodes and expert nodes. Specifically, the first set of computing nodes 110 and the second set of computing nodes 120 can correspond to different types of computing devices.
[0050] As an example, the attention nodes are memory-intensive, performing memory access most of the time and requiring a large amount of storage space to store the KV cache. Thus, the deployment system 110 can use GPUs with higher per-unit-cost memory bandwidth and larger per-unit-cost memory capacity for the attention nodes. In contrast, for the compute-intensive expert nodes, the deployment system 100 can use GPUs with higher per-unit-cost computing power.
[0051] Example Communication Framework
[0052] The following will refer to Figure 3A and Figure 3B to describe an example communication framework according to embodiments of the present disclosure. Figure 3A FIG. shows a schematic block diagram of a sending component 300A according to some embodiments of the present disclosure, Figure 3B FIG. shows a schematic block diagram of a receiving component 300B according to some embodiments of the present disclosure.
[0053] As Figure 3A shown, the sending component 300A can be associated with a graphics processing unit (GPU) 310 and a central processing unit (CPU) 320.
[0054] Specifically, the CPU 320 can start the sending control kernel in the GPU 310. When the compute kernel in the GPU 310 finishes processing, the sending control kernel can update the sending flag (also referred to as the first identifier). For example, when the GPU 310 generates attention information, the sending flag can be updated from a first value to a second value.
[0055] Taking the sending azimuth attention node as an example, after the attention node finishes generating the first attention information, the sending control kernel can update the sending flag from a first value to a second value.
[0056] Further, the M2N transmitter in the CPU 320 can poll whether the value of the transmission flag is updated. After detecting that the transmission flag is updated to the second value, the CPU 320 can use the transmission queue to send a preset portion of the first attention information in the tensor memory (also referred to as the first memory) in the GPU 310 to the recipient corresponding to the transmission queue.
[0057] Specifically, as Figure 3A shown, after polling that the transmission flag is updated, the M2N transmitter can call M2NSEND to trigger the kernel transmitter to perform data transmission. Specifically, the kernel transmitter can call the SENDTON(N) function to send the data in the tensor memory (also referred to as the first memory) to N QPs (Queue Par, queue pair, that is, including a transmission queue and a reception queue). Each QP can correspond to a corresponding recipient. For example, QP 1 corresponds to recipient 1, and QP 2 corresponds to recipient 2.
[0058] Further, after completing the transmission process of the QP, the corresponding completion queue (CQ, Complete Queue) can be updated accordingly. Thus, the kernel transmitter can use POLL_CQ to determine that all N queues have completed transmission. Correspondingly, the M2N transmitter can poll through POLLCQ that all N queues have completed transmission, thereby completing the data transmission process of the transmission component 300A.
[0059] The data reception process of the reception component 300B will be further described below with reference to Figure 3B to describe the data reception process of the reception component 300B.
[0060] Specifically, the reception component 300B can also be associated with the GPU 330 and the CPU 340. As Figure 3B shown, after the reception component 300B receives data from N senders using N QPs, such data can be written into a temporary reception buffer. In addition, the kernel receiver can use POLL_CQ to determine that the data corresponding to the N QPs has been received.
[0061] Correspondingly, the M2N receiver can poll through POLLCQ that the data corresponding to the N QPs has been received and can update the reception flag (also referred to as the second identifier). Specifically, the reception control core in the GPU 330 can poll whether the value of the reception flag is updated. After detecting that the value of the reception flag is updated from the third value to the fourth value, the GPU 330 can copy the received data (for example, a preset portion of the first attention data) from the temporary reception buffer to the tensor memory (also referred to as the second memory). After completing the data copy, the computing core can use the data stored in the tensor memory to perform a corresponding computing process.
[0062] Further, when all the received data (e.g., a preset part of the first attention data) is copied to the tensor storage, the receiving control core can update the value of the copy flag (also referred to as the third identifier) to indicate that the temporary receiving buffer is allowed to be reused. For example, the CPU 340 can support writing the subsequently received data into the temporary receiving buffer, or the CPU can release the temporary receiving buffer, etc.
[0063] Based on the communication framework described above, the embodiments of the present disclosure can eliminate unnecessary GPU-to-CPU data copy operations, avoiding the latency and performance loss caused by frequent data transmission between different processing units in the traditional communication method. In addition, this communication framework also removes the complex group initialization and processing overhead, thus significantly improving the communication efficiency.
[0064] Example Process
[0065] Figure 4 A flowchart of an example process 400 of information processing according to some embodiments of the present disclosure is shown. The process 400 can be implemented at the deployment system 100.
[0066] As Figure 4 shown, at block 410, the deployment system 100 receives an input request for a model, where the attention processing unit of the model is deployed in a first group of computing nodes, and the feed-forward processing unit of the model is deployed in a second group of computing nodes different from the first group of computing nodes.
[0067] At block 420, the deployment system 100 processes the first data using the attention processing unit deployed in the first group of computing nodes to determine the first attention information, where the first data is determined based on the input request.
[0068] At block 430, the deployment system 100 provides the first attention information to the second group of computing nodes via communication between the first group of computing nodes and the second group of computing nodes.
[0069] At block 440, the deployment system 100 processes the first attention information using the feed-forward processing unit deployed in the second group of computing nodes to generate the second data.
[0070] At block 450, the deployment system 100 generates response content for the input request based on the second data.
[0071] In some embodiments, the process 400 further includes: providing the second data from the second group of computing nodes to the first group of computing nodes via communication.
[0072] In this way, embodiments of the present disclosure add a step of the second set of computing nodes feeding back second data to the first set of computing nodes. This two-way communication mechanism enables the system to achieve more complex interactions and collaborative processing. For example, the first set of computing nodes can further adjust the processing strategy or perform subsequent processing based on the second data returned by the second set of nodes, thereby enhancing the flexibility and adaptability of the system and better coping with complex input requests and dynamic workload changes.
[0073] In some embodiments, process 400 further includes: determining data of a target batch based on an input request; and dividing the data of the target batch into multiple micro-batches of data, where the multiple micro-batches of data include first data.
[0074] In this way, embodiments of the present disclosure further refine the data processing process by introducing a step of dividing the target batch data into multiple micro-batches. In this way, the system can more flexibly process large-scale data, and at the same time, the micro-batch processing method helps to better utilize computing resources, such as improving the processing speed by processing multiple micro-batches in parallel. In addition, the division of micro-batches also provides a basis for subsequent pipeline processing, enabling the system to prepare the next micro-batch while processing one micro-batch, thereby further improving efficiency.
[0075] In some embodiments, the process of the first set of computing nodes generating second attention information corresponding to third data is at least partially parallel to the process of providing first attention information to the second set of computing nodes, where the third data is the data of the next micro-batch after the first data among the multiple micro-batches of data.
[0076] In this way, embodiments of the present disclosure describe a parallel processing mechanism of the first set of computing nodes when processing micro-batch data. Specifically, while the first set of computing nodes provides first attention information to the second set of computing nodes, it has already started processing the data of the next micro-batch (third data) and generating corresponding second attention information. This parallel processing method significantly improves the processing efficiency of the system, reduces waiting time, and fully utilizes computing resources, enabling the system to respond to input requests more quickly, especially when dealing with large-scale data, the advantage is more obvious.
[0077] In some embodiments, the reception time of the first set of computing nodes receiving the second data is earlier than the completion time of the first set of computing nodes generating attention information corresponding to fourth data, where the fourth data is the data of the last micro-batch among the multiple micro-batches of data.
[0078] In this way, embodiments of the present disclosure further clarify the time sequence for the first set of computing nodes to receive the second data. Specifically, before the first set of computing nodes finish processing all micro-batches (i.e., before processing the last micro-batch), they have already received the second data returned by the second set of computing nodes. This mechanism can hide the communication overhead between nodes through pipeline parallelism.
[0079] In some embodiments, the deployment plan of the model is determined based on hardware information and / or model information of the model. The hardware information is associated with the first set of computing nodes and / or the second set of computing nodes, and the deployment plan at least indicates the number of multiple micro-batches.
[0080] In this way, embodiments of the present disclosure indicate that the deployment plan of the model is determined based on hardware information and / or model information, and at least indicates the number of micro-batches. This means that the deployment of the system is not fixed, but can be dynamically adjusted according to specific hardware configurations and model requirements. For example, according to factors such as the computing power and memory capacity of the hardware, as well as information such as the size and complexity of the model, the system can flexibly determine the number of micro-batches, thereby achieving optimal resource allocation and processing efficiency. This dynamic deployment method enables the system to better adapt to different hardware environments and model requirements, improving the versatility and scalability of the system.
[0081] In some embodiments, the system can support heterogeneous computing environments and can make full use of the advantages of different types of devices. For example, the first set of computing nodes can use devices suitable for processing attention tasks (such as GPUs with high memory bandwidth), while the second set of computing nodes can use devices suitable for processing feed-forward tasks (such as GPUs with high computing power). In this way, the system can more efficiently utilize various computing resources, further improving processing efficiency and performance.
[0082] In some embodiments, the sending process associated with the first attention information includes: in response to the graphics processing unit generating the first attention information, the graphics processing unit updates the first identifier from the first value to the second value; and in response to detecting that the first identifier is updated to the second value, the central processing unit uses the sending queue to send a preset portion of the first attention information in the first memory of the graphics processing unit to the recipient corresponding to the sending queue.
[0083] In this way, embodiments of the present disclosure describe a sending process associated with first attention information. Through clear division of labor and collaborative operations, efficient heterogeneous computing resource management is achieved. The Graphics Processing Unit (GPU) is responsible for generating the first attention information and updating the identifier, while the Central Processing Unit (CPU) is responsible for sending data to the receiver using a sending queue. This division of labor mechanism gives full play to the advantages of the GPU in parallel computing and data generation, while leveraging the efficiency of the CPU in data transmission and queue management, avoiding the switching overhead of a single processing unit between different tasks, and thus significantly improving the efficiency of data sending and the overall performance of the system.
[0084] In some embodiments, the receiving process associated with the first attention information includes: in response to receiving a preset portion of the first attention information from the sender using a receiving queue, the Central Processing Unit updates the second identifier from a third value to a fourth value; and in response to detecting that the second identifier is updated to the fourth value, the Graphics Processing Unit copies a preset portion of the first attention data from a buffer corresponding to the receiving queue to a second memory of the Graphics Processing Unit.
[0085] In this way, embodiments of the present disclosure detail the receiving process associated with the first attention information. Through clear identifier updates and data copying operations, efficient heterogeneous computing resource collaboration is achieved. The Central Processing Unit (CPU) is responsible for the management of the receiving queue and identifier updates, while the Graphics Processing Unit (GPU) is responsible for copying the received data to its memory. This mechanism not only ensures the timeliness and reliability of data reception, but also reduces delays and errors during data transmission through the collaborative work of the CPU and GPU, improving the overall communication efficiency of the system. At the same time, this division of labor also enables the system to better utilize heterogeneous computing resources and give full play to their respective advantages.
[0086] In some embodiments, process 400 further includes: in response to a preset portion of the first attention data being copied to the second memory, the Graphics Processing Unit updates the value of a third identifier to indicate that the buffer is allowed to be reused.
[0087] In this way, embodiments of the present disclosure further optimize resource management during the receiving process. The Graphics Processing Unit (GPU) updates the identifier after completing data copying, clearly indicating that the buffer can be reused. This mechanism enables the buffer resources to be released and reused in a timely manner, avoiding resource idling and waste, and improving the resource utilization rate of the system. Especially when dealing with large-scale data, this efficient resource management method can significantly reduce memory occupancy, improve the throughput and response speed of the system, and further enhance the overall performance and efficiency of the system.
[0088] Example Apparatus and Device
[0089] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 5 A schematic structural block diagram of an example apparatus 500 for information processing according to some embodiments of the present disclosure is shown. The apparatus 500 may be implemented as or included in a deployment system 100. Each module / component in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0090] As Figure 5 shown, the apparatus 500 includes a receiving module 510 configured to receive an input request for a model, wherein an attention processing unit of the model is deployed on a first set of computing nodes, and a feed-forward processing unit of the model is deployed on a second set of computing nodes different from the first set of computing nodes; an attention processing module 520 configured to process first data by using the attention processing unit deployed on the first set of computing nodes to determine first attention information, the first data being determined based on the input request; a communication module 530 configured to provide the first attention information to the second set of computing nodes via communication between the first set of computing nodes and the second set of computing nodes; a feed-forward processing module 540 configured to process the first attention information by using the feed-forward processing unit deployed on the second set of computing nodes to generate second data; and a generating module 550 configured to generate response content for the input request based on the second data.
[0091] In some embodiments, the apparatus 500 further includes a providing module configured to: provide second data from the second set of computing nodes to the first set of computing nodes via communication.
[0092] In some embodiments, the apparatus 500 further includes a data processing module configured to: determine a target batch of data based on the input request; and divide the target batch of data into multiple micro-batches of data, wherein the multiple micro-batches of data include the first data.
[0093] In some embodiments, the process of the first set of computing nodes generating second attention information corresponding to third data is at least partially parallel to the process of providing the first attention information to the second set of computing nodes, the third data being the next micro-batch of data after the first data in the multiple micro-batches of data.
[0094] In some embodiments, the receiving time of the second data by the first set of computing nodes is earlier than the completion time of the first set of computing nodes generating attention information corresponding to fourth data, the fourth data being the last micro-batch of data in the multiple micro-batches of data.
[0095] In some embodiments, the deployment plan of the model is determined based on hardware information and / or model information of the model, the hardware information is associated with the first set of computing nodes and / or the second set of computing nodes, and the deployment plan at least indicates the number of multiple micro-batches.
[0096] In some embodiments, the first set of computing nodes and the second set of computing nodes correspond to different types of computing devices.
[0097] In some embodiments, the sending process associated with the first attention information includes: in response to the graphics processing unit generating the first attention information, the graphics processing unit updates the first identifier from a first value to a second value; and in response to detecting that the first identifier is updated to the second value, the central processing unit uses the sending queue to send a preset portion of the first attention information in the first memory of the graphics processing unit to the recipient corresponding to the sending queue.
[0098] In some embodiments, the receiving process associated with the first attention information includes: in response to receiving a preset portion of the first attention information from the sender using the receiving queue, the central processing unit updates the second identifier from a third value to a fourth value; and in response to detecting that the second identifier is updated to the fourth value, the graphics processing unit copies a preset portion of the first attention data from the buffer corresponding to the receiving queue to the second memory of the graphics processing unit.
[0099] In some embodiments, the apparatus 500 further includes an update module configured to: in response to a preset portion of the first attention data being copied to the second memory, the graphics processing unit updates the value of the third identifier to indicate that the buffer is allowed to be reused.
[0100] The units included in the apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units in the apparatus 500 can be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0101] Figure 6 A block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 6The illustrated electronic device 600 is merely exemplary and should not impose any limitation on the functionality and scope of the embodiments described herein. Figure 6 The illustrated electronic device 600 can be used to implement Figure 1 the deployment system 100.
[0102] As Figure 6 shown, the electronic device 600 is in the form of a general-purpose electronic device. The components of the electronic device 600 can include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing ability of the electronic device 600.
[0103] The electronic device 600 generally includes multiple computer storage media. Such media can be any accessible media available to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, magnetic disks, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.
[0104] The electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 620 can include a computer program product 625 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.
[0105] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented with a single computing cluster or multiple computing machines that are capable of communicating via a communication link. Accordingly, the electronic device 600 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0106] The input device 650 can be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 660 can be one or more output devices such as a display, speakers, printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) as needed via the communication unit 640, the external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 600, or communicate with any device that enables the electronic device 600 to communicate with one or more other electronic devices (e.g., network cards, modems, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0107] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0108] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0109] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0110] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0111] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart, and combinations of blocks in the block diagrams and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.
[0112] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. A method for information processing, comprising: Receiving an input request for a model, wherein an attention processing unit of the model is deployed on a first set of computing nodes, and a feed-forward processing unit of the model is deployed on a second set of computing nodes different from the first set of computing nodes; Processing first data by the attention processing unit deployed on the first set of computing nodes to determine first attention information, the first data being determined based on the input request; Providing the first attention information to the second set of computing nodes via communication between the first set of computing nodes and the second set of computing nodes; Processing the first attention information by the feed-forward processing unit deployed on the second set of computing nodes to generate second data; And Generating response content for the input request based on the second data.
2. The method according to claim 1, further comprising: Providing the second data from the second set of computing nodes to the first set of computing nodes via the communication.
3. The method according to claim 1, further comprising: Determining target batch data based on the input request; And Dividing the target batch data into multiple micro-batch data, wherein the multiple micro-batch data includes the first data.
4. The method according to claim 3, wherein the process of the first set of computing nodes generating second attention information corresponding to third data is at least partially parallel to the process of providing the first attention information to the second set of computing nodes, the third data being the next micro-batch data after the first data in the multiple micro-batch data.
5. The method according to claim 3, wherein the receiving time of the second data by the first set of computing nodes is earlier than the completion time of the first set of computing nodes generating attention information corresponding to fourth data, the fourth data being the last micro-batch data in the multiple micro-batch data.
6. The method according to claim 3, wherein the deployment plan of the model is determined based on hardware information and / or model information of the model, the hardware information being associated with the first set of computing nodes and / or the second set of computing nodes, and the deployment plan at least indicating the number of the multiple micro-batches.
7. The method according to claim 1, wherein the first set of computing nodes and the second set of computing nodes correspond to different types of computing devices.
8. The method according to claim 1, wherein the sending process associated with the first attention information includes: In response to a graphics processing unit generating the first attention information, updating a first identifier from a first value to a second value by the graphics processing unit; And In response to detecting that the first identifier is updated to the second value, using a sending queue by a central processing unit to send a preset portion of the first attention information in a first memory of the graphics processing unit to a receiving party corresponding to the sending queue.
9. The method according to claim 1, wherein the receiving process associated with the first attention information includes: In response to a preset portion of the first attention information being received from a sender using a receive queue, a central processing unit updates a second identifier from a third value to a fourth value; and In response to detecting that the second identifier is updated to the fourth value, a graphics processing unit copies the preset portion of the first attention data from a buffer corresponding to the receive queue to a second memory of the graphics processing unit.
10. The method according to claim 9, further comprising: In response to the preset portion of the first attention data being copied to the second memory, the graphics processing unit updates the value of a third identifier to indicate that the buffer is allowed to be reused.
11. An apparatus for information processing, comprising: a receiving module configured to receive an input request for a model, wherein an attention processing unit of the model is deployed on a first set of computing nodes, and a feed-forward processing unit of the model is deployed on a second set of computing nodes different from the first set of computing nodes; an attention processing module configured to process first data using the attention processing unit deployed on the first set of computing nodes to determine first attention information, the first data being determined based on the input request; a communication module configured to provide the first attention information to the second set of computing nodes via communication between the first set of computing nodes and the second set of computing nodes; a feed-forward processing module configured to process the first attention information by the feed-forward processing unit deployed on the second set of computing nodes to generate second data; and a generating module configured to generate response content for the input request based on the second data.
12. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.