Distributed processing method for tasks, controller, controller cluster and electronic device
By splitting tasks in the controller cluster and utilizing idle BMC resources for distributed computing, the problem of high BMC resource utilization is solved, and the response speed and concurrent processing capability of the question-answering system are improved.
Patent Information
- Application Number
- CN202510971871.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-15
AI Technical Summary
In traditional question-and-answer systems, the baseboard management controller (BMC) has a high resource utilization rate when processing large-scale knowledge data and highly concurrent question-and-answer requests, which affects the operation of the main business, and the single-node computing model is difficult to meet actual needs.
By splitting tasks in the controller cluster and utilizing the idle resources of multiple baseboard management controllers for distributed computing, the tasks are split based on the weight matrices of the multi-head attention layer and feedforward network layer of the large language model, and the subtasks are assigned to the idle controllers for processing, and the results are finally integrated.
It effectively utilizes the computing resources of the BMC cluster, reduces the CPU operating pressure of a single BMC, improves the response speed and concurrent processing capability of the question-and-answer system, and solves the problem of high resource utilization.
Smart Images

Figure CN120492130B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a distributed task processing method, a controller, a controller cluster, and an electronic device. Background Art
[0002] With the rapid development of modern information technology, question-answering systems are increasingly being used in various fields. Whether it's internal enterprise knowledge querying, data center equipment management consulting, or interactive Q&A for intelligent hardware, efficient and accurate Q&A services can significantly improve work efficiency and user experience.
[0003] Traditional question-answering systems often rely on the computing resources of a single device. When processing large amounts of knowledge data and highly concurrent question-answering requests, performance bottlenecks gradually become apparent. Problems such as slow response times and limited processing power often lead to long wait times for users to access question-answering services, and even system crashes. With the continuous growth of data volumes and the increasing diversity of user needs, this single-node computing model is no longer able to meet the needs of real-world applications. Summary of the Invention
[0004] The present application provides a distributed task processing method, controller, controller cluster and electronic device to at least solve the problem in the related art that the resource occupancy rate of the large prediction model in the baseboard management controller is high during operation, which easily affects the main business operation of the baseboard management controller.
[0005] The present application provides a distributed task processing method, including: obtaining a task to be processed; determining an idle controller in a controller cluster, the controller cluster including multiple baseboard management controllers, the idle controller being a baseboard management controller whose task volume processed at the current moment is less than a preset task volume; splitting the task to be processed based on the number of attention heads of the multi-head attention layer of a large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple subtasks to be processed, and allocating each subtask to be processed to the corresponding idle controller, so that the idle controller processes the subtask to be processed to obtain a processed subtask; and integrating all the processed subtasks to obtain a processed task.
[0006] The present application also provides a controller, including: an acquisition unit, for acquiring tasks to be processed; a determination unit, for determining an idle controller in a controller cluster, wherein the controller cluster includes multiple baseboard management controllers, and the idle controller is a baseboard management controller whose task volume processed at the current moment is less than a preset task volume; a splitting unit, for splitting the task to be processed based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple sub-tasks to be processed, and assigning each of the sub-tasks to be processed to the corresponding idle controller, so that the idle controller processes the sub-task to be processed to obtain a processed sub-task; an integration unit, for integrating all the processed sub-tasks to obtain a processed task.
[0007] The present application also provides a controller cluster, including: a controller, used to perform the following steps: obtaining tasks to be processed; determining an idle controller in the controller cluster, the controller cluster including multiple baseboard management controllers, the idle controller being a baseboard management controller whose task volume processed at the current moment is less than a preset task volume; splitting the tasks to be processed based on the number of attention heads of the multi-head attention layer of the large language model and the weight matrix of the feedforward network layer in the baseboard management controller to obtain multiple sub-tasks to be processed, and allocating each of the sub-tasks to be processed to the corresponding idle controller, so that the idle controller processes the sub-task to be processed to obtain a processed sub-task; integrating all the processed sub-tasks to obtain a processed task; at least one idle controller, the idle controller being a baseboard management controller whose task volume processed at the current moment is less than a preset task volume; at least one non-idle controller, the non-idle controller being a baseboard management controller whose task volume processed at the current moment is greater than or equal to the preset task volume, the controller, the idle controller and the non-idle controller are communicatively connected.
[0008] The present application also provides an electronic device comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a distributed processing method for performing any one of the tasks described.
[0009] Through this application, a controller cluster includes multiple baseboard management controllers. When receiving a task to be processed, the task to be processed is split based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer, and each sub-task to be processed is assigned to the corresponding idle controller. The idle controller processes the sub-task to be processed and integrates the processed sub-tasks to obtain the final processed tasks. Based on the above steps, the problem of high resource occupancy during operation of the large prediction model in the BMC in the related technology, which is easy to affect the main business operation of the BMC, is solved. In addition, the idle BMC computing resources in the cluster are utilized. When a BMC receives user input and needs to call AI model calculation, the computing tasks are assigned to other idle BMCs according to the division strategy, thereby reducing the CPU operating pressure of the BMC that the user directly interacts with. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 A flowchart of a distributed task processing method provided in an embodiment of the present application;
[0012] Figure 2 A flowchart of another distributed task processing method provided in an embodiment of the present application;
[0013] Figure 3 A flowchart of a distributed processing method for another task provided in an embodiment of the present application;
[0014] Figure 4 A schematic diagram of the structure of a controller provided in an embodiment of the present application;
[0015] Figure 5 A schematic diagram of a controller cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0018] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the distributed processing method of the task depends, the specific application environment architecture or specific hardware architecture is described here.
[0020] The embodiment of the present application provides a distributed processing method for tasks, such as Figure 1 As shown, including:
[0021] Step S101, obtaining tasks to be processed;
[0022] In a distributed computing system, acquiring pending tasks can refer to an AI model, compute node, or system component receiving data processing, inference, or training tasks from a task pool or upstream components. For example, when a large language model based on the Transformer architecture is running in a distributed environment, a BMC (baseboard management controller) node may need to "acquire pending tasks." This means it receives specific computational tasks from the task dispatcher (the tasks acquired by the user client) during the model inference process, such as processing a portion of a specific input sequence or participating in the parallel computation of a multi-head attention mechanism.
[0023] Step S102: determining an idle controller in a controller cluster, where the controller cluster includes multiple baseboard management controllers. An idle controller is a baseboard management controller whose current processing task volume is less than a preset task volume.
[0024] Among them, the Baseboard Management Controller (BMC), as a key component in the server system, is mainly responsible for monitoring the hardware status of the server, managing system logs, performing remote management and other functions. It usually has an independent processor, memory and network interface, and can run independently outside the server operating system. Recently, our company has a related patented design to embed a large language model into the BMC to create a question-and-answer system designed specifically for the BMC. This AI (Artificial Intelligence) system uses the BMC's local hardware resources for calculations, and users can use this function by visiting the BMC page. However, in a server cluster, each server will have a BMC for management, and users generally only use the AI services of one or several BMCs at the same time. Other idle BMCs only perform hardware monitoring and management-related processes. Therefore, there are a large number of potential computing resources that have not been fully utilized.
[0025] Step S103: Split the task to be processed based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple subtasks to be processed, and assign each subtask to a corresponding idle controller so that the idle controller processes the subtask to obtain a processed subtask;
[0026] Distributed reasoning is a technology that distributes reasoning tasks across multiple computing nodes for parallel execution. By breaking complex reasoning tasks into multiple subtasks and assigning them to different nodes for simultaneous processing, distributed reasoning can significantly improve reasoning efficiency and system processing capabilities. In question-answering systems based on large language models, distributed reasoning can balance the model's computational load across multiple nodes, thereby accelerating question-answering responses and improving the system's concurrent processing capabilities.
[0027] While there has been some research and application of distributed reasoning, most of this is based on specialized high-performance computing clusters. These clusters are not only expensive in hardware but also complex to deploy and maintain. This cost is unbearable for resource-limited applications. However, BMC clusters, widely present in server systems and possessing sufficient computing and communication capabilities, can achieve low-cost and efficient distributed reasoning based on a large language model architecture, meeting the needs of some low-cost and efficient scenarios.
[0028] Step S104: Integrate all processed subtasks to obtain a processed task.
[0029] Currently, lightweight large language models can be run within the BMC, and all calculations are performed using the BMC's hardware resources during model execution. However, since the BMC lacks computational acceleration hardware such as a graphics processing unit (GPU), it can only use the CPU for inference. As a result, the CPU occupancy rate is very high during model inference, which risks impacting the operation of the main business. The above-mentioned embodiments utilize a baseboard management controller (BMC) cluster-based, Transformer-based large language model tensor parallel distributed inference method and system, fully utilizing the computing resources of the BMC cluster and implementing efficient distributed inference of large language models within the BMC through a tensor parallel strategy, thereby improving the performance of the question-answering assistant and reducing hardware costs and deployment complexity.
[0030] In the above embodiment, each BMC is capable of running a lightweight, large language model based on the transformer architecture and storing the model file. Users can access AI services through the BMC's web page and interact with the BMC's built-in AI model. Generally speaking, users will only use the AI services of one (or several) BMCs in the server cluster at a time, leaving the other BMCs with a large amount of idle computing resources. The traditional solution is to directly use the resources of the BMC with which the user interacts to run the model. However, due to the large number of AI model parameters, using the resources of only one BMC for all calculations will result in excessive CPU usage on that BMC. Therefore, the above embodiment optimizes this issue by utilizing the idle BMC computing resources in the cluster. When a BMC receives user input and needs to invoke AI model calculations, the computing task is allocated to other idle BMCs based on a partitioning strategy, reducing the CPU operating pressure on the BMC with which the user directly interacts.
[0031] That is Figure 2 As shown, first, the user client accesses one of the BMCs to process the task, then collects data such as the task amount of each BMC in the BMC cluster, determines the idle BMCs in the BMC cluster, and distributes the computing tasks to each idle BMC based on the collected data.
[0032] It should be noted that the BMC accessed by the user client is now referred to as the target BMC, and the selected idle BMCs include the target BMC. That is, after the user client accesses the target BMC, multiple idle controllers other than the target BMC are identified from the BMC cluster. Tasks are then split and distributed to the target BMC and the idle controllers other than the target BMC. After the target BMC and the idle controllers other than the target BMC complete the distributed tasks, the tasks processed by each BMC are aggregated and consolidated in the target BMC to obtain the final processed tasks.
[0033] That is, through this application, a controller cluster includes multiple baseboard management controllers. When receiving a task to be processed, the task to be processed is split based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer, and each sub-task to be processed is assigned to the corresponding idle controller. The idle controller processes the sub-task to be processed to obtain the processed sub-task and then integrates it to obtain the final processed task. Based on the above steps, the problem of high resource occupancy during operation of the large prediction model in the BMC in the related technology, which is easy to affect the main business operation of the BMC, is solved. In addition, the idle BMC computing resources in the cluster are utilized. When a BMC receives user input and needs to call the AI model calculation, the computing task is assigned to other idle BMCs according to the division strategy, thereby reducing the CPU operating pressure of the BMC that the user directly interacts with.
[0034] In some embodiments, each controller in the controller cluster adopts the Intelligent Platform Management Interface communication protocol (IPMI) or the Extensible Service Interface communication protocol for communication connection, and the method also includes: when each controller in the controller cluster adopts the Intelligent Platform Management Interface communication protocol for communication connection, encapsulating the data to be sent into Intelligent Platform Management Interface information and sending it to the target controller, the Intelligent Platform Management Interface information includes the identification of the target controller, the data type, data length and data content of the data to be sent; when each controller in the controller cluster adopts the Extensible Service Interface communication protocol for communication connection, encapsulating the data to be sent into target format data, and sending the target format data to the target controller through a request, and the target format is a format recognizable by the Extensible Service Interface communication protocol.
[0035] The Redfish protocol is used as the communication protocol for the Scalable Service Interface. Each BMC in the BMC cluster uses the Intelligent Platform Management Interface (IPMI) or Redfish protocol to implement efficient tensor parallel inference. First, the BMC cluster must be initialized and configured, assigning each BMC node a unique identifier and communication address. A communication network based on the IPMI or Redfish protocol is then established to ensure stable and efficient data transmission between nodes.
[0036] That is, the IPMI communication process is that when a node needs to send an intermediate result to other nodes, the result is encapsulated into an IPMI message. The message contains the identifier of the target node, data type, data length and specific data content. The message is sent to the target node through the IPMI message transmission mechanism. After the target node receives the message, it parses the message content and extracts the required data. The Redfish communication process is that if the Redfish protocol is adopted, the sending node encapsulates the intermediate result into JSON data that complies with the RESTful API specification. The data is sent to the corresponding API address of the target node through an HTTP POST or PUT request. After receiving the request, the target node parses the JSON data and obtains the required intermediate result. The above target format data is JSON data that complies with the RESTful API specification.
[0037] For IPMI communication configuration, if using the IPMI protocol, set IPMI communication parameters for each BMC node, including the IP address, subnet mask, and gateway. Configure IPMI authentication information, such as the username and password, to ensure communication security. Test connectivity between nodes using the IPMI command interface to ensure that IPMI messages can be sent and received normally.
[0038] For Redfish communication configuration, if the Redfish protocol is used, assign a unique Restful API address to each BMC node. Configure Redfish authentication mechanisms, such as username and password-based authentication, to ensure communication security. Test whether the Redfish service between nodes is functioning properly by sending HTTP requests.
[0039] IPMI communicates through the BMC's independent network interface, meaning that even if the server's main operating system is inoperative, the server can still be managed through IPMI communication, providing a reliable remote management method for the BMC. IPMI supports password-based authentication and encrypted communication, ensuring the security of BMC communications, which is critical for sensitive server management operations.
[0040] However, IPMI primarily focuses on server hardware monitoring and basic management functions, and its functionality may be limited for more advanced application scenarios, such as distributed computing or complex data exchange. Furthermore, IPMI uses custom binary commands for communication, which may be less efficient than the HTTP-based Redfish protocol for complex task distribution and data transfer, especially when transferring large amounts of data. Due to its relatively static design, IPMI lacks scalability and flexibility, making it difficult to integrate new features and services, which can be a limitation when handling complex distributed computing tasks.
[0041] The Redfish protocol provides direct support for complex management operations and advanced services such as firmware updates, resource allocation, virtual media management, etc., which makes it easy to perform more complex tasks through BMC. Using HTTP and RESTful architecture, the Redfish protocol is highly compatible with modern web services and network architectures, making BMC communication easier to understand and implement. The Redfish protocol supports efficient data transmission. Due to its JSON-based data format and HTTP protocol, it can handle large amounts of data and complex data exchange, making it very suitable for task allocation and result aggregation in distributed computing environments. Redfish provides modern security mechanisms such as TLS encryption and OAuth2 authentication, while its architectural design allows flexible expansion and customization, which provides higher security guarantees and functional flexibility for BMC communication.
[0042] However, Redfish's advanced features and RESTful architecture may require more complex implementation. For devices with limited resources such as BMC, implementing Redfish services may require higher hardware requirements.
[0043] The choice between IPMI and Redfish depends on specific application requirements and scenarios. For basic hardware monitoring and management, IPMI is a mature and reliable option. However, if the goal is to build a BMC communication system that supports advanced services, efficient data processing, and high security standards, Redfish offers more powerful features and superior performance. In distributed computing scenarios, such as distributing large language model inference tasks to a BMC cluster, Redfish's efficient data transmission, scalability, and higher security make it a more preferred communication protocol, and is therefore generally chosen.
[0044] Whether using the IPMI or Redfish protocol, communication parameters and authentication information must be configured to ensure secure and stable inter-node communication. During data exchange, data encapsulation, transmission, and parsing processes are optimized based on protocol characteristics. For example, IPMI communication utilizes meticulously designed message structures, while Redfish communication adheres strictly to RESTful API specifications to reduce communication overhead and ensure that intermediate and final inference results are quickly and accurately transmitted and aggregated between nodes.
[0045] In addition, in some embodiments, hardware information about each BMC node, such as CPU performance, memory capacity, and network bandwidth, can also be collected to provide a basis for subsequent task allocation. This includes screening idle controllers based on CPU performance and memory capacity. The number of tasks handled by the BMC is directly related to CPU performance and memory capacity. When resources are limited, an excessive number of tasks may lead to excessive CPU utilization and insufficient memory on the BMC, thereby affecting the speed and quality of its task execution.
[0046] In some optional embodiments, splitting the tasks to be processed based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer includes the following steps:
[0047] Step S201, converting the format of the task to be processed into a vector format to obtain a task vector;
[0048] In the fields of artificial intelligence and machine learning, especially natural language processing tasks, raw text data needs to be converted into computer-understandable digital vectors for model processing. This conversion process can be accomplished using word embeddings, one-hot encoding, or other vectorization methods. For example, word embeddings can be used to map each word in a text to a fixed-length vector, and then the entire sentence or document can be represented as a sequence or matrix of word vectors.
[0049] Step S202: Based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer, the task vector is split to obtain multiple subtasks to be processed. The large language model is a converter model running in the idle controller. The large language models in each idle controller are the same, and the size of each subtask to be processed is the same.
[0050] Among them, the transformer model is a Transformer model. In the Transformer architecture, the multi-head attention mechanism is one of its core components, which allows the model to focus on different parts of the input in parallel, thereby improving the performance of the model. Each attention head actually performs attention calculations on different subspaces of the input vector. After step S201, the task vector has been converted into a format suitable for model processing. The next step is to split the task vector into multiple subtasks based on the number of attention heads in the BMC cluster, and each subtask corresponds to the calculation of an attention head.
[0051] The weight matrix of a feedforward network layer is often one of the largest parameters in the model, processing numerous matrix multiplication operations. By splitting the weight matrix by rows or columns, the computational load can be distributed across different BMC nodes in the cluster. Each node is then responsible for multiplying a portion of the weight matrix with the input vector, and the results are then aggregated to reconstruct the output of the entire feedforward network layer.
[0052] By splitting the task vector into multiple attention heads, the computation of each attention head can be performed in parallel, which significantly improves the speed of model inference. Because different attention heads can be processed simultaneously, the computing power of multiple nodes in the BMC cluster can be utilized, significantly reducing the overall computation time.
[0053] By splitting the task vectors and processing them in parallel on multiple BMC nodes, the model inference time is significantly reduced. This is because the computational tasks that originally needed to be processed sequentially on one node can now be performed simultaneously on multiple nodes, especially in the processing of multi-head attention mechanisms and feedforward network layers, which takes advantage of tensor parallelism.
[0054] Furthermore, this approach leverages the computing resources of idle BMC nodes in the cluster, avoiding overloading a single node. In a BMC cluster comprised of multiple servers, even if only one node is directly accessed by a user, other nodes can participate in model inference through task splitting, improving overall resource utilization. This also allows for faster responses to user requests, especially in high-load environments. This distributed computing approach effectively improves system processing capabilities and reduces user wait time.
[0055] Generally, BMCs in a BMC cluster have the same hardware performance, such as CPU performance, memory capacity, and network bandwidth. Therefore, the idle controller can be determined by simply determining the current workload of the BMC. That is, when hardware performance is the same, the BMC with the smallest workload is the idler. However, in some cases, BMCs in a BMC cluster have different hardware performance, such as CPU performance, memory capacity, and network bandwidth. Therefore, the idle controller needs to be determined by comprehensively considering the BMC's CPU performance, memory capacity, and current workload. For example, if the BMC has high CPU performance and memory capacity, the preset workload can be set slightly higher; if the BMC has low CPU performance and memory capacity, the preset workload can be set slightly lower.
[0056] In step S202, the sizes of the subtasks to be processed are the same, that is, after the idle BMCs are screened out by the preset task volume, the task to be processed is evenly split into subtasks to be processed and distributed to each idle BMC. This can indeed save a lot of computing power, and for BMCs with the same hardware performance such as CPU performance, memory capacity, and network bandwidth, and when the preset task volume is set to a small level, this mechanism also allows for relatively synchronous task processing. However, if the BMCs in the BMC cluster have different hardware performance such as CPU performance, memory capacity, and network bandwidth, or if the preset task volume is set to a large level, then even though different idle BMCs all meet the condition of being less than the preset task volume, the task volumes currently being processed by any two idle BMCs are different, which can lead to the problem of asynchronous processing speeds of the subtasks to be processed by any two idle BMCs.
[0057] To prevent the above problems, the task to be processed is split based on the number of attention heads in the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple subtasks to be processed, which also includes the following steps:
[0058] Step S301: Obtain performance parameters of an idle controller, where the performance parameters represent the maximum amount of tasks that the idle controller can handle at the current moment.
[0059] Step S302, based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller, the weight matrix of the feedforward network layer of the large language model and the performance parameters of the idle controller, the task to be processed is split to obtain multiple subtasks to be processed, and the size of each subtask to be processed is positively correlated with the performance parameters of the idle controller.
[0060] That is, by splitting tasks that are positively correlated with the performance parameters of idle controllers, the system can ensure that computing resources are fully utilized, avoiding resource waste or underutilization of computing power. By ensuring that each controller receives subtasks of appropriate size based on its actual performance parameters, it avoids the problem of some controllers being overloaded while other controllers' resources are idle, and achieves a balanced load distribution. Because subtasks are effectively assigned to multiple controllers, they can be processed in parallel, reducing waiting time during large language model inference and significantly improving the processing speed of model inference. Reasonable task allocation ensures that individual controllers will not be overloaded, causing system crashes or unpredictable response times, thereby improving the stability and reliability of the entire distributed inference system.
[0061] There are two ways to process the subtasks to be processed. The first is to first perform distributed processing through the multi-head attention layer to obtain the intermediate processing results, then integrate the intermediate processing results, and then distribute the integrated intermediate processing results to the BMC for processing. The second is to perform distributed processing through the multi-head attention layer, and after obtaining the intermediate processing results, continue to process them directly in the BMC without integration.
[0062] The first method is to first assign the subtask to be processed to the multi-head attention layer of the corresponding idle controller for processing. After obtaining the intermediate processing subtask (that is, an attention output matrix), the attention output matrix is integrated to obtain a total output matrix, and then the total output matrix is redistributed to each idle BMC. The idle BMC then performs calculations in the feedforward network layer to finally obtain the processed subtask.
[0063] That is, each pending subtask is assigned to a corresponding idle controller so that the idle controller processes the pending subtask to obtain a processed subtask, including the following steps:
[0064] Step S401: assign each pending subtask to a corresponding idle controller, so that the idle controller processes the pending subtask using a multi-head attention layer of a large language model to obtain an attention output matrix;
[0065] In the multi-head attention mechanism, the input vector is decomposed into three parts: query (Q), key (K), and value (V). These parts are then assigned to different attention heads for processing, with each head independently calculating attention weights. In this step, the model's computational tasks are decomposed, and different subtasks (i.e., the computation of different attention heads) are assigned to different idle controllers based on the idle status of each node in the BMC cluster. In this way, each idle controller is responsible for calculating the attention weights of its assigned attention head and generating an attention output matrix. By parallelizing the computational tasks, each idle controller can perform attention calculations simultaneously, greatly accelerating model inference speed while also balancing the use of computing resources.
[0066] Step S402: Integrate all attention output matrices to obtain a total output matrix;
[0067] Step S403 : Distribute the total output matrix to multiple idle controllers, so that the idle controllers continue to perform calculations according to the total output matrix to obtain processed subtasks.
[0068] The total output matrix is the output of the multi-head attention layer calculation. It is passed as input to subsequent layers of the model, such as the feedforward network layer, for further processing. In this step, the integrated total output matrix is redistributed to the idle controllers participating in the multi-head attention calculation, as well as other idle controllers that may be needed to continue the model's calculation process. Each controller performs calculations based on the assigned subset of the total output matrix, ultimately producing the processed subtask results.
[0069] Through parallel processing and matrix integration, the total model inference time is significantly reduced, improving the system's response speed. By integrating the attention output matrix, the integrity of the model inference is ensured, and the calculation results of each controller are aggregated to provide a unified input for subsequent processing, improving the consistency of calculations and system efficiency.
[0070] Since the format of the pending tasks is converted to a vector format to obtain the task vector, the task splitting actually involves splitting the vectors. For example, if each idle BMC's multi-head attention layer has 16 heads, and there are 4 nodes in the BMC cluster (i.e., 4 idle BMCs), the 16 heads can be evenly divided into 4 groups of 4 heads each. Each node (i.e., 1 idle BMC) is responsible for processing the Q, K, and V tensors corresponding to a group of heads. In other words, each pending subtask obtained by splitting the task vector is calculated using 4 heads, and each idle BMC only uses 4 heads to process the corresponding pending subtask. The intermediate processing subtask obtained by the multi-head attention layer of each idle BMC is actually an attention output matrix. The attention output matrices of all idle BMCs are first integrated into the total output matrix, and then the total output matrix is redistributed to each idle BMC. The idle BMC then performs the calculations of the feedforward network layer to finally obtain the processed subtask.
[0071] In models based on the Transformer architecture, the multi-head attention mechanism is a core component, designed to enable the model to simultaneously focus on different parts of the input sequence and capture dependencies at different levels. Each attention head independently processes the input information, calculating the corresponding attention weights and the weighted sum of the outputs. However, to maintain the overall consistency and effectiveness of the model, these independent attention output matrices need to be integrated. The original intention of the design of multi-head attention is to enable the model to understand the input information from multiple perspectives. Each head may focus on different aspects of the sequence or information of different granularity. Integrating the output matrices of all heads is actually a fusion of information at the model level, ensuring that the model can comprehensively consider the information captured by all attention heads to form a more comprehensive and in-depth understanding.
[0072] In some embodiments, each subtask to be processed is assigned to a corresponding idle controller so that the idle controller uses a multi-head attention layer of a large language model to process the subtask to be processed and obtain an attention output matrix, including: assigning each subtask to be processed to a corresponding idle controller so that the idle controller uses multiple target attention heads of the multi-head attention layer of the large language model to extract different features of the subtask to be processed and obtain multiple task features, and allowing the idle controller to obtain an attention output matrix based on the multiple task features, wherein a target attention head is used to extract a task feature, and multiple target attention heads are part of the attention heads of the multi-head attention layer, and different idle controllers use different target attention heads.
[0073] During model inference, the original input vector or subtask is partitioned and assigned to multiple idle controllers in the cluster. Each controller utilizes a specific attention head (target attention head) within the model to extract distinct features from its assigned subtask. These attention heads form part of a multi-head attention layer, each focusing on extracting features at a different aspect or granularity of the input sequence. Different controllers use different target attention heads, enabling the model to analyze input information from multiple perspectives and dimensions, enhancing its expressiveness and understanding of the input data.
[0074] On each idle controller, its assigned target attention head calculates the corresponding attention weights based on the extracted task features. This process involves a dot product operation between the query, key, and value vectors, as well as softmax normalization of the attention weights. Each controller performs a weighted summation of its value vectors based on the calculated attention weights to generate an attention output matrix associated with that subtask. This attention output matrix reflects the information filtered and weighted by the attention mechanism and is the result of the target attention head processing the subtask.
[0075] Ultimately, all attention output matrices generated by different controllers need to be integrated into a total output matrix to reflect the comprehensive processing results of the multi-head attention layer on the input sequence. This integration ensures the continuity and integrity of information transmission during the model inference process.
[0076] Computational tasks for different target attention heads are rationally distributed among idle controllers in the cluster, balancing the computational load and fully utilizing the cluster's computing resources, avoiding resource waste and single-point overload. Because each controller uses a different target attention head, the features they extract are also unique. This diverse feature set, weighted through the attention mechanism, enriches the model's understanding and representation of the input data and improves its predictive accuracy.
[0077] The specific method steps for determining the target number of attention heads include the following: obtaining the number of idle controllers; obtaining the total number of attention heads of the multi-head attention layer of the large language model; and determining the ratio of the total number of attention heads to the number of idle controllers as the target number of attention heads for each idle controller.
[0078] First, the system automatically detects and counts the number of idle controllers in the current BMC cluster. Idle controllers are BMC devices that are not currently performing primary tasks and have additional computing resources available to assist in model inference. At the same time, the system also needs to know the total number of attention heads in the multi-head attention layer of the large language model. Attention heads are key components in the model used to process different parts of the input sequence in parallel, and their number directly affects the model's parallel processing capabilities. The total number of attention heads is compared with the number of idle controllers to obtain a ratio between the two. This ratio determines the average number of attention heads that each idle controller should process. Based on the calculated ratio, the system determines the target number of attention heads for each idle controller. This number is distributed as evenly as possible to each idle controller to ensure a balanced computing load.
[0079] The query (Q), key (K), and value (V) tensors in the multi-head attention layer are divided by the head dimension based on the number of nodes and computing power of the BMC cluster. For example, if a multi-head attention layer has 16 heads and the BMC cluster has 4 nodes, the 16 heads can be evenly divided into 4 groups of 4 heads each, with each node responsible for processing the Q, K, and V tensors corresponding to a group of heads. The advantage of this division is that the computing tasks of each node are relatively balanced and the parallel computing power between nodes can be fully utilized. At the same time, the index information of the head responsible for each node is recorded to facilitate subsequent calculations and data exchange.
[0080] In Transformer distributed inference, row-wise partitioning is a method for implementing model parallelism. Its core idea is to split the weight matrix in a neural network into different computing devices (such as GPUs / TPUs) by row. This partitioning method is primarily suitable for modules with intensive matrix multiplication operations (such as attention layers and feedforward neural network layers). By distributing the computational load across multiple devices, it overcomes single-device memory bottlenecks and improves computational efficiency.
[0081] The core calculation of the attention layer is: .
[0082] Among them, Q is the query, K is the key, V is the value, d k is the feature dimension. In actual calculation, Q, K, and V are generated by linear projection of the input, such as Q=XW Q, W Q is the weight matrix and X is the input.
[0083] The row segmentation of the weight matrix includes: Q For example, assuming that the device is divided into N devices (i.e. BMC) by row, the original weight matrix W Q After the split, each device holds W i , where W i The number of rows is W Q 1 / N, i=1,2,...,N.
[0084] The distributed computing process includes the input X being broadcast to all devices, and each device calculating , Similarly, the projection matrix W of K and V K 、W V Also split by row, get K i 、V i When calculating the attention score matrix, it is necessary to combine the Q i and K i Aggregation: The local score matrices of all devices are accumulated and summarized to obtain the complete attention score matrix.
[0085] In some embodiments, the total output matrix is distributed to multiple idle controllers so that the idle controllers continue to perform calculations based on the total output matrix to obtain processed subtasks, including: distributing the total output matrix to multiple idle controllers so that the idle controllers calculate the product of the total output matrix and a first weight matrix to obtain processed subtasks, wherein the first weight matrix is obtained based on the weight matrix of the feedforward network layer of the large language model.
[0086] Specifically, the multiplication operation of the total output matrix and the first weight matrix is parallelized, which means that calculations that originally needed to be performed sequentially on a single node can now be performed simultaneously on multiple nodes, significantly improving the calculation speed and reducing the overall time consumption of model inference. By executing the calculation of the feedforward network layer in parallel on multiple idle controllers, the computing power of the BMC cluster is fully utilized, resource waste is avoided, the need for additional hardware investment is reduced, and the overall resource utilization and economic benefits of the system are improved. That is, by parallelizing the calculation process of the product of the total output matrix and the first weight matrix and distributing it to multiple idle controllers in the BMC cluster, not only can the calculation speed and efficiency of model inference be significantly improved, but also the optimal allocation and utilization of resources can be achieved, thereby enhancing the stability and scalability of the system.
[0087] When using a BMC cluster for distributed inference of the feedforward network layer of a large language model, partitioning the weight matrix by rows or columns is one of the core strategies for achieving parallel computing. The first weight matrix is obtained by partitioning the weight matrix of the feedforward network layer of the large language model by each idle controller. The partitioning is performed by rows or columns, and the first weight matrix of each idle controller is different.
[0088] The feedforward network layer usually consists of two linear layers, involving the calculation of the weight matrix and bias vector. The weight matrix can be divided according to the input or output dimension. If the weight matrix dimension of the first linear layer is d_in×d_hidden, it can be divided according to the input dimension d_in, divided into several equal parts, and each part is assigned to a BMC node. For example, if d_in=1024 and there are 4 nodes, each node is responsible for processing the weight matrix part with a dimension of 256×d_hidden. The bias vector is also allocated according to the same division method.
[0089] The weight matrix of the feedforward network layer is partitioned by rows, with each idle controller processing only a subset of the rows. This partitioning approach is particularly suitable for the first linear layer of the feedforward network layer, where the weight matrix is typically represented as a mapping from the input dimension to the hidden dimension. Another approach is to partition by columns, with each controller processing a subset of the columns of the weight matrix. This approach is more suitable for the second linear layer of the feedforward network layer, which represents a mapping from the hidden dimension to the output dimension. After the local computations are completed, the output results of each controller need to be aggregated and integrated to obtain the final feedforward network layer output. If a row-based partitioning strategy is adopted, the output vectors of all controllers are vertically concatenated; if a column-based partitioning strategy is adopted, horizontal concatenation and matrix multiplication operations are required at the upper layer. Therefore, in general, row-based partitioning is chosen to reduce communication overhead.
[0090] Dividing the weight matrix by rows or columns ensures that the computational load of each controller is roughly the same, avoiding situations where some controllers are overloaded while other controller resources are underutilized. The communication cost under this method is mainly reflected in the aggregation of the output vectors. Since the output vector dimension is usually much smaller than the number of rows of the weight matrix, the cost of aggregation communication is relatively low. Although the number of weight matrix columns that need to be transmitted under this method may be large, the overall communication cost is also effectively controlled due to the low data interaction during the calculation process. That is, the strategy of dividing the weight matrix by rows or columns can not only efficiently perform the feedforward network layer calculations of large language models based on the Transformer architecture in a distributed environment, but also ensure the accuracy and consistency of the calculations, while optimizing communication costs and resource usage.
[0091] Splitting large matrices by row reduces single-device memory usage and is suitable for processing high-dimensional models or in memory-constrained scenarios (e.g., BMC scenarios). Row-wise matrix multiplication naturally lends itself to multi-device parallel computing, accelerating inference speed. Row-wise matrix splitting is a common optimization technique when deploying large models on devices with limited video memory, such as consumer-grade GPUs.
[0092] That is, splitting by rows is suitable for large-dimensional models and memory-constrained scenarios, while splitting by columns is suitable for highly parallel computing and communication-sensitive scenarios.
[0093] When the weight matrix is partitioned by row, different BMC nodes are responsible for calculating different parts of the output vector y. For example, if the weight matrix W is divided equally into n parts by row, each BMC node receives a sub-matrix W_i. Each node multiplies its own sub-matrix by the input vector x to obtain the partial result y_i of the output vector y. Finally, the partial results of all nodes are combined to obtain the complete output vector y. In this case, each node's calculations are relatively independent, and it only needs to process its own sub-matrix and input vector. The communication overhead mainly comes from the aggregation of the results.
[0094] The divided tensors are rationally distributed to different BMC nodes, and each node is equipped with corresponding computing tasks and parameters. For each node, it is clear what part of the tensor it needs to process, the calculation steps, and how it interacts with other nodes. For example, in the calculation of the multi-head attention layer, in addition to calculating the attention score of the head it is responsible for, each node also needs to send part of the results to other nodes for merging. At the same time, each node is provided with necessary parameters, such as the type of activation function, the parameters of the normalization layer, etc., to ensure the consistency and accuracy of the calculation. In addition, the data dependencies and communication requirements between each node need to be recorded to provide guidance for subsequent parallel computing and data synchronization.
[0095] In some embodiments, each subtask to be processed is assigned to a corresponding idle controller so that the idle controller processes the subtask to be processed to obtain a processed subtask, including: assigning each subtask to be processed to a corresponding idle controller so that the idle controller uses a multi-head attention layer of a large language model to process the subtask to be processed to obtain an attention output matrix, and allowing each idle controller to continue calculating according to the attention output matrix to obtain a processed subtask.
[0096] In models based on the Transformer architecture, the conventional practice is to combine the output matrices generated by all heads of a multi-head attention mechanism before inputting them into the feedforward network layer for further processing. However, directly using the output matrices generated by the attention heads as independent inputs, skipping the combination step, has unique theoretical considerations and potential benefits. The output matrix of each attention head represents the model's understanding of the input data from different perspectives or levels. Directly inputting these matrices into the feedforward network layer can preserve more diverse information from different heads, which theoretically may help the model learn richer feature representations. Combining the attention output matrices typically requires additional communication between nodes to summarize the outputs of all heads. Not combining means that nodes can directly use their own attention output matrices for further computation, reducing data transmission and communication latency, and improving computational efficiency in distributed environments. Without combining, each idle controller can independently use its own attention output matrix for feedforward network layer computation. This highly parallel computation approach may further improve model inference speed by eliminating the need to redistribute the combined data to all nodes.
[0097] In some embodiments, each subtask to be processed is assigned to a corresponding idle controller so that the idle controller uses the multi-head attention layer of the large language model to process the subtask to be processed, obtains the attention output matrix, and enables each idle controller to continue to calculate according to the attention output matrix to obtain the processed subtask, including: assigning each subtask to be processed to a corresponding idle controller so that the idle controller uses the multi-head attention layer of the large language model to process the subtask to be processed, obtains the attention output matrix, and enables each idle controller to calculate the product of the attention output matrix and the second weight matrix to obtain the processed subtask, wherein the second weight matrix is obtained according to the weight matrix of the feedforward network layer of the large language model.
[0098] Each idle controller directly performs matrix multiplication on the attention output matrix with the submatrix of the second weight matrix assigned to it, resulting in its own "local" processed subtask. This approach avoids the step of consolidating the attention output matrices into a global matrix after the multi-head attention layer, reducing communication and consolidation overhead. After all controllers complete their local computations, their results need to be aggregated to form the final model output.
[0099] By directly multiplying the attention output matrix with the second column-partitioned weight matrix, the computing power of each controller in the BMC cluster can be fully utilized, enabling parallel processing and significantly improving the computational speed of model inference. By avoiding the integration step of the attention output matrix, the number of large-scale data exchanges between nodes is reduced, thereby reducing communication latency and bandwidth consumption, which is particularly important for improving efficiency in distributed computing environments. This strategy simplifies model management and task scheduling in distributed environments. Each controller independently handles subtasks and local weight matrix multiplication, reducing reliance on centralized scheduling and resource management.
[0100] The second weight matrix is obtained by the idle controller dividing the weight matrix of the feedforward network layer of the large language model by columns, and the second weight matrix of each idle controller is different.
[0101] Splitting the second weight matrix by columns and matching the corresponding attention output matrix with it helps achieve a more even distribution of computing tasks in the BMC cluster, avoids local overload of computing resources, and improves the overall stability and reliability of the system.
[0102] For ease of description, some nouns or terms involved in the embodiments of the present application are explained below:
[0103] This embodiment relates to a specific distributed processing method for tasks, such as Figure 3 As shown, the client first inputs a request (for example, a request to call BMC resources to calculate pending tasks). The baseboard management controller (BMC) receives the request sent by the client and calls the large language model (i.e., the Transformer large language model AI module). It then counts the number of idle controllers, determines the partitioning strategy, and splits the computing tasks (i.e., pending tasks). It then uses the Intelligent Platform Management Interface Communication Protocol (IPMI) or the Extensible Service Interface Communication Protocol (Redfish) to send the computing tasks (the split pending tasks) to each BMC. Each BMC calculates the split pending tasks (which may include the synchronization of intermediate results). Finally, it uses the Intelligent Platform Management Interface Communication Protocol (IPMI) or the Extensible Service Interface Communication Protocol (Redfish) to aggregate the computing results and return the results to the user.
[0104] The specific detailed steps of the above embodiment are as follows:
[0105] Data receiving step: Each BMC node receives the assigned tensor data and task instructions through the IPMI or Redfish protocol.
[0106] Parallel computing steps: Based on the algorithmic logic of the large language model, each BMC node performs parallel computing on the tensor portion it is responsible for. In the multi-head attention layer, partial attention scores are calculated and weighted summed; in the feedforward network layer, matrix multiplication and activation function calculations are performed.
[0107] Communication of intermediate results: During the computing process, nodes use the IPMI or Redfish protocol to exchange data as needed to complete cross-node computing tasks. The IPMI communication process is that when a node needs to send intermediate results to other nodes, the result is encapsulated into an IPMI message. The message contains the target node's identifier, data type, data length, and specific data content. The message is sent to the target node through the IPMI message transmission mechanism. After receiving the message, the target node parses the message content and extracts the required data. The Redfish communication process is that if the Redfish protocol is used, the sending node encapsulates the intermediate results into JSON data that complies with the RESTful API specification. The data is sent to the corresponding API address of the target node through an HTTP POST or PUT request. After receiving the request, the target node parses the JSON data and obtains the required intermediate results.
[0108] Intermediate result synchronization (optional): To ensure the accuracy of inference, nodes regularly synchronize intermediate results through IPMI or Redfish protocol.
[0109] IPMI synchronization mechanism: Set a synchronization period. Each node encapsulates its intermediate results into IPMI messages and broadcasts them to other nodes at the synchronization time. After receiving the messages, other nodes update their own copies of the intermediate results.
[0110] Redfish synchronization mechanism: Periodically sends HTTP GET requests to the specified API address of other nodes to obtain their latest intermediate results. The obtained results are compared and updated with the local copy.
[0111] In the above embodiment, by using tensor parallel distributed reasoning on the BMC cluster, the computing resources of each BMC node are fully utilized to parallelize the reasoning tasks of the large language model, significantly improving the response speed and concurrent processing capabilities of the question-answering assistant, providing users with efficient question-answering services. Furthermore, the use of IPMI and Redfish protocols ensures efficient and stable data communication between BMCs, further improving reasoning efficiency.
[0112] Furthermore, leveraging the existing BMC cluster for distributed inference eliminates the need to upgrade the hardware of individual BMCs, reducing hardware costs. Furthermore, BMC clusters of similar products all come pre-installed with a unified image, ensuring environmental consistency. BMC clusters offer excellent flexibility and scalability. Based on actual Q&A needs, the number of BMC nodes participating in inference can be flexibly adjusted, or the cluster can be upgraded and expanded to accommodate Q&A tasks of varying scales. The universality of the IPMI and Redfish protocols enables the system to easily integrate with different types of BMC devices.
[0113] Each BMC node performs parallel computations on the tensors it is responsible for, based on the logic of the large language model algorithm. During critical computational steps, such as the calculation of partial attention scores in the multi-head attention layer and the matrix multiplication and activation function calculations in the feedforward network layer, nodes exchange data on demand via IPMI or Redfish protocols, regularly synchronizing intermediate results. By setting a reasonable synchronization period and using broadcast or directed requests, each node ensures that its computations are based on consistent data, ensuring accurate and efficient reasoning.
[0114] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0115] The embodiment of the present application also provides a controller, such as Figure 4 As shown, it includes an acquisition unit 10, a determination unit 20, a splitting unit 30 and an integration unit 40, wherein the acquisition unit 10 is used to acquire tasks to be processed; the determination unit 20 is used to determine an idle controller in a controller cluster, the controller cluster includes multiple baseboard management controllers, and the idle controller is a baseboard management controller whose task volume processed at the current moment is less than the preset task volume; the splitting unit 30 is used to split the task to be processed based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple subtasks to be processed, and each subtask to be processed is assigned to the corresponding idle controller, so that the idle controller processes the subtask to be processed and obtains the processed subtask; the integration unit 40 is used to integrate all the processed subtasks to obtain the processed task.
[0116] In the above embodiment, the controller cluster includes multiple baseboard management controllers. When receiving a task to be processed, the task to be processed is split based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer, and each sub-task to be processed is assigned to the corresponding idle controller. The idle controller processes the sub-task to be processed to obtain the processed sub-task and then integrates it to obtain the final processed task. Based on the above steps, the problem in the related art that the resource occupancy rate of the large prediction model in the BMC is high during operation, which is easy to affect the main business operation of the BMC. In addition, the idle BMC computing resources in the cluster are utilized. When a BMC receives user input and needs to call the AI model calculation, the computing task is assigned to other idle BMCs according to the division strategy, thereby reducing the CPU operating pressure of the BMC that the user directly interacts with.
[0117] In some embodiments, the splitting unit includes a conversion module and a first splitting module. The conversion module is used to convert the format of the task to be processed into a vector format to obtain a task vector. The first splitting module is used to split the task vector based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple subtasks to be processed. The large language model is a Transformer model running in the idle controller. The large language model in each idle controller is the same, and the size of each subtask to be processed is the same. The computing power of multiple nodes in the BMC cluster is utilized, thereby significantly reducing the overall computing time.
[0118] In some embodiments, the splitting unit includes a first allocation module, a first integration module, and a distribution module. The first allocation module is used to allocate each pending subtask to the corresponding idle controller, so that the idle controller uses the multi-head attention layer of the large language model to process the pending subtask and obtain an attention output matrix; the first integration module is used to integrate all attention output matrices to obtain a total output matrix; and the distribution module is used to distribute the total output matrix to multiple idle controllers, so that the idle controllers continue to calculate based on the total output matrix to obtain processed subtasks. This improves the consistency of calculation and system efficiency.
[0119] In some embodiments, the first allocation module includes a first allocation submodule for allocating each pending subtask to a corresponding idle controller, so that the idle controller uses multiple target attention heads of the multi-head attention layer of the large language model to extract different features of the pending subtask, thereby obtaining multiple task features, and the idle controller obtains an attention output matrix based on the multiple task features, wherein one target attention head is used to extract one task feature, and multiple target attention heads are part of the attention heads of the multi-head attention layer, and different idle controllers use different target attention heads. This diverse feature set enriches the model's understanding and representation of the input data through weighted summation of the attention mechanism, thereby improving the model's prediction accuracy.
[0120] In some embodiments, the device further includes a first acquisition module, a second acquisition module, and a determination module. The first acquisition module is used to obtain the number of idle controllers; the second acquisition module is used to obtain the total number of attention heads of the multi-head attention layer of the large language model; and the determination module is used to determine the ratio of the total number of attention heads to the number of idle controllers as the target number of attention heads for each idle controller. The advantage of this division is that the computing tasks of each node are relatively balanced, and the parallel computing capabilities between nodes can be fully utilized. At the same time, the index information of the head that each node is responsible for is recorded for subsequent calculations and data interaction.
[0121] In some embodiments, the distribution module includes a distribution submodule for distributing the total output matrix to multiple idle controllers, so that the idle controllers calculate the product of the total output matrix and a first weight matrix to obtain a processed subtask, where the first weight matrix is obtained based on the weight matrix of the feedforward network layer of the large language model. By parallelizing the calculation process of the product of the total output matrix and the first weight matrix and distributing it to multiple idle controllers in the BMC cluster, not only can the computational speed and efficiency of model inference be significantly improved, but also optimal resource allocation and utilization can be achieved, thereby enhancing the stability and scalability of the system.
[0122] In some embodiments, the first weight matrix is obtained by partitioning the weight matrix of the feedforward network layer of the large language model by each idle controller. The partitioning is performed by rows or columns, and the first weight matrix of each idle controller is different. This ensures that the computational load of each controller is roughly the same, avoiding situations where some controllers are overloaded while other controllers' resources are underutilized.
[0123] In some embodiments, the splitting unit includes a second allocation module for allocating each pending subtask to a corresponding idle controller, so that the idle controller processes the pending subtask using the multi-head attention layer of the large language model to obtain an attention output matrix, and each idle controller continues to perform calculations based on the attention output matrix to obtain processed subtasks. This highly parallel computing method may further increase the speed of model inference by eliminating the need to redistribute the integrated data to all nodes.
[0124] In some embodiments, the second allocation module includes a second allocation submodule for allocating each pending subtask to a corresponding idle controller, so that the idle controller processes the pending subtask using the multi-head attention layer of the large language model to obtain an attention output matrix, and each idle controller calculates the product of the attention output matrix and a second weight matrix to obtain a processed subtask, wherein the second weight matrix is obtained based on the weight matrix of the feedforward network layer of the large language model. This simplifies model management and task scheduling in a distributed environment, with each controller independently processing subtasks and local weight matrix multiplication, reducing reliance on centralized scheduling and resource management.
[0125] In some embodiments, the second weight matrix is obtained by partitioning the weight matrix of the feedforward network layer of the large language model by columns in the idle controller. Each idle controller has a different second weight matrix. This avoids local overload of computing resources and improves the overall stability and reliability of the system.
[0126] In some embodiments, each controller in the controller cluster uses an intelligent platform management interface communication protocol or an extensible service interface communication protocol for communication connection. The device also includes a first sending module and a second sending module. The first sending module is used to encapsulate the data to be sent into intelligent platform management interface information and send it to the target controller when each controller in the controller cluster uses the intelligent platform management interface communication protocol for communication connection. The intelligent platform management interface information includes the identification of the target controller, the data type, data length and data content of the data to be sent; the second sending module is used to encapsulate the data to be sent into target format data when each controller in the controller cluster uses the extensible service interface communication protocol for communication connection. The target format data is sent to the target controller by request. The target format is a format recognizable by the extensible service interface communication protocol. This ensures that data transmission between nodes can be stable and efficient.
[0127] In some embodiments, the splitting unit includes a third acquisition module and a second splitting module. The third acquisition module is used to obtain the performance parameters of the idle controller, which represent the maximum amount of tasks that the idle controller can handle at the current moment. The second splitting module is used to split the task to be processed based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller, the weight matrix of the feedforward network layer of the large language model, and the performance parameters of the idle controller, to obtain multiple subtasks to be processed, and the size of each subtask to be processed is positively correlated with the performance parameters of the idle controller. This ensures that each controller receives a subtask of an appropriate size based on its actual performance parameters, avoids the problem of some controllers being overloaded while other controller resources are idle, and achieves balanced load distribution.
[0128] For the description of the features in the embodiment corresponding to the controller, please refer to the relevant description of the embodiment corresponding to the distributed processing method of tasks, which will not be repeated here.
[0129] The embodiment of the present application also provides a controller cluster, such as Figure 5 As shown, it includes: a controller, which is used to perform the following steps: obtaining tasks to be processed; determining an idle controller in a controller cluster, the controller cluster includes multiple baseboard management controllers, and the idle controller is a baseboard management controller whose task volume processed at the current moment is less than the preset task volume; splitting the task to be processed based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple sub-tasks to be processed, and allocating each sub-task to be processed to the corresponding idle controller, so that the idle controller processes the sub-task to be processed to obtain a processed sub-task; integrating all the processed sub-tasks to obtain a processed task; at least one idle controller, the idle controller is a baseboard management controller whose task volume processed at the current moment is less than the preset task volume; at least one non-idle controller, the non-idle controller is a baseboard management controller whose task volume processed at the current moment is greater than or equal to the preset task volume, and the controller, the idle controller and the non-idle controller are communicatively connected.
[0130] This controller cluster fully leverages the computing resources of the BMC cluster and implements a tensor parallel strategy to achieve efficient distributed inference of large language models within the BMCs. This improves the performance of the question-answering assistant while reducing hardware costs and deployment complexity. Specifically, this paper focuses on how to achieve efficient tensor parallel inference between BMCs using the Intelligent Platform Management Interface (IPMI) or Redfish protocol.
[0131] An embodiment of the present application also provides an electronic device comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a distributed processing method for performing any one of the tasks.
[0132] The memory stores a computer program, and the processor is configured to run the computer program to perform the steps in the embodiment of the distributed processing method of any one of the above tasks.
[0133] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of the distributed processing method embodiment of any of the above tasks when run.
[0134] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0135] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the distributed processing method embodiment of any of the above tasks are implemented.
[0136] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in the distributed processing method embodiment of any of the above tasks.
[0137] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0138] The above is a detailed introduction to the distributed processing method of a task provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A distributed task processing method, characterized in that: include: Get pending tasks; Determine an idle controller in a controller cluster, the controller cluster including a plurality of baseboard management controllers, the idle controller being a baseboard management controller whose task load processed at a current moment is less than a preset task load; Splitting the task to be processed based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple subtasks to be processed; Allocating each of the to-be-processed subtasks to the corresponding idle controller, so that the idle controller processes the to-be-processed subtask using the multi-head attention layer of the large language model to obtain an attention output matrix; Integrate all the attention output matrices to obtain a total output matrix; Distributing the total output matrix to the plurality of idle controllers so that the idle controllers calculate a product of the total output matrix and a first weight matrix to obtain the processed subtasks, wherein the first weight matrix is obtained according to a weight matrix of a feedforward network layer of the large language model; Integrating all of the processed subtasks to obtain a processed task; The controllers in the controller cluster are connected in communication using an intelligent platform management interface communication protocol or an extensible service interface communication protocol.
2. The distributed task processing method according to claim 1, characterized in that: The tasks to be processed are split based on the number of attention heads of the multi-head attention layer and the weight matrix of the feedforward network layer of the large language model in the baseboard management controller, including: Converting the format of the task to be processed into a vector format to obtain a task vector; Based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer, the task vector is split to obtain a plurality of subtasks to be processed. The large language model is a converter model running in the idle controller. The large language models in each idle controller are the same, and the size of each subtask to be processed is the same.
3. The distributed task processing method according to claim 1, characterized in that: Allocating each of the pending subtasks to the corresponding idle controller, so that the idle controller processes the pending subtask using the multi-head attention layer of the large language model to obtain an attention output matrix, including: Assigning each of the subtasks to be processed to the corresponding idle controller, so that the idle controller uses multiple target attention heads of the multi-head attention layer of the large language model to extract different features of the subtask to be processed, obtain multiple task features, and enable the idle controller to obtain the attention output matrix based on the multiple task features, Among them, one target attention head is used to extract one task feature, multiple target attention heads are part of the attention heads of the multi-head attention layer, and different idle controllers use different target attention heads.
4. The distributed task processing method according to claim 3, characterized in that: The method further comprises: Obtain the number of idle controllers; Get the total number of attention heads in the multi-head attention layer of the large language model; The ratio of the total number of the attention heads to the number of the idle controllers is determined as the target number of attention heads for each of the idle controllers.
5. The distributed task processing method according to claim 1, characterized in that: The first weight matrix is obtained by each idle controller dividing the weight matrix of the feedforward network layer of the large language model, and the division processing is row division or column division, and the first weight matrix of each idle controller is different.
6. The distributed task processing method according to claim 1, characterized in that: Allocating each of the pending subtasks to the corresponding idle controller, so that the idle controller processes the pending subtask using the multi-head attention layer of the large language model to obtain an attention output matrix, including: Assign each of the subtasks to be processed to the corresponding idle controller, so that the idle controller uses the multi-head attention layer of the large language model to process the subtask to be processed, obtain the attention output matrix, and enable each idle controller to continue calculating according to the attention output matrix to obtain the processed subtask.
7. The distributed task processing method according to claim 6, characterized in that: Allocating each of the to-be-processed subtasks to the corresponding idle controller, so that the idle controller uses the multi-head attention layer of the large language model to process the to-be-processed subtask to obtain an attention output matrix, and causing each of the idle controllers to continue calculating according to the attention output matrix to obtain the processed subtask, including: Assigning each of the to-be-processed subtasks to the corresponding idle controller, so that the idle controller uses the multi-head attention layer of the large language model to process the to-be-processed subtask to obtain an attention output matrix, and each of the idle controllers calculates the product of the attention output matrix and a second weight matrix to obtain the processed subtask, The second weight matrix is obtained according to the weight matrix of the feedforward network layer of the large language model.
8. The distributed task processing method according to claim 7, characterized in that: The second weight matrix is obtained by the idle controller dividing the weight matrix of the feedforward network layer of the large language model by columns, and the second weight matrix of each idle controller is different.
9. The distributed task processing method according to claim 1, characterized in that: The method further comprises: When each controller in the controller cluster is communicatively connected using the intelligent platform management interface communication protocol, encapsulating the data to be sent into intelligent platform management interface information and sending it to the target controller, wherein the intelligent platform management interface information includes an identifier of the target controller, a data type, a data length, and data content of the data to be sent; When each controller in the controller cluster adopts the extensible service interface communication protocol for communication connection, the data to be sent is encapsulated into target format data, and the target format data is sent to the target controller through a request. The target format is a format recognizable by the extensible service interface communication protocol.
10. The distributed task processing method according to claim 1, characterized in that: The task to be processed is split based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller and the weight matrix of the feedforward network layer to obtain multiple subtasks to be processed, including: Acquire a performance parameter of the idle controller, where the performance parameter represents a maximum amount of tasks that the idle controller can process at a current moment; Based on the number of attention heads of the multi-head attention layer of the large language model in the baseboard management controller, the weight matrix of the feedforward network layer of the large language model and the performance parameters of the idle controller, the task to be processed is split to obtain a plurality of subtasks to be processed, and the size of each subtask to be processed is positively correlated with the performance parameters of the idle controller.
11. A controller, characterized in that: include: An acquisition unit, used to acquire tasks to be processed; a determining unit, configured to determine an idle controller in a controller cluster, the controller cluster including a plurality of baseboard management controllers, the idle controller being a baseboard management controller whose task load processed at a current moment is less than a preset task load; a splitting unit, configured to split the task to be processed based on the number of attention heads of the multi-head attention layer and the weight matrix of the feedforward network layer of the large language model in the baseboard management controller to obtain a plurality of subtasks to be processed; The splitting unit also includes a first allocation module, a first integration module and a distribution module. The first allocation module is used to allocate each of the to-be-processed subtasks to the corresponding idle controller, so that the idle controller uses the multi-head attention layer of the large language model to process the to-be-processed subtask to obtain an attention output matrix; The first integration module is used to integrate all the attention output matrices to obtain a total output matrix; The distribution module is used to distribute the total output matrix to the plurality of idle controllers, so that the idle controllers calculate the product of the total output matrix and a first weight matrix to obtain the processed subtask, wherein the first weight matrix is obtained according to the weight matrix of the feedforward network layer of the large language model; an integration unit, configured to integrate all of the processed subtasks to obtain a processed task; The controllers in the controller cluster are connected in communication using an intelligent platform management interface communication protocol or an extensible service interface communication protocol.
12. A controller cluster, characterized in that: include: The controller is used to perform the following steps: obtain the pending tasks; Determine an idle controller in the controller cluster, where the controller cluster includes a plurality of baseboard management controllers, and the idle controller is a baseboard management controller whose task volume processed at a current moment is less than a preset task volume; Splitting the task to be processed based on the number of attention heads of the multi-head attention layer of the large language model and the weight matrix of the feedforward network layer in the baseboard management controller to obtain multiple subtasks to be processed, and assigning each of the subtasks to be processed to the corresponding idle controller, so that the idle controller uses the multi-head attention layer of the large language model to process the subtask to be processed, and obtain an attention output matrix; Integrate all the attention output matrices to obtain a total output matrix; Distributing the total output matrix to a plurality of idle controllers so that the idle controllers calculate the product of the total output matrix and a first weight matrix to obtain the processed subtask, wherein the first weight matrix is obtained according to the weight matrix of the feedforward network layer of the large language model; integrating all the processed subtasks to obtain a processed task; and each of the controllers in the controller cluster is communicatively connected using an intelligent platform management interface communication protocol or an extensible service interface communication protocol; At least one idle controller, wherein the idle controller is a baseboard management controller whose task volume processed at the current moment is less than a preset task volume; At least one non-idle controller, where the non-idle controller is a baseboard management controller whose task volume processed at a current moment is greater than or equal to a preset task volume, and the controller, the idle controller and the non-idle controller are communicatively connected.
13. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a distributed processing method for performing the tasks described in any one of claims 1 to 10.