Task information distribution method, server, equipment, medium and program product
By deploying expert models in memory access nodes and optimizing transmission paths, the delay problem caused by cross-node communication of hybrid expert models in the server is solved, and efficient data computing and resource utilization are achieved.
Patent Information
- Application Number
- CN202510758524.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-23
AI Technical Summary
In the existing technology, the hybrid expert model in server task processing results in high latency overhead and memory access delay due to frequent cross-node communication, which reduces system throughput and computing efficiency.
By deploying a specified number of expert models in memory access nodes and establishing a mapping relationship between model identifiers and node identifiers in the routing controller, the target memory access node where the target expert model is located is determined, the task transmission path is optimized, cross-node communication is reduced, and data computing efficiency is improved.
It significantly reduces the communication overhead of multi-node data processing, improves data computing efficiency and resource utilization, and reduces waiting time during task execution.
Smart Images

Figure CN120687211A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method, server, device, medium, and program product for allocating task information. Background Art
[0002] In recent years, with the large-scale development of deep learning models, hybrid expert models (HEMs) have been widely used for server computing tasks due to their high capacity and high efficiency. Existing technologies, when using HEMs for server task processing, can employ a strategy of distributing the same experts across nodes for data computation.
[0003] However, since the server's computing tasks are distributed to different nodes for parallel processing, there may be frequent cross-node communication between multiple nodes, which leads to high latency overhead. In addition, some data may need to be obtained remotely from nodes of other devices, increasing memory access latency, resulting in a decrease in system throughput and low computing efficiency. Summary of the Invention
[0004] The present application provides a task information distribution method, server, device, medium and program product, which can significantly reduce the communication overhead when multiple nodes perform data processing and improve data computing efficiency.
[0005] In a first aspect, the present application provides a method for allocating task information, which is applied to a server, wherein the server includes multiple artificial intelligence processors, each of which has one or more memory access nodes, and each of the memory access nodes is deployed with one or more complete expert models. The method includes: receiving task information to be processed; parsing the task information and determining a target expert model for processing the target data obtained by the parsing; querying the target memory access node where the target expert model is located, and determining a task transmission path from the current memory access node to the target memory access node; and sending the target data to the server where the target memory access node is located based on the task transmission path, so that the target data is processed by the target expert model deployed in the target memory access node.
[0006] In this embodiment, the task information to be processed is parsed by the memory access node, and when the target expert model needs to be called for data processing, the target data is sent to the target memory access node where the target expert model is located for processing. Generally speaking, the target expert model will be split and deployed in multiple memory access nodes, and the target data will be sent to all memory access nodes to execute different computing tasks of the target expert model, which will result in frequent cross-node communication. In this embodiment, by deploying a specified number of expert models in the memory access node, the target data obtained by parsing the task information is accurately sent to the memory access node where the target expert model is located, so that the task can be processed in a targeted manner in one memory access node without sending the target data to all memory access nodes, significantly reducing the communication delay across nodes, reducing the communication overhead when multiple nodes perform data processing, and improving data calculation efficiency. At the same time, by determining an effective task transmission path, the rapid transmission of target data between memory access nodes is ensured, further improving data calculation efficiency. The technical solution provided by this embodiment can significantly reduce the communication requirements across nodes by optimizing the allocation strategy of expert models and task information in memory access nodes, thereby reducing the communication overhead when multiple nodes perform data processing and improving data calculation efficiency.
[0007] In one embodiment, the artificial intelligence processor includes one or more processing cores, wherein each of the processing cores corresponds to its own memory access node.
[0008] In this embodiment, each memory access node in the artificial intelligence processor corresponds to its own processing core and memory, so that each processing core can access the data related to it more quickly and directly, so that the target data can be executed on the processing core of the current memory access node, reducing the waiting time caused by data transmission during task execution and further improving data calculation efficiency.
[0009] In one embodiment, the server also includes a routing controller, which is communicatively connected to each of the artificial intelligence processors in the server; after one or more complete expert models are deployed in the memory access node, the mapping relationship between the model identifier of the expert model and the node identifier of the memory access node is written into the routing controller; querying the target memory access node where the target expert model is located includes: identifying the target model identifier of the target expert model, and sending a positioning request carrying the target model identifier to the routing controller; receiving the target node identifier fed back by the routing controller in response to the positioning request, and determining the memory access node with the target node identifier as the target memory access node where the target expert model is located.
[0010] In this embodiment, by establishing a mapping relationship between the model identifier and the node identifier, after determining the target expert model to which the target data belongs, the target memory access node where the target expert model is located can be directly determined through the above mapping relationship. When transmitting the target data, the correct computing resource can be quickly located. While reducing the communication requirements between nodes, the routing controller ensures the correct transmission of target data between different nodes, thereby improving data processing efficiency.
[0011] In one embodiment, the routing controller also stores topology information of each memory access node; determining the task transmission path from the current memory access node to the target memory access node includes: sending a data transmission request to the routing controller, the data transmission request including the first node identifier of the current memory access node and the second node identifier of the target memory access node; planning the lowest delay path from the first node identifier to the second node identifier through the topology information stored in the routing controller, and determining the lowest delay path as the task transmission path from the current memory access node to the target memory access node.
[0012] In this embodiment, when determining the task transmission path for target data, the lowest latency path can be planned from the current memory access node to the target memory access node based on the topology information in the routing controller. The current memory access node and the target memory access node can be represented by node identifiers, and the routing controller retrieves the topology information of the corresponding nodes through the node identifiers. This topology-based path planning can significantly reduce the delay in the target data transmission process, ensuring the minimum delay in the task transmission path, thereby improving data processing efficiency.
[0013] In one embodiment, the one or more complete expert models are deployed in the memory access node in the following manner: obtaining an expert model set, and determining a first expert model to be deployed in a first memory access node and a second expert model to be deployed in a second memory access node from the expert model set; allocating the deployment weight of the first expert model in the first memory access node, and allocating the deployment weight of the second expert model in the second memory access node, so as to deploy the complete first expert model in the first memory access node and the complete second expert model in the second memory access node.
[0014] In this embodiment, by allocating the deployment weights of each expert model to the same memory access node, the centralized deployment of the expert model is achieved, avoiding the scattered deployment of the expert model among multiple nodes, so that each memory access node can focus on processing a specific expert model task, thereby reducing the need for cross-node communication of target data, improving the utilization efficiency of computing resources, and accelerating the processing speed of target data, thereby improving the processing efficiency of task data.
[0015] In one embodiment, the method further includes: when deploying the corresponding expert model in any memory access node, identifying the parameter quantity of the corresponding expert model, and applying for a static memory space matching the parameter quantity for the corresponding expert model in the corresponding memory access node.
[0016] In this implementation, static memory space matching the expert model's parameters is allocated to each expert model, ensuring full utilization of the memory resources of the memory access nodes and improving system resource utilization efficiency. Furthermore, by allocating dedicated static memory space to each expert model, the expert model can access and process data more efficiently during operation, reducing the need for remote memory access, helping to reduce memory access latency and improve data processing efficiency.
[0017] In one embodiment, the method further includes: if the corresponding expert model receives data to be processed, applying for dynamic memory space in the corresponding memory access node according to the amount of the data to be processed, and releasing the dynamic memory space after processing the data.
[0018] In this implementation, by allocating memory on demand and releasing it promptly, dynamic memory space can be reused by other tasks or expert models, effectively utilizing limited memory space and improving memory resource utilization. Furthermore, this dynamic memory allocation mechanism enables expert models to flexibly process target data of varying sizes, enhancing their adaptability to diverse computing tasks and ultimately improving data processing efficiency.
[0019] In one embodiment, the method further includes: after completing the allocation of the current task information, for the new task information to be allocated, if the information correlation between the new task information and the current task information meets the preset conditions, the new task information is allocated to the server where the target memory access node is located.
[0020] In this implementation, new tasks with a high degree of correlation are assigned to servers located on the same memory access node. Because the data and status of related tasks may be cached in the memory of that server, assigning tasks to the same server based on their degree of correlation reduces context overhead during task switching. New tasks do not need to be reloaded and initialized, thereby speeding up task startup and execution, and thus improving data processing efficiency. When processing complex task scenarios, this task information allocation strategy can also reduce the need for communication between nodes, further improving data processing efficiency.
[0021] In one embodiment, the memory access node also includes a non-expert model and a gating network, wherein the non-expert model is used to parse the task information to obtain the target data; and the gating network is used to determine a target expert model for processing the target data.
[0022] In this embodiment, the non-expert model is used to perform preliminary analysis of the task information of the memory access node, so as to accurately extract the target data and provide clear and accurate input data for subsequent expert model processing. At the same time, the gating network is used to determine the target expert model suitable for processing the current target data based on the parsed task information, so as to transmit the target data to the corresponding target expert model, avoiding the complexity of large-scale data transmission caused by transmitting the target data to all expert models, thereby avoiding additional data processing and access delays and improving data processing efficiency.
[0023] On the other hand, the present application also provides a server, which includes multiple artificial intelligence processors, each of which has one or more memory access nodes, and each of which has one or more complete expert models deployed in the memory access node. The server includes: an information receiving unit for receiving task information to be processed; an expert model determination unit for parsing the task information and determining a target expert model for processing the parsed target data; a path determination unit for querying the target memory access node where the target expert model is located, and determining the task transmission path from the current memory access node to the target memory access node; and a data sending unit for sending the target data to the server where the target memory access node is located based on the task transmission path, so that the target data can be processed by the target expert model deployed in the target memory access node.
[0024] On the other hand, the present application also provides an electronic device, including: a memory storing computer instructions; and at least one processor configured to execute the computer instructions in the memory to perform the method for allocating task information in the above-mentioned embodiment.
[0025] On the other hand, the present application further provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the method for allocating task information in the above-mentioned embodiment.
[0026] On the other hand, the present application further provides a computer program product, including computer instructions, which, when executed by a processor, enable the processor to execute the method for allocating task information in the above embodiment. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0028] Figure 1 A schematic diagram of the distribution of task information in related technologies; Figure 2 A schematic diagram of the steps of a method for allocating task information provided in one embodiment of the present application; Figure 3 A schematic diagram of the structure of an artificial intelligence processor provided in one embodiment of the present application; Figure 4 A schematic diagram of the structure of a server provided in one embodiment of the present application; Figure 5 A schematic diagram of the deployment steps of an expert model provided in one embodiment of the present application; Figure 6 A schematic diagram of the distribution of task information in a specific application example; Figure 7 A schematic diagram of the functional modules of a server provided in one embodiment of the present application; Figure 8 A schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0030] In addition, the descriptions of "first", "second", etc. in this application are for descriptive purposes only and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more. In addition, the use of "based on" or "according to" means openness and inclusiveness, because the process, steps, calculations or other actions "based on" or "according to" one or more of the conditions or values can be based on additional conditions or values beyond the described values in practice.
[0031] See also Figure 1 In the related art, when a hybrid expert model is used to process server tasks, a strategy of allocating the same expert across nodes is usually used for data calculation. Specifically, the server's artificial intelligence processor may include several memory access nodes, and different parts of each expert model are scattered across multiple memory access nodes, and each memory access node only stores a portion of the expert model's weights or parameters. When the expert model needs to perform data calculations, different task information of the same expert model will be allocated to different memory access nodes for processing based on weight division. For any task information, since the expert model corresponding to the computing task is partially configured in multiple memory access nodes, the computing task needs to be sent to multiple memory access nodes. For example Figure 1 As shown in Figure 1, Compute Tasks 1, 2, and 3 are all sent to each memory access node. Furthermore, some expert model data may be stored on the local server, while others may be stored on nodes on other servers. When a local node is performing expert model calculations, the locally stored data may not be sufficient to complete the entire calculation task, requiring it to retrieve the missing data from the model data stored on other remote nodes.
[0032] In practical applications, this cross-node allocation method requires data exchange between memory access nodes to complete computational tasks. This extensive cross-node communication creates significant communication latency, and remote data access further exacerbates memory access latency, resulting in low computational efficiency. Furthermore, since the computational tasks for the same expert model are distributed across different nodes, each node's processing cores load data for computation, resulting in fragmented matrix operations. For smaller computational tasks, the processing cores take a short time to reach a stable operating state after loading data, making it difficult for the cores to operate at full capacity, further reducing the computational efficiency of the task.
[0033] In view of this, an embodiment of the present application provides a method for allocating task information, which can enable target data to be processed in a single memory access node containing a target expert model without sending the target data to all memory access nodes, thereby significantly reducing cross-node communication delays and reducing communication overhead when multiple nodes perform data processing, thereby improving data processing efficiency.
[0034] The task information allocation method provided in this application can be applied to an artificial intelligence processor with task allocation function, specifically any one of GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural Network Processing Unit), DPU (Deep Learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose Graphics Processing Unit).
[0035] See also Figure 2 In one embodiment of the present application, a method for allocating task information is provided, which can be applied to a server. The server includes multiple artificial intelligence processors, each of which has one or more memory access nodes, and each of the memory access nodes is deployed with one or more complete expert models. The method can include the following steps: S1: Receive task information to be processed.
[0036] S3: Parse the task information and determine a target expert model for processing the target data obtained by the parsing.
[0037] S5: Query the target memory access node where the target expert model is located, and determine a task transmission path from the current memory access node to the target memory access node.
[0038] S7: Sending the target data to the server where the target memory access node is located based on the task transmission path, so that the target data can be processed by the target expert model deployed in the target memory access node.
[0039] In this embodiment, the artificial intelligence processor of the server receives the task information to be processed, and sends the task information to be processed to the corresponding memory access node in the artificial intelligence processor for data processing. The information type of the above-mentioned task information is not limited, and the computing tasks represented by it may come from different sources and have different characteristics and processing requirements. For example, in a deep learning model, it may be different types of task information such as image data, text data or voice data. The above-mentioned task information may contain detailed information such as the original data of the task, task identification, processing requirements, etc., and needs to be assigned to a specific expert model for processing. As the system continues to process task information, new task information may continue to be added to the artificial intelligence processor. By judging whether the task information has been sent to the memory access node, it is determined that the unprocessed task information is the task information to be processed, and the task information is sent to a specific memory access node for parsing and processing of the task information.
[0040] In this embodiment, the memory access node performs preliminary parsing of the task information and preprocessing of the data. In order to ensure that the task information can be effectively processed and assigned to the appropriate target expert model, the memory access node needs to parse the task information to extract the target data in the task information, such as voice signals, image pixels, text content, etc., analyze the task requirements to determine the target of data processing, and determine the expert model corresponding to the data processing target. Among them, after parsing, it may be necessary to preprocess the target data, including data cleaning, normalization or other data processing steps, and use the preprocessed data as the target data to provide standardized input for subsequent expert models, such as data cleaning, feature extraction, format conversion, etc.
[0041] For example, in a speech recognition task, a voice file is received as task information, and the voice signal in the voice file is first parsed, and the voice signal is subjected to data preprocessing operations such as removing background noise and calculating MFCC (Mel Frequency Cepstral Coefficients) coefficients. The MFCC feature tensors obtained after data preprocessing are assigned to different expert models as target data. Among them, different expert models are good at processing different target data, such as English speech recognition, Chinese speech recognition or dialect recognition. For example, in an image recognition task, preprocessing operations such as image resizing and contrast enhancement are required. It should be noted that these operations do not involve complex model structures, but only change the basic form of the data and do not need to be processed in the expert model.
[0042] In this embodiment, the target data is sent to the target memory access node where the corresponding target expert model is located, and the target expert model performs data processing. Since each expert model is deployed on a specific memory access node, preferably, by establishing a mapping relationship between the expert model and the memory access node, the target memory access node where the target expert model is located can be quickly queried, and the task transmission path from the memory access node where the target data is currently located to the target access node can be determined. The target data is sent to the target expert model of the target memory access node through the task transmission path, without having to send the target data to all expert models for data processing. This reduces the length of data transmission, avoids excessive communication delays and resource consumption, and thus improves data processing efficiency.
[0043] In this embodiment, the current memory access node and the target access node where the target data is located may be on different servers. In this case, the target data needs to be sent to the server where the target memory access node is located via a task transmission path, and then transmitted to the target expert model corresponding to the target memory access node in the server via the task transmission path for data processing. The above-mentioned task transmission path is preferably the transmission path with the smallest data delay.
[0044] It can be seen that in this embodiment, the task information to be processed is parsed through the memory access node, and when the target expert model needs to be called for data processing, the target data is sent to the target memory access node where the target expert model is located for processing. By deploying a specified number of expert models in the memory access node, the target data obtained by parsing the task information is accurately sent to the memory access node where the target expert model is located, so that the task can be processed in a targeted manner in a memory access node without sending the target data to all memory access nodes, which significantly reduces the communication delay across nodes, reduces the communication overhead when multiple nodes perform data processing, and improves data calculation efficiency. At the same time, by determining an effective task transmission path, the rapid transmission of target data between memory access nodes is ensured, further improving data calculation efficiency.
[0045] In one embodiment, see Figure 3, the artificial intelligence processor includes one or more processing cores, and each of the processing cores corresponds to its own memory access node. The above-mentioned processing core can be understood as the basic unit responsible for graphics processing and computing tasks. The above-mentioned memory access node may include one or more memories, a processing core and a communication interface for managing and transmitting target data to ensure that the target data and the data in the local memory can be used correctly by the processing core. The above-mentioned memory is used to store data and task information. The size and speed of the memory directly affect the processing capacity of the node. The above-mentioned communication interface is used to communicate with other memory access nodes. Among them, the task processing of each memory access node is separated, and the task processing of each memory access node is performed by the processing core. At the same time, a communication interface is provided for data transmission to ensure the orderliness and accuracy of data processing performed by each memory access node. Among them, when the processing core performs a computing task, it needs to read data and instructions from the memory and write the results back to the memory to complete complex graphics processing and computing tasks.
[0046] In this embodiment, each memory access node in the artificial intelligence processor is deployed with its own processing core and memory, so that each processing core can access the data related to it more quickly and directly, so that the target data can be executed on the processing core of the current memory access node, reducing the waiting time caused by data transmission during task execution and further improving data calculation efficiency.
[0047] In one embodiment, see Figure 4 The server also includes a routing controller, which is in communication with each artificial intelligence processor in the server. By establishing a mapping relationship between the node identifier of each memory access node and the model identifier of the corresponding one or more complete expert models in the routing controller, the location of the target memory access node where the target data needs to be transmitted can be quickly located. Specifically, when querying the target memory access node where the target expert model is located, the target model identifier of the target expert model is first identified, and a positioning request carrying the target model identifier is sent to the routing controller. At this time, the routing controller will determine and feedback the target node identifier corresponding to the target model identifier carried by the positioning request based on the mapping relationship between the node identifiers of each memory access node and the model identifiers of the respective expert models. Further, the target node identifier fed back by the routing controller for the positioning request is received, and the memory access node with the target node identifier is determined as the target memory access node where the target expert model is located. The above-mentioned target model identifier can be implemented by extending the identification bit of the positioning request.
[0048] In this embodiment, by establishing a mapping relationship between the model identifier and the node identifier, after determining the target expert model to which the target data belongs, the target memory access node where the target expert model is located can be directly determined through the above mapping relationship. When transmitting the target data, the correct computing resource can be quickly located. While reducing the communication requirements between nodes, the routing controller ensures the correct transmission of target data between different nodes, thereby improving data processing efficiency.
[0049] In one embodiment, after determining the target memory access node for the data to be transmitted, the routing controller dynamically determines the optimal task transmission path to transmit the target data to the target memory access node. The routing controller stores the topology information of each memory access node. The topology information can describe the network structure, connection method, delay and bandwidth of the communication link between the memory access nodes, and characterize the relationship between the nodes and the communication characteristics. In order to determine the task transmission path of the target data, a data transmission request is sent to the routing controller, and the first node identifier of the current memory access node where the target data carried by the data transmission request is located and the second node identifier of the target memory access node are determined. This request mechanism enables the system to clearly define the starting point and focus of the target data transmission and provide identification information for path planning. Furthermore, the routing controller evaluates the delay and bandwidth of each path between the first node identifier and the second node identifier based on the topology information, and plans the lowest delay path between the first node identifier and the second node identifier as the task transmission path. The target data is transmitted to the target memory access node along the task transmission path, thereby ensuring that the data can reach the target memory access node quickly and efficiently.
[0050] In this embodiment, when determining the task transmission path for target data, the lowest latency path can be planned from the current memory access node to the target memory access node based on the topology information in the routing controller. The current memory access node and the target memory access node can be represented by node identifiers, and the routing controller retrieves the topology information of the corresponding nodes through the node identifiers. This topology-based path planning can significantly reduce the delay in the target data transmission process, ensuring the minimum delay in the task transmission path, thereby improving data processing efficiency.
[0051] See also Figure 5 In one embodiment, the one or more complete expert models can be Figure 5 The approach shown is deployed in a memory access node.
[0052] S51: Acquire an expert model set, and determine, from the expert model set, a first expert model to be deployed in a first memory access node and a second expert model to be deployed in a second memory access node.
[0053] S53: Allocate the deployment weight of the first expert model in the first memory access node, and allocate the deployment weight of the second expert model in the second memory access node, so as to deploy the complete first expert model in the first memory access node and deploy the complete second expert model in the second memory access node.
[0054] Specifically, for any expert model to be deployed, the deployment weights of the expert model are all distributed on the same memory access node, rather than being dispersed across various memory access nodes according to the deployment weights. This expert model deployment method concentrates the relevant resources and processing capabilities of the expert model on one node, avoiding a large amount of data transmission between individual node data, thereby reducing the overhead of cross-node communication and achieving the effect of improving data processing efficiency. Among them, when determining the expert model to be deployed from the expert model set, the expert model can be dynamically allocated according to the task requirements and the memory resources of each node to ensure efficient use of memory resources. It should be noted that the above-mentioned expert model set includes one or more expert models to be deployed, and the above-mentioned "first" and "second" descriptions are only for descriptive purposes and are not limited to the expert model set in this application having only two expert models to be deployed. Those skilled in the art can deduce a solution for deploying multiple expert models on memory access nodes from the content disclosed in this specification.
[0055] In this embodiment, the centralized deployment of expert models is achieved by allocating the deployment weights of each expert model to the same memory access node. In the traditional expert model deployment strategy, for the first expert model or the second expert model, each expert model is divided into multiple partial models according to the deployment weight and deployed in multiple memory access nodes. In order to speed up data processing, the target data is processed in parallel, resulting in the need for the target data to be sent to multiple memory access nodes of the target expert model. There is a large amount of cross-node communication overhead, which also brings about a large access delay and low data processing efficiency. In order to avoid the communication delay caused by the decentralized deployment of expert models among multiple nodes, the weights and resources of the expert models are completely allocated to the memory access nodes, so that each memory access node is focused on processing specific expert model tasks, thereby reducing the need for cross-node communication of target data, improving the utilization efficiency of computing resources, and speeding up the processing speed of target data, thereby improving the processing efficiency of task data.
[0056] In one embodiment, when a corresponding expert model is deployed on any memory access node, a static memory space of the memory is allocated for the parameter amount of the expert model. Specifically, each expert model has its own specific parameter amount, and these parameter amounts determine the size of the space required for the model in the memory. Since each expert model is deployed in a corresponding memory access node, it may be necessary to repeatedly call the parameter amount of the expert model when the expert model is executed at the memory access node. At this time, a static memory space that matches the size of the parameter amount is allocated to the parameter amount of the expert model. While efficiently utilizing the memory space, it can ensure that the parameter amount will not be automatically released or reallocated when stored in the memory, avoiding the loss of parameters, thereby ensuring the stability of the expert model deployment and execution.
[0057] It's important to note that while static memory space is allocated for expert models, this does not limit the system's scalability. When adding new expert models or adjusting the parameters of existing models, the static memory space can be reallocated based on the new requirements, or memory space can be reallocated to more needed expert models. This allows the system to flexibly adapt to different task requirements and scale changes, thereby improving system resource utilization efficiency.
[0058] In this implementation, static memory space matching the expert model's parameters is allocated to each expert model, ensuring full utilization of the memory resources of the memory access nodes and improving system resource utilization efficiency. Furthermore, by allocating dedicated static memory space to each expert model, the expert model can access and process data more efficiently during operation, reducing the need for remote memory access, helping to reduce memory access latency and improve data processing efficiency.
[0059] In one embodiment, when a memory access node receives target data, dynamic memory space of the memory is allocated for the target data. In addition to static memory space, the memory configured by the memory access node also includes dynamic memory space. Since the target data does not need to be stored for a long time and repeatedly processed, allocating the target data to the dynamic memory space can further improve the utilization rate of memory resources. Specifically, when the expert model receives the data to be processed, it applies for dynamic memory space that matches the size of the data in the memory access node according to the amount of data to be processed. When the expert model processes the data to be processed, it repeatedly retrieves and stores data from the dynamic memory space. After the data processing is completed, there is no need to store the processed data anymore, and the dynamic memory space is released at this time.
[0060] In this implementation, by allocating memory on demand and releasing it promptly, dynamic memory space can be reused by other tasks or expert models, effectively utilizing limited memory space and improving memory resource utilization. Furthermore, this dynamic memory allocation mechanism enables expert models to flexibly process target data of varying sizes, enhancing their adaptability to diverse computing tasks and ultimately improving data processing efficiency.
[0061] In one embodiment, after completing the allocation of the current task information, if new task information to be allocated is received, the information correlation between the new task information and the current task information is identified, and if the information correlation meets the preset conditions, the new task information is allocated to the server where the target memory access node is located. The above-mentioned correlation may be evaluated based on factors such as task type, data characteristics, and processing requirements. The above-mentioned preset conditions may be a threshold of the correlation. Whether to allocate the new task information to the server where the target memory access node is located is determined by judging whether the correlation of the task information reaches the threshold, thereby ensuring that the new task can utilize the data and resources that have been loaded by the current task. Furthermore, according to the information correlation of the task information, allocating the task information to the same memory access node can reduce the overhead of task switching and improve data processing efficiency.
[0062] For example, in an image processing system, the current task information is to process edge detection of a group of images. If new task information is received at this time to process color enhancement of the same group of images, the system recognizes that both task information involves the same group of images and both belong to image processing tasks. At this time, it is determined that the new task information has a high degree of information correlation with the current task information and meets the preset conditions. The new task information is then assigned to the target memory access node of the current task information. During execution, the new task information can utilize the image data that has been loaded by the current task information, thereby reducing data transmission overhead and improving data processing efficiency.
[0063] In this implementation, new tasks with a high degree of correlation are assigned to servers located on the same memory access node. Because the data and status of related tasks may be cached in the memory of that server, assigning tasks to the same server based on their degree of correlation reduces context overhead during task switching. New tasks do not need to be reloaded and initialized, thereby speeding up task startup and execution, and thus improving data processing efficiency. When processing complex task scenarios, this task information allocation strategy can also reduce the need for communication between nodes, further improving data processing efficiency.
[0064] In one embodiment, the memory access node also includes a non-expert model, which is used to parse task information to obtain target data. Specifically, each memory access node is deployed with a non-expert model. When feature extraction is required during task information parsing, the non-expert model can be used to extract features. For example, when a memory access node receives a speech signal, the non-expert model performs preliminary processing on the speech signal, for example, extracting basic features such as Mel-frequency cepstral coefficients as target data.
[0065] In this embodiment, the memory access node also includes a gating network that is used to determine the target expert model for processing the target data. For example, after extracting the basic features of the Mel-frequency cepstral coefficients as the target data, the gating network determines the speech signal's language, spoken language, or speech rate based on the extracted features, and then sends the target data to the expert model that specializes in that feature.
[0066] In this implementation, a non-expert model performs a preliminary analysis of task information from memory access nodes, accurately extracting the target data and providing clear and accurate input data for subsequent expert model processing. Furthermore, a gating network determines the target expert model suitable for processing the current target data based on the parsed task information, and then transmits the target data to the corresponding target expert model. This avoids the complexity of transferring large amounts of data to all expert models, thus minimizing additional data processing and access delays and improving data processing efficiency.
[0067] For a specific application example, see Figure 6 Each expert model is completely assigned to a memory access node. The memory access node is configured with an expert model, a non-expert model, and a gating network. A memory access node can be configured with multiple expert models. Memory access nodes may be configured in artificial intelligence processors on different servers. Multiple task information to be processed is obtained from the received information set and evenly sent to each memory access model. The task information is parsed by the non-expert model of each memory access node to obtain target data with a specific computing task, and the target data with a specific computing task is sent to the corresponding expert model. If the current memory access unit is not the target memory access node for the target data, the target data needs to be transferred across nodes to the target memory access node.
[0068] For example, Figure 6As shown, task information 1 is sent to memory access node 1, and after being parsed by the non-expert model, the target data is obtained. The gating network determines that the target expert model for the target data is expert model 1. The target data is then directly processed by expert model 1 in memory access node 1. However, task information 2 is parsed by the non-expert model of memory access node 2, and the target expert model determined by the gating network is expert model 2. Expert model 2 is deployed in memory access node 1. In this case, memory access node 2 determines the lowest latency path from memory access node 2 to memory access node 1 as the task transmission path and transmits the target data to expert model 2 in memory access node 1 for data processing.
[0069] In view of this, compared to existing technologies, this solution avoids sending the target data obtained by parsing each task information to all memory access nodes, significantly reducing cross-node communication delays, lowering the communication overhead when processing data on multiple nodes, and improving data computation efficiency. At the same time, because the existing technology integrates the calculation results at each memory access node as the calculation results of each expert model, it further increases the amount of data transmission and reduces computational efficiency. However, the calculation results obtained by each expert model in this solution do not need to be sent to other memory access nodes for data integration, further reducing cross-node data transmission and thus improving data computation efficiency.
[0070] Furthermore, in the prior art, when each expert model in a memory access node processes target data, because each node's expert model is a partial model, it only processes a portion of the target data. When the expert model's weight distribution is small, each memory access node quickly completes its data processing task when loading the target data for processing. This results in a short period of stable operation for the memory access node, making it difficult for the memory access node's computing unit to operate at full load and resulting in low resource utilization. However, in this technical solution, since all the expert model weights are deployed on the same memory access node, the memory access node can fully process the entire target data, maintaining a longer stable operation period and achieving higher resource utilization.
[0071] See also Figure 7 On the other hand, the present application further provides a server, comprising a plurality of artificial intelligence processors, each of which has one or more memory access nodes, and each of which has one or more complete expert models deployed therein. The server comprises: The information receiving unit 100 is used to receive task information to be processed; The expert model determination unit 200 is used to parse the task information and determine a target expert model for processing the target data obtained by the parsing; A path determination unit 300 is configured to query a target memory access node where the target expert model is located, and determine a task transmission path from the current memory access node to the target memory access node; The data sending unit 400 is used to send the target data to the server where the target memory access node is located based on the task transmission path, so that the target data can be processed by the target expert model deployed in the target memory access node.
[0072] In one embodiment, the path determination unit 300 is specifically used to send a data transmission request to the routing controller, wherein the data transmission request includes at least the first node identifier of the current memory access node and the second node identifier of the target memory access node, and plans the lowest delay path from the first node identifier to the second node identifier through the topology information stored in the routing controller, and determines the lowest delay path as the task transmission path from the current memory access node to the target memory access node.
[0073] The further functional description of each of the above units is the same as that of the above corresponding embodiments and will not be repeated here.
[0074] On the other hand, the present application also provides an electronic device, including: a memory storing computer instructions; and at least one processor configured to execute the computer instructions in the memory to perform the method for allocating task information in the above-mentioned embodiment.
[0075] On the other hand, the present application further provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the method for allocating task information in the above-mentioned embodiment.
[0076] On the other hand, the present application further provides a computer program product, including computer instructions, which, when executed by a processor, enable the processor to execute the method for allocating task information in the above embodiment.
[0077] See also Figure 8 , Figure 8 is a structural diagram of an electronic device provided by an optional embodiment of the present invention, such as Figure 8As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 10 is taken as an example.
[0078] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0079] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0080] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0081] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0082] The electronic device further includes a communication interface 30 for the electronic device to communicate with other devices or a communication network.
[0083] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0084] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0085] The present application is described with reference to the flowcharts and / or block diagrams of the methods and systems according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0086] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0088] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0089] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0090] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
[0091] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
Claims
1. A method for allocating task information, characterized in that: The method is applied to a server, wherein the server includes multiple artificial intelligence processors, each of which has one or more memory access nodes, and each of the memory access nodes is deployed with one or more complete expert models. The method includes: Receive pending task information; parsing the task information and determining a target expert model for processing the parsed target data; Querying a target memory access node where the target expert model is located, and determining a task transmission path from the current memory access node to the target memory access node; The target data is sent to the server where the target memory access node is located based on the task transmission path, so that the target data is processed by the target expert model deployed in the target memory access node.
2. The method according to claim 1, characterized in that The artificial intelligence processor includes one or more processing cores, wherein each of the processing cores corresponds to its own memory access node.
3. The method according to claim 1 or 2, characterized in that The server further includes a routing controller, which is in communication with each of the artificial intelligence processors in the server; after one or more complete expert models are deployed in the memory access node, a mapping relationship between the model identifier of the expert model and the node identifier of the memory access node is written into the routing controller; Querying the target memory access node where the target expert model is located includes: Identifying a target model identifier of the target expert model, and sending a positioning request carrying the target model identifier to the routing controller; A target node identifier fed back by the routing controller in response to the positioning request is received, and a memory access node having the target node identifier is determined as the target memory access node where the target expert model is located.
4. The method according to claim 3, characterized in that The routing controller also stores topology information of each memory access node; Determining a task transmission path from the current memory access node to the target memory access node includes: Sending a data transmission request to the routing controller, wherein the data transmission request includes a first node identifier of the current memory access node and a second node identifier of the target memory access node; The lowest delay path from the first node identifier to the second node identifier is planned using the topology information stored in the routing controller, and the lowest delay path is determined as the task transmission path from the current memory access node to the target memory access node.
5. The method according to claim 1, wherein The one or more complete expert models are deployed in the memory access node in the following manner: Obtain an expert model set, and determine, from the expert model set, a first expert model to be deployed in a first memory access node and a second expert model to be deployed in a second memory access node; Allocate the deployment weight of the first expert model in the first memory access node, and allocate the deployment weight of the second expert model in the second memory access node, so as to deploy the complete first expert model in the first memory access node and the complete second expert model in the second memory access node.
6. The method according to claim 1 or 5, characterized in that The method further comprises: When deploying a corresponding expert model in any memory access node, the parameter quantity of the corresponding expert model is identified, and static memory space matching the parameter quantity is applied for the corresponding expert model in the corresponding memory access node.
7. The method according to claim 6, characterized in that The method further comprises: If the corresponding expert model receives data to be processed, it applies for dynamic memory space in the corresponding memory access node according to the amount of the data to be processed, and releases the dynamic memory space after processing the data.
8. The method according to claim 1, characterized in that The method further comprises: After completing the allocation of the current task information, for the new task information to be allocated, if the information correlation between the new task information and the current task information meets the preset conditions, the new task information is allocated to the server where the target memory access node is located.
9. The method according to claim 1, characterized in that The memory access node also includes a non-expert model and a gating network, wherein the non-expert model is used to parse the task information to obtain the target data; and the gating network is used to determine a target expert model for processing the target data.
10. A server, characterized in that: The server includes multiple artificial intelligence processors, each of which has one or more memory access nodes, and each of the memory access nodes is deployed with one or more complete expert models. The server includes: An information receiving unit, configured to receive task information to be processed; an expert model determination unit, configured to parse the task information and determine a target expert model for processing the target data obtained by the parsing; a path determination unit, configured to query a target memory access node where the target expert model is located, and determine a task transmission path from a current memory access node to the target memory access node; A data sending unit is used to send the target data to the server where the target memory access node is located based on the task transmission path, so that the target data can be processed by the target expert model deployed in the target memory access node.
11. An electronic device, characterized in that: include: a memory storing computer instructions; At least one processor is configured to execute the computer instructions in the memory to perform the method according to any one of claims 1-9.
12. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 9.
13. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 9.
Citation Information
Cited By
Computer system, node optimization method, electronic device, and storage medium
CN121217597A