Task processing method, device and equipment based on intelligent computing cluster and storage medium
By dividing processors with different operating frequencies in the intelligent computing cluster and processing pre-filling and decoding tasks, the problem of high energy consumption of artificial intelligence models on the intelligent computing cluster is solved, and the system energy efficiency optimization is achieved.
Patent Information
- Application Number
- CN202411755936.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-02
AI Technical Summary
When running artificial intelligence models on intelligent computing clusters, energy consumption is high, and existing technologies are difficult to save energy consumption at the root cause.
By dividing processors with different operating frequencies among multiple processors in the intelligent computing cluster, the first processor processes the pre-filling task at high frequency, the second processor processes the decoding task at low frequency, and shares data through the intermediate data temporary storage area to optimize energy efficiency.
Without affecting system performance, the energy consumption of the processor is effectively reduced and energy consumption is saved at the root cause.
Smart Images

Figure CN119938245A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of resource scheduling, and in particular to a task processing method, device, equipment and storage medium based on an intelligent computing cluster. Background Art
[0002] An intelligent computing cluster is a computing resource and architecture that combines high-performance computing capabilities and intelligent technologies (such as artificial intelligence, big data analysis, etc.). It aims to improve the utilization of computing resources and accelerate the execution of complex computing tasks through intelligent management and optimization strategies. However, the task of running artificial intelligence models on an intelligent computing cluster is the main source of energy consumption. Taking the GPT3 model as an example, its training process consumes approximately 1287MWh of electricity, which is equivalent to the energy consumed by a medium-sized data center running continuously for several months. Therefore, it is necessary to propose a method to save energy when running artificial intelligence models on an intelligent computing cluster.
[0003] In related technologies, in order to reduce the idle time of the processor and eliminate resource waste, multiple copies of the same artificial intelligence model can be deployed in the intelligent computing cluster to achieve parallel processing of tasks and minimize the waiting time of the processor to avoid energy waste. However, although this method alleviates the problem of energy waste to a certain extent, it cannot reduce the energy consumption of the processor when processing tasks, and cannot save energy at the root. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to propose a task processing method, device, equipment and storage medium based on an intelligent computing cluster, which can save energy consumption at the root.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present application proposes a task processing method based on an intelligent computing cluster, the method comprising:
[0006] Obtaining tasks to be processed and determining an intelligent computing cluster for processing the tasks to be processed;
[0007] According to a preset division rule, a plurality of first processors and a plurality of second processors are determined from a plurality of intelligent computing processors of the intelligent computing cluster, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency;
[0008] Pre-filling the to-be-processed tasks by the plurality of first processors corresponding to the first operating frequency to obtain intermediate features, and transmitting the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor;
[0009] When it is detected that the intermediate data temporary storage area is updated, the intermediate features are obtained from the intermediate data temporary storage area by the multiple second processors corresponding to the second operating frequency, and the intermediate features are decoded to obtain a decoding result of each second processor;
[0010] The decoding results of each second processor are outputted in sequence.
[0011] Accordingly, a second aspect of an embodiment of the present application proposes a task processing device based on an intelligent computing cluster, the device comprising:
[0012] An acquisition module is used to acquire tasks to be processed and determine an intelligent computing cluster for processing the tasks to be processed;
[0013] A determination module, configured to determine, according to a preset division rule, a plurality of first processors and a plurality of second processors from a plurality of intelligent computing processors of the intelligent computing cluster, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency;
[0014] a processing module, configured to perform pre-filling processing on the to-be-processed task by using the plurality of first processors corresponding to the first operating frequency to obtain intermediate features, and transmit the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor;
[0015] a decoding module, configured to, when detecting that the intermediate data temporary storage area is updated, obtain the intermediate feature from the intermediate data temporary storage area through the plurality of second processors corresponding to the second operating frequency, and perform decoding processing on the intermediate feature to obtain a decoding result of each second processor;
[0016] An output module is used to sequentially output the decoding results of each second processor.
[0017] In some implementations, the first processor and the second processor are located in the same cluster node; the determining module is further configured to:
[0018] Determine, from the intelligent computing cluster, a cluster node for processing the task to be processed;
[0019] Determine a first division ratio and a second division ratio for the multiple intelligent computing processors of the cluster node according to a preset division rule;
[0020] Determining a plurality of first processors from the plurality of intelligent computing processors according to the first division ratio;
[0021] According to the second division ratio, a plurality of second processors are determined from the plurality of intelligent computing processors, and the total number of the plurality of first processors and the plurality of second processors is equal to the total number of intelligent computing processors of the cluster node.
[0022] In some implementations, the task processing device based on the intelligent computing cluster further includes an adjustment module for:
[0023] In the process of processing the tasks to be processed, regularly obtaining the average delay of the first characters of the pre-filling processing of the tasks to be processed by the multiple first processors and the average delay of the second characters of the decoding processing of the tasks to be processed by the multiple second processors;
[0024] Compare the first character average delay with the second character average delay to obtain a first comparison result;
[0025] When the first comparison result indicates that the first character average delay is inconsistent with the second character average delay, the number of the plurality of first processors and the number of the plurality of second processors are adjusted according to the first comparison result.
[0026] In some implementations, the adjustment module is further used to:
[0027] When the first comparison result indicates that the first character average delay is greater than the second character average delay, obtaining a first difference between the first character average delay and the second character average delay;
[0028] Determine a first number of first processors to be added according to a first difference range corresponding to the first difference, and adjust the number of the plurality of first processors and the plurality of second processors according to the first number;
[0029] When the first comparison result indicates that the first character average delay is less than the second character average delay, obtaining a second difference between the second character average delay and the first character average delay;
[0030] A second number of added second processors is determined according to a second difference range corresponding to the second difference, and the number of the plurality of first processors and the plurality of second processors is adjusted according to the second number.
[0031] In some implementations, the adjustment module is further used to:
[0032] Obtain the sum of the first character average delay and the second character average delay to obtain a total delay;
[0033] Obtaining a preset delay threshold, and comparing the delay threshold with the total delay to obtain a second comparison result;
[0034] When the second comparison result indicates that the total delay is greater than the delay threshold, determine a third operating frequency corresponding to the second processor, and obtain the intermediate feature from the intermediate data temporary storage area through the multiple second processors corresponding to the third operating frequency, and decode the intermediate feature to obtain a decoding result of each second processor; wherein the third operating frequency is greater than the second operating frequency;
[0035] When the second comparison result indicates that the total delay is less than the delay threshold, the fourth operating frequency corresponding to the second processor is determined, and the intermediate features are obtained from the intermediate data temporary storage area through the multiple second processors corresponding to the fourth operating frequency, and the intermediate features are decoded to obtain the decoding results of each second processor; wherein the fourth operating frequency is less than the second operating frequency.
[0036] In some implementations, the task processing device based on the intelligent computing cluster further includes a prediction module, which is used to:
[0037] Setting the first operating frequencies corresponding to the plurality of first processors as the highest operating frequency;
[0038] Based on the highest operating frequency, predicting the total operating power consumption and total operating delay of the cluster nodes corresponding to the multiple first processors and the multiple second processors at different operating frequencies of the multiple second processors;
[0039] Drawing a power consumption curve based on the total power consumption of the operation, and drawing a performance curve based on the total delay of the operation;
[0040] Based on the power consumption curve and the performance curve, second operating frequencies of the plurality of second processors are determined.
[0041] In some implementations, the task processing device based on the intelligent computing cluster further includes a numbering module, which is used to:
[0042] Divide the to-be-processed task into a plurality of sub-processing tasks, and number the plurality of sub-processing tasks according to the task sequence;
[0043] In the intermediate data temporary storage area, the received intermediate features are stored according to the task number of each sub-processing task, and the intermediate features are sequentially acquired from the intermediate data temporary storage area according to the corresponding task numbers through the multiple second processors.
[0044] Correspondingly, the third aspect of the embodiments of the present application proposes a computer device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the task processing method based on the intelligent computing cluster described in any one of the embodiments of the first aspect of the present application.
[0045] Correspondingly, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the task processing method based on the intelligent computing cluster as described in any one of the embodiments of the first aspect of the present application.
[0046] The embodiment of the present application obtains the tasks to be processed and determines an intelligent computing cluster for processing the tasks to be processed; according to a preset division rule, determines multiple first processors and multiple second processors from multiple intelligent computing processors of the intelligent computing cluster, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency; pre-fills the tasks to be processed through the multiple first processors corresponding to the first operating frequency to obtain intermediate features, and transmits the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor; when an update of the intermediate data temporary storage area is detected, the intermediate features are obtained from the intermediate data temporary storage area through the multiple second processors corresponding to the second operating frequency, and the intermediate features are decoded to obtain a decoding result of each second processor; and the decoding result of each second processor is output in sequence. In this way, different intelligent computing processors can be divided for processing according to the sensitivity of the tasks to be processed to the operating frequency of the intelligent computing processor at different processing stages, wherein the first processor executes the pre-filling task and the second processor executes the decoding task, and the first operating frequency corresponding to the first processor is set to be greater than the second operating frequency corresponding to the second processor. In this way, the pre-filling task that is more sensitive to frequency can be processed by the first processor with a higher frequency to quickly complete data processing; the decoding task that is less sensitive to frequency can be processed by the second processor with a lower frequency, so as to reduce energy consumption as much as possible while maintaining high system performance, so as to effectively reduce the energy consumption of the processor without affecting performance, thereby achieving energy saving at the root. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the hardware architecture of the intelligent computing cluster provided in the embodiment of the present application;
[0048] Figure 2 is a flowchart of a task processing method based on an intelligent computing cluster provided in an embodiment of the present application;
[0049] Figure 3 It is a functional module diagram of a task processing device based on an intelligent computing cluster provided in an embodiment of the present application;
[0050] Figure 4 It is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0054] An intelligent computing cluster is a computing resource and architecture that combines high-performance computing capabilities and intelligent technologies (such as artificial intelligence, big data analysis, etc.). It aims to improve the utilization of computing resources and accelerate the execution of complex computing tasks through intelligent management and optimization strategies. However, the task of running artificial intelligence models on an intelligent computing cluster is the main source of energy consumption. Taking the GPT3 model as an example, its training process consumes approximately 1287MWh (megawatt-hours) of electricity, which is equivalent to the energy consumed by a medium-sized data center running continuously for several months. Therefore, it is necessary to propose a method to save energy when running artificial intelligence models on an intelligent computing cluster.
[0055] In related technologies, in order to reduce the idle time of the processor and eliminate resource waste, multiple copies of the same artificial intelligence model can be deployed in the intelligent computing cluster to achieve parallel processing of tasks and minimize the waiting time of the processor to avoid energy waste. However, although this method alleviates the problem of energy waste to a certain extent, it cannot reduce the energy consumption of the processor when processing tasks, and cannot save energy at the root.
[0056] Based on this, the embodiments of the present application provide a task processing method, device, equipment and storage medium based on an intelligent computing cluster, which can save energy consumption at the source.
[0057] The task processing method, apparatus, device and storage medium based on the intelligent computing cluster provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the hardware architecture of the intelligent computing cluster in the embodiments of the present application is described.
[0058] Please refer to Figure 1 In some implementations, the present application provides a task processing system based on an intelligent computing cluster. The intelligent computing cluster includes multiple intelligent computing processors, which can be used to perform reasoning tasks of large AI models or other complex tasks that require a large amount of computing resources.
[0059] Furthermore, the intelligent computing processor is a unit that performs specific computing tasks. The intelligent computing processor can be divided into a first processor specifically used to process pre-filling tasks, and a second processor specifically used to process decoding tasks. Among them, the first processor can be used for pre-filling tasks, that is, processing the pending tasks input by the user through the terminal and converting them into feature representations that can be understood by the AI large model. The first processor runs at a high frequency to ensure that computing-intensive tasks are completed quickly; the second processor can be used for decoding tasks, that is, generating outputs based on the intermediate feature representations obtained in the pre-filling stage. Since the decoding task is a memory-intensive task and is less sensitive to the processor frequency, the second processor can run at a lower frequency.
[0060] In some embodiments, the intelligent computing server may be installed on a computer device, and the computing device may be responsible for collecting and managing resource information of all intelligent computing clusters, including the load of cluster nodes, available resources, geographic location, allocation of intelligent computing processors contained in each cluster node, etc. The computer device may monitor the delay of pre-filling and decoding tasks, and timely adjust the number of intelligent computing processors used to process pre-filling and decoding tasks to optimize energy efficiency while ensuring performance.
[0061] The task processing method based on the intelligent computing cluster in the embodiment of the present application can be illustrated by the following embodiment.
[0062] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for enabling the normal operation of the embodiment of the present application will be obtained.
[0063] In the embodiment of the present application, the task processing device based on the intelligent computing cluster will be described from the perspective of the task processing device based on the intelligent computing cluster. The task processing device based on the intelligent computing cluster can be integrated into a computer device. Figure 2 , Figure 2 This is a flowchart of the steps of the task processing method based on the intelligent computing cluster provided in the embodiment of the present application. The embodiment of the present application takes the task processing device based on the intelligent computing cluster as an example, which is specifically integrated on a terminal or a server. When the processor on the terminal or the server executes the program instructions corresponding to the task processing method based on the intelligent computing cluster, the specific process is as follows:
[0064] Step 101: Obtain tasks to be processed and determine an intelligent computing cluster for processing the tasks to be processed.
[0065] In some embodiments, in order to understand the user's specific request and ensure efficient completion of the task, pending tasks can be obtained and an intelligent computing cluster used to process the pending tasks can be determined to determine the type and content of the task, ensure that the task is processed correctly, and improve user experience and service quality.
[0066] Among them, the pending tasks can be AI model reasoning tasks that need to be executed on the intelligent computing cluster, including but not limited to natural language processing, image recognition, speech recognition, etc. Specifically, the pending tasks can be requests initiated by users through terminals, such as translating a text, generating a reply, identifying an image, training models, etc.
[0067] Among them, the intelligent computing cluster can be a high-performance computing platform composed of multiple intelligent computing processors (such as AI processors), usually located in a data center or cloud environment, capable of handling large-scale data computing needs. The intelligent computing processor can be a graphics processing unit (GPU), a neural network processor (NPU), etc., which is used to efficiently perform complex AI tasks such as model training and reasoning.
[0068] For example, a user may initiate a request through a terminal (such as a smart phone, a personal computer, etc.), and the request includes specific content of the task to be processed, such as text to be translated, pictures to be recognized, etc. After the terminal sends the request including the task to be processed to the computer device through the network, the computer device may parse the content of the task to be processed, such as the task type, input data, and expected output format.
[0069] Furthermore, the computer device can evaluate the available resources in the intelligent computing cluster, and determine the corresponding cluster nodes and processors to process the pending tasks according to the results of the analysis of the pending tasks. Alternatively, the intelligent computing cluster with a smaller task load can be directly determined as the intelligent computing cluster for processing the pending tasks according to the task load of each intelligent computing cluster.
[0070] In order to reduce the idle time of the processor and eliminate resource waste, multiple copies of the same AI model can be deployed in the intelligent computing cluster to achieve parallel processing of tasks. For example, the same task to be processed can be processed in parallel by multiple intelligent computing processors to minimize the waiting time of the intelligent computing processors and avoid energy waste.
[0071] Through the above steps, it is possible to effectively obtain tasks to be processed and determine the intelligent computing cluster used to process the tasks, thereby ensuring orderly distribution of tasks.
[0072] Step 102: According to a preset division rule, determine multiple first processors and multiple second processors from multiple intelligent computing processors of the intelligent computing cluster, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency.
[0073] In some embodiments, in order to achieve energy efficiency optimization and performance improvement of task processing, different intelligent computing processors can be used to process different processing stages of the task to be processed according to preset division rules, thereby allocating intelligent computing processors with different operating frequencies according to task characteristics, so as to essentially save energy without affecting the performance of the intelligent computing cluster.
[0074] Among them, the preset division rule can be a rule for dividing the number of intelligent computing processors according to a preset standard or algorithm.
[0075] Among them, the intelligent computing processor can be the computing unit that constitutes the intelligent computing cluster. The intelligent computing processor has high-performance computing capabilities and is suitable for executing complex artificial intelligence tasks, including but not limited to graphics processors, neural network processors, etc.
[0076] The first processor may be a group of intelligent computing processors selected from the intelligent computing cluster according to a preset division rule, and is used to execute the tasks to be processed in the pre-filling stage. The first processor is configured to run at a first operating frequency to maximize its computing efficiency.
[0077] The second processor may be a group of intelligent computing processors selected from the intelligent computing cluster according to a preset division rule, and is used to execute the pending tasks in the decoding stage. The second processor is configured to run at a second operating frequency to optimize energy efficiency.
[0078] The first operating frequency may be an operating frequency allocated to the first processor, and may be an upper limit of a maximum frequency at which the first processor operates.
[0079] The second operating frequency may be an operating frequency allocated to the second processor, which is lower than the first operating frequency, for example, 75%, 60%, etc. of the first operating frequency.
[0080] In some embodiments, since the pre-filling process of the task to be processed involves a large number of computing tasks, such as matrix multiplication, convolution operation, etc., the computing power of the intelligent computing processor is required to be high, which is a computing-intensive task. Therefore, the first processor needs to run at a higher frequency (first operating frequency) to ensure computing efficiency. When decoding the intermediate features obtained by the pre-filling process, it mainly involves a large number of memory access tasks, and the computing power requirements of the intelligent computing processor are relatively low. Therefore, the second processor can run at a lower frequency (second operating frequency) to reduce power consumption.
[0081] Exemplarily, the preset division rule may be to divide the multiple intelligent computing clusters according to a preset ratio. For example, if the division ratio of the first processor and the second processor is 50%, when a cluster node has 8 intelligent computing processors, 4 of them may be set as the first processors and the other 4 as the second processors.
[0082] In some implementations, an intelligent scheduling algorithm may be used to analyze the task characteristics of the task to be processed, obtain the task characteristics analysis results, and adjust the number of the first processor and the second processor according to the task characteristics analysis results. For example, for a computationally intensive task, the number of the first processor is increased; for a memory access intensive task, the number of the second processor is increased.
[0083] For example, by using an intelligent scheduling algorithm to analyze the task characteristics of the tasks to be processed, the types of tasks to be processed can be classified, and a database containing typical characteristics of each type of task can be established, with each task type recording the corresponding number of computing operations and memory access times. Then, the first processor and the second processor of the current task to be processed are divided according to the division ratio of the first processor and the second processor of the task type recorded in the database.
[0084] In some embodiments, the number of calculation operations of the task to be processed in the pre-filling stage can also be predicted, and the number of memory accesses of the task to be processed in the decoding stage can be predicted. Afterwards, the ratio of the first processor to the second processor is calculated based on the number of calculation operations and the number of memory accesses. For example, if the number of calculation operations is 10,000 when the task to be processed is in the pre-filling stage, and the number of memory accesses is 5,000 when it is in the decoding stage, then the ratio of the first processor can be 10,000 / (10,000+5,000)≈2 / 3; the ratio of the second processor can be 5,000 / (10,000+5,000)≈1 / 3. Then, when a cluster node has 8 intelligent computing processors, 5 of them can be set as the first processor and the other 3 as the second processor. Furthermore, the prediction of the number of calculation operations of the task to be processed in the pre-filling stage can be predicted by a machine learning algorithm (such as a regression model, a neural network, a decision tree, etc.), or it can be predicted based on historical data, and the embodiment of the present application does not impose specific restrictions on this.
[0085] In some implementations, the division ratio of the first processor to the second processor may be manually set, and the multiple intelligent computing processors of the intelligent computing cluster may be divided accordingly, for example, the first processor and the second processor may be set to correspond to 4 to 5, and so on.
[0086] By running the pre-filling tasks and decoding tasks corresponding to different stages of the tasks to be processed on different intelligent computing processors and setting different operating frequencies for the two, energy consumption in the decoding stage can be effectively saved without affecting system performance. This is of great significance in large-scale cluster systems and is conducive to environmental sustainable development.
[0087] Step 103, pre-filling the tasks to be processed by using multiple first processors corresponding to the first operating frequency to obtain intermediate features, and transmitting the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor.
[0088] In some implementations, in order to improve the efficiency of task processing and the reliability of data, an intermediate buffer area may be set up to centrally manage intermediate features to ensure orderly transmission and integrity of data.
[0089] Among them, pre-filling processing of the task to be processed can be a process of processing input data (such as text, voice, image, etc.) into feature representations suitable for model understanding.
[0090] Among them, the intermediate features can be temporary results or feature representations generated when processing the tasks to be processed in the pre-filling stage. The intermediate features are the data form obtained after preliminary processing of the input data (such as text, speech, image, etc.), which are designed to be suitable for internal processing of the model to facilitate subsequent decoding tasks.
[0091] Among them, the intermediate data temporary storage area can be a storage area located inside the intelligent computing cluster node, which is used to temporarily store the intermediate features generated in the pre-filling stage to support efficient data exchange between the pre-filling task and the decoding task. Since the tasks to be processed are executed by the first processor and the second processor respectively at different stages, the intermediate data temporary storage area can be used to ensure that the intermediate features can be seamlessly transferred from the pre-filling task to the decoding task.
[0092] Exemplarily, in order to increase the computing speed and reduce the delay, the maximum frequency of each first processor can be used as the first operating frequency. Taking a large-scale natural language processing task as an example, the task to be processed includes 100 sentences, each sentence contains 100 words, and the first processor can convert the input text into a vector representation inside the model at a first operating frequency, for example, 1.5 GHz (gigahertz), and generate intermediate features, which are stored in the intermediate data temporary storage area inside the node.
[0093] For example, when the tasks to be processed are processed in the same intelligent computing cluster node, a high-speed cache (such as SRAM, DRAM) can be used as an intermediate data temporary storage area. The intermediate data temporary storage area allows multiple intelligent computing processors to access it simultaneously, providing low-latency data access and reducing the overhead of data copying.
[0094] Furthermore, the intermediate data temporary storage area can be mapped to the address space of each intelligent computing processor (including the first processor and the second processor), so that each intelligent computing processor can directly access the data in the shared intermediate data temporary storage area. Furthermore, multiple first processors can directly write the generated intermediate features into the intermediate data temporary storage area, and multiple second processors can directly read these intermediate features from the intermediate data temporary storage area through the high-speed interconnection inside the node.
[0095] In some embodiments, when the task to be processed needs to be processed across cluster nodes, the intermediate data temporary storage area is a storage area located between different cluster nodes, which is used to temporarily store intermediate features that need to be transmitted across nodes. At this time, a distributed storage system (such as a distributed file system, distributed cache) can be used as an intermediate data temporary storage area to provide high availability and scalability.
[0096] In some embodiments, in order to further improve data processing efficiency and energy efficiency, a multi-level cache strategy can be implemented in the intermediate data temporary storage area shared by the first processor and the second processor. Different levels of cache areas can correspond to storage media with different rates to reduce the number of accesses to the main memory while maintaining a fast data access speed. Exemplarily, in addition to dividing the area for storing intermediate features, three levels of cache areas can be divided in the intermediate data temporary storage area. The first level cache has the fastest access rate but the smallest capacity, and is used to store the data that needs to be accessed most frequently when performing pre-filling tasks, such as the feature representation of the recently processed input data. The second level cache is slower than the first level cache, but has a larger capacity, and is used to store data that may be reused in the pre-filling task, or data that overflows in the first level cache; the third level cache further reduces the access speed than the second level cache to provide a larger storage space for storing data that is accessed less frequently but occasionally required during the pre-filling process, or data that overflows in the second level cache. It should be noted that data consistency should be maintained between caches of different levels to avoid data conflicts. In some embodiments, the sizes of caches of different levels can be dynamically adjusted according to the load and characteristics of the pre-filling task.
[0097] It can be understood that by dividing the three-level cache area in the intermediate data temporary storage area, the data that may be needed can be predicted and loaded into the cache in advance according to the access pattern of the pre-filled task. Through these hierarchical cache strategies, the pre-filled processor's access requirements to the main memory can be effectively reduced, energy consumption can be reduced, and the speed and efficiency of data processing can be improved.
[0098] By using intermediate data temporary storage areas and high-speed interconnections to achieve efficient data transmission, it can be ensured that the intermediate features generated by the first processor can be quickly and efficiently transmitted to the second processor, thereby significantly improving task processing efficiency, data reliability and resource utilization, and optimizing the overall performance and energy efficiency of the system.
[0099] By processing the pre-filling task through the high-frequency first processor, data processing can be completed quickly, task delays can be reduced, and user experience can be improved.
[0100] Step 104, when it is detected that the intermediate data temporary storage area is updated, the intermediate features are obtained from the intermediate data temporary storage area through multiple second processors corresponding to the second operating frequency, and the intermediate features are decoded to obtain decoding results of each second processor.
[0101] In some embodiments, in order to ensure that the decoding processor can obtain the latest intermediate features in a timely manner and reduce the delay in data transmission, when an update of the intermediate data temporary storage area is detected, the second processor corresponding to the second operating frequency can obtain the intermediate features from the intermediate data temporary storage area, thereby reducing power consumption and improving the energy efficiency of the system.
[0102] Among them, decoding the intermediate features can be a process of generating output based on the intermediate feature representation obtained in the pre-filling stage.
[0103] The decoding result can be the final product of the model reasoning process, such as text, labels, classification results, predicted values, etc. The specific form depends on the nature and goal of the task to be processed. For example, in natural language processing tasks, the decoding result may be a translated sentence; in image recognition tasks, the decoding result may be an object label in the image.
[0104] Exemplarily, it can ensure that energy consumption is saved as much as possible without affecting system performance. A lower operating frequency can be set for each second processor, for example, the first operating frequency corresponding to the first processor is 1.5GHz, and the second operating frequency corresponding to the second processor is 1.0GHz. When the intermediate data temporary storage area is detected to be updated, each second processor can read the intermediate features from the intermediate data temporary storage area, perform decoding processing, and generate the final decoding result. Furthermore, each second processor can poll to obtain the corresponding intermediate features, and output the final decoding results in sequence according to the time sequence of the intermediate feature acquisition. For example, for a natural language processing model, the decoding process will be carried out word by word and step by step to generate the final text output.
[0105] Through the above method, it can be ensured that the computation-intensive pre-filling task runs at a high frequency (first frequency) to quickly complete the calculation; while the memory access-intensive decoding task runs at a lower frequency (second frequency) to reduce unnecessary energy consumption. In this way, not only the resource utilization efficiency is improved, but also the impact on the environment is reduced, providing an effective solution for sustainable intelligent computing.
[0106] Step 105, output the decoding result of each second processor in sequence.
[0107] In some embodiments, in order to ensure that the final product of the model reasoning process, such as the translated sentence or the object label in the image, can be output accurately, the decoding result of each second processor can be output sequentially to effectively manage and control the concurrent access of multiple second processors to the intermediate features, ensure the sequentiality of the decoding results, and thus avoid order confusion.
[0108] In some embodiments, when storing the intermediate features in the intermediate data temporary storage area, a unique task number can be assigned to each sub-processing task, and the intermediate features are stored in the intermediate data temporary storage area in the order of the task numbers. After the second processor obtains the intermediate features from the intermediate data temporary storage area in the order of the task numbers for decoding, it outputs the decoding results in sequence according to the task numbers. In this way, it can be ensured that even if multiple second processors work at the same time, they can process the tasks in a predetermined order, thereby ensuring the order of the decoding results.
[0109] In some implementations, a polling queue may be established to queue the intermediate features in the order in which the tasks are submitted. Each second processor sequentially takes out the intermediate features from the polling queue for decoding, thereby ensuring the order of the decoding results.
[0110] In some implementations, a mutex or other synchronization mechanism may be used to control access to the intermediate data temporary storage area, ensuring that only one second processor can obtain intermediate features from the intermediate data temporary storage area at a time, thereby preventing multiple second processors from accessing the intermediate data temporary storage area at the same time, thereby avoiding order confusion.
[0111] In some embodiments, a timestamp or version number may be marked on the intermediate features to indicate the order in which they are processed, and the second processor outputs the decoding results according to the order of the timestamp or version number, so as to ensure that the decoding results can be output in the correct order even in a concurrent environment.
[0112] The embodiment of the present application obtains the tasks to be processed and determines an intelligent computing cluster for processing the tasks to be processed; according to a preset division rule, determines multiple first processors and multiple second processors from multiple intelligent computing processors of the intelligent computing cluster, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency; pre-fills the tasks to be processed through the multiple first processors corresponding to the first operating frequency to obtain intermediate features, and transmits the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor; when an update of the intermediate data temporary storage area is detected, the intermediate features are obtained from the intermediate data temporary storage area through the multiple second processors corresponding to the second operating frequency, and the intermediate features are decoded to obtain a decoding result of each second processor; and the decoding result of each second processor is output in sequence. In this way, different intelligent computing processors can be divided for processing according to the sensitivity of the tasks to be processed to the operating frequency of the intelligent computing processor at different processing stages, wherein the first processor executes the pre-filling task and the second processor executes the decoding task, and the first operating frequency corresponding to the first processor is set to be greater than the second operating frequency corresponding to the second processor. In this way, the pre-filling task that is more sensitive to frequency can be processed by the first processor with a higher frequency to quickly complete data processing; the decoding task that is less sensitive to frequency can be processed by the second processor with a lower frequency, so as to reduce energy consumption as much as possible while maintaining high system performance, so as to effectively reduce the energy consumption of the processor without affecting performance, thereby achieving energy saving at the root.
[0113] In some implementations, it can be determined whether to process the task across cluster nodes based on the size and bandwidth of the task to be processed. Under the condition that bandwidth permits, for large-scale distributed computing tasks (such as big data analysis, machine learning training, etc.), high-performance computing tasks (such as climate simulation, physical simulation, etc.), tasks with strong task dependencies, tasks that require load balancing, etc., while ensuring that cross-cluster node processing can bring performance improvements, the computing resources and storage resources of multiple cluster nodes can be used to improve processing efficiency. That is, the pre-filling stage tasks of the task to be processed are placed on the first cluster node for processing, and the decoding stage tasks of the same task to be processed are placed on the second cluster node for processing, and so on.
[0114] In some implementations, in order to reduce the data transmission delay of the intermediate features from the pre-filling stage to the decoding stage, the pending tasks can be processed in the same cluster node to ensure fast data access and transmission and maintain the efficiency of the overall processing flow. For example, step 102 may include:
[0115] (102.1) Determine a cluster node for processing the task to be processed from the intelligent computing cluster;
[0116] (102.2) determining a first division ratio and a second division ratio for the plurality of intelligent computing processors of the cluster nodes according to a preset division rule;
[0117] (102.3) Determine a plurality of first processors from the plurality of intelligent computing processors according to the first division ratio;
[0118] (102.4) According to the second division ratio, a plurality of second processors are determined from the plurality of intelligent computing processors, and the total number of the plurality of first processors and the plurality of second processors is equal to the total number of intelligent computing processors of the cluster node.
[0119] Among them, a cluster node can be a physical or logical unit in an intelligent computing cluster, which can include multiple intelligent computing processors. Each cluster node is responsible for performing specific computing tasks and can run independently or work in collaboration with other nodes.
[0120] The first division ratio may be a ratio allocated to the first processor when determining the allocation of intelligent computing processors within the cluster node.
[0121] The second division ratio may be a ratio allocated to the second processor when determining the allocation of intelligent computing processors within the cluster node.
[0122] Exemplarily, a cluster node in an idle state may be selected from the intelligent computing cluster to process the task to be processed, or a cluster node with a smaller load may be selected to process the task to be processed.
[0123] Exemplarily, the preset partitioning rule may correspond to a fixed partitioning ratio, for example, setting a first partitioning ratio corresponding to the first processor to 50%, setting a second partitioning ratio corresponding to the second processor to 50%, etc. At this time, when there are 10 intelligent computing processors in the cluster node that are idle, 5 first processors are allocated and 5 second processors are allocated.
[0124] For example, by using an intelligent scheduling algorithm to analyze the task characteristics of the tasks to be processed, the types of tasks to be processed can be classified, and a database containing typical characteristics of each type of task can be established, with each task type recording the corresponding number of computing operations and memory access times. Then, the first processor and the second processor of the current task to be processed are divided according to the division ratio of the first processor and the second processor of the task type recorded in the database.
[0125] In some embodiments, the number of calculation operations of the task to be processed in the pre-filling stage can also be predicted, and the number of memory accesses of the task to be processed in the decoding stage can be predicted. Afterwards, the ratio of the first processor to the second processor is calculated based on the number of calculation operations and the number of memory accesses. For example, if the number of calculation operations is 10,000 when the task to be processed is in the pre-filling stage, and the number of memory accesses is 5,000 when it is in the decoding stage, then the ratio of the first processor can be 10,000 / (10,000+5,000)≈2 / 3; the ratio of the second processor can be 5,000 / (10,000+5,000)≈1 / 3. Then, when a cluster node has 8 intelligent computing processors, 5 of them can be set as the first processor and the other 3 as the second processor. Furthermore, the prediction of the number of calculation operations of the task to be processed in the pre-filling stage can be predicted by a machine learning algorithm (such as a regression model, a neural network, a decision tree, etc.), or it can be predicted based on historical data, and the embodiment of the present application does not impose specific restrictions on this.
[0126] Exemplarily, when the cluster node is in an idle state, the total number of multiple first processors and multiple second processors is equal to the total number of intelligent computing processors of the cluster node; when the cluster node is in a non-idle state, the total number of multiple first processors and multiple second processors is equal to the total number of intelligent computing processors of the cluster node in the idle state.
[0127] In some embodiments, the task type (for example, natural language processing, image processing, speech recognition, etc., which may affect computational intensity and memory access intensity), task scale (input data size of the task, for example, text length, image resolution, etc.), and task complexity (the number of operations required, such as the number of matrix multiplications, the number of convolution operations, etc.) of the task to be processed may be obtained, and clustering may be performed in historical data based on the task type, task scale, and task complexity, and the Euclidean distance between the task type of the task to be processed and the historical task type, the task scale and the historical task scale, and the task complexity and the historical task complexity may be calculated and summed according to their respective weights to find the number of divisions of the first processor and the second processor of the k historical tasks with the smallest Euclidean distance, and the number of divisions of the first processor of the k historical tasks may be averaged to determine a plurality of first processors for processing the task to be processed; and the number of divisions of the second processor of the k historical tasks may be averaged to determine a plurality of second processors for processing the task to be processed.
[0128] By deploying the first processor and the second processor on the same cluster node and dynamically determining their ratio and number according to preset division rules, it is possible to optimize resource allocation, reduce data transmission delay, and reduce energy consumption. At the same time, it improves processing efficiency and the overall performance of the system, ensuring that the intelligent computing cluster is both efficient and energy-saving when processing various computing tasks.
[0129] In some implementations, in order to dynamically balance the workload of the first processor and the second processor, ensure the efficiency matching of the pre-filling and decoding processing stages, and avoid resource waste or bottlenecks caused by uneven processing speeds, by comparing the average delays of different processors and adjusting the number of processors accordingly, ensure that the system can adaptively respond to workload changes, thereby optimizing the efficiency and response time of the overall task processing process. For example, the task processing method based on the intelligent computing cluster may also include:
[0130] (A.1) in the process of processing the task to be processed, regularly obtaining the average delay of the first character of the pre-filling processing of the task to be processed by the first processors and the average delay of the second character of the decoding processing of the task to be processed by the second processors;
[0131] (A.2) comparing the average delay of the first character with the average delay of the second character to obtain a first comparison result;
[0132] (A.3) When the first comparison result indicates that the average delay of the first character is inconsistent with the average delay of the second character, the number of the plurality of first processors and the number of the plurality of second processors are adjusted according to the first comparison result.
[0133] The average delay of the first character may be the average time delay of multiple first processors processing each character during the pre-filling process of the task to be processed. The average delay of the first character may be calculated by counting the total time and the total number of characters processed by all first processors within a certain time interval. The average delay of the first character reflects the processing speed and efficiency of the pre-filling task.
[0134] The second character average delay may be the average time delay for multiple second processors to process each character during the decoding process of the task to be processed. The second character average delay may also be calculated by counting the total time and total number of characters processed by all second processors within a certain time interval. The second character average delay reflects the processing speed and efficiency of the decoding task.
[0135] The first comparison result may be a result obtained by comparing the average delay of the first character with the average delay of the second character. The purpose of the comparison is to evaluate whether the processing speeds of the pre-filling task and the decoding task match. If there is a significant difference between the average delay of the first character and the average delay of the second character, it means that the processing speeds of the two tasks are inconsistent, and the number of processors may need to be adjusted to optimize the overall performance.
[0136] Exemplarily, the first character delay may be determined by the ratio of the sum of the processing time of each first processor processing the pre-filling task within a fixed time period to the total number of characters processed by all the pre-filling tasks.
[0137] Specifically, the calculation formula for the first character delay is as follows:
[0138]
[0139] Among them, T1 i is the first character delay of the i-th first processor in processing the pre-filling task, n is the total number of first processors, and c is the total number of characters processed by all pre-filling tasks.
[0140] The calculation formula for the second character delay is as follows:
[0141]
[0142] Among them, T2 i is the second character delay of the i-th second processor processing the pre-filling task, m is the total number of first processors, and d is the total number of characters processed by all decoding tasks.
[0143] For example, if the average delay of the first character is 10 milliseconds, and the average delay of the second character is 5 milliseconds, the average delay of the first character is compared with the average delay of the second character, and the first comparison result can indicate that the average delay of the first character is inconsistent with the average delay of the second character, and the average delay of multiple first processors processing pre-filling processing is much greater than the average delay of multiple first processors processing decoding tasks, indicating that the workload of the first processor is heavy. In order to balance the load, the number of first processors can be increased, such as increasing the number of first processors from 5 to 6, to reduce the average delay of each first processor, and at the same time reduce the number of second processors, such as from 5 to 4, because the current second processor load is light.
[0144] In some implementations, a prediction model can be introduced to perform a large amount of training based on historical tasks to be processed, predict the delay trend of the tasks to be processed, and adjust the number of the first processor and the second processor in advance to cope with possible load changes and reduce the delay caused by real-time adjustment.
[0145] Through the above methods, resource allocation can be dynamically adjusted according to the actual processing situation to ensure the efficient operation of the intelligent computing cluster, reduce unnecessary waiting time, and improve the overall processing speed and resource utilization.
[0146] In some implementations, in order to dynamically balance the workload of the first processor and the second processor and ensure the efficiency and response speed of the entire task processing flow, the average delay of the first character and the average delay of the second character can be compared, and the number of processors can be adjusted according to the difference, which can effectively reduce processing bottlenecks, avoid resource waste, and improve the throughput and performance of the overall system. For example, "adjusting the number of the plurality of first processors and the plurality of second processors according to the first comparison result" in (A.3) may include:
[0147] (A.3.1) When the first comparison result indicates that the average delay of the first character is greater than the average delay of the second character, obtaining a first difference between the average delay of the first character and the average delay of the second character;
[0148] (A.3.2) determining a first quantity of first processors to be added according to a first difference range corresponding to the first difference, and adjusting the quantity of the plurality of first processors and the plurality of second processors according to the first quantity;
[0149] (A.3.3) When the first comparison result indicates that the average delay of the first character is less than the average delay of the second character, obtaining a second difference between the average delay of the second character and the average delay of the first character;
[0150] (A.3.4) Determine a second number of additional second processors based on a second difference range corresponding to the second difference, and adjust the number of the plurality of first processors and the plurality of second processors based on the second number.
[0151] The first difference may be the difference between the first character average delay and the second character average delay when the first character average delay is greater than the second character average delay. The calculation formula is first difference=first character average delay-second character average delay.
[0152] The first difference range may be a series of predefined threshold intervals for classifying the first difference so as to determine the number of first processors to be added. Each first difference range corresponds to a specific first number. For example, first difference range 1: 50 < first difference ≤ 100 ms; first difference range 2: 100 ms < first difference ≤ 300 ms; first difference range 3: 300 ms < first difference ≤ 150 ms.
[0153] The first number may be the number of first processors that need to be added determined according to a range of the first difference.
[0154] The second difference may be the difference between the second character average delay and the first character average delay when the first character average delay is less than the second character average delay. The calculation formula is second difference = second character average delay - first character average delay.
[0155] The second difference range may be a series of predefined threshold intervals for classifying the second difference so as to determine the number of second processors to be added. Each second difference range corresponds to two specific second quantities. For example, second difference range 1: 50 < second difference ≤ 100ms; second difference range 2: 100ms < second difference ≤ 300ms; second difference range 3: 300ms < second difference ≤ 150ms.
[0156] The second number may be the number of second processors that need to be added determined according to a range of the second difference.
[0157] For example, if the first difference is within the range of 50 < first difference ≤ 100ms, one first processor is added. If the first difference is within the range of 100ms < first difference ≤ 300ms, two first processors are added. If the first difference is within the range of 300ms < first difference, three first processors are added. It is understandable that
[0158] Exemplarily, if the second difference is within the range of 100<second difference≤200ms, one second processor is added. If the second difference is within the range of 200ms<second difference≤300ms, two second processors are added. If the second difference is within the range of 300ms<second difference, three second processors are added.
[0159] In some embodiments, the first difference range, the first quantity, the second difference range, and the second quantity can all be set according to actual conditions. For example, the first difference range can also be in the range of 5<first difference ≤10ms, 10<first difference ≤20ms, 20<first difference ≤50ms, etc. This application does not make any specific limitations on this.
[0160] For example, since the number of intelligent computing processors in the cluster node is fixed, when the number of first processors increases, the number of second processors will decrease by the same amount. For example, the original number of first processors is 5 and the number of second processors is 5. After adding 1 first processor, the number of second processors will decrease by 1. Similarly, when the number of second processors increases, the number of first processors will decrease by the same amount.
[0161] Dynamically adjusting the number of the first processor and the second processor can help achieve optimal resource allocation, improve system processing capabilities, ensure efficient and stable task processing, and reduce operating costs.
[0162] In some implementations, in order to flexibly adjust resource usage according to actual load and performance requirements while ensuring performance, the operating frequency of the second processor may be dynamically adjusted according to the actual performance of the system and preset goals to achieve an optimal balance between performance and energy consumption. For example, after step (A.3), the following may also be included:
[0163] (B.1) Obtain the sum of the average delay of the first character and the average delay of the second character to obtain the total delay;
[0164] (B.2) obtaining a preset delay threshold, comparing the delay threshold with the total delay, and obtaining a second comparison result;
[0165] (B.3) When the second comparison result indicates that the total delay is greater than the delay threshold, determine a third operating frequency corresponding to the second processor, and obtain intermediate features from the intermediate data temporary storage area through multiple second processors corresponding to the third operating frequency, and decode the intermediate features to obtain a decoding result of each second processor; wherein the third operating frequency is greater than the second operating frequency;
[0166] (B.4) When the second comparison result indicates that the total delay is less than the delay threshold, determine the fourth operating frequency corresponding to the second processor, and obtain intermediate features from the intermediate data temporary storage area through multiple second processors corresponding to the fourth operating frequency, and decode the intermediate features to obtain a decoding result of each second processor; wherein the fourth operating frequency is less than the second operating frequency.
[0167] The total delay may be the sum of the average delay of the first character and the average delay of the second character, and the calculation formula is: total delay=average delay of the first character+average delay of the second character.
[0168] The delay threshold may be a preset delay upper limit, which is used to evaluate whether the total delay is within an acceptable range. The delay threshold may be set according to actual application scenarios and performance requirements.
[0169] The second comparison result may be a comparison result of the total delay and the delay threshold, which is used to determine whether the total delay exceeds the allowed range. If the total delay is greater than the delay threshold, it indicates that the current processing speed is slow and the response time needs to be optimized. If the total delay is less than the delay threshold, it indicates that the current processing speed is fast and the energy efficiency can be appropriately reduced to save resources.
[0170] Among them, the third operating frequency can be a higher operating frequency set for the second processor when the total delay is greater than the delay threshold. By increasing the operating frequency, the processing speed of the decoding task can be accelerated to a certain extent, thereby reducing the total delay.
[0171] The fourth operating frequency may be a lower operating frequency set for the second processor when the total delay is less than the delay threshold. By lowering the operating frequency, energy can be saved and energy efficiency can be improved.
[0172] In some embodiments, after the number of the first processor and the second processor is adjusted, the average delay difference between the pre-filling task and the decoding task is within the allowable range. At this time, when the total delay exceeds the delay threshold, it indicates that the current processing speed of the task to be processed is not fast enough. Since the operating frequency of the second processor can speed up the decoding speed to a certain extent, the operating frequency of the second processor can be increased to speed up the decoding speed to a certain extent and reduce the delay; when the total delay is lower than the threshold, it indicates that the processing speed is fast, and the operating frequency of the second processor can be appropriately reduced to reduce energy consumption, so as to achieve a balance between performance and energy consumption.
[0173] Exemplarily, when the sum of the average delay of the first character and the average delay of the second character is 70 milliseconds and the delay threshold is 50 milliseconds, the total delay > the delay threshold. At this time, the operating frequency of each second processor can be increased. For example, if the initial second operating frequency is 1.0 GHz, the third operating frequency can be adjusted to 1.2 GHz.
[0174] Exemplarily, when the sum of the average delay of the first character and the average delay of the second character is 30 milliseconds and the delay threshold is 50 milliseconds, the total delay is less than the delay threshold. At this time, the operating frequency of each second processor can be reduced. For example, if the initial second operating frequency is 1.0 GHz, the fourth operating frequency can be adjusted to 0.8 GHz.
[0175] In some implementations, the amplitude of adjusting the operating frequency of the second processor may be a set granularity, such as adjusting 10% each time, or may be a dynamic calculation result based on a prediction model, and the embodiments of the present application do not impose specific limitations on this.
[0176] When the processing speed is insufficient, the decoding speed is accelerated by increasing the frequency of the second processor, reducing latency and ensuring the timeliness of task processing; when the processing speed is fast, the frequency is reduced to reduce energy consumption and achieve energy efficiency. In this way, not only the system's adaptability to different loads is improved, but also an intelligent balance between performance and energy consumption is achieved, thereby improving the overall system efficiency and responsiveness.
[0177] In some implementations, in order to optimize the energy efficiency ratio of the processors in the intelligent computing cluster, the maximum operating frequency of the first processor can be set to predict the total power consumption and total delay of the system of the second processor at different operating frequencies, and the power consumption and performance curves can be drawn accordingly to determine the optimal operating frequency of the second processor, so as to achieve the best balance between performance and energy consumption while ensuring processing performance. Exemplarily, the second operating frequency can be determined by:
[0178] (C.1) setting the first operating frequencies corresponding to the plurality of first processors as the highest operating frequency;
[0179] (C.2) predicting the total operating power consumption and total operating delay of the cluster nodes corresponding to the plurality of first processors and the plurality of second processors at different operating frequencies of the plurality of second processors based on the highest operating frequency;
[0180] (C.3) Draw a power consumption curve based on the total power consumption of the operation, and draw a performance curve based on the total delay of the operation;
[0181] (C.4) Determine second operating frequencies of the plurality of second processors based on the power consumption curve and the performance curve.
[0182] The maximum operating frequency may be the maximum operating frequency used by the first processor when executing the pre-filling task, and the maximum operating frequency is set in order to maximize computing performance.
[0183] Among them, the total operating power consumption can be the total electrical energy consumed by all intelligent computing processors in the cluster node when executing tasks at different operating frequencies of the second processor. The total operating power consumption can be obtained through prediction and can be expressed in watts (W).
[0184] The total operation delay may be the time required for the cluster node to process the task to be processed until the task is completely completed under different operating frequencies of the second processor, that is, the processing time including the pre-filling task and the decoding task, and may be expressed in milliseconds (ms).
[0185] Among them, the power consumption curve can be a trend chart of the total operating power consumption of the cluster nodes changing with the operating frequency under different operating frequencies of the second processor. By drawing the power consumption curve, the power consumption at different operating frequencies of the second processor can be intuitively seen, which helps to select the optimal second operating frequency.
[0186] The performance curve may be a trend graph showing the total operation delay of the cluster node changing with the operation frequency at different operation frequencies of the second processor. By drawing the performance curve, the performance of the second processor at different operation frequencies can be intuitively seen, which helps to select the optimal second operation frequency.
[0187] Exemplarily, first, the highest operating frequency of the first processor (responsible for the pre-filling process) can be determined as the first operating frequency. For example, if the highest frequency at which the first processor can stably operate is 2.5 GHz, then 2.5 GHz can be used as the first operating frequency of the first processor to ensure that the pre-filling process can be completed at the fastest speed.
[0188] In some implementations, the total power consumption and total delay of the first processor running at the first operating frequency and the second processor running at different operating frequencies can be predicted by a neural network model based on historical data. For example, the total historical power consumption and total historical delay of the historical task at different second operating frequencies under the first operating frequency are determined, each initial second operating frequency can correspond to multiple historical tasks, and from the multiple historical tasks, a target historical task with the highest similarity to the task to be processed is selected.
[0189] Furthermore, when selecting a target historical task with the highest similarity to the task to be processed, the number of first processors and second processors corresponding to the task to be processed, the task type of the task to be processed (for example, natural language processing, image processing, speech recognition, etc., the task type can affect the computational intensity and memory access intensity), task scale (the input data size of the task, for example, text length, image resolution, etc.), and task complexity (the number of operations required, such as the number of matrix multiplications, the number of convolution operations, etc.) can be obtained. Based on the number of first processors and second processors, the task type, the task scale, and the task complexity, the target historical task with the highest similarity to the task to be processed is selected from the historical tasks corresponding to each initial second operating frequency, and the historical total operating power consumption and historical total operating delay corresponding to the target historical task are used as the total operating power consumption and total operating delay of the cluster node when the second processor operates at the second operating frequency.
[0190] Furthermore, the power consumption and latency of the second processor (responsible for decoding processing) at different operating frequencies can be predicted. For example, the total power consumption and latency of the intelligent computing cluster at the second operating frequencies of 1.0 GHz, 1.5 GHz, and 2.0 GHz are predicted respectively, and the power consumption curve and performance curve are drawn based on the predicted results.
[0191] Furthermore, by analyzing the power consumption curve and the performance curve, if when the second operating frequency of the second processor is 1.5 GHz, the power consumption and performance of the second processor reach a good balance point, the power consumption is not increased much compared to the lowest frequency, and the performance (processing speed) is significantly improved, then 1.5 GHz can be used as the second operating frequency.
[0192] By setting the highest operating frequencies of multiple first processors and predicting the total power consumption and latency at different frequencies, a power consumption and performance curve can be drawn to optimize the selection of the best operating frequency of the second processor and achieve the best balance between power consumption and performance.
[0193] In some implementations, in order to ensure the orderly processing of tasks, intermediate features can be stored in the intermediate data temporary storage area according to the task number to ensure the orderliness and consistency of the data. For example, the task processing method based on the intelligent computing cluster can also include:
[0194] (D.1) Divide the task to be processed into multiple sub-processing tasks, and number the multiple sub-processing tasks according to the task sequence;
[0195] (D.2) In the intermediate data temporary storage area, the received intermediate features are stored according to the task number of each sub-processing task, and the intermediate features are sequentially obtained from the intermediate data temporary storage area according to the corresponding task numbers through multiple second processors.
[0196] Among them, the sub-processing task can be a subdivision of the task to be processed to improve the parallelism of processing. It means splitting the task to be processed into multiple smaller, independent processing units. Each sub-processing task can be pre-filled and decoded separately. This division method can help parallel processing and resource optimization, and improve the overall efficiency of task processing.
[0197] The task number may be a unique identifier assigned to each sub-processing task, which is used to distinguish and manage different sub-processing tasks. The task number is usually an increasing integer, which is numbered in the order in which the tasks are processed, so as to maintain the order and consistency of the tasks during the processing process.
[0198] For example, if the task to be processed is a natural language processing task, the input data contains 100 sentences, each sentence contains 100 characters, 5 first processors are used for pre-filling processing, and 3 second processors are used for decoding processing. At this time, 100 sub-processing tasks can be divided, and each sentence is a sub-processing task.
[0199] Furthermore, after the first processor completes processing the corresponding sentence and obtains the intermediate feature, the sub-processing task number can be used to correspond. For example, if the intermediate feature corresponds to sub-processing task 1, then its corresponding task number is also 1, so as to obtain intermediate feature 1, intermediate feature 2,..., intermediate feature 100.
[0200] Furthermore, 100 sub-processing tasks can be assigned to 3 second processors, and the allocation order can be determined according to the order of the serial numbers. For example, intermediate feature 1 is assigned to second processor 1, intermediate feature 2 is assigned to second processor 2, intermediate feature 3 is assigned to second processor 3, and then the cycle starts from second processor 1, intermediate feature 4 is assigned to second processor 1, and so on.
[0201] Furthermore, when a second processor completes decoding and obtains a decoding result, the decoding result of each second processor can be output in sequence, for example, the decoding result of intermediate feature 1 is output first, then the decoding result of intermediate feature 2, and so on.
[0202] In some embodiments, queue-based task scheduling, using tags and flags, managing task scheduling through dependency graphs, etc. can also be used to achieve reasonable task scheduling. For example, pipeline technology can be used to use processors in multiple stages to keep tasks flowing in a fixed order at different stages. After each stage is processed, the task will be passed to the next stage in a predetermined order. The embodiments of the present application do not impose too many restrictions on this.
[0203] Through the above methods, the order and consistency of task processing in the cluster nodes can be ensured, and the effectiveness and predictability of data management can be improved.
[0204] See also Figure 3 The embodiment of the present application further provides a task processing device based on an intelligent computing cluster, which can implement the above-mentioned task processing method based on an intelligent computing cluster. The task processing device based on an intelligent computing cluster includes:
[0205] An acquisition module 31 is used to acquire tasks to be processed and determine an intelligent computing cluster for processing the tasks to be processed;
[0206] A determination module 32, configured to determine a plurality of first processors and a plurality of second processors from a plurality of intelligent computing processors of the intelligent computing cluster according to a preset division rule, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency;
[0207] A processing module 33 is used to perform pre-filling processing on the task to be processed by using multiple first processors corresponding to the first operating frequency to obtain intermediate features, and transmit the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor;
[0208] The decoding module 34 is used to obtain the intermediate features from the intermediate data temporary storage area through the multiple second processors corresponding to the second operating frequency when detecting that the intermediate data temporary storage area is updated, and decode the intermediate features to obtain the decoding results of each second processor;
[0209] The output module 35 is used to sequentially output the decoding results of each second processor.
[0210] The specific implementation of the task processing device based on the intelligent computing cluster is basically the same as the specific implementation of the task processing method based on the intelligent computing cluster, and will not be repeated here. On the premise of meeting the requirements of the embodiment of the present application, the task processing device based on the intelligent computing cluster can also be provided with other functional modules to implement the task processing method based on the intelligent computing cluster in the above embodiment.
[0211] The embodiment of the present application also provides a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned task processing method based on the intelligent computing cluster when executing the computer program. The computer device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0212] See also Figure 4 , Figure 4 The hardware structure of a computer device according to another embodiment is shown, and the computer device includes:
[0213] The processor 41 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0214] The memory 42 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 42 can store an operating system and other applications. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 42, and the processor 41 calls and executes the task processing method based on the intelligent computing cluster in the embodiment of this application;
[0215] Input / output interface 43, used to implement information input and output;
[0216] The communication interface 44 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0217] A bus 45 that transmits information between the various components of the device (e.g., the processor 41, the memory 42, the input / output interface 43, and the communication interface 44);
[0218] The processor 41 , the memory 42 , the input / output interface 43 and the communication interface 44 are connected to each other in communication within the device via a bus 45 .
[0219] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned task processing method based on the intelligent computing cluster is implemented.
[0220] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0221] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0222] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0223] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0224] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0225] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0226] It should be understood that in the present application, "at least one (item)" and "several" refer to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0227] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0228] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0229] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0230] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0231] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A task processing method based on intelligent computing cluster, characterized in that: The method comprises: Obtaining tasks to be processed and determining an intelligent computing cluster for processing the tasks to be processed; According to a preset division rule, a plurality of first processors and a plurality of second processors are determined from a plurality of intelligent computing processors of the intelligent computing cluster, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency; Pre-filling the to-be-processed tasks by the plurality of first processors corresponding to the first operating frequency to obtain intermediate features, and transmitting the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor; When it is detected that the intermediate data temporary storage area is updated, the intermediate features are obtained from the intermediate data temporary storage area by the multiple second processors corresponding to the second operating frequency, and the intermediate features are decoded to obtain a decoding result of each second processor; The decoding results of each second processor are outputted in sequence.
2. The task processing method based on intelligent computing cluster according to claim 1 is characterized in that: The first processor and the second processor are located in the same cluster node; and determining the plurality of first processors and the plurality of second processors from the plurality of intelligent computing processors of the intelligent computing cluster according to a preset partitioning rule comprises: Determine, from the intelligent computing cluster, a cluster node for processing the task to be processed; Determine a first division ratio and a second division ratio for the plurality of intelligent computing processors of the cluster node according to a preset division rule; Determining a plurality of first processors from the plurality of intelligent computing processors according to the first division ratio; According to the second division ratio, a plurality of second processors are determined from the plurality of intelligent computing processors, and the total number of the plurality of first processors and the plurality of second processors is equal to the total number of intelligent computing processors of the cluster node.
3. The task processing method based on intelligent computing cluster according to claim 1 is characterized in that: The method further comprises: In the process of processing the tasks to be processed, regularly obtaining the average delay of the first characters of the pre-filling processing of the tasks to be processed by the multiple first processors and the average delay of the second characters of the decoding processing of the tasks to be processed by the multiple second processors; Compare the first character average delay with the second character average delay to obtain a first comparison result; When the first comparison result indicates that the first character average delay is inconsistent with the second character average delay, the number of the plurality of first processors and the number of the plurality of second processors are adjusted according to the first comparison result.
4. The task processing method based on intelligent computing cluster according to claim 3 is characterized in that: The adjusting the number of the plurality of first processors and the plurality of second processors according to the first comparison result includes: When the first comparison result indicates that the first character average delay is greater than the second character average delay, obtaining a first difference between the first character average delay and the second character average delay; Determine a first number of first processors to be added according to a first difference range corresponding to the first difference, and adjust the number of the plurality of first processors and the plurality of second processors according to the first number; When the first comparison result indicates that the first character average delay is less than the second character average delay, obtaining a second difference between the second character average delay and the first character average delay; A second number of added second processors is determined according to a second difference range corresponding to the second difference, and the number of the plurality of first processors and the plurality of second processors is adjusted according to the second number.
5. The task processing method based on intelligent computing cluster according to claim 3 is characterized in that: When the first comparison result indicates that the first character average delay and the second character average delay are inconsistent, after adjusting the number of the plurality of first processors and the plurality of second processors according to the first comparison result, the method further includes: Obtain the sum of the first character average delay and the second character average delay to obtain a total delay; Obtaining a preset delay threshold, and comparing the delay threshold with the total delay to obtain a second comparison result; When the second comparison result indicates that the total delay is greater than the delay threshold, determine a third operating frequency corresponding to the second processor, and obtain the intermediate feature from the intermediate data temporary storage area through the multiple second processors corresponding to the third operating frequency, and decode the intermediate feature to obtain a decoding result of each second processor; wherein the third operating frequency is greater than the second operating frequency; When the second comparison result indicates that the total delay is less than the delay threshold, the fourth operating frequency corresponding to the second processor is determined, and the intermediate features are obtained from the intermediate data temporary storage area through the multiple second processors corresponding to the fourth operating frequency, and the intermediate features are decoded to obtain the decoding results of each second processor; wherein the fourth operating frequency is less than the second operating frequency.
6. The task processing method based on intelligent computing cluster according to claim 1, characterized in that: The second operating frequency is determined by: Setting the first operating frequencies corresponding to the plurality of first processors as the highest operating frequency; Based on the highest operating frequency, predicting the total operating power consumption and total operating delay of the cluster nodes corresponding to the multiple first processors and the multiple second processors at different operating frequencies of the multiple second processors; Drawing a power consumption curve based on the total power consumption of the operation, and drawing a performance curve based on the total delay of the operation; Based on the power consumption curve and the performance curve, second operating frequencies of the plurality of second processors are determined.
7. The task processing method based on intelligent computing cluster according to claim 1, characterized in that: The method further comprises: Divide the to-be-processed task into a plurality of sub-processing tasks, and number the plurality of sub-processing tasks according to the task sequence; In the intermediate data temporary storage area, the received intermediate features are stored according to the task number of each sub-processing task, and the intermediate features are sequentially acquired from the intermediate data temporary storage area according to the corresponding task numbers through the multiple second processors.
8. A task processing device based on an intelligent computing cluster, characterized in that: The device comprises: An acquisition module is used to acquire tasks to be processed and determine an intelligent computing cluster for processing the tasks to be processed; A determination module, configured to determine, according to a preset division rule, a plurality of first processors and a plurality of second processors from a plurality of intelligent computing processors of the intelligent computing cluster, wherein each first processor corresponds to a first operating frequency, each second processor corresponds to a second operating frequency, and the first operating frequency is greater than the second operating frequency; a processing module, configured to perform pre-filling processing on the to-be-processed task by using the plurality of first processors corresponding to the first operating frequency to obtain intermediate features, and transmit the intermediate features to an intermediate data temporary storage area shared by the first processor and the second processor; a decoding module, configured to, when detecting that the intermediate data temporary storage area is updated, obtain the intermediate feature from the intermediate data temporary storage area through the plurality of second processors corresponding to the second operating frequency, and perform decoding processing on the intermediate feature to obtain a decoding result of each second processor; An output module is used to sequentially output the decoding results of each second processor.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the task processing method based on the intelligent computing cluster as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the task processing method based on an intelligent computing cluster as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Task scheduling method, electronic equipment and computer readable storage medium
CN117632400A
Model reasoning scheduling method and device and server cluster
CN118897736A
Scanning device for coded data
US20040190092A1