Task processing method and device, electronic equipment and storage medium
By fusing the cached data of the large model decoding task with pre-filled data and providing it to decoding nodes with low hardware utilization, the problem of insufficient hardware resource utilization during the large model inference process is solved, and service performance and user experience are improved.
Patent Information
- Application Number
- CN202510186139.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
AI Technical Summary
During the inference process, the large model has high requirements for hardware scheduling, load, computing power, storage and bandwidth, resulting in service overload and affecting user experience and service performance.
By fusing the first decoded cache data of the pending decoding task with the pre-filled cache data, the fused cache data is obtained and provided to the target decoding node with low hardware utilization to continue the pending decoding task.
The load of the decoding node to be optimized is reduced, the computing resources of the target decoding node are fully utilized, the waste of computing resources is avoided, and the service throughput and user experience is improved.
Smart Images

Figure CN120066724A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as chips, large models, and cloud computing. More specifically, the present disclosure provides a task processing method, apparatus, decoding server, electronic device, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, the application of large models is increasing continuously. Large models can include large language models (LLMs), image large models, and audio large models. The number of parameters of large models is relatively large and can reach the order of tens of billions or hundreds of billions. Summary of the Invention
[0003] The present disclosure provides a task processing method, apparatus, device, and storage medium.
[0004] According to one aspect of the present disclosure, there is provided a task processing method, the method including: fusing first decoding cache data of a to-be-processed decoding task with pre-filled cache data for the to-be-processed decoding task to obtain fused cache data of the to-be-processed decoding task, where the to-be-processed decoding task is one of multiple decoding tasks executed by a to-be-optimized decoding node, and the to-be-optimized decoding node is one of multiple decoding nodes; providing the fused cache data of the to-be-processed decoding task to a target decoding node among the multiple decoding nodes, where the target decoding node is configured to continue to execute the to-be-processed decoding task according to the fused cache data.
[0005] According to another aspect of the present disclosure, there is provided a task processing apparatus, the apparatus including: a fusion module configured to fuse first decoding cache data of a to-be-processed decoding task with pre-filled cache data for the to-be-processed decoding task to obtain fused cache data of the to-be-processed decoding task, where the to-be-processed decoding task is one of multiple decoding tasks executed by a to-be-optimized decoding node, and the to-be-optimized decoding node is one of multiple decoding nodes; a providing module configured to provide the fused cache data of the to-be-processed decoding task to a target decoding node among the multiple decoding nodes, where the target decoding node is configured to continue to execute the to-be-processed decoding task according to the fused cache data.
[0006] According to another aspect of the present disclosure, there is provided a decoding server, including a plurality of artificial intelligence acceleration boards, each serving as a plurality of decoding nodes; a first processor configured to: fuse the first decoding cache data of the to-be-processed decoding task with the pre-filled cache data for the to-be-processed decoding task to obtain the fused cache data of the to-be-processed decoding task, where the to-be-processed decoding task is one of a plurality of decoding tasks to be executed by a to-be-optimized decoding node, and the to-be-optimized decoding node is one of the plurality of decoding nodes; and provide the fused cache data of the to-be-processed decoding task to a target decoding node among the plurality of decoding nodes, where the target decoding node is configured to continue to execute the to-be-processed decoding task according to the fused cache data.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided according to the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method provided according to the present disclosure.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 is a schematic diagram of an exemplary system architecture to which a task processing method and apparatus according to an embodiment of the present disclosure can be applied;
[0013] Figure 2 is a flowchart of a task processing method according to an embodiment of the present disclosure;
[0014] Figure 3A and Figure 3B is a schematic diagram of a decoding server according to an embodiment of the present disclosure;
[0015] Figure 4 is a block diagram of a task processing apparatus according to an embodiment of the present disclosure;
[0016] Figure 5 is a schematic block diagram of a decoding server according to an embodiment of the present disclosure; and
[0017] Figure 6 is a block diagram of an electronic device to which a task processing method can be applied according to an embodiment of the present disclosure. Detailed implementation manners
[0018] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.
[0019] As mentioned above, the number of parameters of the large language model is large. During the reasoning process, the number of characters input to the large language model is also large, and the number of characters output by the large language model is also large. Therefore, the application of the large language model has high requirements for hardware scheduling, load, computing power, storage, bandwidth, etc. The following will take the large model being a large language model as an example for illustration.
[0020] The large language model can perform reasoning based on an autoregressive form. That is, new tokens need to be generated by the large language model based on the input prompt text and the previously generated tokens. To reduce the repeated calculation in the attention part, cache data corresponding to the previously generated tokens can be stored. The cache data can be key-value cache (KVcache) data. The reasoning of the large language model includes a prefill stage and a decode stage. In the prefill stage, the large language model receives the input prompt text and generates one or more tokens and the cache data for the prefill stage. In the decode stage, the large model can autoregressively generate tokens based on the cache data in the prefill stage until the end condition is met. The end condition can be that the number of generated tokens reaches a preset threshold, or an end token is generated.
[0021] The performance of large language models can be evaluated based on multiple evaluation metrics. The multiple evaluation metrics can include processing latency and the number of inference requests that can be processed per unit time. To improve the inference performance of large models, multiple inference optimization schemes can be adopted. For example, based on the Continuous Batching technology, inference requests can be dynamically added and removed to fully utilize the computing resources of the artificial intelligence processor. The artificial intelligence processor can be various processors such as general-purpose graphics processing units (GPGPUs), tensor processing units (TPUs), and neural network processing units (NPUs). Another example is that the paged attention mechanism draws on virtual memory and paging technologies in operating systems and can perform chunk management and storage on the above key-value cache data, improving storage utilization. Another example is that based on the architecture that separates pre-filling and decoding, the inference in the pre-filling stage and the inference in the decoding stage can be implemented on different devices respectively. After the inference in the pre-filling stage is completed on one device, the cache data generated in the pre-filling stage can be transferred to the device used for the inference in the decoding stage. The inference in the pre-filling stage is computationally intensive, and the inference in the decoding stage is memory access intensive. By using different devices to execute the inferences in these two stages respectively, different performance optimization strategies can be used for optimization to better utilize computing resources and storage resources.
[0022] For the architecture that separates pre-filling and decoding, the performance bottleneck in the decoding stage lies in memory access. Therefore, the performance optimization goal in the decoding stage can be to increase the batch size per round of inference and insert the requests processed in the pre-filling stage as much as possible, thereby improving the throughput of the service. The decoding stage can continuously generate new tokens. Before multiple inference requests are processed and the occupied memory space is released, the storage resources of the global storage unit for the artificial intelligence processor may be exhausted, resulting in service overload. To continue the inference, one or more inference requests being processed can be forcibly paused to release the occupied memory to ensure that other inference requests can continue. For the forcibly paused inference requests, the inference can be carried out in the following ways: starting the inference again from the pre-filling stage, but this will result in waste of computing resources and storage resources; or, first transferring the cache data of the forcibly paused inference requests to the first storage unit for the first processor, and then migrating the cache data of the forcibly paused requests back to the second storage unit for artificial intelligence processing when the storage resources of the artificial intelligence processor are sufficient to continue the inference, but the data transfer between the first storage unit and the second storage unit will occupy a lot of bandwidth resources and take a long time. Service overload is an inevitable problem in the architecture that separates pre-filling and decoding, which will have a negative impact on user experience and service performance. It can be understood that the first processor can be a central processing unit (CPU).
[0023] To cope with service overload, the load can be evaluated based on service level objectives, and it can be determined whether to receive new requests according to the load evaluation results of the pre-populated resource pool and the load evaluation results of the decoding resource pool. The pre-populated resource pool and the decoding resource pool can each include multiple devices. In the case where the decoding length is unknown, the load of the decoding resource pool may fluctuate. In this regard, the load in the decoding stage can be predicted to determine whether to receive requests. However, predicting the load in the decoding stage requires relying on a large amount of empirical data or statistical data and cannot accurately reflect the actual processing conditions of different inference requests, resulting in difficult full utilization of computing resources.
[0024] To cope with service overload, a prediction model can also be used to predict the number of tokens generated in the decoding stage for an inference request to be processed, and then predict the computing resources and storage resources required to process the inference request. However, using a prediction model for prediction requires additional computing resources.
[0025] Therefore, to effectively cope with service overload, the present disclosure provides a task processing method, which will be described below.
[0026] Figure 1 is a schematic diagram of an exemplary system architecture to which the task processing method and apparatus according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1 The illustration is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0027] As Figure 1 shown, the system architecture according to this embodiment can include a terminal device 11, a network 12, and a server cluster 10. The network 12 is used to provide a medium for a communication link between the terminal device 11 and the server cluster 10. The network 12 can also be used to provide a medium for a communication link within the server cluster 10. The network 12 can include various connection types, such as wired and / or wireless communication links, etc.
[0028] Users can use the terminal device 11 to interact with the server cluster 10 through the network 12 to receive or send messages, etc. For example, the terminal device 11 can send an inference request to the server cluster 10 through the network 12. The inference request can include natural language text as prompt text.
[0029] Various communication client applications can also be installed on the terminal device 11, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0030] The terminal device 11 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop portable computers, desktop computers, and so on.
[0031] The server cluster 10 can be a server that provides various services. For example, it can be a server cluster that provides an inference service for an inference request sent by a user using the terminal device 11 (only for example).
[0032] The server cluster 10 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server, VPS). The server can also be a server of a distributed system or a server combined with a blockchain.
[0033] The server cluster 10 can include one or more pre-population servers and one or more decoding servers. The task processing method can be applied to one or more decoding servers in the server cluster 10. The pre-population server can receive the prompt text passed by the user, execute the pre-population task to complete the inference in the pre-population stage. The pre-population server can provide the pre-population cache data to the decoding server. The decoding server executes the decoding task according to the pre-population buffer data to complete the inference in the decoding stage. As Figure 1 shown, in the server cluster 10, multiple pre-population servers can include a first pre-population server 101_1 and a second pre-population server 101_2, and multiple decoding services can include a first decoding server 100_1 and a second decoding server 100_2. Each pre-population server can include multiple hardware units. Each decoding server can include multiple hardware units.
[0034] It can be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in the server cluster are only illustrative. According to the implementation requirements, there can be any number of terminal devices, the network, and the servers.
[0035] It can be understood that the system architecture of the present disclosure has been described above, and the method of the present disclosure will be described below. It can also be understood that the serial numbers of each operation in the following method are only used as representations of the operation for description and should not be regarded as indicating the execution order of each operation. Unless explicitly stated, the method does not need to be executed exactly in the order shown.
[0036] Figure 2 is a flowchart of a task processing method according to an embodiment of the present disclosure.
[0037] As Figure 2As shown, the method 200 may include operation S210 to operation S220.
[0038] In operation S210, the first decoded cache data of the decoding task to be processed is fused with the pre-filled cache data for the decoding task to be processed, to obtain the fused cache data of the decoding task to be processed.
[0039] In the embodiments of the present disclosure, the decoding server may include multiple artificial intelligence acceleration boards. Each artificial intelligence acceleration board may serve as a decoding node. The decoding node to be optimized may be one of the multiple decoding nodes.
[0040] In the embodiments of the present disclosure, after receiving the prompt text, the pre-filling task may be first executed to complete the inference in the pre-filling stage, to obtain one or more tokens and the pre-filled cache data. According to the pre-filled cache data, the decoding task may be executed to complete the inference in the decoding stage.
[0041] In the embodiments of the present disclosure, the decoding task to be processed is one of the multiple decoding tasks executed by the decoding node to be optimized.
[0042] In the embodiments of the present disclosure, the first decoded cache data may be the decoded cache data generated by the decoding node to be optimized during the execution of the decoding task to be processed. The decoded cache data may be the key-value cache data generated in the decoding stage. For example, in the pre-filling stage, N tokens and the pre-filled cache data for the decoding task to be processed are generated. The N tokens include the first token. N may be an integer greater than or equal to 1. During the execution of the decoding task to be processed by the decoding node to be optimized, M tokens are generated. The M tokens may include the (N + 1)-th token. M may be an integer greater than or equal to 1.
[0043] In operation S220, the fused cache data of the decoding task to be processed is provided to the target decoding node among the multiple decoding nodes.
[0044] In the embodiments of the present disclosure, the target decoding node may continue to execute the decoding task to be processed according to the fused cache data. The target decoding node may be the decoding node other than the decoding node to be optimized among the multiple decoding nodes. For example, the target decoding node may generate K tokens. The K tokens may include the (N + M + 1)-th token. K may be an integer greater than or equal to 1.
[0045] Through the embodiments of the present disclosure, providing the cached data of the decoding node to be optimized to the target decoding node can reduce the load of the decoding node to be optimized and make full use of the computing resources of the target decoding node. Moreover, the cached data provided to the target decoding node is the fused cached data that fuses the first decoded cached data and the pre-filled cached data, enabling the target decoding node to continue to execute the decoding task to be processed, further making full use of the computing resources and avoiding waste of computing resources.
[0046] It can be understood that the method of the present disclosure has been described above. Below, the method of the present disclosure will be further described in conjunction with Figure 3A and Figure 3B to further illustrate the method of the present disclosure.
[0047] Figure 3A and Figure 3B are schematic diagrams of a decoding server according to an embodiment of the present disclosure.
[0048] As Figure 3A shown, the pre-filling server 301 can execute one or more pre-filling tasks. After executing one or more pre-filling tasks, one or more pre-filled cached data are provided to the decoding server 300 so that the decoding server 300 can execute one or more decoding tasks.
[0049] In some embodiments, a pre-filling instance can be constructed based on one or more pre-filling servers. A decoding instance can be constructed based on one or more decoding servers. As Figure 3A shown, a pre-filling instance can be constructed based on the pre-filling server 301. A decoding instance can be constructed based on the decoding server 300. A task scheduling manager 311 can be set for the decoding instance. The decoding server 300 can include a first processor and a first storage unit for the first processor. The first processor can be a central processing unit. The first storage unit can be a main memory. The task scheduling manager 311 can be a system process run by the first processor.
[0050] In some embodiments, the decoding server can include multiple artificial intelligence acceleration boards. The artificial intelligence acceleration board can include a second processor and a second storage unit. The artificial intelligence acceleration board can act as a decoding node to execute one or more decoding tasks. The second processor can be an artificial intelligence processor. The second storage unit can be a global storage unit.
[0051] In some embodiments, the task scheduling manager may determine the hardware utilization rate of each of the multiple decoding nodes. For example, the hardware utilization rate may include memory utilization rate. The data volume of the key-value cache data is positively correlated with the length of the sequence. The sequence may be a sequence of tokens. Without considering memory fragmentation, the capacity of the second storage unit occupied by the key-value cache data may be determined according to the length of the sequence. It can be understood that the memory utilization rate of the decoding node may be the utilization rate of the global storage unit.
[0052] In some embodiments, the task scheduling manager may also perform task scheduling. According to the memory utilization rate of the decoding node, one or more decoding tasks to be executed may be scheduled. If the memory utilization rate of the decoding node is less than the first preset memory utilization threshold, a new decoding task may be inserted into the decoding node. If the memory utilization rate of the decoding node is greater than or equal to the first preset memory utilization threshold and less than or equal to the second preset memory utilization threshold, no new task may be inserted into the decoding node. If the memory utilization rate of the decoding node is greater than the second preset memory utilization threshold, the storage resources occupied by one or more decoding tasks executed by the decoding node may be released. If the memory utilization rates of multiple decoding nodes are all greater than the first preset memory utilization threshold, the decoding tasks to be executed may be added to the task queue to be allocated. The task queue to be allocated for multiple decoding nodes may be stored in the first storage unit. The first preset memory utilization threshold may be 90% for example. The second preset memory utilization threshold may be 95% for example.
[0053] In some embodiments, based on the average memory utilization rate of multiple decoding nodes, the task scheduling manager may determine whether to continue accepting new tasks from the pre-population server. For example, if the average memory utilization rate of multiple decoding nodes is greater than the first preset memory utilization threshold, the decoding server no longer receives new tasks except for the tasks that have already started transmitting pre-populated cache data.
[0054] In some embodiments, the task scheduling manager may manage multiple decoding task queues for multiple decoding nodes. The first storage unit may store multiple decoding task queues respectively for multiple decoding nodes. The first storage unit may also store multiple pre-populated cache data for multiple decoding tasks of the decoding task queues. As Figure 3AAs shown, the first storage unit 312 can store multiple decoding task queues. The multiple decoding task queues include the decoding task queue KVQueue_31 for the artificial intelligence acceleration board 321, the decoding task queue KVQueue_32 for the artificial intelligence acceleration board 322, ……, and the decoding task queue KVQueue_33 for the artificial intelligence acceleration board 323. The decoding task queue KVQueue_31 includes decoding tasks Task_316, Task_322, ……, Task_330. The first storage unit can store multiple pre-filled cache data respectively for the decoding tasks Task_316, Task_322, ……, Task_330. The decoding task queue KVQueue_32 includes decoding tasks Task_311 and Task_319. The first storage unit can store multiple pre-filled cache data respectively for the decoding tasks Task_311 and Task_319. The decoding task queue KVQueue_33 includes decoding tasks Task_320, Task_325, ……, Task_329. The first storage unit can store multiple pre-filled cache data respectively for the decoding tasks Task_320, Task_325, ……, Task_329. It can be understood that Figure 3A The number of the artificial intelligence acceleration boards and the number of the decoding task queues shown are only examples. The number of the artificial intelligence acceleration boards can be 2 or more than 3. The number of the decoding task queues can be 2 or more than 3. Through the embodiments of the present disclosure, the first storage unit stores the pre-filled cache data respectively for multiple decoding tasks. If the decoding tasks in the decoding task queue for a decoding node are migrated to another decoding node, the time cost required to copy the pre-filled cache data for the decoding task from the second storage unit of the decoding node can be reduced, and the bandwidth resources can also be saved, and the task migration efficiency can be improved.
[0055] In some embodiments, the number of decoding tasks included in the decoding task queue is less than or equal to the number of decoding tasks executed by the decoding node. For example, the number of decoding tasks executed by a decoding node can be I, and the number of decoding tasks in the decoding task queue for this decoding node can be J, and I can be greater than or equal to J. I can be an integer greater than or equal to 1, and J can be an integer greater than or equal to 1. That is, the I decoding tasks executed by this decoding node include the J decoding tasks in the decoding task queue and the I-J decoding tasks not in the decoding task queue. The execution duration of any one of the I-J decoding tasks is greater than or equal to the execution duration of any one of the decoding tasks in the decoding task queue. The execution duration can be the duration for the decoding node to execute a decoding task.
[0056] It can be understood that the task scheduling manager of the present disclosure has been described above, and the decoding nodes to be optimized of the present disclosure will be described below.
[0057] In some embodiments, the decoding nodes to be optimized are the decoding nodes among multiple decoding nodes whose hardware utilization rate is greater than or equal to a preset utilization rate threshold. For example, taking the hardware utilization rate as the memory utilization rate as an example, the preset utilization rate threshold can be the second preset memory utilization rate threshold mentioned above. Below, taking the memory utilization rate of the artificial intelligence acceleration board 321 being greater than the second preset memory utilization rate threshold as an example, an explanation will be given. That is, the artificial intelligence acceleration board 321 can be used as a decoding node to be optimized.
[0058] In some embodiments, at least one decoding task to be screened can be determined from multiple decoding tasks executed by the decoding nodes to be optimized. The decoding task to be screened can be a decoding task that satisfies at least one of multiple preset conditions among the multiple decoding tasks executed by the decoding nodes to be optimized. The preset conditions can include a first preset condition and a second preset condition. The first preset condition can be that the execution duration is the minimum among the multiple decoding tasks executed by the decoding nodes to be optimized. The second preset condition can be that the execution duration is less than or equal to a preset duration threshold. For example, the decoding task Task_316 is executed by the artificial intelligence acceleration board 321 for a long time. The artificial intelligence acceleration board 321 has decoded, for example, 20 markers corresponding to the decoding task Task_316. The decoding task Task_322 is also executed by the artificial intelligence acceleration board 321 for a long time. The artificial intelligence acceleration board 321 has decoded, for example, 15 markers corresponding to the decoding task Task_322. Among the decoding tasks Task_316, the decoding task Task_322, and the decoding task Task_330, the decoding task Task_330 is executed by the artificial intelligence acceleration board 321 for a shorter time. The artificial intelligence acceleration board 321 has decoded, for example, 1 marker corresponding to the decoding task Task_330. Among the multiple decoding tasks included in the decoding task queue KVQueue_31, if the execution duration of the decoding task Task_330 by the artificial intelligence acceleration board 321 is the minimum, it can be determined that the decoding task Task_330 satisfies the first preset condition and can be used as a decoding task to be screened. If one decoding task to be screened is determined from multiple decoding tasks, this decoding task to be screened can be used as a decoding task to be processed.
[0059] It can be understood that the present disclosure has been described above by taking the decoding task Task_330 satisfying the first preset condition as an example. However, the present disclosure is not limited thereto. If it is determined that a decoding task satisfies the second preset condition, this decoding task can also be used as a decoding task to be screened.
[0060] It can be understood that if multiple decoding tasks to be screened are determined from multiple decoding tasks, the decoding tasks to be screened can be sequentially used as the decoding tasks to be processed, or a decoding task to be processed can be randomly determined from the multiple decoding tasks to be screened.
[0061] It can be understood that the decoding tasks to be processed in the present disclosure have been described above, and the first decoded cache data of the decoding tasks to be processed will be described below.
[0062] In some embodiments, the first decoded cache data of the decoding task to be processed is stored in the storage space to be processed of the decoding node, and the storage space to be processed is located in the second storage unit of the decoding node. As described above, the decoding task Task_330 can be used as the decoding task to be processed. The first decoded cache data of the decoding task Task_330 can be stored in the storage space to be processed of the artificial intelligence acceleration board 321. This storage space to be processed can be the second storage unit located in the artificial intelligence acceleration board 321. As described above, for example, the artificial intelligence acceleration board 321 has decoded 1 tag corresponding to the decoding task Task_330, and this first decoded cache data can be the key-value cache data corresponding to this 1 tag.
[0063] It can be understood that the first decoded cache data of the present disclosure has been described above, and the fused cache data of the present disclosure will be described below.
[0064] In some embodiments, in some implementation manners of the above operation S210, fusing the first decoded cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed to obtain the fused cache data of the decoding task to be processed may include: providing the first decoded cache data of the decoding task to be processed from the second storage unit of the decoding node to be optimized to the first storage unit. As Figure 3A shown, the first decoded cache data of the decoding task Task_330 in the second storage unit can be provided to the first storage unit 312. Through the embodiments of the present disclosure, when the first storage unit stores the pre-filled cache data for the decoding task to be processed, the first decoded cache data in the second storage unit can be provided to the first storage unit, which can effectively save bandwidth resources.
[0065] In some embodiments, in some implementation manners of the above operation S210, fusing the first decoded cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed to obtain the fused cache data of the decoding task to be processed may further include: in the first storage unit, fusing the first decoded cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed to obtain the fused cache data of the decoding task to be processed. As Figure 3AAs shown, in the first storage unit 312, the first decoded cache data can be added after the pre-filled cache data for the decoding task Task_330 to fuse the first decoded cache data and the pre-filled cache data, resulting in fused cache data.
[0066] In some embodiments, in some implementations of the above operation S210, fusing the first decoded cache data of the decoding task to be processed with the pre-filled cache data of the decoding task to be processed to obtain the fused cache data of the decoding task to be processed may further include: after providing the first decoded cache data of the decoding task to be processed from the second storage unit of the decoding node to be optimized to the first storage unit, releasing the storage space to be processed. As Figure 3A shown, after providing the first decoded cache data of the decoding task Task_330 to the first storage unit 312, the first decoded cache data in the artificial intelligence acceleration board 321 can be deleted to release the storage space to be processed. Through the embodiments of the present disclosure, the hardware utilization rate of the decoding node to be optimized can be reduced, and the load of the cache space to be optimized can be reduced, so that the decoding node to be optimized can continue to efficiently execute one or more decoding tasks other than the decoding task to be processed.
[0067] It can be understood that the fused cache data of the present disclosure has been described above, and the target decoding node of the present disclosure will be described below.
[0068] In some embodiments, the hardware utilization rate of the target decoding node is less than the hardware utilization rate of the decoding node to be optimized. For example, the memory utilization rate of the target decoding node can be less than the first preset memory utilization rate threshold (90%). As Figure 3A shown, if the memory utilization rate of the artificial intelligence acceleration board 322 is less than the memory utilization rate of the artificial intelligence acceleration board 321, the artificial intelligence acceleration board 322 can be used as the target decoding node.
[0069] In some embodiments, in some implementations of the above operation S220, providing the fused cache data of the decoding task to be processed to the target decoding node among multiple decoding nodes may include: in response to determining that there is a target decoding node among the multiple decoding nodes, adding the decoding task to be processed to the decoding task queue for the target decoding node. Writing the fused cache data of the decoding task to be processed into the second storage unit of the target decoding node. As Figure 3BAs shown, taking the AI acceleration board 322 as the target decoding node as an example, the decoding task Task_330 can be added to the decoding task queue KVQueue_32 for the AI acceleration board 322. Moreover, the fused cache data of the decoding task Task_330 can be written into the second storage unit of the AI acceleration board 322. Next, the AI acceleration board 322 can continue to execute the decoding task Task_330 to decode the second token to the last token in the decoding stage. Through the embodiments of the present disclosure, using the decoding node with low hardware utilization to continue executing the pending decoding task can make full use of the computing resources and improve the utilization rate of the computing resources. It can be understood that the key-value cache data corresponding to the second token to the last token can be used as the second decoding cache task of the pending decoding task.
[0070] It can be understood that the above description of the present disclosure is given by taking the existence of a target decoding node among multiple decoding nodes as an example. However, the present disclosure is not limited thereto, and the above-mentioned fused cache data can be obtained at the first moment. Below, taking the non-existence of a target decoding node among multiple decoding nodes at the first moment as an example, an explanation will be given.
[0071] In some other embodiments, in some other implementation manners of the above operation S220, in response to determining that there is no target decoding node among multiple decoding nodes at the first moment, the pending decoding task is added to the pending scheduling task queue as a pending scheduling task. In response to determining that there is a target decoding node among multiple decoding nodes at the second moment, the pending scheduling task at the head of the pending scheduling task queue is provided to the target decoding node, where the second moment is a moment after the first moment. For example, at the first moment, the memory utilization rate of each of the multiple decoding nodes is greater than the first preset memory utilization threshold, and it can be determined that there is no target decoding node among the multiple decoding nodes. In this case, the pending decoding task can be added to the pending scheduling task queue as a pending scheduling task. The pending scheduling task queue can be stored in the first storage unit. At the second moment after the first moment, if a decoding node among the multiple decoding nodes has completed one or more decoding tasks and released the computing resources and storage resources occupied by the completed one or more decoding tasks, and the memory utilization rate of this decoding node is less than the memory utilization rate of the decoding node to be optimized, this decoding node can be used as the target decoding node. Next, the pending scheduling task at the head of the pending scheduling queue can be added to the decoding task queue for the target decoding node, and the fused cache data of the pending scheduling task can be written into the second storage unit of the target decoding node to provide the pending scheduling task to the target decoding node.
[0072] It can be understood that when there is a pending scheduling task in the pending scheduling queue, the decoding server can also receive the pre-filled cache data for a new decoding task. Below, some ways of processing the new decoding task will be described.
[0073] In some embodiments, the first storage unit may further store a task queue to be assigned for multiple decoding nodes. For example, a decoding server may receive multiple pre-filled cache data for multiple new decoding tasks. The new decoding tasks may be added to the task queue to be assigned.
[0074] In some embodiments, the above method may further include: in response to determining that there is no task to be scheduled in the task queue to be scheduled, providing the decoding task in the task queue to be assigned to one of the multiple decoding nodes. For example, in the case where there is a task to be scheduled in the task queue to be scheduled and there is a decoding node among the multiple decoding nodes with a memory utilization rate less than the first preset memory utilization rate threshold, even if there is a new decoding task in the task queue to be assigned, the task to be scheduled may be preferentially provided to the decoding node. In the case where there is no task to be scheduled in the task queue to be scheduled, the decoding task in the task queue to be assigned may be provided to the decoding node.
[0075] It can be understood that the task processing method provided by the present disclosure may be executed by a first processor of a decoding server.
[0076] It can be understood that the above mainly illustrates the present disclosure by taking the hardware utilization rate as the memory utilization rate as an example. However, the present disclosure is not limited thereto, and the hardware utilization rate may also be various utilization rates such as computing power utilization rate and bandwidth utilization rate.
[0077] It can be understood that the method of the present disclosure has been described above, and the apparatus of the present disclosure will be described below.
[0078] Figure 4 is a block diagram of a task processing apparatus according to an embodiment of the present disclosure.
[0079] As Figure 4 shown, the apparatus 411 may include a fusion module 4111 and a providing module 4112.
[0080] The fusion module 4111 is configured to fuse the first decoding cache data of the decoding task to be processed with the pre-filled cache data of the decoding task to be processed, to obtain the fused cache data of the decoding task to be processed. The decoding task to be processed is one of multiple decoding tasks executed by a decoding node to be optimized, and the decoding node to be optimized is one of the multiple decoding nodes.
[0081] The first providing module 4112 is configured to provide the fused cache data of the decoding task to be processed to a target decoding node among the multiple decoding nodes. The target decoding node is configured to continue to execute the decoding task to be processed according to the fused cache data.
[0082] In some embodiments, the decoding node to be optimized is a decoding node among multiple decoding nodes whose hardware utilization rate is greater than or equal to a preset utilization rate threshold, and the hardware utilization rate of the target decoding node is less than that of the decoding node to be optimized.
[0083] In some embodiments, the decoding task to be processed is one of at least one decoding task to be screened among multiple decoding tasks. The decoding task to be screened is a decoding task among multiple decoding tasks that meets at least one of multiple preset conditions. The multiple preset conditions include: the execution duration is the minimum among multiple decoding tasks executed by the decoding node to be optimized; the execution duration is less than or equal to a preset duration threshold.
[0084] In some embodiments, the pre-filled cache data of the decoding task to be processed is stored in a first storage unit for a first processor. The first storage unit is used to store multiple decoding task queues respectively for multiple decoding nodes, and the first storage unit is also used to store multiple pre-filled cache data for multiple decoding tasks in the decoding task queues. The first decoding cache data of the decoding task to be processed is stored in the to-be-processed storage space of the decoding node, and the to-be-processed storage space is located in the second storage unit of the decoding node.
[0085] In some embodiments, the fusion module includes: a first providing sub-module, configured to provide the first decoding cache data of the decoding task to be processed from the second storage unit of the decoding node to be optimized to the first storage unit. A fusion sub-module, configured to fuse the first decoding cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed in the first storage unit to obtain the fused cache data of the decoding task to be processed.
[0086] In some embodiments, the fusion module further includes: a releasing sub-module, configured to release the to-be-processed storage space after providing the first decoding cache data of the decoding task to be processed from the second storage unit of the decoding node to be optimized to the first storage unit.
[0087] In some embodiments, the first providing module includes: a first adding sub-module, configured to add the decoding task to be processed to the decoding task queue for the target decoding node in response to determining that there is a target decoding node among multiple decoding nodes. A writing sub-module, configured to write the fused cache data of the decoding task to be processed into the second storage unit of the target decoding node.
[0088] In some embodiments, the first storage unit is further configured to store a task queue to be scheduled for multiple decoding nodes, and the fused cache data of the decoding task to be processed is obtained at a first moment. The first providing module includes: a second adding sub-module, configured to add the decoding task to be processed as a task to be scheduled to the task queue to be scheduled in response to determining that there is no target decoding node among the multiple decoding nodes at the first moment. A second providing sub-module, configured to provide the task to be scheduled at the head of the task queue to be scheduled to the target decoding node in response to determining that there is a target decoding node among the multiple decoding nodes at a second moment. The second moment is a moment after the first moment.
[0089] In some embodiments, the first storage unit is configured to store a task queue to be assigned for multiple decoding nodes. The apparatus further includes: a second providing module, configured to provide the decoding task in the task queue to be assigned to one of the multiple decoding nodes in response to determining that there is no task to be scheduled in the task queue to be scheduled.
[0090] It can be understood that the apparatus of the present disclosure has been described above, and the decoding server of the present disclosure will be described below.
[0091] Figure 5 It is a schematic block diagram of a decoding server according to an embodiment of the present disclosure.
[0092] As Figure 5 shown, the decoding server 500 may include a first processor 511 and multiple artificial intelligence acceleration boards 520.
[0093] The multiple artificial intelligence acceleration boards 520 may each serve as a plurality of decoding nodes.
[0094] The first processor 511 may be configured to: fuse the first decoding cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed to obtain the fused cache data of the decoding task to be processed. The decoding task to be processed is one of multiple decoding tasks executed by the decoding node to be optimized, and the decoding node to be optimized is one of the multiple decoding nodes. Provide the fused cache data of the decoding task to be processed to the target decoding node among the multiple decoding nodes. The target decoding node is configured to continue to execute the decoding task to be processed according to the fused cache data. For example, the first processor 511 may be configured to execute the above method 200.
[0095] In some embodiments, the decoding node to be optimized is a decoding node among the multiple decoding nodes whose hardware utilization rate is greater than or equal to a preset utilization rate threshold, and the hardware utilization rate of the target decoding node is less than the hardware utilization rate of the decoding node to be optimized.
[0096] In some embodiments, the decoding task to be processed is one of at least one decoding task to be screened out among multiple decoding tasks, and the decoding task to be screened out is a decoding task that satisfies at least one of multiple preset conditions among multiple decoding tasks. The multiple preset conditions include: the execution duration is the minimum among the multiple decoding tasks executed by the decoding node to be optimized. The execution duration is less than or equal to a preset duration threshold.
[0097] In some embodiments, the first processor is further configured to perform the following operations to fuse the first decoded cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed, so as to obtain the fused cache data of the decoding task to be processed: provide the first decoded cache data of the decoding task to be processed from the second storage unit of the decoding node to be optimized to the first storage unit; in the first storage unit, fuse the first decoded cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed to obtain the fused cache data of the decoding task to be processed.
[0098] In some embodiments, the first processor is further configured to perform the following operations to fuse the first decoded cache data of the decoding task to be processed with the pre-filled cache data of the decoding task to be processed, so as to obtain the fused cache data of the decoding task to be processed: after providing the first decoded cache data of the decoding task to be processed from the second storage unit of the decoding node to be optimized to the first storage unit, release the storage space to be processed.
[0099] In some embodiments, the first processor is further configured to perform the following operations to provide the fused cache data of the decoding task to be processed to a target decoding node among multiple decoding nodes: in response to determining that there is a target decoding node among the multiple decoding nodes, add the decoding task to be processed to the decoding task queue for the target decoding node; write the fused cache data of the decoding task to be processed into the second storage unit of the target decoding node.
[0100] In some embodiments, the first storage unit is further used to store a pending task queue for multiple decoding nodes, and the fused cache data of the decoding task to be processed is obtained at a first moment. The first processor is further configured to perform the following operations to provide the fused cache data of the decoding task to be processed to a target decoding node among multiple decoding nodes: in response to determining that there is no target decoding node among the multiple decoding nodes at the first moment, add the decoding task to be processed as a pending task to the pending task queue; in response to determining that there is a target decoding node among the multiple decoding nodes at a second moment, provide the pending task at the head of the pending task queue to the target decoding node, where the second moment is a moment after the first moment.
[0101] In some embodiments, the first storage unit is used to store the task queue to be assigned for multiple decoding nodes. The first processor is further configured to: in response to determining that there is no task to be scheduled in the task queue to be scheduled, provide the decoding task in the task queue to be assigned to one of the multiple decoding nodes.
[0102] In some embodiments, the pre-filled cache data of the decoding task to be processed is stored in the first storage unit for the first processor. The first storage unit is used to store multiple decoding task queues respectively for multiple decoding nodes, and the first storage unit is further used to store multiple pre-filled cache data respectively for the multiple decoding tasks in the decoding task queue. The first decoding cache data of the decoding task to be processed is stored in the storage space to be processed of the decoding node, and the storage space to be processed is located in the second storage unit of the decoding node.
[0103] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0104] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0105] Figure 6 The schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0106] As Figure 6 shown, the device 600 includes a computing unit 601, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 602 or the computer program loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0107] Multiple components in device 600 are connected to I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0108] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the task processing method. For example, in some embodiments, the task processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the task processing method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the task processing method in any other suitable manner (e.g., by means of firmware).
[0109] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0110] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0111] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0112] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0113] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0114] A computer system may include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on respective computers and have a client-server relationship with each other.
[0115] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0116] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A task processing method, comprising: Merging first decoding cache data of a decoding task to be processed with pre-filled cache data for the decoding task to be processed to obtain fused cache data of the decoding task to be processed, wherein the decoding task to be processed is one of multiple decoding tasks executed by a decoding node to be optimized, and the decoding node to be optimized is one of multiple decoding nodes; The fused cache data of the decoding task to be processed is provided to a target decoding node among the plurality of decoding nodes, wherein the target decoding node is used to continue to execute the decoding task to be processed according to the fused cache data.
2. The method according to claim 1, wherein: The decoding node to be optimized is a decoding node whose hardware utilization is greater than or equal to a preset utilization threshold among the multiple decoding nodes, and the hardware utilization of the target decoding node is less than the hardware utilization of the decoding node to be optimized.
3. The method according to claim 1, wherein: The decoding task to be processed is one of at least one decoding task to be screened among the plurality of decoding tasks, and the decoding task to be screened is a decoding task that satisfies at least one of a plurality of preset conditions among the plurality of decoding tasks, and the plurality of preset conditions include: The execution time is the minimum value among the multiple decoding tasks executed by the decoding node to be optimized; The execution duration is less than or equal to a preset duration threshold.
4. The method according to claim 1, wherein: The pre-filled cache data of the decoding task to be processed is stored in a first storage unit for the first processor, the first storage unit is used to store a plurality of decoding task queues respectively used for the plurality of decoding nodes, and the first storage unit is also used to store a plurality of pre-filled cache data for the plurality of decoding tasks in the decoding task queues, The first decoding cache data of the decoding task to be processed is stored in the storage space to be processed of the decoding node, and the storage space to be processed is located in the second storage unit of the decoding node.
5. The method according to claim 4, wherein: The step of fusing the first decoding cache data of the decoding task to be processed with the pre-filled cache data for the decoding task to be processed to obtain the fused cache data of the decoding task to be processed includes: Providing the first decoding cache data of the decoding task to be processed from the second storage unit of the decoding node to be optimized to the first storage unit; In the first storage unit, the first decoding cache data of the decoding task to be processed is merged with the pre-filled cache data for the decoding task to be processed to obtain the merged cache data of the decoding task to be processed.
6. The method according to claim 5, wherein: The step of fusing the first decoding cache data of the decoding task to be processed with the pre-filled cache data of the decoding task to be processed to obtain the fused cache data of the decoding task to be processed further comprises: After the first decoding cache data of the decoding task to be processed is provided from the second storage unit of the decoding node to be optimized to the first storage unit, the storage space to be processed is released.
7. The method according to claim 4, wherein: Providing the fused cache data of the to-be-processed decoding task to a target decoding node among the plurality of decoding nodes comprises: In response to determining that the target decoding node exists among the plurality of decoding nodes, adding the to-be-processed decoding task to a decoding task queue for the target decoding node; The fusion cache data of the decoding task to be processed is written into the second storage unit of the target decoding node.
8. The method according to claim 4, wherein: The first storage unit is further used to store a queue of tasks to be scheduled for the plurality of decoding nodes, wherein the fused cache data of the decoding tasks to be processed is obtained at a first moment; Providing the fused cache data of the to-be-processed decoding task to a target decoding node among the plurality of decoding nodes comprises: In response to determining that the target decoding node does not exist among the plurality of decoding nodes at the first moment, adding the to-be-processed decoding task as a to-be-scheduled task to the to-be-scheduled task queue; In response to determining that the target decoding node exists among the plurality of decoding nodes at a second moment, providing the to-be-scheduled task at the head of the to-be-scheduled task queue to the target decoding node, the second moment being a moment after the first moment.
9. The method according to claim 8, wherein: The first storage unit is used to store a queue of tasks to be assigned to a plurality of the decoding nodes. Also includes: In response to determining that there is no to-be-scheduled task in the to-be-scheduled task queue, providing the decoding task in the to-be-assigned task queue to one of the plurality of decoding nodes.
10. A task processing device, comprising: a fusion module, configured to fuse first decoding cache data of a decoding task to be processed with pre-filled cache data for the decoding task to be processed, to obtain fused cache data of the decoding task to be processed, wherein the decoding task to be processed is one of a plurality of decoding tasks executed by a decoding node to be optimized, and the decoding node to be optimized is one of a plurality of decoding nodes; A module is provided, which is used to provide the fused cache data of the decoding task to be processed to a target decoding node among the multiple decoding nodes, wherein the target decoding node is used to continue to execute the decoding task to be processed according to the fused cache data.
11. A decoding server, comprising: Multiple AI acceleration boards serve as multiple decoding nodes; The first processor is configured as follows: Merging first decoding cache data of a decoding task to be processed with pre-filled cache data for the decoding task to be processed to obtain fused cache data of the decoding task to be processed, wherein the decoding task to be processed is one of multiple decoding tasks executed by a decoding node to be optimized, and the decoding node to be optimized is one of multiple decoding nodes; The fused cache data of the decoding task to be processed is provided to a target decoding node among the plurality of decoding nodes, wherein the target decoding node is used to continue to execute the decoding task to be processed according to the fused cache data.
12. The decoding server according to claim 11, wherein: The pre-filled cache data of the decoding task to be processed is stored in a first storage unit for the first processor, the first storage unit is used to store a plurality of decoding task queues respectively used for a plurality of the decoding nodes, and the first storage unit is also used to store a plurality of pre-filled cache data for a plurality of the decoding tasks in the decoding task queues, The first decoding cache data of the decoding task to be processed is stored in the storage space to be processed of the decoding node, and the storage space to be processed is located in the second storage unit of the decoding node.
13. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.