Task execution method, product and equipment based on hybrid heterogeneous memory resources
By sorting memory in hybrid heterogeneous memory resources and dividing the model stages by computing layer, the efficient deployment of large-scale natural language processing models under resource constraints is solved, and efficient model inference and execution of natural language processing tasks are achieved.
Patent Information
- Application Number
- CN202510856387.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Under resource-constrained hybrid heterogeneous memory resources, it is difficult for the existing technology to efficiently deploy inference tasks of large-scale natural language processing models, resulting in deployment failure, inefficiency and poor performance.
By determining the performance and communication delay of computing resources and memory resources within the server, sorting memory resources, and phase-dividing the model according to the computing layer, the inference tasks of the model phase are deployed into the sorted memory in turn, and the natural language processing tasks are performed using dynamic flow task scheduling.
Under the constraints of resources, efficient deployment of model inference tasks is achieved, which reduces execution time and power consumption and improves the efficiency of natural language processing tasks.
Smart Images

Figure CN120371536A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a task execution method, product, and device based on hybrid heterogeneous memory resources. Background Art
[0002] With the booming development of natural language processing (NLP) technology, large-scale pre-trained models such as GPT (Generative Pre-trained Transformer) and Bert (Bidirectional Encoder Representations from Transformers) have been widely used in fields such as text classification, information retrieval, machine translation, sentiment analysis, medical question answering, speech recognition, and text generation. Moreover, the number of parameters of pre-trained models for natural language processing also shows a significant growth trend and will continue to grow. This growth trend has also greatly increased the demand for computing power and hybrid heterogeneous memory resources required for model training and inference.
[0003] However, currently, when deploying the inference task of a large model, it is usually deployed in a scenario with limited resources on a heterogeneous device cluster (such as hybrid heterogeneous memory). For example, when the memory of a single computing device in a server, such as the memory of an accelerator, is not sufficient to hold the entire model, the inference task cannot be deployed. It can be seen that currently, when deploying the inference task, the hardware resources in the server, such as the hybrid heterogeneous memory resources under the CPU (Central Processing Unit), are not fully utilized, resulting in deployment failure, as well as problems of low deployment efficiency and poor model inference performance. Therefore, how to efficiently deploy the inference task of a large model under hybrid heterogeneous memory resources including accelerator memory, traditional memory, CXL (Compute Express Link, an open interconnect standard) memory, etc. is an issue that still needs to be further solved in this field. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a task execution method, product, and device based on hybrid heterogeneous memory resources, which can achieve efficient deployment of model inference tasks under hybrid heterogeneous memory resources with limited resources, reduce the time for natural language processing tasks to be executed, reduce power consumption, and improve the efficiency of natural language processing tasks. The specific solutions are as follows: In the first aspect, this application discloses a task execution method based on hybrid heterogeneous memory resources, including: Determine the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources within the current server, obtain the device computing power performance and memory storage performance, and respectively count the communication latency between any two memories in the hybrid heterogeneous memory resources; Sort multiple memories in the hybrid heterogeneous memory resources based on the device computing power performance, memory storage performance, and communication latency to obtain the sorted memories; Divide the target natural language processing model into multiple model stages according to the computing layer; wherein, the number of multiple model stages is the same as the number of computing layers in the target natural language processing model, and the memory space required for each model stage is not greater than the size of any memory in the hybrid heterogeneous memory resources; Deploy the inference tasks of each model stage into the sorted memories in sequence to obtain the deployed model, and use the deployed model to execute natural language processing tasks.
[0005] In a second aspect, the present application discloses a computer program product, including a computer program, which when executed by a processor implements the foregoing task execution method based on hybrid heterogeneous memory resources.
[0006] In a third aspect, the present application discloses an electronic device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the foregoing task execution method based on hybrid heterogeneous memory resources.
[0007] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the foregoing task execution method based on hybrid heterogeneous memory resources.
[0008] It can be seen that the present application first determines the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources within the current server, obtains the device computing power performance and memory storage performance, and respectively counts the communication latency between any two memories in the hybrid heterogeneous memory resources; sorts multiple memories in the hybrid heterogeneous memory resources based on the device computing power performance, memory storage performance, and communication latency to obtain the sorted memories; divides the target natural language processing model into multiple model stages according to the computing layer; wherein, the number of multiple model stages is the same as the number of computing layers in the target natural language processing model, and the memory space required for each model stage is not greater than the size of any memory in the hybrid heterogeneous memory resources; deploys the inference tasks of each model stage into the sorted memories in sequence to obtain the deployed model, and uses the deployed model to execute natural language processing tasks.
[0009] The present application can be applied to the deployment of model inference tasks under hybrid heterogeneous memory resources. First, based on the computing power performance of various computing resources in the server, the storage performance of hybrid heterogeneous memory resources, and the communication latency between any two memories, multiple memories are sorted. Then, the inference tasks of multiple model stages obtained after dividing the model are successively deployed to the sorted memories. Since both the computing resources and memory resources in the server are considered simultaneously, and the communication latency between different memories is combined, the hardware resources in the server can be fully utilized, enabling the efficient deployment of model inference tasks even under resource constraints. At the same time, the execution time of natural language processing tasks is reduced, and the power consumption is decreased, thereby improving the execution efficiency of natural language processing tasks. Additionally, since the number of multiple model stages is the same as the number of computing layers in the natural language processing model, and the memory space required for each model stage is not greater than the size of any memory, it can be ensured that the inference tasks of each stage can be successfully deployed. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0011] Figure 1 It is a flowchart of a task execution method based on hybrid heterogeneous memory resources disclosed in the present application; Figure 2 It is a schematic diagram of a specific hybrid heterogeneous memory resource distribution disclosed in the present application; Figure 3 It is a schematic diagram of a specific model stage division disclosed in the present application; Figure 4 It is a schematic diagram of a specific model inference task deployment disclosed in the present application; Figure 5 It is a schematic diagram of a specific model inference task execution disclosed in the present application; Figure 6 It is a schematic diagram of a specific model inference task execution disclosed in the present application; Figure 7 It is a flowchart of a task execution method based on hybrid heterogeneous memory resources disclosed in the present application; Figure 8 It is a schematic diagram of a specific model inference task deployment disclosed in the present application; Figure 9 It is a schematic diagram of a specific model inference task execution disclosed in the present application; Figure 10 A specific schematic diagram of model inference task execution disclosed in this application; Figure 11 A specific schematic diagram of model inference task execution disclosed in this application. Detailed implementation manners
[0012] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0013] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0014] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be made in conjunction with the accompanying drawings and specific implementation manners.
[0015] The embodiment of the present application discloses a task execution method based on hybrid heterogeneous memory resources. Refer to Figure 1 As shown, the method includes: Step S11: Determine the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources in the current server to obtain the device computing power performance and memory storage performance.
[0016] In this embodiment, first, various computing resources and various hybrid heterogeneous memory resources in the current server are collected, and then the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources are respectively obtained to obtain the corresponding device computing power performance and memory storage performance.
[0017] Among them, the computing resources include, but are not limited to, a central processing unit (i.e., CPU), and acceleration computing devices, such as AI (Artificial Intelligence) devices like GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), FPGA (Field-Programmable Gate Array); the hybrid heterogeneous memory resources include, but are not limited to, the local memory of the central processing unit (such as Local dram, local dynamic random access memory), the memory of other central processing units connected to the above central processing unit, that is, the remote memory (such as Remoto dram, remote dynamic random access memory), the open interconnect standard memory (i.e., CXL memory), and the acceleration device memory located in the acceleration computing device (such as GPU memory), etc.
[0018] Specifically, determining the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources within the current server to obtain the device computing power performance and memory storage performance may include: obtaining the computing power performance of the central processing unit and the acceleration computing device within the current server to obtain the processor computing power performance and the acceleration device computing power performance; collecting the sizes of the local memory of the central processing unit, the remote memory connected to the central processing unit, the open interconnect standard memory, and the acceleration device memory located in the acceleration computing device within the server to obtain the corresponding local memory size, remote memory size, interconnect standard memory size, and acceleration device memory size. In this embodiment, the computing power performance of the CPU and the acceleration computing device (such as GPU) within the current server may be collected first to obtain the CPU computing power performance and the GPU computing power performance, and then the sizes of the local memory of the CPU, the remote memory connected to the CPU, the CXL memory, and the acceleration device memory located in the GPU within the server are obtained to obtain the corresponding local memory size, remote memory size, CXL memory size, and GPU memory size. By collecting the performance of various computing resources and various hybrid heterogeneous memory resources within the server, the hardware resource situation of the current server can be understood, so as to make full use of the hardware resources of the server and reasonably deploy the inference tasks of the natural language processing model.
[0019] Specifically, the local memory is connected to the operating system of the server through a first memory, the remote memory is connected to the operating system through a memory access node and a second memory, the open interconnect standard memory is connected to the operating system through an interconnect standard controller, a data transmission bus, and a proxy node, and the acceleration device memory is connected to the operating system through a device controller, a data transmission bus, and a proxy node. For example, see Figure 2As shown in the figure, the operating system (OS) of the server is located within the central processing unit (i.e., CPU), and the local memory of the central processing unit (i.e., CPU) is connected to the operating system through DDR1 (i.e., Double Data Rate Synchronous Dynamic Random Access Memory); moreover, the remote memory of the central processing unit (i.e., CPU) is connected to the operating system through DDR2 and Non-Uniform Memory Access (NUMA) nodes; in addition, the CXL memory is connected to the operating system through a CXL controller, a data transmission bus, such as a PCIe (peripheral component interconnect express, a high-speed serial computer expansion bus standard) bus, and a proxy node (i.e., a home agent Node); furthermore, the accelerated device memory (such as GPU memory) is connected to the operating system through a device controller, a data transmission bus (i.e., a PCIe bus), and a proxy node.
[0020] Step S12: respectively count the communication latency between any two memories in the hybrid heterogeneous memory resources.
[0021] In this embodiment, after determining the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources, the communication latency between any two memories in the hybrid heterogeneous memory resources is respectively counted.
[0022] In this embodiment, respectively counting the communication latency between any two memories in the hybrid heterogeneous memory resources may specifically include: respectively counting the first communication latency between the accelerated device memory and the open interconnect standard memory in the hybrid heterogeneous memory resources, the second communication latency between the accelerated device memory and the local memory, and the third communication latency between the accelerated device memory and the remote memory; wherein, the first communication latency is the sum of the communication latencies of the accelerated device memory, the device controller, the open interconnect standard memory, and the interconnect standard controller; the second communication latency is the sum of the communication latencies of the accelerated device memory, the device controller, the data transmission bus, the proxy node, the local memory, and the first memory; the third communication latency is the sum of the communication latencies of the accelerated device memory, the device controller, the data transmission bus, the proxy node, the remote memory, the memory access node, and the second memory. For example, see Figure 2As shown in the figure, the average communication latency between the acceleration device memory (such as GPU memory) and the open interconnect standard memory (i.e., CXL memory) in the hybrid heterogeneous memory resources = the communication latency of the acceleration device memory (such as GPU memory) (40 nanoseconds) + the communication latency of the device controller (40 nanoseconds) + the communication latency of the open interconnect standard memory (40 nanoseconds) + the communication latency of the interconnect standard controller (40 nanoseconds) = 160 nanoseconds (minimum). The average communication latency between the acceleration device memory (such as GPU memory) and the local memory = the communication latency of the acceleration device memory (40 nanoseconds) + the communication latency of the device controller (40 nanoseconds) + the communication latency of the data transmission bus (140 nanoseconds) + the communication latency of the proxy node (40 nanoseconds) + the communication latency of the local memory (40 nanoseconds) + the communication latency of the first memory (80 nanoseconds) = 380 nanoseconds. The average communication latency between the acceleration device memory and the remote memory = the communication latency of the acceleration device memory (40 nanoseconds) + the communication latency of the device controller (40 nanoseconds) + the communication latency of the data transmission bus (140 nanoseconds) + the communication latency of the proxy node (40 nanoseconds) + the communication latency of the remote memory (40 nanoseconds) + the communication latency of the memory access node (80 nanoseconds) + the communication latency of the second memory (80 nanoseconds) = 460 nanoseconds. By calculating the communication latency between different memories, the communication efficiency between different memories in the current server can be understood, so as to preferentially use the memory with high communication efficiency, which is beneficial to improving the deployment efficiency of subsequent model inference tasks and the execution speed of natural language processing tasks after model deployment.
[0023] Step S13: Sort multiple memories in the hybrid heterogeneous memory resources based on the device computing power performance, memory storage performance, and communication latency to obtain the sorted memories.
[0024] In this embodiment, after obtaining the device computing power performance, memory storage performance, and the communication latency between any two memories, further, sort multiple memories in the hybrid heterogeneous memory resources based on the device computing power performance, memory storage performance, and the communication latency between any two memories, so as to obtain the sorted memories. That is, sort the memories based on the hardware resources of the server and the communication efficiency between different memories at the same time.
[0025] In a specific implementation manner, the memory order in the sorted memories is the acceleration device memory, the cache memory, and the local memory in sequence. It can be understood that, referring to Figure 2As shown, generally, computing devices with acceleration performance have the highest performance and are the preferred devices for large model inference, such as devices like GPUs and TPUs. However, the memory capacity of devices such as GPUs and TPUs is relatively small, and it is often impossible to deploy the inference tasks of the entire model. Therefore, other memories can also be considered to deploy some inference tasks to other memories, such as CXL memory, the local memory of the CPU, etc. Considering that devices such as GPUs and TPUs have high acceleration computing performance, the memory of devices such as GPUs and TPUs can be ranked first, and then CXL memory and the local memory of the CPU can be selected in sequence.
[0026] It can be understood that CXL memory provides a new way for the CPU to access device memory and for the device to access CPU memory. The CPU can access the computing device memory in the same way as accessing local memory, thus greatly improving the efficiency of data exchange between the device and the host, and accelerating the efficiency of deploying the model in the hybrid heterogeneous memory resources.
[0027] In this embodiment, multiple memories in the hybrid heterogeneous memory resources are sorted based on the device computing power performance, memory storage performance, and communication latency to obtain the sorted memory. Specifically, it may include: determining the memory for caching the inference tasks of different model stages from the hybrid heterogeneous memory resources based on the device computing power performance, memory storage performance, and communication latency to obtain the cached memory; sorting the multiple memories in the hybrid heterogeneous memory resources based on the cached memory to obtain the sorted memory. In this embodiment, first, the memory for caching the inference tasks of different model stages is determined from the hybrid heterogeneous memory resources based on the device computing power performance of different computing devices, the memory storage performance of different memories, and the communication latency between different memories to obtain the cached memory, and then the multiple memories in the hybrid heterogeneous memory resources are sorted based on the determined cached memory to obtain the sorted memory. By selecting the memory for caching the inference tasks of different model stages to obtain the cached memory, dynamic scheduling of the inference tasks in multiple memories can be achieved, and the inference tasks to be executed are preferentially cached during the task execution process, thereby improving the execution speed of natural language processing tasks. In a specific implementation manner, the memory with the minimum communication latency with the memory of acceleration devices such as GPUs and TPUs can be used as the cache. Of course, the cached memory can also be selected according to other selection rules.
[0028] Specifically, determining the memory for caching the inference tasks of different model stages from the hybrid heterogeneous memory resources based on the device computing power performance, memory storage performance, and communication latency to obtain the cached memory may include: determining the highest computing power performance among the processor computing power performance and the accelerator device computing power performance to obtain the accelerator device computing power performance; judging whether the accelerator computing device can deploy the entire target natural language processing model based on the size of the accelerator device memory corresponding to the accelerator device computing power performance; if the accelerator computing device cannot deploy the entire target natural language processing model, determining the memory with the smallest communication latency with the accelerator device memory based on the communication latency to obtain the open interconnect standard memory, and using the open interconnect standard memory as the cached memory for caching the inference tasks of different model stages. In this embodiment, first, the highest computing power performance among the processor computing power performance (i.e., CPU computing power performance) and the accelerator device computing power performance (such as GPU computing power performance) is determined to obtain the accelerator device computing power performance (i.e., GPU computing power performance). Then, based on the size of the accelerator device memory (i.e., the capacity of the memory) corresponding to the accelerator device computing power performance (i.e., GPU computing power performance), it is judged whether the corresponding accelerator computing device (i.e., GPU device) can deploy the entire target natural language processing model (such as the GPT model based on the Transformer architecture); if the memory capacity of the accelerator computing device (i.e., GPU device) cannot deploy the entire target natural language processing model, the memory with the smallest communication latency with the accelerator device memory (i.e., GPU device memory) is determined based on the communication latency between different memories to obtain the open interconnect standard memory (such as CXL memory), and then the open interconnect standard memory (i.e., CXL memory) is used as the cached memory for caching the inference tasks of different model stages.
[0029] In this embodiment, it is determined whether the acceleration computing device can deploy the entire target natural language processing model based on the size of the acceleration device memory corresponding to the computing power performance of the acceleration device. Specifically, it may include: counting the data volume of the entire inference task corresponding to the target natural language processing model to obtain the size of the model task data volume; determining whether the size of the acceleration device memory corresponding to the computing power performance of the acceleration device is greater than the size of the model task data volume; if the size of the acceleration device memory is greater than the size of the model task data volume, it is determined that the acceleration computing device can deploy the entire target natural language processing model; if the size of the acceleration device memory is not greater than the size of the model task data volume, it is determined that the acceleration computing device cannot deploy the entire target natural language processing model. In this embodiment, when determining whether the acceleration computing device can deploy the model, the data volume of the entire inference task corresponding to the target natural language processing model can be counted first, and then based on the relationship between the counted size of the model task data volume and the size of the acceleration device memory, it is determined whether the entire model can be deployed. Specifically, if the size of the acceleration device memory is greater than the size of the model task data volume, it is determined that the acceleration computing device (such as a GPU device) can deploy the entire model; if the size of the acceleration device memory is not greater than the size of the model task data volume, it is determined that the acceleration computing device (such as a GPU device) cannot deploy the entire model. Among them, the model task data volume specifically refers to the sum of the parameter quantities of different layers of the model.
[0030] Step S14: Divide the target natural language processing model into multiple model stages according to the computing layer; where the number of the multiple model stages is the same as the number of computing layers in the target natural language processing model, and the memory space required for each model stage is not greater than the size of any one of the hybrid heterogeneous memory resources.
[0031] In this embodiment, referring to Figure 3 As shown, after sorting the multiple memories in the hybrid heterogeneous memory resources, the target natural language processing model can be divided into 8 model stages in the way that one computing layer corresponds to one model stage, and the number of the model stages is the same as the number of computing layers of the target natural language processing model. It should be noted that the memory space required for each model stage is not greater than the capacity size of any one of the hybrid heterogeneous memory resources, that is, the size of the memory space required for each divided model stage is not higher than the size of any one memory. In this way, it can be ensured that any model stage can be deployed to any memory, such as the local memory of the CPU, the CXL memory, the GPU memory, etc.
[0032] Step S15: Sequentially deploy the inference tasks of each model stage to the sorted memory to obtain a deployed model, and use the deployed model to execute the natural language processing task.
[0033] In this embodiment, after the model stage division, the inference tasks corresponding to each model stage can be sequentially deployed to the sorted memory, and then the deployed model is used to execute the natural language processing tasks sent by the user side.
[0034] Among them, the natural language processing tasks include, but are not limited to, text classification, information retrieval, machine translation, intelligent customer service Q&A, sentiment analysis, medical Q&A, voice assistants, search engines, text summarization, text classification, text generation, etc.
[0035] It can be understood that the model consists of multiple layers, and each layer consists of parameters. Therefore, the model deployment is essentially writing the parameters of the corresponding layer into the memory.
[0036] In this embodiment, using the deployed model to execute natural language processing tasks may specifically include: when it is monitored that a natural language processing task needs to be executed, obtaining the natural language data sequence to be processed; sequentially inputting the data in the natural language data sequence into the target natural language processing model to perform pipelining processing on each data in the natural language data sequence in the manner of the first dynamic pipelining task scheduling to obtain the natural language processing result; where the first dynamic pipelining task scheduling method is to dynamically schedule the inference tasks of different model stages to different memories, so that each data in the natural language data sequence is executed in a pipelining mode on different computing devices in the computing resources. See Figure 4 As shown, the deployed model specifically Figure 3 divides the inference tasks of 8 model stages into 3 different memories; among them, the inference tasks corresponding to inference stage 1 are deployed to the accelerator device memory 1 (such as GPU memory 1), and the inference tasks corresponding to inference stage 2 are deployed to the open interconnect standard memory 1 (such as CXL memory 1), and the inference tasks corresponding to inference stages 3 to 8 (i.e., other inference stages) to are deployed to the local memory of the central processing unit 0 (i.e., the local memory of CPU0), and the central processing unit 0 (i.e., CPU0) is also connected to other central processing units (such as central processing unit 1) through UPI (Ultra Path Interconnect), and the local memory of central processing unit 1 can be used as the remote memory of central processing unit 0. That is, after the initial deployment, the first-stage tasks are deployed in the GPU memory, the second-stage tasks are deployed in the CXL memory, and the other stages are deployed in the local memory of the CPU.
[0037] In this embodiment, when it is detected that a natural language processing task (such as a medical Q&A) needs to be executed, first obtain the natural language data sequence to be processed (such as consultation information for a certain disease); then, input the data in the natural language data sequence into the target natural language processing model in turn, so as to perform pipelining processing on each data in the natural language data sequence in the manner of dynamic pipelining task scheduling, and obtain the corresponding natural language processing result. It should be noted that during the data processing process, the inference tasks of different model stages will be dynamically scheduled to different memories, so that each data in the natural language data sequence is executed in a pipelining mode on different computing devices (such as CPU, GPU, etc.) in the computing resources. Specifically, see Figure 4 As shown, process the data 1 in the natural language data sequence through the inference task in the acceleration device memory 1 (such as GPU memory 1), and during the processing process, load and cache the inference task of the next stage, that is, the inference task through the open interconnect standard memory 1 (such as CXL memory 1); then, see As shown, release the used inference task in the acceleration device memory 1 (such as GPU memory 1) to the local memory of the central processing unit 0 (i.e., the local memory of CPU0), and load the inference task cached in the open interconnect standard memory 1 (i.e., CXL memory 1) into the acceleration device memory 1 (such as GPU memory 1) to continue processing the data 1. It should be noted that during the process of preferentially processing the data 1, the corresponding processing operation is also performed on the data 2 through the idle inference task in the local memory of the central processing unit 0 (i.e., the local memory of CPU0). See Figure 5 As shown, process the data 2 through the inference task in the local memory (i.e., the local memory of CPU0). Since the data 1, data 2, etc. in the sequence are input in chronological order, the data input first will be preferentially processed in order. For example, preferentially process the data 1. If there is an idle inference task in the local memory and this inference task happens to be the inference task (such as the inference task required by other data (such as data 2) currently, then process the corresponding data (i.e., data 2). If the inference task required by the current other data (such as data 2) is occupied by the data 1, then preferentially process the data 1, and after the data 1 is processed and released back to the local memory (i.e., the local memory of CPU0), then use the corresponding inference task. Load it into the acceleration device memory 1 (such as GPU memory 1) to continue processing the data 1 operation. It should be noted that during the process of preferentially processing the data 1, the corresponding processing operation is also performed on the data 2 through the idle inference task in the local memory of the central processing unit 0 (i.e., the local memory of CPU0). See Figure 5 As shown, process the data 2 through the inference task in the local memory (i.e., the local memory of CPU0). Since the data 1, data 2, etc. in the sequence are input in chronological order, the data input first will be preferentially processed in order. For example, preferentially process the data 1. If there is an idle inference task in the local memory and this inference task happens to be the inference task (such as the inference task required by other data (such as data 2) currently, then process the corresponding data (i.e., data 2). If the inference task required by the current other data (such as data 2) is occupied by the data 1, then preferentially process the data 1, and after the data 1 is processed and released back to the local memory (i.e., the local memory of CPU0), then use the corresponding inference task. ), then process the corresponding data (i.e., data 2). If the inference task required by the current other data (such as data 2) is occupied by the data 1, then preferentially process the data 1, and after the data 1 is processed and released back to the local memory (i.e., the local memory of CPU0), then use the corresponding inference task.
[0038] In addition, it should be noted that, see Figure 6 As shown, when the acceleration device memory 1 (such as GPU memory 1) executes the inference task of the last stage (i.e., the inference task When it is time to process the next data (such as Data 2), since Data 2 has already undergone processing for 2 inference tasks, the inference task for Stage 3 needs to be executed currently. Therefore, the acceleration device memory 1 (such as GPU memory 1) needs to load the currently pending inference task from the Open Interconnect Standard Memory 1 (i.e., CXL memory 1). To continue processing Data 2, at the same time, the Open Interconnect Standard Memory 1 (i.e., CXL memory 1) will continue to load the inference task for Stage 4. For pipelined processing of Data 2. While processing Data 2, corresponding processing operations are also performed on Data 3 through the idle inference tasks (such as Inference Task ) in the local memory of the central processing unit 0 (i.e., the local memory of CPU0). Further, the above process is repeated until the last data in the natural language data sequence is completely processed, thereby outputting the final natural language processing result. By performing pipelined processing on each data in the natural language data sequence through dynamic pipelined task scheduling, different stage tasks can be efficiently transferred and deployed in different memories, enabling different data to be executed in a pipelined manner on the central processing unit and the acceleration device, thus making full use of each hardware resource in the server and improving the execution speed of the natural language processing task.
[0039] In addition, during the process of sequentially deploying the inference tasks of each model stage to the sorted memory, it specifically further includes: respectively counting the data volumes of the inference tasks corresponding to any two or more adjacent model stages to obtain multiple inference task statistical results; respectively determining whether each inference task statistical result is not greater than the size of any memory in the hybrid heterogeneous memory resources. If not greater, then select the corresponding two or more model stages according to the execution order of the inference tasks, and deploy the selected two or more model stages together to a single memory in the sorted memory to obtain the deployed model. That is, identify multiple adjacent inference tasks and deploy them to a single memory. In this way, the number of inference task scheduling during the subsequent execution of the natural language processing task can be reduced, thereby further improving the execution efficiency of the subsequent natural language processing task.
[0040] It can be seen that the embodiments of the present application can be applied to the deployment of model inference tasks under hybrid heterogeneous memory resources. First, based on the computing power performance of various computing resources in the server, the storage performance of the hybrid heterogeneous memory resources, and the communication latency between any two memories, multiple memories are sorted. Then, the inference tasks of multiple model stages obtained after dividing the model are sequentially deployed to the sorted memories. Since both the computing resources and memory resources in the server are considered, and the communication latency between different memories is combined, the hardware resources in the server can be fully utilized, enabling the efficient deployment of model inference tasks even under resource constraints. At the same time, the execution time of natural language processing tasks is reduced, the power consumption is decreased, and thus the execution efficiency of natural language processing tasks is improved. In addition, since the number of multiple model stages is the same as the number of computing layers in the natural language processing model, and the memory space required for each model stage is not greater than the size of any memory, it can be ensured that the inference tasks of each stage can be successfully deployed.
[0041] The embodiments of the present application disclose a specific task execution method based on hybrid heterogeneous memory resources. Refer to Figure 7 as shown, the method includes: Step S21: Determine the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources in the current server to obtain device computing power performance and memory storage performance; the computing resources include a central processing unit and an acceleration computing device; the hybrid heterogeneous memory resources include the local memory of the central processing unit, the remote memory connected to the central processing unit, the open interconnect standard memory, and the acceleration device memory located in the acceleration computing device.
[0042] Step S22: Statistically calculate the communication latency between any two memories in the hybrid heterogeneous memory resources respectively.
[0043] Step S23: Sort multiple memories in the hybrid heterogeneous memory resources based on the device computing power performance, memory storage performance, and communication latency to obtain sorted memories; the memory order in the sorted memories is acceleration device memory, cache memory, local memory, and remote memory in sequence.
[0044] Step S24: Divide the target natural language processing model into stages according to the computing layer to obtain multiple model stages; wherein, the number of multiple model stages is the same as the number of computing layers in the target natural language processing model, and the memory space required for each model stage is not greater than the size of any memory in the hybrid heterogeneous memory resources.
[0045] Step S25: Sequentially deploy the inference tasks of each model stage to the sorted memories to obtain a deployed model.
[0046] Refer to Figure 8As shown, when initializing and deploying the inference task, according to the server hardware resources and the communication latency between different memories, the inference tasks corresponding to Inference Phase 1 can be deployed to the Accelerator Device Memory 1 (such as GPU Memory 1), and the inference tasks corresponding to Inference Phase 2 can be deployed to the Open Interconnect Standard Memory 1 (such as CXL Memory 1). The inference tasks corresponding to Inference Phases 3 to 8 (i.e., other inference phases) to can be deployed to the local memory of Central Processing Unit 0 (i.e., the local memory of CPU0), and the inference tasks and can be deployed to the remote memory of Central Processing Unit 0.
[0047] Step S26: When it is monitored that a natural language processing task needs to be executed, obtain the multi-batch natural language data sequences to be processed.
[0048] Step S27: Input the multi-batch natural language data sequences into the target natural language processing model to perform parallel processing on the multi-batch natural language data sequences in the manner of the second dynamic pipelining task scheduling, so as to obtain the natural language processing results. Among them, the manner of the second dynamic pipelining task scheduling is to dynamically schedule the inference tasks of different model phases to different memories, so that each data in the same batch of natural language data sequences is executed in a pipeline mode on different computing devices of the computing resources, and different batches of natural language data sequences are executed in parallel on different computing devices of the computing resources.
[0049] In this embodiment, if the data corresponding to the natural language processing task to be processed is a multi-batch natural language data sequence, at this time, in order to improve the data processing speed, the multi-batch natural language data sequences can be input into the target natural language processing model together, so as to perform parallel processing on the multi-batch natural language data sequences in the manner of dynamic pipelining task scheduling, thereby obtaining the natural language processing results. It should be noted that the specific processing process of a single batch of natural language data sequences is the same as that of Figures 4 to 6 , except that the addition of the remote memory is added. Specifically, as shown in Figure 8 , for the processing of different batches of natural language data sequences, a parallel processing method can be adopted. For example, when processing different batches of sequences, Accelerator Device Memory 1 (such as GPU Memory 1) and Open Interconnect Standard Memory 1 (i.e., CXL Memory 1) can be used to process the first batch of natural language data sequences. At the same time, Accelerator Device Memory 2 (such as GPU Memory 2) and Open Interconnect Standard Memory 2 (i.e., CXL Memory 2) can also be used to process the second batch of natural language data sequences.
[0050] In this embodiment, during the parallel processing of multiple batches of natural language data sequences by adopting the second dynamic pipelining task scheduling method, the following steps may also be included: dynamically releasing the inference tasks of the currently completed model stage in the acceleration computing device and loading the inference tasks of the next model stage; when it is detected that the acceleration computing device has calculated the inference tasks of the last model stage, determining the inference tasks of the model stage corresponding to the data to be processed in the next batch of natural language data sequences and loading the inference tasks corresponding to the data to be processed into the cache memory. Specifically, refer to Figure 8 As shown, first, the data 1 in the first batch of natural language data sequences is processed through the inference tasks in the acceleration device memory 1 (such as GPU memory 1), and during the processing, the inference tasks of the next stage, that is, the inference tasks are loaded and cached through the open interconnect standard memory 1 (such as CXL memory 1). Then, refer to As shown, the inference tasks Figure 9 used up in the acceleration device memory 1 (such as GPU memory 1) are released to the remote memory of the central processing unit 0 (i.e., the local memory of CPU1), and the inference tasks cached in the open interconnect standard memory 1 (i.e., CXL memory 1) are loaded into the acceleration device memory 1 (such as GPU memory 1) to continue the processing operation on the data 1. At the same time, the data 2 is also processed through the idle inference tasks (such as the inference tasks in the local memory of the central processing unit 0 (i.e., the local memory of CPU0). It can be understood that the performance of the acceleration computing device (such as the GPU device) is much higher than that of other computing devices (such as the CPU). Therefore, the acceleration computing device (such as the GPU device) will return the completed stage tasks to the CPU memory at one time for processing other data, so as to achieve the effect of parallel processing and reach a higher data processing speed.
[0051] At the same time, refer to Figure 10 As shown, when the acceleration computing device 1 (such as GPU device 1) calculates the inference tasks of the last stage, it is judged which model stage the current batch of data (i.e., data 2) in the CPU has been calculated to, and the next stage (i.e., the inference tasks ) is sent to the open interconnect standard memory 1 (i.e., CXL memory 1) for caching, and the current calculation result (i.e., the intermediate calculation result) of the data 2 is sent to the acceleration computing device 1 (such as GPU device 1) so that the acceleration computing device 1 (such as GPU device 1) can directly receive the inference tasks corresponding to stage 3 from the CXL memory 1 next time., and perform corresponding calculations on the received current calculation results (i.e., intermediate calculation results). Since the communication efficiency between CXL memory and acceleration computing devices (such as GPU devices) is the highest, for example, higher than the memory communication efficiency of the CPU, the data processing flow that interacts with CXL memory is faster.
[0052] In addition, as shown in Figure 10 , when the acceleration device memory 1 (such as GPU memory 1) executes the inference task at the last stage (i.e., stage 10) (i.e., the inference task ), the processing of the next data (such as data 2) is required. Since data 2 has undergone the processing of 2 inference tasks and the inference task of stage 3 needs to be executed currently , the acceleration device memory 1 (such as GPU memory 1) needs to load the currently to-be-executed inference task from the open interconnect standard memory 1 (i.e., CXL memory 1) to continue processing data 2. At the same time, the open interconnect standard memory 1 (i.e., CXL memory 1) will continue to load the inference task of stage 4 , so as to process data 2 in a pipelined mode. And while processing data 2, corresponding processing operations can also be performed on data 3 through the idle inference tasks (such as inference task ) in the local memory of the central processing unit 0 (i.e., the local memory of CPU0). Repeat the above process until all data in the multi-batch natural language data sequence are processed, so as to obtain the final natural language processing result.
[0053] Among them, for the more specific processing processes of the above steps S21 to S24, S26, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0054] It can be seen that in the embodiments of the present application, first, multiple memories in the hybrid heterogeneous memory resources are sorted based on the performance of various hardware resources in the server and the communication latency between different memories. Then, the inference tasks of each model stage are sequentially deployed to the sorted memories. When it is detected that a natural language processing task needs to be executed, multiple batches of natural language data sequences are input into the target natural language processing model, so as to perform parallel processing on the multiple batches of natural language data sequences in a dynamic pipelined task scheduling manner, thereby obtaining the natural language processing result. Since the dynamic pipelined task scheduling method dynamically schedules the inference tasks of different model stages to different memories, each data in the same batch of natural language data sequences is executed in a pipelined mode on different computing devices of the computing resources, and different batches of natural language data sequences are executed in parallel on different computing devices of the computing resources. Therefore, the execution speed of the natural language processing task can be improved. In addition, by completing the scheduling of the divided inference tasks between different memories and the mutual cooperation of different computing devices through the dynamic pipelined task scheduling method, the deployment and execution of the model inference task in the case of extremely few computing resources can be realized, achieving the effects of high resource utilization rate and high inference performance. And the data communication between different memories can be pipelined, which can achieve lower latency, thereby improving the data processing efficiency and further improving the natural language processing efficiency. In addition, based on the performance of various hardware resources and the communication latency between different memories, the precise division of the model tasks can be realized, which can reduce the task execution time and the power consumption of the task execution, thereby solving the problems of low communication efficiency of data-intensive tasks and difficult utilization of heterogeneous memory resources. Moreover, more and more heterogeneous memories and bandwidth can be fully utilized.
[0055] Specifically, as shown in Figure 11 Accelerated computing devices (such as GPUs) and central processing units (i.e., CPUs) can process the data in different batches of natural language data sequences in parallel. After the accelerated computing device (such as GPU) finishes processing the inference task of stage 1 of data 1, it continues to process the inference task of stage 2 of data 1. At the same time, the accelerated computing device (such as GPU) will release the inference task of stage 1 to the central processing unit (i.e., CPU), and the central processing unit (i.e., CPU) starts to process the inference task of stage 1 of data 2. It should be noted that the processing of the same batch of data is preferably selected to be processed on the accelerated computing device (such as GPU). That is, when the accelerated computing device (such as GPU) finishes processing a batch of data, it will dynamically obtain the next batch of data (such as data 2) and the inference task of the current computing stage 3 from the central processing unit (i.e., CPU) memory for processing. Through the above dynamic scheduling strategy, extremely efficient multi-data parallel processing can be realized, thereby improving the task execution efficiency.
[0056] Embodiments of the present application further provide a task execution device based on hybrid heterogeneous memory resources. For the descriptions of the features in the corresponding embodiments of the task execution device based on hybrid heterogeneous memory resources, reference may be made to the relevant descriptions in the corresponding embodiments of the task execution method based on hybrid heterogeneous memory resources, which will not be elaborated herein one by one.
[0057] Embodiments of the present application further provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the task execution method based on hybrid heterogeneous memory resources.
[0058] Embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the task execution method based on hybrid heterogeneous memory resources when running.
[0059] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0060] Embodiments of the present application further provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the task execution method based on hybrid heterogeneous memory resources.
[0061] Embodiments of the present application further provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the task execution method based on hybrid heterogeneous memory resources.
[0062] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0063] The above has introduced in detail the task execution method, product, device, and storage medium based on hybrid heterogeneous memory resources provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A task execution method based on hybrid heterogeneous memory resources, characterized in that, Including: Determine the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources within the current server, and obtain the device computing power performance and memory storage performance; Respectively count the communication latency between any two memories in the hybrid heterogeneous memory resources; Sort multiple memories in the hybrid heterogeneous memory resources based on the device computing power performance, the memory storage performance, and the communication latency, and obtain the sorted memories; Divide the target natural language processing model by calculation layer to obtain multiple model stages; wherein, the number of the multiple model stages is the same as the number of calculation layers in the target natural language processing model, and the memory space required for each model stage is not greater than the size of any memory in the hybrid heterogeneous memory resources; Deploy the inference tasks of each model stage to the sorted memories in sequence to obtain the deployed model, and use the deployed model to execute natural language processing tasks.
2. The task execution method based on hybrid heterogeneous memory resources according to claim 1, wherein The computing resources include a central processing unit and an acceleration computing device; the hybrid heterogeneous memory resources include the local memory of the central processing unit, the remote memory connected to the central processing unit, the open interconnect standard memory, and the acceleration device memory located in the acceleration computing device.
3. The task execution method based on hybrid heterogeneous memory resources according to claim 2, wherein, The determining the computing power performance of various computing resources and the storage performance of various hybrid heterogeneous memory resources within the current server, and obtaining the device computing power performance and memory storage performance includes: Obtain the computing power performance of the central processing unit and the acceleration computing device within the current server, and obtain the processor computing power performance and the acceleration device computing power performance; Collect the sizes of the local memory of the central processing unit, the remote memory connected to the central processing unit, the open interconnect standard memory, and the acceleration device memory located in the acceleration computing device within the server, and obtain the corresponding local memory size, remote memory size, interconnect standard memory size, and acceleration device memory size.
4. The task execution method based on hybrid heterogeneous memory resources according to claim 3, wherein The sorting multiple memories in the hybrid heterogeneous memory resources based on the device computing power performance, the memory storage performance, and the communication latency, and obtaining the sorted memories includes: Determine the memories for caching the inference tasks of different model stages from the hybrid heterogeneous memory resources based on the device computing power performance, the memory storage performance, and the communication latency, and obtain the caching memories; Sort multiple memories in the hybrid heterogeneous memory resources based on the caching memories, and obtain the sorted memories.
5. The task execution method based on hybrid heterogeneous memory resources according to claim 4, wherein The determining the memories for caching the inference tasks of different model stages from the hybrid heterogeneous memory resources based on the device computing power performance, the memory storage performance, and the communication latency, and obtaining the caching memories includes: Determine the highest computing power performance among the processor computing power performance and the acceleration device computing power performance, and obtain the acceleration device computing power performance; Judge whether the acceleration computing device can deploy the entire target natural language processing model based on the size of the acceleration device memory corresponding to the acceleration device computing power performance; If the acceleration computing device cannot deploy the entire target natural language processing model, determine the memory with the minimum communication latency with the acceleration device memory based on the communication latency, obtain the open interconnect standard memory, and use the open interconnect standard memory as the cache memory for caching inference tasks in different model stages.
6. The task execution method based on hybrid heterogeneous memory resources according to claim 5, wherein The determining whether the acceleration computing device can deploy the entire target natural language processing model based on the size of the acceleration device memory corresponding to the computing power performance of the acceleration device includes: Count the data volume of the entire inference task corresponding to the target natural language processing model to obtain the size of the model task data volume; Determine whether the size of the acceleration device memory corresponding to the computing power performance of the acceleration device is greater than the size of the model task data volume; If the size of the acceleration device memory is greater than the size of the model task data volume, determine that the acceleration computing device can deploy the entire target natural language processing model; If the size of the acceleration device memory is not greater than the size of the model task data volume, determine that the acceleration computing device cannot deploy the entire target natural language processing model.
7. The task execution method based on hybrid heterogeneous memory resources according to claim 4, wherein The memory order in the sorted memory is the acceleration device memory, the cache memory, and the local memory in sequence; Alternatively, the memory order in the sorted memory is the acceleration device memory, the cache memory, the local memory, and the remote memory in sequence.
8. The task execution method based on hybrid heterogeneous memory resources according to claim 7, wherein The using the deployed model to perform natural language processing tasks includes: When it is detected that a natural language processing task needs to be performed, obtain the natural language data sequence to be processed; Input the data in the natural language data sequence into the target natural language processing model in sequence, and perform pipelining processing on the data in the natural language data sequence in the first dynamic pipelining task scheduling manner to obtain the natural language processing result; Among them, the first dynamic pipelining task scheduling manner is to dynamically schedule inference tasks in different model stages to different memories, so that each data in the natural language data sequence is executed in a pipeline mode on different computing devices in the computing resources.
9. The task execution method based on hybrid heterogeneous memory resources according to claim 7, wherein The using the deployed model to perform natural language processing tasks includes: When it is detected that a natural language processing task needs to be performed, obtain the multi-batch natural language data sequence to be processed; Input the multi-batch natural language data sequences into the target natural language processing model, and perform parallel processing on the multi-batch natural language data sequences in the second dynamic pipelining task scheduling manner to obtain the natural language processing result; Among them, the second dynamic pipelining task scheduling manner is to dynamically schedule inference tasks in different model stages to different memories, so that each data in the same batch of natural language data sequences is executed in a pipeline mode on different computing devices in the computing resources, and different batches of natural language data sequences are executed in parallel on different computing devices in the computing resources.
10. The task execution method based on hybrid heterogeneous memory resources according to claim 9, wherein, During the process of performing parallel processing on the multi-batch natural language data sequences in the second dynamic pipelining task scheduling manner, it further includes: Dynamically release the inference tasks of the currently completed model stage in the acceleration computing device and load the inference tasks of the next model stage. When it is detected that the acceleration computing device has calculated the inference tasks of the last model stage, determine the inference tasks of the model stage corresponding to the data to be processed in the next batch of the natural language data sequence, and load the inference tasks corresponding to the data to be processed into the cache memory.
11. The task execution method based on hybrid heterogeneous memory resources according to any one of claims 2 to 10, characterized in that, The local memory is connected to the operating system of the server through a first memory, the remote memory is connected to the operating system through a memory access node and a second memory, the open interconnect standard memory is connected to the operating system through an interconnect standard controller, a data transmission bus, and a proxy node, and the acceleration device memory is connected to the operating system through a device controller, the data transmission bus, and the proxy node.
12. The task execution method based on hybrid heterogeneous memory resources according to claim 11, wherein The method of respectively counting the communication delays between any two memories in the hybrid heterogeneous memory resources includes: Respectively count the first communication delay between the acceleration device memory and the open interconnect standard memory, the second communication delay between the acceleration device memory and the local memory, and the third communication delay between the acceleration device memory and the remote memory in the hybrid heterogeneous memory resources; Wherein, the first communication delay is the sum of the communication delays of the acceleration device memory, the device controller, the open interconnect standard memory, and the interconnect standard controller; The second communication delay is the sum of the communication delays of the acceleration device memory, the device controller, the data transmission bus, the proxy node, the local memory, and the first memory; The third communication delay is the sum of the communication delays of the acceleration device memory, the device controller, the data transmission bus, the proxy node, the remote memory, the memory access node, and the second memory.
13. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the task execution method based on hybrid heterogeneous memory resources according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the task execution method based on hybrid heterogeneous memory resources according to any one of claims 1 to 12.
15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the task execution method based on hybrid heterogeneous memory resources according to any one of claims 1 to 12.
Citation Information
Patent Citations
Hybrid heterogeneous memory-oriented tagged data and job scheduling method and system
CN109753246A
Dynamic data planning method
CN111158903A
Multi-source heterogeneous environmental data asset construction method based on computing power network
CN115858829A
Text data reasoning method and device based on hybrid expert model
CN119443279A
Memory allocation method and device for model in heterogeneous system, medium and product
CN119621356A