Inference task processing method, distributed system, computing device, and computer-readable storage medium
By dynamically identifying and switching task processing models in a distributed system, the inefficiency problems caused by resource shortage and model mismatch are solved, efficient utilization and rapid response of resources are achieved, and the overall performance and reliability of the distributed system are improved.
Patent Information
- Application Number
- CN202510330379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In distributed systems, resource shortage and mismatched task inference models lead to inefficient inference tasks processing. How to intelligently schedule task inference models to optimize resource allocation and improve efficiency.
By dynamically identifying the task processing model type of inference requests on nodes of distributed systems, automatically uninstalling the no longer needed models and installing models suitable for the current request, ensuring efficient utilization and rapid response to resources.
Optimize resource allocation, improve the inference task processing efficiency of a single node, and thus improve the overall performance and reliability of the distributed system.
Smart Images

Figure CN119847770B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of artificial intelligence, and particularly to a method for processing inference tasks, a distributed system, a computing device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence and machine learning technologies, using task processing models to execute inference tasks has been widely applied in many fields due to its advantages of efficiently processing complex data, quickly responding, and adapting to diverse application scenarios, such as image recognition, natural language processing, intelligent recommendation systems, etc.
[0003] Currently, with the expansion of the user base and the diversification of application scenarios, deploying task inference models on a distributed system and using multiple nodes and the parallel computing units on the nodes can achieve large-scale concurrent processing and improve the overall performance and reliability.
[0004] However, for the growing and dynamically changing inference requests, the resources of the parallel computing units running the task inference models have become relatively scarce. At the same time, inference requests often correspond to different inference tasks and application scenarios, and using a single or mismatched task inference model is likely to cause resource waste and reduce the efficiency of processing inference tasks. How to intelligently schedule the instances of running task inference models and optimize resource allocation to improve the efficiency of processing inference tasks is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the embodiments of this specification provide a method for processing inference tasks. One or more embodiments of this specification also relate to another method for processing inference tasks, a distributed system, a computing device, another computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0006] According to the first aspect of the embodiments of this specification, a method for processing inference tasks is provided, which is applied to the first node of a distributed system. The first node is any one of at least one node of the distributed system, and a parallel computing unit is deployed on the first node, including:
[0007] In response to an execution instruction of a current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unload the task processing model corresponding to the previous inference request from the parallel computing unit and install the task processing model corresponding to the current inference request;
[0008] Run the task processing model on the parallel computing unit to execute the current inference request.
[0009] According to the second aspect of the embodiments of the present specification, another method for processing inference tasks is provided, which is applied to the management module of a distributed system. The distributed system further includes at least one node, and the method includes:
[0010] Obtain multiple inference requests for different types of task processing models, and distribute the multiple inference requests to at least one node. Among them, the first node is any one of the at least one node, and the following steps are performed on the first node:
[0011] In response to the execution instruction of the current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unload the task processing model corresponding to the previous inference request from the parallel computing unit, and install the task processing model corresponding to the current inference request;
[0012] Run the task processing model on the parallel computing unit to execute the current inference request.
[0013] According to the third aspect of the embodiments of the present specification, a distributed system is provided, including a management module and at least one node. A parallel computing unit is deployed on any node, and the first node is any one of the at least one node;
[0014] The management module is configured to obtain multiple inference requests for different types of task processing models, and distribute the multiple inference requests to at least one node;
[0015] The first node is configured to, in response to the execution instruction of the current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unload the task processing model corresponding to the previous inference request from the parallel computing unit, and install the task processing model corresponding to the current inference request, run the task processing model on the parallel computing unit, and execute the current inference request.
[0016] According to the fourth aspect of the embodiments of the present specification, a computing device is provided, including:
[0017] A memory and a processor;
[0018] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the inference task processing method in the first aspect are implemented.
[0019] According to the fifth aspect of the embodiments of the present specification, a computing device is provided, including:
[0020] A memory and a processor;
[0021] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the inference task processing method in the second aspect above are implemented.
[0022] According to the sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned inference task processing method are implemented.
[0023] According to the seventh aspect of the embodiments of this specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned inference task processing method are implemented.
[0024] In one embodiment of this specification, on the first node, the task processing model type of the current inference request is dynamically identified, which is the same as the task processing model type corresponding to the previous inference request. When different types of task processing models are detected, the task inference models that are no longer needed are automatically uninstalled, and at the same time, the task inference model suitable for the current inference request is installed, thereby ensuring the efficient use of resources and fast response, effectively solving the problem of low efficiency caused by resource shortage and the use of mismatched task inference models, optimizing resource allocation, improving the efficiency of the inference task processing of a single node, and further improving the overall performance and reliability of the distributed system. Description of the Drawings
[0025] Figure 1 is a flowchart of an inference task processing method provided by an embodiment of this specification;
[0026] Figure 2 is a flowchart of another inference task processing method provided by an embodiment of this specification;
[0027] Figure 3 is a flowchart of the processing process of an inference task processing method applied to large language model inference provided by an embodiment of this specification;
[0028] Figure 4 is a schematic structural diagram of a distributed system provided by an embodiment of this specification;
[0029] Figure 5 is a schematic architecture diagram of a distributed system provided by an embodiment of this specification;
[0030] Figure 6 is a structural block diagram of a computing device provided by an embodiment of this specification;
[0031] Figure 7It is a structural block diagram of another computing device provided by an embodiment of this specification. Detailed implementation manners
[0032] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0033] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0034] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0035] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0036] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0037] When a large model is actually applied, only a small number of samples are needed to fine-tune the pre-trained model for application to different tasks. Large models can be widely applied in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0038] First, the noun terms related to one or more embodiments of this specification are explained.
[0039] Inference task: In the fields of machine learning and artificial intelligence, an inference task refers to the process of using a trained model to make predictions or classifications on new data. Different from model training, inference focuses on applying the model to make quick and accurate judgments and is widely used in fields such as image recognition and natural language processing. Efficient inference can reduce latency and improve decision-making quality, which is a key link in realizing intelligent services.
[0040] Distributed system: A system composed of multiple independent computers connected through a network. These computers work together to achieve a common goal. Distributed systems are designed to improve processing speed, increase system reliability and scalability. It can allocate computing and storage resources across multiple nodes to effectively handle large-scale data processing and high-concurrency requests and is suitable for various scenarios such as cloud computing and big data analysis.
[0041] Convolutional Neural Network (CNN): A feedforward neural network where its artificial neurons can respond to surrounding units within a partial coverage range, and it has excellent performance in large-scale image processing. Convolutional neural networks are particularly suitable for identifying visual patterns because they can effectively reduce the dimensionality of image data while retaining key information. Such networks typically consist of one or more convolutional layers and pooling layers, followed by fully connected layers (similar to traditional neural networks). Convolutional neural networks are widely used in fields such as image classification and object detection.
[0042] Recurrent Neural Network (RNN): A type of neural network used to process sequential data, where the network's output depends not only on the current input but also on previous states or previous inputs. The uniqueness of recurrent neural networks lies in their internal state (memory), which allows information to persist, making them very suitable for tasks such as time series analysis, speech recognition, and natural language processing.
[0043] Generative Adversarial Network (GAN): A generative adversarial network consists of two parts - a generator and a discriminator, both of which are deep learning models. The generator attempts to create fake data that appears real, while the discriminator tries to distinguish between real data and the fake data generated by the generator. Through this adversarial process, the two networks train and improve each other, and ultimately the generator can create very realistic synthetic data. Generative adversarial networks are used to generate content such as images, videos, and audio.
[0044] Transformer architecture model: A deep learning model based on the attention mechanism, designed to process sequential data without relying on recurrent neural networks (RNN). It was initially designed to solve natural language processing (NLP) tasks, but its application scope has been extended to include multiple fields such as image processing and speech recognition. The core of the Transformer model is the self-attention mechanism, which allows the model to consider other elements in the entire sequence when processing an element at a certain position. This feature makes the Transformer particularly suitable for tasks that require understanding context, such as machine translation, text summarization, and question-and-answer systems. In addition, the Transformer significantly improves computational efficiency by parallelizing the training process.
[0045] Bidirectional Encoder Representations from Transformers (BERT) model: A pre-trained language representation model based on the Transformer architecture, designed to improve the performance of various natural language processing tasks. Note: Different from traditional unidirectional language models, BERT uses a bidirectional training method, which means it can consider context information on both the left and right sides of a word simultaneously. This enables BERT to perform well in various natural language processing tasks such as sentiment analysis, named entity recognition, and question answering systems.
[0046] Graphics Processing Unit (GPU): A processor designed to accelerate graphics rendering and processing, especially good at performing parallel computing tasks. In addition to graphics processing, GPUs are widely used in scientific computing, machine learning, etc., because they can efficiently handle tasks of a large amount of data parallel computing. Through its numerous cores, the GPU can execute multiple computing threads simultaneously, thus performing well in tasks that require a large amount of computing resources such as deep learning.
[0047] Tensor Processing Unit (TPU): A processor designed specifically to accelerate machine learning workloads, especially for tensor operations in deep learning. It optimizes the execution efficiency of neural network calculations and can more efficiently handle large-scale datasets and complex model training tasks.
[0048] Neural Processing Unit (NPU): A type of processor designed specifically to execute neural network calculations, aiming to accelerate the execution process of deep learning algorithms. The neural processing unit optimizes the training and inference speed of machine learning models, especially suitable for processing complex tasks such as image recognition and natural language processing, improving efficiency and reducing energy consumption.
[0049] In this specification, a method for processing inference tasks is provided. This specification also relates to another method for processing inference tasks, a distributed system, a computing device, another computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0050] See Figure 1 , Figure 1The figure shows a flowchart of a method for processing inference tasks provided by an embodiment of this specification. This method is applied to the first node of a distributed system. The first node is any one of at least one node in the distributed system, and a parallel computing unit is deployed on the first node. The specific steps are as follows:
[0051] Step 102: In response to an execution instruction of the current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unload the task processing model corresponding to the previous inference request from the parallel computing unit and install the task processing model corresponding to the current inference request.
[0052] A distributed system is a system composed of multiple independent computers connected through a network. Each computer works collaboratively to provide an inference service based on a neural network model. The distributed system includes multiple nodes, and independent resources are deployed on any node, including a parallel computing unit, host system memory, persistent storage medium, etc. The distributed system can effectively handle large-scale data processing and multiple high-concurrency inference requests, improving the speed of processing inference tasks, increasing system reliability and scalability. For example, a cloud computing platform, which is a distributed system, includes multiple cloud server nodes. Any cloud server node runs on a corresponding host, and a parallel computing unit is deployed on any host.
[0053] A task processing model is a neural network model used to execute inference tasks, including but not limited to: convolutional neural network, recurrent neural network, generative adversarial network, Transformer architecture model, BERT model, large language model, etc. For example, a large language model applied to inference tasks in natural language.
[0054] Different inference tasks need to use different types of task processing models to execute. The type of the task processing model can be obtained by classification according to at least one of the application scenario, data type, model specification, etc. For example, image recognition may require a convolutional neural network model, while text analysis is more suitable for using a recurrent neural network or a large language model. Another example is that for real-time sentiment analysis in a resource-constrained environment, a large language model of a lightweight version of BERT may be required, while for large-scale text summarization in the cloud, a large language model with a larger scale of the Transformer architecture is more suitable.
[0055] The first node is any node that has distributed multiple inference requests. This node has independent resources, such as a parallel computing unit, host system memory, persistent storage medium, etc., for executing inference tasks based on a neural network model. In a distributed system, the first node is the basic unit for processing inference requests. Each first node can work independently.
[0056] The current inference request is the inference task request that needs to be processed currently, which has been distributed to the first node but has not started execution yet. The task processing model for executing the current inference request has not been loaded into the parallel computing unit, and it is necessary to determine whether it is necessary to dynamically switch the task processing model installed on the parallel computing unit according to the model type. For example, node A has just received a new inference request for sentiment analysis of text.
[0057] The previous inference request is the most recent inference task that was distributed to the first node before the current inference request. The previous inference request has been executed, and the task processing model for executing the previous inference request is installed on the parallel computing unit. For example, node A has just completed the task of object recognition for a picture as the previous inference request.
[0058] The parallel computing unit is a model hardware accelerator deployed on the node that can perform parallel computing, including but not limited to graphics processing units, tensor processing units, and neural network processing units, etc. The parallel computing unit is the core component for executing inference requests. The parallel computing unit reduces latency and improves throughput by efficiently processing data in parallel. Different parallel computing units are optimized for specific types of computing tasks, so selecting the appropriate computing unit plays a role in improving resource utilization. The efficient operation of the task inference model depends on the acceleration support of the parallel computing unit. However, considering factors such as cost, power consumption, and physical space limitations, the parallel computing units deployed in the distributed system are relatively limited, and it is necessary to reasonably plan resource allocation strategies and scheduling algorithms to achieve maximum performance and efficiency with limited hardware resources.
[0059] Optionally, the current inference request can be dynamically determined based on a hybrid scheduling strategy (preemption mode and first-in-first-out queue mode), or it can be determined based on any one mode (preemption mode or first-in-first-out queue mode). Among them, the preemption mode is that high-priority requests can interrupt low-priority tasks and load their models first. The first-in-first-out queue mode is to strictly queue according to the time order to avoid resource contention. Optionally, through the configuration interface, it is allowed to define the mode switching rules through a policy file or an application programming interface (Application Programming Interface, abbreviated as API), such as enabling the preemption mode when the node load is lower than the threshold.
[0060] It is recognized that the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent. An optional method is: it is recognized that the model identifiers of the task processing models corresponding to the previous inference request and the current inference request are inconsistent. For example, if the model hash of the previous inference request is a1b2c3 and the model hash of the current inference request is d4e5f6, then it is determined that the types are inconsistent. Another optional method is: it is recognized that the request metadata corresponding to the previous inference request and the current inference request are inconsistent. For example, the X-Model-Type field (such as model_type=resnet30) is extracted from the Hyper-Text Transfer Protocol (HTTP) header of the current inference request, and it is inconsistent with the X-Model-Type field (such as model_type=resnet50) extracted from the HTTP header of the most recent inference request recorded by the node, then it is determined that the types are inconsistent. Another optional method is: it is recognized that the containers corresponding to the previous inference request and the current inference request are inconsistent. For example, if the previous request is processed by the container pod-1 and the current request needs to be processed by pod-2, then it is determined that the types are inconsistent, which is not limited herein.
[0061] Unload the task processing model corresponding to the previous inference request from the parallel computing unit of the first node. An optional method is: release the unit storage medium of the parallel computing unit of the first node and unload the task processing model corresponding to the previous inference request. Another optional method is: asynchronously migrate the task processing model corresponding to the previous inference request of the parallel computing unit of the first node to the backup storage medium, which is not limited herein.
[0062] Install the task processing model corresponding to the current inference request. An optional method is: load the task processing model corresponding to the current inference request into the unit storage medium of the parallel computing unit of the first node. Another optional method is: asynchronously migrate the task processing model corresponding to the current inference request in the backup storage medium to the unit storage medium of the parallel computing unit of the first node, which is not limited herein.
[0063] Exemplarily, on node A, the current inference request is inference request 1 (real-time chatbot interaction request), and the previous inference request is:
[0064] Inference request 4: A multi-modal image batch processing request submitted by user D, which requires enterprise version model support (preferably large-scale parallel processing).
[0065] The type of the task processing model corresponding to the former is a large language model with a basic version (7B parameters), and the type of the task processing model corresponding to the latter is a large language model with an enterprise version (175B parameters). The model identifiers of the two are inconsistent. Release the video memory of the graphics processing unit of node A, unload the large language model of the enterprise version of inference request 4, and load the large language model of the basic version of inference request 1 into the video memory of the graphics processing unit of node A.
[0066] When different types of task processing models are detected, the task inference models that are no longer needed are automatically unloaded, and at the same time, the task inference models suitable for the current inference request are installed, thus ensuring the efficient use of resources and fast response, effectively solving the problem of low efficiency caused by resource shortage and the use of mismatched task inference models, optimizing resource allocation, and at the same time, providing a basis for computing resources and model environment configuration for the subsequent execution of the current inference request.
[0067] Step 104: Run the task processing model on the parallel computing unit to execute the current inference request.
[0068] Running the task processing model on the parallel computing unit to execute the current inference request, one optional way is: running the task processing model on the parallel computing unit to directly execute the current inference request. Another optional way is: running the task processing model on the parallel computing unit to containerize the execution of the current inference request. Another optional way is: running the task processing model on the parallel computing unit to call the hardware acceleration library to execute the current inference request, which is not limited here.
[0069] It should be noted that the task processing model corresponding to the current inference request can be installed in step 102 to complete the cold start process of the model. The cold start here refers to the process of first loading the model when no relevant task processing models are pre-loaded on the parallel computing unit. This process may involve operations such as reading the model file from the persistent storage medium, initializing the model parameters, and preparing the necessary running environment. By optimizing this process, such as pre-caching commonly used models and using efficient serialization techniques to reduce the loading time, the cold start time can be significantly reduced, so as to directly execute the inference request and improve the system response speed and user experience.
[0070] Hot updates or dynamic adjustments can also be performed after the model is loaded, which means updating or adjusting the model in use according to actual needs without interrupting the service. For example, in a high-concurrency scenario, the system can monitor real-time performance metrics to decide whether to allocate more resources to the currently running task processing model, or seamlessly switch to a new model version when a better version of the model is detected to improve the inference accuracy or efficiency.
[0071] Exemplarily, run the base version of the large language model loaded on the graphics processing unit to directly execute the inference request 1 and complete the real-time chatbot interaction.
[0072] In the embodiments of this specification, on the first node, dynamically identify the type of the task processing model corresponding to the current inference request and the type of the task processing model corresponding to the previous inference request. When detecting different types of task processing models, automatically unload the task inference models that are no longer needed, and at the same time install the task inference models applicable to the current inference request, thereby ensuring the efficient use of resources and fast response, effectively solving the problem of low efficiency caused by resource shortage and the use of mismatched task inference models, optimizing resource allocation, improving the efficiency of the inference task processing of a single node, and further improving the overall performance and reliability of the distributed system.
[0073] In an alternative embodiment of this specification, the following specific steps are further included:
[0074] In response to the execution instruction of the current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are the same, run the task processing model on the parallel computing unit to execute the current inference request.
[0075] Running the task processing model on the parallel computing unit to execute the current inference request, one alternative way is: run the task processing model on the parallel computing unit to directly execute the current inference request. Another alternative way is: run the task processing model on the parallel computing unit to containerize and execute the current inference request. Another alternative way is: run the task processing model on the parallel computing unit to call the hardware acceleration library to execute the current inference request, which is not limited here.
[0076] Exemplarily, on node A, the current inference request is inference request 3 (multilingual translation batch request), and the previous inference request is:
[0077] Inference request 4: A multimodal image batch request submitted by user D, which requires enterprise version model support (preferably large-scale parallel processing).
[0078] The type of the task processing model corresponding to the former is the large language model of the enterprise version (175B parameters), and the type of the task processing model corresponding to the latter is the large language model of the enterprise version (175B parameters), and the model identifiers of the two are the same.
[0079] Run the large language model of the enterprise version on the graphics processing unit to directly execute inference request 3 and complete the language translation batch processing.
[0080] In the embodiments of this specification, by dynamically identifying and switching task processing models in a distributed system, efficient utilization of resources and fast response are achieved. When it is detected that the current inference request uses the same type of model as the previous request, the existing model on the parallel computing unit is directly run for processing, avoiding unnecessary unloading and reloading processes, reducing latency, and improving processing efficiency. This not only optimizes resource allocation but also ensures the high availability and scalability of the distributed system.
[0081] In an optional embodiment of this specification, at least one container corresponding to the inference service of a task processing model runs on the first node, and any container processes inference requests for a specific type of task processing model. The method for identifying whether the types of the task processing models corresponding to the previous inference request and the current inference request are the same includes:
[0082] Identifying whether the previous inference request and the current inference request come from the same container on the first node.
[0083] The inference service of the task processing model is a service that provides inference capabilities based on a specific neural network model. For example, image recognition, text analysis, or speech processing, etc. The inference service is deployed in the node containers of the distributed system to handle large-scale concurrent requests and meet the requirements of efficient processing, and to achieve isolation and flexible scheduling.
[0084] A container is a lightweight operating system virtualization technology that allows developers to package an application and its dependencies into a portable package and run it consistently in any environment. Containers provide functions such as process isolation and resource limitation, but share the kernel of the host operating system. Container technology (such as Docker) enables applications to migrate and execute seamlessly in different environments, greatly improving development efficiency and deployment flexibility. Each container is an independent execution environment with its own file system, memory space, process table, etc., and is more lightweight than traditional virtual machines.
[0085] The container corresponding to the inference service of the task processing model is a container used to carry the inference service of a specific task processing model. This container contains all the dependencies required to run the model (such as library files, configuration files, etc.) and provides an isolated execution environment. Each containerized inference service can focus on processing inference requests for a specific type of task processing model, thereby achieving efficient resource utilization and service isolation. Through container orchestration tools (such as Kubernetes), the management and scheduling of containers can be further optimized.
[0086] Identify whether the previous inference request and the current inference request come from the same container on the first node. An optional way is to identify whether the previous inference request and the current inference request come from the same container on the first node based on the container identifiers of the containers where the previous inference request and the current inference request are located.
[0087] Exemplarily, two containers are running on node A:
[0088] Container pod-1: Deploy the basic large language model (7B parameters), dedicated to processing low-latency requests such as real-time chatbots;
[0089] Container pod-2: Deploy the enterprise large language model (175B parameters), dedicated to processing large-scale batch processing tasks.
[0090] The previous inference request is inference request 4: A multi-modal image batch processing request (requiring an enterprise model) submitted by user D, processed by container pod-2;
[0091] The current inference request is inference request 1: A real-time chatbot interaction request (requiring a basic model) submitted by user A, which needs to be processed by container pod-1.
[0092] By comparing the container identifiers "pod-2" and "pod-1", it is determined that they come from different containers.
[0093] In the embodiments of this specification, by identifying whether the inference requests come from the same container, precise scheduling of the task processing model and optimal resource allocation are achieved, unnecessary model switching and resource reallocation are avoided, processing latency is reduced, and response speed and service stability are improved. At the same time, based on the isolation characteristics of the container, the independence between different types of tasks is ensured, the flexibility and scalability of the system are enhanced, and the distributed system can process large-scale concurrent inference requests more efficiently.
[0094] In an optional embodiment of this specification, in step 102, unloading the task processing model corresponding to the previous inference request from the parallel computing unit of the first node includes the following specific steps:
[0095] Migrate the parameters of the task processing model corresponding to the previous inference request from the unit storage medium of the parallel computing unit of the first node to the host system memory or persistent storage medium of the first node;
[0096] In step 102, installing the task processing model corresponding to the current inference request includes the following specific steps:
[0097] Migrate the parameters of the task processing model corresponding to the current inference request from the host system memory or persistent storage medium of the first node to the unit storage medium of the parallel computing unit of the first node.
[0098] The parameters of the task processing model are the weights and bias values learned by the task processing model during training, as well as any other configuration parameters or hyperparameters, which together determine the running effect of the model. For different task processing models, the quantity and structure of the parameters may vary greatly. During the inference phase, these parameters are loaded into the computing unit to execute specific tasks.
[0099] The unit storage medium of the parallel computing unit is a high-speed storage area directly connected to the parallel computing unit, used to store the currently used task processing model and its parameters for quick access and processing. This storage medium usually has a high read / write speed and low latency, enabling the parallel computing unit to efficiently execute computing tasks. The unit storage medium of the parallel computing unit is the video memory or similar fast storage resources on dedicated hardware accelerators (such as graphics processing units, tensor processing units, and neural network processing units, etc.).
[0100] The host system memory is a high-speed storage area running the host operating system, which provides temporary data storage space to support the computing requirements of the central processing unit (CPU). The host system memory is used to store the currently active application program code and data, as well as the temporary information required by the operating system. Compared with the persistent storage medium, its access speed is faster, but the data will be lost after power-off.
[0101] The persistent storage medium is a storage device that can save data even after the power is turned off, including hard disk drives (HDDs), solid state drives (SSDs), etc. The persistent storage medium is used to store data for a long time. The capacity of the persistent storage medium is large but the access speed is relatively slow.
[0102] From the unit storage medium of the parallel computing unit of the first node, migrate the parameters of the task processing model corresponding to the previous inference request to the host system memory or the persistent storage medium of the first node. An optional method is as follows: adopt the Device-to-Host (D2H for short) access mode to migrate the parameters of the task processing model corresponding to the previous inference request from the unit storage medium of the parallel computing unit of the first node to the host system memory or the persistent storage medium of the first node. Specifically, perform direct memory access or asynchronous data transfer through the PCIe or NVLink hardware interface from the unit storage medium of the parallel computing unit to the host memory.
[0103] From the host system memory or the persistent storage medium of the first node, migrate the parameters of the task processing model corresponding to the current inference request to the unit storage medium of the parallel computing unit of the first node. An optional method is as follows: adopt the Host-to-Device (H2D for short) access mode to migrate the parameters of the task processing model corresponding to the current inference request from the host system memory or the persistent storage medium of the first node to the unit storage medium of the parallel computing unit of the first node. Specifically, perform direct memory access or asynchronous data transfer through the PCIe or NVLink hardware interface to load the model parameters from the host system memory to the unit storage medium of the parallel computing unit.
[0104] Exemplarily, adopt the Device-to-Host access mode and start the asynchronous data transfer from device to host through the hybrid interface of PCIe 4.0 bus and NVLink 3.0: read 17.5 billion parameters of the enterprise version model in chunks from the video memory of the graphics processing unit and migrate the model parameters to the cache area of the host memory. Load 700 million parameters of the basic version model from the cache area of the host memory, perform pipelined transmission of the parameters through the Host-to-Device access mode, load the parameters into the video memory of the graphics processing unit in batches, and at the same time start the pre-initialization of the computing core. Utilize the dedicated hardware acceleration library to optimize the PCIe bandwidth utilization. After the loading is completed, call the deep learning inference optimization engine to perform just-in-time compilation on the model.
[0105] In the embodiments of the present specification, by migrating the task processing model parameters of the previous inference request from the unit storage medium of the parallel computing unit to the host system memory or the persistent storage medium, and migrating the task processing model parameters corresponding to the current inference request to the unit storage medium of the parallel computing unit of the first node, the hardware resources of the parallel computing unit are dynamically scheduled, so that the new task processing model can be quickly loaded into the parallel computing unit for processing. This not only optimizes the usage efficiency of computing resources, reduces the model switching time, improves the response speed and throughput of the overall distributed system, is applicable to application scenarios that require frequent model replacement, but also greatly improves the efficiency and flexibility of inference task processing.
[0106] In an alternative embodiment of the present specification, the distributed system further includes a management module;
[0107] Before step 102, the following specific steps are further included:
[0108] Receive at least one inference request sent by the management module;
[0109] Put at least one inference request into the node task queue of the first node.
[0110] The node task queue corresponding to the node is a request buffering structure pre-allocated for the node, and the sequential processing of requests is completed in a first-in, first-out mode. Optionally, on the basis of first-in, first-out, a hybrid scheduling strategy can be implemented by combining a dynamic priority queue-jumping mechanism. Optionally, the capacity of the node task queue is dynamically adjusted according to the node hardware resources. For example: for a node equipped with a high-computing-power parallel computing unit, the queue length threshold is set to 50-100 inference requests. For an edge computing node, the queue length threshold is limited to 5-10 inference requests.
[0111] Exemplarily, in the inference task queue of node 1, there are the following requests:
[0112] Inference request 5: Real-time video analysis task of user D (priority P1, requires a professional version of the large language model)
[0113] Node 1 receives 2 inference requests sent by the management module and puts the 2 inference requests into the inference task queue of the first node. At this time, the requests sorted by reception time in the inference task queue are as follows:
[0114] Inference request 5: Real-time video analysis task of user D (priority P1, requires a professional version of the large language model)
[0115] Inference request 1: Real-time chatbot interaction of user A (priority P3, requires a basic version of the large language model)
[0116] Inference Request 2: Multilingual Translation Batch Processing for User C (Priority P0, requires large language model of enterprise version)
[0117] In the embodiments of this specification, multiple inference requests are placed into the node task queue corresponding to the node. By adopting the first-in, first-out queue mechanism, efficient task scheduling and resource management are achieved, optimizing the service response time and system throughput.
[0118] See Figure 2 , Figure 2 FIG. shows a flowchart of another inference task processing method provided by an embodiment of this specification. This method is applied to the management module of a distributed system. The distributed system further includes at least one node, and the specific steps are as follows:
[0119] Step 202: Obtain multiple inference requests for different types of task processing models, and distribute the multiple inference requests to at least one node. Among them, the first node is any one of the at least one node. The following steps are executed on the first node: In response to the execution instruction of the current inference request on the first node, in the case where the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unload the task processing model corresponding to the previous inference request from the parallel computing unit, and install the task processing model corresponding to the current inference request; Run the task processing model on the parallel computing unit to execute the current inference request.
[0120] The inference request for the task processing model is a task request that needs to use the task processing model to execute an inference task. The inference request is sent to the distributed system, and the distributed system determines on which node to execute according to factors such as the type, priority, and resource requirements of the inference request. For example, an inference request that requires object recognition of an uploaded picture will be assigned to a node equipped with an efficient graphics processing unit for execution to accelerate the calculation process.
[0121] One optional way to obtain multiple inference requests for different types of task processing models is: Obtain multiple inference requests for different types of task processing models in batches according to the time sequence. Another optional way is: Group and obtain multiple inference requests for different types of task processing models according to the model type. Another optional way is: Obtain multiple inference requests for different types of task processing models in batches according to the priority of the inference request. This is not limited here.
[0122] Distribute multiple inference requests to at least one node. One optional method is to distribute multiple inference requests to at least one node in a round-robin manner. For example, there are three nodes A, B, and C in a distributed system. When the distributed system receives multiple inference requests, it sequentially distributes the inference requests to nodes A, B, and C, and then starts the cycle from A again. Another optional method is to distribute multiple inference requests to at least one node based on the node status parameters of at least one node. For example, there are three nodes A, B, and C in a distributed system. When the distributed system receives multiple inference requests, node A currently has a low load and has more available resources, while nodes B and C are close to full load. , multiple inference requests are preferentially sent to node A. Another optional method is: based on the request parameters of multiple inference requests, multiple inference requests are distributed to at least one node. For example, there are three nodes A, B and C in the distributed system. When the distributed system receives an inference request G1 for performing an image recognition task and an inference request G2 for text analysis, the graphics processing unit of node A is currently loaded with a visual processing model, and the graphics processing unit of node B is currently loaded with a large-scale Transformer architecture model suitable for natural language processing. Inference request G1 will be sent to node A, and inference request G2 will be sent to node B, without limitation here.
[0123] For example, on a cloud computing platform for a large language model, three specifications of large language model instances are provided to the outside world: basic version (7B parameters), professional version (13B parameters), and enterprise version (175B parameters). Users can deploy on-demand elastically scalable reasoning services on this cloud computing platform according to project requirements. The distributed system divides the corresponding resources (such as GPU video memory, number of CPU cores, etc.) on a host cluster equipped with a high-performance graphics processing unit according to the instance specifications subscribed by the user and real-time load monitoring data, builds cloud server nodes, and isolates and runs reasoning services of different model versions in the form of containers on this cloud server. The distributed system obtains three reasoning requests for different types of task processing models in batches according to the time sequence:
[0124] Reasoning request 1: A real-time chatbot interaction request submitted by user A, requiring the use of the basic version of the model (low latency first);
[0125] Reasoning request 2: User B submits a request to generate an academic literature abstract, which requires calling the professional version model (high accuracy is preferred);
[0126] Reasoning request 3: A multi-language translation batch processing request submitted by user C, which requires support from the enterprise version model (large-scale parallel processing is preferred).
[0127] The distributed system polls and distributes these 3 inference requests to 2 nodes (Node 1 and Node 2): Inference Request 1 and Inference Request 3 to Node 1, and Inference Request 2 to Node 2.
[0128] In the embodiments of this specification, in a distributed system, multiple inference requests for different types of task processing models are distributed to at least one node, achieving large-scale concurrent processing, allowing different nodes to independently execute their respective inference tasks, reducing waiting time, and improving processing speed. At the same time, it lays a solid foundation for subsequent dynamic resource adjustment and intelligent scheduling strategies.
[0129] In an optional embodiment of this specification, in step 202, distributing multiple inference requests to at least one node includes the following specific steps:
[0130] Traverse the current node states of at least one node recorded in the polling table, and assign the first inference request to the next available node, where the polling table is a pre-constructed sequential circular list that real-time records the node states of at least one node, and the first inference request is any one of the multiple inference requests;
[0131] Update the first inference request, and return to the step of traversing the current node states of at least one node recorded in the polling table until multiple inference requests are distributed to at least one node.
[0132] The polling table is a pre-constructed sequential circular list that real-time records the node states of at least one node, used to real-time record the node states of at least one node in the distributed system. The polling table is dynamically updated, which can be updated at a preset frequency or triggered by a change in the node state, and is not limited here. The polling table is used to guide how inference requests are distributed among various nodes to ensure load balancing and effective utilization of resources. For example, in a distributed system with three nodes A, B, and C, the polling table is initially initialized as [A, B, C]. When the first inference request arrives, it is assigned to node A; the second request is assigned to node B; the third request is assigned to node C; the fourth request returns to node A again, and so on in a cycle.
[0133] The current node state of a node is the resource usage state of the node at the current moment, which reflects the current working condition, resource usage, and processing capacity of the node. The current node state of a node includes but is not limited to CPU utilization, memory usage, storage space occupancy, the number of current tasks, network bandwidth usage, etc. For example, if the current CPU utilization of a node is 80%, the memory usage is 60%, and 5 inference requests are being executed, then the current node state of this node is "high load, unavailable". On the contrary, if the CPU utilization of another node is only 20%, the memory usage is 30%, and there are no ongoing inference requests, then the current node state of this node is "low load, available".
[0134] An available node is a node that can receive and process new inference requests. An available node has sufficient resources (such as computing power, memory, storage) to ensure that new tasks can be executed smoothly.
[0135] The next available node is the first available node in the sequential loop recorded in the polling table that meets the resource conditions for executing an inference request. For example, an inference request needs to process a large number of matrix operations and requires at least 4 idle CPU cores and a high-performance GPU. If during the polling process, it is found that the current CPU utilization of a certain node has reached 80%, only 2 cores are available, and its GPU is being occupied by other tasks, then this node does not meet the resource conditions for this inference request, and the system will continue to check the next node. Another example, an inference request involves large-scale dataset operations and requires at least 16GB of idle RAM to ensure smooth operation. If during the polling process, it is found that although a certain node has sufficient CPU and GPU resources, its memory usage has reached 90%, and only 8GB is unused, then this node also does not meet the resource conditions for this inference request, and the system will continue to look for the next node. Also, for an inference request that needs to load a large model file (such as an enterprise version large language model), it may require at least 50GB of storage space for temporary data storage. If during the polling process, it is found that a certain node has insufficient hard disk space, even if its computing and memory resources are abundant, it cannot undertake this task and will therefore be skipped. Also, some inference tasks with high real-time requirements (such as real-time video stream analysis) have specific requirements for network bandwidth, and at least 1Gbps of available bandwidth is required to ensure data transmission efficiency. If during the polling process, it is found that most of the current network bandwidth of a certain node has been occupied by other tasks, leaving only 500Mbps of available bandwidth, then it will not be regarded as an available node that meets the resource conditions for this inference request.
[0136] Exemplarily, in a distributed system where a large language model is deployed, the polling table includes a sequential circular list: Node 1, Node 2, Node 3. The current status of each node is marked as "available". When the system receives the following sequence of inference requests:
[0137] Inference request 1 (real-time translation task, requires a basic version model) is assigned to Node 1 in the polling order;
[0138] Inference request 2 (document abstract generation, requires a professional version model) is assigned to Node 2 in the polling order;
[0139] Inference request 3 (image batch processing, requires an enterprise version model) is assigned to Node 3 in the polling order.
[0140] In the embodiments of this specification, by updating the node status in real time through the polling table, the uniform distribution of inference requests to at least one node is completed, different inference requests are scattered, reducing resource conflicts caused by simultaneous access to the same node. On this basis, since at least one node receives inference requests in the order of available status, load balancing between nodes is achieved, avoiding the accumulation of inference requests on the same node, realizing intelligent elastic scheduling, improving resource utilization, reducing request processing latency, and enhancing the efficiency of inference task processing.
[0141] In an alternative embodiment of this specification, after step 202, the following specific steps are further included:
[0142] In the case where the queue length of the node task queue of the first node exceeds a preset threshold, a scaling policy for the first node is triggered, where the first node is any one of the at least one node.
[0143] The preset threshold is the upper limit of the queue length dynamically adjusted according to the hardware resources and / or historical load data of the first node. For example, for a node equipped with a high-performance graphics processing unit, the preset threshold is 50 inference requests; for an edge computing node, the preset threshold is 10 inference requests; when the real-time latency of the node task queue (such as the average waiting time) exceeds 200 ms, the threshold is automatically reduced by 20% to prioritize service quality.
[0144] The scaling strategy for nodes is a strategy that automatically adjusts the allocation of computing resources based on real-time monitoring data. The scaling strategy includes, but is not limited to, the parallel computing unit scaling strategy, the node scaling strategy, etc. For example, the automatic horizontal container scaling (Horizontal Pod Autoscaler, abbreviated as HPA) rule is used for dynamic scaling. When it is detected that the length of the node task queue of the first node exceeds a preset threshold and the real-time latency (such as the average waiting time) exceeds the acceptable range, the distributed system will trigger a scaling operation. Another example is that through the elastic scaling service provided by the cloud computing platform, additional virtual machine or container instances can be quickly deployed to share the load, ensuring that all inference requests can be processed within a reasonable time. Still another example is to upgrade the resources of the first node, upgrade the node specification from 4-core central processing unit + 16GB memory to 8-core central processing unit + 32GB memory, and dynamically adjust the allocation strategy of the parallel computing unit, switching from the exclusive mode of a single graphics processing unit to the multi-task sharing mode.
[0145] Exemplarily, the task queue threshold of node A is 50 requests. When the real-time monitoring shows that the queue length reaches 55 and the video memory occupancy rate of the graphics processing unit exceeds 85%, the following scaling strategy is triggered:
[0146] According to the types of requests in the queue (70% are basic version models and 30% are enterprise version models), calculate the types and quantities of nodes to be scaled, call the application programming interface of the cloud computing platform to start a new cloud server instance, pre-load the container image of the basic version of the large language model on the new node, and mount the model parameters in the persistent storage medium. Register the new container on node A to the load balancer through the service discovery mechanism so that it can receive subsequent inference requests.
[0147] In the embodiments of this specification, when the length of the task queue of the first node exceeds the preset threshold, the scaling strategy is automatically triggered to ensure that the distributed system can process more inference requests without affecting the service quality. Through real-time monitoring and elastic scaling, not only the increase in latency caused by overload is avoided, but also the stability and response speed of the system are improved. It can effectively balance resource utilization and service efficiency, guarantee the user experience while optimizing costs.
[0148] In an alternative embodiment of this specification, when the length of the task queue of the first node exceeds the preset threshold, the scaling strategy for the first node is triggered and executed, including the following specific steps:
[0149] When the length of the task queue of the first node exceeds the preset threshold, obtain idle candidate nodes deployed with parallel computing units from the node pool, and add the candidate nodes to the distributed system.
[0150] A node pool is a set of pre-configured standby computing nodes that can be dynamically added to a distributed system. These nodes are usually in an idle or low-load state, with consistent basic software and hardware configurations, and can quickly respond to the system's expansion requirements to ensure service continuity and efficiency.
[0151] Candidate nodes are standby computing nodes pre-deployed in the node pool and in an idle or low-load state, and their hardware configurations match the current task requirements of the distributed system. Optionally, candidate nodes meet the following conditions: being in an idle state, having installed the host operating system, container environment, hardware drivers, and basic inference service framework, and having a parallel computing unit deployed to ensure that candidate nodes can quickly access the distributed system.
[0152] Exemplarily, the task queue length of node B reaches 60 (threshold 50), where 80% are batch translation requests requiring the enterprise version of the large language model, and 20% are real-time summary requests requiring the professional version of the model. Select 3 idle A100 nodes (nodes D, E, and F) from the node pool, all pre-installed with the enterprise version model container image. At the same time, start the container services of nodes D and E, load the enterprise version model with 175B parameters, and node F deploys the professional version model container to share real-time requests. The load balancer distributes newly arrived batch inference requests to nodes D and E in a round-robin manner, and real-time inference requests are directed to node F. 40 batch inference tasks in the queue of node B can also be migrated to nodes D and E.
[0153] In the embodiments of this specification, when the queue length of the node task queue of the first node exceeds the preset threshold, idle candidate nodes with parallel computing units deployed are obtained from the node pool and added to the distributed system, significantly improving the scalability and flexibility of the system. By pre-configuring candidate nodes that match the task requirements, the processing capacity can be quickly increased during peak loads, effectively avoiding service delays and performance degradation caused by overload. This not only ensures that high-priority and real-time tasks can be processed in a timely manner but also optimizes the resource utilization efficiency and reduces the waiting time of users. In addition, the way of dynamically adjusting and allocating computing resources enables the system to quickly respond to load changes, maintain the efficient and stable operation of the service, and thus improve the user experience and service quality.
[0154] In an alternative embodiment of this specification, after step 202, the following specific steps are further included:
[0155] When the number of nodes corresponding to at least one node task queue being an empty queue exceeds the preset threshold, a second node is released from the distributed system, where the second node is a node whose node task queue is an empty queue among the at least one node.
[0156] The preset threshold is the upper limit of the number of empty nodes in a pre-set distributed system. For example, if the preset threshold is set to 5 idle nodes and the task queues of 8 nodes are empty, the distributed system will select 3 nodes from these 8 nodes for release, thereby reducing unnecessary resource consumption and saving costs.
[0157] The second node is any node with an empty node queue. This node has independent resources, such as parallel computing units, host system memory, persistent storage media, etc., for executing inference tasks based on neural network models. In a distributed system, the second node is the basic unit for processing inference requests. Each second node can work independently.
[0158] Exemplarily, the cloud computing platform currently includes 10 cloud server nodes:
[0159] Nodes 1 - 3: Deploy enterprise version models;
[0160] Nodes 4 - 7: Deploy professional version models;
[0161] Nodes 8 - 10: Deploy basic version models.
[0162] Monitoring shows that the task queues of nodes 1, 2, 4, 5, and 8 are empty and they have been continuously idle for more than 15 minutes. The preset threshold is 3 idle nodes. Release nodes 1 and 2, return them to the node pool of the cloud computing platform, and cancel their service registration information. The load balancer will no longer distribute new inference requests to them.
[0163] In the embodiments of this specification, when the number of nodes with empty node task queues corresponding to at least one node exceeds the preset threshold, release the second nodes from the distributed system, effectively reducing unnecessary resource consumption and lowering the operating cost. This mechanism ensures that the system only maintains a reasonable number of working nodes by dynamically adjusting the number of active nodes, improving resource utilization and economic efficiency. It not only avoids resource waste but also does not affect the overall performance and service quality of the system.
[0164] The following combines the attached Figure 3 , taking the application of the inference task processing method provided in this specification in large language model inference as an example, to further illustrate the inference task processing method. Among them, Figure 3 FIG. shows the processing procedure flowchart of an inference task processing method applied to large language model inference provided in an embodiment of this specification. This method is applied to a cloud computing platform, which includes multiple nodes, and a graphics processing unit is deployed on any node, and the specific steps are as follows:
[0165] Step 302: Obtain M inference requests for different types of large language models, traverse the current node status of multiple nodes recorded in the polling table, and assign the M inference requests to the next available node.
[0166] Step 304: Put the M inference requests into the node task queues corresponding to N available nodes.
[0167] Step 306: Identify whether the previous inference request and the current inference request of the first node come from the same container on the first node, where the first node is any one of the N available nodes.
[0168] Step 308: If so, run the large language model on the graphics processing unit and execute the current inference request.
[0169] Step 310: If not, migrate the parameters of the large language model corresponding to the previous inference request from the video memory of the graphics processing unit of the first node to the host system memory of the first node, and migrate the parameters of the target large language model from the host system memory of the first node to the video memory of the graphics processing unit of the first node, run the large language model on the graphics processing unit, and execute the current inference request.
[0170] Step 312: In the case where the queue length of the node task queue of the first node exceeds a preset threshold, obtain idle candidate nodes with graphics processing units deployed from the node pool and add the candidate nodes to the cloud computing platform to obtain N + 1 available nodes.
[0171] Step 314: In the case where the number of nodes with empty node task queues corresponding to N available nodes exceeds a preset threshold, release the second node from the cloud computing platform to obtain N - 1 available nodes.
[0172] In the embodiments of this specification, a reasoning architecture that combines elastic reasoning and dynamic unloading is proposed, aiming to significantly improve the efficiency of GPU resource utilization and system flexibility. By dynamically adjusting the size of the GPU instance, the capacity is automatically expanded to meet the demand when the number of reasoning requests increases; the capacity is reduced to release resources when the load is low, effectively reducing the cost. In addition, the dynamic parameter migration mechanism between the GPU and the host system memory is used to support time-sharing sharing of GPU resources by different large language models, thereby improving hardware utilization. Specifically, when it is detected that the first node needs to switch to process tasks of different models, the model parameters of the previous task can be quickly migrated to the host memory, and the new model parameters can be loaded into the video memory of the GPU to achieve efficient task switching. At the same time, for nodes with excessive load, the load is dispersed by adding idle nodes from the node pool to ensure system stability and response speed. For nodes that have been idle for a long time, resources are released to further optimize resource allocation, thereby improving the elasticity and efficiency of the reasoning service as a whole and ensuring a high-quality service experience.
[0173] Corresponding to the above method embodiment, this specification also provides a distributed system embodiment. Figure 4 FIG. 1 shows a schematic diagram of a distributed system provided by an embodiment of this specification. Figure 4 As shown, the distributed system 400 includes a management module 410 and at least one node, any node is deployed with a parallel computing unit, and the first node 420 is any one of the at least one node;
[0174] A management module 410, configured to obtain multiple reasoning requests for different types of task processing models, and distribute the multiple reasoning requests to at least one node;
[0175] The first node 420 is used to respond to the execution instruction of the current reasoning request on the first node 420. When the types of task processing models corresponding to the previous reasoning request and the current reasoning request are inconsistent, the task processing model corresponding to the previous reasoning request is uninstalled from the parallel computing unit, and the task processing model corresponding to the current reasoning request is installed, and the task processing model on the parallel computing unit is run to execute the current reasoning request.
[0176] The management module 410 is the core control unit of the distributed system 400, responsible for global resource scheduling, task distribution, node management, and the execution of dynamic scaling policies. Its core functions include: 1. Request perception and routing: Receiving external inference requests, analyzing metadata such as request type, priority, and resource requirements, and dynamically selecting reasonable nodes for distribution. 2. Node status monitoring: Real-time collection of load metrics of each node (such as the utilization rate of the graphics processing unit, memory occupancy, and task queue length), constructing a global resource view to support intelligent scheduling. 3. Elastic scaling decision-making: Triggering scaling-up or scaling-down operations according to preset policies (such as queue overflow thresholds and the number of idle nodes), maintaining the resource efficiency and service stability of the distributed system. 4. Model lifecycle management: Coordinating the hot switching of models between nodes (such as unloading old models and loading new models) to ensure seamless connection of the inference service.
[0177] Among them, the request dispatcher routes the inference requests to appropriate nodes based on load balancing algorithms (such as weighted round-robin and least connections). For example, using Nginx reverse proxy combined with a custom scripting language to achieve dynamic routing; or extending the scheduling logic based on the container orchestration system ingress controller (Kubernetes Ingress Controller).
[0178] Among them, the resource monitoring engine collects node hardware metrics and container status through a proxy (such as Prometheus Exporter) to generate real-time monitoring data. For example, integrating the Prometheus + Grafana monitoring stack and using the GPU MetricsExporter to collect metrics such as video memory utilization rate and computing core occupancy rate.
[0179] Among them, the policy executor parses the scaling rules (such as queue length thresholds and response time service level agreements), and calls the cloud platform application programming interface or orchestration tools to execute operations. For example, implementing automatic scaling points based on the automatic horizontal container scaling of Kubernetes.
[0180] Among them, the metadata repository stores metadata such as model versions, node configurations, and scheduling policies, supporting fast query and version rollback. For example, using Redis as a distributed key-value store to record key information such as model hashes, node IPs, and container image versions.
[0181] In an optional embodiment of this specification, at least one container corresponding to the inference service of the task processing model is running on the first node 420, and any container processes the inference requests of a specific type of task processing model;
[0182] The first node 420 is specifically used to identify whether the previous inference request and the current inference request on the first node 420 come from the same container on the first node 420.
[0183] In an alternative embodiment of this specification, the first node 420 is further configured to place multiple inference requests into the node task queue of the first node 420, and when the queue length of the node task queue of the first node 420 exceeds a preset threshold, send an expansion policy request to the management module 410;
[0184] The management module 410 is further configured to, in response to the expansion policy request, obtain idle candidate nodes deployed with parallel computing units from the node pool and add the candidate nodes to the distributed system.
[0185] In an alternative embodiment of this specification, the management module 410 is further configured to, when the number of nodes corresponding to at least one node task queue being an empty queue exceeds a preset threshold, release a second node from the distributed system, where the second node is a node with an empty node task queue among the at least one node.
[0186] In the embodiments of this specification, in the distributed system, the management module distributes multiple inference requests for different types of task processing models to at least one node, achieving large-scale concurrent processing, allowing different nodes to independently execute their respective inference tasks, reducing waiting time, and improving processing speed. On the first node, dynamically identify the task processing model type of the current inference request and the task processing model type corresponding to the previous inference request, and when detecting different types of task processing models, automatically uninstall the no-longer-needed task inference models, and at the same time install the task inference models applicable to the current inference request, thereby ensuring efficient utilization of resources and fast response, effectively solving the problem of low efficiency caused by resource shortage and the use of mismatched task inference models, optimizing resource allocation, improving the efficiency of the inference task processing of the first node, and further improving the overall performance and reliability of the distributed system.
[0187] The above is a schematic solution of a distributed system in this embodiment. It should be noted that the technical solution of this distributed system and the technical solution of the above inference task processing method belong to the same concept. For the details not described in the technical solution of the distributed system, reference can be made to the description of the technical solution of the above inference task processing method.
[0188] Corresponding to the above Figure 4 content, Figure 5 FIG. shows an architecture schematic diagram of a distributed system provided by an embodiment of this specification, as Figure 5 shown:
[0189] The management module distributes multiple inference requests to Node 1 and Node 2.
[0190] The node task queue of node 1 is used to store the inference requests assigned to node 1, and the node task queue of node 2 is used to store the inference requests assigned to node 2.
[0191] Multiple inference requests are executed in a first-in-first-out order. Any inference request is executed on the container of the corresponding inference service isolated on the node.
[0192] The node identifies whether the task processing model type of the current inference request is the same as that of the previous inference request. If not, the device-to-host mode is adopted to unload the task processing model corresponding to the previous inference request from the video memory of the graphics processing unit to the host system memory, and the host-to-device mode is adopted to install the task processing model corresponding to the current inference request from the host system memory to the video memory of the graphics processing unit.
[0193] The above is a schematic solution of a distributed system according to this embodiment. It should be noted that the technical solution of this distributed system and the technical solution of the above inference task processing method belong to the same concept. For the details not described in the technical solution of the distributed system, reference can be made to the description of the technical solution of the above inference task processing method.
[0194] Figure 6 The structural block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.
[0195] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Controller (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0196] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 other components not shown therein may also be connected to each other, for example, via a bus. It should be understood that Figure 6 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0197] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0198] Wherein, the processor 620 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the above Figure 1Steps of the inference task processing method shown in the embodiment.
[0199] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the above Figure 1 The technical solution of the inference task processing method shown in the embodiment belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the inference task processing method shown in the above Figure 1 embodiment.
[0200] Figure 7 The block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 700 include but are not limited to a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to store data.
[0201] The computing device 700 further includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (for example, a Network Interface Controller (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0202] In an embodiment of this specification, the above components of the computing device 700 and Figure 7 other components not shown in the figure may also be connected to each other, for example, through a bus. It should be understood that Figure 7The block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0203] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0204] Among them, the processor 720 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the inference task processing method shown in the above Figure 2 embodiment are implemented.
[0205] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the inference task processing method shown in the above Figure 2 embodiment belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the inference task processing method shown in the above Figure 2 embodiment.
[0206] An embodiment of this specification also provides a computer-readable storage medium, which stores computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above inference task processing method are implemented.
[0207] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above inference task processing method belong to the same concept. For the details not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the inference task processing method.
[0208] An embodiment of this specification also provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above inference task processing method are implemented.
[0209] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned inference task processing method belong to the same concept. For the details not described in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above-mentioned inference task processing method.
[0210] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0211] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM for short), a random access memory (RAM for short), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0212] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0213] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0214] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present specification, many modifications and variations can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for processing inference tasks, which is applied to a first node of a distributed system. The first node is any one of at least one node of the distributed system, and a parallel computing unit is deployed on the first node. The method includes: In response to an execution instruction of a current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unload the task processing model corresponding to the previous inference request from the parallel computing unit, and install the task processing model corresponding to the current inference request, where the current inference request is an inference task request that has been distributed to the first node but has not started execution yet, and the previous inference request is the most recent inference task request that has been distributed to the first node and has been executed and completed before the current inference request; Run the task processing model on the parallel computing unit to execute the current inference request.
2. The method according to claim 1, further including: In response to an execution instruction of a current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are consistent, run the task processing model on the parallel computing unit to execute the current inference request.
3. The method according to claim 1, wherein At least one container corresponding to the inference service of the task processing model runs on the first node. Any container processes the inference request of a specific type of task processing model. The container corresponding to the inference service of the task processing model is a container for hosting the inference service of a certain specific task processing model. The containerized inference service focuses on processing the inference request of a specific type of task processing model; The manner of identifying whether the types of the task processing models corresponding to the previous inference request and the current inference request are consistent includes: Identify whether the containers corresponding to the previous inference request and the current inference request are the same, and determine whether the types of the task processing models corresponding to the previous inference request and the current inference request are consistent.
4. The method according to claim 1, where unloading the task processing model corresponding to the previous inference request from the parallel computing unit includes: Migrate the parameters of the task processing model corresponding to the previous inference request from the unit storage medium of the parallel computing unit of the first node to the memory of the host system of the first node or a persistent storage medium; Installing the task processing model corresponding to the current inference request includes: Migrate the parameters of the task processing model corresponding to the current inference request from the memory of the host system of the first node or a persistent storage medium to the unit storage medium of the parallel computing unit of the first node.
5. The method according to any one of claims 1-4, wherein, The distributed system further includes a management module; Before, in response to an execution instruction of a current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unloading the task processing model corresponding to the previous inference request from the parallel computing unit and installing the task processing model corresponding to the current inference request, it further includes: Receive at least one inference request sent by the management module; Put the at least one inference request into the node task queue of the first node.
6. An inference task processing method applied to a management module of a distributed system, the distributed system further including at least one node, including: Obtain multiple inference requests for different types of task processing models, and distribute the multiple inference requests to the at least one node, where the first node is any one of the at least one node, and a parallel computing unit is deployed on the first node. The following steps are performed on the first node: In response to an execution instruction of the current inference request on the first node, when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, unload the task processing model corresponding to the previous inference request from the parallel computing unit, and install the task processing model corresponding to the current inference request, where the current inference request is an inference task request that has been distributed to the first node but has not started execution, and the previous inference request is the most recent inference task request that has been distributed to the first node and has been executed before the current inference request; Run the task processing model on the parallel computing unit to execute the current inference request.
7. According to the method of claim 6, the distributing the multiple inference requests to the at least one node includes: Traverse the current node states of the at least one node recorded in the polling table, and assign the first inference request to the next available node, where the polling table is a pre-constructed sequential circular list that records the node states of the at least one node in real time, and the first inference request is any one of the multiple inference requests; Update the first inference request, and return to the step of traversing the current node states of the at least one node recorded in the polling table until the multiple inference requests are distributed to the at least one node.
8. According to the method of claim 6 or 7, after the distributing the multiple inference requests to the at least one node, it further includes: When the queue length of the node task queue of the first node exceeds a preset threshold, trigger the execution of an expansion strategy for the first node.
9. According to the method of claim 8, the triggering the execution of an expansion strategy for the first node when the queue length of the node task queue of the first node exceeds a preset threshold includes: When the queue length of the node task queue of the first node exceeds a preset threshold, obtain idle candidate nodes with a parallel computing unit deployed from the node pool, and add the candidate nodes to the distributed system.
10. According to the method of claim 6 or 7, after the distributing the multiple inference requests to the at least one node, it further includes: When the number of nodes corresponding to the node task queues of the at least one node being empty queues exceeds a preset threshold, release a second node from the distributed system, where the second node is a node among the at least one node whose node task queue is an empty queue.
11. A distributed system, including a management module and at least one node, with a parallel computing unit deployed on any node, and the first node being any one of the at least one node; The management module is configured to obtain a plurality of inference requests for different types of task processing models and distribute the plurality of inference requests to the at least one node; The first node is configured to, in response to an execution instruction of a current inference request on the first node, unload a task processing model corresponding to the previous inference request from the parallel computing unit and install a task processing model corresponding to the current inference request when the types of the task processing models corresponding to the previous inference request and the current inference request are inconsistent, run the task processing model on the parallel computing unit, and execute the current inference request, where The current inference request is an inference task request that has been distributed to the first node but has not started execution yet, and the previous inference request is the most recent inference task request that has been distributed to the first node and has been executed and completed before the current inference request.
12. The distributed system according to claim 11, wherein, At least one container corresponding to the inference service of the task processing model runs on the first node, and any container processes the inference request of a specific type of task processing model. The container corresponding to the inference service of the task processing model is a container used to carry the inference service of a certain specific task processing model, and the containerized inference service focuses on processing the inference requests of specific types of task processing models; The first node is specifically configured to identify whether the containers corresponding to the previous inference request and the current inference request are the same, and determine whether the types of the task processing models corresponding to the previous inference request and the current inference request are the same.
13. The distributed system according to claim 11, wherein the first node is further configured to put the plurality of inference requests into the node task queue of the first node, and when the queue length of the node task queue of the first node exceeds a preset threshold, send an expansion policy request to the management module; The management module is further configured to, in response to the expansion policy request, obtain idle candidate nodes deployed with parallel computing units from a node pool and add the candidate nodes to the distributed system.
14. The distributed system according to claim 11, wherein the management module is further configured to release a second node from the distributed system when the number of nodes whose corresponding node task queues are empty queues exceeds a preset threshold, where The second node is a node among the at least one node whose node task queue is an empty queue.
15. A computing device, including: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
16. A computing device, including: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 6 to 10 are implemented.
17. A computer-readable storage medium that stores computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
18. A computer program product, comprising a computer program / instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
System for novel service platform of loose cloud nodes under cloud computing network environment
CN102821162A
Method for dynamically calling multiple models for deep learning algorithm and service architecture
CN119645607A