An AI model pooling sharing and online inference implementation method

CN122840251APending Publication Date: 2026-09-29BEIJING WANJIE DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611053924.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

若过早回收旧实例的资源,会导致长尾任务被强制中断,客户端无法收到推理结果;若等待固定时长后再回收,则无法精确判断任务是否全部完成,造成资源释放延迟或任务丢失

Benefits of technology

[0045]本发明通过对每个推理请求计算辅助损失函数对中间特征张量的辅助梯度,并在每个实例的缓冲区中持续累积记录,构建推理流形张力场张量并汇总得出共享应力;当共享应力超过预设应力阈值时,自动判定该实例需要分裂;这使得多模型共享冲突的程度能够被精确量化和持续监测,在冲突达到可接受精度损失的临界点之前主动触发结构演化;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840251A_ABST
    Figure CN122840251A_ABST
Patent Text Reader

Abstract

This invention discloses a method for AI model pooling and online inference, relating to the field of machine learning model deployment and online inference. It solves the problems of decreased inference accuracy and traffic interruption during instance splitting caused by feature-driven direction conflicts when multiple AI models share an instance pool. This invention constructs a shared instance pool by extracting the common computational cores of multiple AI models. While processing online inference requests, it calculates auxiliary gradients and stores them in a buffer. Based on the buffer records, it calculates the inference manifold tension field tensor to obtain shared stress. When the shared stress exceeds a preset threshold, it triggers splitting. It freezes traffic by controlling inference inertia flux, temporarily storing new requests in a micro-buffer queue. It monitors the emptying status of old instance tasks through an instance retirement emptying entropy function. After emptying, it reclaims resources and activates new instances, distributing requests in the micro-buffer queue to the new instances for processing. This invention achieves efficient utilization of shared inference resources across multiple models and lossless online instance evolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine learning model deployment and online inference, specifically a method for AI model pooling and sharing and online inference implementation. Background Technology

[0002] With the rapid development of artificial intelligence technology, deep learning models have been widely used in fields such as text processing, image recognition, and speech analysis. In real-world business scenarios, a service system often needs to deploy multiple AI models with similar functions but different task types simultaneously. For example, the same text input may need to be processed by multiple models such as sentiment analysis, topic classification, and named entity recognition. These models typically share the same backbone network structure, differing only in the final output layer.

[0003] In typical online inference deployment architectures, model pooling and sharing are often used to improve hardware resource utilization and reduce operational costs. This involves extracting the common parts of multiple AI models and building a shared pool of inference instances. Each inference instance loads the shared computational portion and the private output portions of all models, with the instance pool uniformly receiving and processing online inference requests from clients. This deployment method allows multiple models to reuse the same computational resources, reducing model loading times and GPU memory usage.

[0004] Online inference services have stringent requirements for response latency and service stability. During the operation of a shared instance pool, different AI models may have different optimization directions in the feature space. When multiple models share the same inference instance, the private output headers of each model exert different pulls on the shared feature representation. Maintaining the inference accuracy of each model while sharing inference resources, and expanding or reorganizing the shared instance online when necessary, while ensuring uninterrupted and packet-free online inference service, are the key technical challenges that need to be addressed in the implementation of AI model pooling and online inference.

[0005] The following problems exist in the existing technology:

[0006] In scenarios where multiple AI models share the same inference instance, existing technologies show that different AI models have different directions of attraction to the shared feature space. The conflict of optimization directions of each model in the feature space cannot be quantitatively perceived, which leads to a gradual decrease in the inference accuracy of some models during the sharing process. Furthermore, there is a lack of technical means to automatically trigger instance splitting or reorganization.

[0007] In existing technologies, when a shared instance is determined to need to be split and replaced, newly arriving online inference requests will be lost after the old instance is emptied if they are still assigned to the old instance that is about to be decommissioned; if they are directly switched to a new instance that is not yet ready, request processing will be interrupted. There is a lack of a technical means to precisely control traffic switching during the instance replacement transition period.

[0008] In existing technologies, long-tail inference tasks may still be running on old instances after new requests are no longer allocated to them. If the resources of old instances are reclaimed too early, the long-tail tasks will be forcibly interrupted, and the client will not receive the inference results. If the resources are reclaimed after a fixed period of time, it will be impossible to accurately determine whether all tasks have been completed, resulting in resource release delays or task loss. Summary of the Invention

[0009] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes an AI model pooling and sharing and online inference implementation method to solve the above-mentioned technical problems.

[0010] The first aspect of this invention provides a method for AI model pooling and sharing and online inference, comprising the following steps:

[0011] S1: Obtain a set of trained AI models, extract the common computing core of this set of AI models, build an instance pool containing multiple instances, each instance is loaded with the common computing core and the private header of this set of AI models, and preheat each instance to a hot standby state;

[0012] S2: In response to the online inference request, select an instance from the instance pool, call the instance's public computing core to calculate the intermediate feature tensor against the input tensor of the online inference request, call the private head of the target AI model specified in the online inference request to calculate the inference result against the intermediate feature tensor, and calculate the auxiliary gradient of the auxiliary loss function of the target AI model's private head on the intermediate feature tensor. Store the intermediate feature tensor, auxiliary gradient, and the identifier of the target AI model as a record in the instance's buffer.

[0013] S3: Based on the records in the buffer of the instance, calculate the inference manifold tension field tensor of the instance, and summarize the shared stress of the instance. When the shared stress exceeds a preset stress threshold, determine that the instance needs to be split.

[0014] S4: Mark the instance that needs to be split as the instance to be split. For the newly arrived online inference request, traverse all active instances in the instance pool as non-triggered split instances. Calculate the inference inertia flux of the newly arrived online inference request from the instance to be split to each non-triggered split instance. Take the minimum value among them. When the minimum value exceeds the preset flux threshold, stop allocating the newly arrived online inference request to the instance to be split and store the newly arrived online inference request in the micro buffer queue.

[0015] S5: Create two new instances and load a common computing core. Divide the private header on the instance to be split and load it into the two new instances. Continuously calculate the instance retirement and emptying entropy function for the instance to be split. When the value of the instance retirement and emptying entropy function drops below a preset critical value and remains stable, reclaim the resources occupied by the instance to be split, activate the two new instances to join the instance pool, and distribute the online inference requests in the micro-buffer queue to the two new instances for processing.

[0016] Preferably, step S1 includes the following steps:

[0017] Obtain a set of trained AI models, perform subgraph isomorphism detection on the inference computation graphs of each AI model, take the longest continuous prefix subgraph that is completely consistent from the input node in the inference computation graphs of each AI model as the common computation kernel, and take the remaining subgraphs as the private headers of each AI model; create a preset number of instances on the specified hardware, and load the common computation kernel and the private headers of each of the AI ​​models in this set of AI models.

[0018] For each instance, a complete forward inference is performed using an empty tensor through a public computation kernel and a private header, triggering the instance to perform memory allocation, operator fusion, and just-in-time compilation of the computation kernel, and enter a hot standby state.

[0019] Preferably, step S2 includes the following steps:

[0020] In response to an online inference request, the number of currently executing inference tasks for each instance in the instance pool is traversed, and the instance with the fewest inference tasks is selected. The online inference request is then assigned to that instance. The input tensor of the online inference request is fed into the common computation kernel of that instance, and after being transformed layer by layer by each operator layer in the common computation kernel, an intermediate feature tensor is output.

[0021] Based on the target AI model identifier in the online inference request, the corresponding private head is found from the private head loaded by the instance. The intermediate feature tensor is sent into the private head, and after being transformed layer by layer by the operators in the private head, the inference result is output and returned.

[0022] An auxiliary loss function is constructed at the output of the private head. The construction method of the auxiliary loss function depends on the task type of the AI ​​model corresponding to the private head. When the task type is classification, the cross-entropy between the probability distribution of the private head output and the uniform distribution is calculated as the auxiliary loss function. When the task type is regression, the mean square error between the predicted value of the private head output and the preset constant value is calculated as the auxiliary loss function. The auxiliary gradient of the auxiliary loss function on the intermediate feature tensor is calculated by automatic differentiation. The backpropagation of the auxiliary gradient terminates at the intermediate feature tensor, and the intermediate feature tensor is used as the termination node of the backpropagation.

[0023] Preferably, in step S3, calculating the inference manifold tension field tensor of the instance based on the records in the buffer of the instance includes the following steps:

[0024] For each record in the buffer of the instance, obtain the intermediate feature tensor z and auxiliary gradient of that record. Calculate the normalization factor ;

[0025] The intermediate feature tensor z is used as the feature point p, and the auxiliary gradient is used as well. Calculate the tensor of the inference manifold tension field at the feature point p. Where k is the identifier of the target AI model, The buffer record is a set of identifiers for the different target AI models involved. for The number of AI models in China.

[0026] Preferably, in step S3, the process of summarizing the shared stress of the instance includes the following steps:

[0027] The intermediate feature tensor in each record participating in the calculation of the inference manifold tension field tensor in the buffer is taken as a feature point, and all feature points constitute a feature point set. ,remember for The total number of feature points in the matrix, for each feature point p∈ Calculate the trace of the tension field tensor T(p) of the inference manifold Then, the shared stress of the instance is calculated. .

[0028] Preferably, step S4 includes the following steps:

[0029] The status flag of the instance to be split in the instance pool is switched from active to pending split, so that it no longer receives new online inference requests; online inference requests that have been allocated to the instance to be split but have not yet been completed continue to be processed;

[0030] For each newly arrived online inference request after the state flag of the instance to be split is switched to the state flag after the split, all instances in the instance pool that are active are traversed as non-split instances. The inference inertia flux of the newly arrived online inference request from the instance to be split to each non-split instance is calculated, and the minimum value is taken.

[0031] Inference inertia flux is the ratio of the context reconstruction cost of the newly arrived online inference request to the computational margin of the non-triggered split instance, where the context reconstruction cost is equal to the product of the shared stress of the instance to be split and the total number of input tensor elements and the normalization coefficient of the newly arrived online inference request, and the computational margin is the maximum number of concurrent processing tasks of the non-triggered split instance minus the number of inference tasks currently being executed.

[0032] When the minimum value exceeds the preset throughput threshold, the allocation of the newly arrived online inference request to the instance to be split is stopped, and it is not allocated to any instance that has not triggered a split. The newly arrived online inference request is stored in the micro buffer queue.

[0033] Preferably, in step S5, the continuous calculation of the instance retirement and emptying entropy function for the instance to be split includes the following steps:

[0034] The runtime of all inference tasks that have not yet been completed on the instance to be split is obtained at a fixed sampling period, and a normalized task runtime distribution p(τ,t) is constructed, where τ is the task runtime and t is the sampling time.

[0035] Obtain the historical average task processing rate γ of the instance to be split, and define the cumulative distribution of task completion probability under ideal empty state as follows: ;

[0036] Calculate the instance decommissioning and emptying entropy function at sampling time t. The formula is:

[0037]

[0038] in, The idle time of the instance to be split since it last received a new task; The rate at which new tasks are assigned to the instance to be split at time s.

[0039] Preferably, in step S5, when the value of the instance retirement entropy function drops below a preset threshold and remains stable, the resources occupied by the instance to be split are reclaimed, two new instances are activated and added to the instance pool, and the online inference requests in the micro-buffer queue are distributed to the two new instances for processing, including the following steps:

[0040] When an instance retires, the entropy function is emptied. When the value is less than the preset threshold and is true for two consecutive fixed sampling periods, it is determined that the instance to be split has completed all task emptying and no new tasks have infiltrated.

[0041] Reclaim the video memory and memory resources occupied by the instance to be split, and remove the instance to be split from the instance pool;

[0042] Mark the two new instances as active and add them to the list of available instances in the instance pool;

[0043] Online inference requests are retrieved one by one from the micro-buffer queue in a first-in, first-out order. Based on the target AI model identifier carried by each online inference request, the online inference request is distributed to a new instance loaded with the corresponding private header for processing.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] This invention calculates the auxiliary gradient of the auxiliary loss function on the intermediate feature tensor for each inference request and continuously accumulates and records it in the buffer of each instance to construct the inference manifold tension field tensor and summarize the shared stress. When the shared stress exceeds a preset stress threshold, the instance is automatically determined to split. This enables the degree of shared conflict among multiple models to be accurately quantified and continuously monitored, and actively triggers structural evolution before the conflict reaches the critical point of acceptable accuracy loss.

[0046] This invention calculates the inference inertia flux of newly arriving online inference requests migrating from the instance to be split to the instance that has not yet triggered a split, and compares the context reconstruction cost with the computing power margin of the target instance; when the inference inertia flux exceeds a preset flux threshold, it stops allocating new requests to the instance to be split and stores the new requests in a micro-buffer queue; this makes the traffic switching during instance replacement have a precise quantitative decision basis, ensuring that new requests are not lost or interrupted;

[0047] This invention obtains the runtime distribution of unfinished tasks on the instance to be split at a fixed sampling period, and continuously calculates the instance retirement emptying entropy function by combining the historical average task processing rate. When the value of this entropy function drops below a preset threshold and remains stable, it is accurately determined that all tasks have been emptyed and no new tasks have entered. At this time, resource reclamation is performed. This allows the timing of old instance reclamation to be determined based on the calculated value of the emptying entropy function, achieving safe retirement with zero packet loss. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0049] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0050] Please see Figure 1 This invention provides a method for AI model pooling and sharing and online inference, comprising the following steps:

[0051] S1: Obtain a set of trained AI models, extract the common computing core of this set of AI models, build an instance pool containing multiple instances, each instance is loaded with the common computing core and the private header of this set of AI models, and preheat each instance to a hot standby state.

[0052] S2: In response to the online inference request, select an instance from the instance pool, call the instance's public computing core to calculate the intermediate feature tensor against the input tensor of the online inference request, call the private head of the target AI model specified in the online inference request to calculate the inference result against the intermediate feature tensor, and calculate the auxiliary gradient of the auxiliary loss function of the target AI model's private head on the intermediate feature tensor. Store the intermediate feature tensor, auxiliary gradient, and the identifier of the target AI model as a record in the instance's buffer.

[0053] S3: Based on the records in the buffer of the instance, calculate the inference manifold tension field tensor of the instance, and summarize the shared stress of the instance. When the shared stress exceeds a preset stress threshold, determine that the instance needs to be split.

[0054] S4: Mark the instance that needs to be split as the instance to be split. For the newly arrived online inference request, traverse all active instances in the instance pool as non-triggered split instances. Calculate the inference inertia flux of the newly arrived online inference request from the instance to be split to each non-triggered split instance. Take the minimum value among them. When the minimum value exceeds the preset flux threshold, stop allocating the newly arrived online inference request to the instance to be split and store the newly arrived online inference request in the micro buffer queue.

[0055] S5: Create two new instances and load a common computing core. Divide the private header on the instance to be split and load it into the two new instances. Continuously calculate the instance retirement and emptying entropy function for the instance to be split. When the value of the instance retirement and emptying entropy function drops below a preset critical value and remains stable, reclaim the resources occupied by the instance to be split, activate the two new instances to join the instance pool, and distribute the online inference requests in the micro-buffer queue to the two new instances for processing.

[0056] Specifically, firstly, a set of trained AI models is acquired, and the common computational core of this set of AI models is extracted. An instance pool containing multiple instances is constructed, with each instance loaded with the common computational core and the private header of the AI ​​model set. Each instance is then preheated to a hot standby state. Once the online inference service begins, in response to an online inference request, an instance from the instance pool is selected. The instance's common computational core is invoked to compute intermediate feature tensors on the input tensor. The private header corresponding to the target AI model is invoked to compute the inference result on the intermediate feature tensor and return it. Simultaneously, the auxiliary gradient of the auxiliary loss function on the intermediate feature tensor is calculated. The intermediate feature tensor, auxiliary gradient, and the identifier of the target AI model are stored as a record in the instance's buffer. Based on the record in the buffer, the inference manifold tension field tensor of the instance is calculated and summarized to obtain the total... When the shared stress exceeds a preset stress threshold, the instance is determined to need to be split. Instances determined to need to be split are marked as instances to be split. For newly arriving online inference requests, the computation request is migrated from the instance to be split to the inference inertia flux of instances in the instance pool that have not triggered splits. When the inference inertia flux exceeds a preset flux threshold, the allocation of the request to the instance to be split is stopped, and the request is stored in the micro-buffer queue. Two new instances are created and a common computation core is loaded. The private header on the instance to be split is divided and loaded into the two new instances. The instance retirement and emptying entropy function is continuously calculated for the instance to be split. When the value of the function drops below a preset critical value and remains stable, the resources occupied by the instance to be split are reclaimed, the two new instances are activated and added to the instance pool, and the online inference requests in the micro-buffer queue are distributed to the two new instances for processing.

[0057] In one embodiment of the present invention, step S1 includes the following steps:

[0058] Obtain a set of trained AI models, perform subgraph isomorphism detection on the inference computation graphs of each AI model, take the longest continuous prefix subgraph that is completely consistent from the input node in the inference computation graphs of each AI model as the common computation kernel, and take the remaining subgraphs as the private headers of each AI model; create a preset number of instances on the specified hardware, and load the common computation kernel and the private headers of each of the AI ​​models in this set of AI models.

[0059] For each instance, a complete forward inference is performed using an empty tensor through a public computation kernel and a private header, triggering the instance to perform memory allocation, operator fusion, and just-in-time compilation of the computation kernel, and enter a hot standby state.

[0060] Specifically, a set of pre-trained AI models is acquired, with at least two models in each set. For example, in a practical application scenario, three pre-trained AI models might be acquired: the first is a BERT-based model for text sentiment classification, the second is a BERT-based model for text topic classification, and the third is a BERT-based model for named entity recognition. The number of AI models acquired is not limited to three; in actual deployment, any number of pre-trained AI models sharing the same basic network structure can be acquired based on business needs.

[0061] Export the inference computation graph for each acquired AI model. The inference computation graph is a directed acyclic graph (DAG). The nodes in the graph represent computation operators, such as matrix multiplication operators, activation function operators, and layer normalization operators. The directed edges in the graph represent the flow direction of tensor data between operators. The inference computation graph completely describes the entire computation process from input tensor to output result.

[0062] After obtaining the inference computation graphs of all AI models, subgraph isomorphism detection is performed on the inference computation graphs of each AI model. Subgraph isomorphism detection is a graph matching technique used to determine whether the structure of one graph is completely contained within another graph. In this step, subgraph isomorphism detection starts from the input nodes of each inference computation graph and compares them one by one along the data flow direction. Specifically, the first operator node of each inference computation graph is first extracted, and its operator type and operator parameter configuration are compared; if the operator type and parameter configuration of the first operator node of all inference computation graphs are completely consistent, the comparison continues to the second operator node, and so on, until an operator node is found to be different in all inference computation graphs, or the end of an inference computation graph is reached.

[0063] During the alignment process, the continuous prefix portion of all inference computation graphs, starting from the input node and completely identical across operators, is determined, and the longest continuous prefix subgraph is selected. "Longest" means that if an inconsistency occurs at the M-th operator node, the subgraph from the input node to the (M-1)-th operator node is the longest continuous prefix subgraph. This longest continuous prefix subgraph is used as the common computation kernel, where M is a positive integer representing the sequence number of the operator node counting from the input node. Taking the three AI models mentioned above as examples, since they are all based on the BERT-base network structure, their embedding layers and multiple Transformer encoder layers constitute the same prefix network. Assuming that after subgraph isomorphism detection, it is found that the inference computation graphs of the three AI models are completely identical from the input node to the output of the 12th Transformer encoder layer, this part is used as the common computation kernel. The common computation kernel accepts an input tensor and outputs an intermediate feature tensor. The input tensor is the text sequence representation after word segmentation and embedding, and the intermediate feature tensor is the high-dimensional feature vector obtained after computation by the shared network layer. For each AI model, the remaining subgraph portion of the inference computation graph after the common computation kernel is treated as the private head of the corresponding AI model. Each private head accepts the intermediate feature tensor output by the common computation kernel as input and outputs the final inference result of that AI model. For example, the private head of the first AI model mentioned above is a linear layer for sentiment classification, with its input being the intermediate feature tensor and its output being the sentiment category probability distribution.

[0064] Creates a preset number of N instances on specified hardware. Specified hardware refers to one or more physical computing devices allocated for running the instance pool. These devices are either GPUs (Graphics Processing Units) or CPUs (Central Processing Units). Each physical computing device contains a certain amount of video memory or RAM for storing model parameters and intermediate calculation results. N is an integer between 2 and 8; for example, N = 4 creates 4 instances. The preset number can be determined based on the ratio of the total video memory or RAM of the hardware to the resources required by a single instance, with a certain margin. An instance is a runtime unit that can independently perform inference computation tasks, formed by loading a common computing core and a set of private headers onto the specified hardware. An instance has independent video memory or RAM space on the hardware to store model parameters and intermediate calculation results, and has an independent computation flow for processing online inference requests allocated to that instance serially or in parallel.

[0065] For each of the N created instances, the same loading operation is performed. First, the network parameters of the common computational kernel are read from the storage medium and loaded into the video memory or memory space allocated to that instance; the network parameters include the weight matrices and bias vectors of each operator in the common computational kernel. Second, the network parameters of the private heads of all AI models in this set of AI models are also read and loaded into the video memory or memory space of that instance. After loading, each instance has a copy of the parameters of the common computational kernel and copies of the parameters of each private head residing in its video memory or memory. At this point, each private head is in a callable state, meaning that intermediate feature tensors can be passed to the private head and the output results can be obtained through function calls.

[0066] A warm-up operation is performed on each instance. This warm-up operation involves performing a complete forward inference on the instance using an empty tensor and discarding the output. Specifically, an empty tensor is used as input to the instance. This empty tensor is not a tensor carrying real data, but rather a tensor with the same shape as the real input tensor, but filled with arbitrary values. The shape of the empty tensor is determined according to the input specifications accepted by the AI ​​model. For example, in a text classification scenario, the shape of the empty tensor is 1 x 512, where 1 represents the batch size and 512 represents the maximum sequence length. This empty tensor is then fed into the common computation kernel of the instance, and forward propagation computation is performed layer by layer. The common computation kernel outputs an intermediate feature tensor. Then, any private head already loaded on the instance is selected, and the intermediate feature tensor is fed into that private head. Forward propagation computation is performed within the private head to obtain the inference output. This inference output is then discarded without any further processing.

[0067] During the warm-up operation, the forward inference of the empty tensor triggers a series of initialization actions for the instance. During the first forward inference, the computation framework allocates the necessary GPU memory or RAM space for each operator based on the tensor's actual shape; this process is called GPU memory allocation. Simultaneously, the framework automatically fuses mergeable operators in the computation graph, such as merging convolution operators, batch normalization operators, and activation function operators into a single computation kernel to reduce GPU memory accesses and kernel startup overhead; this process is called operator fusion. Furthermore, for computation backends supporting just-in-time (JIT) compilation, the first execution of an operator triggers JIT compilation of the computation kernel, compiling the computation logic into optimized machine code for the current hardware and caching it. Subsequent calls to the same operator can directly use the compiled kernel. All of the aforementioned GPU memory allocation, operator fusion, and JIT compilation of the computation kernel are triggered and completed by this warm-up operation. After the warm-up operation is complete, the instance enters a hot standby state, ready to immediately receive and process real online inference requests without requiring time-consuming resource allocation and compilation operations during the first real inference. A preheating operation is performed on each of the N instances; once all N instances have completed preheating and entered hot standby state, these N instances together constitute the instance pool.

[0068] In one embodiment of the present invention, step S2 includes the following steps:

[0069] In response to an online inference request, the number of currently executing inference tasks for each instance in the instance pool is traversed, and the instance with the fewest inference tasks is selected. The online inference request is then assigned to that instance. The input tensor of the online inference request is fed into the common computation kernel of that instance, and after being transformed layer by layer by each operator layer in the common computation kernel, an intermediate feature tensor is output.

[0070] Based on the target AI model identifier in the online inference request, the corresponding private head is found from the private head loaded by the instance. The intermediate feature tensor is sent into the private head, and after being transformed layer by layer by the operators in the private head, the inference result is output and returned.

[0071] An auxiliary loss function is constructed at the output of the private head. The construction method of the auxiliary loss function depends on the task type of the AI ​​model corresponding to the private head. When the task type is classification, the cross-entropy between the probability distribution of the private head output and the uniform distribution is calculated as the auxiliary loss function. When the task type is regression, the mean square error between the predicted value of the private head output and the preset constant value is calculated as the auxiliary loss function. The auxiliary gradient of the auxiliary loss function on the intermediate feature tensor is calculated by automatic differentiation. The backpropagation of the auxiliary gradient terminates at the intermediate feature tensor, and the intermediate feature tensor is used as the termination node of the backpropagation.

[0072] Specifically, after all N instances in the instance pool have entered a hot standby state, they begin receiving online inference requests from external sources. Each online inference request carries an input tensor and a target AI model identifier. The target AI model identifier is used to specify which AI model in a set of AI models will produce the final inference result. The input tensor is a multidimensional array formed after preprocessing of the external raw data, which can be directly fed into the common computing kernel. The generation process of the input tensor is as follows: the external raw data is segmented into a word sequence by a word segmenter. The word segmenter assigns a word ID to each word in the word sequence. The word ID is the index number of the word in the vocabulary. The word sequence is converted into a word ID sequence, which is an integer sequence. The start marker word ID is added to the beginning of the integer sequence, and the end marker word ID is added to the end. The integer sequence is padded or truncated to a preset maximum sequence length to obtain a fixed-length one-dimensional integer array. Then, the one-dimensional integer array is expanded by a batch dimension, and each integer word ID is mapped to a fixed-dimensional floating-point vector through the embedding layer to obtain the input tensor.

[0073] Taking the BERT-base model's reasoning on the text "The plot of this movie is very exciting" as an example. The original text is segmented into a sequence of tokens by the BERT tokenizer. The BERT tokenizer assigns a corresponding token ID to each token, resulting in a sequence of token IDs. A start marker (ID 101) is added to the beginning of the token ID sequence, and an end marker (ID 102) is added to the end. The maximum sequence length of the BERT-base model is 512. If the length of the integer sequence after adding the start and end marker token IDs is less than 512, zero values ​​are padded to the end of the sequence until the length reaches 512; if it exceeds 512, the excess part is truncated. After obtaining a one-dimensional integer array of length 512, a batch dimension is expanded to obtain a two-dimensional integer tensor of shape 1×512, where 1 is the batch size. Then, through an embedding layer mapping, the embedding dimension is 768, resulting in a floating-point tensor of shape 1×512×768 as the input tensor.

[0074] When an online inference request arrives, iterate through all N instances in the instance pool to count the number of currently executing inference tasks. Each instance maintains a task counter, the value of which represents the number of online inference tasks currently being executed but not yet completed and returning inference results. The task counter is incremented by one for each new request assigned to that instance. The task counter is decremented by one after each request is processed and returns a result. Read the current value of the task counter for each instance and find the instance with the smallest task counter value. If multiple instances have the same smallest task counter value, select the instance that ranks highest in the instance pool list. Assign the current online inference request to this selected instance.

[0075] The input tensor of the online inference request is fed into the common computation kernel of the selected instance. Based on the target AI model identifier in the online inference request, the private head corresponding to that identifier is located from the private heads loaded by that instance. During the loading operation, a key-value mapping relationship has been established between each private head and its corresponding target AI model identifier, with each target AI model identifier uniquely mapped to one private head. Through this mapping relationship, the corresponding private head is located using the target AI model identifier as the key. The intermediate feature tensor is then fed into this private head. The private head contains operator layers unique to the AI ​​model, such as a fully connected linear layer followed by a softmax activation function layer. The intermediate feature tensor undergoes layer-by-layer transformation through the operator layers in the private head, and the private head outputs the final inference result. For example, for a sentiment classification task, the private head first performs a linear transformation on the vector corresponding to the starting marker position in the intermediate feature tensor, mapping it to a vector with a dimension equal to the number of categories, and then converts it into a probability distribution vector through a softmax function. Each element of this vector represents the probability that the input text belongs to a certain sentiment category. The inference result is returned to the client that initiated the online inference request via a callback function or network response.

[0076] Construct an auxiliary loss function at the output of this private header. Auxiliary Loss Function It is a loss function that can be calculated using only the output of the private head and a preset reference target, where k is the identifier of the target AI model; The construction depends on the task type of the AI ​​model corresponding to the private head.

[0077] When the AI ​​model corresponding to the private head is a classification task, the probability distribution vector output by the private head is used, and the cross-entropy between this probability distribution vector and a uniform distribution is calculated as the auxiliary loss function. The dimension of the uniform distribution is the same as the dimension of the probability distribution vector, and the probability value of each category in the uniform distribution is the reciprocal of the total number of categories in that dimension. For example, for a sentiment classification task, assuming there are 3 categories, the probability of each category in the uniform distribution is one-third. Let the probability distribution vector output by the private head be... Where C is the total number of categories, and the uniform distribution vector is... Then the auxiliary loss function , where j is a temporary variable for iterating through the category index, and the value of j is an integer from 1 to C.

[0078] When the task type of the AI ​​model corresponding to the private head is a regression task, the predicted value output by the private head is taken. The mean square error between the predicted value and a preset constant value c is calculated as an auxiliary loss function. The default constant value is the mean of the training data labels used by the AI ​​model during the training phase. For example, for a regression model that predicts ratings, assuming the mean of the ratings in the training data is 3.5, the default constant value is 3.5.

[0079] The auxiliary gradient of the auxiliary loss function with respect to the intermediate feature tensor is calculated using automatic differentiation. Automatic differentiation starts with the auxiliary loss function and propagates the gradient back along the computation graph of the private head. Backpropagation terminates when the gradient reaches the input node of the private head, i.e., the intermediate feature tensor. This backpropagation is terminated by specifying the intermediate feature tensor as the termination node when calling the backpropagation function. The auxiliary gradient is a tensor with the exact same shape as the intermediate feature tensor, where each element represents the magnitude and direction of the gradient of the auxiliary loss function with respect to the corresponding feature value in the intermediate feature tensor.

[0080] The intermediate feature tensor, auxiliary gradient, and target AI model identifier generated in this inference are stored as a record in the instance's buffer. The instance's buffer is initialized when the instance is created and is a fixed-capacity circular sliding window data structure. Each record in the buffer contains three fields, storing the intermediate feature tensor, auxiliary gradient, and target AI model identifier, respectively.

[0081] In one embodiment of the present invention, step S3, which involves calculating the inference manifold tension field tensor of the instance based on the records in the buffer of the instance, includes the following steps:

[0082] For each record in the buffer of the instance, obtain the intermediate feature tensor z and auxiliary gradient of that record. Calculate the normalization factor ;

[0083] The intermediate feature tensor z is used as the feature point p, and the auxiliary gradient is used as well. Calculate the tensor of the inference manifold tension field at the feature point p. Where k is the identifier of the target AI model, The buffer record is a set of identifiers for the different target AI models involved. for The number of AI models in China.

[0084] Specifically, during the continuous operation of the S2 online inference service, for each instance in the instance pool that is in an active service state, the calculation of its shared stress is triggered at a fixed evaluation period. The fixed evaluation period is a preset time interval, such as once every 2 seconds. When the evaluation period arrives, the number of records currently stored in the instance's buffer is checked. If the number of records is less than the preset minimum number of records (e.g., 16), the calculation is skipped, and the process waits for the next evaluation period. If the number of records reaches or exceeds the preset minimum number of records, the following calculation steps are performed.

[0085] Read all currently stored records in the buffer of this instance. Each record contains three fields: the intermediate feature tensor z, the auxiliary gradient, and so on. And the target AI model identifier k. Using k as the key, group all records in the buffer by model, and count the different target AI model identifiers involved to obtain a set. , where i is the identifier of the instance, It is a subset of the total set of AI model identifiers loaded by this instance; express The number of different target AI model identifiers in the buffer record, i.e. the number of different AI models actually involved.

[0086] Calculate auxiliary gradient norm The norm is taken as the L2 norm, which is the square root of the sum of squares of all elements in the auxiliary gradient tensor. If... If the value is zero, skip this record and do not participate in subsequent calculations. If it is not zero, then calculate the normalization factor. .

[0087] The intermediate feature tensor z in each record is treated as a feature point p, where p is numerically equivalent to z. Correspondingly, an auxiliary gradient is calculated with z as the independent variable. At feature point p, the corresponding value is... The two are completely identical in value.

[0088] With normalization factor With auxiliary gradient The product of and is used as the weighted gradient The weighted gradient directions of each model reflect their expected adjustment directions for shared features. The degree of difference in these directions characterizes the intensity of feature-pulling conflict among multiple models at point p, which is mathematically measured by calculating the covariance matrix of each weighted gradient. Based on this, the inference manifold tension field tensor T(p) is calculated at feature point p. When the auxiliary gradient directions of all AI models at p are completely consistent, the weighted gradients of each model all point in the same direction. In this case, T(p) is a zero matrix, indicating that there is no multi-model shared conflict at this feature point. T(p) is a symmetric square matrix with both rows and columns equal to the dimension z. Each diagonal element of T(p) represents the square of the deviation of the weighted gradient of each AI model from the average gradient in the corresponding feature dimension, and each off-diagonal element of T(p) represents the correlation between gradient deviations between two different feature dimensions. For each record participating in the calculation in the buffer, a T(p) is calculated at the corresponding feature point p.

[0089] In one embodiment of the present invention, step S3, which involves summarizing the shared stress of the instance, includes the following steps:

[0090] The intermediate feature tensor in each record participating in the calculation of the inference manifold tension field tensor in the buffer is taken as a feature point, and all feature points constitute a feature point set. ,remember for The total number of feature points in the matrix, for each feature point p∈ Calculate the trace of the tension field tensor T(p) of the inference manifold Then, the shared stress of the instance is calculated. .

[0091] Specifically, a feature point set is formed by assembling all feature points p involved in the calculation of the tensor T(p) of the inference manifold tension field. . Each element in the array corresponds to a feature point p. for The total number of feature points, i.e. the number of valid records participating in this shared stress calculation.

[0092] right For each feature point p in the matrix, calculate the trace of its corresponding T(p), denoted as [T(p)]. T(p) is a symmetric square matrix with dimension D. The element in the d-th row and d-th column of T(p) is denoted as... Where d is an integer from 1 to D; This represents the square of the deviation of the weighted gradient of each AI model from the average gradient on the d-th feature dimension. The sum of all diagonal elements of T(p), i.e. The trace is a scalar, reflecting the sum of shared conflict intensity across all feature dimensions for each AI model at point p. For example, if the dimension D of the intermediate feature tensor z is 768, then T(p) is a symmetric matrix of 768 rows and 768 columns, and its trace is a floating-point value obtained by adding the 768 diagonal elements. The traces corresponding to all feature points are calculated as follows. Sum the results and then divide by the total number of feature points. The shared stress of the corresponding instance is obtained. , where i is the identifier of the instance.

[0093] Will Compared with the preset stress threshold Compare. If This indicates that the sharing conflicts among the AI ​​models within this instance are within an acceptable range. The evaluation concludes without triggering a split, and the instance continues to process online inference requests in its original state. If the multi-model sharing conflict within the instance exceeds an acceptable range, then a split is required. The determination method is as follows: After the instance pool is built and warmed up, a set of test data consistent with the distribution of online inference requests is selected. Several rounds of inference are performed on each instance, and the shared stress value of each instance after each round of inference is calculated. The maximum value of the shared stress of each instance under stable operating conditions is taken and multiplied by a preset safety factor as the result. The typical value for the safety factor is 1.2. Its purpose is to provide a margin when there are shared stress fluctuations, so as to avoid unnecessary splitting caused by instantaneous fluctuations.

[0094] In one embodiment of the present invention, step S4 includes the following steps:

[0095] The status flag of the instance to be split in the instance pool is switched from active to pending split, so that it no longer receives new online inference requests; online inference requests that have been allocated to the instance to be split but have not yet been completed continue to be processed;

[0096] For each newly arrived online inference request after the state flag of the instance to be split is switched to the state flag after the split, all instances in the instance pool that are active are traversed as non-split instances. The inference inertia flux of the newly arrived online inference request from the instance to be split to each non-split instance is calculated, and the minimum value is taken.

[0097] Inference inertia flux is the ratio of the context reconstruction cost of the newly arrived online inference request to the computational margin of the non-triggered split instance, where the context reconstruction cost is equal to the product of the shared stress of the instance to be split and the total number of input tensor elements and the normalization coefficient of the newly arrived online inference request, and the computational margin is the maximum number of concurrent processing tasks of the non-triggered split instance minus the number of inference tasks currently being executed.

[0098] When the minimum value exceeds the preset throughput threshold, the allocation of the newly arrived online inference request to the instance to be split is stopped, and it is not allocated to any instance that has not triggered a split. The newly arrived online inference request is stored in the micro buffer queue.

[0099] Specifically, instances exceeding a preset stress threshold are marked as pending split instances, and these instances will no longer receive new online inference requests. Pending split instances maintain a long-tail task counter, recording the number of currently incomplete long-tail tasks; each time a long-tail task completes on this instance and returns an inference result, the long-tail task counter is decremented; all long-tail tasks continue to be processed normally by this pending split instance. The instance's status flag in the instance pool is switched from active to pending split. Each instance in the instance pool maintains a status flag field, with a value of active or pending split. Instances marked as active belong to the available instance list, and the scheduler can allocate new online inference requests to them. After the status flag is switched to pending split, the instance is removed from the available instance list and added to the pending decommissioning instance list; the scheduler will no longer allocate any new online inference requests to it.

[0100] From the moment the state flag of the instance to be split changes to "pending split," all subsequent online inference requests are considered newly arrived online inference requests. When these newly arrived online inference requests enter the scheduling and allocation process, the available instance list no longer includes the instance to be split, and the scheduler needs to re-determine the allocation destination for each newly arrived online inference request. For each newly arrived online inference request, all instances marked as active in the instance pool are traversed, and these instances are treated as instances u that have not yet triggered a split.

[0101] Calculate the inference inertia flux of newly arriving online inference requests migrating from the pending split instance r to each non-triggered split instance u. Where x is the input tensor of the newly arrived online inference request, The cost of context reconstruction for this request, This represents the computing power margin of u.

[0102] Context reconstruction cost Shared stress proportional to r The product of the total number of input tensor elements in the newly arrived online inference request. The total number of input tensor elements is L(x), which is the product of the sizes of each dimension of the input tensor. For example, if the shape of the input tensor is 1×512×768, then L(x) = 1×512×768 = 393216. Shared stress The larger L(x) is, the more severe the conflict between the multi-model feature traction directions on the instance to be split, the more drastic the distortion of the intermediate feature space, and the more implicit context information is lost when migrating the request from that instance. The larger L(x) is, the larger the computational scale of the request, and the greater the migration cost. Context reconstruction cost The preset normalization constant β takes the value of The order of magnitude is determined by calibrating the actual response latency of the deployment environment to ensure that the migration cost and computing power margin are comparable, and can be determined through a limited number of offline experiments.

[0103] u's computing power margin Let u be the current capacity of u to accept new tasks, and its value is the maximum number of concurrent tasks that u can process. Subtract the number of inference tasks it is currently performing. During instance pool initialization, stress testing was used to determine the optimal approach. The number of concurrent requests was gradually increased for each instance, and the number of concurrent requests at which the inference latency began to exceed the preset latency limit was recorded as the baseline. The preset latency limit refers to the maximum tolerance value set by the online inference service for the inference latency of a single request. Exceeding this value is considered as a failure to meet service quality standards. The typical value of the preset latency limit is 100 milliseconds, and its specific value is determined according to the service quality requirements of the online inference service. For example, 50 milliseconds can be used in real-time dialogue scenarios. Provided by a task counter maintained in real time by u. .like ≤0 indicates that u has reached or exceeded its maximum concurrent processing capacity and is not capable of handling new tasks. In this case, let .

[0104] For each instance u that has not triggered a split, calculate... Then, take the minimum value among them, and denote it as... . This indicates the minimum cost to migrate a newly arriving online inference request from the instance r to which a split has not yet been triggered.

[0105] Preset flux threshold Determined through offline testing; during normal instance pool operation, a scenario of instance splitting was simulated, and the inference inertia throughput value for different types of inference requests migrating to different instances was calculated. The maximum throughput value that could be tolerated without causing request timeouts was recorded, and the average value was taken as the threshold. ; The value is 1.0. When A value ≤ 1.0 indicates that the migration cost does not exceed the computing power margin of the target instance, and the migration is feasible. With preset flux threshold Compare. If This indicates that there is at least one instance that has not triggered a split and can accept the newly arrived online inference request at an acceptable cost. The scheduler then allocates the request to the instance that has triggered a split. Take the minimum value of u. If This indicates that all instances that have not triggered a split are unable to accept the newly arrived online inference request at an acceptable cost; at this time, a traffic freeze operation is performed, stopping the allocation of the request to r and also preventing it from being allocated to any instances that have not triggered a split, and storing the newly arrived online inference request in the microbuffer queue.

[0106] The micro-buffer queue is a first-in, first-out (FIFO) temporary request queue used to temporarily store online inference requests that cannot be immediately allocated due to inference inertia throughput exceeding limits. The micro-buffer queue is independent of the inference task queues of each instance and is managed uniformly by the scheduler. The micro-buffer queue has a preset maximum capacity, which is determined based on the available memory resources of the deployment environment and the average data volume of a single request. If the micro-buffer queue is full and there are still new online inference requests that need to be enqueued, the acceptance of new requests is temporarily blocked until there is free space in the micro-buffer queue. During the blocking period, active instances in the instance pool continue to process already allocated requests; once processing is complete, the queue is released, and the blocking is automatically lifted.

[0107] In one embodiment of the present invention, the step of continuously calculating the instance retirement and emptying entropy function for the instance to be split includes the following steps:

[0108] The runtime of all inference tasks that have not yet been completed on the instance to be split is obtained at a fixed sampling period, and a normalized task runtime distribution p(τ,t) is constructed, where τ is the task runtime and t is the sampling time.

[0109] Obtain the historical average task processing rate γ of the instance to be split, and define the cumulative distribution of task completion probability under ideal empty state as follows: ;

[0110] Calculate the instance decommissioning and emptying entropy function at sampling time t. The formula is:

[0111]

[0112] in, The idle time of the instance to be split since it last received a new task; The rate at which new tasks are assigned to the instance to be split at time s.

[0113] Specifically, on the same designated hardware as the instance pool, using the GPU or CPU computing units on that hardware that have not yet been allocated and can be used to create new instances, along with their associated video memory or memory space, two new runtime copies are created as two new instances, denoted as the first new instance and the second new instance, respectively. Both new instances load the same common computing core as the instance r to be split. The loading process includes reading the network parameters of the common computing core from the model repository and copying them to the video memory or memory space allocated to each new instance, and loading the executable code of the common computing core into the computing engine of each new instance.

[0114] Read the buffer used by `r` when calculating shared stress and count the frequency of occurrence of each target AI model identifier within the buffer. Following the principle of balancing the total request frequency of the AI ​​models corresponding to the private heads on the two new instances as much as possible, divide all the private heads originally loaded on `r` into two groups. The division method is as follows: sort all AI models in descending order of frequency, and use a greedy strategy to sequentially assign the private head of each AI model to the group with the smaller current total frequency, until all private heads have been assigned; load the first group of private heads into the first new instance, and the second group of private heads into the second new instance. Perform a warm-up operation on each of the two new instances. After warm-up, the two new instances will not receive any real inference requests.

[0115] For each split instance r, its instance retirement and emptying entropy function is continuously calculated. The calculation process is performed with a fixed sampling period of 100 milliseconds, i.e., 10 samples per second. At each sampling time t, the runtime of all inference tasks on r that have not yet been completed is obtained. For each currently executing task, its runtime τ from the start of execution to the current sampling time t is recorded. The τ values ​​of all currently executing tasks are statistically analyzed to construct a task runtime distribution p(τ,t). Specifically, the range of task runtime values ​​is divided into several equally wide intervals, the number of tasks falling within each interval is counted, and the result is divided by the total number of currently executing tasks to obtain the probability value corresponding to each interval, forming a normalized histogram. This histogram is the discrete estimate of p(τ,t).

[0116] The historical average task processing rate γ of r is obtained by dividing the total number of inference tasks completed by r since its creation by the total service time of r since its creation. Service time is the length of time from when r finishes warming up and begins processing the first online inference request until the current sampling time t. For example, if r has completed 10,000 inference tasks since its creation, and the total service time is 2,000 seconds, then γ = 10,000 / 2,000 = 5 tasks per second. The cumulative distribution of task completion probability under ideal empty state is defined as follows: , where τ is the duration the task has been running.

[0117] r Idle time since the last new task was received The calculation method is to record the moment when r last received a new allocation request, and subtract that moment from the current sampling time t. The allocation rate function λ(s) represents the instantaneous rate at which the scheduler allocates new tasks to r at time s. Its value is obtained by the scheduler counting the number of requests allocated to r per unit time. After the traffic freeze operation in S4, λ(s) is always zero.

[0118] Calculate the instance decommissioning and emptying entropy function at sampling time t. The calculation uses a histogram for discretization: for each interval of the histogram, let the probability value of that interval be... The reference distribution is located at the midpoint of this interval. The value at that location is ,calculate Summing up the results of all intervals and taking the negative value, we get The value of λ(s). Under normal operating conditions where the freeze is complete, λ(s) is always zero. The value is determined solely by the discretization calculation results described above. It decreases monotonically as the existing tasks on the instance to be split are gradually completed, and approaches zero when all long-tail tasks have been processed.

[0119] In one embodiment of the present invention, in step S5, when the value of the instance retirement entropy function drops below a preset threshold and remains stable, the resources occupied by the instance to be split are reclaimed, two new instances are activated to join the instance pool, and the online inference requests in the micro-buffer queue are distributed to the two new instances for processing, including the following steps:

[0120] When an instance retires, the entropy function is emptied. When the value is less than the preset threshold and is true for two consecutive fixed sampling periods, it is determined that the instance to be split has completed all task emptying and no new tasks have infiltrated.

[0121] Reclaim the video memory and memory resources occupied by the instance to be split, and remove the instance to be split from the instance pool;

[0122] Mark the two new instances as active and add them to the list of available instances in the instance pool;

[0123] Online inference requests are retrieved one by one from the micro-buffer queue in a first-in, first-out order. Based on the target AI model identifier carried by each online inference request, the online inference request is distributed to a new instance loaded with the corresponding private header for processing.

[0124] Specifically, preset threshold values The value is 0.05, which indicates that all long-tail tasks on the instance to be split have been basically processed, and the deviation between the actual task backlog distribution and the ideal emptying distribution is negligible. This is calculated in each sampling period. After that, and Compare the results and record them.

[0125] To judge Whether the state remains stably below a preset threshold is determined by setting stability criteria. Less than And when it is true in two consecutive sampling periods, that is, when it is obtained from three samples (including the first triggered sample) within 200 milliseconds. The values ​​are all less than It was determined that the instance to be split had completed all task emptying and no new tasks had infiltrated.

[0126] For example, the first sampling time Calculated =0.03, less than 0.05, recorded as satisfied; second sampling time Calculated =0.02, less than 0.05, satisfied. At this point, the condition is met for two consecutive sampling periods, and the emptying is considered complete. If If the value rises back to 0.08 due to a momentary fluctuation, the condition is no longer met, so the count is reset and the system waits for the next time. Accumulation restarts when the value is less than 0.05.

[0127] After the decommissioning is completed, all resources occupied by the instance to be split, r, are immediately reclaimed. The reclaimed resources include the allocated GPU memory and RAM space. GPU memory stores network parameters for the common computing cores, network parameters for all private headers, and intermediate computation results generated during inference. RAM space stores runtime data structures such as the instance's task counter and buffers. GPU memory is released by calling the GPU memory release interface provided by the computing framework, returning the GPU memory to the GPU's memory pool. RAM space is released by calling the RAM release interface provided by the operating system, returning the RAM to the operating system. The computing engine handle of r is closed; if r is running as an independent process or thread, the corresponding process or thread is terminated. The record for r is removed from the list of instances to be decommissioned in the instance pool, and the instance is completely destructed.

[0128] After the resource reclamation operation of r is completed, the two new instances that have been preheated are marked as active. The status flag is changed by switching the status flag field of the two new instances from preheated to active. These two new instances are then added to the available instance list of the instance pool, which is the list of instances that the scheduler iterates through when allocating online inference requests.

[0129] All online inference requests temporarily stored in the micro-buffer queue are distributed. The distribution process follows a first-in, first-out (FIFO) order, meaning that online inference requests are retrieved one by one starting from the head of the micro-buffer queue, with requests that were stored in the micro-buffer queue earlier being retrieved and allocated first.

[0130] For each online inference request retrieved from the micro-buffer queue, its target AI model identifier is read. Based on this identifier, it is determined which group its corresponding private header was assigned to during the private header partitioning process. If the private header belongs to the first group and is loaded into the first new instance, the online inference request is assigned to that instance. If it belongs to the second group and is loaded into the second new instance, the request is assigned to that instance. If the private header needs to be loaded into two new instances simultaneously for load balancing, the instance with the smaller current task counter value is selected for allocation based on the principle of minimum task count. After all online inference requests in the micro-buffer queue have been distributed, the queue is cleared, and subsequent new online inference requests are then normally distributed by the scheduler to all active instances in the instance pool.

[0131] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A method for implementing AI model pooling and online inference, characterized in that, Includes the following steps: S1: Obtain a set of trained AI models, extract the common computing core of this set of AI models, build an instance pool containing multiple instances, each instance is loaded with the common computing core and the private header of this set of AI models, and preheat each instance to a hot standby state. S2: In response to the online inference request, select an instance from the instance pool, call the instance's public computing core to calculate the intermediate feature tensor against the input tensor of the online inference request, call the private head of the target AI model specified in the online inference request to calculate the inference result against the intermediate feature tensor, and calculate the auxiliary gradient of the auxiliary loss function of the target AI model's private head on the intermediate feature tensor. Store the intermediate feature tensor, auxiliary gradient, and the identifier of the target AI model as a record in the instance's buffer. S3: Based on the records in the buffer of the instance, calculate the inference manifold tension field tensor of the instance, and summarize the shared stress of the instance. When the shared stress exceeds a preset stress threshold, determine that the instance needs to be split. S4: Mark the instance that needs to be split as the instance to be split. For the newly arrived online inference request, traverse all active instances in the instance pool as non-triggered split instances. Calculate the inference inertia flux of the newly arrived online inference request from the instance to be split to each non-triggered split instance. Take the minimum value among them. When the minimum value exceeds the preset flux threshold, stop allocating the newly arrived online inference request to the instance to be split and store the newly arrived online inference request in the micro buffer queue. S5: Create two new instances and load a common computing core. Divide the private header on the instance to be split and load it into the two new instances. Continuously calculate the instance retirement and emptying entropy function for the instance to be split. When the value of the instance retirement and emptying entropy function drops below a preset critical value and remains stable, reclaim the resources occupied by the instance to be split, activate the two new instances to join the instance pool, and distribute the online inference requests in the micro-buffer queue to the two new instances for processing.

2. The AI ​​model pooling and sharing and online inference implementation method according to claim 1, characterized in that, S1 includes the following steps: Obtain a set of trained AI models, perform subgraph isomorphism detection on the inference computation graphs of each AI model, take the longest continuous prefix subgraph that is completely consistent from the input node in the inference computation graphs of each AI model as the common computation kernel, and take the remaining subgraphs as the private headers of each AI model; create a preset number of instances on the specified hardware, and load the common computation kernel and the private headers of each of the AI ​​models in this set of AI models. For each instance, a complete forward inference is performed using an empty tensor through a public computation kernel and a private header, triggering the instance to perform memory allocation, operator fusion, and just-in-time compilation of the computation kernel, and enter a hot standby state.

3. The AI ​​model pooling and sharing and online inference implementation method according to claim 1, characterized in that, S2 includes the following steps: In response to an online inference request, the number of currently executing inference tasks for each instance in the instance pool is traversed, and the instance with the fewest inference tasks is selected. The online inference request is then assigned to that instance. The input tensor of the online inference request is fed into the common computation kernel of that instance, and after being transformed layer by layer by each operator layer in the common computation kernel, an intermediate feature tensor is output. Based on the target AI model identifier in the online inference request, the corresponding private head is found from the private head loaded by the instance. The intermediate feature tensor is sent into the private head, and after being transformed layer by layer by the operators in the private head, the inference result is output and returned. An auxiliary loss function is constructed at the output of the private head. The construction method of the auxiliary loss function depends on the task type of the AI ​​model corresponding to the private head. When the task type is classification, the cross-entropy between the probability distribution of the private head output and the uniform distribution is calculated as the auxiliary loss function. When the task type is regression, the mean square error between the predicted value of the private head output and the preset constant value is calculated as the auxiliary loss function. The auxiliary gradient of the auxiliary loss function on the intermediate feature tensor is calculated by automatic differentiation. The backpropagation of the auxiliary gradient terminates at the intermediate feature tensor, and the intermediate feature tensor is used as the termination node of the backpropagation.

4. The AI ​​model pooling and sharing and online inference implementation method according to claim 1, characterized in that, In step S3, calculating the inference manifold tension field tensor of the instance based on the records in the buffer of the instance includes the following steps: For each record in the buffer of the instance, obtain the intermediate feature tensor z and auxiliary gradient of that record. Calculate the normalization factor ; The intermediate feature tensor z is used as the feature point p, and the auxiliary gradient is used as well. Calculate the tensor of the inference manifold tension field at the feature point p. Where k is the identifier of the target AI model, The buffer record is a set of identifiers for the different target AI models involved. for The number of AI models in China.

5. The AI ​​model pooling and sharing and online inference implementation method according to claim 4, characterized in that, In step S3, the process of summarizing the shared stress of the instance includes the following steps: The intermediate feature tensor in each record participating in the calculation of the inference manifold tension field tensor in the buffer is taken as a feature point, and all feature points constitute a feature point set. ,remember for The total number of feature points in the matrix, for each feature point p∈ Calculate the trace of the tension field tensor T(p) of the inference manifold Then, the shared stress of the instance is calculated. .

6. The AI ​​model pooling and sharing and online inference implementation method according to claim 1, characterized in that, S4 includes the following steps: The status flag of the instance to be split in the instance pool is switched from active to pending split, so that it no longer receives new online inference requests; online inference requests that have been allocated to the instance to be split but have not yet been completed continue to be processed; For each newly arrived online inference request after the state flag of the instance to be split is switched to the state flag after the split, all instances in the instance pool that are active are traversed as non-split instances. The inference inertia flux of the newly arrived online inference request from the instance to be split to each non-split instance is calculated, and the minimum value is taken. Inference inertia flux is the ratio of the context reconstruction cost of the newly arrived online inference request to the computational margin of the non-triggered split instance, where the context reconstruction cost is equal to the product of the shared stress of the instance to be split and the total number of input tensor elements and the normalization coefficient of the newly arrived online inference request, and the computational margin is the maximum number of concurrent processing tasks of the non-triggered split instance minus the number of inference tasks currently being executed. When the minimum value exceeds the preset throughput threshold, the allocation of the newly arrived online inference request to the instance to be split is stopped, and it is not allocated to any instance that has not triggered a split. The newly arrived online inference request is stored in the micro buffer queue.

7. The AI ​​model pooling and sharing and online inference implementation method according to claim 1, characterized in that, In step S5, the continuous calculation of the instance retirement and emptying entropy function for the instance to be split includes the following steps: The runtime of all inference tasks that have not yet been completed on the instance to be split is obtained at a fixed sampling period, and a normalized task runtime distribution p(τ,t) is constructed, where τ is the task runtime and t is the sampling time. Obtain the historical average task processing rate γ of the instance to be split, and define the cumulative distribution of task completion probability under ideal empty state as follows: ; Calculate the instance decommissioning and emptying entropy function at sampling time t. The formula is: in, The idle time of the instance to be split since it last received a new task; The rate at which new tasks are assigned to the instance to be split at time s.

8. The AI ​​model pooling and sharing and online inference implementation method according to claim 7, characterized in that, In step S5, when the value of the instance retirement entropy function drops below a preset critical value and remains stable, the resources occupied by the instance to be split are reclaimed, two new instances are activated and added to the instance pool, and the online inference requests in the micro-buffer queue are distributed to the two new instances for processing. This includes the following steps: When an instance retires, the entropy function is emptied. When the value is less than the preset threshold and is true for two consecutive fixed sampling periods, it is determined that the instance to be split has completed all task emptying and no new tasks have infiltrated. Reclaim the video memory and memory resources occupied by the instance to be split, and remove the instance to be split from the instance pool; Mark the two new instances as active and add them to the list of available instances in the instance pool; Online inference requests are retrieved one by one from the micro-buffer queue in a first-in, first-out order. Based on the target AI model identifier carried by each online inference request, the online inference request is distributed to a new instance loaded with the corresponding private header for processing.