A load balancing decision method, system, device and medium for high-concurrency inference of a multi-modal large model
Patent Information
- Application Number
- CN202610993304.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]因此,本发明提供一种多模态大模型高并发推理的负载均衡决策方法、系统、设备及介质,解决现有技术在处理大模型高并发推理负载均衡时存在请求特征利用不充分、性能预测准确性不足、计算开销大、缓存管理效率低的问题
[0016]与现有技术相比,本发明的有益效果为:本发明通过接收推理请求后提取语义向量并计算硬度因子,构建请求的语义与计算复杂度双重表征,克服传统方法仅依赖token静态特征的局限性;通过为每个服务器维护历史请求记录,并基于当前请求硬度因子与历史硬度因子的相似性计算初始优化利用项,实现对服务器处理同类请求性能的快速预估,避免了重复的在线探测开销;通过引入语义相似度与时间衰减机制计算预期残差补偿因子,并将初始优化利用项与残差补偿因子相加得到最佳利用项,有效修正了硬度相似性估计的系统性偏差,提升了性能预测的准确性;通过计算哈希匹配度评分和实时负载量评分,分别量化服务器缓存的复用潜力和当前任务压力,使决策兼顾历史缓存收益与实时负载均衡;最终通过加权求和得到综合决策得分并选择目标服务器,将多维度信息融合为一个统一决策指标,实现了轻量高效的路由选择。整体上,本发明无需依赖复杂的在线优化或预设固定参数,即可在请求特征分布变化时保持良好的适应性,显著降低高并发场景下的系统开销和决策延迟,有效避免调度模块成为性能瓶颈。
Smart Images

Figure CN122824748A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of inference load balancing technology, and in particular to a load balancing decision-making method, system, device and medium for multimodal large-model high-concurrency inference. Background Technology
[0002] With the rapid development of large language models and multimodal models, the concurrency of inference services is growing exponentially, placing higher demands on the accuracy of request scheduling and system throughput. In high-concurrency inference scenarios, the computational complexity of requests varies greatly, and cache hit rate directly affects inference latency. Therefore, intelligently distributing requests to appropriate servers has become crucial for improving system throughput and reducing response latency.
[0003] Currently, common request scheduling methods include traditional strategies such as round-robin and least connections. These methods do not fully consider the semantic features and computational complexity of requests, making them difficult to adapt to the dynamic characteristics of large-scale model inference. In recent years, some improved methods have attempted to utilize the token features of requests for scheduling, such as determining the similarity between requests by hash matching of token sequences, and then routing similar requests to the same server to utilize its pre-populated cache. However, these methods are based solely on the static similarity of token sequences and fail to fully utilize the semantic information of requests, resulting in a single feature dimension for scheduling decisions and limited prediction accuracy. In addition, some methods introduce a request hardness factor as a metric for computational complexity, but relying solely on the hardness factor for performance estimation ignores the impact of historical prediction bias on decision accuracy, which can easily lead to accumulated errors in high-concurrency scenarios. At the same time, existing methods often use a first-in, first-out (FIFO) strategy for managing historical requests on servers, which can easily discard valuable historical cached information, affecting prediction accuracy. Therefore, there is an urgent need for a load balancing decision method that can integrate multi-dimensional features, is lightweight and efficient, and has adaptive capabilities for multimodal large-scale model high-concurrency inference. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a load balancing decision-making method, system, device, and medium for multimodal large-model high-concurrency inference, which solves the problems of insufficient utilization of request features, insufficient accuracy of performance prediction, large computational overhead, and low cache management efficiency in the existing technology when dealing with load balancing of large-model high-concurrency inference.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a load balancing decision-making method for multimodal large-model high-concurrency inference, comprising: Receive inference requests from multimodal large models, perform lexical sequence analysis on the inference requests, extract semantic vectors, and calculate hardness factors; Maintain historical request records for each server, and calculate the initial optimization utilization for each server based on the similarity between the hardness factor of the current request and the hardness factor of the server's historical requests. Based on the semantic similarity and time decay between the current request and the server's historical requests, the expected residual compensation factor for each server is calculated. The initial optimization utilization term is added to the expected residual compensation factor to obtain the optimal utilization term for each server. Calculate the hash matching score and real-time load score for each server; The optimal utilization, hash matching score, and real-time load score of each server are weighted and summed to obtain a comprehensive decision score. The target server is selected based on the comprehensive decision score, and the current inference request is routed to the selected target server.
[0007] As a preferred embodiment of the load balancing decision-making method for multimodal large-model high-concurrency inference described in this invention, the step of performing lexical sequence analysis on the inference request, extracting semantic vectors, and calculating a hardness factor includes: The original text of the reasoning request is lexicalized to obtain a lexical sequence; Starting from a preset base number, the word sequence is divided into multiple hierarchical subsets according to an exponential growth method, and a hash value is calculated for each hierarchical subset to form the hash value set of the current request; The inference request is mapped into a semantic vector using a pre-trained sentence encoder; The hardness factor is calculated based on the number of input terms in the inference request, the expected output length, and the context complexity coefficient.
[0008] As a preferred embodiment of the load balancing decision-making method for multimodal large-model high-concurrency inference described in this invention, the step of maintaining historical request records for each server includes: Allocate a fixed amount of storage space for historical request records to each server. Each historical request record contains a set of hash values of processed requests, a semantic vector, a hardness factor, the request time, the actual processing time, and the original performance gains. Maintain a fixed-size recent performance gain time window for each server, collect the raw performance gains of all completed requests within the window, and form a recent performance gain set; In response to the storage space being full and the need to add new records, a retention score is calculated for each historical request record based on the timestamp and access frequency. The historical request records are then sorted in descending order according to the retention score. A preset proportion of historical request records are eliminated according to the sorting result, and the new records are then stored in the storage space.
[0009] As a preferred embodiment of the load balancing decision-making method for multimodal large-model high-concurrency inference described in this invention, the calculation of the expected residual compensation factor for each server includes: The adaptive state decay rate of the server is calculated based on the aforementioned recent performance gain set, and the time decay factor is calculated based on the adaptive state decay rate. For each historical request, the residual weight factor is calculated by comprehensively considering the cosine similarity and time decay factor between the semantic vectors of the current request and the historical request. The difference between the original performance gain of each historical request and the initial optimization utilization item corresponding to the server is obtained as the performance prediction residual; The weighted average of the residual weighting factors of all historical requests and the performance prediction residuals is used as the expected residual compensation factor.
[0010] As a preferred embodiment of the load balancing decision-making method for multimodal large-model high-concurrency inference described in this invention, the calculation of the hash matching score and real-time load score for each server includes: The hash values corresponding to each hierarchical subset in the hash value set of the current request are matched with the historical hash values cached by the server, and different weights are assigned according to the hierarchical level of the hierarchical subset corresponding to the successfully matched hash value. The hash matching score is calculated based on the sum of all matching weights and the theoretical maximum matching score. Obtain the request queue currently being processed by the server, and in response to each request in the request queue having a length less than a preset length threshold, assign a fixed penalty value; In response to each request in the request queue having a length greater than or equal to a preset length threshold, a multiplier penalty value is applied by rounding up the ratio of the request length to the base length. The real-time load score is obtained by summing the penalty scores of all requests.
[0011] As a preferred embodiment of the load balancing decision-making method for multimodal large-model high-concurrency inference described in this invention, the step of selecting the target server based on the comprehensive decision score includes: Based on the comprehensive decision score, a greedy strategy is used to select the server with the highest score as the target server. During system operation, a weighted loss function of load imbalance and average latency is calculated based on real-time performance indicators, and the summation weight of the comprehensive decision score is fine-tuned online using the gradient descent method.
[0012] As a preferred embodiment of the load balancing decision-making method for multimodal large-model high-concurrency inference described in this invention, the step of selecting a target server based on the comprehensive decision score and routing the current inference request to the selected target server includes: The server is divided into a pre-filled component set and a decoding component set, and a hit rate threshold is set. If the hash matching score of all servers is lower than the hit rate threshold, the inference request is determined to be a new request and routed to the one with the highest comprehensive decision score in the pre-filled component set. If the hash matching score of the existing server is not lower than the hit rate threshold, the inference request is determined to be a continuation request and routed to the one with the highest comprehensive decision score in the set of decoding components. In response to the existence of a server's hash matching score being approximately equal to the hit rate threshold, a decision is made taking into account the optimal utilization factors.
[0013] Secondly, the present invention provides a load balancing decision system for multimodal large-model high-concurrency inference, comprising: The request feature extraction module is used to receive inference requests from multimodal large models, perform word sequence analysis on the inference requests, extract semantic vectors, and calculate hardness factors. The hardness similarity prediction module is used to maintain historical request records for each server and calculate the initial optimization utilization item for each server based on the similarity between the hardness factor of the current request and the hardness factor of the server's historical requests. The compensation and optimal utilization module is used to calculate the expected residual compensation factor for each server based on the semantic similarity and time decay of the current request and the server's historical requests, and to add the initial optimized utilization term to the expected residual compensation factor to obtain the optimal utilization term for each server. The scoring calculation module is used to calculate the hash matching score and real-time load score for each server. The weighted comprehensive decision module is used to perform a weighted summation of the best utilization item, hash matching score and real-time load score of each server to obtain a comprehensive decision score; The target routing module is used to select a target server based on the comprehensive decision score and route the current inference request to the selected target server.
[0014] Thirdly, the present invention provides an electronic device, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor executes the computer-executable instructions to implement the steps of a load balancing decision method for multimodal large-model high-concurrency inference.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of a load balancing decision method for multimodal large-model high-concurrency inference.
[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention extracts semantic vectors and calculates hardness factors after receiving inference requests, constructing a dual representation of the request's semantics and computational complexity, overcoming the limitations of traditional methods that rely solely on static token features; by maintaining historical request records for each server and calculating initial optimization utilization terms based on the similarity between the current request hardness factor and historical hardness factors, it achieves rapid prediction of server performance in handling similar requests, avoiding redundant online probing overhead; by introducing semantic similarity and time decay mechanisms to calculate the expected residual compensation factor, and adding the initial optimization utilization term to the residual compensation factor to obtain the optimal utilization term, it effectively corrects the systematic bias in hardness similarity estimation and improves the accuracy of performance prediction; by calculating hash matching scores and real-time load scores, it quantifies the server cache reuse potential and current task pressure respectively, enabling decisions to consider both historical cache benefits and real-time load balancing; finally, by weighted summation to obtain a comprehensive decision score and select the target server, it integrates multi-dimensional information into a unified decision index, achieving lightweight and efficient routing selection. Overall, this invention does not rely on complex online optimization or preset fixed parameters, and can maintain good adaptability when the distribution of request characteristics changes. It significantly reduces system overhead and decision latency in high-concurrency scenarios and effectively avoids the scheduling module becoming a performance bottleneck. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the overall process logic of a load balancing decision method for multimodal large-model high-concurrency inference provided in an embodiment of the present invention.
[0019] Figure 2 The flowchart illustrates the process of obtaining the comprehensive decision score for a load balancing decision-making method for multimodal large-model high-concurrency inference, as provided in an embodiment of the present invention.
[0020] Figure 3 This diagram illustrates the target server selection for a load balancing decision-making method for multimodal large-model high-concurrency inference, as provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0022] Example 1, referring to Figure 1 As one embodiment of the present invention, a load balancing decision method for multimodal large-model high-concurrency inference is provided, such as... Figure 1 The specific steps shown are as follows: S100: Receives inference requests from multimodal large models, performs lexical sequence analysis on the inference requests, extracts semantic vectors, and calculates hardness factors; S200: Maintain historical request records for each server, and calculate the initial optimization utilization for each server based on the similarity between the hardness factor of the current request and the hardness factor of the server's historical requests. S300: Based on the semantic similarity and time decay between the current request and the server's historical requests, calculate the expected residual compensation factor for each server, add the initial optimization utilization term to the expected residual compensation factor, and obtain the optimal utilization term for each server. S400: Calculates the hash matching score and real-time load score for each server; S500: The best utilization, hash matching score and real-time load score of each server are weighted and summed to obtain a comprehensive decision score; S600: Select a target server based on the comprehensive decision score and route the current inference request to the selected target server.
[0023] It should be noted that this invention extracts semantic vectors and calculates hardness factors after receiving inference requests, constructing a dual representation of the request's semantics and computational complexity, overcoming the limitations of traditional methods that rely solely on static token features. By maintaining historical request records for each server and calculating initial optimization utilization terms based on the similarity between the current request hardness factor and historical hardness factors, it achieves rapid prediction of server performance in handling similar requests, avoiding redundant online probing overhead. By introducing semantic similarity and time decay mechanisms to calculate the expected residual compensation factor, and adding the initial optimization utilization term to the residual compensation factor to obtain the optimal utilization term, it effectively corrects the systematic bias in hardness similarity estimation and improves the accuracy of performance prediction. By calculating hash matching scores and real-time load scores, it quantifies the server cache reuse potential and current task pressure, respectively, ensuring that decisions consider both historical cache benefits and real-time load balancing. Finally, by weighted summation, a comprehensive decision score is obtained and the target server is selected, integrating multi-dimensional information into a unified decision index, achieving lightweight and efficient routing selection. Overall, this invention does not rely on complex online optimization or preset fixed parameters, and can maintain good adaptability when the distribution of request characteristics changes. It significantly reduces system overhead and decision latency in high-concurrency scenarios and effectively avoids the scheduling module becoming a performance bottleneck.
[0024] Example 2, refer to Figure 2 and Figure 3 Based on the previous embodiment, this embodiment provides a specific implementation method for a load balancing decision-making method for multimodal large-model high-concurrency inference.
[0025] In step S100 of this invention, an inference request from a multimodal large model is received. The inference request refers to natural language text content submitted by the user to the large language model inference service, including but not limited to: dialogue prompts, system instructions, historical dialogue context, and the user's currently input question or instruction. For example, in an intelligent customer service scenario, the inference request could be "Please help me check the order status," and in a code generation scenario, the inference request could be "Write a quicksort function."
[0026] In step S100 of the present invention, the reasoning request is subjected to lexical sequence analysis, semantic vector is extracted, and hardness factor is calculated, including: The original text of the reasoning request is lexicalized to obtain a lexical sequence; Starting from a preset base number, the word sequence is divided into multiple hierarchical subsets according to an exponential growth method, and a hash value is calculated for each hierarchical subset to form the hash value set of the current request; The inference request is mapped to a semantic vector using a pre-trained sentence encoder; The hardness factor is calculated based on the number of input tokens in the inference request, the expected output length, and the context complexity coefficient.
[0027] Specifically, when an LLM inference request reaches the load balancer, the request is first tokenized. Let the original request text be T; after processing by the tokenizer, a token sequence is obtained. , where n is the total number of tokens; Specifically, a hierarchical subset partitioning strategy is adopted, based on a preset cardinality. Starting from the i-th level, the token subsets are divided according to an exponential growth pattern. The number of tokens contained in the i-th level subset is: Calculate the hash value for each subset to form a hash value set. , where m is the total number of subsets; the hash function uses the MurmurHash3 algorithm, which has good distribution uniformity and computational efficiency.
[0028] Specifically, mapping inference requests to semantic vectors using a pre-trained sentence encoder includes: The request text is cleaned and standardized, including removing invalid characters, standardizing the encoding format, and truncating excessively long texts to the maximum length (typically 2048 tokens); the request text is segmented into clause sequences according to punctuation marks such as periods, question marks, and exclamation marks to facilitate the capture of local semantic features; The preprocessed text is input into the pre-trained Sentence-BERT model, which performs the following operations in sequence: converting the text into a sequence of sub-word tokens; extracting semantic representations from each layer; and averaging the hidden states of the last layer to obtain a 768-dimensional original semantic vector. The original semantic vector is L2 normalized to lie on the unit hypersphere, as expressed by the formula: in, This represents a d-dimensional semantic vector; the normalized semantic vector is bound to the request ID and stored in the cache for subsequent similarity calculations.
[0029] Specifically, the hardness factor is calculated based on the number of input tokens in the inference request, the expected output length, and the context complexity coefficient: Where n is the number of input tokens. Where C is the expected output length, and C is the context complexity coefficient. , , For the weight parameters, satisfying Typical values are .
[0030] It should be noted that the above step S100 can more comprehensively reflect the computational complexity and semantic content of the request, providing a rich information foundation for subsequent accurate load balancing decisions, avoiding misjudgments caused by single features, and thus improving the accuracy of route matching.
[0031] In step S200 of the present invention, maintaining historical request records for each server includes: Allocate a fixed amount of storage space for historical request records to each server. Each historical request record contains a set of hash values of processed requests, a semantic vector, a hardness factor, the request time, the actual processing time, and the original performance gains. Maintain a fixed-size recent performance gain time window for each server, collect the raw performance gains of all completed requests within the window, and form a recent performance gain set; In response to the storage space being full and the need to add new records, a retention score is calculated for each historical request record based on the timestamp and access frequency. The historical request records are then sorted in descending order according to the retention score. A preset proportion of historical request records are eliminated according to the sorting result, and the new records are then stored in the storage space.
[0032] Specifically, for each server Maintain historical request record set Each record contains the following field: Request ID: Token hash set: Semantic vector: Hardness factor: Request time: Actual processing time: Original performance gains: The original performance gain is defined as the reciprocal of the processing speed, that is: Specifically, a fixed-size recent performance gain time window (typically 300 seconds) is maintained for each server, and the raw performance gains of all completed requests within the window are collected to form a recent performance gain set. This set is used to calculate the server's adaptive state decay rate.
[0033] Specifically, a fixed-size memory space M is allocated to each server (typically 1024 records). When the number of records exceeds M, a hybrid eviction strategy based on timestamps and access frequency is adopted: calculating the retention score for each historical request record. Here, freq represents the access frequency of the record, that is, the cumulative number of times the cached record has been matched in history. The higher the freq value, the greater the probability that the record is queried repeatedly, and the higher its weight is given in the retention score, making it less likely to be evicted. Records with the lowest retention score (10%) are evicted, and new records are stored in the storage space.
[0034] In step S200 of the present invention, calculating the initial optimization utilization item for each server based on the similarity between the hardness factor of the current request and the hardness factors of historical requests from the server includes: For the current request and server Historical Request Calculate the hardness similarity weights: in, The variance of the historical request hardness factor. To adjust the parameter (typically 2.0), This represents the processing latency of the current request, that is, the time elapsed from when the current request reaches the load balancer until a server response is received. In the weight calculation formula, This is used to normalize the latency differences among servers, allowing servers with lower latency to receive higher allocation weights.
[0035] The initial optimization utilization term is obtained by weighted averaging the raw performance gains of historical requests: Where N is the total number of historical requests.
[0036] It should be noted that step S200 above can quickly predict the expected performance of the server in handling similar hardness requests without the need for complex online probing or additional testing. Since the hardness factor directly reflects the computational cost of the request, this performance estimation based on hardness similarity has high interpretability and reliability, providing a stable benchmark value for subsequent residual correction.
[0037] In step S300 of the present invention, based on the semantic similarity and time decay between the current request and the server's historical requests, the expected residual compensation factor for each server is calculated, including: The adaptive state decay rate of the server is calculated based on the recent performance gain set: in, For performance gain variance, This represents the average performance gain. To prevent division by zero of small constants (typical values) This decay rate reflects the degree of fluctuation in server performance; the greater the fluctuation, the faster the timeliness of historical data decays. For each historical request, the residual weight factor is calculated by comprehensively considering the cosine similarity between semantic vectors and the time decay factor: in, The semantic vector representing the current request is a 768-dimensional vector obtained by Sentence-BERT encoding and L2 normalization of the request text that has reached the load balancer. Semantic vectors used for historical requests Calculate the cosine similarity to determine the semantic relevance between the current request and historical requests; the semantic similarity is calculated using cosine similarity. Among them, time difference In seconds, This indicates the timestamp of the current request arriving at the load balancer, accurate to the millisecond level (Unix timestamp). Used for timestamps in historical requests Difference, calculate time difference To measure the freshness of historical data; The difference between the original performance gain of each historical request and the initial optimization utilization item corresponding to the historical request is obtained as the performance prediction residual, expressed by the formula: in, Request for history The initial predicted value of the optimized utilization term at that time; The weighted average of the residual weighting factors of all historical requests and the performance prediction residuals is used as the expected residual compensation factor, expressed by the formula: In step S300 of this invention, the initial optimized utilization term is added to the expected residual compensation factor to obtain the optimal utilization term for each server: ; It should be noted that step S300 above can capture performance differences between requests with similar hardness but different semantics, and adaptively weaken the weight of outdated historical data through a time decay mechanism. The resulting optimal utilization term significantly improves the accuracy of performance prediction, and is particularly suitable for online inference scenarios where request feature distribution changes over time.
[0038] In step S400 of the present invention, as Figure 2 As shown, the calculation of each server's hash match score and real-time load score includes: The hash values corresponding to each hierarchical subset in the current request's hash value set are matched with the historical hash values cached by the server, and different weights are assigned according to the hierarchical level of the hierarchical subset corresponding to the matched hash value. The hash matching score is calculated based on the sum of all matching weights and the theoretical maximum matching score. Get the request queue that the server is currently processing. If the length of each request in the request queue is less than a preset length threshold, a fixed penalty value is given. If the length of each request in the request queue is greater than or equal to a preset length threshold, a multiplier penalty value is applied by rounding up the ratio of the request length to the base length. The real-time load score is obtained by summing the penalty values of all requests.
[0039] Specifically, calculate the hash set of the current request. Matching historical hash values with the server's cache. For each matching hash value, different weights are assigned based on its corresponding token subset level: in, The number of matching hash values at the i-th layer. For the exponential growth model, the weights are set to [the appropriate weights]. This allows longer subsets of tokens to achieve higher scores; the hash matching score is calculated based on the sum of all matching weights and the theoretical maximum matching score, using the following formula: in, This represents the theoretical maximum matching score.
[0040] Specifically, for servers The current set of requests being processed Calculate the real-time load score: in, To request the length of q, the function f is defined as: in, The length threshold (typically 512), This is the base length (typical value 1024).
[0041] It should be noted that step S400 above quantifies the server cache's reuse potential for the current request by using a hierarchical token hash matching score; it employs a segmented penalty function to score the load of requests currently being processed by the server, giving short requests a fixed penalty and long requests a penalty proportional to their length. This approach can both prioritize routing requests to servers with high cache hit rates to reduce redundant calculations and dynamically balance the workload of each node based on actual load pressure, avoiding local overload or wasted cache resources due to a single-dimensional scoring method.
[0042] In step S500 of the present invention, as Figure 2 As shown, the optimal utilization score, hash matching score, and real-time load score for each server are weighted and summed to obtain the comprehensive decision score: in, , , These are the weighting parameters. A typical configuration is... .
[0043] It should be noted that the S500 steps described above aim to minimize processing time while also considering cache hit rate and system load balancing. The weight parameters can be adaptively adjusted based on system operating status using PSO or gradient descent, enabling the decision model to flexibly adapt to different cluster sizes, hardware heterogeneity, and request patterns. Compared to single-objective or fixed-rule routing strategies, this significantly improves overall throughput and reduces tail latency.
[0044] In step S600 of the present invention, routing is performed based on the comprehensive decision scores of all servers using the following strategy: ① Greedy strategy: Select the server with the highest score ; ②Top-K random strategy: Randomly select from the K servers with the highest scores to increase load distribution; ③ Softmax probabilistic strategy: Calculate the selection probability based on the score. ,in This refers to the temperature parameter.
[0045] In step S600 of the present invention, as Figure 3 The diagram also includes: dividing the server into a pre-populated set of components. and decoding component collection And set a hit rate threshold. ; If the hash matching score of all servers is lower than the hit rate threshold, the inference request is determined to be a new request and routed to the one with the highest overall decision score in the pre-populated component set. If the hash matching score of the existing server is not lower than the hit rate threshold, the inference request is determined to be a continuation request and routed to the one with the highest comprehensive decision score in the decoding component set; The decision is made by taking into account the fact that the hash matching score of the existing server is approximately equal to the hit rate threshold, and the best utilization option is considered.
[0046] In step S600 of the present invention, a particle swarm optimization (PSO) parameter tuning step is also included: The weight parameters are dynamically adjusted using the particle swarm optimization algorithm. Define the particle position as a parameter vector. Speed update formula: Where is the inertia weight. , As a learning factor, , It is a random number. g is the individual optimum, and g is the global optimum. The fitness function is defined as the reciprocal of the system's average response time: In step S600 of the present invention, an adaptive weight adjustment step is also included: the weights are fine-tuned using the gradient descent method based on real-time performance indicators; The loss function is defined as a weighted sum of load imbalance and average delay: Where L is the loss function, which is a weighted sum of load imbalance and average delay, used to guide the direction of weight gradient descent; This is the load imbalance weighting coefficient (typically 0.6), which controls the priority of load balancing optimization. This is the average delay weighting coefficient (typically 0.4), which controls the priority of delay optimization. and The sum of these values is 1; Imbalance is the load imbalance, which measures the relative degree to which the load of each server deviates from the average; Latency is the average latency, which is the arithmetic mean of the response times of all requests, in seconds.
[0047] The weight update rules are as follows: in, This represents the learning rate (typically 0.01).
[0048] It should be noted that, while ensuring that high-performance servers prioritize the processing of complex requests, the above step S600 avoids extreme imbalances caused by traffic skew through random or probabilistic strategies. Under the separate architecture, it can direct new requests to the pre-filling component and redirect cached, reusable continuation requests to the decoding component, maximizing the utilization of the KV cache and reducing redundant pre-filling calculations.
[0049] Example 3 provides an application example of a load balancing decision-making method for multimodal large-model high-concurrency inference, which verifies and illustrates the technical effectiveness of this method.
[0050] In one specific implementation, the system comprises four homogeneous servers processing text generation tasks. A request has arrived containing 256 tokens, with an expected output of 128 tokens. First, request features are extracted: the token sequence length n=256, and the expected output length... =128, context complexity C=1.2, hardness factor D=0.4×256+0.4×128+0.2×1.2=153.84, and a 768-dimensional semantic vector is generated and normalized using Sentence-BERT. Then, token subset partitioning is performed, and cardinality is... =16 and uses an exponential growth method to generate hash values corresponding to 16, 32, 64, 128, and 256 tokens in sequence. to For servers Its historical request count N=50, hardness factor variance For hardness factor Historical requests, hardness similarity weight: Initial optimization utilization items =45.2, adaptive attenuation rate λ=0.0015, expected residual compensation C=2.3, yielding the optimal utilization term. =47.5; Hash matching detected a match. , , Matching, scoring Currently, there are 2 requests being handled. =-4; Overall score Similarly, the scores for S2 were calculated to be 38.6, S3 41.2, and S4 39.8. Finally, a greedy strategy was adopted, selecting S1, which had the highest score, for request forwarding.
[0051] In one specific implementation, the system adopts a pre-filling and decoding separation architecture, comprising 3 pre-filling servers and 5 decoding servers. Upon receiving a new request with a 512-token token, feature extraction yields a hardness factor D=287.6, and a 6-layer hash value set is generated. The hash matching degree of all servers is calculated, and the highest matching degree is found. =35, below the threshold of 60, therefore it is judged as a new request. Subsequently, the composite score is calculated only within the pre-populated server set: P1's =52.3、 =-6, overall score 36.65; P2 =48.7、 =-3, overall score 34.35; P3 =50.1、 =-2, overall score 35.55. The final route is to P1 for pre-filling.
[0052] In one specific implementation, for a request that has already been pre-filled on P1 and now needs to be decoded, the first 256 tokens of the request have been cached. Hash matching is performed on the decoding server: D1 matches the first 4 layers of hash. =85; D2 matches the first 3 layers of hash. =68; D3 matches the first two layers of hash. =42; D4 and D5 have low matching scores. This is because D1 and D2 have... If the threshold of 60 is exceeded, it is determined to be a continued request. Further calculation of the decoding server's overall score: D1. =58.2、 =-8, overall score 53.95; D2 =55.8、 =-5, overall score 47.30. Select D1 for decoding to make full use of its KV buffer and reduce redundant calculations.
[0053] In one specific implementation, after the system has been running for a period of time, the particle swarm optimization (PSO) algorithm is used to dynamically adjust the weight parameters. Initialize 20 particles, with initial positions randomly distributed in the range [0,1] and satisfying normalization constraints. The inertia weight ω = 0.7, and the learning factor... Under test load, 100 requests were run to measure the average response time: the average response time for the initial parameters (0.5, 0.3, 0.2) was 2.3 seconds with a fitness of 0.435; the average response time for particle 5 parameters (0.45, 0.35, 0.20) was 2.1 seconds with a fitness of 0.476; and the average response time for particle 12 parameters (0.40, 0.40, 0.20) was 1.95 seconds with a fitness of 0.513. After 50 iterations, the global optimum converged to (0.42, 0.38, 0.20), and the average response time decreased to 1.85 seconds, representing a performance improvement of 19.6%.
[0054] In one specific implementation, the system comprises servers with varying performance levels: two high-performance A100 GPU servers and four medium-performance V100 GPU servers. First, performance benchmarking is performed, determining the baseline performance gains of the A100 servers. =80, V100 server =50. Introducing a performance normalization factor. This allows different servers to Comparability: Among them, reference performance The median is 60. For a batch of 10 concurrent requests, the allocation result after scoring is as follows: A100 server 1 is allocated 3 requests, A100 server 2 is allocated 3 requests, and V100 servers 1 to 4 are each allocated 1 request. This achieves load distribution proportional to server performance, with high-performance servers undertaking more load.
[0055] In one specific implementation, the system suddenly receives a burst of traffic of 100 concurrent requests under normal load. First, a batch processing mode is used to calculate the semantic vectors and hash values of all requests in parallel, reducing the processing time from 5 seconds serially to 0.8 seconds. Then, a hierarchical routing strategy is executed: in the first round, requests are quickly classified based on hash matching degree, with the 30 high-matching requests prioritized and routed to the decoding server; in the second round, the 70 low-matching requests are routed according to... The data is sorted and distributed sequentially to pre-filled servers; in the third round, a Softmax probability strategy is used to increase the dispersion of queue backlog. Simultaneously, the threshold is temporarily lowered upon detecting high concurrency. The utilization rate of the decoding server was increased from 60 to 45, alleviating the pressure on the pre-filled server. The final performance results were: average response time of 3.2 seconds, P95 latency of 5.8 seconds, and load imbalance of 0.18, representing a 35% reduction in response time compared to the round-robin strategy.
[0056] In one specific implementation, an extremely long text request containing 4096 tokens is processed. A multi-level subset partitioning (8 levels) is employed, generating 8 hash values from 16 to 4096 tokens. A server is then identified through progressive matching. The first 2048 tokens have been cached (the first 7 layers of hashes all match), and the matching score is as follows. Theoretically the largest Hit rate 49.8. Although the hit rate did not reach the threshold of 60%, considering the large amount of caching already in place, the decision was adjusted: the cache benefit was calculated as Benefit = 2048 / 4096 × 100 = 50. Since Benefit > 40, an additional +15 was awarded. =64.8 exceeds the threshold, and is ultimately routed to This saves 50% of the pre-filled calculations.
[0057] Example 4: This example provides a load balancing decision system for multimodal large-model high-concurrency inference, including: The request feature extraction module is used to receive inference requests from multimodal large models, perform lexical sequence analysis on the inference requests, extract semantic vectors, and calculate hardness factors. The hardness similarity prediction module is used to maintain historical request records for each server and calculate the initial optimization utilization item for each server based on the similarity between the hardness factor of the current request and the hardness factor of the server's historical requests. The compensation and optimal utilization module is used to calculate the expected residual compensation factor for each server based on the semantic similarity and time decay between the current request and the server's historical requests. The initial optimized utilization term is added to the expected residual compensation factor to obtain the optimal utilization term for each server. The scoring calculation module is used to calculate the hash matching score and real-time load score for each server. The weighted comprehensive decision module is used to sum the optimal utilization item, hash matching score and real-time load score of each server in a weighted manner to obtain a comprehensive decision score. The target routing module is used to select a target server based on the comprehensive decision score and route the current inference request to the selected target server.
[0058] It should be noted that the technical solution of the multimodal large model high-concurrency inference load balancing decision system is based on the same concept as the technical solution of the multimodal large model high-concurrency inference load balancing decision method described above. For details not described in detail in the technical solution of the multimodal large model high-concurrency inference load balancing decision system in this embodiment, please refer to the description of the technical solution of the multimodal large model high-concurrency inference load balancing decision method described above.
[0059] The above-mentioned unit modules can be embedded in the processor of the electronic device in hardware form or independent of it, or they can be stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of the above modules.
[0060] This embodiment also provides an electronic device, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a load balancing decision-making method for multimodal large-model high-concurrency inference. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0061] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method proposed in the above embodiments.
[0062] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory, random access memory, flash memory, hard disk, or optical disk, and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute the method of the embodiments of the present invention.
[0063] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the present invention.
Claims
1. A load balancing decision-making method for multimodal large-model high-concurrency inference, characterized in that, include: Receive inference requests from multimodal large models, perform lexical sequence analysis on the inference requests, extract semantic vectors, and calculate hardness factors; Maintain historical request records for each server, and calculate the initial optimization utilization for each server based on the similarity between the hardness factor of the current request and the hardness factor of the server's historical requests. Based on the semantic similarity and time decay between the current request and the server's historical requests, the expected residual compensation factor for each server is calculated. The initial optimization utilization term is added to the expected residual compensation factor to obtain the optimal utilization term for each server. Calculate the hash matching score and real-time load score for each server; The optimal utilization, hash matching score, and real-time load score of each server are weighted and summed to obtain a comprehensive decision score. The target server is selected based on the comprehensive decision score, and the current inference request is routed to the selected target server.
2. The load balancing decision-making method for multimodal large-model high-concurrency inference as described in claim 1, characterized in that, The step of performing lexical sequence analysis on the reasoning request, extracting semantic vectors, and calculating the hardness factor includes: The original text of the reasoning request is lexicalized to obtain a lexical sequence; Starting from a preset base number, the word sequence is divided into multiple hierarchical subsets according to an exponential growth method, and a hash value is calculated for each hierarchical subset to form the hash value set of the current request; The inference request is mapped into a semantic vector using a pre-trained sentence encoder; The hardness factor is calculated based on the number of input terms in the inference request, the expected output length, and the context complexity coefficient.
3. The load balancing decision-making method for multimodal large-model high-concurrency inference as described in claim 2, characterized in that, Maintaining historical request records for each server includes: Allocate a fixed amount of storage space for historical request records to each server. Each historical request record contains a set of hash values of processed requests, a semantic vector, a hardness factor, the request time, the actual processing time, and the original performance gains. Maintain a fixed-size recent performance gain time window for each server, collect the raw performance gains of all completed requests within the window, and form a recent performance gain set; In response to the storage space being full and the need to add new records, a retention score is calculated for each historical request record based on the timestamp and access frequency. The historical request records are then sorted in descending order according to the retention score. A preset proportion of historical request records are eliminated according to the sorting result, and the new records are then stored in the storage space.
4. The load balancing decision-making method for multimodal large-model high-concurrency inference as described in claim 3, characterized in that, The calculation of the expected residual compensation factor for each server includes: The adaptive state decay rate of the server is calculated based on the aforementioned recent performance gain set, and the time decay factor is calculated based on the adaptive state decay rate. For each historical request, the residual weight factor is calculated by comprehensively considering the cosine similarity and time decay factor between the semantic vectors of the current request and the historical request. The difference between the original performance gain of each historical request and the initial optimization utilization item corresponding to the server is obtained as the performance prediction residual; The weighted average of the residual weighting factors of all historical requests and the performance prediction residuals is used as the expected residual compensation factor.
5. The load balancing decision-making method for multimodal large-model high-concurrency inference as described in claim 4, characterized in that, The calculation of the hash matching score and real-time load score for each server includes: The hash values corresponding to each hierarchical subset in the hash value set of the current request are matched with the historical hash values cached by the server, and different weights are assigned according to the hierarchical level of the hierarchical subset corresponding to the successfully matched hash value. The hash matching score is calculated based on the sum of all matching weights and the theoretical maximum matching score. Obtain the request queue currently being processed by the server, and in response to each request in the request queue having a length less than a preset length threshold, assign a fixed penalty value; In response to each request in the request queue having a length greater than or equal to a preset length threshold, a multiplier penalty value is applied by rounding up the ratio of the request length to the base length. The real-time load score is obtained by summing the penalty scores of all requests.
6. The load balancing decision-making method for multimodal large-model high-concurrency inference as described in claim 5, characterized in that, The step of selecting a target server based on the comprehensive decision score includes: Based on the comprehensive decision score, a greedy strategy is used to select the server with the highest score as the target server. During system operation, a weighted loss function of load imbalance and average latency is calculated based on real-time performance indicators, and the summation weight of the comprehensive decision score is fine-tuned online using the gradient descent method.
7. The load balancing decision-making method for multimodal large-model high-concurrency inference as described in claim 6, characterized in that, The step of selecting a target server based on the comprehensive decision score and routing the current inference request to the selected target server includes: The server is divided into a pre-filled component set and a decoding component set, and a hit rate threshold is set. If the hash matching score of all servers is lower than the hit rate threshold, the inference request is determined to be a new request and routed to the one with the highest comprehensive decision score in the pre-filled component set. If the hash matching score of the existing server is not lower than the hit rate threshold, the inference request is determined to be a continuation request and routed to the one with the highest comprehensive decision score in the set of decoding components. In response to the existence of a server's hash matching score being approximately equal to the hit rate threshold, a decision is made taking into account the optimal utilization factors.
8. A load balancing decision system for multimodal large-model high-concurrency inference, employing the load balancing decision method for multimodal large-model high-concurrency inference as described in any one of claims 1 to 7, characterized in that, include: The request feature extraction module is used to receive inference requests from multimodal large models, perform word sequence analysis on the inference requests, extract semantic vectors, and calculate hardness factors. The hardness similarity prediction module is used to maintain historical request records for each server and calculate the initial optimization utilization item for each server based on the similarity between the hardness factor of the current request and the hardness factor of the server's historical requests. The compensation and optimal utilization module is used to calculate the expected residual compensation factor for each server based on the semantic similarity and time decay of the current request and the server's historical requests, and to add the initial optimized utilization term to the expected residual compensation factor to obtain the optimal utilization term for each server. The scoring calculation module is used to calculate the hash matching score and real-time load score for each server. The weighted comprehensive decision module is used to perform a weighted summation of the best utilization item, hash matching score and real-time load score of each server to obtain a comprehensive decision score; The target routing module is used to select a target server based on the comprehensive decision score and route the current inference request to the selected target server.
9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store computer-executable instructions, and when the processor executes the computer-executable instructions, it implements the steps of the load balancing decision method for multimodal large-model high-concurrency inference as described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: When the computer-executable instructions are executed by the processor, they implement the steps of the load balancing decision method for multimodal large-model high-concurrency inference as described in any one of claims 1 to 7.