Vector library model heterogeneous scheduling method and system for asymmetric hardware resources

CN122470384BActive Publication Date: 2026-09-01SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610952763.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-01
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

然而,由于异构硬件之间存在的显著的非对称差异,现有调度方法存在以下缺陷:其一,未有效感知硬件非对称特性,调度策略多为静态绑定或简单轮询,导致计算密集型任务与访存密集型任务错配;其二,未能将向量检索任务与大模型推理任务进行细粒度解耦,导致推理效率低下;其三,缺乏动态软硬件协同与预取缓冲机制,海量向量特征在异构存储层间的频繁搬运引发严重I/O阻塞,造成响应延时过高,无法满足海量数据实时检索的业务需求

Benefits of technology

本发明的面向非对称硬件资源的向量库模型异构调度方法,通过构建基于计算访存比的全局非对称特征矩阵,实现了对异构硬件资源的精细化感知与动态分类,并针对向量库场景中计算密集的嵌入生成任务与访存密集的向量检索任务进行细粒度解耦与异构映射,显著避免了任务与硬件特性错配导致的性能瓶颈;同时,针对大内存与小显存构成的极端非对称节点,创新性地引入三级存储架构与双缓冲异步预取机制,使计算与数据搬运完全重叠并结合零拷贝技术,极大消除了PCIe I/O阻塞与数据传输延迟;此外,通过实时监测资源饱和度并触发智能重定向与跨架构降维计算,有效化解了高并发场景下的单点资源过载问题。本发明能够全面提升异构集群在AI大模型推理与海量向量检索混合负载下的资源利用率、任务吞吐量与响应实时性,显著降低系统延时与能耗,为边缘计算及大规模RAG系统提供了高效、稳定且自适应的核心调度解决方案。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470384B_ABST
    Figure CN122470384B_ABST
Patent Text Reader

Abstract

The application discloses a vector library model heterogeneous scheduling method and system facing asymmetric hardware resources, and belongs to the technical field of hardware scheduling. The vector library model heterogeneous scheduling method facing asymmetric hardware resources extracts a computation memory ratio to construct a global asymmetric feature matrix by monitoring physical indexes of nodes in real time, divides the nodes into computation advantage type nodes and capacity advantage type nodes; decouples a vector processing request into reasoning subtasks and retrieval subtasks, and schedules the subtasks to corresponding type nodes based on the global asymmetric feature matrix; for extreme asymmetric nodes, adopts double-buffer asynchronous scheduling based on a three-level storage architecture to completely overlap computation operation and data transfer; monitors resource saturation in real time, and performs subtask cross-node migration and asynchronous compensation computation based on a redirection strategy. The application significantly improves resource utilization and response real-time performance of a heterogeneous cluster under mixed loads of AI reasoning and vector retrieval, reduces transmission delay, and expands model deployment capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of hardware scheduling technology, specifically relating to a heterogeneous scheduling method and system for vector library models oriented towards asymmetric hardware resources. Background Technology

[0002] The statements herein provide only background information in relation to this invention and do not necessarily constitute prior art.

[0003] With the widespread application of large language models and generative AI technologies, feature extraction using embedding models combined with vector databases for knowledge retrieval and enhancement (such as RAG systems) has become a core technical approach to improve the accuracy and real-time performance of artificial intelligence systems. When processing massive amounts of high-dimensional vector data, the system not only needs to perform complex model inference to generate vectors, but also needs to perform high-frequency approximate nearest neighbor (ANN) retrieval in the massive vector set.

[0004] Currently, vector library models are widely deployed on heterogeneous computing platforms composed of CPUs, GPUs, NPUs, etc. However, due to the significant asymmetric differences between heterogeneous hardware, existing scheduling methods have the following drawbacks: First, they fail to effectively perceive the asymmetric characteristics of hardware, and scheduling strategies are mostly static binding or simple round-robin, leading to a mismatch between computationally intensive tasks and memory-intensive tasks; second, they fail to decouple vector retrieval tasks from large model inference tasks in a fine-grained manner, resulting in low inference efficiency; third, they lack dynamic hardware-software collaboration and prefetch buffering mechanisms, and the frequent movement of massive vector features between heterogeneous storage layers causes severe I / O blocking, resulting in excessively high response latency and failing to meet the business needs of real-time retrieval of massive data. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a heterogeneous scheduling method and system for vector library models oriented towards asymmetric hardware resources. After the upper-layer application initiates a request, it can realize the accurate matching and collaborative execution of tasks among hardware with different characteristics, and effectively adapt to the complex scenario of collaborative processing of massive vector retrieval by multiple types of heterogeneous hardware.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: In a first aspect, the technical solution of the present invention provides a heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources, including: Obtain multiple physical metrics and instantaneous load of each heterogeneous node, calculate the computation-to-memory access ratio of each node and construct a global asymmetric feature matrix, and divide each heterogeneous node into computation-dominant nodes and capacity-dominant nodes based on the global asymmetric feature matrix. Based on the received vector processing request, it is parsed and decomposed into inference subtasks and retrieval subtasks. Based on the global asymmetric feature matrix, the optimal execution target node of the subtask is determined. The inference subtask is scheduled to a computationally advantageous node, and the retrieval subtask is scheduled to a capacity-advantaged node. When the optimal execution target node is in an asymmetric architecture with a large system memory and limited accelerator memory, double-buffered asynchronous scheduling based on a three-level storage architecture is performed inside the node, so that the computation operation and data transfer operation completely overlap on the time axis. The system monitors the resource saturation of the optimal execution target node in real time. When the resource saturation exceeds the saturation threshold and the queue is congested, a redirection strategy is triggered to migrate some of the queued subtasks across nodes to another type of node and perform asynchronous compensation calculations.

[0007] In at least one embodiment, multiple physical metrics of each heterogeneous node include theoretical peak computing power, tiered storage capacity, actual data exchange rate between storage tiers, bus communication bandwidth, and communication latency.

[0008] In at least one embodiment, calculating the memory access ratio of each node and constructing a global asymmetric feature matrix specifically includes: calculating the ratio of the theoretical peak computing power of the heterogeneous node to the actual data exchange rate between storage tiers to obtain the memory access ratio of the heterogeneous node; calculating the product of bus communication bandwidth and communication latency to obtain the bandwidth-delay product of the heterogeneous node; constructing asymmetric feature vectors of the heterogeneous node based on the memory access ratio, tier storage capacity, bandwidth-delay product, and instantaneous load of the heterogeneous node; and constructing a global asymmetric feature matrix based on the asymmetric feature vectors of each heterogeneous node.

[0009] In at least one embodiment, heterogeneous nodes are divided into computationally dominant nodes and capacity-dominant nodes based on a global asymmetric feature matrix. Specifically, this includes: based on a first set threshold and a second set threshold, if the computation-to-memory access ratio of a heterogeneous node is greater than the first set threshold, then the heterogeneous node is marked as a computationally dominant node; if the computation-to-memory access ratio of a heterogeneous node does not exceed the first set threshold, then it is further determined whether the hierarchical storage capacity of the heterogeneous node is greater than the second set threshold. If it does, then the heterogeneous node is marked as a capacity-dominant node. Otherwise, based on the bandwidth-delay product and instantaneous load, the heterogeneous node is used as a candidate node for communication-sensitive tasks, or as a backup node for capacity-dominant nodes to participate in subsequent redirection.

[0010] In at least one embodiment, the inference subtask is specifically an embedded feature generation subtask that mainly involves intensive matrix multiplication; the retrieval subtask is specifically a vector graph index retrieval subtask that mainly involves non-contiguous memory access and large-scale candidate vector traversal.

[0011] In at least one embodiment, determining the optimal execution target node for a subtask based on a global asymmetric feature matrix specifically includes: calculating the task requirement vectors for the inference subtask and the retrieval subtask; calculating the cosine similarity between the requirement vector of each subtask and the asymmetric feature vector of each node in the global asymmetric feature matrix; determining the optimal execution target node for the subtask based on the cosine similarity; and scheduling the inference subtask to a computationally dominant node and the retrieval subtask to a capacity-dominant node based on the optimal execution target node.

[0012] In at least one embodiment, the three-level storage architecture is as follows: a large-capacity system memory is designated as the main index and cold parameter layer, a dedicated locked page memory area is opened in the system memory as a shadow buffer pool, and the limited GPU video memory is used as a hot computing cache layer. The dual-buffered asynchronous scheduling works as follows: When the accelerator computing core processes the current batch of model inference in the current buffer of the hot computing cache layer, the scheduling module, in coordination with the CPU, uses the direct memory access controller to asynchronously prefetch the feature blocks or model weights required for the next batch from system memory to the shadow buffer pool and push them into the standby buffer of the hot computing cache layer. After the current batch of computation is completed, the current buffer and the standby buffer switch roles so that the next batch of computation can directly read the prefetched data, ensuring that computation and I / O operations completely overlap on the timeline.

[0013] In at least one embodiment, during asynchronous prefetching, zero-copy technology is used to directly map the address of the page-locked memory to the address space accessible by the accelerator computing core, so as to avoid secondary copying of data between kernel mode and user mode. At the same time, a multi-level semaphore synchronization mechanism is introduced. By monitoring the bus conflict rate in real time, if the PCIe bandwidth contention exceeds the preset warning value, the prefetch frequency is dynamically reduced, and the loading of key weights is prioritized to prevent instruction pipeline bubbling caused by excessive prefetching.

[0014] In at least one embodiment, the redirection strategy specifically includes: intercepting non-real-time tasks and stripping some of the non-real-time retrieval subtasks in the queue, migrating them across nodes to capacity-advantaged nodes; if the capacity-advantaged nodes are unavailable, selecting a backup node that meets the bandwidth-delay product and instantaneous load constraints to take over the communication-sensitive subtasks; after migrating across nodes, calling the SIMD instruction set based on vector quantization or dimensionality reduction operators for asynchronous compensation calculation.

[0015] Secondly, the technical solution of the present invention also provides a heterogeneous scheduling system for vector library models oriented towards asymmetric hardware resources, comprising: The hardware status awareness module is configured to: acquire multiple physical indicators and instantaneous load of each heterogeneous node, calculate the computation-to-memory access ratio of each node and construct a global asymmetric feature matrix, and divide each heterogeneous node into computation-dominant nodes and capacity-dominant nodes based on the global asymmetric feature matrix. The task scheduling module is configured to: based on the received vector processing request, parse and decompose it into inference subtasks and retrieval subtasks; based on the global asymmetric feature matrix, determine the optimal execution target node of the subtasks; schedule the inference subtasks to computationally advantageous nodes; and schedule the retrieval subtasks to capacity-advantaged nodes. The pipeline control module is configured to: when the optimal execution target node is in an asymmetric architecture with a large system memory and limited accelerator memory, execute double-buffered asynchronous scheduling based on a three-level storage architecture inside the node, so that the computation operation and the data transfer operation completely overlap on the time axis; The dynamic monitoring and redirection module is configured to: monitor the resource saturation of the optimal execution target node in real time; when the resource saturation exceeds the saturation threshold and the queue is congested, trigger the redirection strategy to migrate some of the queued subtasks across nodes to another type of node and perform asynchronous compensation calculations.

[0016] The beneficial effects of the above-described technical solution of the present invention are as follows: This invention presents a heterogeneous scheduling method for vector library models targeting asymmetric hardware resources. By constructing a global asymmetric feature matrix based on computation-to-memory access ratio, it achieves fine-grained perception and dynamic classification of heterogeneous hardware resources. Furthermore, it performs fine-grained decoupling and heterogeneous mapping between computationally intensive embedding generation tasks and memory-intensive vector retrieval tasks in vector library scenarios, significantly avoiding performance bottlenecks caused by task-hardware mismatch. Simultaneously, for extremely asymmetric nodes composed of large memory and small video memory, it innovatively introduces a three-level storage architecture and a double-buffered asynchronous prefetch mechanism, allowing computation and data transfer to completely overlap and, combined with zero-copy technology, greatly eliminating PCIe I / O blocking and data transmission latency. In addition, by real-time monitoring of resource saturation and triggering intelligent redirection and cross-architecture dimensionality reduction computation, it effectively resolves the problem of single-point resource overload in high-concurrency scenarios. This invention comprehensively improves the resource utilization, task throughput, and real-time response of heterogeneous clusters under mixed loads of AI large-scale model inference and massive vector retrieval, significantly reducing system latency and energy consumption, and providing an efficient, stable, and adaptive core scheduling solution for edge computing and large-scale RAG systems. Attached Figure Description

[0017] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0018] Figure 1 This is a schematic diagram of the heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources disclosed in Embodiment 1 of the present invention; Figure 2 This is an architecture diagram of the heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources disclosed in Embodiment 1 of the present invention. Detailed Implementation

[0019] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0020] Example 1 In typical vector library deployment and massive data retrieval scenarios, the underlying system is usually composed of computing nodes with differentiated characteristics. For example, an edge computing node may be configured with a large capacity of 128GB of system memory, but only equipped with a graphics processing unit with 8GB of video memory. This extremely asymmetric physical architecture requires the system to have fine-grained resource awareness and scheduling capabilities when performing large-parameter embedded model inference and high-frequency vector retrieval.

[0021] As described in the background section, there are significant asymmetric differences between heterogeneous hardware. For example, GPUs have extremely high floating-point computing power but limited video memory (VRAM) capacity and PCIe communication bandwidth; while CPUs have low computational parallelism but have massive amounts of system memory (RAM); NPUs are extremely energy efficient in specific tensor calculations but perform poorly in general logic branch processing. These asymmetric differences lead to the following significant shortcomings in existing vector library model scheduling methods when facing typical vector library deployment and massive data retrieval scenarios: First, they fail to effectively perceive the asymmetric characteristics of hardware, and scheduling strategies are mostly static binding or based on simple round-robin and load balancing algorithms, resulting in a mismatch between computationally intensive and memory-intensive tasks; Second, the scheduling strategies are too simplistic and fail to decouple vector retrieval tasks from large model inference tasks in a fine-grained manner across heterogeneous computing power, resulting in low model inference efficiency; Third, in high-concurrency retrieval scenarios, due to the lack of dynamic hardware-software collaboration and prefetching buffer mechanisms, the frequent movement of massive vector features between different hardware storage layers causes severe I / O blocking, resulting in extremely low system resource utilization, retrieval lag, and excessively high response latency, failing to meet the business needs of real-time retrieval of massive data.

[0022] To overcome the shortcomings of the prior art, in a typical embodiment of the present invention, such as... Figures 1 to 2As shown, this embodiment discloses a heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources. It is based on a complete heterogeneous scheduling architecture that includes hardware state awareness, task decomposition and mapping, hierarchical prefetching scheduling, and dynamic queue management. Through this architecture, the system can achieve precise matching and collaborative execution of tasks among hardware with different characteristics after a request is initiated by an upper-layer application. This heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources specifically includes the following steps: S1. Obtain multiple physical indicators and instantaneous load of each heterogeneous node, calculate the computation-to-memory access ratio of each node and construct a global asymmetric feature matrix, and divide each heterogeneous node into computation-dominant nodes and capacity-dominant nodes based on the global asymmetric feature matrix. S2. Based on the received vector processing request, parse and decompose it into inference subtasks and retrieval subtasks. Based on the global asymmetric feature matrix, determine the optimal execution target node for the subtasks, schedule the inference subtasks to computationally advantageous nodes, and schedule the retrieval subtasks to capacity advantageous nodes. S3. When the optimal execution target node is in an asymmetric architecture with a large system memory and limited accelerator memory, double-buffered asynchronous scheduling based on a three-level storage architecture is executed inside the node, so that the computation operation and data transfer operation completely overlap on the time axis. S4. Monitor the resource saturation of the optimal execution target node in real time. When the resource saturation exceeds the saturation threshold and the queue is congested, trigger the redirection strategy to migrate some of the queued subtasks across nodes to another type of node and perform asynchronous compensation calculations.

[0023] The following section provides a detailed explanation of the heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources, using specific implementation methods.

[0024] S1. Obtain multiple physical indicators and instantaneous load of each heterogeneous node, calculate the computation-to-memory access ratio of each node and construct a global asymmetric feature matrix. Based on the global asymmetric feature matrix, divide each heterogeneous node into computation-dominant nodes and capacity-dominant nodes.

[0025] In this step, we will focus on heterogeneous clusters. heterogeneous nodes in First, the heterogeneous node is obtained through the underlying driver interface. Theoretical peak computing power Hierarchical storage capacity Actual data exchange rate between storage tiers Bus communication bandwidth and communication latency Physical indicators, and instantaneous load. .in, This represents the total number of heterogeneous nodes. Indicates the node number.

[0026] Secondly, based on heterogeneous nodes Theoretical peak computing power The actual data exchange rate between storage tiers Computing heterogeneous nodes The memory access ratio is calculated as follows:

[0027] In the formula, heterogeneous nodes The memory access ratio is calculated to characterize heterogeneous nodes. The theoretical computing power available at a unit data exchange rate; heterogeneous nodes The measured effective data exchange rate between storage layers such as accelerator memory, paged memory, and system memory.

[0028] Based on bus communication bandwidth With communication delay Computing heterogeneous nodes The bandwidth-delay product is specifically expressed as:

[0029] In the formula, heterogeneous nodes The bandwidth delay product.

[0030] Based on heterogeneous nodes Calculation of memory access ratio Hierarchical storage capacity Bandwidth delay product and instantaneous load Building heterogeneous nodes The asymmetric eigenvectors of are specifically represented as:

[0031] In the formula, heterogeneous nodes Asymmetric eigenvectors; This is the normalized memory access ratio. This represents the normalized hierarchical storage capacity. This is the normalized bandwidth-delay product; This represents the normalized instantaneous load.

[0032] A global asymmetric feature matrix is ​​constructed based on the asymmetric feature vectors of each heterogeneous node, specifically represented as follows:

[0033] In the formula, It is a global asymmetric feature matrix, where each row corresponds to an asymmetric feature vector of a heterogeneous node.

[0034] Then, based on the global asymmetric feature matrix The underlying heterogeneous nodes are dynamically divided into computing-dominant nodes and capacity-dominant nodes, thereby providing an accurate physical view and quantitative basis for subsequent task scheduling.

[0035] Specifically, based on the first set threshold Second set threshold If heterogeneous nodes Calculation of memory access ratio Greater than the first set threshold Then heterogeneous nodes Marked as a computationally dominant node , recorded as If heterogeneous nodes Calculation of memory access ratio Not exceeding the first set threshold Then continue to determine heterogeneous nodes. Hierarchical storage capacity Is it greater than the second set threshold? If the number exceeds the limit, the heterogeneous nodes will be marked as capacity-dominant nodes. , recorded as If heterogeneous nodes Hierarchical storage capacity Not exceeding the second set threshold Then, based on the bandwidth-delay product With instantaneous load This heterogeneous node It can participate in subsequent redirection as a candidate node for communication-sensitive tasks or as a backup node for capacity-superior nodes.

[0036] As a further implementation, a first set threshold is used. Second set threshold The system automatically calibrates based on the micro-benchmark tests during the initialization phase, and follows the global asymmetric eigenma matrix. The periodic updates are used for dynamic weighted adjustment. The dynamic weighted adjustment employs a weighted decay factor. Smooth transient load The impact on scheduling decisions is specifically expressed as follows:

[0037] In the formula, heterogeneous nodes In the current sampling period Smooth eigenvectors; This is the currently measured feature vector; This is the smoothed feature vector from the previous sampling period. The scheduler is based on all nodes. Update the first set threshold Second set threshold To reduce instantaneous load The impact of jitter on node classification results, and the ability of the scheduling scheme to reflect the actual available computing power under dynamic operating conditions. This dynamic calibration mechanism enables the first set threshold to be used. Second set threshold It can adaptively update as hardware ages or the system expands.

[0038] S2. Based on the received vector processing request, parse and decompose it into inference subtasks and retrieval subtasks. Based on the global asymmetric feature matrix, determine the optimal execution target node for the subtasks, schedule the inference subtasks to computationally advantageous nodes, and schedule the retrieval subtasks to capacity advantageous nodes.

[0039] In this step, a complex vector library task flow that is continuously arriving and initiated by the upper-layer application is addressed. Based on the received vector processing request First, the computational graph of the task is decomposed. Vector processing request. This includes requests to vectorize input text, images, or multimodal data, as well as requests to perform approximate nearest neighbor retrieval based on generated vectors.

[0040] During the decomposition process, the scheduling center obtains operator descriptions from the embedded model runtime and the vector indexing execution engine to construct vector processing requests. Corresponding calculation graph ,in, Represents a set of operators. This represents the data dependency edges between operators. The scheduling center processes the computation graph based on operator type, input / output tensor shape, and dependencies. Topological sorting is performed, and the linear layer, attention mechanism operator, normalization operator, and activation operator are grouped into the Transformer-based embedding model inference stage. Vector normalization, inverted list scanning, graph index neighbor expansion, candidate vector distance calculation, and Top-K merging operator are grouped into the vector index retrieval stage. Vector processing requests are grouped using the embedded vector output node as the split boundary. Decoupling into reasoning subtasks and retrieval subtasks Among them, the reasoning subtask The first subtask primarily involves intensive matrix multiplication for embedding feature generation, used to call a Transformer-based embedding model to generate query vectors; the second subtask is the retrieval subtask. It is a vector graph index retrieval subtask that mainly involves non-contiguous memory access and large-scale candidate vector traversal. It is used to perform candidate recall, distance calculation and result merging in HNSW, IVF, PQ or their combined indexes.

[0041] Then, the inference subtask is computed. and retrieval subtasks The task requirement vector. For any subtask Its task requirement vector Represented as:

[0042] In the formula, This represents the normalized theoretical computing power requirement. This represents the normalized peak storage usage. To meet the needs of normalized data transfer; This refers to the normalized task priority or real-time constraints.

[0043] For the reasoning subtask Theoretical computing power requirements The peak storage usage is obtained by summing the floating-point operations of matrix multiplication, attention mechanism operators, and feedforward network operators in the embedded model computation graph. The data transfer requirement is estimated from model weights, activation values, and intermediate tensor peak occupancy. The data volume is estimated based on weighted blocks and activation value dumps. For the retrieval subtask... Theoretical computing power requirements The peak storage usage is estimated by multiplying the number of distance calculations for candidate vectors by the vector dimension. Data movement requirements are estimated based on index sharding, candidate set, and result cache usage. It is estimated from the number of index page reads and the number of candidate vector accesses.

[0044] Calculate the task requirement vector With global asymmetric characteristic matrix Asymmetric eigenvectors of each node The cosine similarity is specifically expressed as:

[0045] The optimal execution target node for a subtask is determined based on cosine similarity, specifically as follows:

[0046] In the formula, This represents the optimal execution target node for the subtask.

[0047] Determine the optimal execution target node When the subtask is a reasoning subtask Then priority is given to computing dominant nodes. The highest cosine similarity and instantaneous load were selected. Nodes that have not exceeded the load threshold; if the subtask is a retrieval subtask. Prioritize nodes with capacity advantages. The node with the highest cosine similarity and available storage capacity that meets the index sharding requirements is selected. Therefore, the theoretical computing power requirements, peak storage usage, and data transfer requirements of the subtask are matched with the underlying asymmetric feature vectors based on similarity, and the optimal execution target node for the subtask is determined based on this similarity. This achieves the initial binding between task logic and hardware physical characteristics.

[0048] As a further implementation, at a finer-grained execution level, the scheduling center employs a topology sorting algorithm to perform in-depth analysis of the model's operator dependencies. For large model inference tasks, the system further identifies intensive and communication-sensitive operators within the task. For edge nodes with extremely limited GPU memory, the system implements an operator offloading strategy: temporarily dumping the activation values ​​of intermediate layers with extremely high GPU memory usage but low computational intensity to the system memory layer. Through this fine-grained operator scheduling, the system can overcome the physical GPU memory limitations of a single hardware device, enabling the deployment of large-parameter models exceeding the physical GPU memory capacity on low-cost heterogeneous devices, thereby greatly expanding the applicability of the device while ensuring inference accuracy.

[0049] S3. When the optimal execution target node is in an asymmetric architecture with a large system memory and limited accelerator memory, double-buffered asynchronous scheduling based on a three-level storage architecture is executed inside the node, so that the computation operation and the data transfer operation completely overlap on the time axis.

[0050] To address the severe I / O blocking problem within extremely asymmetric nodes, this step first determines whether the optimal execution target node is in an asymmetric architecture with large system memory and limited accelerator memory. If so, a double-buffered asynchronous scheduling based on a three-level storage architecture is executed within the optimal execution target node to achieve complete overlap between computation operations and data transfer operations on the time axis.

[0051] Specifically, the three-tier storage architecture refers to allocating a large amount of system memory as the primary index and cold parameter layer, and creating a dedicated pinned memory area within it as a shadow buffer pool. Simultaneously, the limited GPU memory is used as a hot computing cache layer. Based on this three-level storage architecture, a dual-buffer switching mechanism is adopted. When the accelerator computing core is processing the current batch of model inference in the current buffer of the hot computing cache layer, the scheduling module, in coordination with the CPU multi-threading, asynchronously prefetches the feature blocks or model weights required for the next batch from system memory to the shadow buffer pool through the direct memory access (DMA) controller. The data is then pushed into the spare buffer of the hot computing cache layer. After the current batch of computation is completed, the current buffer and the spare buffer switch roles, allowing the next batch of computation to directly read the prefetched data. This ensures that computation and I / O operations are completely overlapped on the timeline, reducing transmission latency. This deeply overlapped scheduling method of communication transmission and computing power execution can mask the high transmission latency between heterogeneous hardware. In extremely asymmetric environments, the system uses a parallel execution of a dual-buffered asynchronous scheduling mechanism based on a three-level storage architecture. While the core operators are processing the current batch of computation, the prefetching operation of the next batch of weights is executed synchronously, thereby decoupling computation and transmission.

[0052] As a further implementation, in order to ensure the shadow buffer pool To ensure the atomicity of data exchange between the DMA controller and the hot computing cache layer, when the DMA controller performs asynchronous prefetch operations, the system utilizes zero-copy technology to directly map the addresses in the locked page memory to the address space accessible by the accelerator computing core, thus avoiding the secondary copying of data between kernel mode and user mode in traditional scheduling. Simultaneously, a multi-level semaphore synchronization mechanism is introduced. By monitoring the bus conflict rate in real time, if PCIe bandwidth contention exceeds a preset warning value, the scheduler will dynamically reduce the prefetch frequency and prioritize the loading of critical weights, thereby preventing instruction pipeline bubbling caused by excessive prefetching and ensuring the ultimate smoothness of pipeline scheduling in heterogeneous environments.

[0053] S4. Monitor the resource saturation of the optimal execution target node in real time. When the resource saturation exceeds the saturation threshold and the queue is congested, trigger the redirection strategy to migrate some of the queued subtasks across nodes to another type of node and perform asynchronous compensation calculations.

[0054] To ensure system throughput stability under high-concurrency retrieval scenarios, this step introduces a multi-level task queuing and dynamic conflict resolution mechanism. The scheduling center maintains multi-level feedback queues for various tasks. Real-time online retrieval requests enter a high-priority queue, while offline batch data entry tasks, index rebuilding tasks, and background pre-computation tasks enter a low-priority queue, ensuring that the execution priority of real-time online retrieval requests is higher than that of offline batch data entry tasks. During this process, if a computing-advantaged node is detected... The memory bandwidth utilization has reached the saturation threshold. Furthermore, if high-priority queues become congested, the scheduler will automatically trigger a dimensionality reduction or redirection strategy.

[0055] Specifically, real-time monitoring of the optimal execution target node. resource saturation With bandwidth usage .

[0056] resource saturation Based on the video memory utilization rate, hot computing cache layer utilization rate, and instantaneous load, the specific calculation is expressed as follows:

[0057] In the formula, This refers to the video memory usage rate; This refers to the occupancy rate of the hot computing cache layer. The normalized instantaneous load; , and These are the weighting coefficients, and .

[0058] Queue congestion status is determined by the length of the highest priority queue. Queue waiting time and bandwidth usage The detection, specifically, is as follows:

[0059] In the formula, This is the queue length threshold; This is the waiting time threshold; This is the bandwidth usage threshold; The OR operation indicates that the queue is in a congested state if any one of the criteria is met.

[0060] When the optimal execution target node resource saturation greater than the saturation threshold Furthermore, when congestion is detected in a high-priority queue, a redirection strategy is triggered. The system prioritizes intercepting non-real-time tasks. Furthermore, some non-real-time retrieval subtasks in the queue are stripped out and migrated across nodes to nodes with capacity advantages. If a capacity-advantage node is unavailable, a backup node that meets the constraints of bandwidth-delay product and instantaneous load will be selected to take over the communication-sensitive subtask.

[0061] After cross-node migration, the system invokes the SIMD instruction set to perform asynchronous compensation computation based on vector quantization or dimensionality reduction operators to compensate for the accuracy of the approximate retrieval results generated by the migration subtasks. Finally, it aggregates the execution results of all subtasks. At the same time, the system will record the instantaneous load of each node. Resource saturation and bandwidth usage Write back to the global asymmetric eigenmatrix The load status field of the corresponding node provides real-time basis for the next round of scheduling decisions. Through deep collaboration between hardware and software and dynamic load balancing, the system bottleneck caused by the exhaustion of single-point physical resources is effectively resolved, ensuring the efficient and stable operation of the vector library in an asymmetric environment.

[0062] Example 2 In a typical embodiment of the present invention, this embodiment discloses a heterogeneous scheduling system for vector library models oriented towards asymmetric hardware resources, comprising: The hardware status awareness module is configured to: acquire multiple physical indicators and instantaneous load of each heterogeneous node, calculate the computation-to-memory access ratio of each node and construct a global asymmetric feature matrix, and divide each heterogeneous node into computation-dominant nodes and capacity-dominant nodes based on the global asymmetric feature matrix. The task scheduling module is configured to: based on the received vector processing request, parse and decompose it into inference subtasks and retrieval subtasks; based on the global asymmetric feature matrix, determine the optimal execution target node of the subtasks; schedule the inference subtasks to computationally advantageous nodes; and schedule the retrieval subtasks to capacity-advantaged nodes. The pipeline control module is configured to: when the optimal execution target node is in an asymmetric architecture with a large system memory and limited accelerator memory, execute double-buffered asynchronous scheduling based on a three-level storage architecture inside the node, so that the computation operation and the data transfer operation completely overlap on the time axis; The dynamic monitoring and redirection module is configured to: monitor the resource saturation of the optimal execution target node in real time; when the resource saturation exceeds the saturation threshold and the queue is congested, trigger the redirection strategy to migrate some of the queued subtasks across nodes to another type of node and perform asynchronous compensation calculations.

[0063] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources, characterized in that, include: Multiple physical metrics and instantaneous loads of each heterogeneous node are obtained. The computation-to-memory access ratio and bandwidth-to-latency product of each node are calculated. Based on the computation-to-memory access ratio, tiered storage capacity, bandwidth-to-latency product, and instantaneous load of the heterogeneous nodes, an asymmetric feature vector of the heterogeneous nodes is constructed. Based on the asymmetric feature vectors of each heterogeneous node Construct a global asymmetric feature matrix; Based on the global asymmetric feature matrix, each heterogeneous node is divided into computationally dominant nodes and capacity-dominant nodes. Specifically, based on a first set threshold and a second set threshold, if the computation-to-memory access ratio of a heterogeneous node is greater than the first set threshold, then the heterogeneous node is marked as a computationally dominant node. If the computation-to-memory ratio of the heterogeneous node does not exceed the first set threshold, then it is further determined whether the hierarchical storage capacity of the heterogeneous node is greater than the second set threshold. If it is, the heterogeneous node is marked as a capacity-dominant node. Otherwise, based on the bandwidth-delay product and instantaneous load, the heterogeneous node is used as a candidate node for communication-sensitive tasks, or as a backup node for the capacity-dominant node to participate in subsequent redirection. Among them, the first set threshold Second set threshold The system automatically calibrates based on the micro-benchmark tests during the initialization phase, and dynamically adjusts the weights as the global asymmetric eigenvalue matrix is ​​periodically updated. Specifically: In the formula, heterogeneous nodes In the current sampling period Smooth eigenvectors; This is the currently measured feature vector; This is the smoothed feature vector of the previous sampling period; The weighted decay factor; the scheduler is based on all nodes. Update the first set threshold Second set threshold ; Based on the received vector processing request, it is parsed and decomposed into inference subtasks and retrieval subtasks. Based on the global asymmetric feature matrix, the optimal execution target node of the subtask is determined. The inference subtask is scheduled to a computationally advantageous node, and the retrieval subtask is scheduled to a capacity-advantaged node. When the optimal execution target node is in an asymmetric architecture with a large amount of system memory and limited accelerator memory, a double-buffered asynchronous scheduling based on a three-level storage architecture is executed inside the node, so that the computation operation and the data transfer operation completely overlap on the time axis. Specifically, the three-level storage architecture is as follows: the large amount of system memory is designated as the main index and cold parameter layer, a dedicated page-locked memory area is opened in the system memory as a shadow buffer pool, and the limited GPU memory is used as a hot computing cache layer. The dual-buffered asynchronous scheduling works as follows: When the accelerator computing core processes the current batch of model inference in the current buffer of the hot computing cache layer, the scheduling module, in coordination with the CPU, uses the direct memory access controller to asynchronously prefetch the feature blocks or model weights required for the next batch from system memory to the shadow buffer pool and push them into the standby buffer of the hot computing cache layer; after the current batch of computation is completed, the current buffer and the standby buffer switch roles so that the next batch of computation can directly read the prefetched data, ensuring that computation and I / O operations completely overlap on the timeline; The system monitors the resource saturation of the optimal execution target node in real time. When the resource saturation exceeds the saturation threshold and the queue is congested, a redirection strategy is triggered to migrate some of the queued subtasks across nodes to another type of node. After the cross-node migration, the system calls the SIMD instruction set to perform asynchronous compensation calculation based on vector quantization or dimensionality reduction operators to compensate for the accuracy of the approximate retrieval results generated by the migrated subtasks.

2. The heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources as described in claim 1, characterized in that, Several physical metrics for each heterogeneous node include theoretical peak computing power, tiered storage capacity, actual data exchange rate between storage tiers, bus communication bandwidth, and communication latency.

3. The heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources as described in claim 2, characterized in that, The computational access ratio of each node is calculated and a global asymmetric feature matrix is ​​constructed. Specifically, this includes: calculating the ratio of the theoretical peak computing power of heterogeneous nodes to the actual data exchange rate between storage tiers to obtain the computational access ratio of heterogeneous nodes; calculating the product of bus communication bandwidth and communication latency to obtain the bandwidth-latency product of heterogeneous nodes; constructing asymmetric feature vectors of heterogeneous nodes based on their computational access ratio, tiered storage capacity, bandwidth-latency product, and instantaneous load; and constructing a global asymmetric feature matrix based on the asymmetric feature vectors of each heterogeneous node.

4. The heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources as described in claim 1, characterized in that, The reasoning subtask is specifically an embedded feature generation subtask that mainly involves intensive matrix multiplication; the retrieval subtask is specifically a vector graph index retrieval subtask that mainly involves non-contiguous memory access and large-scale candidate vector traversal.

5. The heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources as described in claim 4, characterized in that, Based on the global asymmetric feature matrix, the optimal execution target node for each subtask is determined. Specifically, this includes: calculating the task requirement vectors for the inference subtask and the retrieval subtask; calculating the cosine similarity between the requirement vector of each subtask and the asymmetric feature vector of each node in the global asymmetric feature matrix; determining the optimal execution target node for each subtask based on the cosine similarity; and scheduling the inference subtask to a computationally advantageous node and the retrieval subtask to a capacity-advantaged node based on the optimal execution target node.

6. The heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources as described in claim 1, characterized in that, During asynchronous prefetching, zero-copy technology is used to directly map the address of the paged memory to the address space accessible by the accelerator computing core, so as to avoid secondary copying of data between kernel mode and user mode. At the same time, a multi-level semaphore synchronization mechanism is introduced. By monitoring the bus conflict rate in real time, if the PCIe bandwidth contention exceeds the preset warning value, the prefetch frequency is dynamically reduced, and the loading of key weights is prioritized to prevent instruction pipeline cavitation caused by excessive prefetching.

7. The heterogeneous scheduling method for vector library models oriented towards asymmetric hardware resources as described in claim 1, characterized in that, The redirection strategy specifically includes: intercepting non-real-time tasks, stripping some non-real-time retrieval subtasks in the queue, migrating them across nodes to capacity-advantaged nodes, and if capacity-advantaged nodes are unavailable, selecting backup nodes that meet the constraints of bandwidth-delay product and instantaneous load to take over communication-sensitive subtasks.

8. A heterogeneous scheduling system for vector library models oriented towards asymmetric hardware resources, characterized in that, include: The hardware status awareness module is configured to: acquire multiple physical metrics and instantaneous load of each heterogeneous node; calculate the compute-to-memory access ratio and bandwidth-to-latency product of each node; and construct an asymmetric feature vector of the heterogeneous nodes based on the compute-to-memory access ratio, tiered storage capacity, bandwidth-to-latency product, and instantaneous load. Based on the asymmetric feature vectors of each heterogeneous node Construct a global asymmetric feature matrix; Based on the global asymmetric feature matrix, each heterogeneous node is divided into computationally dominant nodes and capacity-dominant nodes. Specifically, based on a first set threshold and a second set threshold, if the computation-to-memory access ratio of a heterogeneous node is greater than the first set threshold, then the heterogeneous node is marked as a computationally dominant node. If the computation-to-memory ratio of the heterogeneous node does not exceed the first set threshold, then it is further determined whether the hierarchical storage capacity of the heterogeneous node is greater than the second set threshold. If it is, the heterogeneous node is marked as a capacity-dominant node. Otherwise, based on the bandwidth-delay product and instantaneous load, the heterogeneous node is used as a candidate node for communication-sensitive tasks, or as a backup node for the capacity-dominant node to participate in subsequent redirection. Among them, the first set threshold Second set threshold The system automatically calibrates based on the micro-benchmark tests during the initialization phase, and dynamically adjusts the weights as the global asymmetric eigenvalue matrix is ​​periodically updated. Specifically: In the formula, heterogeneous nodes In the current sampling period Smooth eigenvectors; This is the currently measured feature vector; This is the smoothed feature vector of the previous sampling period; The weighted decay factor; the scheduler is based on all nodes. Update the first set threshold Second set threshold ; The task scheduling module is configured to: based on the received vector processing request, parse and decompose it into inference subtasks and retrieval subtasks; based on the global asymmetric feature matrix, determine the optimal execution target node of the subtasks; schedule the inference subtasks to computationally advantageous nodes and the retrieval subtasks to capacity advantageous nodes. The pipeline control module is configured to: when the optimal execution target node is in an asymmetric architecture with a large amount of system memory and limited accelerator memory, execute double-buffered asynchronous scheduling based on a three-level storage architecture within the node, so that the computation operation and data transfer operation completely overlap on the time axis; the three-level storage architecture is specifically: the large amount of system memory is designated as the main index and cold parameter layer, a dedicated page-locked memory area is opened in the system memory as a shadow buffer pool, and the limited GPU memory is used as a hot computing cache layer; The dual-buffered asynchronous scheduling works as follows: When the accelerator computing core processes the current batch of model inference in the current buffer of the hot computing cache layer, the scheduling module, in coordination with the CPU, uses the direct memory access controller to asynchronously prefetch the feature blocks or model weights required for the next batch from system memory to the shadow buffer pool and push them into the standby buffer of the hot computing cache layer; after the current batch of computation is completed, the current buffer and the standby buffer switch roles so that the next batch of computation can directly read the prefetched data, ensuring that computation and I / O operations completely overlap on the timeline; The dynamic monitoring and redirection module is configured to: monitor the resource saturation of the optimal execution target node in real time; when the resource saturation exceeds the saturation threshold and the queue is congested, trigger the redirection strategy to migrate some of the queued subtasks across nodes to another type of node; after the cross-node migration, the system calls the SIMD instruction set to perform asynchronous compensation calculation based on vector quantization or dimensionality reduction operators to perform accuracy compensation on the approximate retrieval results generated by the migrated subtasks.

Citation Information

Patent Citations

  • Data prefetching processing method and multi-level cache processor architecture

    CN121501700A

  • Large model reasoning intelligent scheduling and memory management method and device

    CN122261768A