Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

538 results about "Request distribution" patented technology

Request For Distribution. Following is/are the item (s) or service (s) for which a distribution is requested. A receipt, purchase order or like must be attached. If additional information is needed, requester will be notified.

Cache management method and device, storage medium and electronic equipment

The invention provides a cache management method, a cache management device, a computer storage medium and electronic equipment, and relates to the technical field of computers. The method comprises the steps of receiving a reasoning task request and distributing the reasoning task request to a target storage page; key value cache information of the first round of reasoning task is stored in a hard disk cache, and when the second round of reasoning task is executed, key value cache information generated before the second round of reasoning task is preloaded layer by layer from the hard disk cache; when the last round of reasoning task is received, storing first target key value cache information correspondingly generated by the last round of reasoning task into the matched target physical block; and performing hybrid grouping compression on key cache information and value cache information in the first target key value cache information to obtain second target key value cache information after quantization compression. According to the invention, triple balance of video memory-calculation performance-precision can be realized.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Micro-service dynamic weight load balancing method and related equipment

The invention discloses a micro-service dynamic weight load balancing method and related equipment. The method comprises the following steps: asynchronously acquiring a multi-dimensional real-time load index of a micro-service instance, obtaining a load prediction value by virtue of a model, fusing the real-time index, the prediction value and a service health degree, calculating a dynamic weight by depending on service load data, distributing a request according to the weight, and feeding back an adjustment coefficient. According to the method, by asynchronously collecting the multi-dimensional real-time indexes, an accurate basis is provided for weight calculation, so that the request allocates the actual load of the matched instance, and the situation that the response of a high-load instance is slow due to too many requests is avoided; by introducing a load prediction value, load change can be pre-judged in advance, a weight strategy is adjusted in advance, overload when an instance load suddenly increases is prevented, and the fluctuation resistance of the system is improved; the weight is calculated by fusing multi-dimensional data, so that the comprehensive load of the instance can be comprehensively reflected, and evaluation deviation caused by a single index is avoided; a feedback mechanism can enable weight calculation to continuously adapt to a system state, and load unevenness caused by weight fixing or adjustment lag is avoided.
Owner:创优数字科技(广东)有限公司

Intelligent scheduling method and system for load balancing of server cluster

The invention relates to the technical field of computers, discloses an intelligent scheduling method and system for server cluster load balancing, and aims to solve the defects of the existing server cluster load balancing technology in response lag, non-uniform resource utilization rate, service quality guarantee, global optimization capability, fine-grained state perception and scheduling decision. The method comprises the following steps: collecting server state and request feature data, constructing a cluster state and service capability model, and predicting a load trend; and generating an optimal routing strategy by using deep reinforcement learning and multi-objective optimization, and issuing adjustment request distribution. The system comprises a data acquisition module, an application request feature acquisition module, a state sensing and modeling module, a load prediction module, an intelligent scheduling decision module and an instruction execution module. By adopting the technical scheme, the resource utilization rate can be improved, the response time can be reduced, the throughput can be improved, the system stability and elasticity can be enhanced, and the operation cost and energy consumption can be reduced.
Owner:LIANYUNGANG DONGLING TECHNOLOGY CO LTD

Zero-copy data transmission method and device, medium and product

The invention discloses a zero-copy data transmission method and device, a medium and a product, and belongs to the technical field of data transmission, and the method comprises the steps that an application program requests a memory buffer area from a kernel module; the kernel module allocates a physical memory buffer area according to the request, and returns a specially designed descriptor (including a handle corresponding to an address pointer); the application program submits a descriptor to the kernel module and requests an address pointer of the memory buffer area; after the kernel module verifies the descriptor, the physical memory buffer area is mapped to a virtual address space of the application program, and a corresponding address pointer is returned; the application program writes data needing to be sent according to the address pointer, and after writing is completed, the descriptor is submitted to the kernel module and the written data is requested to be sent; after the kernel module verifies the descriptor, a DMA controller is configured and started, and the written data is directly read from the physical memory buffer area and pushed to a network card to be sent. According to the invention, the security and certainty can be improved while the transmission performance is ensured.
Owner:JITAI AVIATION TECH (SUZHOU) CO LTD

PD separation reasoning framework optimization method oriented to large language model

The invention relates to the technical field of reasoning optimization, in particular to a PD separation reasoning framework optimization method oriented to a large language model, which comprises the following steps: S1, reconstructing a memory structure of a KV Cache, adjusting an original discrete storage structure allocated according to a model layer into a continuous storage structure allocated according to blocks, and changing the memory structure from [layer, k / v, block id, numhead, head, block size] into [block id, layer, k / v, numhead, head size, block size]; s2, dividing a short sequence, a medium sequence and a long sequence according to the length of the cue word, combining the short sequence into a Batch group, and preferentially extruding the short sequence and then processing the long sequence; and S3, deploying a hybrid throughput node cluster, and according to the input prompt word length dynamic allocation request, allocating a short request to a low throughput TP node of the throughput, and allocating a long request to a high throughput TP node of the throughput. The method is more suitable for a PD separated system architecture, and the transmission efficiency of the KV cache and the calculation efficiency of the GPU are improved, so that the throughput of the system is improved, and the maximum resource utilization is achieved.
Owner:PIO CLOUD COMPUTING (SHANGHAI) CO LTD

Server hardware test terminal and method

The invention discloses a server hardware testing terminal and method, and relates to the technical field of server hardware testing, and the method comprises the steps: obtaining server hardware configuration parameters and storage equipment specification information, collecting the number of processor cores, memory capacity and storage equipment read-write speed reference data through a system monitoring interface, building a hardware performance baseline file, and storing the baseline file in a server; obtaining a complete hardware resource list and a performance index range; the method comprises the following steps: constructing a multi-dimensional workload model according to information retrieval scene features, creating query request sets of different scales by adopting a random number generator, simulating a mixed load mode of full-text retrieval and database query, and determining resource consumption weights and execution time distribution of various query operations; and analyzing query request distribution characteristics in the workload model through a dynamic load balancing algorithm, if query requests are concentrated in a specific time period, adjusting a load distribution strategy, obtaining a uniformly distributed test load sequence, and judging the impact degree of a load peak value on hardware resources.
Owner:BEIJING DISCOVERY INTELLIGENT MFG TECH CO LTD

Multi-node dynamic switching method and system of MongoDB

The invention provides a multi-node dynamic switching method and a multi-node dynamic switching system for a MongoDB (MongoDB). The method comprises the following steps: acquiring configuration information of a plurality of MongoDB nodes, and establishing a data access layer based on the configuration information; through a data access layer, acquiring index data marks of the MongoDB nodes about a plurality of operation indexes according to a hierarchical frequency acquisition task configured by each MongoDB node; in combination with a node weight preset value of the MongoDB node, converting the node performance index into a weight value of the MongoDB node; based on a connection pool management mechanism and the node available state data, a request distribution strategy is determined according to the weight value, and the request distribution strategy is used for indicating the distribution proportion of the MongoDB requests distributed to the MongoDB nodes; and performing state detection on the MongoDB node through a data access layer, marking the detected MongoDB node as an unavailable state when detecting that the MongoDB node meets an abnormal state condition, and updating available state data of the node. The problems that in the prior art, manual intervention is needed, the real-time load condition of the node cannot be sensed, and the code coupling degree is high are solved.
Owner:BEIJING YULORE INNOVATION TECH

Multi-speech synthesis model bearing method and device based on virtual GPU

The invention provides a multi-speech synthesis model bearing method and device based on a virtual GPU, and relates to the technical field of graphics processing units, and the method comprises the steps: carrying out the virtualization processing of a physical graphics processing unit, dividing the physical graphics processing unit into a plurality of virtual processing units with independent video memories and calculation quotas, and combining a resource scheduling mechanism, and deploying the speech synthesis language model instances in a plurality of service containers, and constructing a plurality of speech synthesis model bearing units. After a voice synthesis request is accessed, the scheduling module carries out load balancing according to the request connection number of each bearing unit, the request is distributed to a target bearing unit with the minimum connection number, and a voice generation task is completed by a virtual processing unit bound with the target bearing unit. According to the invention, resource division can be carried out on the physical graphic processing unit, and efficient operation of the multi-speech synthesis model is realized.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Network access method and device, computer equipment and storage medium

The invention belongs to the technical field of computers, and relates to a network access method and device, computer equipment and a storage medium, and the method comprises the steps: receiving an access domain name input by a user in a target network environment; calling a target load balancer corresponding to the target network environment; analyzing the access domain name into a virtual I P address of the target load balancer based on a preset DNS server, and returning the virtual I P address to the user; receiving an access request sent by a user to the virtual IP address; based on the target load balancer, screening out a target server from a plurality of servers corresponding to the target network environment, and distributing the access request to the target server; based on the target server, performing identity verification on the user by using a preset identity verification strategy; and if the user passes the identity verification, performing corresponding response processing on the access request based on a preset access control strategy. According to the network access processing method provided by the invention, the security and availability of network access can be effectively improved.
Owner:HUNAN DATA IND GRP CO LTD

Internet of Things industry intelligent customer service supervision and control system based on artificial intelligence

The invention discloses an Internet of Things industry intelligent customer service supervision and control system based on artificial intelligence, and belongs to the technical field of customer service supervision. Comprising an omni-channel intelligent access and intention understanding module, an AI intelligent center and complex decision module, an intelligent supervision and ethical regulation and control module, a man-machine cooperation and humanistic care module, a data-driven optimization and edge intelligent module and a security privacy and controllable treatment module, and the omni-channel intelligent access and intention understanding module is used for integrating multiple channels. Unified request distribution is achieved, data input by a user are analyzed through voice recognition, NLU natural language understanding and image recognition technologies, composite intentions are accurately captured, interaction strategies are dynamically adjusted in combination with device state data, user historical behaviors and real-time positions, and pacified talking skills or manual intervention are triggered through voiceprint / text emotion analysis. On the basis of realizing customer service supervision and regulation, all-channel fusion and intention accurate analysis can be realized, and ethical safety integrated design can be realized.
Owner:YANCHENG XINZHIRUN INTELLIGENT TECHNOLOGY CO LTD

Load balancing processing method, device, storage medium, and system

PCT designated stageWO2025177049A1TransmissionApplication serverEngineering
Embodiments of the present disclosure provide a load balancing processing method, a device, a storage medium, and a system. The method comprises: obtaining a first number of Internet of Things devices corresponding to each of a plurality of application servers, wherein the first number of Internet of Things devices corresponding to any application server is the number of Internet of Things devices allowed to be accessed when a resource consumption of the application server reaches a set upper limit value; in response to a persistent connection establishment request of a target Internet of Things device, obtaining a second number of Internet of Things devices actually accessed in the plurality of application servers; determining load weights of the plurality of application servers on the basis of the first number of Internet of Things devices and the second number of Internet of Things devices corresponding to each of the plurality of application servers; and on the basis of the load weights of the plurality of application servers, determining a target application server to which the persistent connection establishment request is allocated, so as to establish a persistent connection between the target Internet of Things device and the target application server.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

DNS (Domain Name Server) fault diagnosis method and device based on statistical verification and storage medium

The invention discloses a DNS fault diagnosis method and device based on statistical verification, and a storage medium. The method comprises the following steps: extracting a domain name query request and response message information from a communication message; matching the suffixes of the domain names to determine an authoritative resolution server; according to the response code and the response time delay of the response message information, identifying an abnormal analysis request and respectively counting the total request quantity and the abnormal request quantity of each domain name and each authoritative analysis server; acquiring a forwarding path topology, and performing multi-layer aggregation on the total request quantity and the abnormal request quantity based on the forwarding path topology to generate abnormal request distribution feature data; and matching the abnormal request distribution feature data with the fault hypothesis model to identify a target analysis service device with a fault. According to the method, the authoritative resolution server is accurately positioned through the inverted-order dictionary tree, and the abnormal distribution characteristics are generated in combination with multi-layer aggregation statistics, so that automatic and accurate positioning and influence range quantification of the DNS fault are realized, and the accuracy and processing efficiency of fault diagnosis are remarkably improved.
Owner:CHINA MERCHANTS BANK

Memory operation method and device, equipment, storage medium and program product

The invention provides a memory operation method and device, equipment, a storage medium and a program product, and the method can comprise the steps that a memory operation request for a memory is obtained, the memory operation request is used for requesting memory allocation or requesting memory release, the memory comprises a plurality of memory blocks, and the memory blocks correspond to at least two memory granularities; determining memory block information corresponding to at least one memory granularity, wherein the memory block information comprises a memory granularity value, a memory block state and memory block use information; and processing the memory operation request according to the memory block information corresponding to the at least one memory granularity so as to allocate or release the memory. And the flexibility of operating the memory is improved.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Server load balancing method and device, medium and program product

The invention provides a server load balancing method and device, a medium and a program product, and a specific implementation mode of the application comprises the steps that after a load balancer receives a task request, attributes such as basic types (authentication, authorization and charging), priorities and protocol versions of the request are analyzed firstly; meanwhile, the resource utilization rate, response time, health state and other information of the server cluster are monitored in real time, a dynamic strategy decision tree is constructed based on the data, in addition, the load balancer sends a health check request to the server regularly, whether the server breaks down or not is judged according to the state information returned by the server, and the service life of the server is prolonged. And if a fault exists, removing the fault from the request allocation list to ensure that the request is always allocated to the available server. According to the method, the dynamic adjustment of the load balancing strategy and the update of the request allocation list are realized, and the resource utilization rate of the server is improved so as to cope with the rapidly changed request load and flow mode.
Owner:ZUOYEBANG EDUCATION TECH (BEIJING) CO LTD

Buffer, instruction processing method, processor and computer equipment

The invention discloses a buffer, an instruction processing method, a processor and computer equipment, and relates to the technical field of chips. The buffer is an L2-level buffer in a processor and comprises a request distribution circuit and a buffer circuit, the request distribution circuit is connected with the cache circuit; the cache circuit comprises a cache unit and an operation unit; the operation unit is connected with the cache unit; the request distribution circuit is used for receiving the cache access request and sending the cache access request to the cache circuit; the cache unit is used for reading source data corresponding to the cache access request under the condition that the cache access request is a write request corresponding to the atomic operation, and inputting the read source data into the operation unit; and the arithmetic unit is used for executing atomic operation corresponding to the cache access request on the source data to obtain operation result data and writing the operation result data into the cache unit.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Dynamic hotspot data migration method of distributed cache system

The invention belongs to the technical field of distributed cache, and relates to a dynamic hotspot data migration method of a distributed cache system. According to the method, the multi-dimensional operation state data is acquired in real time, and the real-time state data set is generated, so that the hysteresis quality and inadaptability of a traditional static fragmentation strategy in coping with hotspot data are overcome, and the sensing precision of system load distribution is improved; a dynamic linkage mechanism among hot spot prediction, migration cost simulation and node resource allocation is established, and a data migration path and a synchronization strategy are adjusted according to the difference of access trend parameters and node instantaneous processing capacity variation. The optimal matching between the hotspot data distribution and the node resource utilization rate and the dynamic balance between the service performance and the stability are realized; read-write request distribution in the data migration process is mastered in real time through background increment synchronization and flow cooperative control, a migration scheme is continuously optimized according to objective indexes, and the problems of service jitter and resource competition in the migration period are avoided.
Owner:JILIN AGRI SCI & TECH COLLEGE

Method and apparatus for allocating data storage space

A method for allocating a data storage space includes detecting one allocation request of an operating system for a continuous storage space for a target program; extracting a feature of the allocation request; determining, based on the feature of the allocation request, a fault tolerance requirement corresponding to the allocation request; and allocating a storage space of a corresponding fault tolerance level to the allocation request based on the fault tolerance requirement corresponding to the allocation request.
Owner:HUAWEI TECH CO LTD

Service request processing method and device, medium, equipment and program product

The invention discloses a service request processing method and device, a medium, equipment and a program product, and relates to the technical field of smart home / smart home. According to the method, a first service request is received, the first service request comprises at least one target resource, a lock identifier is generated for the first service request, the lock identifier comprises a serial number used for identifying the receiving sequence of the first service request, and the lock identifier of the first service request is added into distributed queues corresponding to the at least one target resource. Under the condition that the lock identifier of the first service request is arranged at the head in the distributed queue corresponding to the at least one target resource, requesting to allocate a respective resource lock of the at least one target resource for the first service request; according to the method, the uniform lock identifier containing the serial number is generated for the service request, and resource lock coordination is performed based on the distributed queue, so that the consistency of a multi-resource locking sequence is ensured, deadlock is effectively avoided, and the stability and reliability of service processing in a high-concurrency scene are improved.
Owner:QINGDAO HAIER TECH +2

OPC-based energy priority satellite edge service device

The invention relates to an OPC (optical proximity correction)-based energy priority satellite edge service device, which comprises a constellation-level orbit perception request distribution module (ORD), a satellite-level pulse task scheduling module (PTS) and a task-level energy perception runtime management module (ERM). And high-throughput, low-delay and low-energy-consumption intelligent services in a multi-satellite and multi-task environment are realized.
Owner:SHANGHAI JIAOTONG UNIV

Service node processing method and device, equipment and medium

The invention relates to the technical field of computers, and discloses a service node processing method, device and equipment and a medium, which are applied to a service node scheduling system and comprise a scheduling module, a calculation module, a processing module and a state management module. Obtaining an effective weight and a current weight of each service node in the current service cluster; calculating a target scheduling priority value of the service node based on the current weight and the effective weight through a calculation module; determining the service node with the highest target scheduling priority value as a target service node through a processing module, and processing the target service node to obtain a processing result; and if the state management module determines that the processing result is in a successful state, the effective weight of the target service node is increased. The method can be applied to business management program systems such as financial science and technology, medical health, old-age care and the like, and the accuracy of service node request distribution and the stability of system operation can be improved.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Distributed large model reasoning optimization method and device based on dynamic micro-batch scheduling

The invention discloses a distributed large model reasoning optimization method and device based on dynamic micro-batch scheduling, and the method comprises the steps: (1) initializing a system, and building a large model distributed reasoning assembly line; (2) evaluating the computing power and the network state of each node and summarizing the computing power and the network state to a head node; (3) determining the number of Micro-Batches and the scheduling quota of each Micro-Batch according to the request distribution condition, the computing power of each node and the network state; and (4) sequentially scheduling the Micro-Batch by adopting a Control Batch strategy and a Chunked Prefill strategy, and starting to execute the Micro-Batch. According to the method, by dynamically adjusting the number of Micro-Batch, the serious assembly line cavitation problem in a distributed large model reasoning system is effectively solved, the GPU utilization rate and the system throughput are remarkably improved, and meanwhile key indexes such as TTFT (first token time delay) and TPOT (inter-token time delay) in the large model reasoning field are also improved. The method has good adaptability, can adaptively adjust the dynamic scheduling strategy under different hardware equipment, network conditions and request loads, is suitable for different distributed large model deployment scenes, and has wide application value.
Owner:ZHEJIANG UNIV

Industrial identification flow intelligent scheduling method and system

The invention relates to the technical field of industrial network identification traffic scheduling, and discloses an industrial identification traffic intelligent scheduling method and system, and the method comprises the steps: S1, receiving an identification analysis request sent by an industrial network terminal, extracting a service emergency degree feature of an application layer and a transmission reliability feature of a transmission layer in the identification analysis request; s2, analyzing a dynamic coupling influence relationship between the service emergency degree feature and the transmission reliability feature, and generating a priority feature tag corresponding to the request based on the dynamic coupling influence relationship in combination with a preset balance strategy; s3, distributing the identification analysis request to a preset high-throughput processing path or a low-delay processing path based on the priority feature tag; according to the invention, the problems that the conflict between the service timeliness and the transmission stability is difficult to balance and the dynamic scheduling requirement of diversified services in an industrial scene is difficult to meet in the prior art can be solved.
Owner:ZHONGKE ZHENGTONG (JINAN) INFORMATION TECHNOLOGY CO LTD

Secure access method and device for processor resources, equipment, medium and product

The invention relates to the technical field of embedded systems, in particular to a secure access method and device for processor resources, equipment, a medium and a product, and the method comprises the steps: distributing each resource access request to a corresponding target virtual domain according to a matching result of a main equipment process identifier of each resource access request and a corresponding preset mask, injecting the distributed domain identifier into each request to obtain a domain marking request; intercepting and analyzing the domain marking requests to obtain a target address, an operation type and a domain identifier corresponding to each domain marking request; and verifying the operation type of the domain identifier corresponding to each domain marking request on the target address based on preset security policy information, and intercepting the abnormal access request. Therefore, the technical problems of large performance loss, insufficient security isolation granularity and lack of system-level protection in the existing intra-core resource isolation scheme are solved, the real-time performance and security of the system are remarkably improved, and the troubleshooting capability is enhanced.
Owner:INALFA ZHILIAN TECH (BEIJING) CO LTD

Direct memory access system with read reassembly circuit

A direct memory access (DMA) system includes a plurality of read circuits and a switch coupled to a plurality of data port controllers configured to communicate with one or more data processing systems. The DMA system includes a read scheduler circuit coupled to the plurality of read circuits and the switch. The read scheduler circuit is configured to receive read requests from the plurality of read circuits, request allocation of entries of a data memory for the read requests, and submit the read requests to the one more data processing systems via the switch. The DMA system includes a read reassembly circuit coupled to the plurality of read circuits, the switch, and the read scheduler circuit. The read reassembly circuit is configured to reorder read completion data received from the switch for the read requests and provide read completion data, as reordered, to the plurality of read circuits.
Owner:XILINX INC

High-concurrency request processing method and device for reducing Linux kernel resource loss

The invention relates to a high-concurrency request processing method for reducing Linux kernel resource loss. The high-concurrency request processing method comprises the following steps: receiving a user request and extracting a routing key; the requests are distributed to one fragment in the fixed N fragments through a fragment event executor, each fragment is bound with a single thread and a bounded queue, and the same routing key requests are executed in the same fragment in a strongly sequenced mode according to the enqueue sequence; calling a coprogram state machine to push a business process in the fragment execution thread, and returning a state result of migrating to a next state, keeping a current state or terminating the process through a state processing function; dynamically triggering self-adaptive back pressure based on queue depth or waiting time delay, and adjusting a request submission strategy according to a queue filling rate and exponentially weighted moving average waiting time delay; and submitting the long blocking task to an asynchronous task subsystem, and performing isolated execution through an independent bounded queue and a small concurrent thread pool. The same-key sequence is guaranteed through fragmentation single-thread execution, and thread switching and lock contention are remarkably reduced.
Owner:BEIJING MICO WORLD TECH CO LTD

Intelligent router for distributing requests to different generative ai instances

PCT designated stageWO2026035354A1Biological modelsRouting modelEngineering
An intelligent router for generative artificial intelligence (GAI) model instances optimizes request routing to reduce latency. The system predicts output lengths using a trained response-length predictor and assesses the state of multiple GAI instances, including prompt and decode distributions. It estimates the workload mixing impact of routing requests to each instance and determines selection probabilities using a machine-learning routing model. The router either assigns the request to the most suitable instance or delays routing if conditions are suboptimal. This approach improves end-to-end latency, Time-To-First-Token (TTFT), and Time-Between-Tokens (TBT) by considering the distinct characteristics of GAI workload phases.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Virtual consistency multi-node message passing interface expansion method, device and equipment

The invention relates to a virtual consistency multi-node message passing interface expansion method, device and equipment. The method comprises the following steps: deploying a distributed file system client at each computing node of a multi-node cluster, creating a unified virtual directory mounting point, transparently converting a file operation request under the unified virtual directory mounting point into a corresponding network protocol, and docking a global unified data storage pool, the unified virtual directory mounting point is used for providing consistent file views for all the computing nodes; pointing a file read-write path prefix of the message passing interface process to the unified virtual directory mounting point; and in response to the received user operation request, allocating computing nodes to the user operation request, ensuring that all the allocated computing nodes are mounted to the distributed file system, and starting a message passing interface process. By adopting the method, the consistency of the cross-node file view can be solved from the system level on the premise of not changing the communication logic.
Owner:SHANG HAI ZHANG JIANG SHU XUE YAN JIU YUAN

Cloud service-based model inference service system and method

The application relates to a cloud service-based model inference service system and method capable of on-demand scaling. The system comprises a gateway module, a model inference service management module and a request distribution module. The gateway module is used for receiving a user request and extracting heterogeneous features representing the request's demand for computing resources. The model inference service management module is used for generating resource scheduling instructions for each inference service instance based on the multi-dimensional indicators of the deployed inference service instances, in combination with the resource cost and availability information of each cloud resource pool. The request distribution module is used for distributing the user request to the inference service instance that is adapted to the heterogeneous features according to the resource scheduling instructions. The system can solve the problems of low utilization rate of computing resources, poor request adaptability and lack of cross-cloud scheduling in the prior art.
Owner:ZHEJIANG LAB