Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

412 results about "Request distribution" patented technology

Request For Distribution. Following is/are the item (s) or service (s) for which a distribution is requested. A receipt, purchase order or like must be attached. If additional information is needed, requester will be notified.

Cache management method and device, storage medium and electronic equipment

The invention provides a cache management method, a cache management device, a computer storage medium and electronic equipment, and relates to the technical field of computers. The method comprises the steps of receiving a reasoning task request and distributing the reasoning task request to a target storage page; key value cache information of the first round of reasoning task is stored in a hard disk cache, and when the second round of reasoning task is executed, key value cache information generated before the second round of reasoning task is preloaded layer by layer from the hard disk cache; when the last round of reasoning task is received, storing first target key value cache information correspondingly generated by the last round of reasoning task into the matched target physical block; and performing hybrid grouping compression on key cache information and value cache information in the first target key value cache information to obtain second target key value cache information after quantization compression. According to the invention, triple balance of video memory-calculation performance-precision can be realized.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Micro-service dynamic weight load balancing method and related equipment

The invention discloses a micro-service dynamic weight load balancing method and related equipment. The method comprises the following steps: asynchronously acquiring a multi-dimensional real-time load index of a micro-service instance, obtaining a load prediction value by virtue of a model, fusing the real-time index, the prediction value and a service health degree, calculating a dynamic weight by depending on service load data, distributing a request according to the weight, and feeding back an adjustment coefficient. According to the method, by asynchronously collecting the multi-dimensional real-time indexes, an accurate basis is provided for weight calculation, so that the request allocates the actual load of the matched instance, and the situation that the response of a high-load instance is slow due to too many requests is avoided; by introducing a load prediction value, load change can be pre-judged in advance, a weight strategy is adjusted in advance, overload when an instance load suddenly increases is prevented, and the fluctuation resistance of the system is improved; the weight is calculated by fusing multi-dimensional data, so that the comprehensive load of the instance can be comprehensively reflected, and evaluation deviation caused by a single index is avoided; a feedback mechanism can enable weight calculation to continuously adapt to a system state, and load unevenness caused by weight fixing or adjustment lag is avoided.
Owner:创优数字科技(广东)有限公司

Intelligent scheduling method and system for load balancing of server cluster

The invention relates to the technical field of computers, discloses an intelligent scheduling method and system for server cluster load balancing, and aims to solve the defects of the existing server cluster load balancing technology in response lag, non-uniform resource utilization rate, service quality guarantee, global optimization capability, fine-grained state perception and scheduling decision. The method comprises the following steps: collecting server state and request feature data, constructing a cluster state and service capability model, and predicting a load trend; and generating an optimal routing strategy by using deep reinforcement learning and multi-objective optimization, and issuing adjustment request distribution. The system comprises a data acquisition module, an application request feature acquisition module, a state sensing and modeling module, a load prediction module, an intelligent scheduling decision module and an instruction execution module. By adopting the technical scheme, the resource utilization rate can be improved, the response time can be reduced, the throughput can be improved, the system stability and elasticity can be enhanced, and the operation cost and energy consumption can be reduced.
Owner:LIANYUNGANG DONGLING TECHNOLOGY CO LTD

Server hardware test terminal and method

The invention discloses a server hardware testing terminal and method, and relates to the technical field of server hardware testing, and the method comprises the steps: obtaining server hardware configuration parameters and storage equipment specification information, collecting the number of processor cores, memory capacity and storage equipment read-write speed reference data through a system monitoring interface, building a hardware performance baseline file, and storing the baseline file in a server; obtaining a complete hardware resource list and a performance index range; the method comprises the following steps: constructing a multi-dimensional workload model according to information retrieval scene features, creating query request sets of different scales by adopting a random number generator, simulating a mixed load mode of full-text retrieval and database query, and determining resource consumption weights and execution time distribution of various query operations; and analyzing query request distribution characteristics in the workload model through a dynamic load balancing algorithm, if query requests are concentrated in a specific time period, adjusting a load distribution strategy, obtaining a uniformly distributed test load sequence, and judging the impact degree of a load peak value on hardware resources.
Owner:BEIJING DISCOVERY INTELLIGENT MFG TECH CO LTD

Multi-node dynamic switching method and system of MongoDB

The invention provides a multi-node dynamic switching method and a multi-node dynamic switching system for a MongoDB (MongoDB). The method comprises the following steps: acquiring configuration information of a plurality of MongoDB nodes, and establishing a data access layer based on the configuration information; through a data access layer, acquiring index data marks of the MongoDB nodes about a plurality of operation indexes according to a hierarchical frequency acquisition task configured by each MongoDB node; in combination with a node weight preset value of the MongoDB node, converting the node performance index into a weight value of the MongoDB node; based on a connection pool management mechanism and the node available state data, a request distribution strategy is determined according to the weight value, and the request distribution strategy is used for indicating the distribution proportion of the MongoDB requests distributed to the MongoDB nodes; and performing state detection on the MongoDB node through a data access layer, marking the detected MongoDB node as an unavailable state when detecting that the MongoDB node meets an abnormal state condition, and updating available state data of the node. The problems that in the prior art, manual intervention is needed, the real-time load condition of the node cannot be sensed, and the code coupling degree is high are solved.
Owner:BEIJING YULORE INNOVATION TECH

Multi-speech synthesis model bearing method and device based on virtual GPU

The invention provides a multi-speech synthesis model bearing method and device based on a virtual GPU, and relates to the technical field of graphics processing units, and the method comprises the steps: carrying out the virtualization processing of a physical graphics processing unit, dividing the physical graphics processing unit into a plurality of virtual processing units with independent video memories and calculation quotas, and combining a resource scheduling mechanism, and deploying the speech synthesis language model instances in a plurality of service containers, and constructing a plurality of speech synthesis model bearing units. After a voice synthesis request is accessed, the scheduling module carries out load balancing according to the request connection number of each bearing unit, the request is distributed to a target bearing unit with the minimum connection number, and a voice generation task is completed by a virtual processing unit bound with the target bearing unit. According to the invention, resource division can be carried out on the physical graphic processing unit, and efficient operation of the multi-speech synthesis model is realized.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Internet of Things industry intelligent customer service supervision and control system based on artificial intelligence

The invention discloses an Internet of Things industry intelligent customer service supervision and control system based on artificial intelligence, and belongs to the technical field of customer service supervision. Comprising an omni-channel intelligent access and intention understanding module, an AI intelligent center and complex decision module, an intelligent supervision and ethical regulation and control module, a man-machine cooperation and humanistic care module, a data-driven optimization and edge intelligent module and a security privacy and controllable treatment module, and the omni-channel intelligent access and intention understanding module is used for integrating multiple channels. Unified request distribution is achieved, data input by a user are analyzed through voice recognition, NLU natural language understanding and image recognition technologies, composite intentions are accurately captured, interaction strategies are dynamically adjusted in combination with device state data, user historical behaviors and real-time positions, and pacified talking skills or manual intervention are triggered through voiceprint / text emotion analysis. On the basis of realizing customer service supervision and regulation, all-channel fusion and intention accurate analysis can be realized, and ethical safety integrated design can be realized.
Owner:YANCHENG XINZHIRUN INTELLIGENT TECHNOLOGY CO LTD

DNS (Domain Name Server) fault diagnosis method and device based on statistical verification and storage medium

The invention discloses a DNS fault diagnosis method and device based on statistical verification, and a storage medium. The method comprises the following steps: extracting a domain name query request and response message information from a communication message; matching the suffixes of the domain names to determine an authoritative resolution server; according to the response code and the response time delay of the response message information, identifying an abnormal analysis request and respectively counting the total request quantity and the abnormal request quantity of each domain name and each authoritative analysis server; acquiring a forwarding path topology, and performing multi-layer aggregation on the total request quantity and the abnormal request quantity based on the forwarding path topology to generate abnormal request distribution feature data; and matching the abnormal request distribution feature data with the fault hypothesis model to identify a target analysis service device with a fault. According to the method, the authoritative resolution server is accurately positioned through the inverted-order dictionary tree, and the abnormal distribution characteristics are generated in combination with multi-layer aggregation statistics, so that automatic and accurate positioning and influence range quantification of the DNS fault are realized, and the accuracy and processing efficiency of fault diagnosis are remarkably improved.
Owner:CHINA MERCHANTS BANK

Server load balancing method and device, medium and program product

The invention provides a server load balancing method and device, a medium and a program product, and a specific implementation mode of the application comprises the steps that after a load balancer receives a task request, attributes such as basic types (authentication, authorization and charging), priorities and protocol versions of the request are analyzed firstly; meanwhile, the resource utilization rate, response time, health state and other information of the server cluster are monitored in real time, a dynamic strategy decision tree is constructed based on the data, in addition, the load balancer sends a health check request to the server regularly, whether the server breaks down or not is judged according to the state information returned by the server, and the service life of the server is prolonged. And if a fault exists, removing the fault from the request allocation list to ensure that the request is always allocated to the available server. According to the method, the dynamic adjustment of the load balancing strategy and the update of the request allocation list are realized, and the resource utilization rate of the server is improved so as to cope with the rapidly changed request load and flow mode.
Owner:ZUOYEBANG EDUCATION TECH (BEIJING) CO LTD

Dynamic hotspot data migration method of distributed cache system

The invention belongs to the technical field of distributed cache, and relates to a dynamic hotspot data migration method of a distributed cache system. According to the method, the multi-dimensional operation state data is acquired in real time, and the real-time state data set is generated, so that the hysteresis quality and inadaptability of a traditional static fragmentation strategy in coping with hotspot data are overcome, and the sensing precision of system load distribution is improved; a dynamic linkage mechanism among hot spot prediction, migration cost simulation and node resource allocation is established, and a data migration path and a synchronization strategy are adjusted according to the difference of access trend parameters and node instantaneous processing capacity variation. The optimal matching between the hotspot data distribution and the node resource utilization rate and the dynamic balance between the service performance and the stability are realized; read-write request distribution in the data migration process is mastered in real time through background increment synchronization and flow cooperative control, a migration scheme is continuously optimized according to objective indexes, and the problems of service jitter and resource competition in the migration period are avoided.
Owner:JILIN AGRI SCI & TECH COLLEGE

Method and apparatus for allocating data storage space

A method for allocating a data storage space includes detecting one allocation request of an operating system for a continuous storage space for a target program; extracting a feature of the allocation request; determining, based on the feature of the allocation request, a fault tolerance requirement corresponding to the allocation request; and allocating a storage space of a corresponding fault tolerance level to the allocation request based on the fault tolerance requirement corresponding to the allocation request.
Owner:HUAWEI TECH CO LTD

Service request processing method and device, medium, equipment and program product

The invention discloses a service request processing method and device, a medium, equipment and a program product, and relates to the technical field of smart home / smart home. According to the method, a first service request is received, the first service request comprises at least one target resource, a lock identifier is generated for the first service request, the lock identifier comprises a serial number used for identifying the receiving sequence of the first service request, and the lock identifier of the first service request is added into distributed queues corresponding to the at least one target resource. Under the condition that the lock identifier of the first service request is arranged at the head in the distributed queue corresponding to the at least one target resource, requesting to allocate a respective resource lock of the at least one target resource for the first service request; according to the method, the uniform lock identifier containing the serial number is generated for the service request, and resource lock coordination is performed based on the distributed queue, so that the consistency of a multi-resource locking sequence is ensured, deadlock is effectively avoided, and the stability and reliability of service processing in a high-concurrency scene are improved.
Owner:QINGDAO HAIER TECH +2

OPC-based energy priority satellite edge service device

The invention relates to an OPC (optical proximity correction)-based energy priority satellite edge service device, which comprises a constellation-level orbit perception request distribution module (ORD), a satellite-level pulse task scheduling module (PTS) and a task-level energy perception runtime management module (ERM). And high-throughput, low-delay and low-energy-consumption intelligent services in a multi-satellite and multi-task environment are realized.
Owner:SHANGHAI JIAOTONG UNIV

Secure access method and device for processor resources, equipment, medium and product

The invention relates to the technical field of embedded systems, in particular to a secure access method and device for processor resources, equipment, a medium and a product, and the method comprises the steps: distributing each resource access request to a corresponding target virtual domain according to a matching result of a main equipment process identifier of each resource access request and a corresponding preset mask, injecting the distributed domain identifier into each request to obtain a domain marking request; intercepting and analyzing the domain marking requests to obtain a target address, an operation type and a domain identifier corresponding to each domain marking request; and verifying the operation type of the domain identifier corresponding to each domain marking request on the target address based on preset security policy information, and intercepting the abnormal access request. Therefore, the technical problems of large performance loss, insufficient security isolation granularity and lack of system-level protection in the existing intra-core resource isolation scheme are solved, the real-time performance and security of the system are remarkably improved, and the troubleshooting capability is enhanced.
Owner:INALFA ZHILIAN TECH (BEIJING) CO LTD

High-concurrency request processing method and device for reducing Linux kernel resource loss

The invention relates to a high-concurrency request processing method for reducing Linux kernel resource loss. The high-concurrency request processing method comprises the following steps: receiving a user request and extracting a routing key; the requests are distributed to one fragment in the fixed N fragments through a fragment event executor, each fragment is bound with a single thread and a bounded queue, and the same routing key requests are executed in the same fragment in a strongly sequenced mode according to the enqueue sequence; calling a coprogram state machine to push a business process in the fragment execution thread, and returning a state result of migrating to a next state, keeping a current state or terminating the process through a state processing function; dynamically triggering self-adaptive back pressure based on queue depth or waiting time delay, and adjusting a request submission strategy according to a queue filling rate and exponentially weighted moving average waiting time delay; and submitting the long blocking task to an asynchronous task subsystem, and performing isolated execution through an independent bounded queue and a small concurrent thread pool. The same-key sequence is guaranteed through fragmentation single-thread execution, and thread switching and lock contention are remarkably reduced.
Owner:BEIJING MICO WORLD TECH CO LTD

Intelligent router for distributing requests to different generative ai instances

PCT designated stageWO2026035354A1Biological modelsRouting modelEngineering
An intelligent router for generative artificial intelligence (GAI) model instances optimizes request routing to reduce latency. The system predicts output lengths using a trained response-length predictor and assesses the state of multiple GAI instances, including prompt and decode distributions. It estimates the workload mixing impact of routing requests to each instance and determines selection probabilities using a machine-learning routing model. The router either assigns the request to the most suitable instance or delays routing if conditions are suboptimal. This approach improves end-to-end latency, Time-To-First-Token (TTFT), and Time-Between-Tokens (TBT) by considering the distinct characteristics of GAI workload phases.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Virtual consistency multi-node message passing interface expansion method, device and equipment

The invention relates to a virtual consistency multi-node message passing interface expansion method, device and equipment. The method comprises the following steps: deploying a distributed file system client at each computing node of a multi-node cluster, creating a unified virtual directory mounting point, transparently converting a file operation request under the unified virtual directory mounting point into a corresponding network protocol, and docking a global unified data storage pool, the unified virtual directory mounting point is used for providing consistent file views for all the computing nodes; pointing a file read-write path prefix of the message passing interface process to the unified virtual directory mounting point; and in response to the received user operation request, allocating computing nodes to the user operation request, ensuring that all the allocated computing nodes are mounted to the distributed file system, and starting a message passing interface process. By adopting the method, the consistency of the cross-node file view can be solved from the system level on the premise of not changing the communication logic.
Owner:SHANG HAI ZHANG JIANG SHU XUE YAN JIU YUAN

Cloud service-based model inference service system and method

The application relates to a cloud service-based model inference service system and method capable of on-demand scaling. The system comprises a gateway module, a model inference service management module and a request distribution module. The gateway module is used for receiving a user request and extracting heterogeneous features representing the request's demand for computing resources. The model inference service management module is used for generating resource scheduling instructions for each inference service instance based on the multi-dimensional indicators of the deployed inference service instances, in combination with the resource cost and availability information of each cloud resource pool. The request distribution module is used for distributing the user request to the inference service instance that is adapted to the heterogeneous features according to the resource scheduling instructions. The system can solve the problems of low utilization rate of computing resources, poor request adaptability and lack of cross-cloud scheduling in the prior art.
Owner:ZHEJIANG LAB

Elevator dispatch arbitration method and system

The present application relates to elevator dispatching technical field, disclose a kind of elevator dispatching arbitration method, comprising the following steps: step 1: receiving multi-source call request, corresponding identity fingerprint is assigned to each call request and monitors the waiting time of each call request;Step 2: obtain the current load of target elevator and car remaining space, whether the current load of target elevator and car remaining space meets corresponding call request is judged;When the current load of target elevator or car remaining space does not meet corresponding call request, then cancel the call request;When the current load of target elevator and car remaining space both meet corresponding call request, then carry out step 3;Step 3: using dispatching score model to calculate call request, obtain the score of call request;Step 4: using shadow dispatcher to send shadow instruction to target elevator, shadow instruction triggers the corresponding stop layer action of target elevator.Simultaneously, a kind of elevator dispatching arbitration system is also disclosed.
Owner:GUANGZHOU ROBUSTEL CO LTD

Cloud platform computing resource allocation method and system, terminal and storage medium

The application provides a cloud platform computing resource allocation method, system, terminal and storage medium, comprising: collecting idle resource rates of each host in a cloud platform cluster; converting each idle resource rate of each host into a corresponding recommended reference value by using a normalization exponential function; calculating the product of each recommended reference value of each host respectively, and accumulating the product of all hosts as a recommended coefficient; calculating a host idle degree value according to the recommended coefficient and the product of each recommended reference value of the host; and allocating a corresponding host to a virtual machine creation request according to the principle of preferentially allocating the highest idle degree value. The application converts data into relative recommended probability by analyzing the idle resource conditions of all hosts in the cluster through a softmax function, and then obtains the recommended value of the host by using DS evidence theory for information fusion, so as to allocate appropriate hosts for user requests. The application can reasonably allocate resources when planning multiple user creation of virtual machines, thereby improving the efficiency of the system.
Owner:JINAN INSPUR DATA TECH CO LTD

Charging method and device

The invention discloses a charging method and equipment. In the method, a first control node receives a first request for requesting to allocate resources; the first control node obtains a charging policy and charging information, the charging policy comprises a charging policy corresponding to at least one charging type, the charging information comprises charging information corresponding to at least one charging type, and the charging type comprises one or more of the following types: a function type, a value type, a resource type and a scenario type; and the first control node sends the charging strategy and the charging information to at least one execution node, so that the at least one execution node reports resource use information according to the charging strategy and the charging information, and the at least one execution node is an execution node which is allocated by the first control node and used for providing resources. According to the method, a single charging strategy and charging information are not adopted any more, different charging strategies and charging information can be adopted for different charging types, and diversified charging requirements of different services can be met.
Owner:HUAWEI TECH CO LTD

Memory management method, electronic equipment and vehicle

The invention relates to a memory management method, electronic equipment and a vehicle. The method comprises the steps of receiving a memory allocation request; wherein the memory allocation request indicates the size of the memory requested to be allocated; determining the size of a to-be-allocated memory and a memory address linked list corresponding to the to-be-allocated memory according to the size of the memory requested to be allocated and indicated by the memory allocation request and the size of a preset memory unit; wherein the size of the preset memory unit indicates the size of the pre-divided memory unit; the memory address linked list is used for storing memory addresses; according to the determined size of the to-be-allocated memory, acquiring a memory address of the to-be-allocated memory from a memory address linked list corresponding to the determined to-be-allocated memory; and allocating the to-be-allocated memory based on the determined memory address of the to-be-allocated memory. On the basis of improving the reasonability of confirming the address of the memory to be allocated, the reasonability and user experience of memory management based on the determined address of the memory to be allocated are further improved.
Owner:CHONGQING CHANGAN TECH CO LTD

Industrial identification flow intelligent scheduling method and system

The application relates to the technical field of industrial network identification flow scheduling, and discloses an intelligent industrial identification flow scheduling method and system.The method comprises the following steps: S1, receiving an identification analysis request sent by an industrial network terminal, and extracting a service urgency feature of an application layer and a transmission reliability feature of a transmission layer in the identification analysis request; S2, analyzing a dynamic coupling influence relationship between the service urgency feature and the transmission reliability feature, and generating a priority feature label corresponding to the request based on the dynamic coupling influence relationship and a preset balancing strategy; and S3, distributing the identification analysis request to a preset high-throughput processing path or a low-latency processing path based on the priority feature label.The application can solve the problems that the prior art is difficult to balance the conflict between service timeliness and transmission stability and difficult to meet the dynamic scheduling demand of diversified services in an industrial scene.
Owner:ZHONGKE ZHENGTONG (JINAN) INFORMATION TECHNOLOGY CO LTD

Multi-cluster resource processing method and device

The invention discloses a multi-cluster resource processing method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the steps that in response to a resource processing request sent by a client, at least one target cluster is determined according to the resource processing request, and the resource processing request comprises a resource type and an operation type; determining a request forwarding path according to the resource type and the operation type; according to the request forwarding path, allocating the resource processing request to at least one target cluster for processing to obtain a processing result of each target cluster; and aggregating the processing results of the target clusters to obtain a resource processing result. According to the embodiment, resource aggregation processing across multiple clusters can be realized in a small-scale and simple multi-cluster scene or a complex multi-cluster scene, the scene adaptability is high, and the implementation is simple; and meanwhile, different resource aggregation processing is carried out on different resource types and operation types, so that the flexibility of resource aggregation processing is improved.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Request processing methods and apparatus, electronic devices and storage media

This disclosure provides a request processing method and apparatus, an electronic device, and a storage medium. The request processing method is applied to an artificial intelligence processor and includes: receiving a first request; assigning an identifier to the first request to identify the first request; determining, based on the target storage address of the first request, whether a second request exists in a set including pending requests and requests in execution, wherein the second request is a request having the same target storage address as the first request; in response to the existence of the second request, establishing a dependency relationship between the first request and the second request, wherein the dependency relationship indicates that the first request will be executed after the second request is completed; removing the dependency relationship after the second request is executed; and executing the first request. This method can establish dependencies between requests, enabling the serialized execution of requests.
Owner:SHANGHAI BIREN TECH CO LTD

Method for determining an abnormal request

The embodiment of the application provides a kind of determination method of abnormal request, this method includes: through the tracker deployed in node, corresponding entry is allocated for each request to node and the entry identification of generating entry identification;According to the generation order of entry identification, generate entry identification chain, wherein, entry identification chain includes entry identification and the validity state of entry identification;Through the polling detection of request in timer in tracker, determine timeout request, and the request corresponding to the first entry identification of chain is determined as abnormal request;Wherein, the first entry identification of chain is the earliest generation and the validity state is valid in entry identification chain.The technical problem that related art cannot accurately detect and locate abnormal request under the condition of low resource consumption is solved by the embodiment of the application.
Owner:SANECHIPS TECH CO LTD

Technologies for isolating network functions

Examples described herein include shared reserved memory regions providing communications among network functions for isolation among network slices. In some examples, circuitry is configured to: based on receipt of a first request, allocate a first region of one or more memory regions of a memory to store data reserved for access by a first network function, wherein the first network function comprises an Open Radio Access Network (ORAN) Control Unit (CU) of a radio access network (RAN) and wherein an ORAN Distributed Unit (DU) is to provide the data and based on receipt of a second request, allocate a second region of one or more memory regions of the memory to store second data reserved for access by a second network function, wherein the second network function comprises a second ORAN CU of the RAN, the DU is to provide the second data, and the DU is shared among the first CU and the second CU.
Owner:INTEL CORP

Large model reasoning service dynamic connection number limiting method based on service attributes

PendingCN121585724AEnergy efficient computingTransmissionData setThree dimensionality
The invention relates to the technical field of computers, and discloses a large model reasoning service dynamic connection number limiting method based on service attributes, which comprises the following steps of: 1, carrying out initialization setting on service attributes, resource parameters and a weight model; 2, collecting multi-dimensional data, wherein the collected multi-dimensional data comprises load data, service data and resource data; step 3, carrying out data preprocessing and data fusion on the collected multi-dimensional data to form a three-dimensional fusion data set; step 4, calculating a three-dimensional dynamic weight based on the three-dimensional fusion data set; 5, generating a connection number threshold value according to the three-dimensional dynamic weight, and carrying out connection request distribution and over-limit processing; and step 6, forming closed-loop feedback through real-time monitoring, historical backtracking and algorithm tuning. By utilizing the scheme of the invention, the refined management and control of the connection number in the large model reasoning service are realized, and the resource cost is optimized while the high-priority service quality is ensured.
Owner:ASPIRE TECH (SHENZHEN) LTD

Dynamically assigning user devices to workload clusters

ActiveUS12717649B2User deviceCluster systems
Systems and methods described herein relate to the assignment of user devices to workload clusters. Resource utilization on a plurality of user devices is monitored. A device agent on each user device may be used to monitor the resource utilization. A workload execution request identifies resource requirements of a cluster workload. The workload execution request is assigned to a cluster based on the resource requirements of the cluster workload and the resource utilization on the plurality of user devices. The cluster comprises the plurality of user devices. The workload execution request is caused to be executed on the cluster. Each user device executes user workloads and a respective portion of the cluster workload.
Owner:SAP SE