Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20449results about "Resource allocation" patented technology

AI-Optimized Memory Fabric for Large Contexts and Multimodal Workloads

A coherent, intelligent, packet-switched memory fabric enables predictive, cache-coherent access across distributed compute, accelerator, and memory resources using a Memory-Fabric Transaction Layer Protocol (MF-TLP). MF-TLP defines routable packet formats for read, write, vectorized, atomic, reduction, collective, and predictive-prefetch transactions executed by memory-centric network interface controllers (MC-NICs). Each MC-NIC performs packet parsing, address translation, coherence management, and near-memory arithmetic or tensor operations while coordinating with MF-TLP-aware switches providing hierarchical directory control, multi-path routing, and in-network aggregation. Vectorized and multimodal packets encode multiple addresses or tensor offsets to reduce scatter / gather overhead, and programmable caching and quality-of-service modules manage tiered memory and tenant fairness. MF-TLP supports extension headers for predictive prefetch, collective coordination, and tenant governance, operating across hierarchical leaf-spine topologies using Ultra-Ethernet Transport, InfiniBand, or CXL fabrics. The system delivers scalable, low-latency, memory-centric orchestration for large-language-model training, multimodal AI, and data-intensive analytics.
Owner:QOMPLX INC

AI Serving Hardware and Software Frontier Enhancements

A computer system implements a unified framework integrating an adaptive elastic funnel (AEF) with a convergent intelligence fabric (CIF) for multi-agent AI collaboration. The system provides a universal multi-modal key-value subsystem for sharing partial computations, implements hybrid placement strategies for dynamic memory management, and incorporates quantum-resistant secure enclaves. The architecture integrates hardware acceleration through GPU-FPGA hybrid caching and neuromorphic processors, applies adaptive energy and thermal management across hardware generations, and implements autonomous flash resource orchestration with multi-dimensional wear management. The system orchestrates tensor workflows using hierarchical scheduling, enables cross-agent collaboration with privacy preservation, and supports continuous learning without catastrophic forgetting. This integration delivers unprecedented computational efficiency and security in high-dimensional decision-making environments while supporting incremental adoption through modular interfaces.
Owner:QOMPLX INC

Adaptive Real-Time Multi-Modal Compression System with Dynamic Resource Allocation

A system and method for adaptive real-time multi-modal compression with dynamic resource allocation provides intelligent compression optimization based on continuously monitored device conditions. The system monitors battery level, CPU utilization, and memory availability while classifying incoming multi-modal data streams comprising image, audio, text, and sensor data to determine processing priorities. Multi-objective optimization balances compression efficiency, reconstruction quality, and energy consumption using evolutionary algorithms that generate optimal parameters for an adaptive variational autoencoder. The autoencoder features dynamically selectable processing complexity, adjustable latent space dimensionality, and modality-specific processing layers. The system automatically switches between operational modes including emergency mode triggered by resource constraints, which applies maximum compression settings and intelligent data triage. Continuous learning adapts compression parameters based on observed performance outcomes, improving future optimization decisions. The system enables homomorphic operations on compressed data and provides enhanced compression performance under varying resource constraints across diverse edge computing applications.
Owner:ATOMBEAM TECH INC

Heterogeneous resource computing power intelligent scheduling method and system

The invention relates to the technical field of computing power scheduling, and discloses a heterogeneous resource computing power intelligent scheduling method and system. According to the method, real-time state monitoring is conducted on heterogeneous computing resources, and resource state parameters such as the computing unit utilization rate and the memory occupancy rate are obtained; task attributes and user request parameters of the task queue are collected, historical task data are processed based on the genetic algorithm optimization model to execute task demand prediction, and predicted demand parameters are generated. A dependency graph containing resource unit nodes and communication link roadsides is constructed through a resource topology analysis tool, predicted demand parameters are input into a scheduling priority classifier trained by a graph neural network, and an actual scheduling priority is identified. And executing resource conflict prediction based on the priority, inputting task feature vectors into a conflict resolution module of a fuzzy logic decision maker, outputting actual conflict resolution parameters, and finally integrating to generate a scheduling scheme containing a resource allocation sequence and an execution time table.
Owner:BEIJING WEICHENG TECHNOLOGY CO LTD

Computing power scheduling method and system based on dynamic load prediction and resource priority ranking

The invention discloses a computing power scheduling method and system based on dynamic load prediction and resource priority ranking. The computing power scheduling method comprises the following steps: collecting historical load data, task submission data and resource state data of each node in a computing power cluster; on the basis of the preprocessed multi-dimensional load feature data set, constructing an improved hybrid prediction model, optimizing model parameters through training, and predicting the load change trend of each computing power node in a future preset time period by using the trained model to obtain a node load prediction result; extracting a service level protocol parameter, a resource demand type and historical execution efficiency data of a to-be-scheduled task, and establishing a multi-dimensional resource priority evaluation index system; according to the computing power scheduling method, the problems of low resource utilization rate and high task response delay caused by low load prediction precision and mismatching of resource allocation and task priority in a traditional computing power scheduling method are solved, and the overall operation efficiency and service quality of a computing power cluster are improved.
Owner:SHAOGUAN DATA IND RESEARCH INSTITUTE

Multi-agent cooperative task processing method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to service scenes such as pension service, financial science and technology and medical health, and discloses a multi-agent cooperative task processing method, device and equipment and a medium, and the method comprises the steps: obtaining a task instruction, analyzing a core target, and decomposing the core target into a plurality of subtasks; obtaining environment information, dividing task areas, and generating a cooperation framework in combination with agent capability and area weight; real-time states of the agents are obtained, and the optimal agents are matched based on the cooperation framework to generate a task allocation table; a task distribution table is issued to control the intelligent agent to execute the task and upload execution information; monitoring an execution process, and performing dynamic adjustment and updating a task allocation table when detecting path conflicts or equipment faults; and after the subtask is completed, obtaining environment completion state data, and comparing the data with a preset standard model for acceptance. According to the method, efficient task decomposition and intelligent distribution are realized by fusing task semantics, environment information and intelligent agent capability, and the cooperation stability is improved by introducing a real-time state perception and self-adaptive mechanism.
Owner:平安科技(上海)有限公司

Task complexity driven graph semantic multi-agent collaborative decision-making method and system

The invention belongs to the field of natural language processing, and provides a task complexity driven graph semantic multi-agent collaborative decision-making method and system.The task complexity driven graph semantic multi-agent collaborative decision-making method comprises the steps that a task text is obtained and subjected to semantic coding to obtain a task semantic vector, evaluation is conducted based on the task semantic vector to obtain a complexity vector, and a task complexity score of the complexity vector is calculated; the task semantic vector and the complexity vector are fused to obtain a task representation vector, an agent capability relation graph is constructed, the participation probability of each agent node is obtained according to the task representation vector and the agent capability relation graph, and a dynamic agent combination scheme is formed; and performing task decomposition according to the agent combination scheme, constructing a sub-task dependency graph, scheduling the execution sequence of the sub-tasks through topological sorting, realizing cooperative execution of the agents, and generating a task result. According to the method, precise matching and efficient cooperation of the agent combination are realized, and the capability of processing complex tasks and the resource utilization efficiency of the multi-agent system are remarkably improved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +3

Artificial intelligence (AI) agents orchestration

An AI orchestration system dynamically manages multiple artificial intelligence (AI) agents within a cloud computing environment to efficiently process user requests. A model orchestration subsystem determines whether a request is handled locally using a domain-specific database or by invoking one or more AI agents. The system maintains AI agents in active and inactive states, provisioning computing resources for inactive agents as needed. Real-time model metrics guide the selection of target AI agents, and if a degrading performance trend is detected, the system preemptively spins up additional AI instances. The system provisions processor cycles, memory, and network bandwidth through a cloud-based resource manager, instantiates containerized execution environments or virtual machines, and performs automated load balancing among AI instances.
Owner:PROACTIVE AI LAB INC

Cloud edge cooperative computing framework for multi-modal data stream fusion processing and processing method

The invention relates to a cloud edge cooperative computing framework and processing method for multi-modal data stream fusion processing, and the method comprises the following steps: S1, carrying out the noise suppression based on an original data stream collected by an edge computing node through employing an improved Wiener filtering algorithm, achieving the signal denoising through the adaptive threshold wavelet transformation, and obtaining a cloud edge data stream; and a timestamp alignment technology is utilized to solve the problem of time delay difference of multi-modal data, and a space-time alignment purified data stream is generated. Through combination of the improved Wiener filtering algorithm and the adaptive threshold wavelet transform, the noise suppression efficiency of the original data stream is significantly improved, the timestamp alignment technology effectively solves the time delay difference of the multi-modal data, the generation of the space-time alignment purified data stream ensures that the subsequent processing has a unified time sequence benchmark, and the efficiency of noise suppression of the original data stream is improved. The space-time attention fusion network adopts a collaborative architecture effect of a bidirectional gating circulation unit and a lightweight 3D convolutional network.
Owner:NANJING NANDA SIWEI TECHNOLOGY DEVELOPMENT CO LTD

Cloud computing resource optimization method based on intelligent scheduling

The invention discloses a cloud computing resource optimization method based on intelligent scheduling, and belongs to the technical field of cloud computing resource processing. The method comprises the steps of obtaining real-time operation data of target data in a data optimization detection range, collecting historical resource scheduling records and task execution logs, and constructing a multi-dimensional resource state data set; according to the method, multi-objective optimization, simulation verification and reinforcement learning feedback in the step S5 are carried out, a perception-prediction-scheduling-monitoring-optimization closed-loop mechanism is constructed, the resource utilization rate, the response time and the energy consumption cost of a multi-objective optimization function are balanced, and a particle swarm optimization algorithm is combined with simulation verification to generate a global optimal strategy; and reinforcement learning dynamically adjusts model parameters by taking the execution deviation as a reward signal, continuously updates a resource perception dimension and a prediction model, realizes continuous iterative upgrade of a resource optimization effect, and performs optimization processing on cloud computing resource optimization based on intelligent scheduling.
Owner:ZHONGHUI YIGUAN (JIANGSU) CLOUD COMPUTING TECHNOLOGY CO LTD

Agricultural information management system and method based on big data platform

The invention relates to the technical field of agricultural information management, and particularly discloses an agricultural information management system and method based on a big data platform, and the method comprises the steps: firstly deploying a multi-source data collection module at an edge calculation node, and obtaining and standardizing the soil moisture content, meteorological environment and equipment operation data in real time; secondly, constructing a local dynamic irrigation strategy model, and realizing multi-objective optimization through a reinforcement learning algorithm; establishing a federated learning framework at the cloud, dynamically distributing node weights by adopting an attention mechanism, and realizing model aggregation of privacy protection in combination with secure multi-party computing; an optimal irrigation instruction is generated through a multi-source data fusion engine, and a three-level response exception handling mechanism is established; and finally, a closed-loop feedback system containing short-term incremental learning and long-term architecture optimization is formed. The corresponding management system comprises six functional modules, namely a data acquisition module, a local modeling module, a federated learning module, a real-time decision-making module, an abnormal monitoring module and a closed-loop optimization module.
Owner:BEIJING XINGHENG TECH CO LTD

Multi-agent scheduling method and system based on task intention matching

The invention relates to a multi-agent scheduling method and system based on task intention matching. The method comprises the following steps that S1, tasks are received and analyzed; s2, intelligent agent matching candidates are obtained; s3, selecting an optimal agent by a scoring mechanism; s4, task assignment, execution monitoring and result acquisition; s5, performing reflection evaluation and strategy updating: dynamically updating an intelligent agent score, and adjusting a task label matching rule or other scheduling parameters so as to optimize a distribution decision of a subsequent task; and S6, result combination and output: when the task comprises a plurality of sub-tasks executed by a plurality of intelligent agents, the results are verified, sorted and combined. Compared with the prior art, the method has the advantages that the task target can be dynamically analyzed, the optimal agent can be intelligently matched to execute the task, and the scheduling strategy is continuously optimized through feedback after the task is completed, so that the task execution efficiency and effect of the multi-agent system are remarkably improved, and the accuracy, the adaptability and the intelligent level of task allocation are improved.
Owner:CHINA NUCLEAR EQUIP TECH RES (SHANGHAI) CO LTD

Real-time computing system resource coordination and decision engine system, method and equipment based on large language model

The invention discloses a real-time computing system resource coordination and decision engine system based on a large language model. According to the system, a large language model is innovatively used as a central strategic decision engine, and a'decision-coordination-execution 'three-layer architecture is constructed. According to the system, macroscopic strategy generation and microscopic real-time control are decoupled by introducing a hierarchical decision-making mechanism (a strategic layer, a tactical layer and an execution layer), so that the core contradiction between LLM high reasoning delay and the microsecond / millisecond-level real-time requirement of the system is effectively solved, and the method is suitable for local computing equipment and a cloud data center. The system comprises a predictive strategy preloading system, and transient response can be achieved. Meanwhile, the system adopts an asynchronous event-driven decision-making mechanism for continuous intelligent optimization. According to the method, the top-down, semantic understanding-based and global collaborative intelligent management of the computing resources is realized, and the resource utilization efficiency, the system automation degree and the overall energy efficiency in a complex and dynamic computing environment are remarkably improved.
Owner:SHENZHEN LANRUN TECH CO LTD

Structured decision-making method based on multi-agent collaborative decision-making and reinforcement learning

The invention discloses a structured decision-making method based on multi-agent collaborative decision-making and reinforcement learning, and relates to the technical field of natural language processing, knowledge engineering and agent collaboration, and the method comprises the steps: receiving an original rule document, analyzing the document type, complexity and constraint conditions, and defining a task target and a success standard; and according to the task target, matching and scheduling the intelligent agent from the registered intelligent agent library, and further analyzing the capacity configuration of the intelligent agent for standby. Through the multi-agent cooperation and reinforcement learning technology, full-process automation of rule documents from input to structured analysis is realized, document types, complexity evaluation and constraint condition analysis can be automatically identified, and a clear task target and a success standard are generated; and the large language model generates a structured workflow according to task requirements and agent capabilities, so that the performability is ensured through logic verification, manual intervention is greatly reduced, and the processing efficiency and the system intelligence degree are improved.
Owner:SHANGHAI XUEDA BIOMEDICAL TECHNOLOGY CO LTD

Distributed storage resource intelligent scheduling method and device

The invention provides a distributed storage resource intelligent scheduling method and device, and relates to the technical field of data processing, and the method comprises the steps: carrying out the feature splicing of feature vectors of different modes, so as to obtain a multi-mode feature vector; predicting the user satisfaction based on the multi-modal feature vector and the resource adjustment parameter; establishing a resource demand priority mapping table under different service scenes according to the user satisfaction; constructing a multi-objective optimization model according to the resource demand priority mapping table; on the basis of the multi-objective optimization model, predicting the resource state and the load condition of each node in the distributed storage system; and according to the resource state and the load condition of each node, lightweight rule scheduling is carried out to obtain a preliminary scheduling scheme. According to the invention, intelligent scheduling management of distributed storage resources is realized.
Owner:CCTV INT NETWORK CO LTD

Automated workflow creation

A computer-implemented method generates workflow definitions in workflow definition language using one or more large language models, LLMs. The method includes receiving a natural language description of an automated workflow and generating a plan generation prompt including the natural language description and plan generation instructions. The plan generation prompt is input to one of the LLMs and in response a structured plan comprising a plurality of actions are received. For each action, a corresponding segment of workflow definition language is generated to provide a plurality of segments of workflow definition language. The segments are combined to form a workflow definition corresponding to the natural language description.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Complex scene-oriented AI large model lightweight deployment method

The invention provides a complex scene-oriented AI large model lightweight deployment method, and relates to the technical field of edge computing, and the method comprises the steps: carrying out the structured pruning of a pre-trained Transform network based on the attention head importance score, carrying out the dynamic sparsification of the activation state of a feedforward network according to the input tensor entropy value, employing the dynamic mixing precision quantization, and carrying out the reconstruction of an AI large model. Obtaining network parameters after pruning quantization; deploying the pruned and quantized network parameters to an edge computing device, distributing a feature extraction operator to a neural network processor through a heterogeneous computing scheduler, and unloading a classification operator to a multi-core central processing unit; and managing an on-chip memory in combination with a virtual memory paging mechanism, realizing zero-copy data transmission by utilizing a direct memory access controller, and outputting a reasoning result tensor. According to the method, efficient and reliable operation of the large model at the resource-constrained edge node is realized.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD

Information scheduling scenarized service agent system

The invention relates to the technical field of information management, in particular to an information scheduling scenarized service agent system. The system comprises an information scheduling acquisition and analysis module, a scene classification and service matching module, a service scheme instruction generation module and a service intelligent execution module, user behavior data, environment perception data and business system data corresponding to information scheduling can be obtained, and an original information set is generated; the original information set is preprocessed, an information scheduling scene classifier is constructed, the service matching degree is calculated in combination with the real-time service resource state, and a candidate service scheme set is generated; performing optimization sorting on the candidate service scheme set to generate an optimal service scheme, and disassembling the optimal service scheme into an executable service instruction sequence; receiving a service instruction sequence and generating a service evaluation report; and if the display does not reach the preset service quality threshold value, triggering a dynamic adjustment mechanism to realize self-iteration upgrading of the information scheduling scenarized service. According to the invention, intelligent scheduling and service of information in a complex scene can be realized.
Owner:BEIJING SGITG ACCENTURE INFORMATION TECH CO LTD

Computing power resource multi-dimensional scheduling method and system based on dynamic weight

The invention relates to the technical field of computers, and discloses a computing power resource multi-dimensional scheduling method and system based on dynamic weight, and the method comprises a data perception step, a weight generation step, an intelligent decision-making step and a scheduling optimization step. The system corresponds to the method. The method comprises the following steps: a data sensing step: collecting multi-dimensional state parameters of computing power nodes and carrying out feature modeling to construct a global feature space; a weight generation step: dynamically adjusting the weight of each dimension based on a machine learning model and a rule engine; an intelligent decision-making step of screening candidate nodes in the global feature space and evaluating priorities, generating an optimal node cluster and performing resource dynamic slice distribution; and a scheduling optimization step: monitoring an execution effect and performing closed-loop feedback so as to iteratively optimize a weight strategy and decision logic. The problems that in the prior art, the sensing dimension is single, and decision-making weight is rigid are solved, and multi-dimensional accurate sensing, dynamic weight decision making and elastic resource allocation of computing power resources are achieved.
Owner:GLORYVIEW TECH INC

Allocating resources among autonomous artificial intelligence agents within a distributed computational network

Systems and methods disclosed herein automatically evaluate, select, and coordinate artificial intelligence (AI)-based agents for collaborative distributed task execution based on dynamic, multi-attribute scoring and resource allocation models. The system obtains a task specification request defining a computational requirement set, a performance metric set, and an available resource set for one or more tasks to be executed by a network of AI-based agents. A first AI model set generates domain-specific test datasets and validates prospective agents by comparing agent-generated fingerprints against predetermined hash values stored on a distributed or federated ledger. A second AI model set constructs a multi-dimensional scoring data structure for each agent by using historical performance metrics to compute weighted composite scores. The system selects a subset of AI-based agents, ranks the agents, and allocates resources proportional to each agent's composite score. A third AI model set coordinates and executes distributed computer-executable workflows across the selected agents.
Owner:CITIBANK N A

Computing power resource elastic allocation monitoring system

The invention relates to the technical field of computing power resources, and discloses a flexible allocation monitoring system for computing power resources, which defines a clear resource use boundary for different types of tasks through a container-level QoS strategy of a resource isolation unit, and adjusts resource allocation in real time in combination with a dynamic partition management module, thereby avoiding resource waste in a traditional fixed allocation mode, and improving the resource allocation efficiency. The dynamic preemption unit accurately screens low-priority tasks for resource recovery through a multi-factor decision engine and a preemption cost evaluation algorithm, in 50 concurrent task scenes, the resource preemption response time is shortened to 20-35 ms and is improved by 50%-70% compared with 70-100 ms of Docker / K8s, the blocking duration of high-priority tasks is reduced to 15-25 ms from 80-120 ms, and the core service interruption risk is greatly reduced.
Owner:HEBEI MINGWEI DIGITAL TECHNOLOGY CO LTD

Intelligent agent collaborative optimization data center management system based on knowledge graph driving

The invention discloses an agent collaborative optimization data center management system based on knowledge graph driving, and relates to the technical field of data center management. The system specifically comprises the following modules: a modeling and entity management module, a cross-regional global scheduling optimization module, an agent game negotiation optimization module, an agent trust management module, a game negotiation conflict identification module, a reasoning evidence management module and a cross-regional consistency verification module. By establishing a complete knowledge graph model, systematic modeling of resource attributes, constraints, historical decisions and strategy preferences is realized, a unified data basis is provided for negotiation among multiple agents, the problem of a suboptimal solution caused by information asymmetry is effectively solved, a hierarchical reasoning mechanism is adopted, and the probability of resource disruption is reduced. In combination with global-region-node three-level reasoning and multi-round game negotiation, recursive optimization from global to local is realized, the reasoning complexity is effectively reduced, and the negotiation efficiency is improved.
Owner:北京紫翰科技有限公司

Information integration management method and system supporting real-time updating

PendingCN121326983ADatabase updatingResource allocationService adaptationData access
The invention belongs to the technical field of computers, and particularly relates to an information integration management method and system supporting real-time updating, and the method comprises the steps: accessing a heterogeneous data source through a distributed gateway, carrying out the adaptive analysis of a protocol, and standardizing a structure; adding a millisecond timestamp and a unique identifier, and generating a semantic data cluster through streaming engine sliding window association analysis; after multi-dimensional consistency verification is executed, writing and storing are carried out by adopting a double-buffering and snapshot isolation mechanism; and dynamically assembling the customized view according to the user role and the context and safely returning the customized view. The system correspondingly comprises a protocol analysis module, a time sequence marking module, a streaming association module, a semantic verification module, a double-buffer control module, an index updating module, a view assembling module and a secure transmission module. According to the invention, millisecond-level real-time convergence, conflict-free fusion and on-demand view output are realized, and the data access flexibility, the quality guarantee capability and the service adaptation flexibility are comprehensively improved.
Owner:HARBIN HONGCHENG NETWORK TECHNOLOGY CO LTD

Task-aware migration-based dynamic allocation method for cloud edge-end cooperative computing resources

The invention relates to the technical field of cloud side end computing, and discloses a cloud side end cooperative computing resource dynamic allocation method based on task-aware migration. The method comprises the following steps: acquiring real-time load characteristics and resource demand characteristics of calculation tasks in a cloud side end system, and dividing task priority queues in combination with task type identifiers; extracting historical execution records of tasks at cloud, edges and terminal nodes, constructing a task execution feature library, and generating a resource demand prediction model in combination with real-time load features; analyzing network transmission time delay characteristics of a cloud end and edge nodes, measuring real-time calculation capability fluctuation data of terminal equipment, and establishing an inter-node resource collaboration degree evaluation matrix; generating an initial migration strategy according to the prediction model and the evaluation matrix, monitoring actual resource occupancy deviation of the task, forming a final decision in combination with a node resource state correction strategy, triggering cross-node migration, and synchronously updating the priority queue and the evaluation matrix.
Owner:ZHONGKE SUANWANG TECH CO LTD

Resource scheduling control method and system for big data server

The invention provides a resource scheduling control method and system for a big data server, and the method comprises the steps: constructing a multi-dimensional resource portrait module, collecting the CPU, memory, network, storage I / O load and task queue length of each node in real time, and predicting a resource demand trend through a time sequence algorithm; extracting characteristics such as calculation intensity, data dependence, memory requirements, network transmission quantity and the like; adjusting the weight coefficients of the resource utilization rate, the task completion time and the energy consumption efficiency according to the system load and the historical effect; establishing a bipartite graph model by taking a resource trend as a node feature and a task vector as an edge feature, and calculating a matching score through graph convolution and a multi-objective optimization function; the scheduling scheme is synchronized by adopting a consistency algorithm; automatic rollback and reallocation are carried out when resources are detected to be insufficient; and optimizing a weight coefficient and a network parameter through reinforcement learning. Through the method, the system resource utilization rate can be improved, the task execution efficiency is improved, the overall scheduling effect stability is improved, and the system fault recovery time is shortened.
Owner:SHANGHAI HONGXING INFORMATION TECH CO LTD

Dynamic, topology-aware load balancing in a distributed ML inference system

ActiveDE202025104872U1Resource allocationTransmission
A dynamic, topology-aware load balancing in a distributed ML inference system (100), consisting of: a topology and monitoring module configured to detect, map, and dynamically update the network topology by capturing real-time latency, bandwidth, and connectivity metrics across distributed nodes; a module for capturing node performance, configured to evaluate and record the computing power, hardware configuration, accelerator availability, and storage capacity of each node; a workload characterization module configured to analyze incoming inference requests with respect to computational complexity, latency sensitivity, memory requirements, and model type; a dynamic load balancing engine configured to assign inference tasks to the optimal nodes based on topology data, node capabilities, and workload characteristics; an adaptive feedback and optimization module configured to capture performance metrics for task execution and refine task allocation strategies based on historical results; a fault tolerance and recovery module configured to detect errors, redirect tasks, and maintain uninterrupted service; and an orchestration and integration interface module configured to provide APIs and orchestration hooks for seamless deployment in heterogeneous computing environments; The system dynamically adjusts the load distribution to respond to changes in network conditions, fluctuations in resource availability, and workload requirements to improve efficiency, minimize latency, and maximize throughput in distributed ML inference environments.
Owner:SINGH RAMKINKER CAMPBELL

System and method for cost-aware autoscaling of artificial intelligence workloads using predictive queuing models

The present invention relates to a system and computer implemented method for cost-aware autoscaling of artificial intelligence workloads using predictive queueing models, designed to achieve proactive and economically optimized scaling of computational resources across cloud and edge environments. The invention introduces a predictive queueing-based technique that anticipates future workload congestion by modeling dynamic task arrivals and service times using a stochastic queueing process. A cost estimation unit computes the total projected operational cost of potential scaling actions by integrating real-time infrastructure pricing data, predicted delay penalties derived from service-level objectives, and estimated energy consumption. A scaling decision unit applies reinforcement learning-based optimization to select the scaling action that minimizes total cost while ensuring compliance with latency and throughput constraints. The system includes a hardware-integrated autoscaling controller device comprising a predictive computation processor, cost-decision processor, and scaling actuation interface configured for real-time execution of predictive and scaling operations.
Owner:MIRZA MAHAMOOD HUSSAIN +3

Multi-modal big language model reasoning optimization method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical health, and discloses a reasoning optimization method, device, equipment and medium for a multi-modal large language model.The method comprises the steps that an input long context sequence is obtained, and key value projection is conducted on the long context sequence to generate an initial key value cache; for each attention layer of the multi-modal large language model, calculating an attention matrix of the attention layer according to the vector dimension of the long context sequence and the initial key value cache; calculating a cross-modal attention entropy according to the attention matrix, and determining a cache size of an attention layer according to the cross-modal attention entropy; optimizing the initial key value cache based on a cumulative attention scoring mechanism and a window strategy to obtain a target key value cache; and reasoning the long context sequence according to the cache size and the target key value cache to generate a long context reasoning result. And the reasoning efficiency and the reasoning accuracy are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method for compatible operation of Android camera HAL in container based on memory access virtualization

The invention discloses a compatible operation method for an Android camera HAL in a container based on memory access virtualization, and the method comprises the steps: taking a DMA-Buf memory heap of a Linux kernel host system as a target memory heap, taking an ION memory heap of an Android container system as a source memory heap, building a process exclusive memory management context, an FD cache table and a synchronous fence pool, and carrying out the execution of the process exclusive memory management context, the FD cache table and the synchronous fence pool; kernel registration and node binding of virtual ION equipment, pre-allocation of a first memory pool and access hook registration of kernel layer equipment are completed, and when an ION file descriptor is obtained in an HAL process, legality of a container process is verified, and context binding is initialized to the file descriptor; intercepting a memory allocation request of the HAL process, analyzing and adapting parameters, and preferentially multiplexing first memory pool resources to obtain an ION handle; when a data sharing request is processed, matching an FD cache or generating a new FD through an ION handle; when the HAL process releases the memory, resources are recycled according to a memory source, and when the HAL process exits, the context is cached or destroyed, so that cross-system memory operation compatibility is realized.
Owner:北京麟卓信息科技有限公司

Distributed computing power scheduling method and system

The invention relates to the technical field of computing power scheduling, provides a distributed computing power scheduling method and system, breaks through the limitation that a traditional computing power scheduling system depends on a single index or a static strategy by constructing an integrated architecture of data perception, load prediction and intelligent scheduling, and constructs a self-adaptive and closed-loop optimized intelligent scheduling system. The scheduling system obtains multi-source data in real time and fuses the multi-source data into a perception vector to accurately describe the running state of the system, accurate pre-judgment of load changes is achieved through hierarchical load prediction, and resources can be efficiently allocated in different scenes in combination with double scheduling paths and execution feedback optimization; according to the scheduling system, the resource utilization rate of the distributed system is effectively improved, the energy consumption cost is reduced, the task execution delay is reduced, the system stability and the fault-tolerant capability are enhanced, and an innovative computing power scheduling solution is provided for a large-scale distributed computing scene.
Owner:GUIYANG YIYI TECH CO LTD