Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

88 results about "Distributed reasoning" patented technology

Multi-modal agent RAG-ReAct double-engine cooperative training method

The invention relates to the technical field of artificial intelligence, in particular to a multi-mode agent RAG-ReAct double-engine cooperative training method. The method comprises the following steps: acquiring multi-modal data, and converting the multi-modal data into a high-dimensional vector; constructing a knowledge graph based on the high-dimensional vector; encoding the semantic relationship of the knowledge graph into a model fine adjustment gradient direction by using a dynamic distillation technology; designing a distributed architecture based on the high-dimensional vector; performing hybrid retrieval based on a distributed architecture to obtain hybrid retrieval data; executing distributed reasoning based on a preset recursive reflection mechanism and the mixed retrieval data to generate distributed reasoning data; performing data parallel detection according to the distributed reasoning data to obtain data parallel parameters; and performing node resource scheduling based on the data parallel parameters so as to obtain node load balancing data. Based on the artificial intelligence technology, the reasoning accuracy and the resource utilization rate of the multi-modal agent in a complex task environment are effectively improved.
Owner:BEIJING ZHONGSHURUIZHI TECH CO LTD

Distributed reasoning industrial Internet of Things cloud edge collaboration method

The invention relates to an industrial Internet of Things cloud edge collaboration method based on distributed reasoning, and the method comprises the steps: segmenting a deep learning reasoning model into a plurality of sub-models according to a reasoning task, and packaging each sub-model into a lightweight WASM model through format conversion, operator detection and structured pruning in combination with containerized deployment and WasmEdge operation, the memory and CPU occupation of the edge node is obviously reduced, the loading and reasoning time delay is shortened, and the network and computing resource utilization rate is improved; and then, distributing each Docker container mirror image to each edge node according to a resource distribution strategy to perform distributed reasoning, thereby realizing multi-node load balancing and reliable scheduling, and improving the throughput and the overall reasoning rate of the industrial Internet of Things.
Owner:NINGBO UNIV

Distributed reasoning method and device based on hybrid expert architecture, equipment and medium

The invention relates to the field of artificial intelligence, and provides a distributed reasoning method, device and equipment based on a hybrid expert architecture, and a medium, and the distributed reasoning method based on the hybrid expert architecture applied to an edge computing power gateway comprises the steps: determining a plurality of target intelligent vehicles participating in a reasoning task if the reasoning task is received; obtaining an expert reasoning model, and deploying the expert reasoning model to each target intelligent vehicle; wherein the expert reasoning model comprises a plurality of expert reasoning modules, and at least one expert reasoning module is deployed in each target intelligent vehicle; sending the reasoning task to each target intelligent vehicle, so that each target intelligent vehicle executes the reasoning task based on the expert reasoning module to obtain a reasoning result; and aggregating the reasoning results of the target intelligent vehicles to obtain a target reasoning result. Through the technical scheme provided by the invention, the computing power resource of the idle intelligent vehicle is fully utilized.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Distributed reasoning task allocation method for edge computing large model

The invention provides a distributed reasoning task allocation method for an edge computing large model, which belongs to the technical field of mobile edge computing and comprises the following steps of: considering a large language model distributed deployment scene, constructing a workflow for cooperatively finishing a reasoning task by a plurality of edge servers, and adopting accurate secondary reconstruction by considering the generation type characteristic of the reasoning task; the uncertainty of reasoning task resource occupancy is solved, mathematical model expression is constructed for the edge reasoning task allocation problem in the scene, the aim is to maximize the service provider income, the constructed mathematical model is reconstructed into a combination selection problem, an approximate solution with theoretical guarantee is given by adopting a dual method, and the problem of edge reasoning task allocation is solved. And the reasoning task distribution decision of the service provider is optimized, so that the income of the service provider is maximized.
Owner:NANJING UNIV OF POSTS & TELECOMM

Large model distributed reasoning acceleration method based on multi-modal feature fusion and dynamic weight optimization

The invention relates to the technical field of large models, in particular to a large model distributed reasoning acceleration method based on multi-modal feature fusion and dynamic weight optimization, and the method comprises the following steps: S1, carrying out real-time semantic analysis through a reasoning context analysis module embedded in a load balancer to obtain semantic features; s2, calculating a cache adaptation degree according to the semantic features; s3, querying a global cache directory service to obtain a matched node list, and making a decision by using a multi-fusion decision algorithm in combination with various parameters of cache adaptation condition, node load condition and network quality; and S4, deploying a global cache directory service in the load balancing layer according to the decision, maintaining a local cache pool at each computing node, and synchronizing the cache metadata to the global cache directory service in real time. The method solves the problems that when a conventional load balancing strategy is adopted for multi-node deployment, waste of computing resources is easily caused, the overall energy consumption of the system is increased, and the request processing speed is reduced.
Owner:CHINA UNICOM (GUANGDONG) IND INTERNET CO LTD

Neural network pipeline distributed reasoning acceleration method for multi-NPU heterogeneous platform

The invention discloses a neural network pipeline distributed reasoning acceleration method for a multi-NPU heterogeneous platform, and the method comprises the steps: dividing a neural network model into a plurality of sub-models according to the theoretical calculation time of each operator in the neural network model needing distributed reasoning; generating a first splitting strategy and a second splitting strategy according to the plurality of sub-models; establishing an NPU execution time delay model, and calculating the load when the first splitting strategy and the second splitting strategy are executed on the NPU; according to the load, selecting a better one from the first splitting strategy and the second splitting strategy as a current splitting strategy, and performing fine adjustment on the current splitting strategy through iteration; and respectively deploying each sub-model in the obtained optimal splitting strategy to each NPU, and scheduling by the main control CPU to carry out pipeline distributed reasoning of the input image. According to the invention, a pipeline parallel architecture is formed among multiple NPUs, so that each NPU performs reasoning on different input images in different time units, and parallel computing of multiple NPU cores is realized.
Owner:XIDIAN UNIV

Inference method and apparatus for large language model

An inference method and apparatus for a large language model, which method and apparatus are used for improving the inference calculation efficiency of a large language model. The method comprises: receiving an inference task, which carries a morpheme number, wherein the morpheme number is used for identifying a morpheme to be subjected to inference calculation, the morpheme comprising one or more sub-words determined on the basis of input text of a large language model (201); on the basis of the morpheme number corresponding to the inference task, querying a global index tree to determine a target inference node, wherein the global index tree comprises sub-trees corresponding to a plurality of inference nodes in a distributed inference cluster, and a sub-tree corresponding to the target inference node is a sub-tree in the global index tree, the matching degree of similarity of which sub-tree with the morpheme number is greater than a threshold, the matching degree of similarity indicating the number of pieces of reusable key-value cache data during inference calculation (202); and on the basis of the target inference node, executing the inference task (203).
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Istio-based inference service global activation connection number limiting method

The invention relates to the technical field of micro-services, and discloses an Istio-based inference service global activation connection number limiting method, which comprises the following steps: step 1, when a client request is received, extracting label information in a request header, and carrying out label analysis; 2, after label analysis is completed, threshold value query is conducted, and interface response content is obtained; 3, acquiring real-time load data of the target service, and calculating a real-time dynamic threshold value; 4, performing three-level quota verification from the tenant level, the service level and the instance level; and 5, after three-level quota verification, connection processing is carried out, a connection processing result is reported, and closed-loop feedback is carried out. According to the scheme, a global and intelligent connection number limiting mechanism is constructed in a micro-service architecture, the problems of service overload, resource competition and system stability caused by sudden increase of the connection number in distributed reasoning service are solved, and dynamic sensing, intelligent decision making and accurate control of the reasoning service connection number are achieved.
Owner:ASPIRE TECH (SHENZHEN) LTD

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Big language model reasoning method and device

The embodiment of the invention discloses an inference method and device for a large language model, which are used for improving the inference calculation efficiency of the large language model. The method comprises the steps that an inference task is received, the inference task carries morpheme numbers, the morpheme numbers are used for identifying morphemes to be subjected to inference calculation, and each morpheme comprises one or more sub-words determined based on an input text of a large language model; a global index tree is inquired based on morpheme numbers corresponding to the reasoning tasks, target reasoning nodes are determined, the global index tree comprises sub-trees corresponding to multiple reasoning nodes in a distributed reasoning cluster, and the sub-trees corresponding to the target reasoning nodes are sub-trees with the similarity matching degree with the morpheme numbers larger than a threshold value in the global index tree; the similarity matching degree indicates the number of reusable key value cache data in reasoning calculation. And executing the reasoning task based on the target reasoning node.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Large model distributed reasoning acceleration method and device based on speculation sampling

The embodiment of the invention provides a large model distributed reasoning acceleration method and device based on speculation sampling, and relates to the field of artificial intelligence. And inputting the terminal prefix information into a draft small model for content generation to obtain a plurality of candidate tokens and corresponding small model probability distributions, and transmitting the parameter sequence, each candidate token and a small model probability value corresponding to the candidate token to a base station, so that the base station generates a large model probability distribution of each candidate token by using a target large model, and the candidate tokens are divided into accepted tokens and / or rejected tokens. And if the terminal receives the large model probability distribution corresponding to the first rejection token, resampling the first rejection token to obtain a resampling token, writing the resampling token into the parameter sequence, updating terminal prefix information according to the resampling token, and performing iterative execution until a reasoning result is obtained. The rejection / receiving process of the candidate tokens is executed on the base station, and the resampling process after rejection is executed on the terminal equipment, so that the cooperative reasoning efficiency is improved.
Owner:PENG CHENG LAB

Multi-modal model distributed reasoning optimization method for dynamic intelligent Internet of Things

The embodiment of the invention provides a multi-modal model distributed reasoning optimization method for a dynamic intelligent internet of things. The method comprises the following steps: constructing a multi-modal model distributed reasoning optimization system for the dynamic intelligent internet of things; the system comprises a plurality of Internet of Things devices for generating and uploading DNN reasoning tasks and a plurality of servers for receiving and processing the DNN reasoning tasks, determining a DNN reasoning task data parallel scheduling strategy for the system; determining system reasoning time and system energy consumption by using a DNN reasoning task data parallel scheduling strategy; according to the system reasoning time and the system energy consumption, establishing a minimum long-term time delay under the constraint of long-term energy consumption, and taking the minimum long-term time delay as an optimization target to establish a long-term distributed reasoning stability optimization model; and solving the model by adopting a heuristic algorithm to realize distributed reasoning optimization of the multi-modal model. In this way, distributed reasoning tasks of the multi-modal model can be effectively managed, and reasoning time and energy consumption can be optimized.
Owner:TONGJI UNIV

Cloud edge collaborative diffusion model reasoning method and system based on block chain and reinforcement learning

The invention provides a cloud edge collaborative diffusion model reasoning method and system based on a block chain and reinforcement learning, and the method comprises the steps: obtaining a current text prompt word submitted by a user in a block chain network composed of a cloud server and an edge server, and calling a semantic matching model through an intelligent contract, and judging whether a historical intermediate result can be reused or not; obtaining a server environment state, generating a collaborative reasoning strategy by using a pre-trained multi-agent attention actor-commentator model, and determining cloud edge denoising step number distribution; if the image cannot be reused, the cloud server executes partial denoising to generate an intermediate result and stores the intermediate result to the block chain, and the edge server continues to complete residual denoising to generate a final image; and if the edge server can be reused, the edge server directly completes denoising based on the historical intermediate result. According to the method, historical intermediate results can be effectively reused, dynamic intelligent task allocation is realized, the diffusion model reasoning efficiency and the image generation quality are improved, and the credibility and the traceability of a distributed reasoning process are guaranteed.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Methods and apparatus for support of vertical federated learning at network data analytics function (NWDAF)

Apparatus and methods to perform vertical federated learning (VFL) in a network, e.g. 5G New Radio network. At step 13, a VFL server, such as a Network Data Analytics Function (NWDAF) hosting VFL functionality, receives a message from a service consumer such as a consumer network function (NF). The message may be a subscription to analytics or a request for analytics. At step 14, a distributed VFL inference process is triggered by the VFL server invoking a Machine Learning (ML) model inference request. The request is transmitted to a participant, i.e. a VFL client entity which may also be a NWDAF. At step 15, based on the request, the VFL client entity obtains intermediate inference results by performing local inference computation using a local ML model. At step 16, if the VFL client entity is the last participant, the intermediate inference results are transmitted to the VFL server. At step 17, the VFL server performs further inference computation on the received intermediate inference results and derives the requested analytics which may be delivered to the NF consumer at step 18. The VFL server may discover or select the VFL client entities to participate in the VFL procedure.
Owner:SAMSUNG ELECTRONICS CO LTD

Distributed large model reasoning optimization method and device based on dynamic micro-batch scheduling

The invention discloses a distributed large model reasoning optimization method and device based on dynamic micro-batch scheduling, and the method comprises the steps: (1) initializing a system, and building a large model distributed reasoning assembly line; (2) evaluating the computing power and the network state of each node and summarizing the computing power and the network state to a head node; (3) determining the number of Micro-Batches and the scheduling quota of each Micro-Batch according to the request distribution condition, the computing power of each node and the network state; and (4) sequentially scheduling the Micro-Batch by adopting a Control Batch strategy and a Chunked Prefill strategy, and starting to execute the Micro-Batch. According to the method, by dynamically adjusting the number of Micro-Batch, the serious assembly line cavitation problem in a distributed large model reasoning system is effectively solved, the GPU utilization rate and the system throughput are remarkably improved, and meanwhile key indexes such as TTFT (first token time delay) and TPOT (inter-token time delay) in the large model reasoning field are also improved. The method has good adaptability, can adaptively adjust the dynamic scheduling strategy under different hardware equipment, network conditions and request loads, is suitable for different distributed large model deployment scenes, and has wide application value.
Owner:ZHEJIANG UNIV

Cross-GPU video memory dynamic scheduling system and method based on equipment demand perception

The invention discloses a cross-GPU video memory dynamic scheduling system and method based on equipment demand perception, belongs to the field of artificial intelligence of computer science, and is suitable for a distributed reasoning scene of a large language model. The method comprises the following steps: aiming at the problems of unbalanced utilization of video memory resources and cross-device communication bottleneck in a PD separation architecture, a system constructs a video memory exchange area in a pre-filling instance, when the video memory occupation of a decoding instance is critical, part of KV cache is migrated to the exchange area and returns when the resources are allowed, so that the idle video memory resources of the pre-filling instance are fully utilized, and the resource utilization rate of the decoding instance is improved. And the video memory pressure in the decoding stage is relieved. The system realizes the conversion from the low-efficiency KV cache migration process transferred by the original CPU to the GPU direct access mode through a high-speed communication path established by direct connection between the GPUs. A layered asynchronous transmission strategy is adopted in the migration process, communication overhead is hidden in a calculation assembly line, migration delay is reduced, and cross-GPU video memory sharing multiplexing and efficient scheduling are achieved.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1

Orchestration of distributed inference operations

Approaches presented herein provide for the management of resources to be used to process a request, such as may involve orchestration of nodes for an inference request. Upon receiving an inference request, an orchestrator can determine a sequence of context nodes and inference nodes to be used to process the inference request, based in part upon a determined class of inferencing to be performed. The orchestrator can append metadata to the inference request that identifies the sequence, and can transmit the appended request to one or more first nodes in the sequence. If the nodes have a network programmable device, or similar capability, the request can be forwarded to the nodes in sequence without having to go back to the orchestrator between nodes.
Owner:NVIDIA CORP

Ensemble communication unloading method, system, equipment and medium

ActiveCN121979690AResource allocationInference methodsCollective communicationComputer network
The invention discloses a set communication unloading method, system and device and a medium, and is applied to the technical field of computers, and the method comprises the steps: in a model deployment stage, a distributed reasoning controller generates a communication primitive blueprint based on a computational graph description file of a tensor parallel reasoning model and issues the communication primitive blueprint to a DPU; the DPU establishes a hardware-level communication context semantic environment based on the communication primitive blueprint; in the model reasoning stage, the GPU / NPU sends a trigger signal to the DPU when calculating to a communication boundary; and the DPU executes a DMA data pulling assembly line and an RDMA data sending assembly line in parallel based on a hardware-level communication context semantic environment, performs aggregation calculation on all tensor fragment data to be synchronized, and writes an aggregation calculation result back to the GPU / NPU. A hardware-level communication context semantic environment is established in advance to a model deployment stage, and double assembly lines are executed in the DPU in parallel, so that end-to-end communication delay is remarkably reduced, and zero participation of a host CPU is realized.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD +1

A distributed inference communication method and system of a super-large-scale graph neural network

PendingCN122450656AGlobal topologyIndependent set
A distributed inference communication method of a super-large-scale graph neural network comprises the following steps: dividing a global topology to be processed into a plurality of local sub-domains matched with the number of GPUs, and determining a local master node and a boundary master node allocated to each GPU; calculating coloring priorities of nodes, and performing parallel graph coloring based on the coloring priorities to construct a global independent set sequence; mapping the global independent set sequence to the local sub-domains of the GPUs to generate a local task slice sequence, which comprises boundary node slices and internal node slices, evaluating communication load and calculation load of the slices, and taking the internal node slices as calculation fillers to be executed in the communication waiting period of the boundary node slices; and each GPU executes an inference task by using a multi-stream asynchronous pipeline constructed locally according to the local task slice sequence. The method can improve the throughput of distributed inference.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Distributed cluster construction method, distributed reasoning method and resource scheduler

The invention provides a distributed cluster construction method, a distributed reasoning method and a resource scheduler, which can be applied to the technical field of computers. The distributed cluster construction method comprises: in response to a submission operation for a system arrangement file, binding a virtual node created based on the system arrangement file to a physical node equipped with a hardware accelerator, the system arrangement file indicating an expected construction state for a distributed cluster; calling a plurality of physical nodes, pulling a target reasoning mirror image from a mirror image warehouse, and starting respective container instances of the plurality of physical nodes by using the target reasoning mirror image, the target reasoning mirror image being obtained by packaging a development kit of a hardware accelerator of the physical node and a distributed reasoning framework; calling an accelerator plug-in of the physical node, and allocating resources of a hardware accelerator on the physical node to the container instance; and constructing a distributed cluster according to the plurality of container instances.
Owner:BEIJING KECHENG TECH DEV CO LTD

Task reasoning method and device and electronic equipment

The embodiment of the invention provides a task reasoning method and device and electronic equipment. The method is applied to a distributed reasoning architecture, the distributed reasoning architecture comprises a first node, a second node and a shared memory pool, and the method comprises the steps that the first node obtains at least one first to-be-reasoned task; the first node retrieves the first to-be-reasoned task to obtain at least one first retrieval result related to the first to-be-reasoned task; the first node stores the at least one first retrieval result in a shared memory pool; the second node reads at least one first retrieval result from the shared memory pool; and the second node takes the at least one first to-be-reasoned task and the at least one first retrieval result corresponding to the first to-be-reasoned task as input of a model, and executes the first to-be-reasoned task by utilizing the model. By adopting the mode, the delay problem caused by cross-node transmission can be avoided, so that the overall efficiency of the reasoning task is improved.
Owner:XFUSION DIGITAL TECH CO LTD

Model distributed reasoning test system based on multiple machines and multiple cards

The invention provides a model distributed reasoning test system based on multiple machines and multiple cards. The system comprises a model reasoning test middle table, a server matching module, a test environment construction module, a reasoning test module, an index evaluation module and a test efficiency evaluation module. Obtaining a target multi-card server; constructing a multi-machine multi-card environment; starting a distributed reasoning test in a multi-machine multi-card environment to obtain model reasoning indexes of each model test strategy in different reasoning test stages; performing evaluation according to the model reasoning indexes of each model test strategy in different reasoning test stages to obtain a calculation speed-up ratio curve of each model test strategy in different reasoning test stages; and performing test efficiency evaluation according to the calculation speed-up ratio curve of each model test strategy in different reasoning test stages to obtain an optimal test strategy. According to the invention, the test efficiency of the model in the distributed environment is improved.
Owner:GUANGDONG POWER GRID CO LTD +1

A heterogeneous cluster distributed inference collaboration method and system based on VLLM

The application discloses a VLLM-based heterogeneous cluster distributed reasoning cooperation method and system, and the method comprises the following steps: S1, generating a heterogeneous device performance image signal through a performance probe arranged in each computing unit in the cluster; S2, performing non-uniform task division according to the heterogeneous device performance image signal and combining a calculation graph structure of a to-be-reasoned model to form an optimal task scheduling strategy signal for a heterogeneous environment; S3, generating a task allocation and loading instruction signal based on the optimal task scheduling strategy signal; S4, coordinating each heterogeneous computing unit to perform reasoning calculation in a distributed reasoning execution process to generate a distributed cooperative reasoning signal; and S5, integrating and post-processing the distributed cooperative reasoning signal to output a final reasoning result. The VLLM-based heterogeneous cluster distributed reasoning cooperation method and system can solve the problems of low resource utilization and poor reasoning efficiency when a large model is run on a heterogeneous cluster.
Owner:ENTERPRISE ONLINE (BEIJING) NETWORK CO LTD

Online inference method, service system, and device based on an artificial intelligence geospatial data cube

The disclosure pertains to the field of data analysis and services, specifically to an online inference method, service system, and device based on an artificial intelligence geospatial data cube. The method involves constructing a cube organizational model based on a spatiotemporal grid, where the cube organizes and manages geospatial data and GeoAI models in a unified manner. It performs task-oriented model matching through a combination of explicit and implicit matching, where explicit matching converts user inference requests into multidimensional query conditions to retrieve candidate models, and implicit matching determines the optimal model by computing the feature similarity between inference data tiles and candidate models. Based on a distributed inference framework, an efficient inference workflow is executed in parallel across multiple computing nodes. The disclosure aims to address the limitations of existing technologies in multi-scenarios or large-scale spatial scenarios, such as low inference accuracy and slow inference speed.
Owner:WUHAN UNIV

Distributed reasoning task arrangement method based on semantic driving

The invention discloses a distributed reasoning task arrangement method based on semantic driving. The method comprises the following steps: constructing a domain ontology model and establishing semantic mapping; analyzing task requirements; executing data source retrieval of multi-level semantic fusion; generating an execution path based on multi-objective collaborative optimization; tasks are distributed, and transmission overhead is reduced by adopting model differential updating; and performing adaptive migration and network partition fault tolerance based on a check point and a two-stage commit protocol. According to the method, the technical problems of heterogeneous data semantic gaps, static scheduling and dynamic environment conflicts and state loss caused by coarse-grained fault tolerance are solved, efficient distributed reasoning of data immobility and calculation flow is achieved, and the method is suitable for the multi-source heterogeneous data fields of marine science, meteorological monitoring, industrial Internet of Things and the like.
Owner:SUZHOU ABYSS MATRIX TECHNOLOGY CO LTD

Inference request scheduling method and device, equipment, medium and distributed inference system

The invention relates to a reasoning request scheduling method, device and equipment, a medium and a distributed reasoning system, and relates to the technical field of intelligent calculation. The method comprises the following steps: in response to a reasoning request of a user side, if intermediate cache data corresponding to the reasoning request does not exist in local storage and storage nodes of a distributed shared storage pool, determining a corresponding reasoning service demand according to the reasoning request; determining the demand matching degree corresponding to each reasoning processing node combination according to the reasoning service demand and the communication link attribute between the nodes in each reasoning processing node combination; and according to the demand matching degree, scheduling the reasoning request to a pre-filling node and a decoding node contained in a target reasoning processing node combination in each reasoning processing node combination for reasoning processing. By adopting the method, the memory use efficiency can be improved, the memory overflow risk can be reduced, the high-concurrency reasoning service requirement can be met, and the performance bottleneck in a high-load scene can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Method and apparatus for scheduling for distributed inference

There is provided methods, apparatus and systems for scheduling for distributed inference. According to embodiments, based on collected device capabilities from a plurality of devices that are to perform a distributed inference, a central node can divide an inference task into multiple inference subtasks and assign these tasks to the plurality of devices. Furthermore, according to embodiments, also based on the collected device capabilities and divided inference subtasks, a central node can configure the distributed inference network or network topology associated with the devices, such that these devices can work jointly to perform the distributed inference task. According to embodiments, the plurality of devices associated with the distributed inference task can be transmitted their respective configuration information from the central node. The configuration information can include details regarding the inference subtask being performed and information indicative of reception of input for the inference subtask and the transmission of output upon completion of the inference subtask.
Owner:HUAWEI TECH CO LTD

Hardware resource dynamically adaptive edge-side large model distributed inference method

The application provides an edge large model distributed inference method for dynamic self-adaptation of hardware resources, comprising the following steps: a master node collects operation characteristics related to a large model inference task, including data of the large model inference task itself, computing capacity of each computing node, memory of each computing node, and network bandwidth between devices, obtains memory consumption of each layer task according to an activation value tensor size of each layer of the large model inference task, and calculates execution time of each layer on different computing nodes; according to the operation characteristics and the execution time, transmission overhead and migration overhead between the computing nodes are modeled to obtain a model division strategy for minimizing total execution time of inference; the master node distributes model parameters and computing tasks to the computing nodes according to the model division strategy, and each computing node pauses inference every interval of a preset time, and the performance profiling step is executed again until the entire large model inference task is completed.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Routing method and device for distributed reasoning task, equipment, medium and product

The invention discloses a routing method, device and equipment for a distributed reasoning task, a medium and a product. The method comprises the following steps: acquiring the distributed reasoning task; according to the distributed reasoning task, determining a task characteristic parameter set, a slice state set of the network slices meeting task requirements and a node state set of the node set; based on the task feature parameter set, the slice state set and the node state set, performing multi-dimensional evaluation and screening on nodes in the node set to obtain an effective candidate node set; and according to the effective candidate node set and the multi-dimensional joint decision-making condition, carrying out elastic routing decision-making to obtain an optimal node. Effective candidate nodes are screened by combining the states of tasks, slices and nodes with a multi-dimensional evaluation mode, and then elastic routing decision is carried out according to a multi-dimensional joint decision condition to obtain an optimal node, so that dynamic adaptation of a distributed reasoning task, an optimal calculation node and slice resources is realized, the delay violation rate is remarkably reduced, and the resource utilization rate is improved.
Owner:PURPLE MOUNTAIN LAB