Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

75 results about "Distributed reasoning" patented technology

Inference method and apparatus for large language model

An inference method and apparatus for a large language model, which method and apparatus are used for improving the inference calculation efficiency of a large language model. The method comprises: receiving an inference task, which carries a morpheme number, wherein the morpheme number is used for identifying a morpheme to be subjected to inference calculation, the morpheme comprising one or more sub-words determined on the basis of input text of a large language model (201); on the basis of the morpheme number corresponding to the inference task, querying a global index tree to determine a target inference node, wherein the global index tree comprises sub-trees corresponding to a plurality of inference nodes in a distributed inference cluster, and a sub-tree corresponding to the target inference node is a sub-tree in the global index tree, the matching degree of similarity of which sub-tree with the morpheme number is greater than a threshold, the matching degree of similarity indicating the number of pieces of reusable key-value cache data during inference calculation (202); and on the basis of the target inference node, executing the inference task (203).
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Istio-based inference service global activation connection number limiting method

The invention relates to the technical field of micro-services, and discloses an Istio-based inference service global activation connection number limiting method, which comprises the following steps: step 1, when a client request is received, extracting label information in a request header, and carrying out label analysis; 2, after label analysis is completed, threshold value query is conducted, and interface response content is obtained; 3, acquiring real-time load data of the target service, and calculating a real-time dynamic threshold value; 4, performing three-level quota verification from the tenant level, the service level and the instance level; and 5, after three-level quota verification, connection processing is carried out, a connection processing result is reported, and closed-loop feedback is carried out. According to the scheme, a global and intelligent connection number limiting mechanism is constructed in a micro-service architecture, the problems of service overload, resource competition and system stability caused by sudden increase of the connection number in distributed reasoning service are solved, and dynamic sensing, intelligent decision making and accurate control of the reasoning service connection number are achieved.
Owner:ASPIRE TECH (SHENZHEN) LTD

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Big language model reasoning method and device

The embodiment of the invention discloses an inference method and device for a large language model, which are used for improving the inference calculation efficiency of the large language model. The method comprises the steps that an inference task is received, the inference task carries morpheme numbers, the morpheme numbers are used for identifying morphemes to be subjected to inference calculation, and each morpheme comprises one or more sub-words determined based on an input text of a large language model; a global index tree is inquired based on morpheme numbers corresponding to the reasoning tasks, target reasoning nodes are determined, the global index tree comprises sub-trees corresponding to multiple reasoning nodes in a distributed reasoning cluster, and the sub-trees corresponding to the target reasoning nodes are sub-trees with the similarity matching degree with the morpheme numbers larger than a threshold value in the global index tree; the similarity matching degree indicates the number of reusable key value cache data in reasoning calculation. And executing the reasoning task based on the target reasoning node.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Cloud edge collaborative diffusion model reasoning method and system based on block chain and reinforcement learning

The invention provides a cloud edge collaborative diffusion model reasoning method and system based on a block chain and reinforcement learning, and the method comprises the steps: obtaining a current text prompt word submitted by a user in a block chain network composed of a cloud server and an edge server, and calling a semantic matching model through an intelligent contract, and judging whether a historical intermediate result can be reused or not; obtaining a server environment state, generating a collaborative reasoning strategy by using a pre-trained multi-agent attention actor-commentator model, and determining cloud edge denoising step number distribution; if the image cannot be reused, the cloud server executes partial denoising to generate an intermediate result and stores the intermediate result to the block chain, and the edge server continues to complete residual denoising to generate a final image; and if the edge server can be reused, the edge server directly completes denoising based on the historical intermediate result. According to the method, historical intermediate results can be effectively reused, dynamic intelligent task allocation is realized, the diffusion model reasoning efficiency and the image generation quality are improved, and the credibility and the traceability of a distributed reasoning process are guaranteed.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Methods and apparatus for support of vertical federated learning at network data analytics function (NWDAF)

Apparatus and methods to perform vertical federated learning (VFL) in a network, e.g. 5G New Radio network. At step 13, a VFL server, such as a Network Data Analytics Function (NWDAF) hosting VFL functionality, receives a message from a service consumer such as a consumer network function (NF). The message may be a subscription to analytics or a request for analytics. At step 14, a distributed VFL inference process is triggered by the VFL server invoking a Machine Learning (ML) model inference request. The request is transmitted to a participant, i.e. a VFL client entity which may also be a NWDAF. At step 15, based on the request, the VFL client entity obtains intermediate inference results by performing local inference computation using a local ML model. At step 16, if the VFL client entity is the last participant, the intermediate inference results are transmitted to the VFL server. At step 17, the VFL server performs further inference computation on the received intermediate inference results and derives the requested analytics which may be delivered to the NF consumer at step 18. The VFL server may discover or select the VFL client entities to participate in the VFL procedure.
Owner:SAMSUNG ELECTRONICS CO LTD

Distributed large model reasoning optimization method and device based on dynamic micro-batch scheduling

The invention discloses a distributed large model reasoning optimization method and device based on dynamic micro-batch scheduling, and the method comprises the steps: (1) initializing a system, and building a large model distributed reasoning assembly line; (2) evaluating the computing power and the network state of each node and summarizing the computing power and the network state to a head node; (3) determining the number of Micro-Batches and the scheduling quota of each Micro-Batch according to the request distribution condition, the computing power of each node and the network state; and (4) sequentially scheduling the Micro-Batch by adopting a Control Batch strategy and a Chunked Prefill strategy, and starting to execute the Micro-Batch. According to the method, by dynamically adjusting the number of Micro-Batch, the serious assembly line cavitation problem in a distributed large model reasoning system is effectively solved, the GPU utilization rate and the system throughput are remarkably improved, and meanwhile key indexes such as TTFT (first token time delay) and TPOT (inter-token time delay) in the large model reasoning field are also improved. The method has good adaptability, can adaptively adjust the dynamic scheduling strategy under different hardware equipment, network conditions and request loads, is suitable for different distributed large model deployment scenes, and has wide application value.
Owner:ZHEJIANG UNIV

Cross-GPU video memory dynamic scheduling system and method based on equipment demand perception

The invention discloses a cross-GPU video memory dynamic scheduling system and method based on equipment demand perception, belongs to the field of artificial intelligence of computer science, and is suitable for a distributed reasoning scene of a large language model. The method comprises the following steps: aiming at the problems of unbalanced utilization of video memory resources and cross-device communication bottleneck in a PD separation architecture, a system constructs a video memory exchange area in a pre-filling instance, when the video memory occupation of a decoding instance is critical, part of KV cache is migrated to the exchange area and returns when the resources are allowed, so that the idle video memory resources of the pre-filling instance are fully utilized, and the resource utilization rate of the decoding instance is improved. And the video memory pressure in the decoding stage is relieved. The system realizes the conversion from the low-efficiency KV cache migration process transferred by the original CPU to the GPU direct access mode through a high-speed communication path established by direct connection between the GPUs. A layered asynchronous transmission strategy is adopted in the migration process, communication overhead is hidden in a calculation assembly line, migration delay is reduced, and cross-GPU video memory sharing multiplexing and efficient scheduling are achieved.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1

Orchestration of distributed inference operations

Approaches presented herein provide for the management of resources to be used to process a request, such as may involve orchestration of nodes for an inference request. Upon receiving an inference request, an orchestrator can determine a sequence of context nodes and inference nodes to be used to process the inference request, based in part upon a determined class of inferencing to be performed. The orchestrator can append metadata to the inference request that identifies the sequence, and can transmit the appended request to one or more first nodes in the sequence. If the nodes have a network programmable device, or similar capability, the request can be forwarded to the nodes in sequence without having to go back to the orchestrator between nodes.
Owner:NVIDIA CORP

Ensemble communication unloading method, system, equipment and medium

ActiveCN121979690AResource allocationInference methodsCollective communicationComputer network
The invention discloses a set communication unloading method, system and device and a medium, and is applied to the technical field of computers, and the method comprises the steps: in a model deployment stage, a distributed reasoning controller generates a communication primitive blueprint based on a computational graph description file of a tensor parallel reasoning model and issues the communication primitive blueprint to a DPU; the DPU establishes a hardware-level communication context semantic environment based on the communication primitive blueprint; in the model reasoning stage, the GPU / NPU sends a trigger signal to the DPU when calculating to a communication boundary; and the DPU executes a DMA data pulling assembly line and an RDMA data sending assembly line in parallel based on a hardware-level communication context semantic environment, performs aggregation calculation on all tensor fragment data to be synchronized, and writes an aggregation calculation result back to the GPU / NPU. A hardware-level communication context semantic environment is established in advance to a model deployment stage, and double assembly lines are executed in the DPU in parallel, so that end-to-end communication delay is remarkably reduced, and zero participation of a host CPU is realized.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD +1

A distributed inference communication method and system of a super-large-scale graph neural network

PendingCN122450656AGlobal topologyIndependent set
A distributed inference communication method of a super-large-scale graph neural network comprises the following steps: dividing a global topology to be processed into a plurality of local sub-domains matched with the number of GPUs, and determining a local master node and a boundary master node allocated to each GPU; calculating coloring priorities of nodes, and performing parallel graph coloring based on the coloring priorities to construct a global independent set sequence; mapping the global independent set sequence to the local sub-domains of the GPUs to generate a local task slice sequence, which comprises boundary node slices and internal node slices, evaluating communication load and calculation load of the slices, and taking the internal node slices as calculation fillers to be executed in the communication waiting period of the boundary node slices; and each GPU executes an inference task by using a multi-stream asynchronous pipeline constructed locally according to the local task slice sequence. The method can improve the throughput of distributed inference.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Distributed cluster construction method, distributed reasoning method and resource scheduler

The invention provides a distributed cluster construction method, a distributed reasoning method and a resource scheduler, which can be applied to the technical field of computers. The distributed cluster construction method comprises: in response to a submission operation for a system arrangement file, binding a virtual node created based on the system arrangement file to a physical node equipped with a hardware accelerator, the system arrangement file indicating an expected construction state for a distributed cluster; calling a plurality of physical nodes, pulling a target reasoning mirror image from a mirror image warehouse, and starting respective container instances of the plurality of physical nodes by using the target reasoning mirror image, the target reasoning mirror image being obtained by packaging a development kit of a hardware accelerator of the physical node and a distributed reasoning framework; calling an accelerator plug-in of the physical node, and allocating resources of a hardware accelerator on the physical node to the container instance; and constructing a distributed cluster according to the plurality of container instances.
Owner:BEIJING KECHENG TECH DEV CO LTD

Task reasoning method and device and electronic equipment

The embodiment of the invention provides a task reasoning method and device and electronic equipment. The method is applied to a distributed reasoning architecture, the distributed reasoning architecture comprises a first node, a second node and a shared memory pool, and the method comprises the steps that the first node obtains at least one first to-be-reasoned task; the first node retrieves the first to-be-reasoned task to obtain at least one first retrieval result related to the first to-be-reasoned task; the first node stores the at least one first retrieval result in a shared memory pool; the second node reads at least one first retrieval result from the shared memory pool; and the second node takes the at least one first to-be-reasoned task and the at least one first retrieval result corresponding to the first to-be-reasoned task as input of a model, and executes the first to-be-reasoned task by utilizing the model. By adopting the mode, the delay problem caused by cross-node transmission can be avoided, so that the overall efficiency of the reasoning task is improved.
Owner:XFUSION DIGITAL TECH CO LTD

A heterogeneous cluster distributed inference collaboration method and system based on VLLM

The application discloses a VLLM-based heterogeneous cluster distributed reasoning cooperation method and system, and the method comprises the following steps: S1, generating a heterogeneous device performance image signal through a performance probe arranged in each computing unit in the cluster; S2, performing non-uniform task division according to the heterogeneous device performance image signal and combining a calculation graph structure of a to-be-reasoned model to form an optimal task scheduling strategy signal for a heterogeneous environment; S3, generating a task allocation and loading instruction signal based on the optimal task scheduling strategy signal; S4, coordinating each heterogeneous computing unit to perform reasoning calculation in a distributed reasoning execution process to generate a distributed cooperative reasoning signal; and S5, integrating and post-processing the distributed cooperative reasoning signal to output a final reasoning result. The VLLM-based heterogeneous cluster distributed reasoning cooperation method and system can solve the problems of low resource utilization and poor reasoning efficiency when a large model is run on a heterogeneous cluster.
Owner:ENTERPRISE ONLINE (BEIJING) NETWORK CO LTD

Online inference method, service system, and device based on an artificial intelligence geospatial data cube

The disclosure pertains to the field of data analysis and services, specifically to an online inference method, service system, and device based on an artificial intelligence geospatial data cube. The method involves constructing a cube organizational model based on a spatiotemporal grid, where the cube organizes and manages geospatial data and GeoAI models in a unified manner. It performs task-oriented model matching through a combination of explicit and implicit matching, where explicit matching converts user inference requests into multidimensional query conditions to retrieve candidate models, and implicit matching determines the optimal model by computing the feature similarity between inference data tiles and candidate models. Based on a distributed inference framework, an efficient inference workflow is executed in parallel across multiple computing nodes. The disclosure aims to address the limitations of existing technologies in multi-scenarios or large-scale spatial scenarios, such as low inference accuracy and slow inference speed.
Owner:WUHAN UNIV

Distributed reasoning task arrangement method based on semantic driving

The invention discloses a distributed reasoning task arrangement method based on semantic driving. The method comprises the following steps: constructing a domain ontology model and establishing semantic mapping; analyzing task requirements; executing data source retrieval of multi-level semantic fusion; generating an execution path based on multi-objective collaborative optimization; tasks are distributed, and transmission overhead is reduced by adopting model differential updating; and performing adaptive migration and network partition fault tolerance based on a check point and a two-stage commit protocol. According to the method, the technical problems of heterogeneous data semantic gaps, static scheduling and dynamic environment conflicts and state loss caused by coarse-grained fault tolerance are solved, efficient distributed reasoning of data immobility and calculation flow is achieved, and the method is suitable for the multi-source heterogeneous data fields of marine science, meteorological monitoring, industrial Internet of Things and the like.
Owner:SUZHOU ABYSS MATRIX TECHNOLOGY CO LTD

Inference request scheduling method and device, equipment, medium and distributed inference system

The invention relates to a reasoning request scheduling method, device and equipment, a medium and a distributed reasoning system, and relates to the technical field of intelligent calculation. The method comprises the following steps: in response to a reasoning request of a user side, if intermediate cache data corresponding to the reasoning request does not exist in local storage and storage nodes of a distributed shared storage pool, determining a corresponding reasoning service demand according to the reasoning request; determining the demand matching degree corresponding to each reasoning processing node combination according to the reasoning service demand and the communication link attribute between the nodes in each reasoning processing node combination; and according to the demand matching degree, scheduling the reasoning request to a pre-filling node and a decoding node contained in a target reasoning processing node combination in each reasoning processing node combination for reasoning processing. By adopting the method, the memory use efficiency can be improved, the memory overflow risk can be reduced, the high-concurrency reasoning service requirement can be met, and the performance bottleneck in a high-load scene can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Method and apparatus for scheduling for distributed inference

There is provided methods, apparatus and systems for scheduling for distributed inference. According to embodiments, based on collected device capabilities from a plurality of devices that are to perform a distributed inference, a central node can divide an inference task into multiple inference subtasks and assign these tasks to the plurality of devices. Furthermore, according to embodiments, also based on the collected device capabilities and divided inference subtasks, a central node can configure the distributed inference network or network topology associated with the devices, such that these devices can work jointly to perform the distributed inference task. According to embodiments, the plurality of devices associated with the distributed inference task can be transmitted their respective configuration information from the central node. The configuration information can include details regarding the inference subtask being performed and information indicative of reception of input for the inference subtask and the transmission of output upon completion of the inference subtask.
Owner:HUAWEI TECH CO LTD

Hardware resource dynamically adaptive edge-side large model distributed inference method

The application provides an edge large model distributed inference method for dynamic self-adaptation of hardware resources, comprising the following steps: a master node collects operation characteristics related to a large model inference task, including data of the large model inference task itself, computing capacity of each computing node, memory of each computing node, and network bandwidth between devices, obtains memory consumption of each layer task according to an activation value tensor size of each layer of the large model inference task, and calculates execution time of each layer on different computing nodes; according to the operation characteristics and the execution time, transmission overhead and migration overhead between the computing nodes are modeled to obtain a model division strategy for minimizing total execution time of inference; the master node distributes model parameters and computing tasks to the computing nodes according to the model division strategy, and each computing node pauses inference every interval of a preset time, and the performance profiling step is executed again until the entire large model inference task is completed.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Routing method and device for distributed reasoning task, equipment, medium and product

The invention discloses a routing method, device and equipment for a distributed reasoning task, a medium and a product. The method comprises the following steps: acquiring the distributed reasoning task; according to the distributed reasoning task, determining a task characteristic parameter set, a slice state set of the network slices meeting task requirements and a node state set of the node set; based on the task feature parameter set, the slice state set and the node state set, performing multi-dimensional evaluation and screening on nodes in the node set to obtain an effective candidate node set; and according to the effective candidate node set and the multi-dimensional joint decision-making condition, carrying out elastic routing decision-making to obtain an optimal node. Effective candidate nodes are screened by combining the states of tasks, slices and nodes with a multi-dimensional evaluation mode, and then elastic routing decision is carried out according to a multi-dimensional joint decision condition to obtain an optimal node, so that dynamic adaptation of a distributed reasoning task, an optimal calculation node and slice resources is realized, the delay violation rate is remarkably reduced, and the resource utilization rate is improved.
Owner:PURPLE MOUNTAIN LAB

Distributed reasoning method and device based on hybrid expert architecture, equipment and medium

The application relates to the field of artificial intelligence and provides a distributed reasoning method, device and equipment based on a hybrid expert architecture and a medium, wherein the distributed reasoning method based on the hybrid expert architecture applied to an edge computing gateway comprises the following steps: if a reasoning task is received, determining a plurality of target intelligent vehicles participating in the reasoning task; acquiring an expert reasoning model and deploying the expert reasoning model into each target intelligent vehicle; wherein the expert reasoning model comprises a plurality of expert reasoning modules, and at least one expert reasoning module is deployed in each target intelligent vehicle; sending the reasoning task to each target intelligent vehicle, so that each target intelligent vehicle executes the reasoning task based on the expert reasoning module to obtain a reasoning result; and aggregating the reasoning results of the target intelligent vehicles to obtain a target reasoning result. Through the technical scheme provided by the application, the computing power resources of idle intelligent vehicles are fully utilized.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

System and method for distributed reasoning in autonomous systems

What is disclosed is: A system to implement distributed reasoning for autonomous operation in a team of members, wherein the team of members comprise a first member, wherein the first member comprises a first team autonomous subsystem and first member autonomous subsystem coupled to each other, the first member executes a first hierarchical plan comprising a team plan and a first member plan, wherein the team plan is related to one or more team goals, and the first member plan is related to one or more self-goals for the first member, further wherein the execution comprises the first team autonomous system executing the team plan, and the first member autonomous system executing the first member plan.
Owner:FOUR DROBOTICS CORP

Distributed reasoning energy efficiency optimization method and device based on energy storage, equipment and medium

The invention discloses a distributed reasoning energy efficiency optimization method and device based on energy storage, equipment and a medium. The method is characterized by comprising the following steps: for a dynamic reasoning task set corresponding to each task execution time slot, obtaining model set parameters, energy storage system parameters and power supply parameter information of distributed computing nodes; distributing an optimal reasoning model in a current task execution time slot for each reasoning task of the dynamic reasoning task set based on the model set parameters, the energy storage system parameters and the power supply parameter information, and determining a task model distribution decision and a power supply decision corresponding to the current task execution time slot; and performing model reasoning on each reasoning task based on the task model distribution decision and the power supply decision. According to the method, the operation cost of the distributed computing nodes can be effectively reduced, and the energy efficiency and the operation stability of reasoning task processing are improved; based on the task model allocation decision and the power supply decision, the service quality balance is improved, the power supply resource utilization efficiency is improved, the power control precision is improved, and the energy storage effect is improved.
Owner:INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +1

Insurance-oriented AIGC compliance detection method, apparatus and device, and medium

The invention relates to the technical field of artificial intelligence, provides an insurance-oriented AIGC compliance detection method and device, equipment and a medium, is applied to financial and medical health care service scenes, and can collect insurance field knowledge based on a retrieval enhancement generation mechanism to construct a dynamically-updated knowledge base so as to improve the accuracy of the AIGC compliance detection. Comprehensive, real-time and accurate knowledge support is provided for model judgment, and misjudgment of the model due to knowledge blind areas is avoided; low-rank adaptive fine tuning is performed on the large language model according to the knowledge base, so that the identification capability of the model on specific risk points of the industry can be further improved while the cost is reduced, and the accuracy of compliance detection is improved; the task fragmentation is performed based on the distributed reasoning framework, and the sub-tasks are distributed to the plurality of detection nodes, so that the problem of insufficient computing resources of a single node can be solved, the task processing throughput is improved, and second-level response can still be realized in a high-concurrency scene, thereby improving the efficiency of compliance detection.
Owner:PING AN TECH (SHENZHEN) CO LTD

Information processing method and device, electronic equipment and storage medium

PendingCN120851213ABiological modelsInference methodsDistributed information processingEngineering
The invention provides an information processing method, and relates to the technical field of artificial intelligence, in particular to the technical fields of large models, reinforcement learning, deep learning, distributed training, distributed reasoning and the like. According to the specific implementation scheme, at least one target training sample is determined according to at least one piece of problem information and at least one initial response result used for the at least one piece of problem information, and the initial response result is obtained through reasoning by a reasoning service layer through a to-be-trained model according to the problem information; providing the at least one target training sample to a training service layer, wherein the training service layer is used for determining weight update data of the to-be-trained model according to the at least one target training sample; and providing the weight update data to an inference service layer. The invention further provides a distributed information processing system and device, electronic equipment and a storage medium.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Cloud native computing power network machine learning distributed reasoning method and system

The invention discloses a server-free distributed deep neural network inference system and method based on computing power network cloud side-end collaboration, a three-layer computing structure of an edge layer, a fog layer and a cloud layer is designed, OpenFaaS or Fission server-free platform hierarchical inference function deployment is customized, cross-layer communication is completed by adopting a customized KafkaHTTP asynchronous connector developed based on TypeScript / Deno, and the cloud side-end collaboration based server-free distributed deep neural network inference system and the cloud side-end collaboration based server-free distributed deep neural network inference method based on TypeScript / Deno. The method has the following innovations: 1) a double-function collaborative architecture is proposed for the first time, a funnel function is responsible for Kafka connection, an inference function is responsible for loading DDDN hierarchical calculation by a TensorFlow model, and the cold start delay is reduced by 40% after separation deployment; 2) dynamic partition optimization is carried out, and a BranchyNet architecture early exit strategy (confidence coefficient gt; a result can be returned when the result is equal to 95%), an ARM / x86 heterogeneous hardware template is supported, and a DDNN model is supported in a resource-constrained device (CPUlt; the length is equal to 500m, and the memory is lt; = 512 MB); 3) a connector technology is innovated, a kafkatohttp connector adopts a non-blocking IO model, a 1000 + req / s high throughput proxy is realized, and the problem of resource leakage of a native connector of a platform is solved; and 4) an intelligent capacity expansion and contraction engine is adopted, the OpenFaaS uses Prometheus alarm to trigger copy dynamic adjustment (120 copies), the Fission uses Kubernetes HPA to carry out copy accurate control, and the delay is optimized by 55%, 318-494 s compared with a KafkaML framework under five client loads. The system has millisecond-level reasoning delay, can operate on various platforms such as a RaspberryPi-GPU cluster and the like, and has the characteristics of automatic expansion and contraction and the like.
Owner:ZHEJIANG UNIV

Model function alignment method in communication system and communication apparatus

A model function alignment method in a communication system and a communication apparatus. The method comprises: a terminal acquires a first function supported by a network device; the terminal sends first information to the network device, wherein the first information indicates a second function and / or a third function, the second function and the third function are associated with a first task, and the first function comprises the second function and does not comprise the third function; and the terminal acquires a computation result of the first task, the computation result being determined on the basis of AI models of the second function and the third function. In the method, distributed reasoning of functions can be realized, and functions limited by a capability of a terminal side are implemented on a network device side, thereby improving the user experience.
Owner:HUAWEI TECH CO LTD

A communication scheduling method and system based on PON physical layer in-order aggregation

The application discloses a communication scheduling method and system based on physical layer ordered aggregation of PON, and belongs to the technical field of optical access network communication. The application realizes full-isomorphic mapping of distributed reasoning task logical topology to physical network by constructing an optical distribution network physical network diagram and establishing an algorithm power collaborative cluster. Each network terminal device embeds a real-time state vector in an extended XGEM frame header of an uplink data frame, and an optical line terminal constructs a dynamic panoramic ledger based on the real-time state vector, predicts a logical fragment completion time, generates an authorized window sequence satisfying phase locking and ordered constraints, and broadcasts the authorized window sequence to each device. The device uploads an execution result in a corresponding window, and an aggregation point records and generates an audit report according to an actual arrival order. The application does not need to cache rearrangement, realizes physical layer ordered aggregation, and improves the scheduling efficiency and real-time performance of a distributed reasoning task.
Owner:SICHUAN VOCATIONAL & TECH COLLEGE OF POSTS & TELECOMM

Distributed multi-modal training reasoning method and system for heterogeneous nodes and medium

The invention relates to the technical field of distributed computing, and discloses a distributed multi-modal training reasoning method and system for heterogeneous nodes and a medium, and the method comprises the steps: analyzing a distributed deployment topological structure of a multi-modal model, obtaining a connection relation between modal branches, and generating a modal interaction dependency graph; applying the unified precision configuration of the precision cooperation group to an input end of a cross-modal interaction layer, configuring a precision alignment operator, and outputting interactive input data after precision alignment; obtaining the hardware precision capability of each heterogeneous node in the reasoning stage, and generating a training-reasoning precision migration mapping table; and based on the training-reasoning precision migration mapping table, generating precision configuration and a distributed reasoning scheduling scheme in a reasoning stage. The technical effects of guaranteeing the stability of the model training numerical value and improving the convergence quality are achieved.
Owner:NANJING HUIZHI INTERACTIVE ENTERTAINMENT NETWORK TECH CO LTD