Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

53 results about "Distributed reasoning" patented technology

Istio-based inference service global activation connection number limiting method

The invention relates to the technical field of micro-services, and discloses an Istio-based inference service global activation connection number limiting method, which comprises the following steps: step 1, when a client request is received, extracting label information in a request header, and carrying out label analysis; 2, after label analysis is completed, threshold value query is conducted, and interface response content is obtained; 3, acquiring real-time load data of the target service, and calculating a real-time dynamic threshold value; 4, performing three-level quota verification from the tenant level, the service level and the instance level; and 5, after three-level quota verification, connection processing is carried out, a connection processing result is reported, and closed-loop feedback is carried out. According to the scheme, a global and intelligent connection number limiting mechanism is constructed in a micro-service architecture, the problems of service overload, resource competition and system stability caused by sudden increase of the connection number in distributed reasoning service are solved, and dynamic sensing, intelligent decision making and accurate control of the reasoning service connection number are achieved.
Owner:ASPIRE TECH (SHENZHEN) LTD

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Cloud edge collaborative diffusion model reasoning method and system based on block chain and reinforcement learning

The invention provides a cloud edge collaborative diffusion model reasoning method and system based on a block chain and reinforcement learning, and the method comprises the steps: obtaining a current text prompt word submitted by a user in a block chain network composed of a cloud server and an edge server, and calling a semantic matching model through an intelligent contract, and judging whether a historical intermediate result can be reused or not; obtaining a server environment state, generating a collaborative reasoning strategy by using a pre-trained multi-agent attention actor-commentator model, and determining cloud edge denoising step number distribution; if the image cannot be reused, the cloud server executes partial denoising to generate an intermediate result and stores the intermediate result to the block chain, and the edge server continues to complete residual denoising to generate a final image; and if the edge server can be reused, the edge server directly completes denoising based on the historical intermediate result. According to the method, historical intermediate results can be effectively reused, dynamic intelligent task allocation is realized, the diffusion model reasoning efficiency and the image generation quality are improved, and the credibility and the traceability of a distributed reasoning process are guaranteed.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Orchestration of distributed inference operations

Approaches presented herein provide for the management of resources to be used to process a request, such as may involve orchestration of nodes for an inference request. Upon receiving an inference request, an orchestrator can determine a sequence of context nodes and inference nodes to be used to process the inference request, based in part upon a determined class of inferencing to be performed. The orchestrator can append metadata to the inference request that identifies the sequence, and can transmit the appended request to one or more first nodes in the sequence. If the nodes have a network programmable device, or similar capability, the request can be forwarded to the nodes in sequence without having to go back to the orchestrator between nodes.
Owner:NVIDIA CORP

Ensemble communication unloading method, system, equipment and medium

ActiveCN121979690AResource allocationInference methodsCollective communicationComputer network
The invention discloses a set communication unloading method, system and device and a medium, and is applied to the technical field of computers, and the method comprises the steps: in a model deployment stage, a distributed reasoning controller generates a communication primitive blueprint based on a computational graph description file of a tensor parallel reasoning model and issues the communication primitive blueprint to a DPU; the DPU establishes a hardware-level communication context semantic environment based on the communication primitive blueprint; in the model reasoning stage, the GPU / NPU sends a trigger signal to the DPU when calculating to a communication boundary; and the DPU executes a DMA data pulling assembly line and an RDMA data sending assembly line in parallel based on a hardware-level communication context semantic environment, performs aggregation calculation on all tensor fragment data to be synchronized, and writes an aggregation calculation result back to the GPU / NPU. A hardware-level communication context semantic environment is established in advance to a model deployment stage, and double assembly lines are executed in the DPU in parallel, so that end-to-end communication delay is remarkably reduced, and zero participation of a host CPU is realized.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD +1

A distributed inference communication method and system of a super-large-scale graph neural network

PendingCN122450656AGlobal topologyIndependent set
A distributed inference communication method of a super-large-scale graph neural network comprises the following steps: dividing a global topology to be processed into a plurality of local sub-domains matched with the number of GPUs, and determining a local master node and a boundary master node allocated to each GPU; calculating coloring priorities of nodes, and performing parallel graph coloring based on the coloring priorities to construct a global independent set sequence; mapping the global independent set sequence to the local sub-domains of the GPUs to generate a local task slice sequence, which comprises boundary node slices and internal node slices, evaluating communication load and calculation load of the slices, and taking the internal node slices as calculation fillers to be executed in the communication waiting period of the boundary node slices; and each GPU executes an inference task by using a multi-stream asynchronous pipeline constructed locally according to the local task slice sequence. The method can improve the throughput of distributed inference.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Distributed cluster construction method, distributed reasoning method and resource scheduler

The invention provides a distributed cluster construction method, a distributed reasoning method and a resource scheduler, which can be applied to the technical field of computers. The distributed cluster construction method comprises: in response to a submission operation for a system arrangement file, binding a virtual node created based on the system arrangement file to a physical node equipped with a hardware accelerator, the system arrangement file indicating an expected construction state for a distributed cluster; calling a plurality of physical nodes, pulling a target reasoning mirror image from a mirror image warehouse, and starting respective container instances of the plurality of physical nodes by using the target reasoning mirror image, the target reasoning mirror image being obtained by packaging a development kit of a hardware accelerator of the physical node and a distributed reasoning framework; calling an accelerator plug-in of the physical node, and allocating resources of a hardware accelerator on the physical node to the container instance; and constructing a distributed cluster according to the plurality of container instances.
Owner:BEIJING KECHENG TECH DEV CO LTD

Task reasoning method and device and electronic equipment

The embodiment of the invention provides a task reasoning method and device and electronic equipment. The method is applied to a distributed reasoning architecture, the distributed reasoning architecture comprises a first node, a second node and a shared memory pool, and the method comprises the steps that the first node obtains at least one first to-be-reasoned task; the first node retrieves the first to-be-reasoned task to obtain at least one first retrieval result related to the first to-be-reasoned task; the first node stores the at least one first retrieval result in a shared memory pool; the second node reads at least one first retrieval result from the shared memory pool; and the second node takes the at least one first to-be-reasoned task and the at least one first retrieval result corresponding to the first to-be-reasoned task as input of a model, and executes the first to-be-reasoned task by utilizing the model. By adopting the mode, the delay problem caused by cross-node transmission can be avoided, so that the overall efficiency of the reasoning task is improved.
Owner:XFUSION DIGITAL TECH CO LTD

A heterogeneous cluster distributed inference collaboration method and system based on VLLM

The application discloses a VLLM-based heterogeneous cluster distributed reasoning cooperation method and system, and the method comprises the following steps: S1, generating a heterogeneous device performance image signal through a performance probe arranged in each computing unit in the cluster; S2, performing non-uniform task division according to the heterogeneous device performance image signal and combining a calculation graph structure of a to-be-reasoned model to form an optimal task scheduling strategy signal for a heterogeneous environment; S3, generating a task allocation and loading instruction signal based on the optimal task scheduling strategy signal; S4, coordinating each heterogeneous computing unit to perform reasoning calculation in a distributed reasoning execution process to generate a distributed cooperative reasoning signal; and S5, integrating and post-processing the distributed cooperative reasoning signal to output a final reasoning result. The VLLM-based heterogeneous cluster distributed reasoning cooperation method and system can solve the problems of low resource utilization and poor reasoning efficiency when a large model is run on a heterogeneous cluster.
Owner:ENTERPRISE ONLINE (BEIJING) NETWORK CO LTD

Online inference method, service system, and device based on an artificial intelligence geospatial data cube

The disclosure pertains to the field of data analysis and services, specifically to an online inference method, service system, and device based on an artificial intelligence geospatial data cube. The method involves constructing a cube organizational model based on a spatiotemporal grid, where the cube organizes and manages geospatial data and GeoAI models in a unified manner. It performs task-oriented model matching through a combination of explicit and implicit matching, where explicit matching converts user inference requests into multidimensional query conditions to retrieve candidate models, and implicit matching determines the optimal model by computing the feature similarity between inference data tiles and candidate models. Based on a distributed inference framework, an efficient inference workflow is executed in parallel across multiple computing nodes. The disclosure aims to address the limitations of existing technologies in multi-scenarios or large-scale spatial scenarios, such as low inference accuracy and slow inference speed.
Owner:WUHAN UNIV

Distributed reasoning task arrangement method based on semantic driving

The invention discloses a distributed reasoning task arrangement method based on semantic driving. The method comprises the following steps: constructing a domain ontology model and establishing semantic mapping; analyzing task requirements; executing data source retrieval of multi-level semantic fusion; generating an execution path based on multi-objective collaborative optimization; tasks are distributed, and transmission overhead is reduced by adopting model differential updating; and performing adaptive migration and network partition fault tolerance based on a check point and a two-stage commit protocol. According to the method, the technical problems of heterogeneous data semantic gaps, static scheduling and dynamic environment conflicts and state loss caused by coarse-grained fault tolerance are solved, efficient distributed reasoning of data immobility and calculation flow is achieved, and the method is suitable for the multi-source heterogeneous data fields of marine science, meteorological monitoring, industrial Internet of Things and the like.
Owner:SUZHOU ABYSS MATRIX TECHNOLOGY CO LTD

Inference request scheduling method and device, equipment, medium and distributed inference system

The invention relates to a reasoning request scheduling method, device and equipment, a medium and a distributed reasoning system, and relates to the technical field of intelligent calculation. The method comprises the following steps: in response to a reasoning request of a user side, if intermediate cache data corresponding to the reasoning request does not exist in local storage and storage nodes of a distributed shared storage pool, determining a corresponding reasoning service demand according to the reasoning request; determining the demand matching degree corresponding to each reasoning processing node combination according to the reasoning service demand and the communication link attribute between the nodes in each reasoning processing node combination; and according to the demand matching degree, scheduling the reasoning request to a pre-filling node and a decoding node contained in a target reasoning processing node combination in each reasoning processing node combination for reasoning processing. By adopting the method, the memory use efficiency can be improved, the memory overflow risk can be reduced, the high-concurrency reasoning service requirement can be met, and the performance bottleneck in a high-load scene can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Hardware resource dynamically adaptive edge-side large model distributed inference method

The application provides an edge large model distributed inference method for dynamic self-adaptation of hardware resources, comprising the following steps: a master node collects operation characteristics related to a large model inference task, including data of the large model inference task itself, computing capacity of each computing node, memory of each computing node, and network bandwidth between devices, obtains memory consumption of each layer task according to an activation value tensor size of each layer of the large model inference task, and calculates execution time of each layer on different computing nodes; according to the operation characteristics and the execution time, transmission overhead and migration overhead between the computing nodes are modeled to obtain a model division strategy for minimizing total execution time of inference; the master node distributes model parameters and computing tasks to the computing nodes according to the model division strategy, and each computing node pauses inference every interval of a preset time, and the performance profiling step is executed again until the entire large model inference task is completed.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Routing method and device for distributed reasoning task, equipment, medium and product

The invention discloses a routing method, device and equipment for a distributed reasoning task, a medium and a product. The method comprises the following steps: acquiring the distributed reasoning task; according to the distributed reasoning task, determining a task characteristic parameter set, a slice state set of the network slices meeting task requirements and a node state set of the node set; based on the task feature parameter set, the slice state set and the node state set, performing multi-dimensional evaluation and screening on nodes in the node set to obtain an effective candidate node set; and according to the effective candidate node set and the multi-dimensional joint decision-making condition, carrying out elastic routing decision-making to obtain an optimal node. Effective candidate nodes are screened by combining the states of tasks, slices and nodes with a multi-dimensional evaluation mode, and then elastic routing decision is carried out according to a multi-dimensional joint decision condition to obtain an optimal node, so that dynamic adaptation of a distributed reasoning task, an optimal calculation node and slice resources is realized, the delay violation rate is remarkably reduced, and the resource utilization rate is improved.
Owner:PURPLE MOUNTAIN LAB

Distributed reasoning method and device based on hybrid expert architecture, equipment and medium

The application relates to the field of artificial intelligence and provides a distributed reasoning method, device and equipment based on a hybrid expert architecture and a medium, wherein the distributed reasoning method based on the hybrid expert architecture applied to an edge computing gateway comprises the following steps: if a reasoning task is received, determining a plurality of target intelligent vehicles participating in the reasoning task; acquiring an expert reasoning model and deploying the expert reasoning model into each target intelligent vehicle; wherein the expert reasoning model comprises a plurality of expert reasoning modules, and at least one expert reasoning module is deployed in each target intelligent vehicle; sending the reasoning task to each target intelligent vehicle, so that each target intelligent vehicle executes the reasoning task based on the expert reasoning module to obtain a reasoning result; and aggregating the reasoning results of the target intelligent vehicles to obtain a target reasoning result. Through the technical scheme provided by the application, the computing power resources of idle intelligent vehicles are fully utilized.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

System and method for distributed reasoning in autonomous systems

What is disclosed is: A system to implement distributed reasoning for autonomous operation in a team of members, wherein the team of members comprise a first member, wherein the first member comprises a first team autonomous subsystem and first member autonomous subsystem coupled to each other, the first member executes a first hierarchical plan comprising a team plan and a first member plan, wherein the team plan is related to one or more team goals, and the first member plan is related to one or more self-goals for the first member, further wherein the execution comprises the first team autonomous system executing the team plan, and the first member autonomous system executing the first member plan.
Owner:FOUR DROBOTICS CORP

Distributed reasoning energy efficiency optimization method and device based on energy storage, equipment and medium

The invention discloses a distributed reasoning energy efficiency optimization method and device based on energy storage, equipment and a medium. The method is characterized by comprising the following steps: for a dynamic reasoning task set corresponding to each task execution time slot, obtaining model set parameters, energy storage system parameters and power supply parameter information of distributed computing nodes; distributing an optimal reasoning model in a current task execution time slot for each reasoning task of the dynamic reasoning task set based on the model set parameters, the energy storage system parameters and the power supply parameter information, and determining a task model distribution decision and a power supply decision corresponding to the current task execution time slot; and performing model reasoning on each reasoning task based on the task model distribution decision and the power supply decision. According to the method, the operation cost of the distributed computing nodes can be effectively reduced, and the energy efficiency and the operation stability of reasoning task processing are improved; based on the task model allocation decision and the power supply decision, the service quality balance is improved, the power supply resource utilization efficiency is improved, the power control precision is improved, and the energy storage effect is improved.
Owner:INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +1

Model function alignment method in communication system and communication apparatus

A model function alignment method in a communication system and a communication apparatus. The method comprises: a terminal acquires a first function supported by a network device; the terminal sends first information to the network device, wherein the first information indicates a second function and / or a third function, the second function and the third function are associated with a first task, and the first function comprises the second function and does not comprise the third function; and the terminal acquires a computation result of the first task, the computation result being determined on the basis of AI models of the second function and the third function. In the method, distributed reasoning of functions can be realized, and functions limited by a capability of a terminal side are implemented on a network device side, thereby improving the user experience.
Owner:HUAWEI TECH CO LTD

A communication scheduling method and system based on PON physical layer in-order aggregation

The application discloses a communication scheduling method and system based on physical layer ordered aggregation of PON, and belongs to the technical field of optical access network communication. The application realizes full-isomorphic mapping of distributed reasoning task logical topology to physical network by constructing an optical distribution network physical network diagram and establishing an algorithm power collaborative cluster. Each network terminal device embeds a real-time state vector in an extended XGEM frame header of an uplink data frame, and an optical line terminal constructs a dynamic panoramic ledger based on the real-time state vector, predicts a logical fragment completion time, generates an authorized window sequence satisfying phase locking and ordered constraints, and broadcasts the authorized window sequence to each device. The device uploads an execution result in a corresponding window, and an aggregation point records and generates an audit report according to an actual arrival order. The application does not need to cache rearrangement, realizes physical layer ordered aggregation, and improves the scheduling efficiency and real-time performance of a distributed reasoning task.
Owner:SICHUAN VOCATIONAL & TECH COLLEGE OF POSTS & TELECOMM

Distributed multi-modal training reasoning method and system for heterogeneous nodes and medium

The invention relates to the technical field of distributed computing, and discloses a distributed multi-modal training reasoning method and system for heterogeneous nodes and a medium, and the method comprises the steps: analyzing a distributed deployment topological structure of a multi-modal model, obtaining a connection relation between modal branches, and generating a modal interaction dependency graph; applying the unified precision configuration of the precision cooperation group to an input end of a cross-modal interaction layer, configuring a precision alignment operator, and outputting interactive input data after precision alignment; obtaining the hardware precision capability of each heterogeneous node in the reasoning stage, and generating a training-reasoning precision migration mapping table; and based on the training-reasoning precision migration mapping table, generating precision configuration and a distributed reasoning scheduling scheme in a reasoning stage. The technical effects of guaranteeing the stability of the model training numerical value and improving the convergence quality are achieved.
Owner:NANJING HUIZHI INTERACTIVE ENTERTAINMENT NETWORK TECH CO LTD

Orchestration of distributed reasoning operations

The invention relates to orchestration of distributed reasoning operations. The approaches presented herein provide for management of resources to be used to process requests, such as possibly involving orchestration of nodes for inference requests. Upon receiving an inference request, an orchestrator may determine a sequence of context nodes and inference nodes to be used to process the inference request based in part on the determined category of inference to be performed. The orchestrator may append metadata identifying the sequence to the inference request, and may transmit the appended request in order to one or more first nodes. If the nodes have network programmable devices or similar capabilities, the requests may be forwarded to the nodes in order without having to return back to orchestrators between the nodes.
Owner:NVIDIA CORP

Distributed reasoning method and electronic device

The application relates to the computer technical field, in particular to a distributed reasoning method and an electronic device, the method comprises the following steps: performing parallel processing on an input sequence of a current reasoning request to obtain an initial key-value cache and an initial hidden state, then obtaining a first hidden state and expert distribution metadata, and performing matrix multiplication on the first hidden state based on the expert distribution metadata to generate a second hidden state; updating the initial key-value cache based on the second hidden state to obtain an updated key-value cache, taking the updated key-value cache as the initial key-value cache, taking the second hidden state as the initial hidden state, and re-executing the above steps until a preset reasoning end condition is reached to obtain a final reasoning result. Therefore, the problem of resource mismatch and low utilization caused by the isomorphism deployment of the attention module and the feedforward network module in the related art is solved, the hardware utilization is improved, and the reasoning cost is reduced.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Distributed reasoning acceleration method based on chained structure model on edge device

The invention belongs to the technical field of computers, and provides a distributed reasoning acceleration method based on a chained structure model on edge equipment, which comprises the following steps of: firstly, performing hierarchical analysis on the model, dynamically and uniformly dividing the model into a plurality of sub-models according to a calculated amount and a dependency relationship, and deploying the sub-models to an edge cluster, the minimum input data range is calculated through coordinate mapping for a convolutional layer, local dependence cutting is achieved, parallel decomposition is conducted on a linear layer through matrix partitioning and sparsification, a model is further divided into a plurality of calculation stages, in-stage parallel execution and inter-stage serial execution, and calculation, repetition and communication overhead are balanced through an optimization function, so that the calculation efficiency is improved. And finally, abstracting the sub-models into computational graph nodes, and realizing multi-thread task scheduling and pipeline data transmission by adopting an Actor process mechanism. According to the method, efficient distributed reasoning can be realized on the edge cluster, the end-to-end delay is remarkably reduced, the throughput rate and the resource utilization rate are improved, and meanwhile, redundant calculation and communication overhead are reduced.
Owner:HOHAI UNIV

Fractional processing capacity allocation based on estimated load

A method (1100) for artificial intelligence (AI) inference workload allocation includes receiving (1102), at a computing device (124) of a distributed AI inference platform (108), an estimated hint load (132) and an estimated generated load (134) of an AI inference workload (116) to be implemented by a processing unit (104) of a computing node (100) of the distributed AI inference platform (108). Based at least in part on the estimated cue load (132) and the estimated generated load (134), an inference unit (IU) processing load (136) is estimated, the IU processing load (136) to be applied to the processing unit (104) when implementing the AI inference workload (116). Based at least in part on the IU processing load (136), a fractional processing capacity (138A) of the processing unit (104) is allocated for implementing the AI inference workload (116).
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Distributed cluster construction method, distributed inference method and resource scheduler

This application provides a distributed cluster construction method, a distributed inference method, and a resource scheduler, which can be applied to the field of computer technology. The distributed cluster construction method includes: in response to a commit operation on a system orchestration file, binding virtual nodes created based on the system orchestration file to physical nodes equipped with hardware accelerators; the system orchestration file indicates the desired construction state for the distributed cluster; invoking multiple physical nodes to pull a target inference image from an image repository, and using the target inference image to launch container instances on each of the multiple physical nodes; the target inference image is obtained by encapsulating the hardware accelerator development kit and distributed inference framework of the physical nodes; invoking the accelerator plugins of the physical nodes to allocate the resources of the hardware accelerators on the physical nodes to the container instances; and constructing a distributed cluster based on the multiple container instances.
Owner:BEIJING KECHENG TECH DEV CO LTD

Communication method and device

A communication method and apparatus in a distributed reasoning scenario, taking the method applied to a first device as an example, in the method, the first device receives first indication information, the first indication information indicates execution of a first sub-task, the first sub-task is a sub-task of a reasoning task, the first device further receives second indication information, and the second indication information indicates execution of a second sub-task of the reasoning task. The second indication information indicates that the first subtask is completed, and then the first device skips execution of the first subtask. By adopting the method, when the first device is allocated to execute the first sub-task but notified that the first sub-task is completed, the first device can skip execution of the sub-task, so that the time delay of the reasoning task can be reduced, and the utilization efficiency of the computing power of each distributed reasoning participant can be improved.
Owner:HUAWEI TECH CO LTD

A method, medium, and system for network threat assessment based on dynamic Bayesian networks

This invention provides a network threat assessment method, medium, and system based on dynamic Bayesian networks, belonging to the field of network security technology. The invention uses a parameter adaptive learning unit with a sliding time window mechanism and a Bayesian online change point detection algorithm to update the conditional probability table parameters. It inputs the topology and parameters into a threat temporal inference model that integrates a liquid neuron layer and a hypergraph matching layer, receives real-time security event evidence, and outputs the posterior probability distribution of nodes. Based on the posterior probability distribution, it calculates the neuron activation regulation function value and dynamically adjusts the liquid neuron parameters. When the posterior probability of a threat node exceeds a preset threat threshold, it executes an iterative belief propagation algorithm through a hierarchical distributed inference collaboration module to generate an attack path probability graph and output the network security assessment result. This solves the technical problem of not being able to simultaneously and accurately model the causal dependencies and high-order collaborative modes of network threat events.
Owner:WUZHONG POWER SUPPLY COMPANY STATE GRID NINGXIA ELECTRIC POWER

Method and system for multi-device load balancing of mixed expert large language model inference

The application discloses a multi-device load balancing method and system for mixed expert large language model reasoning, wherein the multi-device load balancing method for mixed expert large language model reasoning comprises the following steps: constructing a reference data set of expert activation features to form expert activation frequency distribution data covering a plurality of preset task scenarios; performing initial grouping on all experts based on the number of a target device cluster and a preset parallel degree to form an initial storage allocation scheme; calculating the initial delay of the target device cluster by using a quantitative evaluation mechanism for the initial grouping and the reference data set, wherein the factors considered by the quantitative evaluation mechanism include the calculation delay caused by the calculation time consumption of the activated experts; and adopting an iteration optimization strategy of two-by-two exchange of experts between groups to make the final delay reach a preset range and obtain a final expert storage allocation scheme. The application can significantly improve the utilization rate of computing resources and is helpful to the industrial application of distributed reasoning of mixed expert large language models.
Owner:SEMICON TECH INNOVATION CENT(BEIJING) CORP +1

Transmission data gradient sparse method based on distributed reasoning

InactiveCN121457534AInference methodsNeural learning methodsData packSparse methods
The invention relates to the technical field of data gradient sparseness, in particular to a transmission data gradient sparseness method based on distributed reasoning, which comprises the following steps: confirming an original neural network model of a distributed computational node, performing gradient iterative computation on the original neural network model to obtain a plurality of model update gradient value sets, and performing gradient value aggregation on the plurality of model updating gradient value sets based on the to-be-trained model to obtain a local updating gradient value set, performing data packaging on the local updating gradient value set to obtain computational node gradient data, and updating the to-be-trained model in the reasoning main node by using the computational node gradient data to obtain a global updating model. According to the method, the training efficiency and the resource utilization rate of the global model in federated learning can be improved, and the communication efficiency between the main node and the computing node is improved.
Owner:ENTERPRISE ONLINE (BEIJING) NETWORK CO LTD