Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

396 results about "Computation graph" patented technology

In simple terms, a computation graph is a DAG in which nodes represent variables (tensors, matrix, scalars, etc.) and edge represent some mathematical operations (for example, summation, multiplication). The computation graph has some leaf variables.

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Heterogeneous computing power cooperative scheduling system and method for mixed precision training

The invention discloses a heterogeneous computing power cooperative scheduling system and method for mixed precision training, and belongs to the technical field of artificial intelligence computing. The system comprises a computational graph analysis and operator portrait module which is used for analyzing and dividing a model computational graph and extracting operator features; the heterogeneous hardware capability sensing and matching module is used for managing performance files and real-time states of heterogeneous hardware in the cluster and matching optimal execution hardware for each calculation partition; and the data flow coordination and pipeline parallel controller is used for generating a global execution plan, managing cross-device data dependence and communication and calculating overlapping optimization execution efficiency through communication. According to the method, the problem of low scheduling efficiency of mixed precision training in a heterogeneous environment is solved, automatic and accurate mapping from a calculation task to heterogeneous hardware is realized, the training speed is remarkably improved, the training cost is reduced, and the overall resource utilization rate of a cluster is improved.
Owner:HANHOU (BEIJING) TECH CO LTD

End side malicious behavior security detection method

The invention relates to an end-side malicious behavior security detection method, which comprises the following steps: acquiring multi-source data, carrying out standardization processing on the multi-source data, generating a unified quintuple sequence, and constructing a quaternary dynamic behavior graph; respectively extracting time sequence features, semantic features and context features by utilizing a quintuple sequence, calculating confidence coefficients of cross-modal association, mapping the confidence coefficients to a quaternary dynamic behavior graph, and generating an individual behavior graph, a current behavior graph and a group baseline graph; and performing projection processing on the individual behavior graph, the current behavior graph and the group baseline graph to calculate a deviation degree score and a structural anomaly contribution degree of graph-level comprehensive deviation characterization, and triggering a grading response strategy by taking the deviation degree score and the structural anomaly contribution degree as double indexes of safety detection to solve an anomaly root. According to the method, internal and external threats, especially abnormal behaviors and potential risks which are difficult to find by traditional security defense means, are detected by analyzing behavior modes of users and entities.
Owner:SHAOGUAN DATA IND RESEARCH INSTITUTE

Heterogeneous computing power adaptive compiling method and system for large model

The invention provides a large-model-oriented heterogeneous computing power adaptive compiling method, which comprises the following steps that: a user inputs a trained large model through a system interface, and a system front-end conversion module analyzes a computational graph of the model and converts the computational graph into an intermediate representation based on a unified operator description language (UDL); the system hardware sensing module automatically detects and extracts hardware feature fingerprints of at least one piece of target hardware; based on the unified operator description language UDL intermediate representation and the hardware feature fingerprint, automatically generating an optimization adaptation rule oriented to at least one piece of target hardware; wherein the basis of parameterized filling comprises specific parameters of hardware feature fingerprints and optimized attribute tags carried in an intermediate representation of a unified operator description language (UDL); and generating and deploying multiple back-end codes. The method has the beneficial effects that intelligent compiling based on hardware features can be realized, and the deployment efficiency and the operation performance of a large model in a complex heterogeneous computing power cluster are remarkably improved.
Owner:SHENZHEN XINGSHENG DIGITAL TECH CO LTD

Smart home distributed heterogeneous computing power collaborative reasoning method

The invention relates to the technical field of smart home distributed heterogeneous computing power cooperative reasoning methods, and particularly discloses a smart home distributed heterogeneous computing power cooperative reasoning method. The objective of the invention is to solve the problems that heterogeneous equipment in an existing smart home system is uneven in computing power utilization, depends on a central scheduling node, is difficult to ensure privacy security and is insufficient in dynamic adaptability. The method comprises the following steps: constructing a dynamic computing power portrait of household equipment and updating the dynamic computing power portrait in real time; performing semantic analysis and computational graph decomposition on the intelligent reasoning task to generate a fine-grained reasoning unit with computing power and delay constraints; based on a decentralized broadcast protocol and a weighted Hungary algorithm, the reasoning unit is matched to the optimal local device; executing cross-device collaborative reasoning through an encrypted publishing-subscribing mechanism; and finally aggregating a result and returning a log for optimization. According to the technical scheme, efficient cooperation of family heterogeneous computing power can be achieved, end-side reasoning delay is lower than 200 milliseconds, the comprehensive utilization rate of resources exceeds 85%, and meanwhile data privacy and system robustness are guaranteed.
Owner:NINGXIA HUIWAN NETWORK TECH CO LTD

Dynamic operator-oriented incremental compiling optimization method and system

The invention discloses a dynamic operator-oriented incremental compilation optimization method and system, belongs to the technical field of deep learning, and aims to solve the technical problem of how to provide a fine-grained, reusable and extensible compilation strategy for dynamic operator-oriented incremental compilation optimization. Comprising the following steps: extracting attributes of a dynamic operator and constructing a dynamic operator hash signature according to the extracted attributes; dividing the calculation graph into a static sub-graph and a dynamic sub-graph, establishing a mapping table of a dynamic operator hash signature and a compiling kernel, and caching the mapping table in an incremental compiling cache; solving and instantiating the dynamic sub-graph through a dynamic specifying engine, generating a specialized sub-graph, querying whether a matched compiling result exists in an incremental cache or not, and updating the incremental cache; and for the static sub-graph which is not changed, directly multiplexing the last compiling result of the static sub-graph, executing cross-sub-graph fusion optimization work between the dynamic sub-graph and the static sub-graph which are subjected to specialized processing, and submitting a fusion execution plan to a runtime engine.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Inference method and device, equipment and storage medium

The invention provides an inference method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, large language models and the like. The method comprises the steps of determining target metadata of a target operator through a predefined function interface in response to an identified extended dependency package corresponding to the reference target operator in an inference framework; a calculation graph is generated according to the target metadata and native metadata, and the native metadata is metadata of other operators except the target operator in the reasoning framework; and reasoning on the target hardware according to the input data and the calculation graph to obtain a reasoning result. According to the method, the operator and the frame are decoupled, meanwhile, the fracture condition of the computational graph is avoided, and the compatibility and reasoning efficiency of the reasoning frame are improved.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Space-time constraint end-to-end automatic driving system and method based on unified VLA model

The invention relates to the technical field of artificial intelligence and automatic driving, and discloses a space-time constraint end-to-end automatic driving system and method based on a unified VLA model, and the system comprises a unified encoder, a computing power management unit, a core neural network and an output decoder. A gating operator embedded into a core neural network is used for calculating characteristic information entropy and computing power consumption in real time, and through topology reconstruction logic based on entropy and computing power joint constraint, computational graph collapse operation is dynamically triggered at an instruction level, so that network reasoning depth is adaptively expanded along with physical time constraint, and the network reasoning accuracy is improved. According to the method, isomorphic mapping of the topological structure of the calculation model and the physical time constraint is established, the deterministic delay upper limit under the extreme working condition is determined, the problem of dimension mismatching of an existing static calculation graph under the hard real-time constraint is solved, and optimal bit allocation of calculation resources in the time dimension is achieved.
Owner:GUANGZHOU SMART BODY TECH CO LTD

OpenVX framework system for NPU acceleration

The invention provides an OpenVX framework system for NPU acceleration, which comprises an application layer, an OpenVX framework layer, an NPU runtime layer and an NPU hardware layer, and is characterized in that the OpenVX framework layer comprises an NPU perception graph optimizer. According to the invention, a set of OpenVX framework for completing the visual task on the NPU acceleration chip is designed, the calculation graph of the visual task constructed by the standard OpenVX API can be optimized according to the hardware characteristics of the NPU, and the memory management of the hardware layer of the NPU is designed, so that the visual task can be completed on the NPU in an accelerated manner, and the calculation efficiency of the visual task is improved.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

End-to-end multi-task learning method and system based on Swin Transform

The invention discloses an end-to-end multi-task learning method and system based on Swin Transform, and belongs to the technical field of artificial intelligence and computer vision crossing, and the method comprises the steps: collecting low-quality paper document images which have physical degradation characteristics and cause substantial obstacles to information recognition, carrying out the corresponding text labeling of each image, and constructing a target data set; based on the target data set, adopting a backbone network capable of extracting multi-scale hierarchical features as a shared image encoder, and taking an image enhancement task and a text recognition task as two parallel downstream branches to construct an end-to-end multi-task learning network architecture; based on the output of the image enhancement task and the text recognition task, image enhancement loss and text recognition loss are calculated respectively, a joint loss function is constructed, and multi-task collaborative optimization is realized through end-to-end training; the robust recognition capability of fuzzy, continuous and low-quality handwritten characters is improved, and a recognition-friendly high-quality image is generated.
Owner:INFORMATION CENT OF YUNNAN POWER GRID CO LTD

Hardware-aware ONNX graph optimization method and hardware-aware graph optimization and compiling engine

The invention relates to the field of artificial intelligence, and provides a hardware-aware ONNX graph optimization method and a hardware-aware graph optimization and compilation engine, which break through the barrier between hardware-independent optimization and hardware-related compilation, and improve the optimization efficiency through a unified and closed-loop optimization framework. An original ONNX model calculation graph is optimized and compiled based on hardware characteristics of an AI chip, deep collaborative optimization from the ONNX calculation graph to AI chip codes is achieved, and therefore the performance of the AI chip is played to the maximum extent.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Multimodal large language model-oriented compiling method and system for generating hardware accelerator executable codes, terminal and storage medium

The invention relates to the technical field of data processing, and discloses a compiling method and system for generating hardware accelerator executable codes facing a multi-mode large language model, a terminal and a storage medium. The method comprises the following steps: constructing a computational graph according to the structure of a to-be-deployed large language model and operators supported by a hardware platform; traversing the computational graph to deduce output tensor type information of each operator, and verifying the type consistency between adjacent operators; based on the verified type information, executing static memory allocation to determine a static storage address of the weight data, and executing dynamic memory allocation to determine a dynamic storage address of the state data; then combining the verified type information and the static and dynamic storage addresses to generate an intermediate representation instruction sequence for configuring a hardware acceleration platform register; and finally, converting the intermediate representation instruction sequence into a target code file which can be directly loaded and executed by a hardware acceleration platform. The compiling efficiency, the deployment reliability and the resource utilization rate are improved.
Owner:SHENZHEN MAITEXIN TECH CO LTD

Deep learning task resource allocation method and device, equipment and medium

The invention relates to a deep learning task resource allocation method and device, equipment and a medium. The method comprises the following steps: respectively analyzing a computational graph structure and a historical resource monitoring log corresponding to a deep learning task, and generating a tensor dependency graph and a resource use time sequence matrix; performing time-varying demand prediction on the basis of the matrix, generating a time-phased resource constraint table, performing memory allocation processing on the basis of a tensor dependency graph, and generating a tensor memory partitioning scheme and an inter-partition communication cost matrix; and performing static resource pre-allocation based on the time-phased resource constraint table and the tensor memory partitioning scheme, generating pre-allocated resource configuration, and performing resource scheduling and outputting real-time resource configuration through a deep reinforcement learning model according to the time-phased resource constraint table, the inter-partition communication cost matrix and the pre-allocated resource configuration. According to the method, by means of dynamic resource allocation, cross-partition communication cost optimization, reinforcement learning optimization and the like, the resource utilization rate and task execution efficiency of a deep learning task are remarkably improved.
Owner:FUZHOU IND & COMMERCIAL UNIV +1

Calculation graph memory layout automatic optimization method and system based on double-layer intermediate representation

The invention discloses a calculation graph memory layout automatic optimization method and system based on double-layer intermediate representation, and relates to the technical field of artificial intelligence model memory layout optimization. Aiming at the defect that the memory layout optimization effect in the existing AI compiler is limited, the scheme adopted by the invention comprises the following steps of: constructing a computational graph based on a given deep learning model; the calculation graph is converted into logic intermediate representation, and graph optimization is executed; converting the optimized logic intermediate representation into a physical intermediate representation, and adding a memory layout descriptor to each tensor; candidate execution configuration is generated for the whole computational graph through a memory optimizer, an optimal scheme is decided, and physical intermediate representation is reconstructed; and analyzing the reconstructed physical intermediate representation, automatically inserting a memory release operation after calculating the final use position of the tensor in the graph, and finally compiling the physical intermediate representation to generate an executable code of the target hardware. The method is used for realizing automatic optimization of the memory layout of the computational graph.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

System and computer-implemented method for preserving model confidentiality during graph optimizations

A system and method are described which provide a unique obfuscation mechanism for conducting performance optimization of deep neural network (DNN) computational graphs. The method obfuscates performance optimization in three steps. First, an obfuscation step where the original computation graph is obfuscated such that an adversary cannot feasibly identify the original model, thus providing confidentiality. Second, the optimization step is carried out flexibly and independently by the optimizer party on the obfuscated computational graph, providing performance speedups. Finally, the de-obfuscation step where the original model is retrieved by the model owner in its optimized form.
Owner:CENTML AI INC

Network security situation real-time evaluation method and system based on urban rail transit

The invention discloses a network security situation real-time evaluation method and system based on urban rail transit, and the method comprises the steps: collecting original data from a plurality of heterogeneous data sources of an urban rail transit network, carrying out the standardization processing, and generating a standardized security event; constructing a digital twin map; calculating the risk probability of nodes in the graph by using a time-space diagram neural network model, predicting an attack path and evaluating the service influence; a stream processing engine is adopted to consume a standardized security event in real time, and model reasoning, node risk probability updating, attack path prediction and service influence evaluation are performed on a local sub-graph related to the current event in the digital twin graph; and performing real-time parameter fine tuning on the space-time diagram neural network model through an online learning mechanism, and updating the space-time diagram neural network model through an incremental learning mechanism. According to the technical scheme, real-time dynamic prediction and active defense of urban rail network security threats are achieved, and the situation awareness capability and the autonomous protection level of key infrastructures are remarkably improved.
Owner:CASCO SIGNAL LTD

AI compiler-oriented optimization sensitive test system and method

The invention discloses an optimization sensitive test system and method for an AI compiler, and belongs to the technical field of software testing. The method mainly comprises the following steps: firstly, automatically extracting an IR sub-graph mode capable of triggering specific optimization Pass from an existing test of a compiler through dynamic instrumentation; secondly, through a strategy of multiplexing compatible nodes or creating new nodes, the optimization modes are intelligently synthesized into a diversified seed calculation graph context, and a large number of test cases capable of effectively triggering compiler optimization are generated; and finally, efficiently revealing the defects of the AI compiler in the optimization stage by taking crash detection and reasoning consistency comparison as test predictions. The defect that an existing random generation or grammar generation method is insufficient in test coverage rate in the optimization stage is overcome, the deep optimization defect in the AI compiler can be automatically and efficiently detected, and the method has the high error detection rate and good universality.
Owner:TIANJIN UNIV

Method for compiling computational graph, and related product

A method for compiling a computational graph, and a related product. The method comprises: acquiring a computational graph to be compiled that is expressed by a second intermediate representation, performing forward inference of a shape, and on the basis of the forward inference and tensor data splitting information, obtaining complete shape information; using the complete shape information to determine whether the tensor data splitting information needs to be adjusted; on the basis of a determination result, determining the tensor data splitting information that meets requirements; on the basis of the tensor data splitting information that meets the requirements, performing memory access pattern derivation on operators in the computational graph; on the basis of a derived memory access pattern, determining address-domain-related parameters of instructions involved in loops in code logic of the computational graph; performing pipeline scheduling on the instructions in the loops; and on the basis of the address-domain-related parameters of the instructions involved in the loops and a pipeline scheduling result of the instructions, compiling the computational graph, so as to obtain a binary file recognizable by an intelligent processor.
Owner:SHANGHAI CAMBRICON INFORMATION TECH CO LTD

Large model heterogeneous reasoning engine method, device and equipment and storage medium

The invention relates to a large-model heterogeneous reasoning engine method, device and equipment and a storage medium. The method comprises the following steps: acquiring an end side model generated by performing model conversion on a neural network model by a host end; constructing a calculation graph at the equipment end according to the end side model; a model reasoning task request is received, a task type is obtained, the computational graph is analyzed, differentiated task distribution is conducted on all nodes of the analyzed computational graph based on the task type, and all the nodes are bound with heterogeneous computing units executing different tasks; and according to a task allocation result, executing a model reasoning task through the heterogeneous computing unit bound with the allocated node. According to the method, the nodes are bound with the heterogeneous computing units executing different tasks, differentiated task distribution is performed on the nodes in the computing graph according to the task types, the computing power utilization rate of the heterogeneous computing units can be maximized when the end-side large model reasoning task is executed, the high throughput and low time delay requirements during reasoning task execution are ensured, and the reasoning efficiency of the end-side large model reasoning task is improved. The execution efficiency is improved.
Owner:FIBOCOM WIRELESS

Decentralized graph neural network architecture for beam forming of multi-cell multi-user MIMO (Multiple Input Multiple Output) communication system

The invention provides a decentralized graph neural network architecture for multi-cell multi-user MIMO communication system beam forming, and the construction steps of the architecture comprise: step 1, in a method for optimizing the multi-cell multi-user MIMO communication system beam forming by using a graph neural network, constructing a local loss function of a base station, and realizing computational graph decoupling; step 2, constructing a heterogeneous graph model based on the local loss function of the base station proposed in the step 1, and realizing decoupling of a graph model level; and step 3, carrying out distributed training and reasoning on the heterogeneous graph model to obtain a decentralized graph neural network architecture oriented to multi-cell multi-user MIMO communication system beam forming.
Owner:NANJING UNIV

Systems and methods of preconfiguring coherency protocol for computing systems

A multi-processor computing system (e.g., a system-on-chip) can store, in a shared memory, (i) a reservation table that is accessible by the one or more workload processors, and (ii) a scheduling program. The system can further execute the scheduling program to schedule execution of a set of workloads by one or more workload processors in accordance with an optimized compute graph, an optimized data positioning graph, and a coherence protocol that is precomputed based on the optimized compute graph and the optimized data positioning graph.
Owner:MERCEDES BENZ GROUP AG

Buffer area address allocation and SPILL scheduling method under multi-level cache architecture

The invention discloses a buffer area address allocation and SPILL scheduling method under a multi-level cache architecture, which comprises the following steps of: initializing a free block list for each cache type of an NPU (Network Processing Unit), allocating physical address offset for each buffer area in a computational graph, and executing free block merging and dynamic threshold adjustment; detecting a cache space distribution state, and screening an optimal SPILL victim from the candidate buffer area; reconstructing a dependency relationship between the SPILL operation and original buffer area nodes, and updating the calculation graph; an address multiplexing dependency graph is established and maintained, the life cycle overlapping condition of the buffer area is detected, and address multiplexing constraint and buffer area use time sequence constraint are maintained; and carrying out statistics on the extra data carrying amount brought by the SPILL operation, minimizing the data carrying amount except the total amount, and outputting a final cache allocation scheme and an SPILL operation set. According to the method, on the premise that all hardware constraints and execution constraints of the NPU of the SIMD architecture are met, the cache fragmentation degree is remarkably reduced.
Owner:GUIZHOU UNIV +2

Dynamic-length stateful tensor array

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for efficiently processing dynamic length tensors of a machine learning model represented by a computational graph. A program is received that specifies a dynamic, iterative computation that can be performed on input data for processing by a machine learning model. A directed computational graph representing the machine learning model is generated that specifies the dynamic, iterative computation as one or more operations using a tensor array object. Input is received for processing by the machine learning model and the directed computational graph representation of the machine learning model is executed with the received input to obtain output.
Owner:GOOGLE LLC

Ensemble communication unloading method, system, equipment and medium

ActiveCN121979690AResource allocationInference methodsCollective communicationComputer network
The invention discloses a set communication unloading method, system and device and a medium, and is applied to the technical field of computers, and the method comprises the steps: in a model deployment stage, a distributed reasoning controller generates a communication primitive blueprint based on a computational graph description file of a tensor parallel reasoning model and issues the communication primitive blueprint to a DPU; the DPU establishes a hardware-level communication context semantic environment based on the communication primitive blueprint; in the model reasoning stage, the GPU / NPU sends a trigger signal to the DPU when calculating to a communication boundary; and the DPU executes a DMA data pulling assembly line and an RDMA data sending assembly line in parallel based on a hardware-level communication context semantic environment, performs aggregation calculation on all tensor fragment data to be synchronized, and writes an aggregation calculation result back to the GPU / NPU. A hardware-level communication context semantic environment is established in advance to a model deployment stage, and double assembly lines are executed in the DPU in parallel, so that end-to-end communication delay is remarkably reduced, and zero participation of a host CPU is realized.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD +1

Cloud-edge collaborative map feature library and similar case retrieval method and system

The invention discloses a cloud-edge collaborative map feature library and similar case retrieval method and system, and belongs to the technical field of artificial intelligence and big data processing. According to the method, multi-modal feature extraction, weighted fusion and quantitative coding are completed at an edge end, and a local map is constructed; and the cloud carries out cross-version node alignment, fine-grained difference calculation, graph influence propagation and index increment updating, and establishes an audit verification closed loop. According to the method, the edge cloud features are consistent, the retrieval result is credible, the change is traceable, the communication load is effectively reduced, and the real-time performance and interpretability of the retrieval system are improved.
Owner:GUANGZHOU CITY UNIV OF TECH

Large model operation environment adaptive deployment method based on strategy optimization

The invention discloses a large model operating environment adaptive deployment method based on strategy optimization, and the method specifically comprises the steps: S1, collecting the computing power capacity, the video memory capacity, the communication bandwidth and the communication time delay of a computing node, and constructing an operating environment diagram structure; s2, constructing a node set and an edge set, configuring state feature vectors, and forming a topological calculation graph; s3, performing multi-scale filtering and persistent coherence calculation on the topology calculation graph to generate a topology gating vector; s4, generating a candidate calculation structure set under the limitation of the topology gating vector; s5, configuring different rank parameters for the weight tensor according to topology complexity distribution, and executing tensor decomposition; s6, performing tensor re-parameterization on the candidate calculation structure to form a large model calculation structure; and S7, mapping the large model calculation structure to a calculation node to complete deployment execution. According to the method, topological gating and heterogeneous rank tensor decomposition are introduced, and self-adaptive deployment of the structure constrained by the environment is achieved.
Owner:NANJING TECHN COLLEGE OF SPECIAL EDUCATION

Systems, methods, and devices for preventing credential passing attacks

PendingUS20260205484A1TicketData pack
A system and method for the detection and mitigation of Kerberos golden ticket, silver ticket, and related identity-based cyberattacks by passively monitoring and analyzing Kerberos and authentication operations within the network. The system and method provide real-time detections of identity attacks using time-series data and data pipelines, and by transforming the stateless Kerberos protocol into stateful protocol. A packet capturing agent is deployed on the network where captured time-series Kerberos and related event and log information is processed in distributed computational graph (DCG) stages where declarative rules determine if an attack is being carried out and what type of attack it is.
Owner:SENTINELONE INC

A video target recognition method based on multi-model hot switching

The application discloses a video target recognition method based on multi-model hot switching, and belongs to the technical field of computer vision and video processing. The method comprises an arbitration module, a switching control module and a feature adaptation and buffer module. The arbitration module generates a switching preparation signal by extracting multi-dimensional indexes such as optical flow mean, local variance, target density and scene confidence in real time through a lightweight channel independent of main reasoning. The switching control module performs atomic replacement of a computation graph pointer in a vertical blanking period, and realizes millisecond-level hot switching with zero frame loss by combining an asynchronous pre-copy and a chasing mechanism of a double buffer. The feature adaptation and buffer module solves tensor shape mismatch between heterogeneous models through a pre-compiled adaptation layer. The application also provides optimization schemes such as multi-index nonlinear fusion decision, zero-copy memory management, local slice focus reasoning and edge-cloud hierarchical unloading, significantly reduces switching delay, guarantees continuous recognition of a video stream, and is suitable for edge computing scenes with limited resources.
Owner:SICHUAN BAICHUAN SIWEI INFORMATION TECH CO LTD

Streaming based generative artificial intelligence (AI) workload execution

An apparatus and method for efficiently performing efficient data storage and data transfer of machine learning data. In various implementations, a host processing circuit of a computing system executes a machine learning (ML) application. The application includes a computational graph that indicates the computational order of the ML nodes, layers, and stages of the ML model. The host processing circuit translates function calls in the application to commands particular to an accelerator circuit. The accelerator circuit preloads weights to be used by ML nodes of the ML model by retrieving compressed weights from a storage device different from system memory. The accelerator circuit uses a streaming application programming interface (API) and bypasses the host processing circuit to retrieve the compressed weights from the storage device. The accelerator circuit decompresses the retrieved weights and executes the ML node using the decompressed weights.
Owner:ADVANCED MICRO DEVICES INC

Executing a compute graph on multiple reconfigurable dataflow processors

A method for a reconfigurable computing system includes receiving a compute graph for execution on multiple RDPs interconnected with a ring network having R interconnected RDPs. A compute graph with a node specifying a reduction operation for a first and second tensor is detected. Executing the compute graph on the multiple RDPs.
Owner:SAMBANOVA SYSTEMS INC