Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

321 results about "Network processor" patented technology

A network processor is an integrated circuit which has a feature set specifically targeted at the networking application domain. Network processors are typically software programmable devices and would have generic characteristics similar to general purpose central processing units that are commonly used in many different types of equipment and products.

Complex scene-oriented AI large model lightweight deployment method

The invention provides a complex scene-oriented AI large model lightweight deployment method, and relates to the technical field of edge computing, and the method comprises the steps: carrying out the structured pruning of a pre-trained Transform network based on the attention head importance score, carrying out the dynamic sparsification of the activation state of a feedforward network according to the input tensor entropy value, employing the dynamic mixing precision quantization, and carrying out the reconstruction of an AI large model. Obtaining network parameters after pruning quantization; deploying the pruned and quantized network parameters to an edge computing device, distributing a feature extraction operator to a neural network processor through a heterogeneous computing scheduler, and unloading a classification operator to a multi-core central processing unit; and managing an on-chip memory in combination with a virtual memory paging mechanism, realizing zero-copy data transmission by utilizing a direct memory access controller, and outputting a reasoning result tensor. According to the method, efficient and reliable operation of the large model at the resource-constrained edge node is realized.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD

Power grid dispatching strategy optimization method and system

The invention provides a power grid dispatching strategy optimization method and system, and the method comprises the steps: deploying a device based on a quantum entanglement technology between transformer substations, and collecting the state information of a power grid, and encrypting and transmitting the state information to a control center through a quantum network; a pulse neural network processor is used in a control center to extract spatial-temporal characteristics, and a quantum game theory model is used to generate an optimized scheduling strategy. Multi-modal verification data is collected, including visual deformation, abnormal sound monitoring, and operational resistance data. A genetic algorithm and deep reinforcement learning are utilized to optimize a decision model, and a dynamic weight distribution mechanism is included. And through a photon-quantum hybrid computing architecture execution model, a scheduling strategy is collaboratively optimized, and a result is fed back to a physical power grid. According to the method, the quantum technology and the spiking neural network are combined, new energy fluctuation is captured in real time, the multi-target weight is dynamically adjusted through the quantum game theory model, and the scheduling strategy robustness is improved. And a photon-quantum hybrid computing architecture is adopted, so that the feature extraction speed and the optimization efficiency are improved.
Owner:SICHUAN PROVINCE AIRPORT GRP CO LTD

Message uploading method, data processing unit and network processor

The invention relates to a message uploading method, a data processing unit and a network processor, and the method comprises the steps: obtaining the number of messages currently cached in the data processing unit in a target queue, so as to determine the number of descriptors to be read, obtaining the descriptors in batches through one PCIe reading operation, and sending the descriptors to the target queue; the messages of the currently cached target queue are written into host side caches pointed by the descriptors, and every time one message is successfully uploaded, one is subtracted from the message count of the currently cached target queue; and if the currently read descriptor is used up, but the target queue has unprocessed messages, repeating the above steps, determining and reading the required descriptor according to the number of the remaining messages of the currently cached target queue again, continuing to write the messages, and repeating the above steps. And all the messages of the currently cached target queue are successfully carried to the host side for caching. According to the method and the device, the message uploading efficiency and the bandwidth utilization rate are improved by optimizing the descriptor reading mode.
Owner:SHENZHEN JAGUAR MICROSYSTEMS CO LTD

Neural network processor based on SIMT and task execution method thereof

The invention provides an SIMT-based neural network processor and a task execution method thereof, and the method comprises the steps that a general processor queries a state register of a coprocessor, and the state register stores the resource condition of the coprocessor; the universal processor completes splitting from a thread block to a thread bundle according to the resource condition, and the universal processor forwards a thread bundle instruction to a thread bundle distributor of the coprocessor; and the thread beam distributor decodes instructions in the thread beams in sequence and schedules the instructions to instruction queues of different calculation cores, and the thread beams sequentially perform thread beam scheduling, instruction emission and instruction execution according to the sequence, so that all calculations of a neural network task are completed in parallel, and a running result of the neural network task is obtained. According to the method, a general processor for task splitting is introduced in front of the thread beam scheduler or a specific compiler is directly used, so that dynamic scheduling of the threads can be realized, and a scheme for dynamically expanding the number of the threads according to data precision becomes a feasible architecture option.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Heterogeneous trusted execution environment architecture construction method and device and processor

The invention relates to a heterogeneous trusted execution environment architecture construction method and device, and the method comprises the following steps: constructing a tensor analyzer which is located in a memory control unit of a central processing unit of a heterogeneous trusted execution environment architecture; and providing memory protection with uniform tensor granularity for data interaction between the neural network processor of the heterogeneous trusted execution environment architecture and the central processing unit through the tensor analyzer.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Coagulant adding control system and method based on multi-scale alumen ustum characteristic analysis

The invention provides a coagulant addition control system and method based on multi-scale alumen ustum characteristic analysis, and relates to the technical field of coagulant addition control. Comprising an underwater camera device, an embedded prediction module, an integrated neural network processor NPU, a dosing device, a feature extraction module and a time sequence feature coding module. According to the method, underwater camera devices are arranged in a flocculation area and a settlement area respectively, spatial multi-dimensional feature analysis is carried out on images acquired at double view angles, spatial multi-scale parameters are obtained, and real-time prediction of coagulant dosage is realized through an improved deep learning model; the improved deep learning model is embedded into a controller of an integrated neural network processor NPU, after the controller receives a model prediction result, regulation and control signals of PAC and PAM are generated in combination with the current operation state of the dosing pump, and dosing equipment is driven to execute actions; meanwhile, a regulation and control result is fed back to the feature extraction module in real time, and accurate matching of the coagulant dosage and dynamic evolution of alumen ustum is ensured.
Owner:NORTHEASTERN UNIV CHINA

Intelligent control method, device, equipment and system for automatic stamping

The invention relates to the technical field of office automation, in particular to an intelligent control method, device, equipment and system for automatic stamping, and aims to recognize and semantically analyze the recognized file content of a paper material through a multi-modal stamping feature recognition model, realize intelligent determination and calibration of a stamping position, improve the stamping accuracy and improve the stamping efficiency. The stamping method can adapt to different file formats and is flexible without presetting a fixed template, so that the applicability of the stamping method is improved; edge AI computing power is deployed in the stamping equipment, and a neural network processor is locally integrated, so that data processing localization is realized, data security is improved, data response delay in the stamping processing process is reduced, and network dependence is reduced; and based on local powerful computing power support of the stamping equipment, large-batch stamping processing can be achieved, the stamping efficiency is improved, and therefore the working efficiency is improved, and the human resource cost is reduced.
Owner:GUANGDONG PLANNING & DESIGNING INST OF TELECOMM

Code compiling method and device based on P4 assembler, assembler and medium

The invention provides a code compiling method and device based on a P4 assembler, the assembler and a medium. The code compiling method comprises the steps that a to-be-processed assembly code file is obtained; splitting and compiling the assembly code file to obtain a plurality of code segment compiling results; according to the auxiliary information, by taking Parser components, Mat components and Deparser components contained in the Ingress unit and the Egress unit as classification types, classifying the code segment compiling results to obtain code segment groups; linking the P4 table item data and the machine instruction code contained in the code segment group, and endowing a link result with a global instruction identifier and code segment header information to obtain a code segment binary code; and generating a firmware file based on the code segment binary code. The compiling process of the network processor can be simplified, so that the development efficiency is improved.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD

Image perspective conversion display apparatus and method, and storage medium

Provided in the present invention are an image perspective conversion display apparatus, an image perspective conversion display method, and a computer-readable storage medium. The image perspective conversion display apparatus comprises: a digital signal processor, which is configured to perform warping processing on a first video image from a camera's perspective and depth data corresponding thereto, so as to obtain a second video image from a user's human eye perspective; a network processor, which is configured to perform target detection on the first video image so as to determine a first image mask indicating an occluded region in the second video image, and which is configured to perform an inpainting operation on the basis of the second video image and the first image mask so as to obtain an inpainted image that fills the value of at least one pixel in the occluded region; and an application-specific integrated circuit, which is configured to perform correction and layer blending on the first video image and the inpainted image so as to obtain a third video image to be displayed.
Owner:GRAVITYXR ELECTRONICS & TECH CO LTD

Quantization and inverse quantization method in large language model and neural network processor

The invention relates to a quantization method, an inverse quantization method and a neural network processor in a large language model. A core (namely a quantization unit) in a neural network processor is arranged to execute quantization processing in a large-scale language model reasoning process, and the quantization unit has a data partitioning function, so that the quantization processing efficiency is greatly improved compared with an existing quantization processing core which is limited by the data size (such as block wise) when receiving data to be quantized. And frequent interaction with a cache is not needed, so that online high-efficiency large-model quantification processing is realized. And performing an inverse quantization process in a large language model inference process by setting a core (i.e., an inverse quantization unit) in a neural network processor, and the inverse quantization unit having a data partitioning function, a matrix multiplication function, and a high precision accumulation function (e.g., multiplying by a corresponding quantization parameter), the inverse quantization processing does not need to frequently carry data in a plurality of cores, and the inverse quantization processing efficiency of a large model is improved.
Owner:北京凌川科技有限公司

Method and apparatus for executing instruction by neural network processor, storage medium, and device

Disclosed are a data access method and apparatus, a storage medium, and an electronic device. The method includes: determining a first instruction to be executed by the neural network processor; determining, based on requested storage space information corresponding to the first instruction, storage space address information meeting the requested storage space information, the storage space address information comprising first storage space address information of a first storage space and second storage space address information of a second storage space; and executing, based on the storage space address information, the first instruction through the neural network processor.
Owner:BEIJING HORIZON INFORMATION TECH CO LTD

Flexible and scalable thermal test vehicle design for electronics cooling solutions

The density and power consumption of modern integrated circuits, such as Graphic Processing Units (GPUs), Central Processing Units (CPUs), and Network Processing Units (NPUs) is growing rapidly, which necessitates designing advanced cooling systems. Existing solutions for characterizing and validating these cooling system are inadequate. A flexible, scalable Thermal Test Vehicle (TTV) is disclosed which is based on an array of power transistors, measurement / control circuitry, and onboard computer. The TTV is configured for characterizing the performance of electronic cooling solutions under a variety of operating conditions.
Owner:RGT UNIV OF CALIFORNIA +1

Neural network processing based on subgraph recognition

Systems and methods for providing executable instructions to a neural network processor are provided. In one example, a system comprises a database that stores a plurality of executable instructions and a plurality of subgraph identifiers, each subgraph identifier of the plurality of subgraph identifiers being associated with a subset of instructions of the plurality of executable instructions. The system further includes a compiler configured to: identify a computational subgraph from a computational graph of a neural network model; compute a subgraph identifier for the computational subgraph, based on whether the subgraph identifier is included in the plurality of subgraph identifiers, either: obtain, from the database, first instructions associated with the subgraph identifier; or generate second instructions representing the computational subgraph; and provide the first instructions or the second instructions for execution by a neural network processor to perform computation operations for the neural network model.
Owner:AMAZON TECH INC

Model task scheduler, computing core, neural network processor and method

The invention provides a model task scheduler, a computing core, a neural network processor and a method, and the model task scheduler comprises a model task queue unit which is configured to receive model tasks issued by a host and cache the model tasks to corresponding task queues according to task priorities, outputting the model tasks in the task queue to a model task scheduling unit according to task priorities; wherein the model task queue unit comprises a plurality of task queues, and each task queue corresponds to one task priority; and the model task scheduling unit is configured to distribute the model tasks to one or more computing cores corresponding to the task types in the neural network processor according to the task types of the model tasks output by the model task queue unit.
Owner:SANECHIPS TECH CO LTD

Software and hardware collaborative optimization method of hybrid in-memory architecture

The invention discloses a software and hardware collaborative optimization method for a hybrid in-memory architecture, which comprises the following steps of: performing joint feature representation on a target AI algorithm and an in-memory computing architecture, extracting context features, and parameterizing in-memory computing unit configuration, a neural network processor assembly line, a multi-core interconnection topology and a storage level interface of the in-memory computing architecture; an off-line reference data set is constructed, a predictive agent model is trained, discrete architecture parameters are processed by the model by adopting an embedding method, feature association is learned by applying an encoder with a self-attention mechanism, joint prediction of multi-dimensional PPA indexes is realized through a parallel prediction network, and a feasible region constraint learning mechanism is introduced in training; the trained agent model is embedded into a multi-objective evolutionary algorithm, the energy efficiency ratio, the average computing power utilization rate, the model execution delay and the like serve as optimization objectives, chip area efficiency and power consumption constraints are met at the same time, and a Pareto optimal in-memory computing architecture configuration set is searched.
Owner:SOUTH CHINA UNIV OF TECH

Tensor data acceleration processing device, tensor data acceleration processing system and integrated circuit

The invention discloses a tensor data acceleration processing device, a tensor data acceleration processing system and an integrated circuit. The device is used for post-processing tensor data output by a neural network processor and comprises an interface unit, a cache unit, a data processing unit and an output unit. And the interface unit is connected with the neural network processor and is used for transmitting the to-be-processed tensor data output by the neural network processor. The cache unit is connected with the interface unit and is used for storing tensor data to be processed in processing; and the data processing unit is connected with the interface unit and the cache unit and is used for processing the tensor data to be processed. And the output unit is connected with the data processing unit and transmits the processed tensor data to be processed to an external memory. According to the device, the tensor post-processing efficiency is improved, the time delay is reduced, and the bandwidth consumption in the data reading and writing process is saved.
Owner:XINXIN HANGTU (SUZHOU) TECHNOLOGY CO LTD

High-speed multi-port cache arbiter with variable word length

The invention discloses a high-speed multi-port cache arbiter with a variable word length. Along with the high-speed development of the modern network technology, storage management and scheduling take up most of time in the processing of data packets by network equipment, so that the low-speed caching capability of most memories becomes a bottleneck for limiting the further improvement of a network processor. Therefore, by managing the shared cache of not less than 32 blocks of 256Kbit SRAM units and supporting simultaneous cache writing of a plurality of ports through an arbitration mechanism, write scheduling and read scheduling can be realized without mutual influence, and concurrent read-write conflicts are avoided; meanwhile, each port supports eight priority queues, so that the data transmission bandwidth of each port can reach 1G bps; by processing a data packet, caching and data scheduling according to packets are supported, the length of the data packet is 64-1024 bytes variable, and the influence of the length of the data packet on memory resources can be avoided; and the purpose of storing resources can be achieved by dynamically adjusting space saving through memory recovery, and the data storage efficiency is improved.
Owner:NANJING UNIV

Reconfigurable multiply-accumulate unit for neural network processor and processor

According to the reconfigurable multiply-accumulate unit for the neural network processor and the processor, an efficient reconfigurable multiplier is constructed by organizing a large number of low-precision multiplication units. A peripheral reconfigurable computing circuit is designed to form a multiply-accumulate unit, so that the multiply-accumulate unit can flexibly support computing requirements of various precisions, various computing types and various data paths, and the compatibility and efficiency of hardware to a neural network are remarkably improved.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Deep Learning for Four-Dimensional (4D) Modeling of Glioblastoma Multiforme with Tumor Treating Fields (TTFields) Therapy

PendingUS20250299340A1Image enhancementImage analysisGlioblastomaFollow up examination
The technology disclosed relates to deep learning for four-dimensional (4D) modeling of glioblastoma multiforme with tumor treating fields (TTFields) therapy. In particular, the technology disclosed relates to a system comprising memory and a neural network processor. The memory stores input image data characterizing a current spatial distribution of glioblastoma multiforme (GBM). The current spatial distribution of the GBM is detected at a precursor examination of a patient receiving tumor treating fields (TTFields) therapy. The neural network processor, is in communication with the memory, and is configured to cause a neural network to process the input image data and, in response, generate output probability data characterizing a future spatial distribution of the GBM at a follow-up examination of the patient receiving the TTFields therapy. The neural network determines the future spatial distribution based in part on a time interval between the precursor examination and the follow-up examination.
Owner:RGT UNIV OF CALIFORNIA

Configurable decompression circuit supporting COO and Bitmap compression algorithms

The invention relates to a configurable decompression circuit supporting COO and Bi tmap compression algorithms, and belongs to the field of integrated circuits. According to the hybrid compression method suitable for the circuit, after sparseness analysis is carried out on weight matrixes of all layers of a neural network, COO compression based on a coordinate type sparse matrix or Bitmap compression based on bitmap masks is selected according to the sparseness characteristics of different layers, so that the storage space is optimized; comprising a control module, a first selector, a second selector, a bitmap description memory, a coordinate index memory, a numerical memory, a converter, a COO decoder and a decompression data storage module, decoding of two compression formats of COO and Bi tmap can be supported, and the operation efficiency and flexibility of a neural network processor are improved. The method effectively reduces the weight storage demand of the neural network, enhances the adaptability of the compression algorithm, and is suitable for hardware implementation of an efficient neural network model.
Owner:BEIJING MXTRONICS CORP +1

Deep learning model reasoning acceleration method and system based on heterogeneous computing architecture

The invention relates to the technical field of deep learning reasoning acceleration, and discloses a deep learning model reasoning acceleration method and system based on a heterogeneous computing architecture. The method comprises the steps of obtaining a calculation intensity index by analyzing a calculation operation type of a model layer, and dividing the model into a front calculation area, a dynamic calculation area and a rear calculation area according to the calculation intensity index. And allocating the calculation-intensive and regular front and rear areas to a neural network processor, and allocating the dynamic sparse area to a central processing unit. The neural network processor executes preposition calculation, generates a preposition activation value vector and transmits the preposition activation value vector to the central processing unit; after completing dynamic calculation, the central processing unit transmits a generated dynamic activation value vector back to the neural network processor; and the neural network processor executes post calculation and outputs a result. According to the method, the utilization efficiency of heterogeneous hardware resources is improved and the overall reasoning time delay is reduced through calculation feature guided fine task division and cooperative scheduling between processors.
Owner:STORAGEX TECH INC

Dynamic slimmable neural network for sequential data processing

There is disclosed an apparatus (100, 200, 300, 400, 500) comprising: an input interface (101) to receive an input segment (102) of an input sequential data; a neural network, NN, processor (110) to derive an output result (104) by processing the input segment (102) through a NN having a number of layers (111) from a first layer to a last layer, the NN using a predetermined number of deactivatable units which are selectively deactivatable; a gating module (120) configured to deactivate at least one deactivatable unit based on the input segment (102) and / or on at least one intermediate output segment (112, 123), an output interface (103) configured to provide an output (104) derived from the last layer of the NN.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Apparatus, method and computer program for processing an audio signal using feature segmentation and feature combination

An apparatus for processing an information signal has: a feature extractor for extracting a set of features having a first dimension; a feature segmenter for segmenting into a first subset having a second dimension and a second subset having a third dimension, which overlap, both being lower than the first dimension; a neural network processor for processing the first and second subsets using a first and a second neural network to obtain a first and a second result, respectively; a feature combiner for combining the first and second results using a third neural network, having a third complexity lower than a first or a second complexity of the first and second neural network to obtain a result set of features having a result dimension; and an output post-processor for post-processing the result set of features to obtain a processed information signal.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Apparatus for quick detection

An apparatus for quick detection of a blood pressure level is described. The apparatus comprises a module, a device and a network processor. The module is for capturing data of a heart sound. The device is for transforming the data into a plurality of time-frequency spectrograms having a plurality of image features. The network processor is trained to analyze the image features of the time-frequency spectrograms, for giving the blood pressure level. The network processor is equipped with a convolutional neural network trained to analyze the image features of the time-frequency spectrograms, by focusing on a second heart sound, for detecting the hypertension. The network processor is also trained to identify the time-frequency spectrograms within a frequency band, to distinguish the long-term hypertension and the exercise-induced transient hypertension.
Owner:CHEN TING-JU +2

Subtask storage for streaming convolutions in neural network processor

Embodiments relate to streaming convolution operations in a neural processor circuit that includes a neural engine circuit and a neural task manager. The neural task manager obtains multiple task descriptors and multiple subtask descriptors. Each task descriptor identifies a respective set of the convolution operations of a respective layer of a set of layers. Each subtask descriptor identifies a corresponding task descriptor and a subset of the convolution operations on a portion of a layer of the set of layers identified by the corresponding task descriptor. The neural processor circuit configures the neural engine circuit for execution of the subset of the convolution operations using the corresponding task descriptor. The neural engine circuit performs the subset of the convolution operations to generate output data that correspond to input data of another subset of the convolution operations identified by another subtask descriptor from the list of subtask descriptors.
Owner:APPLE INC

Photonic NPU embedded neural network processor

The invention discloses a photon NPU (Network Processing Unit) embedded neural network processor, which comprises (1) an electromagnetic wave bus, (2) a multiply-add module, (3) an activation function module, (4) a two-dimensional data operation module and (5) a decompression module, the NPU embedded neural network processor is divided into independent modules with different sizes according to functions and purposes, each functional module is provided with an independent input and output end, and each functional module is provided with an independent input end and an independent output end. All input and output ends are connected with a transmitting end and a receiving end of an electromagnetic wave bus, the bandwidth advantage of electromagnetic waves is utilized, a processing terminal is further formed through interconnection, and the electromagnetic wave bus is used for replacing a control bus, a data bus, an address bus and all replaceable circuits to transmit needed control signals, data and data addresses. An electromagnetic wave bus is used as a carrier of data needing to be processed, a data address and control signal transmission, and a photon (NPU) embedded neural network processor processes and outputs a result after passing through a radio frequency chip and a baseband chip.
Owner:刘国栋

Image processing method and chip

Embodiments of the present application provide an image processing method. The method comprises: an image signal processor in a chip processing data received from a camera to generate a target image; the image signal processor writing a first patch in the target image into a system cache in the chip; a neural network processor of the chip reading the first patch from the system cache; the neural network processor processing the first patch based on a neural network model to obtain a first output patch; and the neural network processor writing the first output patch into the system cache. The processing includes at least one of color interpolation or high dynamic range (HDR) processing. The technical solution provided by the present application saves the time-consuming of image processing and saves the system power consumption.
Owner:HUAWEI TECH CO LTD

Method of pruning weights in convolutional layer of neural network

A method includes: receiving N sets of weights of a convolutional layer of a neural network, each set of the weights having a same number of weights and corresponding to one of a sequence of output channels (OCs) of the convolutional layer; and performing a pruning process to prune M sets of the weights among the N sets of the weights such that each of the M sets of the weights has a same number of non-zero weights, M being smaller than or equal to N, M being equal to a number of active OCs to be processed in parallel in a neural network processor.
Owner:MEDIATEK INC

Server information collection, updating method and device, system, medium, product and equipment

This application discloses a method, apparatus, system, medium, product, and device for server information collection and updating. The method includes: acquiring a target image; loading the acquired target image into a target server, wherein the target server includes a neural network processor (NPU); and collecting hardware information of the target server through the loaded target image, wherein the hardware information includes the network interface card (NIC) information of the NPU. This improves the efficiency and accuracy of server hardware information collection, thereby increasing the construction efficiency and scalability of large-scale clusters and reducing the operational complexity of the cluster.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1