Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Network processor" patented technology

A network processor is an integrated circuit which has a feature set specifically targeted at the networking application domain. Network processors are typically software programmable devices and would have generic characteristics similar to general purpose central processing units that are commonly used in many different types of equipment and products.

A lightweight neural network processor storage architecture co-optimization method

ActiveCN121809563BComputer hardwareNetwork processor
The present application relates to neural network inference hardware technical field, particularly to a kind of lightweight neural network processor storage architecture collaborative optimization method, comprising: step 1: the difference of on-chip data on bandwidth demand and access mode is analyzed, and the differential single-port storage setting is carried out to the buffer storage of each data in neural network processor;Step 2: the GEMM module of neural network processor is based on semi-pulsating array setting, and ALU operation fusion processing module is used, to reduce the storage bandwidth and bit width of data;Step 3: the storage structure is optimized by introducing the dynamic adjustment strategy based on timing analysis, to output the optimized neural network processor.The present application guarantees the computing performance, significantly improves the area efficiency and access efficiency of storage system, realizes the collaborative optimization of bandwidth, area and power consumption.
Owner:ZHEJIANG UNIV

Neural network processor and processing element

A processor for executing neural network inference includes a CPU, weight and state memory, and a fabric comprising a plurality of programming elements. Each programming element comprises an ALU, having a controller and a core, and a sequencer configured to receive a sequencer program comprising at least one sequencer instruction, the at least one sequencer instruction comprising a compressed representation of one or more arithmetic ALU op codes; decompress the sequence instruction to produce the one or more ALU op codes; and provide the one or more ALU op codes thus produced to the ALU controller. The ALU controller is configured to output a control signal to configure a data path of the ALU for executing one or more operations in performing inference, for each op code received from the sequencer. The CPU controls execution of sequencer programs by the plurality of programming elements.
Owner:APPL BRAIN RES INC

Three-dimensional integrated and substrate interconnected in-memory neural network processor

The application relates to a memory-compute integrated neural network processor based on three-dimensional integration and substrate interconnection heterogeneous storage media and a control method thereof, which comprises an off-chip volatile storage chip connected with a logic chip and used for information interaction with the logic chip, an off-chip non-volatile storage chip connected with the logic chip and used for information interaction with the logic chip, and the logic chip configured to realize a combination of any of the following functions: executing a computing task, performing data caching and realizing information interaction; the off-chip volatile storage chip and the logic chip are integrated into a three-dimensional integrated module in a vertical direction by using a three-dimensional stacking technology; the three-dimensional integrated module and the off-chip non-volatile storage chip are integrated by using a substrate interconnection technology; in the three-dimensional integrated module, the off-chip volatile storage chip and the logic chip are bonded in a flip-chip bonding mode; and the high-speed memory access and high-bandwidth data interaction requirements of the off-chip volatile storage chip are met.
Owner:HANG ZHOU NANO CORE CHIP ELECTRONIC TECH CO LTD

Crossbar circuit for unaligned memory access in neural network processor

ActiveUS12675679B2Computer hardwareCrossbar switch
Embodiments of the present disclosure relate to an unaligned memory access in a neural processor circuit. The neural processor circuit includes a crossbar circuit and a neural engine circuit coupled to the crossbar circuit. During each operating cycle of the neural processor circuit, the crossbar circuit receives a portion of input data, and re-aligns or bypasses the portion of input data. The neural engine circuit receives at least a portion of the re-aligned or bypassed portion of the input data, and performs a convolution operation on the received portion of re-aligned or bypassed portion of input data to generate output data.
Owner:APPLE INC

Functional safety test circuits and methods, neural network processors and storage media

PendingJP2026086368ADetecting faulty hardware using neural networksDigital circuit testingCircuit under testSafety testing
This application discloses functional safety test circuits and methods, neural network processors, and storage media, relating to the field of functional safety technology. [Solution] The functional safety test circuit includes a configurator for generating and outputting configuration information for testing an integrated circuit under test; a plurality of test data generators for generating and outputting first test data corresponding to each of the plurality of test data generators based on the configuration information; an integrated circuit under test for processing the plurality of first test data and acquiring and outputting a plurality of second test data; and a first comparator for comparing the consistency of the plurality of second test data and acquiring a first test result of the integrated circuit under test.
Owner:BEIJING HORIZON INFORMATION TECH CO LTD

Method of transporting data, direct memory access device and computer system

Disclosed are a data moving method, a direct memory access device and a computer system. The data moving method is used for a neural network processor including at least one processing unit array, and the method includes: receiving a first instruction, wherein the first instruction indicates address information of target data to be moved, and the address information of the target data is obtained based on a mapping relationship between the target data and at least one processing unit in the processing unit array; generating a data moving request according to the address information of the target data; and moving the target data for the neural network processor according to the data moving request.
Owner:BEIJING ESWIN COMPUTING TECH CO LTD

Image rendering methods, apparatuses, electronic devices, storage media and software products

This application provides an image rendering method, apparatus, electronic device, storage medium, and program product, belonging to the field of image processing technology. The method includes: receiving input data; inputting the input data into an image rendering model on a neural network processor, and obtaining the next frame image of the input image output by the image rendering model; the image rendering model includes a diffusion model module and a diffusion forced control module; the diffusion forced control module is used to determine the noise level and the number of denoising steps after compression according to the rendering instructions; the diffusion model module is used to combine the input image with noise generated according to the noise level to form an intermediate state, and under the control of the diffusion forced control module, perform autoregressive denoising processing on the intermediate state based on the number of denoising steps after compression to obtain the next frame image. This application utilizes a neural network processor and a diffusion forced control module to improve the image generation speed and solves the problem of difficulty in balancing rendering performance, power consumption, and image quality.
Owner:ZHONGGUANCUN INSTITUTE OF ARTIFICIAL INTELLIGENCE

Image convolution calculation method and device, computer program product and terminal equipment

The invention belongs to the technical field of neural networks, and particularly relates to an image convolution calculation method and device, a computer program product and terminal equipment. The method comprises the following steps: quantizing a target initial weight of a target convolution kernel to obtain a target quantized weight; performing convolution kernel transformation on the target quantization weight by using a neural network processor to obtain a target transformation weight; and based on the target transformation weight, performing convolution calculation on the target image blocks by using a neural network processor to obtain a target convolution calculation result. According to the method, the quantization process of the weight can be set before the convolution kernel transformation, so that the convolution kernel transformation of the weight can be carried out under the condition that the precision of the weight is kept, and the transformation weight with relatively high precision is obtained; according to the transformation weight with high precision, accurate convolution calculation can be carried out on the image blocks of any initial pixel, and the availability of the Winograd algorithm can be improved.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD +1

Neural network apparatus, neural network processor, and method of operating neural network processor

A neural network processor and method include a fetch controller configured to receive input feature information, indicating whether each of a plurality of input features of an input feature map includes a non-zero value, and weight information, indicating whether each of a plurality of weights of a weight map includes a non-zero value, and configured to determine input features and weights to be convoluted, from among the plurality of input features and the plurality of weights, based on the input feature information and the weight information. The neural network processor and method also include a data arithmetic circuit configured to convolute the determined weights and input features to generate an output feature map.
Owner:SAMSUNG ELECTRONICS CO LTD +1

A low power multi-core shared floating point unit structure

This invention discloses a low-power multi-core shared floating-point unit structure, relating to the field of digital signal processing technology. It addresses the technical problem that existing floating-point units occupy a large silicon wafer area, leading to increased manufacturing costs and reduced yields after configuring one floating-point unit for each core. The low-power multi-core shared floating-point unit structure of this invention includes a processor core, shared floating-point units, and a multi-level logarithmic interconnect network. The processor core and the shared floating-point units communicate via an auxiliary processing unit interface designed for tightly coupled accelerators. The processor core executes application streams and dispatches floating-point calculation instructions to the shared floating-point unit cluster. The shared floating-point units perform floating-point operations as a shared computing resource time-division multiplexed by multiple processor cores. The multi-level logarithmic interconnect network dynamically and scalably connects all processor cores and shared floating-point units to form a communication path.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Real-time gesture interaction method based on NPU edge computing platform and safety helmet

PendingCN122363513AAlgorithmEngineering
A real-time gesture interaction method based on an NPU edge computing platform, steps comprising: S1 image acquisition facing the edge device use scene; S2 gesture recognition model construction; S3 depth model optimization facing the characteristics of a neural network processor NPU; S4 model conversion and end side deployment.In S3, for the inference characteristics of the NPU-based edge computing platform, the gesture recognition model obtained in S2 is adjusted in structure and compressed in parameters, the method being: firstly, by analyzing the inference delay and source occupation of each network layer of the gesture recognition model network structure on the target hardware, the network structure of the gesture recognition model is pruned and reconstructed, and the network layer with large calculation overhead is preferentially optimized; then, parameter compression and operator optimization are performed.The present application realizes a lightweight gesture recognition model with high precision and low delay through a technical path of layered performance query and iterative pruning and other hardware and software collaborative optimization techniques, and has the characteristics of lightweight, high inference efficiency and the like.
Owner:NANJING TECH UNIV

Communication apparatus and method

ActiveCN116939716BComputer hardwareInformation function
A communication device and method can be used in 802.11 series protocols, such as 802.11be and 802.11bf. The device comprises an input interface, a neural network processor and an output interface, the neural network processor storing a neural network comprising an input layer and an output layer; the input interface is configured to acquire at least one of input indication information, function indication information, output indication information or parameter indication information; the neural network processor is configured to execute a function indicated by the function indication information according to the input layer and the output layer; and the output interface is configured to output a result after the neural network processor executes the function. The input layer structure and the output layer structure of the neural network can be determined based on the information acquired by the input interface. Through the device, the function implemented by the MAC layer can be implemented based on the neural network, the complexity of the neural network is simplified, and hardware acceleration is facilitated.
Owner:HUAWEI TECH CO LTD

Passive suppression method of relative carrier envelope phase noise in dual optical comb spectral system

PendingCN122385526ANerve networkPhase noise
The present application relates to the technical field of double optical comb spectrum, and specifically discloses a passive suppression method for relative carrier envelope phase noise in a double optical comb spectrum system, which comprises the following steps: constructing a hardware optical path, generating a double optical comb signal through a common pump unit and a differential dispersion compensation unit, and performing phase matching at a physical level by using a passive optical delay matching unit; inputting the double optical comb signal into a field programmable logic gate array core of a customized system-level chip, executing synchronous acquisition and passive suppression processing of phase noise of the double optical comb signal, and obtaining a suppressed signal. In the actual industrial field measurement process, the present application first completes the basic phase constraint of the double optical comb signal at the physical level through common pumping and optical delay matching, then directly uses the field programmable logic gate array core inside the system-level chip to perform real-time passive noise suppression, and seamlessly hands over to the neural network processor core to execute multi-spectrum interference elimination.
Owner:HEFEI QINGXIN SENSING TECH CO LTD

Matrix multiplication conversion method, apparatus, medium, device and product

The application discloses a matrix multiplication conversion method and device based on systolic array convolution calculation, a medium, equipment and a program product, and belongs to the technical field of neural network processors. The method mainly comprises the following steps: according to the input channel number of an input feature map of a systolic array, cutting matrix row data of a first matrix, and inputting the cut matrix row data of the same row into the systolic array in the form of row data; according to the input channel number and the output channel number of the weight of the systolic array, cutting matrix column data of a second matrix, and inputting the cut matrix column data of the same row into the systolic array in the form of column data; and calculating the product of the first matrix and the second matrix by using the systolic array. The application improves the data reuse rate of matrix multiplication operation, reduces memory access, improves the cache utilization rate, and optimizes data locality.
Owner:CCORE TECH CO LTD

An AI application development platform that integrates convolutional networks

This invention relates to the field of artificial intelligence technology and discloses an AI application development platform that integrates convolutional networks. The platform includes: a model parsing module for parsing the model and calculating the spatial importance weights and stability coefficients of the channels; a policy generation module for generating a policy table containing load levels and binary masks based on the weights and coefficients; and an application building module for encapsulating the computing power control unit to generate the target application. During program execution, the computing power control unit determines the policy level based on the hardware state and uses masks to control the data flow to either enter the neural network processor for convolution or enter the graphics processing unit for motion vector multiplexing. This invention achieves heterogeneous hardware collaboration and dynamic policy fusion through the development platform, solving the high power consumption problem of frame-by-frame full computation and improving the inference efficiency of the application and the device's battery life.
Owner:YISHU TECHNOLOGY (TIANJIN) CO LTD

Instruction generation apparatus and method, device, storage medium, and computer program product

An instruction generation apparatus includes: a retrieval instruction transmitting module configured to acquire a retrieval decoding signal corresponding to a target neural network, and generate a retrieval instruction according to the retrieval decoding signal, wherein the retrieval instruction is configured to control a neural network processor to acquire an input characteristic diagram and a convolution kernel; a matrix instruction transmitting module configured to acquire a matrix decoding signal, and generate a matrix computation instruction according to the matrix decoding signal, wherein the matrix computation instruction is configured to control the neural network processor to perform an operation of a convolution computation on the input characteristic diagram and the convolution kernel; a storage instruction transmitting module configured to acquire a storage decoding signal, and generate a storage instruction according to the storage decoding signal, wherein the storage instruction is being configured to control the neural network processor to store an output characteristic diagram.
Owner:GUANGZHOU PWR SPLY BUR OF GNGDNG PWR GRD CO LTD

A pulse neural network processor with unified architecture

The application relates to the technical field of pulse neural network processing, in particular to a pulse neural network processor with a unified architecture, which comprises an LIF neuron module, a pulse collector module, an SRAM controller module, an SRAM module and a configuration module; the SRAM module comprises a weight SRAM, an entry SRAM and a neuron SRAM; the configuration module is used for realizing configuration of a chip architecture, and the configuration of the chip architecture comprises a sparse architecture and a dense architecture. The sparse architecture and the dense architecture can be simultaneously supported under the same chip architecture, storage and pulse driving calculation of the sparse architecture are used, lower task energy consumption can be realized when a sparse pulse neural network is processed, storage and matrix operation of the dense architecture are used, lower task energy consumption can be realized when a dense pulse neural network is processed, and the advantages of the two architectures can be combined in one chip.
Owner:FUDAN UNIVERSITY

Hardware accelerators, chips, computer devices suitable for machine learning

This invention belongs to the field of NPU, specifically relating to a hardware accelerator suitable for machine learning, its corresponding neural network processor chip, and a computer device. The hardware accelerator includes: a data computation module, a data storage module, a data read / write module, a data allocation module, and a computation control module. The data computation module contains all operators applicable to a specified machine learning algorithm. The data storage module includes multiple internal buffers. The data read / write module includes two DMAs for accessing external memory. The data allocation module is used to preprocess feature maps according to acquired configuration information and transfer data between internal and external memory. The computation control module manages the operation of the data computation module according to network configuration and parameters. The solution of this invention can improve the computational efficiency of processing data processing tasks involving machine learning algorithms in a computer system; overcoming the performance shortcomings of existing computers using CPUs or GPUs.
Owner:ANHUI UNIV

System and method for training multimodal behavior prediction model

A system and a method for training a multimodal behavior prediction model. The method is performed in a computing device that includes a processor and a neural processor. The processor retrieves multiple types of sensor data generated by one or more untrusted sensors and trusted sensors, and the neural processor uses multiple types of models corresponding to the multiple types of sensor data to predict behaviors of a user. The sensor data generated by the trusted sensors can be used to train the sensor data that are generated by the one or more untrusted sensors at the same time so as to train one or more prediction models. Therefore, the neural processor uses the trained prediction models and the trusted model to jointly establish the multimodal behavior prediction model that can be used to predict behaviors of the user and send a reminder for a specific event.
Owner:REALTEK SEMICON CORP

A heterogeneous system and method for implementing attention mechanisms in large language models

This invention discloses a heterogeneous system for implementing an attention mechanism in a large language model, comprising: a controller, a computing unit, and a memory; the computing unit includes a neural network processor and a graphics processor; the memory includes a local high-speed storage module; the controller controls the computing unit through an access interface and a data bus, and is connected to the memory via the computing unit; the computing unit and the local high-speed storage module are bidirectionally connected via a storage interface and a data bus; the heterogeneous system performs data transmission and exchange with an external global storage module; the local high-speed storage module and the global storage module are connected and interact with each other via a direct memory access channel. This invention also discloses an inference optimization method implemented through the above heterogeneous system, which has broad application value.
Owner:SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD

An instruction buffer and a control method thereof, a network processor and a chip

This application relates to an instruction buffer and its control method, a network processor, and a chip. It includes at least a storage module, a read pointer address management module, and a controller. The storage module comprises contiguously addressed memory cells, each storing an entry containing a valid flag, a thread identifier, and instruction data. The controller reads the instruction according to the read pointer address and sends it to the back-end pipeline, simultaneously invalidating the entry flag. The read pointer address management module manages the read pointer. After each read, it searches backward within a preset range from the current address for the first valid flag memory cell. If found, the read pointer is updated to that address; otherwise, it is updated to the address following the end of the range. The instruction buffer of this application only needs to update the read pointer address after reading the instruction, without physically shifting data. Furthermore, the read pointer address can skip invalid instructions, thereby reducing power consumption and minimizing invalid instructions entering the pipeline.
Owner:SHENZHEN JAGUAR MICROSYSTEMS CO LTD

Cluster intralayer safety mechanism in an artificial neural network processor

ActiveUS12670233B2Data streamSpatial mapping
Novel and useful system and methods of functional safety mechanisms for use in an artificial neural network (ANN) processor. The mechanisms can be deployed individually or in combination to provide a desired level of safety in neural networks. Multiple strategies are applied involving redundancy by design, redundancy through spatial mapping as well as self-tuning procedures that modify static (weights) and monitor dynamic (activations) behavior. The NN processor incorporates several functional safety concepts which reduce its risk of failure that occurs during operation from going unnoticed. The mechanisms function to detect and promptly flag and report the occurrence of an error with some mechanisms capable of correction as well. The safety mechanisms cover data stream fault detection, software defined redundant allocation, cluster interlayer safety, cluster intralayer safety, layer control unit (LCU) instruction addressing, weights storage safety, and neural network intermediate results storage safety.
Owner:HAILO TECH LTD

A system on chip, electronic device and vehicle

This application provides a system-on-a-chip (SoC), an electronic device, and a vehicle, relating to the field of automotive electronic control technology. The SoC includes: a main processor module containing multiple processor cores; at least one digital signal processor (DSP) module connected to the main processor module; and at least one neural network processor (NN) module connected to both the main processor module and the DSP module. At least one first processor core in the main processor module allocates tasks to be processed to at least one of the other processor cores, the DSP module, or the NN module for execution. The DSP module, the NN module, and the main processor module interact with each other via a shared memory area. This application achieves precise matching of computing power for general control, signal processing, and artificial intelligence inference tasks through a heterogeneous multi-core architecture and a dynamic allocation mechanism based on task characteristics, thereby improving processing performance, energy efficiency, and functional safety levels.
Owner:CHIPSEA TECH SHENZHEN CO LTD +1

Network packet processing device and network packet forwarding method

PendingCN122073571ANetwork connectionsData streamNetwork processor
The invention discloses a network data packet processing device and a network data packet forwarding method. The device comprises a parallel processing circuit, a data packet distribution circuit and a data packet order-preserving processing circuit. The parallel processing circuit comprises a plurality of data packet processing circuits which are used for carrying out parallel processing on different data packets. Each packet processing circuit includes a network processor core. The data packet distribution circuit respectively distributes a plurality of data packets to the plurality of data packet processing circuits. The data packet order-preserving processing circuit performs order-preserving transmission on a plurality of processed data packets generated by the parallel processing circuit, and the plurality of processed data packets comprise a first processed data packet and a second processed data packet which respectively correspond to the first data packet and the second data packet in the plurality of data packets. And the sequence of the first processed data packet and the second processed data packet in the output data stream is the same as the sequence of the first processed data packet and the second processed data packet in the input data stream.
Owner:AIROHA TECH (SUZHOU) LTD

Sparse data processing method and device of neural network processor

The present disclosure provides a sparse data processing method and device of a neural network processor, the neural network processor comprising a basic computing unit, the method comprising: obtaining a plurality of groups of weight sub-vectors, wherein the weight sub-vectors are obtained by sparse processing of a to-be-computed weight vector; determining a to-be-computed feature vector corresponding to the to-be-computed weight vector; performing combination processing on the plurality of groups of weight sub-vectors to obtain a combined weight vector, wherein the combined weight vector is the same as an information unit supported by the basic computing unit; and controlling the basic computing unit to perform shift-and-add computation on the combined weight vector and the to-be-computed feature vector to obtain a sparse data processing result. Through the present disclosure, the distribution of weights can be fully utilized, variable-precision sparse computation can be supported, and the hardware cost required for sparse data processing can be effectively reduced while ensuring the accuracy of sparse data processing.
Owner:AXERA TECH (BEIJING) CO LTD

Data processing apparatus, data processing method and related devices

ActiveCN116362304Bwrite lessreduce readComputer hardwareDatasheet
The application provides a data processing device, a data processing method and related devices, comprising a neural network processor, first, according to the channel number n of the first data, the first data is divided into n first sub-data, the first data represents data to be stored, n is a positive integer less than or equal to M / 2; each first sub-data is written into a circular cache space composed of INT(M / n) ring cache space groups, and any one circular cache space is used for storing any one first sub-data. The area of the cache control circuit in the neural network processor can be reduced by a specific storage space architecture, and the power consumption of data writing and reading is reduced.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Operator compiling method and device, equipment and storage medium

PendingCN122072555AProgram code adaptionComputer simulationsNetwork processorCode transformation
The invention discloses an operator compiling method and device, equipment and a storage medium, which are used for improving the development efficiency of operators. The operator compiling method comprises the following steps: acquiring a high-level programming language code of a to-be-compiled operator; the advanced programming language code is converted into a virtual instruction of a to-be-compiled operator, the virtual instruction is an instruction in a virtual instruction set, a mapping relation exists between a virtual instruction set and a hardware instruction set of the neural network processor, and the virtual instruction set is constructed based on a hardware instruction type of the hardware instruction set; and converting the virtual instruction into a hardware instruction of the operator to be compiled. According to the scheme, during compiling, the advanced programming language code of the operator to be compiled is converted into the virtual instruction, and then the virtual instruction is further converted into the hardware instruction of the operator to be compiled, so that the negative influence caused by the difference between the hardware instruction sets can be avoided; therefore, the high-level programming language code of the operator to be compiled can be quickly compiled into the hardware instruction, and the development efficiency of the operator is improved.
Owner:HUAWEI TECH CO LTD +1