Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

324 results about "Neural processing" patented technology

Neural processing originally referred to the way the brain works, but the term is more typically used to describe a computer architecture that mimics that biological function. In computers, neural processing gives software the ability to adapt to changing situations and to improve its function as more information becomes available.

Graph neural network execution on neural processing unit

Workloads for executing a graph neural network (GNN) may be divided among various processing units, such as a central processing unit (CPU) and a neural processing unit (NPU). The NPU may include a data processing unit (DPU) and a digital signal processor (DSP). The CPU may perform precomputation, model optimization, hardware optimization, and compilation. For example, the CPU may precompute a parameter matrix and use the parameter matrix as internal parameters of a GNN. The CPU may also perform node padding, approximation computation, or transfer of DSP operations to DPU to optimize the GNN. The CPU may also perform sparsity data compute and storage, vertical fusion of DSP operations and DPU operations, or data quantization to optimize performance of the NPU. The compiled GNN may be provided to the NPU, and the DPU and DSP may perform the operations in the compiled GNN to produce a prediction of the GNN.
Owner:INTEL CORP

Neuro-Generative Adversarial System for real-time detection and combating of malware morphing in high-density edge networks

ActiveDE202025106911U1Platform integrity maintainanceData packEmbedded security
A system for real-time detection and mitigation of morphing malware in high-density edge networks, consisting of: a data acquisition unit configured to receive, normalize, and encode multimodal telemetry data streams originating from at least one of the following domains: network traffic, process behavior, system call sequences, binary instruction traces, and control flow graphs; the data acquisition unit is further configured to compute feature embeddings over sliding time windows and apply privacy-preserving redactions prior to storage; a generative neural processor that is operationally coupled to the data acquisition unit and configured to generate synthetic morphing malware variants by learning probabilistic transformations of previously observed malicious data representations, maintaining semantic functionality while varying structural and behavioral features; a discriminative neural processor trained adversarially with the generative neural processor, wherein the discriminative neural processor is configured to detect morphing malware by evaluating a probability distribution over multimodal telemetry embeddings and classifying anomalous process and flow behaviors in real time; a coordination processor that is communicatively connected to both the generative neural processor and the discriminative neural processor and is configured to orchestrate adversarial co-training, regulate detection thresholds, calculate reinforcement-based penalties for false negative results, and trigger countermeasures as soon as a detection confidence level exceeds a predefined adaptive threshold; a secure, system-integrated inference and enforcement unit configured to perform low-latency countermeasures at the network edge, including selective packet filtering, flow isolation, process interruption, or system microsegmentation, based on instructions from the coordinating processor; and a hardware-embedded security enclave that is embedded in the system and configured to store cryptographic keys, neural model parameters, and integrity affirmation data to ensure the confidentiality, authenticity, and immutability of model artifacts and policy configurations.
Owner:ANAJAVADIDHODDI RAMACHANDRA NAIK CHAYAPATHI BENGALURU +7

Systems and methods for heterogeneous large language model encoder and decoder processing

Systems and methods are disclosed for efficient memory allocation for processing large language model encoders and decoders based on an attention model. The system can utilize a plurality of two types of processors suitable for different types of LLM processing. These include Neural Processor Units and Graphic Processing Units. Each NPU processor has dedicated DDR memory coupled to each NPU. The DDR memory caches the neural network weights used in the generation of neural network activations. A plurality of GPUs provides KVQ token processing. LLM tokens processing can be performed in parallel batches or sub-batches to utilize idle NPU processors within the neural network. In some embodiments, the NPUs are structured in a matrix with a bus between adjacent processors. In another embodiment, NPUs provide both KVQ processing and neural network processing. The system can be integrated on a silicon substrate using chiplets in a 2.5 or 3-D architecture.
Owner:EXPEDERA INC

System integrated machine-learning co-processing

The technology described adds a ML inference to the output of an image signal processor (ISP) associated with a camera. The combined image and ML inference may be described herein as an augmented image. Once generated, the augmented image may be communicated to other components of a computing system associated with the camera and / or ISP. The initial inference may be generated by a neural processing unit (NPU) associated with the ISP. The ISP may communicate a generated image to the NPU prior to communicating the image to a computing system. In an aspect, the NPU inference is combined with the image using image steganography. Once communicated from the camera to the computing device, the augmented image may be separated into a base image and inference by a camera driver or other component associated with the image management.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Neural processing unit including post-processing unit

According to one example of the present disclosure, the neural processing unit may comprise a processing element array configured to perform operations of a neural network model and a post-processing unit configured to process data output from the processing element array. The post-processing unit includes a first computation circuit that extracts a subset of classes for each bounding box by comparing class scores of classes and a second computation circuit configured to extract one or more bounding boxes by comparing a class confidence score of each bounding box with a threshold confidence score.
Owner:DEEPX CO LTD

Edge device with built-in compiler for neural network models

A system includes a substrate on which a first memory, a neural processing unit (NPU) including a plurality of processing elements (PEs) with multiplier-accumulator circuits, a controller, and a second memory, and a central processing unit (CPU) are disposed. The CPU may be configured to execute a universal compiler to perform a conversion for a particular neural network model into a machine code executable by the NPU and store the machine code in the first memory or the second memory. When the particular neural network model, generated by one among a plurality of machine learning frameworks that are incompatible with each other, is received and stored in the first memory, the universal compiler may perform the conversion based on mapping information indicating mapping between elements of machine learning frameworks and functions or operations executable by the CPU or NPU.
Owner:DEEPX CO LTD

Reconfigurable heterogeneous radar data calculation module and calculation device

The invention discloses a reconfigurable heterogeneous radar data computing module and computing device, and relates to the technical field of computing devices, the reconfigurable heterogeneous radar data computing module comprises a ZYNQ processor, a digital signal processor (DSP) and a neural processing unit (NPU), the ZYNQ processor is connected with the DSP and the NPU; the ZYNQ processor comprises an ARM processor and an FPGA (Field Programmable Gate Array) logic resource; the FPGA logic resource is used for carrying out parallel class algorithm processing on the radar data; the DSP is used for carrying out serial algorithm processing on the radar data; the NPU is used for performing artificial intelligence algorithm processing on the radar data; the ARM processor is used for scheduling FPGA logic resources according to the calculation tasks, and the DSP and the NPU execute corresponding algorithms. According to the reconfigurable heterogeneous radar data calculation module, the high-performance calculation requirement can be met.
Owner:BEIJING DONGYUAN RUNXING TECH CO LTD

Updating of parameters of neural network model for efficient execution on neural processing unit

Embodiments relate to converting functions or function call instructions of a first neural network (NN) model into graph module. The relationship between one or more inputs and one or more outputs of the graph modules are analyzed. A second neural network (NN) model in a form of a directed acyclic graph (DAG) including using the graph modules is generated by mapping inputs and outputs of the graph modules based on the relationship. Markers are added to the graph modules in the second NN model. First calibration data is generated by collecting input values and output values of each of the graph modules using the markers. An adjustment value for outlier alleviation for each of the graph modules is generated based on the first calibration data. For each graph module of the second NN model, an input parameter and a weight parameter are updated based on the adjustment value.
Owner:DEEPX CO LTD

Neural processing unit that reuses feature maps and its operation method

A method of reusing feature maps that is operable on a neural processing unit (NPU) includes calculating, for a first layer of at least one artificial neural network (ANN) model, a space cost based on sizes of an input feature map, an output feature map, and a weight of a subsequent layer of the first layer; calculating a caching value for an operation of the first layer; determining a caching profit for the operation of the first layer based on the space cost and the caching value; determining a caching entry having maximum caching profit among caching candidate entries; and storing the caching entry in a variable memory. A system including the NPU includes a main memory to store a portion of the ANN model; a variable memory to selectively store a caching entry of the portion of the ANN model; and a controller to perform the method.
Owner:DEEPX CO LTD

Server cable management method and device, equipment and storage medium

The invention discloses a server cable management method and device, equipment and a storage medium, and relates to the technical field of fault detection, and the method comprises the steps: constructing a cable connection relation graph of a target cable of a target server based on a preset graph neural network; the target cable is a cable embedded in the controller and the target sensor in advance, and a neural processing unit is integrated in the controller; generating a cable code of the target cable, and determining a cable state of the target cable through the target sensor by using the controller; when the cable state is abnormal, determining a fault path of the target server according to the cable code and the cable connection relation graph; and displaying the fault path by using a positioning display device of the target cable so as to manage the target server. The controller and the sensor are implanted into the cable body, the cable connection relation graph is constructed in combination with the neural network model, cable troubleshooting and life management are further achieved through cable coding, the troubleshooting cost is greatly reduced, and the intelligent management process of a server is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Device and method for providing information using on-device ai recognition and server interworking

A mobile artificial intelligence (AI) device includes a camera configured to capture a video of a product, and a neural processing unit (NPU) configured to perform inference with an AI recognition model to recognize product information from the captured video on the mobile AI device. A battery supplies power to the NPU. A transceiver is configured to transmit only the recognized product information to a remote server without transmitting the captured video, and to receive additional commercial information corresponding to the recognized product information. A display is configured to present the additional commercial information.
Owner:DEEPX CO LTD

Consciousness perception quantum gravitational nerve acceleration unified computing platform and method based on formalized verification

The invention discloses an awareness perception quantum gravitational nerve acceleration unified computing platform and method based on formalized verification, and belongs to the field of high-performance computing, medical AI and quantum information processing. A quantum gravitational coupling neural processing unit (QGCNPU); a consciousness information integral calculation core (phi ICC); formalizing a receipt chain verifier (FRCV); and the mirror reflection symmetry correction module (MRSCM) is used for eliminating the continuous projection mistake and the discrete Boolean mistake through topological equivalence verification of a local observer and a global GodCam visual angle. The system adopts an FPGA + ASIC + quantum processor three-layer heterogeneous architecture, and supports real-time pathological modeling (precision gt; gt) of neurodegenerative diseases such as Alzheimer's disease, Parkinson's disease and ALS; 99.7%), drug target discovery (acceleration by 50 times) and consciousness state monitoring (time resolution lt; and meanwhile, the mathematical proving performance and the clinical auditing performance of the calculation process are ensured, and the requirements of the global 152 billion dollar neural science and technology market and the 89 billion dollar quantum calculation industry are met.
Owner:GUANGZHOU KINGPIN IND CO LTD

System for processing heterogeneous generative artificial intelligence models

In accordance with the present disclosure, an apparatus is provided. The apparatus includes: a first memory having a first capacity configured to store a first generative neural network model including a first parameter; and a first neural processing unit configured to generate a response corresponding to the input query using the first generative neural network model stored in the first memory; and wherein the first neural processing unit may be configured to store first execution code of the first generative neural network model, the first execution code being compiled to process the speculative decoding.
Owner:DEEPX CO LTD

Accelerating artificial neural networks using hardware-implemented lookup tables

The invention is notably directed to a hardware system (1) designed to implement an artificial neural network (ANN). The hardware system basically includes a neural processing apparatus (15), e.g., involving as crossbar array structure, one or more lookup table circuits (17), and one or more processing units (18). The neural processing apparatus is configured to implement M artificial neurons, where M≥1. The lookup table circuits are configured to implement a lookup table (LUT). The system further includes M′ processing units, where M≥M′≥1. Each processing unit is connected by at least one neuron, in order to be able to access a first value outputted by each connected neuron. In addition, each processing unit is connected to a LUT circuit, in order to efficiently access parameter values of a set of parameters from the LUT. Finally, each processing unit is configured to output a second value, corresponding to a value of a mathematical function taking said first value as argument. The mathematical function is otherwise determined by the set of parameters, the parameter values of which are accessed by each processing unit from the LUT, in operation. I.e., the mathematical function is defined (and thus determined) by a set of parameters, the values of which are efficiently retrieved from the hardware-implemented LUT. This results in a substantial acceleration of the computations of the function outputs, beyond the acceleration that may already be achieved within the neural processing apparatus and the processing units themselves. As a result, the neuron outputs can be more efficiently processed, prior to being passed to a next neuron layer. The invention is further directed to a method of operating such a hardware system.
Owner:AXELERA AI BV

Nonlinear tensor compression and decompression for neural networks

Devices and techniques are generally described for nonlinear tensor compression for neural networks. In various examples, a first tensor associated with a first layer of a neural network may be determined. One or more neural processing units of accelerator hardware may generate a first compressed tensor by applying a nonlinear compression function to the first tensor. The first compressed tensor may be stored in a first memory of the one or more computer-readable media. A first operation associated with a second layer of the neural network may be determined, where the first operation uses output of the first layer. The first operation may be performed based on the first compressed tensor.
Owner:AMAZON TECH INC

Dock-based neural processing unit handoff to client computing device on disconnect of the client computing device

A docking station includes a processor and a memory coupled to the processor. The docking station may be configured to establish a sideband connection between an information handling system and the docking station subsequent to detecting the information handling system docking at the docking station. The docking station may also execute a workload of the information handling system, subsequent to the establishment of the sideband connection. In addition, the docking station may transmit a notification to resume the workload via the sideband connection subsequent to detecting that the information handling system is undocked from the docking station while performing the executing of the workload.
Owner:DELL PROD LP

Apparatus, method, and system for deploying neural network model

A method may comprise receiving a first neural network (NN) model including one or more functions; generating a second NN model in a form of directed acyclic graph (DAG) including one or more graph modules by converting the one or more functions; calculating one or more scale values by obtaining maximum and minimum values of parameters input to the one or more graph modules; updating the parameters based on the one or more scale values; and generating a third NN model, in a form of machine code executable on a particular neural processing unit, including the updated parameters.
Owner:DEEPX CO LTD

Processing heterogeneous generative artificial intelligence models

According to the present disclosure, a device is provided. The device includes a first memory of a first capacity configured to store a first generative neural network model comprising a first parameters, and a first neural processing unit configured to generate a response corresponding to an input query utilizing the first generative neural network model stored in the first memory, and wherein the first neural processing unit may be configured to store a first execution code of the first generative neural network model compiled to process speculative decoding.
Owner:DEEPX CO LTD

Peer Configuration of Compute Resources Among Data Storage Devices

Example storage systems, storage devices, and methods provide peer configuration of hardware compute resources among peer storage devices. Data storage devices may include hardware circuits, such as graphics processor units, media processor circuits, compute accelerator circuits, and neural processing units, configured for compute operations targeting host data associated with that storage device. One of the storage devices is configured to act as a master storage device for determining firmware configurations for the other storage devices and sending an indication of the firmware image to be used for some operating period. The storage device receiving the firmware image loads the firmware image for the hardware circuit and reboots the hardware circuit to configure it with the firmware image.
Owner:SANDISK TECHNOLOGIES LLC

Heterogeneous nerve processing system of generative AI model

A heterogeneous neural processing system includes a first processor configured to perform encoding and decoding operations of an autoencoder, and a second processor configured to perform neural network operations of a particular task, the operations having iterative processing. The processors perform compute tasks through synchronous data exchange to implement a generative AI model. The first processor processes a feature map divided into data rows, caches the data into active memory using row-based depth-first scheduling, and selects deeper operations in the network hierarchy when processing branch inputs, outputs, and residual connections. And H reuses boundary pixels between the cache storage space segments, so that convolution and element-level operation can be executed simultaneously. A neural network adjusts the device analysis model to identify layer dependencies, applies search space constraints, performs an iterative search to generate a fusion plan, and selects an optimal plan based on external memory access and execution delays.
Owner:MEDIATEK INC

Processor, computing node, and computing cluster

Disclosed in the embodiments of the present application are a processor, a computing node, and a computing cluster, which are used for simplifying the network architecture of the computing cluster. The processor comprises: a general-purpose protocol port, which is used for supporting a first communication protocol and a second communication protocol, wherein the second communication protocol is different from the first communication protocol; the first communication protocol is a communication protocol of a first network, the second communication protocol is a communication protocol of a second network, the first network is used for supporting communication between a plurality of processors in one computing node including the processor, the second network is used for supporting communication between the processor and other processors comprised in other computing nodes, and the processor is any one of a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU) and a general-purpose graphics processing unit (GPGPU).
Owner:HUAWEI TECH CO LTD

System and method for fine-tuning rotated outlier-free large language models for effective weight-activation quantization

A computing device includes at least one processor, one or more non-transitory computer-readable storage media, a system for fine-tuning a large language model under low-bit weight-activation quantization. The computing device further comprises a graphics processing unit (GPU), a neural processing unit (NPU), or a tensor processing unit (TPU). The hardware interface module of the system is configured to load a low-bit model representation from the memory module and transmit the model representation to the GPU, NPU, or TPU for inference execution.
Owner:THE HONG KONG UNIV OF SCI & TECH

Neural processing device and method for using shared page table thereof

A neural processing device and a method for using shared page table thereof are provided. The neural processing device including at least one neural processor, a shared memory shared by the at least one neural processor, and a global interconnection configured to exchange data between the at least one neural processor and the shared memory, comprises at least one processing unit each of which included in each of the at least one neural processor and configured to provide logical addresses, a memory management unit configured to receive and translate the logical addresses into physical addresses, and a physical memory accessible by the physical addresses, wherein the memory management unit comprises a shared page table that has translation information between the logical addresses and the physical addresses and is shared by at least one process with each other.
Owner:REBELLIONS INC

Neural processing unit operable in multiple modes to approximate activation function

A neural processing unit may be provided. The neural processing unit may comprise a controller circuit configured to select an activation function processing method among a first method or a second method, according to an activation function included in a neural network model, a programmed activation function execution unit (PAFE unit) configured to execute a programmed activation function (PAF) that approximate the activation function and output a first activation value, and a converter circuit configured to convert the first activation value and output a second activation value. In the first method, only the PAFE unit may operate. In the second method, both the PAFE unit and the converter may operate.
Owner:DEEPX CO LTD

Compressed weight distribution in networks of neural processors

A neural inference chip includes a global weight memory; a neural core; and a network connecting the global weight memory to the at least one neural core. The neural core comprises a local weight memory. The local weight memory comprises a plurality of memory banks. Each of the plurality of memory banks is uniquely addressable by at least one index. The neural inference chip is adapted to store in the global weight memory a compressed weight block comprising at least one compressed weight matrix. The neural inference chip is adapted to transmit the compressed weight block from the global weight memory to the core via the network. The core is adapted to decode the at least one compressed weight matrix into a decoded weight matrix and store the decoded weight matrix in its local weight memory. The at core is adapted to apply the decoded weight matrix to a plurality of input activations to produce a plurality of output activations.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Processing device and method for managing tasks thereof

A neural processing device and a method for managing tasks thereof are provided. The neural processing device includes a neural core configured to perform a task and generate a completion signal for completion of the task, a core global configured to transfer task information for the task to the neural core and receive the completion signal of the task from the neural core, and a task manager configured to generate and transmit the task information to the core global, receive the completion signal from the core global, generate a completion report, and transmit the completion report.
Owner:REBELLIONS INC

Subtask storage for streaming convolutions in neural network processor

Embodiments relate to streaming convolution operations in a neural processor circuit that includes a neural engine circuit and a neural task manager. The neural task manager obtains multiple task descriptors and multiple subtask descriptors. Each task descriptor identifies a respective set of the convolution operations of a respective layer of a set of layers. Each subtask descriptor identifies a corresponding task descriptor and a subset of the convolution operations on a portion of a layer of the set of layers identified by the corresponding task descriptor. The neural processor circuit configures the neural engine circuit for execution of the subset of the convolution operations using the corresponding task descriptor. The neural engine circuit performs the subset of the convolution operations to generate output data that correspond to input data of another subset of the convolution operations identified by another subtask descriptor from the list of subtask descriptors.
Owner:APPLE INC

Neural core, neural processing device including same, and method for loading data of neural processing device

A neural core, a neural processing device including same and a method for lauding data of a neural processing device are provided. The neural core comprises a processing unit configured to perform operations, an L0 memory configured to store input data and an LSU configured to perform a load task and a store task of data between the processing unit and the L0 memory, wherein the LSU comprises a local memory load unit configured to transmit the input data in the L0 memory to the processing unit, and the local memory load unit comprises a target decision module configured to identify and retrieve the input data in the L0 memory, a transformation logic configured to transform the input data and thereby generate transformed data and an output FIFO configured to receive the transformed data and transmit the transformed data to the processing unit in the received order.
Owner:REBELLIONS INC

Neural processing unit for performing RMS norm operation and control method thereof

A neural processing unit for performing inference operations of a large-scale language model based on an artificial neural network is disclosed. The neural processing unit according to the present disclosure includes a processing element core configured to perform an attention mechanism-based operation based on input data in vector format to output an operation result, a special function unit comprising a plurality of arithmetic circuits including at least one vector-dedicated arithmetic circuit that exclusively performs vector operations and at least one mixed arithmetic circuit capable of performing both vector and scalar operations, and configured to perform a special function operation on the operation result, and a controller configured to, upon receiving an RMS normalization operation execution command, activate at least one of the plurality of arithmetic circuits to control the special function unit to perform an operation of converting at least one of the operation result or the input data into a normalized vector whose magnitude is adjusted based on a root mean square (RMS), wherein the operation result may include an attention score for the input data.
Owner:DEEPX CO LTD