Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

50 results about "Neural processing" patented technology

Neural processing originally referred to the way the brain works, but the term is more typically used to describe a computer architecture that mimics that biological function. In computers, neural processing gives software the ability to adapt to changing situations and to improve its function as more information becomes available.

Nonlinear tensor compression and decompression for neural networks

Devices and techniques are generally described for nonlinear tensor compression for neural networks. In various examples, a first tensor associated with a first layer of a neural network may be determined. One or more neural processing units of accelerator hardware may generate a first compressed tensor by applying a nonlinear compression function to the first tensor. The first compressed tensor may be stored in a first memory of the one or more computer-readable media. A first operation associated with a second layer of the neural network may be determined, where the first operation uses output of the first layer. The first operation may be performed based on the first compressed tensor.
Owner:AMAZON TECH INC

Neural processing unit for performing RMS norm operation and control method thereof

A neural processing unit for performing inference operations of a large-scale language model based on an artificial neural network is disclosed. The neural processing unit according to the present disclosure includes a processing element core configured to perform an attention mechanism-based operation based on input data in vector format to output an operation result, a special function unit comprising a plurality of arithmetic circuits including at least one vector-dedicated arithmetic circuit that exclusively performs vector operations and at least one mixed arithmetic circuit capable of performing both vector and scalar operations, and configured to perform a special function operation on the operation result, and a controller configured to, upon receiving an RMS normalization operation execution command, activate at least one of the plurality of arithmetic circuits to control the special function unit to perform an operation of converting at least one of the operation result or the input data into a normalized vector whose magnitude is adjusted based on a root mean square (RMS), wherein the operation result may include an attention score for the input data.
Owner:DEEPX CO LTD

Model level debugging of machine learning designs on neural processing units

PendingUS20260186951A1EngineeringProcessing element
Model level debugging of a machine learning design includes compiling the machine learning design for execution on target hardware using a compiler. Metadata for the machine is generated. The metadata specifies a mapping of buffers of the machine learning design to a plurality of memory levels of a memory architecture of the target hardware correlated with boundaries of the machine learning design. While running the machine learning design, debug data is dumped from the plurality of memory levels of the memory architecture based on the boundaries. The debug data is correlated with the boundaries of the machine learning design based on the metadata.
Owner:XILINX INC

Signal processing device and vehicle display device comprising same

PCT designated stageWO2026105966A1Neural learning methodsDisplay deviceEngineering
A signal processing device and a vehicle display device comprising same, according to one embodiment of the present disclosure, comprise: at least one neural processor; and a central processor for controlling the neural processor, wherein the central processor divides an artificial intelligence model to be executed in the neural processor into a plurality of groups and controls execution or suspension of the respective groups. Therefore, the neural processor can be efficiently operated.
Owner:LG ELECTRONICS INC

Executing floating-point model through integer datapath in neural processing unit

A neural processing unit (NPU) may perform computations in integer domains to execute a floating-point model. The NPU may include an input delivery unit (IDU), a processing engine, and a post-processing engine. The IDU may convert floating-point weights to integers, e.g., by normalizing the floating-point weights, mapping the normalized floating-point values to normalized integer values using a look-up table, and scaling up the normalized integer values into integer values. The IDU may load the integer values into the processing engine through an integer datapath in the NPU. The processing engine may compute an output tensor of a neural network operation using the integer values. The post-processing engine may perform per-channel quantization of the output tensor using channel-specific quantization parameters stored in configuration registers of the post-processing engine. The look-up table and configuration registers may be programmed with configuration parameters determined by a compiler.
Owner:INTEL CORP

Graphics processing

Graphics processor (2) comprises a programmable execution unit (65) executing programs to perform graphics processing and a machine learning (ML) processing circuit (78) (e.g. neural processing accele
Owner:ARM LTD

Crossbar circuit for unaligned memory access in neural network processor

ActiveUS12675679B2Computer hardwareCrossbar switch
Embodiments of the present disclosure relate to an unaligned memory access in a neural processor circuit. The neural processor circuit includes a crossbar circuit and a neural engine circuit coupled to the crossbar circuit. During each operating cycle of the neural processor circuit, the crossbar circuit receives a portion of input data, and re-aligns or bypasses the portion of input data. The neural engine circuit receives at least a portion of the re-aligned or bypassed portion of the input data, and performs a convolution operation on the received portion of re-aligned or bypassed portion of input data to generate output data.
Owner:APPLE INC

Semiconductor package for npu

PendingUS20260191099A1Memory chipSemiconductor package
Neural processing unit (NPU) semiconductor package products and devices are provided. According to one embodiment, the semiconductor package comprises a substrate, at least one NPU chip mounted on the substrate and disposed at a first rotated orientation relative to a side or reference axis of the substrate, and at least one memory chip mounted adjacent to the at least one NPU chip on the substrate and disposed at a second rotated orientation, wherein the first orientation of the at least one NPU chip and the second orientation of the at least one memory chip are configured in a rotated layout such that overall dimensions of the substrate conform to a predetermined form factor smaller than a form factor of a standard non-rotated layout.
Owner:DEEPX CO LTD

Neural processing unit operable in multiple modes to approximate activation function

A neural processing unit may be provided. The neural processing unit may comprise a controller circuit configured to select an activation function processing method among a first method or a second method, according to an activation function included in a neural network model, a programmed activation function execution unit (PAFE unit) configured to execute a programmed activation function (PAF) that approximate the activation function and output a first activation value, and a converter circuit configured to convert the first activation value and output a second activation value. In the first method, only the PAFE unit may operate. In the second method, both the PAFE unit and the converter may operate.
Owner:DEEPX CO LTD

Function approximation unit configured to approximate nonlinear functions in a neural processing unit and operating method thereof

ActiveUS12675695B1AlgorithmControl signal
Methods and devices of a function approximation unit configured to approximate a nonlinear function within a neural processing unit are described. According to one embodiment, the method includes storing an input value through an input register of the function approximation unit, transmitting the input value to a selected one of a plurality of preprocessing circuits of the unit according to a control signal, generating a preprocessing result corresponding to the input value by the selected one of the preprocessing circuits, transmitting the preprocessing result to a programmable function approximation circuit and a selected one of a plurality of post-processing circuits of the unit, generating an approximated function output based on the preprocessing result in the programmable function approximation circuit, and generating a final output value by post-processing the preprocessing result or the approximated function output in the selected one of the post-processing circuits.
Owner:DEEPX CO LTD

Reconfigurable memory architecture for artificial intelligence models

A neural processing apparatus is disclosed. The apparatus includes a memory array including a plurality of physically distinct memory banks, configurable interconnect circuitry coupled to the plurality of memory banks, and a controller. The controller is configured to dynamically reallocate the plurality of memory banks into variable-sized memory regions corresponding to different data types required for an inference operation of a neural network model. The controller adjusts a number of memory banks allocated to each data type for a current layer of the neural network model based on configuration information derived from a structure of the neural network model. The controller is further configured to enable data access between the reallocated memory banks and a processing engine via the configurable interconnect circuitry.
Owner:DEEPX CO LTD

Method and device for controlling multiple neural processing units

A method for operating an electronic device, the electronic device including a plurality of neural processing unit (NPU) devices and an NPU management device configured to control the plurality of NPU devices, includes: determining whether input / output data of an artificial neural network application is shared by at least one NPU device configured to perform the artificial neural network application; determining whether a shared memory is available, wherein the NPU management device is further configured to control the shared memory; and storing the input / output data in the shared memory in a case that the input / output data of the artificial neural network application is shared by the at least one NPU device and that the shared memory is available, wherein the shared memory is shared by at least one neural network accelerator of the at least one NPU device.
Owner:SAMSUNG ELECTRONICS CO LTD

Neural processing unit capable of performing runtime test

ActiveUS12646584B2Detecting faulty hardware using neural networksDigital computer detailsProcessing elementNeural processing
A neural processing unit (NPU) is capable of testing a component of the NPU in a running system, i.e., during runtime. The NPU includes a plurality of functional components, each of which includes an electronic circuit; at least one wrapper connected to at least one of the functional components; and an in-system component tester (ICT). The ICT performs a selection of one of the at least one functional component, in an idle state, as a component under test (CUT) and performs a test, via the at least one wrapper, of the selected functional component. The ICT may monitor states of the plurality of the functional components via the at least one wrapper, stop the test based on a detection of a collision due to an access to the selected functional component, and return a connection of the selected functional component to the at least one wrapper according to the stop.
Owner:DEEPX CO LTD

Secure execution of an ai model on a neural processing unit of a client device

PendingUS20260189536A1Operational systemEngineering
Techniques are described herein that are capable of securely executing an AI model on a neural processing unit (NPU) of a client device. The NPU runs the AI model. The NPU encrypts data, which includes an AI prompt, using a cryptographic key. The NPU provides the encrypted data to a cloud-based security service via a utility in an operating system that executes on the computing system. The NPU receives a response indicator from the cloud-based security service via the utility. The response indicator represents a result of an analysis of a decrypted representation of the encrypted data. The response indicator suggests an alternative response in lieu of an AI response, which is received from the AI model as a result of the AI prompt, as a response to the AI prompt. The NPU provides the alternative response as the response to the AI prompt.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Neural processing unit

Neural processing units are provided. The neural processing units can be reconfigured to process fine-grained structured weight sparsity arrangements selected from N:M = 1:4, 2:4, 2:8, and 4:8 fine-grained structured weight sparsity arrangements. A weight buffer stores weight values, and a weight multiplexer array outputs one or more weight values stored in the weight buffer as first operand values based on the selected fine-grained structured weight sparsity arrangement. An activation buffer stores activation values, and an activation multiplexer array outputs one or more activation values stored in the activation buffer as second operand values based on the selected fine-grained structured weight sparsity arrangement, where each respective second operand value and respective first operand value form an operand value pair. A multiplier array outputs a product value for each operand value pair.
Owner:SAMSUNG ELECTRONICS CO LTD

Training and fine-tuning of neural networks on neural processing units

This disclosure relates to training and fine-tuning of neural networks on a neural processing unit. A core on a neural processing unit can perform matrix multiplication (MatMul) on tensors of different dimensions. A neural network can be trained through forward and backward operations, both of which can be offloaded to the core. For the forward operation, the core can perform a layer by performing MatMul on an input tensor and a weight tensor and generate an output tensor. A loss can be computed. For the backward operation, the core can compute a weight gradient of the loss by performing MatMul on a gradient of the output tensor and the input tensor, and compute an input gradient of the loss by performing MatMul on the gradient of the output tensor and the weight tensor. The gradient of the output tensor can be computed by an automatic differentiation module. The weight tensor can be updated based on the input gradient and the weight gradient.
Owner:INTEL CORP

Artificial intelligence device and its neural processor and operating method

PendingCN122311312AEngineeringTerm memory
This invention provides an artificial intelligence device, a neural processor, and an operating method. The artificial intelligence device includes a host circuit, memory, and a neural processor. The neural processor is coupled to the host circuit and the memory. The neural processor establishes a transmission connection to the host circuit. The model weight set of the artificial intelligence model includes a first weight subset and a second weight subset. During the initialization period before the neural processor executes the artificial intelligence model, the host circuit preloads the first weight subset into memory. During the execution of the artificial intelligence model, the neural processor receives the first weight subset from memory and the second weight subset from the host circuit to execute the artificial intelligence model.
Owner:NOVATEK MICROELECTRONICS CORP

Multimodal behavior prediction model training system and method

The present application provides a multi-modal behavior prediction model training system and method. The method is executed in a computing device provided with a processor and a neural processor. The processor obtains sensing data generated by one or more untrusted sensors and trusted sensors. The neural processor applies corresponding models to the sensing data to predict the user's behavior. The sensing data generated by the trusted sensors is used to train the sensing data generated by the untrusted sensors at the same time to obtain a prediction model. Thus, the neural processor uses the trained prediction model and the trusted model to establish a multi-modal behavior prediction model to predict the user's behavior and send a reminder for a specific event.
Owner:REALTEK SEMICON CORP

Dynamic scaling of neural network based on input sequence length

PCT designated stageWO2026142734A1Parallel computingProcessing element
A system may facilitate dynamic scaling of multi-head attention (MHA) layers in transformer networks. The system may include a compiler and a neural processing unit (NPU). The compiler may generate workload descriptors that define a plurality of workloads for performing an operation in an MHA layer. The compiler may generate the workload descriptors based on the maximum sequence length (i.e., the maximum number of tokens) that the transformer network supports. The workloads may correspond to various portions of the maximum sequence length, respectively. The NPU may dynamically scale the MHA layer during runtime. The input sequence length for an execution of the transformer network may be less than the maximum sequence length. The NPU may select one or more workloads from the plurality of workloads based on the sequence length and the workload descriptors pre-generated by the compiler. The NPU may execute the selected workload(s) and skip the other workload(s).
Owner:INTEL CORP

Fine-grained preemption of a data flow architecture based neural processing unit

PendingUS20260140764A1Program initiation/switchingProgram saving/restoringData streamExecution control
Fine-grained preemption of a data flow architecture based neural processing unit (NPU) includes executing, by a controller, control-code that implements a first context in the NPU. In response to the controller detecting a preemption opcode in the control-code, detecting, by the controller, a second context awaiting execution by the neural processing unit. The second context has a priority that is greater than a priority of the first context. In response to detecting the second context, the NPU switches from executing the first context to implementing the second context.
Owner:XILINX INC

Method and apparatus for utilizing external neural processor from graphics processor

ActiveUS12682226B2Computer hardwareGraphics
Aspects of the disclosure are directed to concurrent tensor processing with multiple processing engines. In accordance with one aspect, an apparatus including a common memory unit; a first processing engine coupled to the common memory unit, wherein the first processing engine is configured to access a portion of an input tensor and a portion of a kernel tensor from the common memory unit; and a second processing engine coupled to the common memory unit, wherein the first processing engine is further configured to send the portion of the input tensor and the portion of the kernel tensor to the second processing engine and wherein the second processing engine is configured to generate a portion of an output tensor based on the portion of the input tensor and on the portion of the kernel tensor.
Owner:QUALCOMM INC

Electronic device including electronic component for processing multi-modality and operation method and storage medium therefor

An embodiment of the present disclosure may provide an electronic device. The electronic device may comprise: a printed circuit board; a processor disposed on the printed circuit board and having multiple processing units embedded therein, the processing units including a neural processing unit; a first memory stacked on the processor in a package-on-package (PoP) manner; a second memory which is disposed on the printed circuit board, has a side surface positioned alongside one side surface of the processor, and includes multiple sub-memories, the second memory being a high bandwidth memory (HBM) stacked via a through-silicon via (TSV) on the printed circuit board on which the processor is disposed; and a memory management accelerator for optimizing bank allocation to the first memory and the second memory. Various other embodiments are possible.
Owner:SAMSUNG ELECTRONICS CO LTD

Portable keratitis real-time screening device based on lightweight edge computing

This invention provides a portable real-time keratitis screening device based on lightweight edge computing, belonging to the fields of medical auxiliary diagnostic equipment and computer vision technology. The device includes an optical acquisition unit, a control unit, and an interaction unit. The interaction unit receives image acquisition commands. The optical acquisition unit includes a macro camera module, which, based on the image acquisition commands, calls the macro camera module to acquire eye images of the target object. The control unit includes an edge development board, on which a lightweight image recognition algorithm is deployed on a neural processing unit hardware accelerator. The control unit performs real-time keratitis screening on the eye images based on the lightweight image recognition algorithm to obtain the keratitis screening results for the target object. This invention significantly improves inference efficiency, achieving millisecond-level keratitis screening, blocking intercepted or tampered paths, thereby mitigating the risk of data leakage.
Owner:XIAN UNIV OF POSTS & TELECOMM

Dual-sparse neural processing unit with multi-dimensional routing of non-zero values

A general matrix-matrix (GEMM) accelerator core includes first and second buffers, a control logic circuit, and a first processing element (PE). The first buffer receives a elements of a first matrix A of activation values. The second buffer receives b elements of a second matrix B of weight values. The control logic circuit replaces a zero-valued a element in a first column of the first buffer with a nonzero-valued a element that is within a maximum borrowing distance of a location of the zero-valued a element in the first column of the first buffer. The PE receives a elements from the first column of the first buffer including the nonzero-valued element a selected to replace the zero-valued a element and receives b elements from locations in the second buffer that correspond to locations in the first buffer from where the a elements have been received by the PE.
Owner:SAMSUNG ELECTRONICS CO LTD

Scheduling inferencing tasks on hardware resources

An apparatus and method for efficiently scheduling inference tasks for balancing performance and power consumption. In various implementations, a computing system includes a host processing circuit, a neural processing circuit, and an inferencing accelerator. Each of the neural processing circuit and the inferencing accelerator executes a respective machine learning data model. The inferencing accelerator includes less functionality and performance than the neural processing circuit while also consuming less power. When the operating mode requires lower power consumption, the host processing circuit compiles a low power consumption version of a first task and assigns it to the inferencing accelerator, rather than the neural processing circuit. If a second task has a single version that requires high performance, then the host processing circuit compiles the second task and assigns it to the neural processing circuit despite the operating mode indicating low power consumption.
Owner:ADVANCED MICRO DEVICES INC

Branching operations for neural processor circuits

PendingCN122311314AEngineeringTask segmentation
This disclosure relates to branching operations for neural processor circuitry. A neural processor includes a neural engine for performing a convolution operation on input data corresponding to one or more tasks to generate output data. The neural processor circuitry also includes data processor circuitry coupled to the one or more neural engines. The data processor circuitry receives the output data from the neural engines and generates branching commands from the output data. The neural processor circuitry also includes a task manager coupled to the data processor circuitry. The task manager receives the branching command from the data processor circuitry. The task manager enqueues one of two or more segmented branches according to the received branching command. The two or more segmented branches are after a pre-branching task segmentation that includes a pre-branching task. The task manager transfers the task from the selected segmented branch among these segmented branches to the data processor circuitry for execution of the task.
Owner:APPLE INC

Fine-grained preemption of a data flow architecture based neural processing unit

PCT designated stageWO2026111986A1Physical realisationProgram saving/restoringData streamExecution control
Fine-grained preemption of a data flow architecture based neural processing unit (NPU) includes executing, by a controller, control-code that implements a first context in the NPU. In response to the controller detecting a preemption opcode in the control-code, detecting, by the controller, a second context awaiting execution by the neural processing unit. The second context has a priority that is greater than a priority of the first context. In response to detecting the second context, the NPU switches from executing the first context to implementing the second context.
Owner:ADVANCED MICRO DEVICES INC +1

System for real-time decomposition of electrophysiological signals and classification of cardiac arrhythmias using adaptive neural signal processing

UndeterminedDE202026102265U1Medical automated diagnosisSensorsMicrocontrollerInstrumentation amplifier
A system for real-time decomposition of electrophysiological signals and arrhythmia classification using adaptive neural signal processing, comprising: a plurality of electrodes configured to acquire electrophysiological signals from a subject; an analog input stage consisting of an instrumentation amplifier, an impedance matching circuit, and an anti-aliasing filter for processing the acquired electrophysiological signals; an analog-to-digital converter unit operationally coupled to the analog input stage and configured to digitize the processed electrophysiological signals at a programmable sampling rate; a processing unit consisting of a microcontroller and a digital signal processor operationally coupled to each other via a communication bus; and a neural processing unit operationally coupled to the processing unit.a storage unit that stores executable instructions and trained neural parameters; wherein the processing unit and the neural processing unit are configured to jointly perform an adaptive decomposition of the digitized electrophysiological signals into a variety of intrinsic signal components based on dynamically updated baseline representations, extract temporal and morphological features from the intrinsic signal components, and classify the extracted features in real time into one or more arrhythmia categories.
Owner:EASWARI ENG COLLEGE +3

NPU and apparatus for transceiving feature map in a bitstream format

A neural processing unit (NPU) for decoding video or feature map is provided. The NPU may comprise at least one processing element (PE) to perform an inference using an artificial neural network. The at least one PE may be configured to receive and decode data included in a bitstream. The data included in the bitstream may comprise data of a base layer. Alternatively, the data included in the bitstream may comprise data of the base layer and data of at least one enhancement layer. The data of the base layer included in the bitstream may include a first feature map. The data of the at least one enhancement layer included in the bitstream may include a second feature map.
Owner:DEEPX CO LTD

Machine learning model security at a processor

PendingUS20260187254A1PathPingProcessing element
A processor protects a machine learning model (MLM) from unauthorized access. The processor employs a neural processing unit (NPU) to execute the MLM and implements decryption and encryption processes to decrypt the MLM and re-encrypt the MLM at different points along MLM storage and execution paths. Furthermore, the processor executes the encryption and decryption processes at different processing units and processing engines, thereby reducing the ability of malicious software to access the MLM. In addition, the processor protects buffers of the NPU from unauthorized access.
Owner:ADVANCED MICRO DEVICES INC +1