Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

214 results about "Hardware acceleration" patented technology

In computing, hardware acceleration is the use of computer hardware specially made to perform some functions more efficiently than is possible in software running on a general-purpose CPU. Any transformation of data or routine that can be computed, can be calculated purely in software running on a generic CPU, purely in custom-made hardware, or in some mix of both. An operation can be computed faster in application-specific hardware designed or programmed to compute the operation than specified in software and performed on a general-purpose computer processor. Each approach has advantages and disadvantages. The implementation of computing tasks in hardware to decrease latency and increase throughput is known as hardware acceleration.

Universal interface system and its operation method

This invention provides a universal interface system and its operation method. The universal interface system includes: an upper-layer application algorithm layer (4-1), configured to call functional modules and corresponding interfaces to complete a technical task; a fine-grained interface layer (4-2), configured to provide an abstract and unified fine-grained interface for the upper-layer application algorithm layer and define the input and output standards of each fine-grained interface; a chip adaptation layer (4-3), configured to carry out mapping between the interface specifications of the fine-grained interface and the hardware capabilities of the lower chip execution layer, and to implement the various functions defined by the fine-grained interface layer; and the chip execution layer (4-4; 4-5), configured to call chip hardware modules to perform hardware acceleration according to the technical task and return the execution results to the upper layer. This invention thus achieves the technical advantage of “one-time development, multi-chip adaptation”.
Owner:XEROPTIX INFINITY LTD

A hardware accelerator and acceleration method based on a vision transformer neural network

A hardware accelerator and acceleration method based on a VisionTransformer neural network, the accelerator is to deploy the VisionTransformer neural network on a ZYNQ development platform; the acceleration method is: an ARM processor stores a feature picture into a DDR memory, the read data is dispersed to an input cache and a weight cache, the processed feature picture is input to an on-chip cache unit, the processed data is sent to a PL end, the hardware IP of the PL end is configured, and read-write operation is performed at the same time, the final calculation result is obtained, written into the DDR memory, the data in the DDR memory is taken out and probability operation is completed, the probability operation result is transmitted to a PC, a PetaLinux operating system is transplanted to the hardware accelerator system of the VisionTransformer neural network, and a prediction result of inference is obtained from the PC; the neural network structure is optimized through the parallel method of multiple input and multiple output channels, the calculation speed is fast, the hardware resource occupation is low, the recognition accuracy is high, and the image classification task can be efficiently completed.
Owner:XIDIAN UNIV

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Method and device for evaluating full-system performance of hardware accelerator oriented to high-level synthesis

The embodiment of the application discloses a hardware accelerator whole system performance evaluation method and device facing high-level synthesis, which comprises the following steps: constructing a cycle-accurate and event-driven hardware accelerator simulation model based on the intermediate representation and static timing report generated by HLS; executing the hardware accelerator simulation model in the whole system simulator through the accelerator simulation object, so as to convert the timing and operation defined by the hardware accelerator simulation model into simulation events in the event queue of the whole system simulator and schedule and execute; collecting cycle-accurate simulation statistical data generated in the running process of the hardware accelerator simulation model in response to the scheduling and execution of the simulation events; and generating a performance evaluation report and a function verification result of the hardware accelerator based on the simulation statistical data and the function execution result of the hardware accelerator simulation model. The embodiment of the application can efficiently and accurately evaluate the performance of the hardware accelerator generated by HLS, so as to realize true hardware-software co-design and optimization.
Owner:UNIV OF SCI & TECH OF CHINA

Graph drive program-based super-division method, apparatus and device, medium and product

PendingCN122089573AAutomatically determineglobal optimization callGeometric image transformationProcessor architectures/configurationVideo memoryGraphics
The invention provides a super-division method based on a graph driving program, a super-division device based on the graph driving program, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: determining a target super-resolution parameter under the condition of determining to start image super-resolution processing based on a submission time interval of a final frame in computer equipment; based on the hardware operation information of the computer equipment, determining a target execution device used for executing image super-division processing in the computer equipment and a corresponding target super-division algorithm; the hardware operation information represents the rendering time consumption of the computer equipment, the occupation condition of a video memory and the use condition of a hardware acceleration unit; and performing image super-division processing on the submitted final frame based on the target super-division parameter, the target execution device and the target super-division algorithm to obtain a super-division frame.
Owner:MOORE THREADS TECH CO LTD

Satellite communication baseband processing method with flexible configuration function

The application discloses a satellite communication baseband processing method with flexible configuration function, relates to the technical field of satellite communication, and specifically a baseband processing method process is as follows: S1, interface adaptive adaptation scheduling; S2, downlink receiving scheduling; S3, uplink sending scheduling; S4, configurable hardware acceleration scheduling.The application can greatly improve the adaptive flexibility of baseband processing and the hardware multiplexing rate by executing four-step cooperative scheduling of downlink receiving scheduling, uplink sending scheduling, configurable hardware acceleration scheduling and interface adaptive adaptation scheduling, and all signal data are uniformly stored and scheduled with the memory as the core, compared with the fixed hardware pipeline architecture in the prior art, and therefore the technical problems of the prior satellite communication baseband processing, such as the need of modifying hardware circuits when switching satellite constellations or replacing radio frequency devices, high adaptation cost and low efficiency, can be solved.
Owner:QINGHUI ZHITONG (BEIJING) TECHNOLOGY CO LTD

A superconducting quantum bit quantum state reading method, device, equipment and medium

The present application relates to the field of quantum bit reading, in particular to a superconducting quantum bit quantum state reading method, device, equipment and medium, a reading driving signal is sent to a superconducting quantum chip; The reading driving signal is a variable frequency signal comprising the |0> state resonance frequency and the |1> state resonance frequency of the superconducting quantum bit; The reading feedback signal is collected from the superconducting quantum chip; According to the reading feedback signal, the pre-stored |0> state response signal and |1> state response signal corresponding to the reading driving signal, the quantum state of the superconducting quantum bit is determined by correlation analysis. The present application uses a variable frequency signal as the reading driving signal, which can reduce the sensitivity of the system to frequency-independent noise, improve the signal-to-noise ratio of the signal, and the correlation calculation does not need to be converted to the IQ plane, reducing the calculation steps, and the correlation analysis algorithm can usually use hardware acceleration, reducing the resource consumption of the solution, improving the operation efficiency.
Owner:YANGTZE DELTA IND INNOVATION CENT OF QUANTUM SCI & TECH

Hardware accelerator for matrix multiplication operations

A hardware accelerator is provided, including a data acquisition module and a matrix multiplication computation module. The data acquisition module acquires a first matrix forming a first data queue in a first dimension and a second dimension, and a second matrix forming a second data queue in the first dimension and a third dimension. The matrix multiplication computation module includes a three-dimensional array including a plurality of processing elements. Positions of the processing elements are determined based on first dimension values, second dimension values and third dimension values of the processing elements. The processing elements obtain corresponding first data from the first data queue based on the first dimension values and the second dimension values, obtain corresponding second data from the second data queue based on the first dimension values and the third dimension values, and perform matrix multiplication calculation based on the first data and the second data.
Owner:GLENFLY TECH CO LTD

Data processing methods, apparatus, and computer equipment based on butterfly networks

PendingCN122088578Areduce resource requirementsReduce storage overheadPhysical realisationData controlComplete data
This disclosure provides a data processing method, apparatus, and computer device based on butterfly networks, relating to the fields of artificial intelligence and hardware acceleration technology. The solution is as follows: In a switchable data processing architecture including an inverse butterfly network, a preprocessing module, a post-processing module, and a control signal generation module, when switching to aggregation operation, the preprocessing module inputs the first input data into the inverse butterfly network, the control signal generation module outputs the first control signal to complete data processing, and the post-processing module directly outputs the aggregation result. When switching to distribution operation, the preprocessing module performs input position mapping and filling on the second input data to obtain the third input data, the control signal generation module outputs the adapted second control signal to complete data processing, and the post-processing module performs rearrangement on the intermediate results to output the distribution result. This reduces hardware resource consumption, lowers control complexity and processing latency, and improves the hardware deployment adaptability of neural network sparse computing while ensuring functional integrity.
Owner:MOFFETT AI TECHNOLOGY SHENZHEN CO LTD

An FPGA-based DFM Pattern Match hardware accelerator and method

The application provides a DFM Pattern Match hardware accelerator based on FPGA and a method thereof, which comprises a data distribution module, a hash calculation module, a hash comparison module, a candidate FIFO, a target area generation module, and an edge screening module.The data distribution module is used for receiving and transmitting different groups of edge data to the hash calculation module in parallel.The hash calculation module is used for performing direction bucket division and calculating corresponding hash values in parallel for each edge.The hash comparison module is used for performing parallel comparison between the hash values of each direction bucket and the hash values of a target pattern, and generating candidate items.The candidate FIFO is used for decoupling the pre-screening stage and the accurate geometric verification stage.The target area generation module is used for obtaining a target search area.The edge screening module is used for screening all edges located in the target search area from a layout.The accurate matching module is used for mapping the edges output by the edge screening module to a target pattern coordinate system, and performing parallel comparison between the edges and each edge of the target pattern, so as to determine whether the target search area is consistent with the target pattern.The application can efficiently perform pattern matching on a circuit layout in integrated circuit physical verification.
Owner:SHANGHAI LIXIN SOFTWARE TECH CO LTD

High-density single-fingered eDRAM memory-computing integrated macro with bit-level sparse perception and kernel-level weight update

The invention belongs to the field of artificial intelligence hardware acceleration, and discloses a high-density single-fingered eDRAM storage and calculation integrated macro with bit-level sparse perception and kernel-level weight update, which comprises an input buffer, a weight buffer, a word line decoder, a bit line driver, a single-fingered eDRAM array and an output buffer, the weight buffer is used for writing weights from the off-chip and writing the weights into the single-finger eDRAM array line by line, and each single-finger eDRAM storage unit only stores one bit; the input buffer is used for writing an input activation value and transmitting a multi-bit input activation value into the single-fingered eDRAM array bit by bit from a low bit to a high bit, each row of single-fingered eDRAM is equipped with a bit saliency perception analog-to-digital converter BSA-ADC and is used for executing matrix vector multiplication and transmitting a calculation result to a calculation line CL, the BSA-ADC digitalizes an accumulation result on the calculation line, and the calculation line CL is connected with the input buffer. The result is transmitted to the output buffer; the output buffer is used for outputting the calculation result. The integrated macro has the advantages of high density, low energy consumption and high throughput.
Owner:INSTITUTE FOR ADVANCED STUDY OF THE UNIVERSITY OF MACAU IN HENGQIN GUANGDONG-MACAU DEEP COOP ZONE (INSTITUTE FOR ADVANCED STUDY OF THE UNIVERSITY OF MACAU IN HENGQIN)

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Controllers in data processing engine columns

Embodiments herein describe a hardware accelerator with an array of data processing engines (DPEs) which includes a controller (e.g., a microcontroller) for multiple columns of the array. The controllers can be hardened circuitry that executes software code (or firmware) that controls the hardware accelerator. In one embodiment, the task of the controller is to control and orchestrate the functions performed by the hardware accelerator.
Owner:XILINX INC

A reconfigurable baseband processing SoC architecture based on a RISC-V extended instruction set and an instruction control method thereof

The present application relates to a kind of reconfigurable baseband processing SoC architecture based on RISC-V extension instruction set and its instruction control method, belong to communication chip and processor design technical field.For the instruction set flexibility insufficient, data handling inefficiency and dynamic reconfiguration and high energy efficiency difficult to consider of existing 5G / 6G baseband processor, the present application proposes a kind of heterogeneous computing architecture, including the RISC-V processor core cluster of supporting RV64IMAFDCV instruction set, extension instruction execution unit, vector coprocessor and reconfigurable hardware acceleration module group.By adding special instructions such as complex matrix multiplication, optimizing data path design, using double buffering mechanism and hierarchical reconfigurable technology, efficient execution of baseband processing is realized.At the same time, dynamic voltage frequency regulation system can adjust power supply parameters in real time according to work load.The scheme significantly improves the performance and energy efficiency of baseband processing, and is suitable for physical layer signal processing of 5G / 6G communication system.
Owner:GUANGDONG COMM & NETWORKS INST +1

Low-power convolutional neural network hardware acceleration method based on stochastic computing

This invention discloses a low-power convolutional neural network hardware acceleration method based on random computation. First, a convolutional neural network model is constructed and trained. Then, during the inference phase, the trained model is deployed to a hardware computing unit, and random computation is used to partially or completely replace the convolutional computations in the model. The random computation includes three stages: encoding, computation, and decoding. In the encoding stage, the input feature map is scaled, and the convolutional kernel is quantized. The gray values ​​of the convolutional kernel and each point in the input feature map are normalized into probability values, and an encoding sequence is generated for each probability value. In the computation stage, the two encoding sequences of the convolutional kernel and corresponding points in the input feature map are subjected to a traversal AND operation to generate an AND operation sequence. In the decoding stage, the AND operation sequence is decoded to obtain the computation result. Without significantly affecting model performance, this method reduces computational complexity and hardware resource consumption, making it suitable for scenarios with limited power consumption and resources.
Owner:HEBEI UNIV OF TECH

A transform domain based eigenvector similarity measure method and apparatus

The present application relates to the technical field of feature vector merging in large model inference acceleration, in particular to a feature vector similarity measurement method and device based on transformation domain, by constructing a new type of calculation path of "transformation-measurement", using an orthogonal transformation matrix to project the feature vector to the transformation domain and then calculating the absolute value error sum, effectively extracting the internal structured features, solving the defect of insufficient accuracy of traditional lightweight measurement indicators, designing a multiplierless hardware acceleration architecture supporting the algorithm, by strictly constraining the element values of the orthogonal transformation matrix within the integer set of {-1, 0, 1}, completely equivalent to the matrix multiplication into simple addition and subtraction operation, completely avoiding multiplication, division and square root operation, greatly reducing the chip area and power consumption overhead, in addition, the pareto optimality of model accuracy and hardware efficiency is also realized, and the configuration capability of searching and customizing the transformation matrix for specific models is also provided.
Owner:NINGBO DOU ZHUAN XIN SHI TECHNOLOGY CO LTD

Resource scheduling method and electronic device

The application provides a resource scheduling method and an electronic device, which can be applied to the technical field of computers. The resource scheduling method comprises the following steps: in response to a resource request for a target task, obtaining resource utilization scores of nodes and hardware accelerators in a cluster according to resource usage amounts of the hardware accelerators, resource allocation amounts of the nodes, and a resource demand amount carried by the resource request; determining a target node in the nodes and a target hardware accelerator of the target node according to the resource utilization scores and a scheduling strategy carried by the resource request, the scheduling strategy indicating a scheduling rule matched with resource integration requirements and / or load balancing requirements of the target task; virtually dividing the target hardware accelerator according to the resource demand amount to obtain a virtual hardware accelerator; and binding the virtual hardware accelerator to a container of the target node to execute the target task by using the container.
Owner:DAWNING INT INFORMATION IND CO LTD +1

Transform hardware acceleration method and accelerator based on hybrid precision quantization and huffman coding

This invention discloses a hardware acceleration method and accelerator for Transformer based on mixed-precision quantization and Huffman coding. The acceleration method includes: using a genetic algorithm to obtain several configuration schemes for mixed-precision quantization of Transformer network layers; performing mixed-precision quantization on each Transformer network layer based on each quantization configuration scheme to obtain a corresponding KL divergence; training a multilayer perceptron to obtain a quantization configuration prediction network using the quantization configuration scheme and the corresponding KL divergence as the output label and input feature, respectively; receiving a user-set target KL divergence value, using the quantization configuration prediction network to obtain the corresponding quantization configuration scheme, and performing mixed-precision quantization on each network layer based on the quantization scheme; and using Huffman coding to encode and compress all quantization weights before on-chip storage. This invention can reduce storage and computational overhead while maintaining model accuracy.
Owner:HUNAN NORMAL UNIVERSITY

Computing dot products at hardware accelerator

A computing device, including a hardware accelerator configured to train a machine learning model by computing a first product matrix including a plurality of first dot products. Computing the first product matrix may include receiving a first matrix including a plurality of first vectors and a second matrix including a plurality of second vectors. Each first vector may include a first shared exponent and a plurality of first vector elements. Each second vector may include a second shared exponent and a plurality of second vector elements. For each first vector, computing the first product matrix may further include computing the first dot product of the first vector and a second vector. The first dot product may include a first dot product exponent, a first dot product sign, and a first dot product mantissa. Training the first machine learning model may further include storing the first product matrix in memory.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

SLAM back-end optimization method based on sparse matrix decomposition and hardware accelerator architecture

The invention relates to an SLAM back-end optimization method based on sparse matrix factorization and a hardware accelerator architecture, and the method comprises the steps: constructing a Hessian matrix through employing a binary sparse coding mechanism according to the actual observation coordinates of feature points on an image plane and the predicted coordinates of road sign points projected to a camera coordinate system; performing Schel elimination processing on the Hessian matrix based on binary sparse coding to obtain a linear equation set only containing a camera pose variable; and carrying out extended Kalman filtering on a linear equation set only containing a camera pose variable to realize incremental updating of a state pose and a road sign point coordinate. According to the method, the overall delay of SLAM rear-end optimization can be remarkably reduced, and the real-time requirement in a high-speed moving scene is met.
Owner:SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI

A method and system for verifying the trusted state of a virtual machine based on TIPU

PendingCN122365514AMemory addressTerm memory
This invention relates to the intersection of cloud computing security, trusted computing, and hardware acceleration technologies, and discloses a virtual machine trusted state verification method and system based on TIPU. The method includes: configuring a measurement task on the host side of the TIPU, the measurement task containing the memory address information of the target virtual machine and its corresponding expected hash value; the TIPU directly reading the memory data of the target virtual machine via DMA based on the memory address information to generate an actual hash value; the TIPU comparing the actual hash value with the expected hash value to determine whether the target virtual machine is in a trusted state; when the virtual machine is determined to be in an untrusted state, the TIPU calls a built-in root of trust to sign the verification result and generate a remote proof report. This invention obtains the mapping table from the client's physical address to the host's physical address through the virtualization platform interface via a host agent and pre-configures it to the TIPU. The TIPU only accesses the target virtual machine's memory, achieving virtual machine context awareness and effectively avoiding cross-virtual machine information leakage.
Owner:TIANFU JIANGXI LAB

Computer-implemented method and system for memory planning for in-memory computing-based hardware accelerators

The invention relates to a computer-implemented method for memory planning for in-memory computing-based hardware accelerators, comprising the steps of: providing (S1) a neural network model (10) with multiple layers (12), wherein each layer (12) is allocated memory buffers (14) for input and output data; sorting (S2) the memory buffers (14) to minimize data transfer between an S-RAM memory (16) and a D-RAM memory (18), wherein the memory buffers (14) are sorted according to transfer requirements, with memory buffers (14) with high transfer requirements being prioritized; allocating (S3) the memory buffers (14) to the S-RAM memory (16) or the D-RAM memory (18) based on a sorted memory buffer list priority; and optimizing (S4) a memory allocation such that access to the D-RAM memory (18) is reduced and utilization of the S-RAM memory (16) is increased.The invention further relates to a system (1) for memory planning for in-memory computing-based hardware accelerators.
Owner:ROBERT BOSCH GMBH

A hardware-accelerated embedded and vector search integrated SSD storage device for a retrieval-enhanced generation system

The application relates to the cross field of SSD storage technology and retrieval enhancement generation technology, in particular to a hardware acceleration embedding and vector retrieval integrated SSD storage device for a retrieval enhancement generation system, which comprises an interface module, a hardware Embedding module, a vector storage module, a hardware retrieval module, a control module and a NAND flash storage module, the interface module, the hardware Embedding module, the vector storage module, the hardware retrieval module, the control module and the NAND flash storage module are coupled and connected through a host internal bus, the interface module adopts a PCIe5.0 interface and is used for high-speed data interaction with a host, receives preprocessed text data blocks and user query texts sent by the host, and feeds back candidate text data or a storage address of the candidate text data obtained through retrieval to the host. The device overcomes the defects of the prior art, such as dependence of RAG system text vectorization and vector retrieval on a host software, low efficiency, high delay and insufficient interface rate of a traditional storage device.
Owner:SHENZHEN HUANYIN TECH CO LTD

Hardware acceleration of relational operations

This disclosure describes an implementation of a computing system that utilizes accelerator hardware configured to perform parallel computation of matching pairs of a join operation of a database program. This computation is performed at least in part by, for each thread that determines a matching pair, computing an offset in an output tuple by adding a global rank of a respective block, an intra block rank of a respective warp, and an intra warp rank of a respective thread. This computation is further performed by storing index values for a primary key and a foreign key in a matching pair at the computed offset location in an output tuple, and outputting the output tuple.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A multi-source picture rendering system serving a clouded collaborative application

PendingCN122457587AComputer hardwareRendering hardware
The application provides a multi-source picture rendering system for cloud-based collaborative application, which comprises an application front end, a front-end rendering engine, one or more cloud virtual machines, a cloud-end rendering engine and a rendering scheduling module. The application front end is used for receiving user operation instructions, providing interfaces and interactive components to users; the front-end rendering engine is used for completing rendering of model data from the cloud virtual machine received by the application front end based on local rendering hardware acceleration technology; the cloud virtual machine is used for running an application instance and receiving user operation instructions for execution; the cloud-end rendering engine is used for transmitting the rendered three-dimensional model to the application front end in the form of a video stream after completing rendering based on video streaming technology; and the rendering scheduling module is used for intelligently deciding which rendering mode is used to complete rendering of each part of the current model. The application adopts a dual-mode overall fusion rendering architecture, combines video streaming and local rendering hardware acceleration, and thus realizes intelligent scheduling rendering in multiple application scenarios.
Owner:SHIWEN (ZHEJIANG) TECHNOLOGY CO LTD

A hardware acceleration system and method supporting matrix and tensor computation

PendingCN122433821AComputer hardwareData stream
The application discloses a kind of hardware acceleration system and method supporting matrix and tensor calculation, belong to high-performance computing and artificial intelligence chip design field.The system is based on the computing core array of annular interconnection, integrated uniform scheduling and variable unit USTU and configurable data flow controller CDC.USTU is responsible for the real-time analysis conversion of high-level tensor operation to optimized matrix multiplication GEMM subtask and completes uniform fine-grained scheduling;CDC dynamically configures the data flow mode of computing core array, storage access mode and inter-core communication logic according to the characteristics of subtask.The application realizes the adaptive mapping of neural network operator to bottom hardware data flow, can efficiently execute various tensor calculations, significantly improves hardware utilization, energy efficiency ratio and application flexibility.
Owner:SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD

A 2-group signed tensor computing circuit structure based on 6-bit approximate full adder

ActiveCN115840556BBinary multiplierNeural network hardware
The application discloses a 2-group signed tensor calculation circuit structure based on a 6-bit approximate full adder, relates to the field of neural network hardware acceleration, and comprises a 6-bit approximate full adder module, a signed 8*8 approximate multiplier circuit structure based on the 6-bit approximate full adder and the 2-group signed tensor calculation circuit structure based on the 6-bit approximate full adder. Since the neural network accelerator can sacrifice part of the accuracy of data in exchange for the optimization of delay, area and power consumption of a circuit structure, the 6-bit approximate full adder module ignores part of the carry, thereby reducing the circuit area and lowering the circuit power consumption, the signed 8*8 multiplier calculation process and the 2-group signed tensor calculation process are optimized by using the 6-bit approximate full adder module, approximate calculation is introduced at some positions, and in exchange for the improvement of area and power consumption of the circuit structure, part of the accuracy is lost.
Owner:SOUTHEAST UNIV

A joint control SoC architecture integrating multi-precision neural network compensation and online learning engine

This invention relates to the fields of integrated circuit design, robot motion control, and embodied intelligent hardware acceleration technology. Specifically, it is a joint control SoC architecture integrating multi-precision neural network compensation and an online learning engine. This architecture includes a main control subsystem, a real-time kernel subsystem, an intelligent compensation subsystem, a data preprocessing unit, and an online learning engine integrated within a single chip via a dedicated interconnect bus. A RISC-V processor is used for high-level control, a hardware-embedded FOC engine achieves high-frequency real-time closed-loop control, a lightweight NPU performs nonlinear compensation inference, the data preprocessing unit achieves nanosecond-level alignment of multi-precision data, and the online learning engine achieves non-blocking hardware-autonomous weight updates through shadow weight SRAM. This invention significantly improves the real-time performance of robot joint control, solves the precision degradation problem caused by mechanical wear during long-term operation, and achieves efficient collaboration between AI algorithms and industrial control algorithms. It is suitable for integrated joint control of humanoid robots and other intelligent devices.
Owner:SHANGHAI SHANYI MICROELECTRONICS TECHNOLOGY CO LTD

General-purpose data partition hardware accelerator

PendingCN122285082AComputer hardwareData set
This invention discloses a general-purpose data partitioning hardware accelerator, tightly coupled to a processor core, for partitioning an input dataset containing multiple data records into multiple output partitions. The accelerator includes a length-aware read and distribute module and multiple parallel processing engines. The length-aware read and distribute module reads the data stream and decoupled metadata stream in parallel, parses variable-length record boundaries in real time, and distributes records. The processing engines receive records, calculate partition keys using a variable-length key pipeline hash unit to determine the target partition, and use a length-aware write merging unit to merge multiple records pointing to the same partition on-chip into cache line-aligned data blocks before writing them to the storage system. Therefore, this invention natively supports variable-length data, improving system throughput and versatility in a tightly coupled environment.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI