Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

725 results about "Hardware implementations" patented technology

Power domain control software architecture and controller thereof

The invention provides a power domain control software architecture and a controller thereof, the software architecture comprises an application layer developed based on an AUTOSAR standard, a runtime environment and a basic software layer, the application layer comprises a plurality of modularized software components, the software components package service logic functions of the controller to be integrated, the software components are provided with an AUTOSAR-based ARXML file definition interface, and the AUTOSAR-based ARXML file definition interface is used for defining the service logic functions of the controller to be integrated. According to the method, the interface compatibility can be improved, seamless integration of application layers across controllers can be realized, and the development parallelism is enhanced by kernel deployment of independently packaged software components according to the real-time performance of functions. The runtime environment generates an inter-core communication agent code, constructs a virtual bus and routes a service request of an application layer component to the basic software layer, and the basic software layer provides system-level service for the application layer through a virtual bus interface of the runtime environment, so that the application layer does not need to directly operate hardware, software and hardware decoupling is realized, and the service performance of the application layer is improved. And smooth realization of software fusion and integration is promoted.
Owner:ZHEJIANG FARIZON ZHIXIN TECHNOLOGY CO LTD +3

Signal sampling method and system for multi-channel signal converter

The invention discloses a signal sampling method for a multi-channel signal converter, and the method comprises the steps: carrying out the synchronous and parallel sampling of an analog signal source according to a multi-channel signal collection module, so as to obtain multi-channel original signal sequence data; performing inter-channel time delay calibration and amplitude normalization processing on the multichannel original signal sequence data to obtain a calibrated multichannel signal data set; the invention relates to the technical field of signal processing. According to the signal sampling method for the multi-channel signal converter, the multi-channel signal converter improves the data consistency and accuracy through the technologies of synchronous sampling, time delay calibration, amplitude normalization and the like, meanwhile, precision adjustment of self-adaptive filtering, dynamic threshold segmentation and entropy driving is adopted, signal interference and abnormal data are reduced, and the signal quality is improved. The non-uniform resampling and compression technology is adopted to optimize the storage and transmission efficiency, key data is ensured to be transmitted preferentially, hardware is accelerated through an FPGA, and the real-time performance and the processing capacity of the system are improved.
Owner:JINING MINING GRP HAINA TECH ELECTROMECHANICAL CO

Wafer level chip system design space construction and rapid parameter search method

The invention discloses a wafer level chip system design space construction and rapid parameter search method. The method comprises the steps of task graph definition, prefabricated member parameter quantification, joint model construction based on task graph features and prefabricated member features, and Bayesian optimization implementation strategy. According to the technical scheme, construction and search of the design space can be guided, so that the number of times of calling a simulator is reduced, and construction of the design space of the wafer-level chip system architecture, efficient parameter search and performance evaluation are achieved. The method can be used for agile design and hardware implementation of a multi-chip integrated wafer-level chip system architecture.
Owner:XI AN JIAOTONG UNIV

TFHE-oriented programmable bootstrap hardware implementation system

The invention discloses a TFHE-oriented programmable bootstrap hardware implementation system, and the system comprises a vector calculation module which is used for executing an analog-to-digital switching step in a TFHE-oriented programmable bootstrap algorithm on an LWE ciphertext; the rotary register is used for caching and rearranging the LWE ciphertext; the function calculation module is used for executing a blind rotation step in a TFHE-oriented programmable bootstrap algorithm on the rearranged LWE ciphertext and the analog-to-digital switching result to obtain a blind rotation result; the vector calculation module is also used for executing a sample extraction step and a key switching step in a TFHE-oriented programmable bootstrap algorithm on the blind rotation result to obtain a decrypted ciphertext; when the function calculation module executes the blind rotation step, the polynomial in the rearranged LWE ciphertext and the analog-digital switching result is decomposed into a plurality of sub-polynomials through a number theory conversion unit, and the sub-polynomials are subjected to parallel calculation. The method can reduce the use amount of hardware units, and is high in decryption success rate, high in performance and small in area.
Owner:TSINGHUA UNIVERSITY

Optimizations for analog hardware realization of trained neural networks

Systems and methods are provided for analog hardware realization of neural networks. The method includes obtaining a neural network topology and weights of a trained neural network. The method also includes transforming the neural network topology to an equivalent analog network of analog components including operational amplifiers and resistors. Each operational amplifier represents an analog neuron of the equivalent analog network, and each resistor represents a connection between two analog neurons. The method also includes computing a weight matrix based on the weights of the trained neural network. The method also includes generating a resistance matrix for the weight matrix. The method also includes pruning the equivalent analog network to reduce the number of operational amplifiers or the resistors, based on the resistance matrix, to obtain an optimized analog network of analog components.
Owner:POLYN TECHNOLOGY LIMITED

Equipment access method, device and equipment based on address space identifier

The invention discloses an equipment access method, device and equipment based on address space identification. The method is applied to simplifying an equipment memory management unit and comprises the steps of establishing each single-level range table; acquiring an equipment access request, reading an address space identifier and a virtual address from the equipment access request, and matching a corresponding target range table from each single-level range table based on the address space identifier; and calculating a target physical address according to the target range table and the virtual address, and establishing equipment access based on the target physical address. By replacing a traditional multi-level page table with a single-level structure, the data structure design of address mapping is simplified, the dynamic configuration complexity is reduced, and hardware formal verification is facilitated. Through direct matching of the address space identifier, the traditional steps of analyzing the equipment identity identifier and querying the external table item are omitted, and the external memory access frequency is reduced. By calculating the target physical address and establishing access, the hardware address conversion logic is simplified, the hardware implementation cost and power consumption are reduced, and the data transmission efficiency is improved.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

SoC interrupt processing method based on hardware Sequence

The invention relates to the technical field of system-on-chip service interrupt processing, in particular to a hardware Sequence-based SoC interrupt processing method, which comprises the following steps of: pre-programming an interrupt processing instruction sequence into an instruction random access memory (RAM), and configuring an instruction set, a sequence processor and an interrupt priority strategy of a task scheduler Sequence; receiving the interrupt signal by a task scheduler, and selecting a corresponding sequence processor according to an interrupt type; for the interrupt with low priority, the Sequencer directly executes the pre-stored instruction sequence to complete the interrupt response; and for high-priority interruption, the Sequencer caches the state of the register and sends a message packet to notify the CPU to perform cooperative processing. According to the method, interrupt autonomous response is achieved through hardware Sequence, zero CPU intervention is achieved in low-priority interrupt, a lightweight message mechanism is adopted in high-priority interrupt, and bandwidth occupation and power consumption are reduced while efficiency is improved.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Method and apparatus for supporting distributed graphics and compute engines and synchronization in multi-dielet parallel processor architectures

This disclosure describes supporting distributed graphics and compute engines in a multi-dielet processor, such as, for example, a multi-dielet graphics processing unit (GPU), architectures and synchronization in such architectures. Each multi-dielet processor includes a hardware-implemented remapping capability and / or a hardware-implemented memory barrier capability.
Owner:NVIDIA CORP

Heterogeneous multi-core chip power consumption state control system and heterogeneous multi-core chip

According to the heterogeneous multi-core chip power consumption state control system and the heterogeneous multi-core chip provided by the invention, the power consumption state request of each sub power supply domain is received through the set system hardware power consumption management module, and the driving signal is sent to the corresponding sub power supply domain according to the preset hardware time sequence, so that the corresponding sub power supply domain enters the target state; generating a PMU control signal; and the power management module drives an internal voltage regulator to execute corresponding regulation operation according to the PMU control signal. A special control processor (SCP) is not needed, the system hardware power consumption management module realized by pure hardware is adopted to directly control the power consumption state of each sub power supply domain in the heterogeneous multi-core SoC, the software processing delay is eliminated, the response speed is high, the static power consumption of the power consumption management module is remarkably reduced, and the system is particularly suitable for an ultra-low power consumption scene and has a wide application prospect. And hardware-level surge current protection is realized in a multi-core power-on process.
Owner:SHANGHAI WU QI MICROELECTRONICS CO LTD +1

Random calculation processing unit and method based on adaptive compensation mechanism

ActiveCN120653304AMachine execution arrangementsStochastic computingOperand
The invention relates to a random calculation processing unit and method based on a self-adaptive compensation mechanism, and the method employs a self-adaptive compensation core formed by solidifying a lightweight neural network to replace a conventional solidification compensation rule, and the self-adaptive compensation core can carry out the random calculation according to two operands participating in multiplication. And an optimal compensation parameter is dynamically predicted in real time for compensation. According to the scheme, the calculation error of random calculation is accurately, continuously and individually compensated in a data driving mode, and hardware implementation is performed by adopting a constant coefficient multiplier technology, so that the self-adaptive compensation core has extremely low area and power consumption overhead in hardware. Compared with the prior art, the method has the advantages that the fidelity of random calculation can be greatly improved on the premise that the hardware cost is not remarkably increased, and a new effective way is provided for constructing a high-performance and high-energy-efficiency random calculation neural network accelerator.
Owner:NAT UNIV OF DEFENSE TECH

Multi-cause regulation AI large model training method and intelligent decision-making system

The invention relates to the technical field of multi-modal feature fusion, in particular to a training method of a multi-cause regulation AI large model and an intelligent decision making system. Through data preprocessing, feature expression optimization, class center construction of support data, efficient distance calculation, dynamic parameter adjustment and a self-adaptive hyper-parameter and field self-adaptive strategy, the multi-cause regulation AI large model is obtained; according to the technical scheme, the robustness and accuracy of model training are remarkably improved. Meanwhile, the modularized hardware implementation scheme ensures the high efficiency and stability of the overall operation of the system, and can adapt to the requirements of the data size and the field change, thereby providing an efficient, stable and intelligent neural network model training platform for an intelligent decision-making system.
Owner:周明全

Floating-point number index maximum value searching circuit and in-memory computing chip

The invention discloses a floating-point number index maximum value searching circuit and an in-memory computing chip. The searching circuit comprises an SRAM (Static Random Access Memory) unit matrix which comprises a plurality of rows and columns of SRAM units, wherein each unit comprises NMOS (N-channel Metal Oxide Semiconductor) tubes N1-N5 and PMOS (P-channel Metal Oxide Semiconductor) tubes P1 and P2; n1, N2, P1 and P2 are in reverse cross coupling to form a pair of storage nodes Q and QB; the grid electrode of the N5 is connected with a node QB; the grid electrode of the N3 in the ith row is connected with a word line WLLlt; igt; the grid electrode of the N4 is connected with a word line WLRlt; igt; ; the drain electrode of each column of N5 is connected in series with the source electrode of the next column of N5, and the source electrode of the first column of N5 is connected with a signal SELlt; igt; the N5 drain electrode of the last column is connected with a signal PRE; in the jth column, N3 is a corresponding node Q and a bit line BLlt; jgt; the transmission tube N4 is a corresponding node QB and a bit line BLBlt; jgt; and a transmission tube. According to the invention, the number of searching cycles is saved, the calculation time is shortened, the subsequent calculation amount is reduced, the efficiency is improved, the power consumption is reduced, and the defects of relatively high power consumption, complex hardware implementation, relatively long calculation delay and the like of an existing circuit are overcome.
Owner:ANHUI UNIV

Hardware implementations of activation functions in neural networks

Circuitry for performing neural-network calculations includes a plurality of compute circuits, arranged in parallel with respective inputs and outputs, to receive function arguments for a node of a neural network on their respective inputs, compute values of a plurality of activation functions using the function arguments, and provide the values on their respective outputs. Each compute circuit of the plurality of compute circuits is to compute the values of a respective activation function of the plurality of activation functions. The circuitry also includes a multiplexor to select between the respective outputs of the plurality of compute circuits and to provide the values on a selected output as activation-function values for the node of the neural network, based on an activation-function selection signal.
Owner:NATARAJ BINDIGANAVALE S

Instruction pipeline processing method of processor and processor

The invention provides an instruction pipeline processing method of a processor and the processor. A processor includes: a control module; the control module is set to execute the branch instruction speculatively according to the branch prediction direction when detecting that the first instruction subjected to initial decoding is the branch instruction, and execute the branch instruction if detecting that the second instruction subjected to initial decoding is the function call instruction or the function return instruction before the speculation execution result of the first instruction is generated. If yes, pausing all operations after the instruction processing assembly line performs initial decoding on the second instruction, and blocking the instruction fetching operation of the instruction processing assembly line on the next instruction until a speculation execution result of the first instruction is obtained; wherein the operation after the initial decoding comprises the step of carrying out a push-in or push-out operation of a return address stack (RAS) according to the second instruction. According to the technical scheme, the pollution risk caused by speculative execution of the branch instruction to the return address stack can be shielded, so that the hardware implementation logic of the return address stack is simplified, and the circuit area and power consumption are saved.
Owner:SUNMMIO SCIENCE & TECHNOLOGY (BEIJING) CO LTD

Ethernet-based CXL protocol extension method, device and system

The invention discloses a CXL protocol extension method, device and system based on the Ethernet. The method comprises the following steps of: carrying out long-distance expansion interconnection on a CXL (Compute Express Link) protocol by using Ethernet as a physical transmission medium; by bearing the CXL protocol on the Ethernet, the range of memory pooling is expanded from the rack level to the whole data center (kilometer level), and the potential of the CXL technology is thoroughly released; through protocol translation, label management and flow control realized by pure hardware, an operating system kernel and a software protocol stack are bypassed, so that the end-to-end memory access delay is far lower than that of any software-based network storage scheme, and a reliable request-response tracking and timeout retransmission mechanism is realized through a hardware label manager. By distinguishing the control flow and the data flow, differentiated QoS guarantee is provided for different types of memory flows, and the stability of the system under high load is ensured.
Owner:SHENZHEN UNIVERSITY OF ADVANCED TECHNOLOGY

Improved plantard algorithm-based lattice cipher modular multiplication method and modular multiplier

The invention discloses a lattice cryptographic modular multiplication method and a modular multiplier based on an improved plantard algorithm, which eliminate the post-processing overhead of the traditional modular multiplication method by preprocessing twiddle factors in the level of the modular multiplication method. In hardware implementation, shift-addition is adopted to replace a complex multiplier, a dynamic bit width truncation technology is combined, the bit width of intermediate data is compressed from 39 bits to 21 bits, and logic resource consumption is greatly reduced, so that modular multiplication operation in a lattice password digital signature scheme can be quickly and efficiently completed.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Test device for modular multi-level converter

Disclosed is a test device to test the operation of an MMC by fabricating only one or several SMs constituting the MMC testing under real-time operating conditions. The test device for a modular multi-level converter (MMC) includes a simulation model of an MMC having at least one arm to which sub modules (SMs) are serially connected, at least one test target SM among the serially connected SMs being replaced with a dependent voltage source; an arm current simulation circuit including an equivalent SM that implements the test target SM as actual physical hardware and an inverter that supplies current to the equivalent SM; and a control unit configured to control the arm current simulation circuit to correspond to an operation of the simulation model and set a voltage corresponding to a charge / discharge voltage of the equivalent SM of the arm current simulation circuit to the dependent voltage source.
Owner:HONGIK UNIV IND ACAD COOP FOUND

Hybrid speculative decoding system with models on silicon

A speculative decoding system may include integrated circuits (ICs), a router, and a processing unit. The ICs may implement different models that can perform different types of tasks. The router may route an input prompt, which may include one or more input tokens, to an IC based on the task to be performed using the input prompt. The IC may include hardware implementations of operators in a model. The IC may generate speculative token(s) from the input prompt by running the operators in the model. The speculative token(s) may be drafted to the processing unit. The processing unit may validate the speculative token(s) and generate output token(s) by executing another model, which may be larger than the model executed by the IC. The processing unit may validate multiple speculative tokens in parallel. Key-value pairs generated by the IC may be used by the processing unit for executing the other model.
Owner:INTEL CORP

GPU system verification method and device, electronic equipment, storage medium and program product

The invention relates to a GPU system verification method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: obtaining a target random test excitation sequence; respectively inputting the target random test excitation sequence into the GPU function model and the register transfer level model; wherein the GPU function model is a software model used for simulating functions of the GPU; the register transfer level model is a model written based on a hardware description language, and the register transfer level model is used for describing hardware implementation and behaviors of the GPU; obtaining a function simulation result output by the GPU function model in response to the target random test excitation sequence and a simulation result output by the register transfer level model in response to the target random test excitation sequence; and comparing the function simulation result with the simulation result to obtain a first random verification result of the GPU system. According to the method and the device, the GPU system verification flexibility can be remarkably improved.
Owner:MOORE THREADS TECH CO LTD

Structured channel pruning target recognition lightweight method and device and medium

The invention discloses a structured channel pruning target recognition lightweight method, which is used for lightweight reconstruction of a target detection algorithm model, and comprises the following steps executed by computer equipment: S1, in a network model training process, taking a scaling factor of a batch normalization layer as an importance degree standard for measuring a convolution channel, performing linear transformation on each output feature image pixel of the convolutional layer to accelerate model network convergence; s2, calculating a cutting threshold according to the maximum value of the scale factor of the batch normalization layer, and deleting redundant weight connections of which the model performance influence is lower than the cutting threshold through a channel structure pruning compression operation so as to reduce the model parameter quantity and the calculation quantity; and S3, through network re-training fine tuning, model lightweight is realized. According to the method, coding and other operations are not needed, software and hardware implementation and deployment are facilitated, the detection precision can be almost not affected, the comprehensive performance of the model is improved, and the effect of model lightweight is achieved.
Owner:CHINESE PEOPLES LIBERATION ARMY ARMY ARTILLERY & AIR DEFENSE ACAD

A method and apparatus for P-frame and / or B-frame image block level rate control

This invention discloses a bitrate control method at the image block level for P-frames and / or B-frames. Step S1: Obtain the inter-frame coding mode candidates and their prediction costs for each image block to be encoded. Step S2: Select the minimum prediction cost to characterize the inter-frame coding complexity of the image block to be encoded; use the sum of the inter-frame coding complexities of all image blocks within a video frame as the inter-frame coding complexity of that video frame. Step S3: Calculate the target number of coding bits for each image block to be encoded within the P-frame or B-frame to be encoded. Step S4: Calculate the Lagrange multipliers for the image block to be encoded within the P-frame or B-frame to be encoded. Step S5: Perform video encoding on the image block to be encoded, and then adjust the target number of coding bits for the next image block to be encoded within the P-frame or B-frame to be encoded. This invention has low hardware overhead and low implementation cost; it does not introduce a large number of complex floating-point operations, reducing the difficulty and cost of hardware implementation.
Owner:ASR MICROELECTRONICS CO LTD

Blind detection result acquisition method and apparatus, storage medium, and electronic device

Embodiments of the present application provide a blind detection result acquisition method and apparatus, a storage medium, and an electronic device. The method comprises: determining estimation value groups respectively corresponding to a plurality of synchronization signal blocks to be subjected to blind detection, wherein each estimation value group comprises a plurality of channel estimation values corresponding to one synchronization signal block; for each synchronization signal block among the plurality of synchronization signal blocks, on the basis of noise power and reference signal received power corresponding to the estimation value group of each synchronization signal block, determining the signal-to-noise ratio of each synchronization signal block so as to obtain the signal-to-noise ratios respectively corresponding to the plurality of synchronization signal blocks; and on the basis of the signal-to-noise ratios respectively corresponding to the plurality of synchronization signal blocks, determining the synchronization signal block having the maximum signal-to-noise ratio among the plurality of synchronization signal blocks, and determining the synchronization signal block having the maximum signal-to-noise ratio as a blind detection result. The problems in the related art of complex algorithms, high resource overhead during implementation by means of hardware, etc., in existing issb blind detection solutions are solved.
Owner:SHANGHAI CYGNUS SEMICON CO LTD

Matrix decomposition device and method based on memristor cross array

The invention relates to a matrix decomposition device and method based on a memristor cross array, which are suitable for efficient hardware implementation of singular value decomposition (SVD), the memristor cross array receives a conductance value mapped by a Grubrum matrix constructed by a to-be-decomposed matrix and a bit line input voltage converted by a bit line column vector, and outputs a corresponding current; converting the corresponding current into a voltage vector; performing normalization processing on the voltage vector to obtain a normalized vector; judging whether convergence occurs or not; performing normalization processing on the steady-state voltage vector to obtain a final output vector; and according to the final output vector, main characteristic values are extracted based on a first normalization circuit, and main singular values and corresponding singular matrixes are calculated. Compared with the prior art, the high parallel computing characteristic of the memristor cross array is utilized, efficient operation and hardware acceleration of matrix decomposition are achieved, the method has the advantages of being low in power consumption, high in speed and good in expansibility, and the method is suitable for application scenes such as artificial intelligence, signal processing and large-scale matrix operation.
Owner:SOUTHEAST UNIV

3D Gaussian splash rendering optimization method in VR based on Unity engine

The invention discloses a 3D Gaussian splash rendering optimization method based on a Unity engine in VR, and relates to the technical field of computer graphics. PLY point cloud files are converted into SOG high-compression formats containing multi-level LOD information, after low-transparency Gaussian points are removed, Morton coding sorting is carried out on the center coordinates of the Gaussian points, and then the 3D Gaussian splash rendering optimization method based on the Unity engine in VR is realized. The method comprises the following steps: constructing an octree spatial index containing a spatial bounding box and multiple LOD data, removing Gaussian points in a grading manner according to a camera distance, selecting an LOD hierarchy, carrying out deep barrel sorting on visible Gaussian points, executing drawing by taking a barrel as a unit, and combining near-to-far sorting, an improved mixed equation and screen space edge removal and amplification operation to obtain a target image. According to the method, through multi-dimensional optimization such as loading, space elimination and rendering, the loading efficiency of the VR equipment is improved, the rendering burden is reduced, Tile-Based hardware is adapted, and smooth operation of 3D Gaussian splashing on the VR all-in-one machine is realized.
Owner:YUANMENG SPACE DIGITAL TECHNOLOGY (CHENGDU) CO LTD

Hardware implementation method and device for communication between CPUs with low time delay

The invention discloses a low-delay hardware implementation method and device for communication between CPUs, and relates to the technical field of computers. The method comprises the following steps: receiving initialization configuration information from a sending CPU (Central Processing Unit), including an initial address and a space size of a memory of the sending CPU and an initial address and a space size of a memory of a receiving CPU; and monitoring a sending tail pointer updated by the sending CPU, and determining whether to read the to-be-sent data from the internal memory of the sending CPU to the internal cache of the hardware communication module based on the initialization configuration information, the sending tail pointer and a sending head pointer maintained by the hardware communication module. And monitoring a receiving head pointer updated by the receiving CPU, and determining whether to write the to-be-written data in the internal cache into the memory of the receiving CPU based on the initialization configuration information, the receiving head pointer and a receiving tail pointer maintained by the hardware communication module. And the hardware communication module transmits data between the sending CPU memory and the receiving CPU memory in a direct memory access mode.
Owner:PENG TI STORAGE TECH (NANJING) CO LTD

Software and hardware combined PCIE address translation and networking design device and method

The invention belongs to the field of switch chips, and relates to a hardware and software combined PCIE address translation and networking design device and method, the hardware and software combined PCIE address translation and networking design device comprises an address translation unit, a built-in CPU and firmware and a TLP transceiving unit, the address translation unit comprises a control register and an address or BDF translation module for implementing TLP address or BDF translation; the firmware is used for completing the simulation of the iEP, and the simulation method is that software is used for completing the response to a received configuration request TLP, so that a DSP and an EP are simulated; and the TLP receiving and transmitting unit connects the built-in CPU to a switching network, so that the built-in CPU receives and transmits a TLP data packet. According to the PCIe switching chip, a tree topology structure of a traditional PCIe switching chip can be broken through, and any needed network topology structure can be flexibly connected through software configuration; a plurality of hosts and devices can be connected at any port as required; the required iEP is realized by using internal firmware, the function of the iEP can be realized by more flexible programming, and the design cost is reduced compared with the method for realizing the iEP by using hardware.
Owner:SHANGHAI DUXIN INTEGRATED CIRCUIT DESIGN CO LTD

Accelerating artificial neural networks using hardware-implemented lookup tables

The invention is notably directed to a hardware system (1) designed to implement an artificial neural network (ANN). The hardware system basically includes a neural processing apparatus (15), e.g., involving as crossbar array structure, one or more lookup table circuits (17), and one or more processing units (18). The neural processing apparatus is configured to implement M artificial neurons, where M≥1. The lookup table circuits are configured to implement a lookup table (LUT). The system further includes M′ processing units, where M≥M′≥1. Each processing unit is connected by at least one neuron, in order to be able to access a first value outputted by each connected neuron. In addition, each processing unit is connected to a LUT circuit, in order to efficiently access parameter values of a set of parameters from the LUT. Finally, each processing unit is configured to output a second value, corresponding to a value of a mathematical function taking said first value as argument. The mathematical function is otherwise determined by the set of parameters, the parameter values of which are accessed by each processing unit from the LUT, in operation. I.e., the mathematical function is defined (and thus determined) by a set of parameters, the values of which are efficiently retrieved from the hardware-implemented LUT. This results in a substantial acceleration of the computations of the function outputs, beyond the acceleration that may already be achieved within the neural processing apparatus and the processing units themselves. As a result, the neuron outputs can be more efficiently processed, prior to being passed to a next neuron layer. The invention is further directed to a method of operating such a hardware system.
Owner:AXELERA AI BV

Coding mode determination method and device, electronic equipment, storage medium and program product

The invention relates to the field of video coding, and provides a coding mode determination method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: in response to acquiring coefficients of a group included in a transformation block in a video frame, starting to perform coding cost calculation on the group through all the coefficients when all the coefficients included in any group are acquired, and performing coding cost calculation of a plurality of groups in the transformation block in parallel; wherein the transform block comprises a plurality of groups, each group comprises a plurality of coefficients, and the coding cost of the group comprises the cost that the group is not coded and the cost that the group is coded in at least one coding mode; and according to the calculated coding cost of the plurality of groups, determining the groups for coding in the transform block and the coding mode of the groups for coding. According to the method disclosed by the embodiment of the invention, the parallel computing degree of hardware is improved, and the complexity and time delay of hardware implementation are reduced.
Owner:MOORE THREADS TECH CO LTD

Hardware implementation method and device of IPSec protocol processor based on ASCON lightweight encryption algorithm

The invention belongs to the technical field of network communication. The invention provides a hardware implementation method and device of an IPSec protocol processor based on an ASCON lightweight encryption algorithm. According to the embodiment of the invention, the ASCON lightweight encryption algorithm and the optimized security policy matching algorithm are introduced, the high throughput rate is ensured, the resource consumption is reduced, the delay is reduced, the hardware implementation of the ASCON algorithm is optimized for the application scene of IPSec protocol processing, and the whole flow of the ASCON algorithm is divided into two stages of pre-calculation and message encryption, so that the processing efficiency of the IPSec protocol is improved. By pre-calculating and storing the initialization of each security alliance and the encryption intermediate state of the associated data, repeated calculation is avoided, and the throughput is improved by executing multiple turns of transformation in parallel in a single cycle in the hardware implementation of the message plus module.
Owner:XIDIAN UNIV

Signal detection method, device, equipment, medium and product

The invention relates to a signal detection method, device and equipment, a medium and a product. The method comprises the following steps: acquiring a target receiving signal to be detected and a Monte Carlo tree corresponding to a transmitting signal; under the condition that the current node in the Monte Carlo tree is completely expanded, for each child node of the current node, shifting processing is carried out based on the number of access times of the child node, the exploration value of the child node is determined, and the optimal child node of the current node is determined based on the reward value and the exploration value of the child node. Adding the optimal child node into a current search path corresponding to the current iterative search process, taking the optimal child node as a new current node until the current search path is searched, and determining a reward value corresponding to each node in the current search path; and determining a target transmitting signal corresponding to the target receiving signal based on the reward value of each node in the Monte Carlo tree when a preset iteration end condition is satisfied. By adopting the method, the hardware implementation complexity can be reduced.
Owner:PURPLE MOUNTAIN LAB