Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41 results about "Processor design" patented technology

Processor design is the design engineering task of creating a processor, a key component of computer hardware. It is a subfield of computer engineering (design, development and implementation) and electronics engineering (fabrication). The design process involves choosing an instruction set and a certain execution paradigm (e.g. VLIW or RISC) and results in a microarchitecture, which might be described in e.g. VHDL or Verilog. For microprocessor design, this description is then manufactured employing some of the various semiconductor device fabrication processes, resulting in a die which is bonded onto a chip carrier. This chip carrier is then soldered onto, or inserted into a socket on, a printed circuit board (PCB).

Hybrid branch predictor capable of dynamically selecting different prediction methods and application

The invention discloses a hybrid branch predictor capable of dynamically selecting different prediction methods and application, and relates to the technical field of high-performance processor design, the hybrid branch predictor mainly comprises a branch direction predictor and a branch target predictor; the branch direction predictor is used for predicting whether a branch instruction jumps or not; and the branch target predictor is used for predicting a specific jump target address when the branch direction is predicted to be jump. By implementing the hybrid branch predictor capable of dynamically selecting different prediction methods and the application provided by the invention, the branch prediction accuracy can be improved, and the hardware resource overhead can be reduced.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719 +1

Tensor transpose processor

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. An example processor includes an input tensor shift buffer, a staging buffer, and an output tensor shift buffer. The input tensor shift buffer reads an input tensor from input memory and performs multiple cycles of input tensor shifting. The shifted tensor data is then written into the staging buffer. The output tensor shift buffer reads the shifted tensor data from the staging buffer and performs multiple cycles of output tensor shifting. Finally, the result is written to the output memory. This configuration facilitates efficient handling and transformation of tensor data, optimizing the computational processes required in machine learning tasks.
Owner:MOFFETT TECH CO LTD

Instruction processing method and device, terminal equipment and program product

The invention is suitable for the technical field of computer processor design, and provides an instruction processing method and device, terminal equipment and a program product, and the method comprises the steps: obtaining instruction information of a to-be-processed instruction set in a processor; generating an index vector corresponding to the to-be-processed instruction set based on the original effective vector; each index value in the index vector is subjected to one-hot code conversion, a one-hot code matrix corresponding to the instruction set to be processed is generated, and each row of one-hot code vector in the one-hot code matrix corresponds to each path of input instruction; and according to the original effective vector and the one-hot code matrix, sorting operation codes of each path of input instruction to obtain processed operation code information, so that a processor executes corresponding operation based on the processed operation code information. The method can meet the requirement that the high-performance processor completes instruction processing in a single cycle, the processing delay is low, and therefore the instruction execution efficiency of the processor is improved.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

Rag-based few-sample processor design ppa prediction method

The application discloses a kind of based on RAG's few-sample processor design PPA prediction method.The technical scheme is: the processor design paradigm, EDA tool module, RAG module, the input module of processor design to be measured and PPA prediction module are formed based on RAG's few-sample processor design PPA prediction system.RAG module carries out sampling to the design space of processor design paradigm, obtains processor design sample, the feature vector coding of sample, utilizes EDA tool module to obtain the PPA value of design sample, constructs the knowledge base required for PPA prediction;RAG module detects the knowledge matched with the processor design code of the PPA value to be predicted in knowledge base, generates prompt word based on the matched knowledge.PPA prediction module and prompt word interaction, obtain the PPA prediction value of the PPA value processor design to be predicted.The application can be based on a few samples to carry out accurate PPA prediction, shortens design cycle.
Owner:NAT UNIV OF DEFENSE TECH

Motor controller and engineering machine

This invention provides a motor controller and engineering machinery. The motor controller includes: a driver connected to a motor; and an MPU (Multi-Core Processor), the MPU including a multi-core processor. Each core of the multi-core processor communicates with each other via a communication link for data exchange. The multi-core processor is connected to the driver and is used to detect motor operating status information and control the motor based on the operating status information. This invention adopts a single-chip MPU multi-core processor design, which can realize the motor controller function with a single MPU, meeting the requirements of small size and compact design of the controller, reducing equipment cost, and conforming to the ISO 13849 functional safety PLD / CAT2 class architecture.
Owner:NANJING HENGLI INTELLIGENT TECH CO LTD

A reconfigurable baseband processing SoC architecture based on a RISC-V extended instruction set and an instruction control method thereof

The present application relates to a kind of reconfigurable baseband processing SoC architecture based on RISC-V extension instruction set and its instruction control method, belong to communication chip and processor design technical field.For the instruction set flexibility insufficient, data handling inefficiency and dynamic reconfiguration and high energy efficiency difficult to consider of existing 5G / 6G baseband processor, the present application proposes a kind of heterogeneous computing architecture, including the RISC-V processor core cluster of supporting RV64IMAFDCV instruction set, extension instruction execution unit, vector coprocessor and reconfigurable hardware acceleration module group.By adding special instructions such as complex matrix multiplication, optimizing data path design, using double buffering mechanism and hierarchical reconfigurable technology, efficient execution of baseband processing is realized.At the same time, dynamic voltage frequency regulation system can adjust power supply parameters in real time according to work load.The scheme significantly improves the performance and energy efficiency of baseband processing, and is suitable for physical layer signal processing of 5G / 6G communication system.
Owner:GUANGDONG COMM & NETWORKS INST +1

Cooperative verification debugging method and device for processor core, and storage medium

PendingCN122064549AFunctional testingTrace fileResource consumption
The invention provides a processor core co-verification debugging method, which comprises the following steps of: in a co-simulation running process of a processor core to be tested, capturing an output verification event, performing standardized storage, and generating a trace file and an interface specification file; automatically generating an independent verification framework compatible with the to-be-tested processor core interface according to the interface specification file; and loading the trace file to drive the verification framework to run, and separating from the processor core to be tested to carry out independent debugging and iteration. According to the invention, the efficiency and reliability of large-scale processor design verification are improved, and the resource overhead is reduced.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

DATA PROCESSING DEVICE

According to various embodiments, a data processing device is described, comprising: a processor designed to execute a computer program, and persistent memory, wherein the computer program includes instructions that effect an update of a signature of the program execution at predetermined signature update phases of the program execution, wherein at least one of the instructions effects an update of the signature depending on a configuration value stored in the persistent memory of the data processing device, wherein the configuration value specifies one or more operating settings of the data processing device, and wherein the processor is designed, for each of one or more predetermined integrity check phases of the program execution, to perform further processing after the predetermined integrity check phase of the program execution depending on thatwhether the signature matches one or more predetermined reference signatures.
Owner:INFINEON TECHNOLOGIES AG

RADIO FREQUENCY INTERFERENCE REDUCTION

A device may include: a memory and a processor designed to: determine a utilization state of a wireless communication link designed to communicate within a radio frequency channel; determine an operating frequency of a radio frequency interference generating device; and instruct the radio frequency interference (RFI) generating device to adjust the operating frequency based on the utilization state of the wireless communication link.
Owner:INTEL CORP

Instruction processing method and system, electronic equipment and storage medium

The embodiment of the invention discloses an instruction processing method and system, electronic equipment and a storage medium, relates to the technical field of computer processor design, and can save write port resources of a register file. The method comprises the following steps of: analyzing a read-write dependency relationship of each instruction in a current instruction group on a flag bit register so as to establish a flag bit data stream dependency chain; according to the read-write dependency relationship and a preset elimination rule, judging whether each instruction indicating to be written into the flag bit register in the current instruction group needs to be written into the flag bit register or not; and if the judgment result shows that the instructions indicating to be written into the flag bit register in the current instruction group do not need to be written into the flag bit register, canceling the operation of writing the instructions indicating not to be written into the flag bit register back to the flag bit register. The method and the device are suitable for eliminating invalid flag bit registers.
Owner:HYGON INFORMATION TECH CO LTD

DATA PROCESSING DEVICE AND METHOD FOR PROTECTING A BOOLEAN MASKED ADDITION FROM ERRORS

According to various embodiments, a data processing device is described which includes a Boolean masked addition processor designed to perform Boolean masked addition and a fault protection circuit designed to generate a checksum for a first part, generate a second checksum part, compare the first checksum part with the second checksum part, and trigger a safety action in response to a mismatch between the first checksum part and the second checksum part.
Owner:INFINEON TECHNOLOGIES AG

Synchronous verification method for asynchronous interrupt event

The invention provides a synchronous verification method for an asynchronous interrupt event, which comprises the following steps of: establishing a system-level verification platform designed by an instruction-level synchronous general processor based on UVM (Universal Verification Methodology), and loading an excitation binary file to enable the excitation binary file to work normally so as to realize instruction-level synchronous verification; creating an asynchronous interrupt event excitation, wherein the asynchronous interrupt event excitation comprises asynchronous interrupt event triggering time; in the normal working process of the system-level verification platform, an asynchronous interrupt event is injected into a to-be-tested product through the UVM environment framework, and the asynchronous interrupt event is injected into the reference model when the asynchronous interrupt event is detected to occur, so that the reference model synchronously triggers the same asynchronous interrupt event, and the system-level verification platform is verified. And instruction-level synchronous verification when an asynchronous interrupt event occurs is realized. According to the synchronous verification method for the asynchronous interrupt event, the synchronization relation of instruction submission is established between the to-be-tested product and the reference model, and synchronous verification can be carried out for the asynchronous interrupt event.
Owner:SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI

Processor, method, and system for accelerating tensor transpose for machine learning

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. It features an instruction decoder that decodes tensor transpose instructions for an input tensor, distinguishing between transposing and stationary axes. An inner transpose engine, equipped with multiple tensor buffer units of varying sizes, transposes the tensor by reading data by rows and writing by columns, efficiently managing memory when the buffer is full. Additionally, an address scheduler determines the target memory addresses in the output tensor memory for the tensor data, improving the handling and transformation of tensors in machine learning environments. This processor significantly enhances the efficiency and speed of tensor operations critical to advanced machine learning applications.
Owner:MOFFETT TECH CO LTD

VEHICLE CONTROL DEVICE, VEHICLE CONTROL SYSTEM, VEHICLE LEARNING DEVICE AND VEHICLE LEARNING PROCEDURE

Vehicle control device, which includes: a processor; and a memory (46) wherein the memory (46) stores relationship-defining data to define a relationship between a state of a vehicle and an action variable, which is a variable relating to operations of a transmission installed in the vehicle, the processor is designed to execute: an investigative processing method to determine the condition of the vehicle based on a sensor reading, an operational processing to control the transmission based on a value of the action variable determined by the state of the vehicle as determined in the investigation processing and the relationship-defining data, a reward calculation process to give a larger reward when the vehicle's characteristics meet a reference, instead of failing to meet the reference, based on the vehicle's state determined in the discovery processing, an update processing to update the relationship-defining data with the state of the vehicle determined in the discovery processing, the value of the action variables used in the operation of the transmission, and the reward corresponding to the operation, as inputs into a pre-set update map, a counting process for counting an update counter by the update processor, and a bounding process to limit downwards a range used by the operation processing, in which a value other than a value that maximizes an expected utility with respect to the reward, of values ​​of the action variable that specify the relationship-defining data, when the update counter is relatively large, and The processor is designed to output updated relationship-defining data, thus increasing the expected benefit when the gearbox is operated according to the relationship-defining data, based on an update map.
Owner:TOYOTA JIDOSHA KK

Multi-thread processor pipeline architecture system, scheduling method and equipment

The invention relates to the technical field of processor design, discloses a multi-thread processor pipeline architecture system, a scheduling method and equipment, and designs an instruction prefetching scheduling algorithm based on mixed priorities, and the algorithm realizes load balancing of thread instruction queues through a static level decision mechanism and a dynamic level decision mechanism. The static level preferentially processes an empty queue thread to avoid starvation, and the dynamic level dynamically adjusts instruction fetching priority according to each queue depth, thread validity and a branch risk state to ensure that an instruction stream is continuously and stably supplied to a subsequent stage; a coupled polling emission scheduling algorithm is provided, so that high complexity of full-permutation search is avoided, and structural conflicts and data conflicts are effectively reduced; a special hardware architecture is designed around the algorithm, a multi-program counter and an instruction queue are integrated at the front end, a double-transmitting channel, a multifunctional unit and a thread private register file are configured at the rear end, and the performance of the single-core CPU is improved by optimizing the hardware architecture.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Fuzzy test and model detection combined hardware verification method and system

ActiveCN121301101ADetecting faulty computer hardwareComputer hardwareProcessor design
The invention discloses a fuzzy testing and model detection combined hardware verification method and system, and belongs to the technical field of hardware testing. The method comprises the steps that multiple rounds of fuzzy testing guided by the coverage rate are executed on a processor based on an initial seed set, snapshots generated in the fuzzy testing process are recorded, value scoring is conducted on the snapshots to construct a snapshot pool, and the processor is designed and generated based on a processor to be verified; selecting an optimal snapshot from the snapshot pool, and initializing a processor design based on the optimal snapshot; under the condition that points which are not covered exist in the processor design, a new test seed is generated through bounding model detection; performing fuzzy testing on the processor design based on the new test seed to obtain a verification result of the optimal snapshot; and synthesizing verification results of a plurality of optimal snapshots in the optimal snapshot set to obtain a hardware verification result of the to-be-verified processor design. According to the method, the efficiency and the applicability of the BMC can be remarkably improved on the premise of not sacrificing the verification precision.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Control device for semiconductor switches of an inverter output stage

The present invention relates to a control device (1) for semiconductor switches (4a, 5a, 4b, 5b, 4c, 5c) of an inverter output stage (2), the control device having a processor (12) and having drive circuits (6a, 7a, 6b, 7b, 6c, 7c) respectively provided for each semiconductor switch (4a, 5a, 4b, 5b, 4c, 5c) of the inverter output stage (2), wherein the processor (12) is designed to communicate via a communication interface (1 4.15) The phase current parameters during the switching process of one of the semiconductor switches (4a, 5a, 4b, 5b, 4c, 5c) in the inverter output stage (2) are transmitted to the corresponding drive circuits (6a, 7a, 6b, 7b, 6c, 7c) of the semiconductor switch (4a, 5a, 4b, 5b, 4c, 5c). c) has a computing unit (23) designed to calculate, based on parameters transmitted by the processor (12), the current sinusoidal phase current curves of the phase currents for the switching process of the semiconductor switches (4a, 5a, 4b, 5b, 4c, 5c), and wherein the driving circuits (6a, 7a, 6b, 7b, 6c, 7c) of the semiconductor switches (4a, 5a, 4b, 5b, 4c, 5c) have current curve storage devices (25) designed to calculate, based on parameters transmitted by the processor (12), the current curves of the phase currents for the switching process of the semiconductor switches (4a, 5a, 4b, 5b, 4c, 5c ... The calculation unit (23) of the driving circuit (6a, 7a, 6b, 7b, 6c, 7c) calculates the sinusoidal phase current curve and / or selects the current curve stored therein according to the voltage and / or according to the temperature, and provides it to the driving circuit (6a, 7a, 6b, 7b, 6c, 7c) of the semiconductor switch (4a, 5a, 4b, 5b, 4c, 5c) for performing the switching process in the semiconductor switch (4a, 5a, 4b, 5b, 4c, 5c).
Owner:ROBERT BOSCH GMBH

Apparatus and method for phase aware reconfigurable processor design for optimized prefill and decode cluster performance

Apparatus and method for a reconfigurable processor for prefill and decode operations. An example processor comprises: compute circuitry to perform compute operations associated with prefill phase of an LLM workload in which first tokens of an input prompt are processed in parallel and a decode phase in which response tokens are generated sequentially; a memory controller; an input / output (I / O) controller; an interconnect fabric; and a management controller to select between a first plurality of operational modes responsive to detecting the prefill phase and a second plurality of operational modes responsive to detecting the decode phase, wherein the first plurality of operational modes are selected to enhance performance of the compute circuitry and the second plurality of operational modes are selected to enhance performance of the memory controller and the interconnect fabric.
Owner:INTEL CORP

Processor, method, and system for accelerating tensor transpose for machine learning

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. It features an instruction decoder that decodes tensor transpose instructions for an input tensor, distinguishing between transposing and stationary axes. An inner transpose engine, equipped with multiple tensor buffer units of varying sizes, transposes the tensor by reading data by rows and writing by columns, efficiently managing memory when the buffer is full. Additionally, an address scheduler determines the target memory addresses in the output tensor memory for the tensor data, improving the handling and transformation of tensors in machine learning environments. This processor significantly enhances the efficiency and speed of tensor operations critical to advanced machine learning applications.
Owner:MOFFETT TECH CO LTD

DISTANCE MEASURING DEVICE

UndeterminedDE102025106927A1Pulse sequenceProcessor design
According to various embodiments, a distance measuring device is described, comprising: a receiver designed to receive first pulses of a first pulse train sent by a transmitter with a first interval between the pulses, and to receive second pulses of a second pulse train sent by the transmitter with a second interval between the pulses, wherein the first interval is different from the second interval;and a processor designed to produce a first processing result by analyzing the timing of one or more peaks in a first received signal received by the receiver upon receiving the first pulses, to produce a second processing result by analyzing the timing of one or more peaks in a second received signal received by the receiver upon receiving the second pulses, and to determine a distance between the transmitter and the distance measuring device as a function of a comparison of the first processing result and the second processing result.
Owner:INFINEON TECHNOLOGIES AG

A light and small inter-satellite laser communication terminal

The utility model provides a kind of light inter-orbital laser communication terminal of small and light, including optical head and processor module of light connection;Optical head includes first plane mirror, main telescope, second plane mirror, fast mirror, dichroic mirror, first narrowband filter, beam splitter, beam receiving collimator, second narrowband filter, tracking collimating camera, signal light emission collimator, pre collimating mirror, shutter and corner reflector;Processor module includes power module, integrated communication module and optical amplifier module in turn stacked in sequence.The utility model reduces the quality of whole machine by reasonable layout;The key components of optical machine main body are all used composite material with high specific stiffness, while reducing quality, ensure force thermal stability;Integrated processor design is carried out, integrated communication module, optical amplifier and secondary power supply are designed as a whole, reduce metal shielding material, reduce processor weight on the basis of ensuring radiation resistance performance.
Owner:BEIJING RES INST OF TELEMETRY

Live FX

1. The name of the design product: Live effecter (MIDIFOX Live FX). 2. The use of the design product: The design product is a portable audio processor designed for mobile live, KTV, podcast and electric pipe instrument performance. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view 1.
Owner:深圳市宝安区小狐狸数码电子商行(个体工商户)

Graphics processor instruction processing method and device based on risc-v instruction set

ActiveCN116342368BGraphicsInstruction set design
The application discloses a kind of based on RISC-V instruction set's graphic processor instruction processing method and device, comprising: the current instruction is decoded, and initial decoding result is obtained;According to the initial decoding result, it is judged whether the current instruction is register expansion instruction;If the current instruction is not the register expansion instruction, then according to the effective bit information of target temporary register, the instruction processing category is determined;Wherein, the instruction processing category includes RISC-V standard decoding category or splicing decoding category;For the splicing decoding category, target expansion data is read from the target temporary register;Wherein, the target expansion data is the expansion data corresponding to last register expansion instruction;The target expansion data and the initial decoding result are spliced and handled, and target decoding result is generated.The application can improve the instruction processing performance of graphic processor designed by using RISC-V instruction set.
Owner:INTERNATIONAL INNOVATION CENTER OF TSINGHUA UNIVERSITY SHANGHAI +1

Processor verification method and system, electronic equipment and readable storage medium

The embodiment of the invention provides a processor verification method and system, electronic equipment and a readable storage medium. The method comprises the following steps: in response to any verification event triggered by a to-be-tested design of a processor, determining whether the verification event belongs to a target verification event; the target verification event is a non-deterministic event of a to-be-tested design of the processor; under the condition of determining that the verification event belongs to the target verification event, taking the verification event as a target event, and sending the target event and a verification label corresponding to the target event to the receiving end; the verification label is used for representing a verification sequence of the target event; and the receiving end is used for receiving each verification event sent by the sending end, and executing each target event based on the verification sequence represented by the verification label of each target event in the process of executing each verification event. And the verification efficiency of the processor is improved.
Owner:BEIJING INSTITUTE OF OPEN SOURCE CHIP

Vehicle-mounted full-localization heterogeneous acceleration device

PendingCN121957793AGeneral Computing Performance ImprovementImprove general computing capabilitiesProgram initiation/switchingResource allocationOperational systemIn vehicle
The invention relates to a vehicle-mounted full-localization heterogeneous acceleration device, and belongs to the field of computer hardware technology and computer application. The device comprises a case, an ATCA (advanced telecom computing architecture) blade (4), a GPU (graphics processing unit) acceleration blade (5), an intelligent acceleration blade (6) and three power panels (1, 2 and 3). Compared with a traditional VPX isomorphic computing service platform based on a domestic CPU and a double-path S5000C / 32 processor design adopted under an ATCA architecture, the universal computing power is doubled, in quality computing application of a multi-target (the number of targets is larger than 4000) track, the designed heterogeneous acceleration device has the advantages that the computing time is greatly shortened compared with CPU multi-thread computing, and the speed-up ratio exceeds 100; nationwide software and hardware design, a domestic chip, firmware and an operating system are adopted, and the risk that foreign technologies depend on is avoided.
Owner:BEIJING INST OF COMP TECH & APPL

On-site backtracking method and device for instruction execution error of processor and storage medium

PendingCN121681248ADetecting faulty computer hardwareData streamProcessor design
The invention provides a field backtracking method and device for instruction execution errors of a processor and a storage medium, and belongs to the technical field of computers. The method comprises the following steps of: performing data analysis and instruction backtracking on a captured data stream on a bus through a preset check point; and comparing the execution results of the preset simulator and the preset simulator, and backtracking the field with the instruction execution error of the processor. According to the method, the site of instruction execution errors can be accurately traced back, and the root of the problem can be quickly positioned; according to the method, the data volume needing to be recorded is reduced, the backtracking efficiency is improved, the accuracy of fault positioning is ensured, and powerful support is provided for processor design and verification; the difficulty and cost in the processor verification process are reduced, and continuous progress and development of the processor technology are promoted.
Owner:HAIGUANG INTEGRATED CIRCUIT DESIGN (BEIJING) CO LTD

Intra-core bus parameterization design method and device, electronic equipment and storage medium

The invention provides an intra-core bus parameterization design method and device, electronic equipment and a storage medium, and relates to the technical field of digital processor design. The method comprises the following steps: reading a data file containing interface description and function channel configuration, analyzing to obtain a corresponding relationship between a client interface and an internal uniform interface, and automatically generating a corresponding logic structure according to an enabling state and an attribute parameter of a function sub-module, so that adaptive logic can be generated based on interface mapping; and a bus circuit description file capable of being synthesized is generated in combination with configuration parameters, so that integrated generation of interface adaptation, function cutting and structure construction is realized. According to the technical scheme, the flexible configurability of the bus structure can be kept when different project requirements are met, repeated labor caused by manual coding can be reduced, the design risk caused by inconsistent modification can be reduced, and the development efficiency and maintainability of the bus can be improved.
Owner:MOORE THREADS TECH CO LTD

Automatic superscalar processor design method based on learning data dependency relationship

According to the automatic superscale processor design method based on the learning data dependency relationship, the superscale processor supporting instruction-level parallelism is automatically designed by learning data dependency between instructions, and the defect that dynamic dependency cannot be processed in the prior art is overcome. According to the technical scheme, dependency prediction is achieved based on a hardware-friendly machine learning model, and the most reusable state is selected from a high-dimensional processor state space through a state selector and stored in a small buffer area; and the state speculator is used for generating a hardware predictor by utilizing the selected state high-precision prediction dependence data and integrating the hardware predictor into the superscalar processor. The predictor is obtained by training a machine learning model S-BSD and comprises a state selector and a state speculator. According to the scheme, low-delay and high-precision prediction is achieved under the condition that hardware resources are limited, and parallel execution of multiple instructions is supported.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A configurable hardware architecture and bit manipulation method based on Banyan networks

This invention provides a configurable hardware architecture and bit manipulation method based on Banyan networks, belonging to the fields of digital logic circuits and high-performance computing technology. The invention aims to solve the problems of high hardware overhead, high power consumption, and long latency caused by the need for separate hardware modules—a Banyan network and a barrel shifter—to achieve combined bit-gathering and cyclic shifting operations in existing technologies. The architecture of this invention includes a Banyan network switching module and a control signal generation module. The control signal generation module generates an offset prefix count by superimposing the cyclic shift amount *r* with the original prefix count based on a bit mask, and uses this count to generate control signals for each stage of the Banyan network via an LROTC circuit. Through innovative modifications to the Banyan network control logic, this invention enables a single network to simultaneously perform both advanced bit operations, bit-gathering and cyclic shifting. This invention offers significant advantages such as hardware resource savings, reduced power consumption, improved performance, flexible design, and easy expansion, and can be widely applied in processor design, cryptographic acceleration, digital signal processing, and other fields.
Owner:FALCON TECHNOLOGY (GUANGZHOU) CO LTD

A multi-precision / hybrid-precision fused multiply-add computing unit circuit

The present application relates to the technical field of digital integrated circuit and processor design, in particular to a multi-precision / mixed-precision fusion multiply-add calculation unit circuit.The present application starts from four aspects: a Booth multiplier array with bit width adaptive mapping, a reconfigurable compressor array, a configurable barrel shifter with high hardware multiplexing, and an asymmetric reconfigurable binary tree composed of a leading zero detector;and the four modules are used to realize multi-precision data type support from low precision (INT4, INT8, FP8, FP16, BF16) to high precision (FP32), and the calculation precision is maximally preserved during mixed-precision calculation;compared with the traditional FMA unit designed separately for different precisions, the present application greatly optimizes the hardware multiplexing rate and energy efficiency ratio on the basis of supporting multi-precision / mixed-precision calculation mode, improves the energy efficiency ratio and calculation throughput, and is suitable for artificial intelligence accelerators, general-purpose processors and high-performance computing chips.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA