Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

202 results about "Register allocation" patented technology

In compiler optimization, register allocation is the process of assigning a large number of target program variables onto a small number of CPU registers. Register allocation can happen over a basic block (local register allocation), over a whole function/procedure (global register allocation), or across function boundaries traversed via call-graph (interprocedural register allocation). When done per function/procedure the calling convention may require insertion of save/restore around each call-site.

Data transmission method and system, chip, electronic equipment and storage medium

The invention provides a data transmission method and system, a chip, electronic equipment and a storage medium, the method is applied to a processing module of the data transmission system, and the method comprises the following steps: under the condition that an interrupt request is received, obtaining a data transmission instruction, and performing data analysis on the data transmission instruction to obtain an interrupt request; obtaining a data operation type, an interface type and a data transmission parameter; acquiring register configuration data corresponding to the interface type, and performing configuration corresponding to the register configuration data and the data transmission parameters on the configuration register through the internal bus; under the condition that configuration of the configuration register is completed, data to be transmitted are obtained through a transmission interface corresponding to the register configuration data, and data transmission corresponding to the data operation type and the data transmission parameter is conducted on the data to be transmitted based on the control register and the encryption and decryption unit. The transmission interface is one of a bus master interface, a bus slave interface and a private interface.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

FPGA (Field Programmable Gate Array) implementation method for 16-path CAN (Controller Area Network) and 16-path serial port data transceiving

The invention provides an FPGA (Field Programmable Gate Array) implementation method for receiving and transmitting 16-path CAN (Controller Area Network) and 16-path serial port data. The FPGA implementation method comprises the following steps of: constructing a hardware architecture suitable for efficiently receiving and transmitting the 16-path CAN and 16-path serial port data by utilizing the characteristic of cooperative work of a processing system PS of XC7Z045FFG900-2 and a programmable logic PL; a CAN bus controller on the PL side is configured; configuring a serial port transceiver chip at a PL side; a ping-pong buffering function on a PL side is realized; establishing register mapping; and a double-AXI interface collaborative interaction mechanism of the PL side and the PS side is realized. The method has the beneficial effects that the ping-pong buffer technology and the parallel processing capability of the FPGA are adopted, so that efficient concurrent processing of 16-path CAN and 16-path serial port data is realized; through reasonable hardware design and register configuration, the stability of communication between the CAN bus and the serial port is ensured; based on the programmable characteristic of the FPGA, the system is flexibly configured and adjusted according to different application requirements; the advantages of an XC7Z045FFG900-2 hardware platform are fully exerted, and the resources of the PL side and the PS side are reasonably distributed and utilized.
Owner:TIANJIN XUNLIAN TECH CO LTD

Compiler optimization method and device and storage medium

The invention discloses a compiler optimization method and device and a storage medium, and belongs to the technical field of computers. The method comprises the steps of obtaining a vector length corresponding to an access operation in a compiling language; the vector length is related to the number of elements included in the memory address to be accessed; under the condition that the vector length is not the power of the preset numerical value and the vector length is greater than a preset threshold value, splitting the access operation according to the vector length to obtain a plurality of sub-access operations; and according to the plurality of sub-access operations, optimizing repeated operations in the compilation language to obtain a compilation optimization result. By standardizing the access operation in the intermediate language of the compiler and splitting the access operation into a plurality of sub-access operations, the repeated sub-access operations in the intermediate language can be eliminated; moreover, in the instruction scheduling process after register allocation, based on the instructions corresponding to the multiple sub-access operations, access conflicts specific to hardware can be eliminated, and the accuracy of compiler optimization results is improved.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Object routing node, routing method and network-on-chip

The invention discloses an object routing node, a routing method and a network-on-chip. The object routing node comprises a first register and a second register, and the first register is configured to respond to the condition that a first message of a first type and a second message of a second type are received at the same time, select to store the first message and send the second message to the object routing node under the condition that a next routing node of the object routing node meets a scheduling condition. Providing the first message for a next routing node; the second register is configured to acquire the second message in response to the fact that the second message is not received by the first register and the anti-starvation mechanism is triggered, and the anti-starvation mechanism is configured to indicate that after the first register is vacant, the first register does not receive the first type of message until the second message is scheduled to the next routing node through the first register. According to the object routing node, the arbitration logic is reduced, and the resource overhead caused by the arbitration logic is greatly reduced.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Register overflow optimization method and device and storage medium

The embodiment of the invention provides a register overflow optimization method and device and a storage medium, and is applied to the technical field of chips. In the method, for each virtual register in a target program, based on a physical register type supported by an instruction operand where the virtual register is located, the virtual register is optimized; selecting a corresponding target register class from N candidate register classes, wherein N is greater than 1; allocating a first physical register in the target register class to the virtual register; when register overflow occurs, a target register is selected from the allocated first physical registers, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes, and compared with the mode that the instruction operand overflows to the memory to generate read-write operation on the memory, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes; according to the method, different types of physical registers are overflowed, read-write operation aiming at the physical registers is generated, the pressure of the registers is relieved, the performance overhead of the overflowed memories is reduced, and the register distribution efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Simulation software-based instruction set conversion method for GPU heterogeneous environment

The invention relates to the technical field of instruction set conversion, and discloses an instruction set conversion method of a GPU heterogeneous environment based on simulation software. According to the method, the key instruction stream in the simulation task is accurately captured through the dynamic instrumentation method, and a solid foundation is provided for subsequent processing; the method comprises the following steps: dividing an original instruction stream into a plurality of instruction blocks to be converted through a classification mechanism according to instruction semantic features and hardware suitability, and creating conditions for parallel conversion; a lightweight front-end translator is adopted to efficiently convert classified instruction blocks into architecture-independent intermediate representations, the degree of parallelism and the data dependency relationship between instructions are reserved, the intermediate representation instructions are deeply optimized, and the method comprises the key steps of instruction selection, register allocation, instruction recombination, SIMT mapping, instruction coding and the like. And generating a native instruction of the target GPU architecture through accelerated translation of the target GPU thread block. According to the method, the instruction conversion period is shortened, and the correctness and the execution efficiency of the conversion result are improved.
Owner:TAIHANG NATIONAL LABORATORY

Register allocation for multi-phase task

A method of rendering in a graphics processing system comprising compiling a program for a dual phase fragment task by a compiler, the first phase of the program being executed at a fragment rate and the second phase or the program being executed at a sample rate with the compiler being configured to determine data comprising the number of registers required per fragment in the first phase, the number of registers common between the first and second phase per fragment and the number of registers required per sample for the second phase to a processor. The compiled program and data are provided to a processor comprising the number of registers required per fragment in the first phase, the number of registers common between the first and second phase per fragment and the number of registers required per sample for the second phase. The processor obtains a fragment shading rate value and uses this together with least the number of registers required in the compiled program to compute the number of registers needed per fragment.
Owner:IMAGINATION TECH LTD

Register allocation method and electronic device

The present application discloses a register allocation method and an electronic device. The method comprises: an electronic device may acquire a plurality of instructions, the plurality of instructions comprising at least one first instruction and at least one second instruction, the first instruction operating consecutive registers, and the at least one second instruction being an instruction other than the at least one first instruction among the plurality of instructions; allocating, to an i-th parameter in the first instruction, a target register corresponding to the i-th parameter, the first instruction comprising n parameters, n being a positive integer greater than 1, i being a positive integer not greater than n, and each parameter among the n parameters uniquely corresponding to a target register; and, after allocating the register to the parameter in the at least one first instruction, allocating a register to a parameter in at least one second instruction. According to the method, the number of move instructions generated in the process of allocating the register to the parameter in the first instruction (i.e., a range instruction) is reduced, so that the volume of an application installation package generated on the basis of the first instruction is reduced.
Owner:HUAWEI TECH CO LTD

Multistage fault-tolerant starting method for spaceborne computer

The invention relates to the technical field of spacecraft electronic systems, and discloses a satellite-borne computer multi-stage fault-tolerant starting method which comprises the following steps: step 1, a management FPGA (Field Programmable Gate Array) reads a starting source identification bit of a BIOSBOOT register, and selects a main BIOS (Basic Input / Output System) storage address or a standby BIOS storage address according to an identification bit value; 2, the management FPGA starts a BIOS level watchdog timer corresponding to the BIOS storage address, and the duration of the timer is configured by a BIOSWDT register; and 3, when the BIOS-level watchdog timer is overtime and does not receive the reset signal, the management FPGA modifies the starting source identification bit of the BIOSBOOT register, and a system reset signal is generated. According to the method, the management FPGA is taken as a core, and a double-BIOS redundant architecture and a multi-disk hot standby mechanism are combined, so that the problem of system startup single-point failure caused by BIOS firmware code bit flipping or startup disk physical damage is solved, autonomous fault-tolerant recovery of firmware and a storage medium is realized, and the defect that an existing single-BIOS single-storage-medium architecture cannot self-repair hardware failure is overcome.
Owner:BEIJING ZHONGKE TIANSUAN TECHNOLOGY CO LTD

Data processing method, multi-core heterogeneous system and computer readable storage medium

The invention discloses a data processing method, a multi-core heterogeneous system and a computer readable storage medium, and belongs to the technical field of data processing.The method comprises the steps that in the initialization stage of a hardware system, a physical storage space is allocated based on a predefined physical address boundary, and the physical storage space comprises a shared space; writing predefined configuration data corresponding to the shared space into a configuration register of the storage management module and latching the configuration register, wherein the configuration data comprises an access control strategy and a storage attribute; in the starting stage of the operating system, an address mapping relation is obtained, the address mapping relation comprises a target mapping relation, and under the condition that the operating system runs, a storage management module is triggered based on an instruction of a target operation to execute the target operation based on the address mapping relation, configuration data and a virtual address corresponding to the instruction, the target operation is read operation or write operation. The calculation amount of the processor is greatly reduced, and the time consumption for reading and writing data is reduced on the whole.
Owner:SHENZHEN PANGO MICROSYST CO LTD

A SATA data DMA transmission system and method based on FPGA

The present invention discloses an FPGA-based SATA data DMA transmission system and method. The system utilizes a register configuration module, a first FIFO storage module, a second FIFO storage module, a DMA state machine module, a SATA controller module, and an MB bus to AXI bus conversion module. The register configuration module is configured to send startup and read / write operation instructions to the DMA state machine module and receive a "busy" signal; send startup and read / write operation instructions and user data length to the SATA controller module and receive a "SATA_done" signal; the first FIFO storage module is configured to store discrete user memory space base addresses and corresponding space sizes; the DMA state machine module is configured to monitor the status of the DMA state machine; the second FIFO storage module is configured to store SATA read / write data; and the SATA controller module is configured to implement SATA interface functions. The present invention improves transmission rate and CPU performance.
Owner:ZETIAN ZHIHANG ELECTRONIC TECHNOLOGY (SICHUAN) CO LTD

Processors, components, devices, and methods for filtering processing of iir filters

ActiveCN115913176Bshort timefully accelerated effectComputer hardwareIir filtering
The application relates to a processor, component, device and method for filter processing of an IIR filter. The processor comprises a first configuration register, a second configuration register, a general register and a matrix multiplication accumulation unit. The general register sequentially reads and stores each row of coefficients of a coefficient matrix of each order of the IIR filter. The matrix multiplication accumulation unit obtains the same input element corresponding to the stored current row of coefficients when the first configuration register is configured in a broadcast mode, copies the stored current row of coefficients when the first configuration register is configured in a copy mode, multiplies the current row of coefficients with the corresponding input element in parallel to obtain corresponding products, and sequentially accumulates the product results of each row of coefficients to obtain a final output value, so that at least four output variables of sequential sampling moments are obtained each time. In this way, the time consumption of IIR filter operation can be significantly shortened, and sufficient acceleration effect is provided.
Owner:BESTECHNIC SHANGHAI CO LTD

LabVIEW test method for multi-channel programmable output clock buffer

The application belongs to the technical field of automatic test program, and particularly relates to a LabVIEW test method of a multi-channel programmable output clock buffer. The method comprises the following steps: initializing the clock buffer after power-on; resetting the FSWP spectrum analyzer and clearing the previous interface settings; switching the radio frequency switch to channel 1 and opening the channel 1 output; setting the signal source input frequency; setting the signal source input amplitude; starting the signal source input, and outputting the set frequency and amplitude to the input end of the clock buffer; starting the register configuration module, and transmitting the 174 24-bit register values to the SPI in sequence through the serial port, and then transmitting the SPI to the clock buffer; reading the data on the FSWP spectrum analyzer; and the method automatically measures the key indicators of the 14-channel clock buffer, saves the data and judges the advantages and disadvantages. The method reduces the burden of the test personnel, reduces the test time cost, and avoids causing data errors.
Owner:58TH RES INST OF CETC

Register allocation for multi-stage tasks

The invention discloses register allocation of multi-stage tasks. Within a graphics processing system, a plurality of different shading programs may be executed by a single processor on multiple threads. For each shading program, a plurality of registers are used to store data for the respective shading program. Accordingly, for a plurality of shading programs executing on a plurality of threads, a plurality of registers are allocated to each program or thread being executed. However, there is a limited number of registers available, so efficient allocation of the registers optimizes performance. Generally, an unnecessary number of registers is assigned to each shading program, but the present invention provides a method of assigning a correct number of registers based on the size of a segment being shaded.
Owner:IMAGINATION TECH LTD

Phase offset configuration method based on device unique identifier RCA

The invention discloses a phase offset configuration method based on a device unique identifier RCA, and belongs to the technical field of storage device management. Comprising the following steps: S1, a host activates all connected eMMC equipment through a standard eMMC initialization command; s2, the host allocates a unique relative device address RCAi for each eMMC device; s3, the host counts the total number n of the currently active devices; s4, calculating a phase deviation angle delta phi i for each eMMC device based on the relative device address RCAi and the total number n of the devices; s5, the host writes the phase deviation angle delta phi i into an extension register of each eMMC device through an extension register configuration command; and S6, the host activates the phase deviation of the target device, and data transmission is carried out in a specified phase window, so that each eMMC device only drives the data bus in the exclusive time slot corresponding to the phase deviation angle delta phi i of the eMMC device. Through innovative combination of protocol layer dynamic configuration and physical layer time isolation, the problems of signal conflict, delay and power consumption in a multi-eMMC equipment system are thoroughly solved.
Owner:JIANGSU XINSHENG INTELLIGENT TECH CO LTD

Compiler automatic debugging method and system for VLIW and SIMD architecture

The application discloses a kind of compiler automatic debugging method and system for VLIW and SIMD architecture, the method of the application includes the semantic correctness verification for the program to be checked to judge whether the program to be checked exists semantic error relative to source program, if semantic correctness verification finds that there is semantic error, then determine that debugging is not passed, otherwise the physical register verification for the program to be checked to judge whether the program to be checked exists physical register allocation error, if there is physical register allocation error, then determine that debugging is not passed, otherwise determine that debugging is passed;When determining that debugging is not passed, then generate feedback verification report.The application can automatically check the semantic correctness of program in the process of compiling, and provide accurate verification report (including error occurrence position and type etc.) for developer, improve the efficiency of compiler development, reduce the burden of developer.
Owner:NAT UNIV OF DEFENSE TECH

SoC chip-oriented efficient DDR4 memory debugging system and debugging method

The invention relates to the field of integrated circuits and the field of chip debugging, and provides an efficient DDR4 memory debugging system and method for an SoC chip, and the system comprises a parameter debugging module which is used for adjusting to-be-adjusted parameters in a register configuration file according to DIMM information, the register configuration file and working mode information, after the register parameter configuration is completed, a DDR4 IP is started to execute a training task according to the training parameter setting and the training control setting; the test module is used for checking whether the adjusted register parameters are correct or not and monitoring whether the training result of the DDR4 IP for executing the training task is correct or not, performing read-write test on the DRAM, and selecting to start the operating system when it is determined that checking and training are passed and the read-write test is passed. The debugging error rate can be effectively reduced, the debugging progress is accelerated, and the debugging efficiency is improved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

FPGA-based PCM collector dynamic framing IP core and method

The invention discloses a PCM (Pulse Code Modulation) collector dynamic framing IP core and method based on an FPGA (Field Programmable Gate Array). The IP core comprises a register access module, a DNA authorization module, an acquisition frame processing module, a framing control module and a PCM coding module. Comprising the following steps: performing parameter configuration on an IP core register; the DNA authorization module verifies the authorization key; the acquisition frame processing module receives and caches the acquisition frames in parallel and assembles PCM frame data; and performing PCM coding according to the code pattern configured by the register. The method has the beneficial effects that the frame format is high in universality, high in flexibility and high in deployment efficiency; predefinition and rapid dynamic switching of three frame formats are supported; parallel framing of 32 paths of acquisition frames is supported; common PCM code pattern output is supported; all functions can be accessed through a register, and extra logic is not needed; the register interface adopts a standard AXI-Lite interface; and a DNA authorization function is integrated.
Owner:TIANJIN XUNLIAN TECH CO LTD

Method for improving control precision of MCU (Microprogrammed Control Unit) on crystal oscillator

The invention discloses a method for improving the precision of an MCU for controlling a crystal oscillator, and relates to the technical field of microcontrollers. Setting a DAC output precision improvement multiple N, and calculating a register configuration range after improvement; calculating a theoretical register configuration value of a voltage needing to be output after the output precision of the DAC is improved; rounding off the calculated theoretical register configuration value to obtain an actual configuration value; creating an N-bit array for storing an actual DAC register value, dividing an actual configuration value by a lifting multiple N, rounding down, and assigning all elements in the array to obtain a rounding down result; calculating a result k obtained by taking a remainder from the lifting multiple N by an actual configuration value, adding 1 to the first k number in the array and then randomly disorganizing; creating timer interruption, and writing an array value into a DAC register every 1 / N seconds; and after all the data are written in, the DAC module outputs a voltage effective value with the control precision improved by N times within one second. According to the invention, high-precision voltage output of the voltage-controlled pin of the crystal oscillator is realized under the condition that hardware configuration is not changed.
Owner:CHENGDU JINNUOXIN HIGH-TECH CO LTD

Shared operation method, system and equipment of network acceleration unit and medium

The invention belongs to the technical field of hardware design, and provides a sharing operation method, system and equipment of a network acceleration unit and a medium, and the method comprises the step of configuring a shared general register, an exclusive general register and an exclusive special register group for multiple hosts in a single network acceleration unit. Each host accesses the register through the low-speed bus based on the unique identifier. And the acceleration unit performs data interaction with each host through a high-speed bus according to the configuration of the register. When the host operation is processed, the management message and the data transmission message are distinguished by judging whether the basic communication resource register group is accessed or not, and high-order modification and resource mapping are carried out on a load data address in the data transmission message. When an external data packet is received, a target host is determined through analysis and address matching, and resource identifier mapping and data access are carried out. According to the invention, each host is ensured to have independent communication resources and data channels, the complexity and the cost of hardware implementation are reduced, and the resource utilization rate and the system expansibility are improved.
Owner:XIAN AVIATION COMPUTING TECH RES INST OF AVIATION IND CORP OF CHINA

Register Allocation of Uniformized Multi-Core Programs

A computer-implemented method for allocating registers for multi-core programs. The method includes generating core-specific programs for execution units of cores, wherein each program contains core-specific configurations. The method includes analyzing the core-specific programs to identify core-specific differences. The method includes creating a uniformized program that consolidates semantics of all programs for the execution units while retaining operations to account for core-specific differences. The method includes performing live range analysis of the uniformized program to identify active intervals for variables and creating segmented live ranges to partition the intervals into global and core-specific segments. The method includes allocating registers to the execution units using the segmented live ranges and multi-casting the uniformized program to the execution units of the cores.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Register resource allocation method, computer equipment, readable storage medium and program product

The invention relates to a register resource allocation method, computer equipment, a readable storage medium and a program product. The method comprises the following steps: constructing a function group through a calling relationship between functions, wherein the function group comprises a coroutine kernel function and at least one coroutine subfunction; creating a resource record table corresponding to the function group, wherein the resource record table is used for recording register resources allocated to coroutine kernel functions and / or coroutine sub-functions in the function group; and carrying out register allocation on the coroutine kernel function and the coroutine sub-function in the function group based on the resource record table. By adopting the method, the requirements of local variable register resource isolation and shared variable register resource consistency in a collaborative mode can be met, so that the parallel computing efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Staged multi-policy instruction scheduling method and system for VLIW architecture

This invention discloses a staged multi-policy instruction scheduling method and system for VLIW architecture. The method includes: basic step S1. Receiving the symbolic assembly structure (SAS); S2. Configuring three types of scheduling vision interfaces and registering corresponding scheduling policies, including a global vision interface, a loop vision interface, and a linear vision interface; S3. Executing loop vision scheduling, traversing all loop blocks, and concurrently calling the loop vision interface to generate candidate scheduling schemes; S4. Executing linear vision scheduling, after all loop blocks have been scheduled, traversing the remaining unscheduled basic blocks, and concurrently calling the linear vision interface to generate candidate schemes; S5. Performing competitive selection and register allocation on the candidate scheduling schemes generated in each stage; S6. Outputting the optimized SAS. This invention can efficiently adapt to various VLIW processor architectures, improving instruction-level parallelism and code execution efficiency.
Owner:NAT UNIV OF DEFENSE TECH

Programmable logic device dynamic register configuration device and control method thereof

The invention provides a programmable logic device dynamic register configuration device and a control method thereof. The configuration device for the dynamic register of the programmable logic device comprises an external controller used for providing a configuration instruction and configuration data; the main controller is used for receiving and analyzing the configuration data, generating a gating signal and recombining a data packet; at least one multiplexing module, each multiplexing module being configured to include a register block; and one or more stages of sub-controllers are arranged between the main controller and the multiplexing module and are used for further analyzing the data packet and controlling the target register group to complete read-write operation. In addition, the invention further provides a control method corresponding to the configuration device. Some embodiments of the invention aim to solve the problems of high configuration switching time consumption and uncertain configuration switching time in the prior art on the premise of not increasing too much complexity and area overhead.
Owner:SHANGHAI XINLU TECH CO LTD

Interrupt controller isolation access control method and system based on virtual machine monitor

The invention relates to an interrupt controller isolation access control method based on a virtual machine monitor, and the method comprises the steps: analyzing an access instruction of a virtual machine to an interrupt controller, so as to obtain an operation attribute and a target address; querying a predefined mapping relationship based on the virtual machine identifier, the processor core information and the target address, and generating an interrupt nucleophilicity parameter; dynamically generating a register operation mask according to the interrupt nucleophilicity parameter and interrupt register configuration information corresponding to the target address; verifying the access authority of the virtual machine to the target address based on the register operation mask; and in response to the fact that the access permission passes the verification, executing the access of the access instruction to the interrupt controller in an atomic operation mode by applying the register operation mask.
Owner:SAIC GENERAL MOTORS +1

Processor, graphics card, computer device, register allocation method and apparatus

The application discloses a processor, a display card, a computer device, a register allocation method and device, and belongs to the technical field of register management. The processor comprises a plurality of registers, a register management unit and a task assembling unit. The register management unit is used for dividing the plurality of registers into at least one configuration block set, each configuration block set in the at least one configuration block set comprising at least one configuration block, and the at least one configuration block comprising at least two registers in the plurality of registers. The task assembling unit is used for allocating a first configuration block set to a first thread group, and the first configuration block set being one of the at least two configuration block sets. The processor allocates register resources in the form of configuration blocks, and can improve the allocation and release efficiency of register resources.
Owner:MOORE THREADS TECH CO LTD

Register configuration system, method, and electronic device

The present disclosure provides a register configuration system, method, and electronic device. The register configuration system includes a display processing unit, memory, and a hardware automatic control unit. The display processing unit includes multiple registers. The memory includes at least one pre-stored configuration instruction, wherein the configuration instruction is configured to indicate target registers that need to be configured in the display processing unit. The hardware automatic control unit is connected to a register configuration interface of the display processing unit. The hardware automatic control unit is configured to receive a startup command and, based on the startup command, retrieve and parse the configuration instructions from the memory. After parsing the configuration instructions, the hardware automatic control unit configures the target registers in the display processing unit. The startup command includes the storage address of the configuration instructions.
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD +1

Coalescing operand register file for graphical processing units

A system and method for register coalescing is described. The system comprises a CORF, a coalescing-aware register file design for GPUs that simultaneously reduces the leakage and dynamic access power, while improving the overall performance of the GPU. CORF achieves these properties by enabling the reads to multiple operands that are packed together to be coalesced, reducing the number of reads to the RF, and improving dynamic energy and performance. CORF combines compiler-assisted register allocation with a reorganized register file (CORF++) in order to maximize operand coalescing opportunities.
Owner:RGT UNIV OF CALIFORNIA

SNN learning accelerator for ECG monitoring

The invention discloses an SNN learning accelerator for ECG monitoring, which comprises a register configuration module, an asynchronous global control module, an asynchronous neural network control module, an asynchronous binary classification CNN module, an asynchronous quadruple classification SNN module, an asynchronous weight updating module and a memory management module, and is characterized in that firstly, a circuit corresponding to the accelerator constructs a double-layer dynamic network; layered dynamic power consumption management: selectively activating a high-precision secondary anomaly detection four-classification SNN network through a primary anomaly detection two-classification CNN network; secondly, the accelerator supports on-chip reasoning and efficient learning, and effectively eliminates ECG feature differences of different individuals on the premise of ensuring privacy security of user data; and finally, the circuit is controlled by using a pulse neural network and an asynchronous logic circuit, so that compared with a synchronous network, the power consumption is greatly reduced.
Owner:ZHEJIANG UNIV