Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Single cycle" patented technology

A single cycle processor is a processor that carries out one instruction in a single Clock cycle. MIPS architecture, MIPS-32 architecture. DLX, a very similar architecture designed by John L. Hennessy (creator of MIPS) for teaching purposes.

Instruction processing method and device, terminal equipment and program product

The invention is suitable for the technical field of computer processor design, and provides an instruction processing method and device, terminal equipment and a program product, and the method comprises the steps: obtaining instruction information of a to-be-processed instruction set in a processor; generating an index vector corresponding to the to-be-processed instruction set based on the original effective vector; each index value in the index vector is subjected to one-hot code conversion, a one-hot code matrix corresponding to the instruction set to be processed is generated, and each row of one-hot code vector in the one-hot code matrix corresponds to each path of input instruction; and according to the original effective vector and the one-hot code matrix, sorting operation codes of each path of input instruction to obtain processed operation code information, so that a processor executes corresponding operation based on the processed operation code information. The method can meet the requirement that the high-performance processor completes instruction processing in a single cycle, the processing delay is low, and therefore the instruction execution efficiency of the processor is improved.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

A storage coherent hub chip based on core particle integration, a storage coherent arbitration device and an adaptive control method

The present application relates to the technical field of multi-core heterogeneous computing, and discloses a storage coherent hub chip based on core integration, a storage coherent arbitration device and an adaptive control method, aiming to solve the technical problems of high cross-core memory access latency, low coupling degree of cache coherence maintenance and memory scheduling, and slow response of operation strategy adjustment under the existing multi-core architecture. The present application takes an independently packaged storage coherent core as a globally consistent unique maintenance node, integrates a single-cycle static addressing architecture, adopts a cache coherence state machine and a memory scheduling controller with deep fusion of logic layers, realizes automatic switching between robust mode and aggressive mode through a pure hardware MHM monitoring unit, and the atomic withdrawal process is transparent to the upper layer. The present application can reduce the cross-core memory access latency by more than 40%, improve the storage access throughput by 25%, while guaranteeing 99.999% operation reliability, adapting to various heterogeneous interconnection protocols, and being applicable to various application scenarios such as servers, high-frequency financial transactions, AR / VR wearable devices, edge computing, etc.
Owner:胡青

Processing circuit architecture supporting multi-calculation precision dynamic switching

The invention relates to the technical field of integrated circuit design, in particular to a processing circuit architecture supporting multi-calculation precision dynamic switching, which comprises a control port module, a partial product generation module, a symbol compression module, an addition compression tree module and a final adder module. And single-cycle dynamic switching of various precisions is realized. The partial product generation module adopts a Booth coding algorithm to split an operation vector, and cooperates with boundary symbol selection logic to solve the problem of symbol expansion and auxiliary bit overlapping; the symbol compression module compresses the redundant extension bits through a preset coding mode; the addition compression tree module is formed by cascading multiple stages of compressors and inserting carry blocking logic; the final adder module is composed of a plurality of carry lookahead adders, and the output bit width is dynamically controlled through blocking logic. The architecture optimizes the parallel operation efficiency and the resource utilization rate, adapts to the deep learning training and reasoning full scene, and has the advantages of real-time performance and low power consumption.
Owner:GUANGDONG INST OF INTELLIGENT SCI & TECH

Micro-controller chip containing multi-protocol communication interface peripheral and operation method thereof

Disclosed are a micro-controller chip containing a multi-protocol communication interface peripheral and an operation method thereof. The micro-controller chip comprises a multi-protocol communication interface peripheral, the multi-protocol communication interface peripheral is connected to a system bus, the multi-protocol communication interface peripheral is connected with an I / O port, the multi-protocol communication interface peripheral comprises an exclusively used RISC instruction set micro-kernel, a code memory and a code program stored on the code memory and executable by the RISC instruction set micro-kernel, the code program at least comprises two bit operation instructions of 1 setting and 0 clearing, the instructions are single-cycle instructions, and when the RISC instruction set micro-kernel executes the code program, the I / O port outputs 1 or 0.
Owner:NANJING QINHENG MICROELECTRONICS CO LTD

Scheduling control method, scheduling control device, electronic equipment and storage medium

The invention provides a scheduling control method, a scheduling control device, electronic equipment and a storage medium, and relates to the technical field of intelligent warehouse management, and the method comprises the steps: generating a plurality of groups of operation basic sequences in a full arrangement form based on Monte Carlo random sampling; aiming at each group of operation basic sequences, generating an initial operation sequence of each crane corresponding to the group of operation basic sequences, and replacing tail end tasks of each initial operation sequence; performing discrete event simulation on each operation sequence after replacement by taking a rigid demand period of a power generation side as a step length, and performing screening by taking single-period consumption meeting the power generation side in the whole process as a constraint condition to obtain a candidate operation sequence meeting the constraint condition; and selecting a sequence meeting a preset optimization index from the candidate operation sequences as a final scheduling sequence, and controlling each crane to operate based on the final scheduling sequence. According to the method and the device, the final scheduling sequence is generated to coordinate the operation of the multiple cranes, so that the congestion is avoided, and the working efficiency is improved.
Owner:BEIJING SHIDAI CHONGSHU TECHNOLOGY CO LTD

32-bit processor based on RISC-V instruction set architecture

The invention discloses a 32-bit processor based on an RISC-V instruction set architecture, and belongs to the field of computer system structures and microprocessor design. According to the technical scheme adopted by the invention, the 32-bit processor based on the RISC-V instruction set architecture comprises an instruction set extension module, two new instructions of'aggresgate 'and'disaggresgate' are introduced, an R-type coding format is adopted, and the 32-bit processor is used for realizing grouping and splitting operation of data bits; the data path unit comprises a general register group, an arithmetic logic unit ALU and a special adder and is used for executing various data operations and processing; the control unit is used for generating a control signal according to an instruction decoding result and controlling operation of the data path unit and the memory; according to the method, by adding the special instruction and optimizing the data flow control mechanism, the complex bit operation is completed in a single period, the cache pressure is reduced, and the method has the advantages of improving the bit operation efficiency, reducing the instruction cache pressure and reducing the dynamic power consumption.
Owner:济南晶谷研究院 +1

A multi-cycle modulation driving component, method, system, and product

ActiveCN120660069BControl engineeringSingle cycle
This invention proposes a multi-cycle modulation driving component, method, system, and product. The multi-cycle modulation driving component includes a state division component, a synchronous multi-cycle configuration component, and a single-cycle modulation component. The state division component performs functional decoding based on the device's functional mode and generates timing-matched trigger signals according to the specific functional mode. The synchronous multi-cycle configuration component generates various types of periodic / aperiodic enable signals at different clock frequencies based on configuration information, realizing multi-cycle timing drive. The single-cycle modulation component can serve as a general-purpose modulation unit; each unit can generate periodic composite on-chip timing drive signals based on configuration information and multi-cycle drive signals, thereby meeting the driving requirements of digital and analog readout units of different device chips. This invention can greatly simplify the overall timing structure of the driving chip.
Owner:NANJING VPS SEMICONDUCTOR TECHNOLOGY CO LTD

Debugging method and system for RISCV system memory

The application discloses a kind of RISCV system memory debugging method and system, its method includes: the interrupt request function of RISCV kernel is configured, enter debugging action and exit debugging action do not trigger the exception mechanism of RISCV processor;With host computer connection RISCV system, enter debugging state;The single-cycle valid signal triggered by host computer execution enter debugging action is converted into long-time valid signal;When long-time valid signal is valid, judge whether RISCV kernel has read-write, if yes, determine that the read-write signal of debugging module is invalid, set the state flag bit of interactive register, notify host computer system busy;If no, determine that the read-write signal of debugging module is valid, debugging module executes read-write to memory;Debugging module completes memory read-write and receives host computer exit debugging signal after exit debugging action, and make long-time valid signal invalid.The application utilizes the gap that RISCV kernel does not read nor write, realizes the read-write of debugging module to memory, realizes no interrupt memory debugging.
Owner:XIAMEN XINSIWANG INTEGRATED CIRCUIT TECH CO LTD

In-memory computing device, in-memory computing method, processing device, tile module and accelerator

The application discloses a memory-computing integrated device, a memory-computing method, a processing device, a tile module and an accelerator, and relates to the technical field of electronic circuits, and comprises: each memory-computing integrated array realizes parallel computation within the array, the memory-computing integrated array can simultaneously perform multiplication operation on single-bit input pulses at each moment in the current cycle buffer and corresponding weights, and synchronously generate membrane potential increment values at each moment; and the parallel accumulation of the increment values by a summation tree forms an efficient computing link of "parallel multiplication + parallel accumulation". Compared with the delay caused by the step-by-step waiting of the row-by-row serial processing of a computing task, the parallel architecture greatly shortens the processing time of the "multiplication-accumulation" whole process in a single cycle under the premise of ensuring high computing precision, further improves the computing efficiency under unit energy consumption due to the reduction of repeated data scheduling and state switching overhead in serial computation, and finally realizes the double optimization of computing delay reduction and energy efficiency ratio.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

System and method for single cycle dynamic hysteresis error control

A signal quantization module comprising: a combining module configured to combine an input signal and a feedback signal to generate a combined signal; an integrator module coupled to the combining module and configured to generate an integrated signal using the combined signal; and a quantizer configured to generate an output signal based at least in part on the integrated signal, wherein the output signal is fed back to provide the feedback signal to the combining module and the integrator module, and wherein the quantizer is further configured to, when the quantizer is triggered at a first point in time, compare an input to the quantizer to a threshold hysteresis, and in response to the input to the quantizer reaching the threshold hysteresis, adjust the threshold hysteresis by an amount that characterizes a difference between a target hysteresis and the input to the quantizer at the first point in time.
Owner:JAMES HAMMOND PTY LTD

Dynamic load balancing mapping method and system based on interrupt transaction combination

The invention discloses a dynamic load balancing mapping method and system based on interrupt transaction combination. The method comprises the following steps: an interrupt synchronization processing step; an interruption mark setting step; a load state acquisition step; a dynamic combination control step; a channel mapping step; a target source processing step; and a closed loop feedback step. The invention also comprises a system for implementing the method. According to the method, the transformation of an interrupt scheduling normal form is realized through'dynamic load perception + hardware closed-loop architecture + elastic scheduling strategy ', firstly, post statistics is replaced by real-time quantitative load perception, and the problem that a scheduling decision and a real-time demand are disjointed is solved; secondly, due to full-hardware implementation, software intervention overhead is eliminated, and scheduling delay is reduced to a single-cycle level; and finally, an elastic combination mechanism and a priority collaboration strategy ensure the deterministic response of the key interruption.
Owner:HUNAN GREAT WALL GALAXY TECH CO LTD

Bounded-Carry Fixed-Delay FP8 Accumulator for Pipelined Mixed-Precision Processing Elements

PendingUS20260203021A1Computer architectureCarry propagation
A bounded-carry partial-sum adder for FP8 accumulation in a pipelined mixed-precision processing element is disclosed. The adder limits carry propagation to a predetermined depth, ensuring deterministic fixed delay independent of operand magnitude. The accumulator forms the second stage of a two-stage fused multiply-add pipeline and supports a one-cycle initiation interval while accumulating products of FP4 weight operands and FP8 activation operands. The architecture enables uniform high-frequency timing across dense arrays of processing elements.
Owner:SILVEBROOK KIA

Encoder, Encoding Method, and Chip

PendingUS20260205142A1AlgorithmForward error correction
An encoder includes: a feedforward module configured to receive to-be-encoded data with a parallelism of n symbols per cycle, and perform calculation in a finite field based on a symbol of the to-be-encoded data in a current cycle, to obtain a single-cycle polynomial corresponding to the current cycle; and a feedback module configured to receive the single-cycle polynomial corresponding to the current cycle output by the feedforward module, and perform calculation in the finite field based on the single-cycle polynomial corresponding to the current cycle and a first polynomial indicating a symbol received in a historical cycle, to obtain a second polynomial, where the second polynomial is used to determine a target polynomial indicating the to-be-encoded data, and the target polynomial is to generate a check sequence of the to-be-encoded data, to obtain a codeword obtained by encoding the to-be-encoded data based on a forward error correction (FEC) encoding scheme.
Owner:HUAWEI TECH CO LTD

Performing multi-point table lookups in a single cycle in a system on a chip

The present disclosure relates to performing multi-point table lookups in a single cycle in a system on a chip. In various examples, a VPU and associated components can be optimized to improve VPU performance and throughput. For example, the VPU can include a min / max collector, an auto store prediction function, a SIMD data path organization that allows inter-channel sharing, a transpose load / store with a stride parameter function, a load with permute and zero insertion function, a hardware, logic, and memory layout function to allow two-point and two-point by one lookups, and a per-memory bank load cache function. Further, a decoupled accelerator can be used to offload VPU processing tasks to improve throughput and performance, and a hardware sequencer can be included in a DMA system to reduce programming complexity of the VPU and DMA system. The DMA and VPU can perform a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
Owner:NVIDIA CORP

Delay-Isomorphic and Mirror-Symmetric Physical Layout for Pipelined Low-Precision Floating-Point Processing Elements

PendingUS20260187024A1Propagation delayPERQ
A delay-isomorphic and mirror-symmetric physical layout for pipelined low-precision floating-point processing elements is described. Transistors and interconnects are arranged as mirrored halves about a central axis, with routing geometries matched to equalize propagation delay across corresponding functional paths. The layout supports one-result-per-cycle throughput in pipelined fused multiply-add units operating on formats including FP4 weights and FP8 activations. Power, ground, clock, and operand nets are routed symmetrically to maintain uniform timing and electrical characteristics. Arrays of such processing elements may use alternating mirrored orientation to reduce cumulative skew. The approach achieves deterministic multi-gigahertz operation suitable for large-scale inference architectures.
Owner:SILVEBROOK KIA

In a system-on-a-chip, performing multi-point table lookups in a single cycle.

To provide a VPU and associated components that are optimized to improve VPU performance and throughput.SOLUTION: A VPU includes a min / max collector, an automatic store predication function, a SIMD data path configuration allowing inter-lane sharing, a transposed load / store including a stride parameter function, a load including a permute and zero insertion function, a hardware, logic device, and memory layout function allowing two point and two by two point lookups, and per memory bank load caching ability. Decoupled accelerators are used to offload VPU processing tasks, and a hardware sequencer is included in a DMA system. The DMA and VPU execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for executing dynamic region-based data movement operations.SELECTED DRAWING: Figure 9A
Owner:NVIDIA CORP

A memory control method, a storage device, a medium and a computer device

The application discloses a memory control method, a storage device, a medium and computer equipment, and relates to the technical field of memory control, and the method comprises the following steps: acquiring the maximum allowed time consumption of a single I / O cycle based on the worst demand performance and the single write data length of a system preset, calculating the dynamic time budget that can be allocated to foreground garbage collection by the host write time consumption and the maximum allowed time consumption, and acquiring a theoretically optimal garbage collection intensity coefficient; monitoring the host write queue depth in real time to adjust the feedforward factor, and monitoring the number of idle blocks in real time to adjust the feedback factor; calculating the currently implemented garbage collection intensity, and allocating the time slice resource of the foreground garbage collection based on the currently implemented garbage collection intensity to control the execution intensity of the garbage collection. The application quantifies the abstract performance target of the worst demand performance into the time budget of a single I / O cycle, introduces a feedforward and feedback compound control mechanism, and realizes accurate and adaptive adjustment of the garbage collection intensity.
Owner:SHENZHEN XINGHUO SEMICON TECH CO LTD

Circuit implementation method for improving decoding speed of RISC-V architecture compression instruction

The invention relates to a circuit implementation method for improving the decoding speed of an RISC-V architecture compression instruction, and belongs to the technical field of decoding. Instruction boundaries are obtained from original data of memory / cache, then the compression instruction boundaries are extracted, RVC compression instruction fragments are extracted, and each compression instruction fragment is directly decoded, so that non-compression instruction boundaries are extracted, and the decoding speed of the RISC-V architecture compression instruction is improved. And extracting a standard 32-bit non-compression instruction, executing direct decoding on each non-compression instruction fragment, and finally performing complete decoding on the compression instruction or the non-compression instruction. The traditional steps 2 and 3 are combined and then are parallel to the step 4, and finally, a decoding result is selectively output through a decoding either-or circuit, so that the overall combinatorial logic level depth can be greatly reduced. In addition, in the parallel decoding process of the compressed instruction, rapid pre-decoding of the branch instruction can be achieved at the same time, the output speed of the branch prediction circuit is increased, the next program pointer can be output through single-cycle branch prediction, and therefore the performance of zero-bubble branch prediction is greatly improved.
Owner:SHENZHEN EVOLUTION CHUANGXIN ELECTRONIC TECHNOLOGY CO LTD

A method and system for ultra-parallel alignment

The method of single cycle super parallel comparison is designed by using FPGA, programmable logic or TCAM chip, which can complete the bitwise comparison between the key item and multiple table rows in a single logic cycle, output the address of the matched table row, and output the statistics data and position information of the same and different points. The algorithm supports table reconfiguration, same and different point processing, filter filtering, table mapping, one-dimensional array, two-dimensional data and multi-dimensional data comparison; the system includes a comparator array, reconfigurable logic, same and different point processor, mapping memory, filter, communication interface. It can form an independent comparison server and PCIE acceleration card. The method can speed up more than 10 9 orders of magnitude compared with the fastest CPU von Neumann computer comparison algorithm when comparing 10M table rows.
Owner:丁贤根

Memory control method, storage device, medium and computer equipment

The invention discloses a memory control method, a storage device, a medium and computer equipment, and relates to the technical field of storage control, and the method comprises the following steps: obtaining the maximum allowable time consumption of a single I / O period based on the worst demand performance preset by a system and the single write-in data length, calculating a dynamic time budget which can be allocated to foreground garbage collection through host write-in time consumption and maximum allowable time consumption, and obtaining a theoretical optimal garbage collection strength coefficient; monitoring the host write-in queue depth in real time to adjust a feed-forward factor, and monitoring the number of free blocks in real time to adjust a feedback factor; and calculating the currently implemented garbage recycling strength, and distributing time slice resources of foreground garbage recycling based on the currently implemented garbage recycling strength so as to control the execution strength of garbage recycling. According to the method, the abstract performance target of the worst demand performance is quantified into the time budget of a single I / O period, a feedforward and feedback composite control mechanism is introduced, and precise and self-adaptive adjustment of garbage collection strength is achieved.
Owner:SHENZHEN XINGHUO SEMICON TECH CO LTD