Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

31 results about "Superscalar" patented technology

A superscalar processor is a CPU that implements a form of parallelism called instruction-level parallelism within a single processor. In contrast to a scalar processor that can execute at most one single instruction per clock cycle, a superscalar processor can execute more than one instruction during a clock cycle by simultaneously dispatching multiple instructions to different execution units on the processor. It therefore allows for more throughput (the number of instructions that can be executed in a unit of time) than would otherwise be possible at a given clock rate. Each execution unit is not a separate processor (or a core if the processor is a multi-core processor), but an execution resource within a single CPU such as an arithmetic logic unit.

Caching method and system for access unit of superscalar processor

The invention belongs to the field of integrated circuits and computer system structures, and provides a caching method and system for a memory access unit of a superscalar processor, and the method comprises the steps: receiving a plurality of memory access instructions in the same period, and determining a corresponding Bank in a to-be-accessed cache through the memory access instructions; after the memory access instruction obtains a cache access permission, if cache line missing occurs, generating a missing request, merging all the missing requests, and performing parallel prefetching training on the merged missing requests by utilizing a mode of fusing a constant step length prefetching mode and a complex step length prefetching mode to obtain a prefetching request and a prefetching cache address corresponding to the prefetching request; requesting a missing cache line from the first-level cache to the second-level cache based on the missing queue, and writing the missing cache line back to the cache line of the corresponding data cache in the first-level cache; and storing the bus consistency request by using the sniffing queue, judging whether the data in the multi-core cache are consistent or not by using the consistency request, and performing consistency modification according to a judgment result. The cache hit rate and the bandwidth utilization rate are improved.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Processor program address buffering method and device

The invention provides a processor program address buffering method and device, and belongs to the technical field of processors, and the processor program address buffering device comprises the following components: an instruction fetching unit, a program address comparison unit, a program address buffer area, a program address buffer area tail value register, an instruction transmitting and executing unit and a reordering buffer area; the instruction fetching unit is in control connection with the program address comparison unit, and the program address comparison unit is in control connection with the program address buffer area, the program address buffer area tail value register, the instruction transmitting and executing unit and the reordering buffer area. In order to solve the problem that the disadvantage of the method is obvious in design iteration of instruction fetching and decoding width increasing and assembly line stage increasing of the superscalar processor, the invention provides a judgment basis for whether a program address is written into a program address buffer area or not through a program address comparison unit; the program addresses of all instructions are not written into the program address buffer area, so that the number of implementation table entries of the program address buffer area can be smaller.
Owner:SHANGHAI YIHUA TECHNOLOGY CO LTD

Instruction scheduling system and method and electronic equipment

The invention discloses an instruction scheduling system and method and electronic equipment, and relates to the technical field of computers, a dispatch module writes a to-be-scheduled instruction and a renamed register index combination into a dispatch queue, and a dependency check module constructs a dependency linked list according to the dispatch queue so as to determine an instruction execution sequence and send the instruction according to the instruction execution sequence; according to the method, the sending sequence and the execution sequence of the instruction with RAW dependency are ensured to be matched, so that the technical problems of low RAW dependency processing efficiency, long instruction waiting time and waste of pipeline instruction storage resources in the instruction scheduling of the superscale out-of-order processor can be solved, and the execution delay of the instruction is reduced.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Pseudo out-of-order instruction scheduling method based on branch jump

The invention provides a pseudo out-of-order instruction scheduling method based on branch jump, belongs to the technical field of processors, is applied to a dispatching and transmitting unit in a superscalar processor, and comprises the following steps: if a currently received instruction is a branch instruction, dispatching the current branch instruction to a branch instruction transmitting queue; judging whether the current branch instruction needs to be speculated and awakened or not; if it is determined that the current branch instruction needs to be speculated and awakened, whether the oldest branch instruction in the branch instruction transmitting queue can be awakened in the two clock periods or not is judged; and if it is determined that the oldest branch instruction in the branch instruction transmitting queue cannot be awakened in the two clock periods, suspending the operation of transmitting other instructions, after the oldest branch instruction, in the branch instruction transmitting queue and the non-branch instruction transmitting queue to the execution unit. Through conditional pause of the branch instruction, the cost of branch prediction errors caused by out-of-order execution is avoided, and the power consumption of a processor is remarkably reduced.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Multi-transmitting channel architecture optimization method and system based on superscalar processor

The invention provides a multi-transmitting-channel architecture optimization method and system based on a superscalar processor, and belongs to the technical field of processors, and the method comprises the steps: reading an instruction by adopting a data capturing and transmitting mechanism, judging whether a register is ready, introducing a feedforward data cache region, optimizing a data path, reading and obtaining a value of the register, and obtaining a multi-transmitting-channel architecture of the superscalar processor; the read values and instructions are stored in a transmitting queue; loading storage instruction transmission optimization is carried out on the instruction, a storage instruction is divided into a storage address and storage data, the storage data is allowed to be transmitted in advance, and table items are established in a loading storage unit; and reconstructing the transmitting queues, reconstructing the LSIQ transmitting queues into a Load instruction queue and a Store instruction queue, arbitrating the transmitting queues respectively, and waiting for execution. According to the method, the access pressure of the register file can be effectively relieved, the data feedforward efficiency is improved, RAW address conflicts are reduced, and the instruction transmitting efficiency and the overall processor performance are improved.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

A method and device for processing instruction address misalignment exception

The present invention discloses a method and device for processing an instruction fetch address misalignment exception. The method of the present invention splits an unconditional jump instruction into a first operation of calculating a link address and writing the result into a destination register, and a second operation of calculating a branch target address and completing the jump, and executes them in an integer calculation unit and a branch unit respectively. The conditional branch instruction only includes the first operation of calculating a branch target address and determining whether to jump, which is executed in the branch unit. The branch unit only implements one calculation module and does not set a result write bus. A completion field and a misalignment field are set in the ROB, and these two fields are updated when the operation is dispatched and executed. After the operation is executed, the first operation checks the misalignment field of its corresponding item in the ROB to determine whether to report an instruction fetch address misalignment exception. The present invention aims to provide a method for processing an instruction fetch address misalignment exception with low area overhead for an out-of-order superscalar processor of the RISC‑V architecture.
Owner:NAT UNIV OF DEFENSE TECH

A fault handling method based on a dual-core lockstep processor

This invention discloses a fault handling method based on a dual-core lockstep processor. The dual-core lockstep processor includes a master processor and a slave processor. The dual-core lockstep processor receives data from the last four stages of the processor's pipeline via a fault detection module, outputting a fault flag signal and a fault program pointer. A front-end fault handling module receives the fault flag signal and the fault program pointer, performing a pipeline-level rollback. A back-end fault handling module controls the handshake and selection signals of the last four stages of the pipeline and executes a virtual write-back fault-tolerance mechanism, managing the operating status of the last four stages of the pipeline and the instruction fetch stage pipeline. This prevents data conflicts during superscalar processor fault handling, thereby achieving full coverage of fault handling while maintaining low time and area overhead, and simultaneously achieving high performance and high reliability. This invention can be widely applied in the field of dual-core lockstep processors.
Owner:SUN YAT SEN UNIV

Energy-offset-resistant chassis lining rigidity parallel optimization method and system

InactiveCN121256932AGeometric CADDesign optimisation/simulationTransfer path analysisNoise
The invention discloses a chassis lining rigidity parallel optimization method and system capable of resisting energy deviation, and relates to the technical field of whole vehicle NVH optimization, and the method comprises the steps: carrying out the modeling of a road noise CAE simulation model, carrying out the analysis of a transmission path, calculating the contribution amount, arranging contribution paths in a descending order, and generating a parameterized lining file according to the linings of the contribution paths; performing dynamic polycondensation on the non-optimized region to form a CMS super-unit, generating a super-unit file, constructing a super-unit hybrid architecture based on the super-unit file and the parameterized lining file, and performing precision verification; and calculating a sound pressure value of a target point in the vehicle, converting the sound pressure value into an A weighting sound pressure level, and constructing an anti-energy offset optimization objective function based on a dynamic weighting superscale function of a superstandard frequency point and an adjacent frequency point correlation penalty function.
Owner:LIUZHOU RAILWAY VOCATIONAL TECHN COLLEGE

A cache method and system for superscalar processor memory access unit

ActiveCN120780659BCache accessCache hit rate
The application belongs to the field of integrated circuits and computer architecture, and provides a cache method and system for a superscalar processor memory access unit, which receives multiple memory access instructions in the same cycle, determines the corresponding Bank in the cache to be accessed by using the memory access instruction, generates a missing request when the cache access permission of the memory access instruction is obtained, merges all missing requests, uses a constant step prefetch mode and a complex step prefetch mode fusion method to perform parallel prefetch training on the merged missing requests, obtains a prefetch request and a corresponding prefetch cache address, requests a missing cache line from a level 1 cache to a level 2 cache based on a missing queue, and writes back the cache line in the corresponding data cache of the level 1 cache. The bus coherence request is stored in a sniffing queue, the data in the multi-core cache is judged to be consistent or not by using the coherence request, and the consistency is modified according to the judgment result. The application improves the cache hit rate and bandwidth utilization.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Method for defending ghost attacks by mshr-based superscalar risc-v processor hardware

The present application belongs to the field of integrated circuit design, and specifically relates to a method for preventing a spectre attack on a superscalar RISC-V processor based on a missing state holding register (MSHR). The method comprises modifying the MSHR module with the original pipeline logic of the Berkeley open source processor BOOM; specifically, adding judgment logic to the MSHR module to prevent non-safe instruction information from being written into the data cache before leaving the error speculation path based on the original MSHR state machine; and adding timeout logic to solve the deadlock problem under certain conditions; and combining the above two points to prevent the spectre attack. The present application can successfully prevent attack variants such as spectre without making large-scale changes to the pipeline architecture, and the performance loss is relatively small.
Owner:FUDAN UNIVERSITY

A verification method and system of a RISC-V out-of-order superscalar processor

PendingCN122450760AVerificationSimics
The application belongs to the technical field of processors, and provides a verification method and system of a RISC-V out-of-order superscalar processor, a test program is generated based on a random instruction generator of a RISC-V instruction set, and the test program is run in parallel by an out-of-order superscalar processor and an instruction set simulator; in the running process of the out-of-order superscalar processor, retired instruction information is sequentially extracted from a reorder buffer thereof, and corresponding instruction execution results are read from a physical register file according to the information of the retired instruction, a check information queue table item of sequential execution is generated, and is stored into a check information queue; through a preset interface, the instruction set simulator is controlled to run in single step, and sequential submission results corresponding to the check information queue table item are extracted and stored into a submission result queue; the table items in the check information queue and the submission result queue are compared one by one, if the comparison is consistent, the verification is continued, and if the comparison is inconsistent, an error instruction is located. The efficiency of verification is improved.
Owner:SHANDONG UNIV

Operation result cache management method and superscalar processor

The invention provides an operation result cache management method and a superscalar processor, and the method comprises the steps: in a transmission period of a first instruction, creating a target cache entry in a data undetermined state in a result cache module, and forwarding a target physical register index of the target cache entry to a transmission queue; the target cache entry and a target physical register used for storing a target operation result have a binding relationship, and the target operation result is obtained after the first instruction is executed; after the transmitting queue receives the target physical register index, screening out a second instruction depending on the target operation result, waking up the second instruction and marking the second instruction as a ready state; after the first instruction is executed, the target operation result is written into a target cache entry, and the target cache entry is marked as a valid state; and after the target cache entry is marked as a valid state, transmitting a second instruction in a ready state. According to the invention, delay bubbles on a dependency chain are eliminated.
Owner:CIX TECH (SHANGHAI) CO LTD

High-order matrix multiplication calculation system and method based on FPGA platform

The invention provides a high-order matrix multiplication calculation system and method based on an FPGA platform, and belongs to the technical field of computers.The system comprises a superscalar pipeline multiplication calculation module used for obtaining row data and column data of matrixes which are input in a serial mode and stored through a BRAM, the matrixes comprise the first matrix and the second matrix, and the first matrix and the second matrix are used for obtaining the row data and the column data of the matrixes; one BRAM stores a row of data in the first matrix; the superscale streamline multiplication calculation module is also used for carrying out multiplication calculation on each element in rows and columns of the matrix through a multiplier in a superscale serial streamline mode, and outputting each multiplication calculation result in parallel; and the superscale assembly line accumulation calculation module is used for performing assembly line processing on the superscale multiplication calculation result after receiving the superscale multiplication calculation result, performing accumulation calculation, and outputting a matrix multiplication result in parallel. According to the high-order matrix multiplication calculation system and method based on the FPGA platform, a matrix multiplication calculation structure capable of reducing the hardware resource consumption and guaranteeing the calculation speed can be provided.
Owner:BEIJING INST OF REMOTE SENSING EQUIP

RISC-V superscalar processor performance optimization method and system based on Gem5 platform

The invention provides an RISC-V superscale processor performance optimization method and system based on a Gem5 platform, and relates to the technical field of processor structure optimization. An instruction set architecture and an RISC-V superscale processor overall architecture are set on the Gem5 platform; according to the pipeline characteristics of the overall architecture of the RISC-V superscalar processor, each pipeline stage is divided into a front-end module, a middle-end module and a rear-end module; setting performance test parameters for the front end, the middle end and the rear end respectively, performing performance test on the front end, the middle end and the rear end based on the performance test parameters, and positioning a performance bottleneck and a bottleneck part of the processor according to a test result; according to the performance bottleneck and the specific bottleneck part, local architecture optimization adjustment is carried out, a progressive multi-level optimization process is set, the performance bottleneck and the bottleneck part of the RISC-V superscale processor are subjected to modular and overall optimization work, and the optimization modeling process of the RISC-V superscale processor is completed.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Pseudo out-of-order instruction scheduling method based on branch jump

The present invention provides a pseudo-out-of-order instruction scheduling method based on branch jumps, which belongs to the field of processor technology and is applied to a dispatch and emission unit in a superscalar processor. The method comprises: if the currently received instruction is a branch instruction, dispatching the current branch instruction to a branch instruction emission queue; determining whether the current branch instruction needs to be speculatively awakened; if it is determined that the current branch instruction needs to be speculatively awakened, determining whether the oldest branch instruction in the branch instruction emission queue can be awakened within two clock cycles; if it is determined that the oldest branch instruction in the branch instruction emission queue cannot be awakened within two clock cycles, pausing the operation of transmitting other instructions in the branch instruction emission queue and the non-branch instruction emission queue that follow the oldest branch instruction to the execution unit. By conditionally pausing branch instructions, the present invention avoids the cost of branch prediction errors caused by out-of-order execution and significantly reduces the power consumption of the processor.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Automatic circuit design method of TAGE branch predictor

The invention provides an automatic circuit design method of a TAGE branch predictor, which comprises the following steps: constructing a superscalar front-end simulator comprising a branch prediction unit, an instruction cache, an instruction queue and a prediction information cache, and simulating processor front-end behaviors through parameterized configuration; a time sequence and combinatorial logic decoupling normal form is introduced into the branch prediction unit, the branch prediction unit is decomposed into a time sequence component and a combinatorial logic component, the time sequence component comprises a global historical register, a folding historical register and a multi-stage prediction table, and the combinatorial logic component comprises a metadata generation unit, a direction prediction unit, a table item updating unit and a historical maintenance unit; obtaining input and output data of the branch prediction unit in the execution process; and based on the input and output data, generating a hardware description language code of the combinatorial logic component by using an automatic circuit design algorithm so as to realize automatic circuit design of the TAGE branch predictor. Therefore, the design efficiency and the performance of automatically designing the branch prediction component can be improved.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Superscalar delay optimization with divided issue queue

ActiveUS12360772B2Concurrent instruction executionEngineeringSuperscalar
Superscalar Delay Optimization with Divided Issue Queue, wherein an issue queue is divided into three types, namely, ready queue, wait 1 queue, and wait 2 queue; a length of each type of queue is one-third of a length of the issue queue, so that the delay of scanning the entire issue queue from beginning to end per clock cycle is reduced to one-third.
Owner:XU XIUQUAN

Register look-ahead renaming and static resource allocation method and device

The embodiment of the invention discloses an out-of-order superscale processor look-ahead renaming and static resource allocation method and a related device. The method is suitable for various out-of-order superscale processors. According to the embodiment of the invention, the blocking drive of the renaming level is expanded; whether the distribution level and the emission level respectively have enough resources to complete distribution and emission is monitored in advance for the head instruction of the renaming buffer area to be renamed; and if the distributed or transmitted resources are insufficient, renaming of the head instruction of the renaming buffer area is prohibited. The flow line blocking preposition measure relieves the bandwidth pressure of each stage after the renaming stage, and meaningless dynamic power consumption is reduced. On the other hand, according to the embodiment of the invention, on the basis of the probability distribution, obtained through real-time monitoring and statistics, of the blocking sources, the resources are pertinently and continuously adjusted, and the regression test is iterated until out-of-order scheduling does not become the performance bottleneck of the whole out-of-order superscalar processor.
Owner:XINGAOQIAO (SHANGHAI) SEMICONDUCTOR TECHNOLOGY CO LTD

Superscalar execution using pipelines that support different precisions

Techniques are disclosed relating to scheduling instructions for floating-point execution units with different capabilities. In some embodiments, a first pipeline is configured to execute a first type of floating-point operation on operands having up to a first precision and a second pipeline is configured to execute the first type of floating-point operation on operands having up to a second, greater precision. In some embodiments, round circuitry is configured to round results from an output precision of the second pipeline to an output precision of the first pipeline. Scheduling circuitry may select operations for issuance for a given cycle from multiple ready threads. This may include to prioritize a determined highest-precision operation of the first type from ready operations and assign the determined operation to a lowest-precision pipeline, of the multiple pipelines, that is configured to perform the first type of operation according to the operand precision of the determined operation.
Owner:APPLE INC

Superscalar field programmable gate array (FPGA) vector processor

The present disclosure relates to a vector processor implemented on programmable hardware (e.g., a field programmable gate array (FPGA) device). The vector processor includes a plurality of vector processor channels, wherein each vector processor channel includes a vector register file having a plurality of register file libraries and a plurality of execution units. Implementations described herein include features for optimizing resource availability on programmable hardware units and enabling superscalar execution when coupled with Time Single Instruction Multiple Data (SIMD).
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Single-Layer Perceptron Branch Prediction Method and System in Asynchronous Superscalar Processors

The present invention discloses a single-layer perceptron branch prediction method and system in an asynchronous superscalar processor. In the asynchronous superscalar processor architecture, after receiving a prediction request signal and a data packet sent by the instruction fetch module, the top-level control module of branch prediction sends prediction data to the perceptron branch prediction module. The perceptron branch prediction module loads the historical data of branch instructions and stores branch instruction information before instruction execution, performs weighted calculation using the branch instruction information, and determines whether the predicted branch instruction jumps according to the weighted result. Meanwhile, the historical record and weight of the branch instruction information are updated, and the prediction information is returned to the instruction fetch module for bidirectional instruction fetch. The correction information of various types of instructions in branch prediction is sent by the out module. The present invention is based on asynchronous circuit design. By introducing a dynamic branch prediction method based on perceptron and a bidirectional addressing mechanism, the instruction processing speed and prediction accuracy in complex and dynamic branch patterns are improved, and the power consumption is reduced.
Owner:LANZHOU UNIV

Automatic superscalar processor design method based on learning data dependency relationship

According to the automatic superscale processor design method based on the learning data dependency relationship, the superscale processor supporting instruction-level parallelism is automatically designed by learning data dependency between instructions, and the defect that dynamic dependency cannot be processed in the prior art is overcome. According to the technical scheme, dependency prediction is achieved based on a hardware-friendly machine learning model, and the most reusable state is selected from a high-dimensional processor state space through a state selector and stored in a small buffer area; and the state speculator is used for generating a hardware predictor by utilizing the selected state high-precision prediction dependence data and integrating the hardware predictor into the superscalar processor. The predictor is obtained by training a machine learning model S-BSD and comprises a state selector and a state speculator. According to the scheme, low-delay and high-precision prediction is achieved under the condition that hardware resources are limited, and parallel execution of multiple instructions is supported.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Method and equipment for processing misalignment abnormity of fetch addresses

The invention discloses an instruction fetch address misalignment exception processing method and device, and the method comprises the steps: dividing an unconditional jump instruction into a first operation for calculating a link address and writing a result into a destination register and a second operation for calculating a branch target address and completing jump, and executing the first operation and the second operation in an integer calculation part and a branch part respectively; the conditional branch instruction only comprises that a first operation for calculating a branch target address and judging whether to jump is executed on the branch component, the branch component only realizes one calculation module and does not set a result write bus, a completion field and an unaligned field are set in the ROB, and the two fields are updated when the operation is dispatched and executed. And after the operation is executed, the first operation checks the misalignment field of the corresponding item in the ROB to determine whether to report the misalignment abnormity of the fetch address. The invention aims to provide the processing method for the misalignment exception of the fetch address, which is small in area overhead, for the out-of-order superscale processor of the RISC-V architecture.
Owner:NAT UNIV OF DEFENSE TECH

An extension instruction interface of a superscalar RISC-V processor

The application belongs to the technical field of integrated circuit design, and particularly relates to an extended instruction interface based on a superscalar RISC-V processor pipeline. The extended instruction interface can be connected with the superscalar RISC-V processor pipeline, and normal processing and execution of an extended instruction in a RISC-V instruction set in the processor can be completed. The application comprises original pipeline modification logic and an extensible instruction management module of an open-source adamantium processor. The pipeline modification logic completes processing logic of the extensible instruction in the pipeline, including: extensible instruction decoding logic, extensible instruction emission logic, extensible instruction write-back and retirement logic. The extensible instruction management module is responsible for behavior management of the extensible instruction between the main processor and the coprocessor, including: an instruction state management module, an instruction information management module and an instruction completion management module. Through information interaction between the extensible instruction management module and the main pipeline, the running of a program containing the extended instruction in the processor can be realized.
Owner:FUDAN UNIVERSITY

RISC-V-based general-purpose neural network processor microarchitecture

This disclosure reveals a general-purpose neural network AI processor microarchitecture based on RISC-V and dedicated extended instruction sets, including a processor front-end unit, an instruction decoding and dispatch unit, a scalar execution unit, a vector-matrix execution unit, and a multi-level data storage unit. This microarchitecture employs a Turing-complete fine-grained instruction set to implement arbitrary algorithms and utilizes dedicated vector and matrix instructions for efficient computation of neural network operators, thus balancing computational power and flexibility for neural network inference. The microarchitecture adopts a superscalar out-of-order issue architecture in its hardware architecture, enabling concurrent execution of scalar, vector, and matrix instructions, optimizing the microarchitecture for deep neural network inference to ensure accelerator execution efficiency.
Owner:XI AN JIAOTONG UNIV

Method for instruction fusion with superscalar processors and related devices

The application provides a method for instruction fusion using a superscalar processor and related equipment. A decoding unit obtains first instruction control signals of multiple instructions to be executed, and sends the first instruction control signals to a fusion decoding unit. The fusion decoding unit determines, for any one of at least one instruction pair, that a first instruction in the any one instruction pair is an I-type instruction, a second instruction in the any one instruction pair is an I-type instruction or an R-type instruction, and judges whether there is an instruction pair matching the any one instruction pair in a plurality of preset instruction pairs. If there is, instruction fusion is performed under the condition that a source operand of the first instruction is equal to a destination operand of the first instruction and the destination operand of the first instruction is equal to at least one source operand of the second instruction, to obtain a fusion instruction. A fusion instruction execution unit executes the fusion instruction according to an operation logic of the fusion instruction.
Owner:SHENZHEN UNIV +1

Computing chip and instruction processing method to access source operands in private registers using a relative distance index

Embodiments of this application provide example computing chips and instruction processing method related to the field of integrated circuit technologies. One example computing chip uses a superscalar processor architecture, and includes an instruction processing unit and a plurality of registers that are separately coupled to the instruction processing unit. The plurality of registers include a general purpose register and a plurality of private registers that are separately coupled to the general purpose register. The general purpose register is configured to store an execution result of a microinstruction that is in a plurality of microinstructions of a computing task and that is executed before a jump instruction and whose execution result is referenced by a microinstruction that is executed after the jump instruction. Each private register in the plurality of private registers is configured to store an execution result of any microinstruction in the plurality of microinstructions.
Owner:HUAWEI TECH CO LTD

Arbiter for superscalar processor and superscalar processor

The invention provides an arbiter for a superscalar processor and the superscalar processor, and can be applied to the technical field of arbiters. The arbiter comprises M first arbitration subunits and (M-1) second arbitration subunits, wherein M is an integer greater than 1; the mth input sequence is input to the mth first arbitration subunit to obtain the mth output sequence, the mth input sequence is input to the mth second arbitration subunit to obtain the mth intermediate output sequence, the mth intermediate output sequence serves as the (m + 1) th input sequence, m is larger than 0 and smaller than M, and m is an integer; the first arbitration subunit is configured to output an input bit of which the first value is a target value in an input sequence as the target value, and output other input bits except the input bit of which the first value is the target value in the input sequence as negation values of the target value; and the second arbitration subunit is configured to output an input bit of which the first value is the target value in the input sequence as a negation value of the target value.
Owner:INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD

Superscalar Execution Using Pipelines That Support Different Precisions

Techniques are disclosed relating to scheduling instructions for floating-point execution units with different capabilities. In some embodiments, a first pipeline is configured to execute a first type of floating-point operation on operands having up to a first precision and a second pipeline is configured to execute the first type of floating-point operation on operands having up to a second, greater precision. In some embodiments, round circuitry is configured to round results from an output precision of the second pipeline to an output precision of the first pipeline. Scheduling circuitry may select operations for issuance for a given cycle from multiple ready threads. This may include to prioritize a determined highest-precision operation of the first type from ready operations and assign the determined operation to a lowest-precision pipeline, of the multiple pipelines, that is configured to perform the first type of operation according to the operand precision of the determined operation.
Owner:APPLE INC

A fine-grained, lockstep-tolerant, superscalar out-of-order processor design method and system

ActiveCN117667477BAvoid backend failuresEasy to detectNon-redundant fault processingEnergy efficient computingLockstepReservation station
This invention discloses a fine-grained, lock-step-tolerant, superscalar out-of-order processor design method and system. The method includes: reading the current instruction address and performing branch prediction to obtain the fetch address of the next instruction; decoding the instruction fetch to extract the opcode, source operand register number, and destination operand register number; renaming the instruction decoding result; processing the renamed instruction decoding result through a reservation station module; performing arithmetic and logical operations on the arbitrated instruction decoding result; comparing and checking the results of executed instructions and writing them into the system; reordering and buffering the written results of executed instructions before committing them to construct the superscalar out-of-order processor. This invention improves the fault detection and recovery capabilities of superscalar out-of-order processors. As a fine-grained, lock-step-tolerant, superscalar out-of-order processor design method and system, this invention can be applied to the field of fault-tolerant processor design technology.
Owner:SUN YAT SEN UNIV