Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

59 results about "Instruction decoder" patented technology

Instruction decoder. The instruction decoder of a processor is a combinatorial circuit sometimes in the form of a read-only memory, sometimes in the form of an ordinary combinatorial circuit. Its purpose is to translate an instruction code into the address in the micro memory where the micro code for the instruction starts.

Data processor, method, electronic device and storage medium

The invention provides a data processor and method, electronic equipment and a storage medium, the data processor comprises an instruction scheduler and a copy engine, the instruction scheduler comprises a first tensor access register, and the first tensor access register is used for storing sub-tensor description information transmitted externally; the replication engine comprises a first queue structure, a second tensor access register and an instruction decoder, the second tensor access register is used for storing sub-tensor description information received from the first queue structure, and the instruction decoder is used for analyzing an instruction received from the first queue structure and generating a control signal to drive data handling; wherein the instruction scheduler is configured to dynamically detect a change parameter item of the sub-tensor description information, and write the change parameter item into the first queue structure, so that the copy engine incrementally updates the sub-tensor description information in the second tensor access register based on the change parameter item, therefore, the data transmission redundancy can be reduced, and the data transmission efficiency is improved. And the instruction processing efficiency is improved.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

System and method for processing different instructions on same assembly line

The invention provides a system and method for processing different instructions on the same assembly line. The system comprises a control unit and an arithmetic unit, wherein the control unit is electrically connected with the arithmetic unit; the control unit comprises an instruction acquisition subunit, an instruction judgment subunit, an instruction register, an instruction decoder and a pipeline management subunit; the instruction registers include a first instruction register and a second instruction register. According to the method, instruction pipeline fusion is adopted for the two different types of instructions, and the two different instructions can run on the same pipeline through pipeline fusion, so that it is guaranteed that single-core hardware can process the two instructions on the same pipeline in a mixed mode, hardware resources are saved, and the hardware cost needed by the system is reduced; and meanwhile, the system has the advantages of high processing speed for the instruction of the first instruction type and high processing efficiency for the instruction of the second instruction type.
Owner:HUNAN ADVANCECHIP ELECTRONICS TECH CO LTD

Processor, method, device and storage medium for data processing

According to an embodiment of the present disclosure, a processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand and a target operand. The target opcode indicates the vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading data to be processed. The target operand specifies a target storage location in the memory for writing a processing result. The processor also includes an arithmetic logic unit coupled to the instruction decoder and the memory. The arithmetic logic unit is configured to: read data to be processed from a source storage location in the memory; perform an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed; and write the processing result to a target storage location in the memory. In this way, the efficiency of vector calculations can be improved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Processor, method, device and storage medium for data processing

A processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading to-be-processed data. The target operand specifies a target storage location in the memory for writing a processed result. The processor further includes an arithmetic logic unit configured to: read the to-be-processed data from the source storage location of the memory; perform, on the to-be-processed data, an arithmetic logic operation associated with the vector operation specified by the target instruction; and write the processed result to the target storage location of the memory.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Large-scale matrix restructuring and matrix-scalar operations

Embodiments of apparatuses and methods for copying and operating on matrix elements are described. In embodiments, an apparatus includes a hardware instruction decoder to decode a single instruction and execution circuitry, coupled to hardware instruction decoder, to perform one or more operations corresponding to the single instruction. The single instruction has a first operand to reference a base address of a first representation of a source matrix and a second operand to reference a base address of second representation of a destination matrix. The one or more operations include copying elements of the source matrix to corresponding locations in the destination matrix and filling empty elements of the destination matrix with a single value.
Owner:INTEL CORP

Data processing device, method and equipment

The invention discloses a data processing device, method and equipment, and belongs to the field of computers. The device comprises; the instruction decoder is used for decoding a four-word loading instruction, a first memory address and a target register are specified in the four-word loading instruction, and the processing circuit is used for responding to the four-word loading instruction and determining a next register of the target register; and the data loading module is used for reading continuous four-word data from the first memory address and orderly loading the continuous four-word data to the target register and the next register. Therefore, on the premise that the instruction format of the processor architecture is followed, the four-word instruction is loaded, and conflicts with existing instructions of the architecture are avoided. Based on similar principles, the device can store four-word instructions.
Owner:BEIJING ESWIN COMPUTING TECH CO LTD

Scheduling Tasks Using Swap Flags

A method of activating scheduling instructions within a parallel processing unit is described. The method comprises decoding, in an instruction decoder, an instruction in a scheduled task in an active state and checking, by an instruction controller, if a swap flag is set in the decoded instruction. If the swap flag in the decoded instruction is set, a scheduler is triggered to de-activate the scheduled task by changing the scheduled task from the active state to a non-active state.
Owner:IMAGINATION TECH LTD

Monitoring memory tag load performance

Techniques for monitoring memory tag load performance are described. In an embodiment, an apparatus includes instruction decoder circuitry to decode a first instruction, the first instruction including an instruction format, the instruction format having a field for a value to indicate that execution of the first instruction is to include one or more memory tag checking operations; execution circuitry coupled to the instruction decoder circuitry, the execution circuitry to perform the one or more memory tag checking operations in response to the first instruction, wherein the one or more memory tag checking operations include a memory tag load operation; and performance monitoring circuitry to count occurrences of an event associated with the memory tag load operation.
Owner:INTEL CORP

Efficient implementation of floating point exponential functions in processor

A processor includes an instruction decoder configured to provide at least a floating point instruction control signal; performing floating point calculation on a data path; a floating point custom instruction control logic block coupled to the floating point compute data path; a control and status register coupled to the floating point computational data path and the floating point custom instruction control logic block; a first multiplexer configured to provide a floating-point instruction control signal or a custom instruction control signal to the floating-point computational data path based on a state of a selection control signal; and a second multiplexer configured to provide a floating point operand or a custom operand to the floating point computing data path based on a state of the selection control signal. The floating point custom instruction control logic block asserts a select signal while directing the floating point computational data path to assist it in executing custom instructions. The custom instruction may be a floating point exponential function.
Owner:NXP BV

Code prefetch instruction

Embodiments of apparatuses, methods, and systems for code prefetching are described. In an embodiment, an apparatus includes an instruction decoder, load circuitry, and execution circuitry. The instruction decoder is to decode a code prefetch instruction. The code prefetch instruction is to specify a first instruction to be prefetched. The load circuitry to prefetch the first instruction in response to the decoded code prefetch instruction. The execution circuitry is to execute the first instruction at a fetch stage of a pipeline.
Owner:INTEL CORP

Vector processor and method of executing arithmetic operation in vector processor

A vector processor includes a mask register configured to hold mask values, an instruction decoder configured to set dependency information included in instruction execution information when a decoded instruction is a subsequent instruction having data dependency with one or more previous instructions, and to set all-set information included in instruction execution information when a decoded instruction sets all of the mask values, a vector processing unit configured to execute vector arithmetic operations based on the instruction execution information, and to store in the data register a result of an arithmetic operation of a vector element corresponding to each mask value that is in a set state, and a dependency reset unit configured to reset the dependency information corresponding to a destination operand of the subsequent instruction and the mask register, when the all-set information is set for the mask register and the mask register is designated by the subsequent instruction.
Owner:FUJITSU LTD

Load chunk instruction and store chunk instruction

Processing circuitry (16) and an instruction decoder (9) supports a load chunk instruction and a store chunk instruction which can be useful for implementing memory copy functions and other library functions for manipulating or comparing blocks of memory. Number of bytes to load or store in response to these instructions is determined based on an implementation specific condition. As well as loading or storing bytes of data, the load chunk instruction and (10) store chunk instruction also designated a load / store length value as data corresponding to an architecturally visible register, which provides an indication of a number of bytes loaded or stored.
Owner:ARM LTD

ALU architecture low-power-consumption optimization system and method based on RISC-V

The invention is applicable to the technical field related to processors, and provides an ALU architecture low-power optimization system and method based on RISC-V. The system comprises a register file, an instruction decoder, a control state machine, an improved multiplier and an improved divider. The register file is used for storing operands; the instruction decoder is used for identifying a multiplication or division instruction and outputting a control signal; the control state machine is used for freezing or releasing the assembly line according to the control signal; the improved multiplier and the improved divider are used for executing low-power-consumption arithmetic operation; through the improvement, the power consumption and the area overhead of a chip are effectively reduced while the multiplier and the divider guarantee high-speed operation, and the method is particularly suitable for an SoC system with high requirements for energy efficiency and performance and shows remarkable advantages in application scenes such as a mobile terminal and artificial intelligence.
Owner:济南晶谷研究院 +1

CPU capable of precise delay control and precise delay control method

The present invention discloses a CPU capable of precise delay control and a precise delay control method, comprising: an instruction decoder for decoding a valid delay instruction and outputting a delay request and instruction information; an adder having a first delay value and a second delay value inputted into its input terminal and a counter connected to its output terminal; a counter that, upon receiving a delay request, reloads the output value of the adder, counts, and outputs the count result to a delay controller; a delay monitor for receiving instruction matching information, matching according to detected pipeline instructions, and transmitting the matching result to the delay controller; and a delay controller for receiving the delay request and instruction matching information outputted by the instruction decoder, the counting result outputted by the counter, and the matching result outputted by the delay monitor, and generating or canceling a pause request to a CPU pipeline controller. The present invention has high delay accuracy, deterministic results, high processor utilization, and low power consumption.
Owner:NANJING QINHENG MICROELECTRONICS CO LTD

Histogram operation

The invention relates to histogram operation. A digital data processor (100) includes: an instruction memory (121) storing instructions each specifying a data processing operation and at least one data operation digit segment; an instruction decoder (113) coupled to the instruction memory for sequentially invoking instructions from the instruction memory and determining the data processing operation and the at least one data operand; and at least one arithmetic unit (110) coupled to the data register file (123) and to an instruction decoder to perform a data processing operation on at least one operand corresponding to an instruction decoded by the instruction decoder and to store a result of the data processing operation. The arithmetic unit is configured to incrementing a histogram value in response to a histogram instruction by incrementing a bin entry at a specified location in at least one histogram of a specified number.
Owner:TEXAS INSTRUMENTS INC

Processing of iterative operations

Processing of iterative operations is disclosed. An apparatus has processing circuitry to perform an iterative operation in response to a decoding of an iterative operation instruction by an instruction decoder, the iterative operation including at least two iterations of processing in which an iteration depends on an operand produced in a preceding iteration. Initial information producing circuitry performs an initial portion of the processing of a given iteration to produce initial information. Result producing circuitry performs a remaining portion of the processing of the given iteration to produce a result value using the initial information. For iterations other than a final iteration, forwarding circuitry forwards the result value as an operand for a next iteration of the iterative operation. The initial information producing circuitry begins performing an initial portion of the next iteration in parallel with the result producing circuitry completing the remaining portion of the current iteration to improve performance.
Owner:ARM LTD

Computer system and method for cache writeback and invalidation based on specified keys

Computer systems and methods for cache writeback and invalidation of specified keys. In one embodiment, in response to a first instruction of an instruction set architecture for writeback and invalidation of a hierarchical cache based on a single specified key identification, a decoder translates at least one microinstruction. According to the at least one microinstruction, a writeback and invalidation request is supplied to an in-core cache via a memory order buffer, and in turn by the in-core cache to a last level cache. In response to the writeback and invalidation request, the last level cache finds all matching cache lines that match the specified key identification, writes back to system memory those matching cache lines that have been modified and do not exist in a superior cache, and invalidates all matching cache lines found regardless of whether a state change has occurred. For a second instruction of an instruction set architecture for writeback and invalidation of a hierarchical cache based on multiple specified key identifications, the present application implements multiple writeback and invalidation requests.
Owner:VIA ALLIANCE SEMICON CO LTD

Processor, method, and system for accelerating tensor transpose for machine learning

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. It features an instruction decoder that decodes tensor transpose instructions for an input tensor, distinguishing between transposing and stationary axes. An inner transpose engine, equipped with multiple tensor buffer units of varying sizes, transposes the tensor by reading data by rows and writing by columns, efficiently managing memory when the buffer is full. Additionally, an address scheduler determines the target memory addresses in the output tensor memory for the tensor data, improving the handling and transformation of tensors in machine learning environments. This processor significantly enhances the efficiency and speed of tensor operations critical to advanced machine learning applications.
Owner:MOFFETT TECH CO LTD

Efficient tag checking for dynamically repeating memory accesses

PendingUS20260252243A1Computer architectureTerm memory
Techniques for tag checking for dynamically repeating memory accesses are described. In an embodiment, an apparatus includes instruction decoder circuitry to decode a single instruction, the single instruction having a format including an opcode field, the single instruction having a first opcode value in the opcode field; and execution circuitry coupled to the instruction decoder circuitry, the execution circuitry to perform operations in response to decoding of the single instruction, the operations including performing repeating memory tag checking operations in connection with repeating memory access operations.
Owner:INTEL CORP

Processor for controlling pipeline processing based on jump instruction, and program storage medium

Provided is a processor controlling pipeline processing to avoid the occurrence of pipeline bubbles as much as possible even when executing jump instruction. The present processor executes pipeline processing in which: an instruction fetcher fetching a machine-language instruction based on a memory address set in a program counter; a decoder decoding the machine-language instruction output from the instruction fetcher into control information; and an executer executing the control information output from the decoder, are connected, and the present processor comprises: a table describing a head address and a head machine-language instruction for each destination of jump; and a pipeline controller setting, when the executer executing a control information of a jump, an address specifying a second machine-language instruction at a destination of the jump into the program counter by using the table, while to input a head machine-language instruction at the destination of the jump to the decoder.
Owner:TAKEOKA LAB +1

Techniques for efficient complex vector multiplication

Processing circuitry to perform a vector operation and instruction decoder circuitry to decode an instruction from an instruction set to control the processing circuitry to perform the vector operation specified by the instruction are provided. An array storage device having storage elements for storing data blocks is used to store at least one two-dimensional array of data blocks, the processing circuitry being accessible to the at least one two-dimensional array of data blocks when the vector operation is performed. The instruction set includes a complex numerical outer product instruction specifying a first source operand, a second source operand, and a destination operand, where each of the first source operand and the second source operand is a vector operand including a plurality of source data elements, each source data element being a complex number formed by a real part and an imaginary part, and the destination operand identifies a given two-dimensional array of data blocks within the array storage device. The processing circuitry, in response to the complex value outer product instruction, performs an outer product operation using the source data element of the first source operand and the source data element of the second source operand to produce a plurality of result data elements, where each result data element is a complex number formed by a real part and an imaginary part, and wherein each real part and each imaginary part of each result data element are associated with one of the data blocks in the given two-dimensional array of data blocks and are used to update the value of the associated data block.
Owner:ARM LTD

Chip model and chip model expansion method

The invention provides a chip model and an expansion method of the chip model. The chip model comprises an RISC-V function model (a core function module and an expansion control module). The core function module at least comprises an instruction decoder, an instruction executor, a register block and an external interrupt processor; the expansion control module at least comprises an instruction code expansion interface, an instruction execution expansion interface, a register expansion interface, a peripheral interface and an external interrupt interface; the instruction code expansion interface and the instruction execution expansion interface respectively register functions of custom instructions to the instruction decoder and the instruction executor based on respective corresponding preset interface specifications; the register expansion interface is used for registering a self-defined register in the register group based on a corresponding preset interface specification; the peripheral interface interacts with other function models through a transaction-level modeling protocol; and the external interrupt interface transmits interrupt signals sent by other function models and received through the TLM to an external interrupt processor in the core function module.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

Techniques for efficient multiplication of complex vectors

A processing circuit is provided for performing vector operations, and an instruction decoder circuit is used to decode instructions from a set of instructions in order to control the processing circuit for performing the vector operations specified by the instructions. Array storage is used, having storage elements for storing data blocks, and when performing vector operations, it stores at least one two-dimensional array of data blocks accessible to the processing circuit. The set of instructions includes a complex-valued cross product instruction specifying a first source operand, a second source operand, and a destination operand, each of which the first and second source operands are vector operands containing multiple source data elements, each source data element being a complex number formed from a real part and an imaginary part, and the destination operand identifies a given two-dimensional array of data blocks in the array storage. The processing circuit, in response to a complex-valued cross product instruction, performs a cross product operation using the source data elements of the first source operand and the source data elements of the second source operand to generate multiple result data elements. Each result data element is a complex number formed from a real part and an imaginary part, and each real part and imaginary part of each result data element is associated with one of the data blocks in a given two-dimensional array of data blocks and used to update the value of that associated data block.
Owner:ARM LTD

Encoding special values in anchored data elements

An apparatus comprising processing circuitry to perform data processing; and an instruction decoder to control the processing circuitry to perform an anchored data processing operation to generate an anchored data element. The anchored data element has an encoding comprising type information indicating whether the anchored data element represents a portion of a bit of a two's complement number that corresponds to a given range of significant values that can be represented using the anchored data element; or represents a special value other than the portion of the bit of the two's complement number.
Owner:ARM LTD

Data processing

Data processing apparatus comprising: processing circuitry to apply a processing operation to one or more data items of a linear array, the linear array comprising a plurality of n data items at respective positions in the linear array, the processing circuitry configured to access an array of n x n storage locations, where n is an integer greater than one, the processing circuitry comprising: instruction decoder circuitry to decode program instructions; and instruction processing circuitry to execute instructions decoded by the instruction decoder circuitry; wherein the instruction decoder circuitry, in response to an array access instruction, controls the instruction processing circuitry to access a set of n storage locations arranged in an array direction as a linear array, the array direction being selected from a set of candidate array directions, the set of candidate array directions comprising at least a first array direction and a second array direction different from the first array direction, under control of the array access instruction.
Owner:ARM LTD

Processor, method, and system for accelerating tensor transpose for machine learning

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. It features an instruction decoder that decodes tensor transpose instructions for an input tensor, distinguishing between transposing and stationary axes. An inner transpose engine, equipped with multiple tensor buffer units of varying sizes, transposes the tensor by reading data by rows and writing by columns, efficiently managing memory when the buffer is full. Additionally, an address scheduler determines the target memory addresses in the output tensor memory for the tensor data, improving the handling and transformation of tensors in machine learning environments. This processor significantly enhances the efficiency and speed of tensor operations critical to advanced machine learning applications.
Owner:MOFFETT TECH CO LTD

Vector arithmetic processor and arithmetic execution method of vector arithmetic processor

To suppress a deterioration in processing performance due to a delay in execution of a subsequent instruction to use a mask register in the case of executing all set instructions to set all mask values of a mask register.SOLUTION: A vector arithmetic processor includes a mask register, an instruction decoder, a dependency resetting part, a scheduler and a vector arithmetic unit. The dependency resetting part resets dependency information showing data dependency with a preceding instruction that respectively corresponds to the mask register and a destination operand of a subsequent instruction to be transferred from the instruction decoder to the scheduler in the case that full set information is set by the instruction decoder that decodes all set instructions in correspondence with the mask register designated by the subsequent instruction. This can suppress a deterioration in processing performance due to a delay in execution of the subsequent instruction.SELECTED DRAWING: Figure 3
Owner:FUJITSU LTD

Electronic circuit and method for firmware anomaly detection

An electronic circuit comprising: a processing unit comprising or connected to a memory containing a firmware with executable instructions; the processing unit comprising: an instruction decoder, instruction registers and an arithmetic Logical Unit; wherein the instruction decoder is configured for receiving instructions, and for converting them into operation codes and operands; and being further configured for generating one or more groupcode corresponding to the received instruction, thereby forming a stream of groupcodes during execution of said firmware by said processing unit; wherein the electronic circuit further comprises a groupcode analyser configured for receiving said stream of groupcodes, and for detecting anomalies in the firmware execution by analysing said stream of groupcodes. A method of detecting firmware anomalies.
Owner:MELEXIS TECHNOLOGIES SA

Memory interface

A memory interface circuit includes an instruction decoder configured to receive an instruction from a processor to generate a corresponding control code. An execution circuit is configured to receive the control code from the instruction decoder and access a memory and generate an arithmetic result according to the control code.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Efficient tag checking for dynamically repeated memory accesses

Techniques for tag checking for dynamically repeated memory accesses are described. In an embodiment, an apparatus comprises instruction decoder circuitry to decode a single instruction having a format including an opcode field, the single instruction having a first opcode value in the opcode field, and execution circuitry coupled with the instruction decoder circuitry, the execution circuitry to perform an operation in response to the decoding of the single instruction, the operation including performing a repeated memory tag check operation in association with a repeated memory access operation.
Owner:INTEL CORP