Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

27 results about "Instruction decoder" patented technology

Instruction decoder. The instruction decoder of a processor is a combinatorial circuit sometimes in the form of a read-only memory, sometimes in the form of an ordinary combinatorial circuit. Its purpose is to translate an instruction code into the address in the micro memory where the micro code for the instruction starts.

Data processor, method, electronic device and storage medium

The invention provides a data processor and method, electronic equipment and a storage medium, the data processor comprises an instruction scheduler and a copy engine, the instruction scheduler comprises a first tensor access register, and the first tensor access register is used for storing sub-tensor description information transmitted externally; the replication engine comprises a first queue structure, a second tensor access register and an instruction decoder, the second tensor access register is used for storing sub-tensor description information received from the first queue structure, and the instruction decoder is used for analyzing an instruction received from the first queue structure and generating a control signal to drive data handling; wherein the instruction scheduler is configured to dynamically detect a change parameter item of the sub-tensor description information, and write the change parameter item into the first queue structure, so that the copy engine incrementally updates the sub-tensor description information in the second tensor access register based on the change parameter item, therefore, the data transmission redundancy can be reduced, and the data transmission efficiency is improved. And the instruction processing efficiency is improved.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Monitoring memory tag load performance

Techniques for monitoring memory tag load performance are described. In an embodiment, an apparatus includes instruction decoder circuitry to decode a first instruction, the first instruction including an instruction format, the instruction format having a field for a value to indicate that execution of the first instruction is to include one or more memory tag checking operations; execution circuitry coupled to the instruction decoder circuitry, the execution circuitry to perform the one or more memory tag checking operations in response to the first instruction, wherein the one or more memory tag checking operations include a memory tag load operation; and performance monitoring circuitry to count occurrences of an event associated with the memory tag load operation.
Owner:INTEL CORP

ALU architecture low-power-consumption optimization system and method based on RISC-V

The invention is applicable to the technical field related to processors, and provides an ALU architecture low-power optimization system and method based on RISC-V. The system comprises a register file, an instruction decoder, a control state machine, an improved multiplier and an improved divider. The register file is used for storing operands; the instruction decoder is used for identifying a multiplication or division instruction and outputting a control signal; the control state machine is used for freezing or releasing the assembly line according to the control signal; the improved multiplier and the improved divider are used for executing low-power-consumption arithmetic operation; through the improvement, the power consumption and the area overhead of a chip are effectively reduced while the multiplier and the divider guarantee high-speed operation, and the method is particularly suitable for an SoC system with high requirements for energy efficiency and performance and shows remarkable advantages in application scenes such as a mobile terminal and artificial intelligence.
Owner:济南晶谷研究院 +1

Computer system and method for cache writeback and invalidation based on specified keys

Computer systems and methods for cache writeback and invalidation of specified keys. In one embodiment, in response to a first instruction of an instruction set architecture for writeback and invalidation of a hierarchical cache based on a single specified key identification, a decoder translates at least one microinstruction. According to the at least one microinstruction, a writeback and invalidation request is supplied to an in-core cache via a memory order buffer, and in turn by the in-core cache to a last level cache. In response to the writeback and invalidation request, the last level cache finds all matching cache lines that match the specified key identification, writes back to system memory those matching cache lines that have been modified and do not exist in a superior cache, and invalidates all matching cache lines found regardless of whether a state change has occurred. For a second instruction of an instruction set architecture for writeback and invalidation of a hierarchical cache based on multiple specified key identifications, the present application implements multiple writeback and invalidation requests.
Owner:VIA ALLIANCE SEMICON CO LTD

Processor, method, and system for accelerating tensor transpose for machine learning

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. It features an instruction decoder that decodes tensor transpose instructions for an input tensor, distinguishing between transposing and stationary axes. An inner transpose engine, equipped with multiple tensor buffer units of varying sizes, transposes the tensor by reading data by rows and writing by columns, efficiently managing memory when the buffer is full. Additionally, an address scheduler determines the target memory addresses in the output tensor memory for the tensor data, improving the handling and transformation of tensors in machine learning environments. This processor significantly enhances the efficiency and speed of tensor operations critical to advanced machine learning applications.
Owner:MOFFETT TECH CO LTD

Efficient tag checking for dynamically repeating memory accesses

PendingUS20260252243A1Computer architectureTerm memory
Techniques for tag checking for dynamically repeating memory accesses are described. In an embodiment, an apparatus includes instruction decoder circuitry to decode a single instruction, the single instruction having a format including an opcode field, the single instruction having a first opcode value in the opcode field; and execution circuitry coupled to the instruction decoder circuitry, the execution circuitry to perform operations in response to decoding of the single instruction, the operations including performing repeating memory tag checking operations in connection with repeating memory access operations.
Owner:INTEL CORP

Processor for controlling pipeline processing based on jump instruction, and program storage medium

Provided is a processor controlling pipeline processing to avoid the occurrence of pipeline bubbles as much as possible even when executing jump instruction. The present processor executes pipeline processing in which: an instruction fetcher fetching a machine-language instruction based on a memory address set in a program counter; a decoder decoding the machine-language instruction output from the instruction fetcher into control information; and an executer executing the control information output from the decoder, are connected, and the present processor comprises: a table describing a head address and a head machine-language instruction for each destination of jump; and a pipeline controller setting, when the executer executing a control information of a jump, an address specifying a second machine-language instruction at a destination of the jump into the program counter by using the table, while to input a head machine-language instruction at the destination of the jump to the decoder.
Owner:TAKEOKA LAB +1

Chip model and chip model expansion method

The invention provides a chip model and an expansion method of the chip model. The chip model comprises an RISC-V function model (a core function module and an expansion control module). The core function module at least comprises an instruction decoder, an instruction executor, a register block and an external interrupt processor; the expansion control module at least comprises an instruction code expansion interface, an instruction execution expansion interface, a register expansion interface, a peripheral interface and an external interrupt interface; the instruction code expansion interface and the instruction execution expansion interface respectively register functions of custom instructions to the instruction decoder and the instruction executor based on respective corresponding preset interface specifications; the register expansion interface is used for registering a self-defined register in the register group based on a corresponding preset interface specification; the peripheral interface interacts with other function models through a transaction-level modeling protocol; and the external interrupt interface transmits interrupt signals sent by other function models and received through the TLM to an external interrupt processor in the core function module.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

Techniques for efficient multiplication of complex vectors

A processing circuit is provided for performing vector operations, and an instruction decoder circuit is used to decode instructions from a set of instructions in order to control the processing circuit for performing the vector operations specified by the instructions. Array storage is used, having storage elements for storing data blocks, and when performing vector operations, it stores at least one two-dimensional array of data blocks accessible to the processing circuit. The set of instructions includes a complex-valued cross product instruction specifying a first source operand, a second source operand, and a destination operand, each of which the first and second source operands are vector operands containing multiple source data elements, each source data element being a complex number formed from a real part and an imaginary part, and the destination operand identifies a given two-dimensional array of data blocks in the array storage. The processing circuit, in response to a complex-valued cross product instruction, performs a cross product operation using the source data elements of the first source operand and the source data elements of the second source operand to generate multiple result data elements. Each result data element is a complex number formed from a real part and an imaginary part, and each real part and imaginary part of each result data element is associated with one of the data blocks in a given two-dimensional array of data blocks and used to update the value of that associated data block.
Owner:ARM LTD

Data processing

Data processing apparatus comprising: processing circuitry to apply a processing operation to one or more data items of a linear array, the linear array comprising a plurality of n data items at respective positions in the linear array, the processing circuitry configured to access an array of n x n storage locations, where n is an integer greater than one, the processing circuitry comprising: instruction decoder circuitry to decode program instructions; and instruction processing circuitry to execute instructions decoded by the instruction decoder circuitry; wherein the instruction decoder circuitry, in response to an array access instruction, controls the instruction processing circuitry to access a set of n storage locations arranged in an array direction as a linear array, the array direction being selected from a set of candidate array directions, the set of candidate array directions comprising at least a first array direction and a second array direction different from the first array direction, under control of the array access instruction.
Owner:ARM LTD

Processor, method, and system for accelerating tensor transpose for machine learning

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. It features an instruction decoder that decodes tensor transpose instructions for an input tensor, distinguishing between transposing and stationary axes. An inner transpose engine, equipped with multiple tensor buffer units of varying sizes, transposes the tensor by reading data by rows and writing by columns, efficiently managing memory when the buffer is full. Additionally, an address scheduler determines the target memory addresses in the output tensor memory for the tensor data, improving the handling and transformation of tensors in machine learning environments. This processor significantly enhances the efficiency and speed of tensor operations critical to advanced machine learning applications.
Owner:MOFFETT TECH CO LTD

Vector arithmetic processor and arithmetic execution method of vector arithmetic processor

To suppress a deterioration in processing performance due to a delay in execution of a subsequent instruction to use a mask register in the case of executing all set instructions to set all mask values of a mask register.SOLUTION: A vector arithmetic processor includes a mask register, an instruction decoder, a dependency resetting part, a scheduler and a vector arithmetic unit. The dependency resetting part resets dependency information showing data dependency with a preceding instruction that respectively corresponds to the mask register and a destination operand of a subsequent instruction to be transferred from the instruction decoder to the scheduler in the case that full set information is set by the instruction decoder that decodes all set instructions in correspondence with the mask register designated by the subsequent instruction. This can suppress a deterioration in processing performance due to a delay in execution of the subsequent instruction.SELECTED DRAWING: Figure 3
Owner:FUJITSU LTD

Memory interface

A memory interface circuit includes an instruction decoder configured to receive an instruction from a processor to generate a corresponding control code. An execution circuit is configured to receive the control code from the instruction decoder and access a memory and generate an arithmetic result according to the control code.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Efficient tag checking for dynamically repeated memory accesses

Techniques for tag checking for dynamically repeated memory accesses are described. In an embodiment, an apparatus comprises instruction decoder circuitry to decode a single instruction having a format including an opcode field, the single instruction having a first opcode value in the opcode field, and execution circuitry coupled with the instruction decoder circuitry, the execution circuitry to perform an operation in response to the decoding of the single instruction, the operation including performing a repeated memory tag check operation in association with a repeated memory access operation.
Owner:INTEL CORP

Data processor, method, electronic device and storage medium

The present disclosure provides a data processor, a method, an electronic device and a storage medium, the data processor comprising an instruction scheduler and a replication engine, the instruction scheduler comprising a first tensor access register configured to store sub-tensor description information transmitted externally; the replication engine comprising a first queue structure, a second tensor access register configured to store sub-tensor description information received from the first queue structure, and an instruction decoder configured to parse instructions received from the first queue structure and generate control signals to drive data transfer; wherein the instruction scheduler is configured to dynamically detect a change parameter item of the sub-tensor description information, write the change parameter item into the first queue structure, so that the replication engine incrementally updates the sub-tensor description information in the second tensor access register based on the change parameter item. Thus, the embodiments of the present disclosure can reduce data transmission redundancy and improve instruction processing efficiency.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Automatic operation method and device

The invention provides an automatic operation method and device, and relates to the technical field of artificial intelligence, in particular to the technical field of natural language processing, image processing, computer vision and deep learning. A specific embodiment of the method comprises the following steps: inputting multi-modal information into a corresponding multi-modal encoder for feature extraction, and outputting a multi-modal feature sequence; inputting the multi-modal feature sequence into a cross fusion layer for feature fusion, and outputting a fused feature sequence; inputting the fused feature sequence into a large decision model for reasoning decision, and outputting an action decision result; the action decision result is input into an instruction decoder for decoding conversion, and an operation instruction sequence is output; and inputting the operation instruction sequence into the small model actuator for action execution, and outputting an action execution result sequence.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD +1

CPU capable of quickly processing memory copy instruction, and method using same

Disclosed in the present invention are a CPU capable of quickly processing a memory copy instruction, and a method using same. The CPU comprises: an instruction decoder, a general-purpose register, a memory copy controller, a bus interface, a buffer, an adder, and a comparator. The memory copy controller comprises a state machine, and the state machine comprises an idle state, a read state, and a write state, wherein in the idle state, the memory copy controller waits to receive a valid memory copy instruction; in the read state, the memory copy controller reads data of a source address by means of the bus interface, and temporarily stores the data in the buffer; and in the write state, the memory copy controller writes the data, which is temporarily stored in the buffer, to a destination address by means of the bus interface. The adder is used for updating an address. The comparator is used for determining the end of copying. The present invention can maximize the utilization of a memory bandwidth to greatly improve the memory copy efficiency, can support an arbitrary alignment mode and support interruption, can reduce the power consumption of instruction fetching, and has a simple structure.
Owner:NANJING QINHENG MICROELECTRONICS CO LTD

An all-optical matrix computing unit and chip based on a Chinese character ternary instruction set

The application discloses a kind of full light matrix computing unit and chip based on Chinese character ternary instruction set, belong to photonic computing and computer architecture technical field.The full light matrix computing unit includes: Chinese character instruction decoder, for decoding Chinese character ternary instruction into ternary light field control signal;Coherent light source array, for generating N-way coherent light beam;Tri-state light modulator array receives the ternary light field control signal, and each coherent light beam is independently modulated as positive phase state, zero phase state or negative phase state, correspondingly three states of ternary logic;Optical matrix multiplication network is formed by cascaded directional coupler and phase shifter network, and matrix multiplication operation is carried out on the modulated multi-channel light beam;Balanced photodetector array is used to convert the optical signal of operation result into electrical signal and output.The application realizes ternary matrix computing architecture with Chinese character as native instruction and photon as computing carrier, and the calculation delay is only limited to photon transit time, and the energy efficiency of matrix multiplication operation is greatly improved.
Owner:林延明

Circuitry and method

Circuitry comprises instruction decoder circuitry to decode instructions for execution; processing circuitry to execute instructions decoded by the instruction decoder circuitry; interface circuitry defining an interface for data communication with data compression circuitry; in which the processing circuitry is responsive to one or more instructions of an instruction set defined for the processing circuitry to provide to the interface: input data to be processed by the data compression circuitry; and identification data identifying a compression system for use by the data compression circuitry to process the input data; and in which the processing circuitry is configured to receive from the interface: status data indicating whether data compression circuitry connected to the interface can process data using the compression system identified by the identification data; and, when the status data indicates that the data compression circuitry can process data using the compression system identified by the identification data, output data which has been processed from the input data by the data compression circuitry using the compression system identified by the identification data.
Owner:ARM LTD

User-level interprocessor interrupts

Processors, methods, and systems for user-level interprocessor interrupts are described. In an embodiment, a processing system includes a memory and a processing core. The memory is to store an interrupt control data structure associated with a first application being executed by the processing system. The processing core includes an instruction decoder to decode a first instruction, invoked by a second application, to send an interprocessor interrupt to the first application; and, in response to the decoded instruction, is to determine that an identifier of the interprocessor interrupt matches a notification interrupt vector associated with the first application; set, in the interrupt control data structure, a pending interrupt flag corresponding to an identifier of the interprocessor interrupt; and invoke an interrupt handler for the interprocessor interrupt identified by the interrupt control data structure.
Owner:INTEL CORP

Memory interface

A memory interface circuit includes an instruction decoder configured to receive an instruction from a processor to generate a corresponding control code. An execution circuit is configured to receive the control code from the instruction decoder and access a memory and generate an arithmetic result according to the control code
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

System and method of checking integrity of an instruction decoder of a processing system

A checker pipeline for checking integrity of an instruction decoder of a primary processor pipeline of a processing system including an instruction fetch checker and an instruction decoder checker. The processor pipeline includes an instruction fetch stage that receives an instruction with fields and the instruction decoder stage that decodes the instruction into instruction field values. The instruction fetch checker receives and converts instruction correction information provided with the instruction into instruction byte parity information. The instruction decoder checker includes a parity converter that converts the instruction byte parity information and instruction field information into predicted field parity information used to check the integrity of the instruction decoder. The instruction correction information is ECC bits or the like which are converted into instruction byte parity bits. The parity converter combines instruction byte parity bits with corresponding instruction bits using a logic operation into the predicted field parity information.
Owner:NXP USA INC

Technique for handling data elements stored in an array storage

An apparatus is provided comprising processing circuitry to perform operations, instruction decoder circuitry to decode instructions to control the processing circuitry to perform the operations specified by the instructions, and array storage comprising storage elements to store data elements. The array storage is arranged to store at least one two dimensional array of data elements accessible to the processing circuitry when performing the operations, each two dimensional array of data elements comprising a plurality of vectors of data elements, where each vector is one dimensional. The instruction decoder circuitry is arranged, in response to a move and zero instruction that identifies one or more vectors of data elements of a given two dimensional array of data elements within the array storage, to control the processing circuitry to move the data elements of the one or more identified vectors from the array storage to a destination storage and to set to a logic zero value the storage elements of the array storage that were used to store the data elements of the one or more identified vectors.
Owner:ARM LTD

Neural network accelerator, hybrid convolution-vector operation processing system and computer implementation method

Neural network accelerators, hybrid convolution-vector operation processing systems, and computer-implemented methods are described. The neural network accelerator includes an instruction decoder configured to decode a neural network computing instruction from a processor into a weight load control signal, an activation load control signal, and a computing control signal; a plurality of weight selectors configured to obtain a weight according to a weight load control signal indicating whether the weight is obtained from the weight cache or the weight generator; a plurality of activation input interfaces configured to obtain an activation or a vector from the memory in accordance with an activation load control signal indicating whether to obtain the activation or the vector; and a plurality of circuit channels. The instruction decoder is further configured to, in response to the weights having a mode, instruct the plurality of weight selectors to obtain weights from the weight generator, rather than from the weight cache, to reduce memory access.
Owner:MOZI INT CO LTD

Apparatus and method for managing deprecated instruction set architecture (ISA) features

ActiveUS12717603B2VirtualizationEngineering
An apparatus and method for implementing a new virtualized execution environment while supporting instructions and operations of a legacy virtualized execution environment. For example, one embodiment of a processor comprises: instruction processing circuitry to process instructions in accordance with a microarchitecture, the instruction processing circuitry comprising: instruction fetch circuitry to fetch the instructions; a decoder to decode the instructions; and execution circuitry to execute the instructions based on the microarchitecture; wherein the microarchitecture including hardware support for a virtual execution environment including a virtual machine monitor (VMM) and a first type of virtual machine, wherein both the VMM and the first type of virtual machine are implemented by instructions directly supported by the microarchitecture; and wherein the VMM is to support a second type of virtual machine, the second type of virtual machine including legacy instructions not fully supported by the microarchitecture, the VMM comprising a plurality of emulators, each emulator configured to emulate execution of a different type of the legacy instructions.
Owner:INTEL CORP

Tracker for individual branch misprediction cost

An embodiment of an integrated circuit may comprise a branch prediction unit to predict branches for an instruction decoder and circuitry coupled to the branch prediction unit, the circuitry to track a performance metric for an individual branch misprediction. Other embodiments are disclosed and claimed.
Owner:INTEL CORP

In-memory computing architecture for hybrid AI load

The invention belongs to the technical field of in-memory computing, and discloses an in-memory computing architecture for mixed AI load, which is characterized in that two functional regions for deeply optimizing different computing tasks are divided and integrated in a CIM macro cell, the two regions are physically a whole, but are divided into two heterogeneous partitions in logic function, and the two functional regions are divided into two heterogeneous partitions; comprising a high-density static page area and a two-dimensional writable dynamic page area, the high-density static page area is specially designed for memory access intensive tasks, the two-dimensional writable dynamic page area is specially designed for calculation intensive tasks, and the two areas share a part of peripheral circuits such as an instruction decoder and an instruction decoder. Therefore, the integration level and the resource utilization rate which are higher than those of two independent macros are realized. On the premise that the area and the power consumption efficiency are not sacrificed, high-throughput data reading capacity can be provided for memory access intensive tasks, and high-flexibility data processing and computing capacity can be provided for computing intensive tasks.
Owner:SEMICON TECH INNOVATION CENT(BEIJING) CORP +1