Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

36 results about "Instruction prefetch" patented technology

Cache prefetching is a technique used by computer processors to boost execution performance by fetching instructions or data from their original storage in slower memory to a faster local memory before it is actually needed (hence the term 'prefetch'). Most modern computer processors have fast and local cache memory in which prefetched data is held until it is required. The source for the prefetch operation is usually main memory. Because of their design, accessing cache memories is typically much faster than accessing main memory, so prefetching data and then accessing it from caches is usually many orders of magnitude faster than accessing it directly from main memory.

Prefetch instruction management method, system and equipment

The invention belongs to the technical field of computers, particularly relates to a prefetch instruction management method, system and equipment, and aims to solve the problem of low instruction fetch efficiency of an instruction prefetch technology. The method comprises the following steps: receiving branch prediction information from a branch prediction unit, and performing label comparison on table entries in an instruction fetching target queue and the branch prediction information; under the condition that the table item is hit, querying an instruction fetching address corresponding to the branch prediction information in the hit table item; under the condition that the table item is not hit, a new storage table item is allocated to the branch prediction information, and an instruction fetching address carried by the branch prediction information is determined through the storage table item; in response to a received prefetching request sent by the prefetching unit, determining a target table item based on the prefetching request, and returning an instruction fetching address queried in the target table item to the prefetching unit; and storing the stable cache line in the target table item to a flow buffer area. According to the method, multiple prediction requests can be processed in parallel, the front-end throughput is improved, and the instruction fetching efficiency is improved.
Owner:SHANDONG UNIV +1

Data storage instruction processing method, electronic equipment and storage medium

The invention provides a data storage instruction processing method, electronic equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps that a block address of a target high-speed cache block corresponding to a target cache line is sent into a data storage instruction prefetching queue, all target data storage instructions cached by the target cache line are not fully written in the target cache line, and a data storage instruction address of the target data storage instruction comprises the block address of the target high-speed cache block; and executing data storage address hit check of the data cache on a target data storage instruction in the data storage instruction prefetching queue, and initiating a prefetching request and executing a block address access lower-layer storage system based on a target cache block under the condition that the data storage address is invalid, and the data of the target cache block is obtained and backfilled to the data cache, and the number storage instruction data included in the target number storage instruction which initiates the pre-fetching request and is successfully pre-fetched is written into the target cache block in the data cache, so that the pre-fetching accuracy and the pre-fetching timeliness of the number storage instruction can be improved.
Owner:BEIJING VCORE TECH CO LTD

Instruction cache and prefetching method thereof

The invention provides an instruction cache and a prefetching method thereof. The instruction cache comprises a hit detection module, an instruction prefetching module and an operation module. The hit detection module checks whether a current instruction fetching request of a current thread bundle is hit. The instruction prefetching module dynamically determines an instruction prefetching range corresponding to the current instruction fetching request according to the execution progress of the current thread bundle. The operation module accesses the external storage to prefetch the instruction data corresponding to the instruction prefetch range. And in response to the fact that the hit detection module judges that the current instruction fetching request is hit, the operation module feeds back instruction data corresponding to the current instruction fetching request to the current thread bundle. And in response to the fact that the hit detection module judges that the current instruction fetching request is missed, the operation module accesses external storage so as to feed back instruction data corresponding to the current instruction fetching request to the current thread bundle.
Owner:SHANGHAI BIREN TECH CO LTD

Intelligent session flow scheduling method under converged communication architecture

The invention relates to the technical field of cloud fusion application operation support platforms, and discloses a session flow intelligent scheduling method under a fusion communication architecture, which comprises the following steps: monitoring a session flow logic sequence number of a current node; constructing an asynchronous mirror instance at the target node, and obtaining a computing power comparison parameter; adjusting an asynchronous mirror instance instruction execution rate to execute serial number chasing; after the sequence number difference value enters a synchronous threshold value, processor assembly line branch prediction data is injected, and instruction prefetching sequence filling is driven; according to the instruction level feature injection and execution phase alignment method, hardware execution momentum imbalance in a heterogeneous computing power environment is eliminated through instruction level feature injection and execution phase alignment, and progress connection and logic consistency of session streams at a migration interface are guaranteed.
Owner:SHENZHEN JINCHENGKE INFORMATION TECH CO LTD

Instruction prefetching method, processor and electronic equipment

The embodiment of the invention discloses an instruction prefetching method, a processor, electronic equipment and a computer readable storage medium. The method comprises the following steps: predicting and generating an instruction fetching block of a target instruction based on a branch prediction unit, and sending the instruction fetching block into an instruction fetching target queue; if the fetch block does not carry the first group of index information and the first path information of the target instruction, sending a prefetch request to an instruction cache unit based on the fetch target queue so as to cache the target instruction into the instruction cache unit; based on an instruction fetch unit, reading a fetch block from the fetch target queue so as to read a target instruction from an instruction cache unit; if the fetch block carries the first group of index information and the first path information, the fetch block is read from the fetch target queue based on the instruction fetch unit to obtain the first group of index information and the first path information, and the target instruction is read from the instruction cache unit according to the first group of index information and the first path information, so that redundant pre-fetch requests are reduced, and the operation efficiency is improved. And the power consumption of the processor is reduced.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

Sparse triangular linear system solving accelerator based on FPGA

The invention discloses a sparse triangular linear system solving accelerator based on an FPGA (Field Programmable Gate Array), and belongs to the technical field of calculation, reckoning or counting. The accelerator comprises an HBM, an instruction prefetching and decoding module, an input cache module, an FIFO module, a processing unit array, a partial sum register file module and a data decoding register file module. A super-long instruction word architecture instruction calculation scheduling and data information generation are adopted, FIFO and a register file are used as cache modules for node conversion and node multiplexing, periodic automatic alignment is realized without synchronization cost, a full interconnection structure realizes temporary storage of multiplexing data and intermediate results in the register file, and non-reused calculation results are directly written back. According to the method, a double-floating-point calculation operation granularity capable of fully reusing nodes and parallelism is adopted, a multi-PE parallel high-reusability operation mode is fully exerted, calculation control is simplified through scheduling of software and hardware, and an implementation scheme for efficient parallelism of SpTRSV on an FPGA is provided.
Owner:SOUTHEAST UNIV

Calculation scheduling method and device for AI chip, electronic equipment and storage medium

The invention provides a calculation scheduling method and device for an AI chip, electronic equipment and a storage medium. The method comprises the steps that a task distribution unit sends information of an instruction needing to be loaded in advance to a prefetching unit; in response to the information, received by the prefetching unit, of the instruction needing to be loaded in advance, the prefetching unit sends an instruction prefetching request corresponding to the instruction needing to be loaded in advance to the memory based on the information of the instruction needing to be loaded in advance; in response to the instruction prefetching request received by the memory, the memory sends an instruction needing to be loaded in advance to the prefetching unit; and in response to the instruction which needs to be loaded in advance and is received by the prefetching unit, the prefetching unit sends the instruction which needs to be loaded in advance to the plurality of calculation cores so as to cache the instruction which needs to be loaded in advance to the buffer corresponding to each calculation core.
Owner:BEIJING WEIFAN INTELLIGENT TECHNOLOGY CO LTD

Code prefetch instruction

Embodiments of apparatuses, methods, and systems for code prefetching are described. In an embodiment, an apparatus includes an instruction decoder, load circuitry, and execution circuitry. The instruction decoder is to decode a code prefetch instruction. The code prefetch instruction is to specify a first instruction to be prefetched. The load circuitry to prefetch the first instruction in response to the decoded code prefetch instruction. The execution circuitry is to execute the first instruction at a fetch stage of a pipeline.
Owner:INTEL CORP

Data prefetching method and data prefetching device

A data prefetch method and a data prefetch device are provided. The data prefetch method includes: receiving a first memory access instruction, the address corresponding to the first memory access instruction belongs to a first storage area of ​​the memory; determining prefetch data according to a historical memory access mode of the first storage area, the historical memory access mode is used to indicate the historical access data of the first storage area, and the prefetch data includes part or all of the historical access data; according to the prefetch data, sending a prefetch instruction to the memory controller, the prefetch instruction is used to store the prefetch data in a prefetch cache area. The embodiment of the present application stores part or all of the historical access data of the area to which the address corresponding to the memory access instruction belongs in advance in the cache area, so that the memory access instruction to be executed can obtain the data to be used from the cache area. This method can make the prefetched data cover the data required for the instruction to be executed as much as possible with the least number of operations, thereby helping to reduce the power consumption of the memory.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

A sub-real-time processor, a real-time processor and a system-on-a-chip

ActiveCN117369870BControl signalControl cell
This invention discloses a sub-real-time processor, a real-time processor, and a system-on-a-chip (SoC). The sub-real-time processor includes: a control unit for generating corresponding control signals and sending them to other units; an instruction prefetching unit for acquiring instructions to be executed and sending them to an execution unit; an execution unit for acquiring a set interval time corresponding to the instruction to be executed and sending it to a comparison unit, and executing the instruction to be executed upon receiving a start execution signal; an interval time counting unit for counting pulses based on a time reference and acquiring the current interval time count value and sending it to the comparison unit; and a comparison unit for generating a start execution signal and sending it to the execution unit when it detects that the set interval time is equal to the current interval time count value. The technical solution of this embodiment, by employing an interval time counting unit, can achieve timed execution of instructions, thereby improving the timing accuracy and real-time performance of instruction execution.
Owner:MORNINGCORE HLDG CO LTD

Compiler-generated KILO-instructions deep runahead

Methods and apparatus for a runahead process are provided to prevent frontend stalls when executing a computer program. Methods and apparatus profile and analyze the computer program when it is compiled to extract meta-data defining hyperblocks for the computer program. The hyperblocks each encompass a respective series of basic blocks having transitions that meet a specified threshold. When the computer program is executed, the runahead process is performed for program branches. In this process, a future path of hyperblocks is predicted from each branch and the instructions corresponding to those hyperblocks are prefetched to a memory cache so that they can be readily fetched. Cycles or large call stacks are removed to enable deep runaheads, which may span about a thousand instructions.
Owner:HUAWEI TECH CO LTD

Instruction Address Translation and Instruction Prefetch Engine

Techniques are provided for performing an instruction fetch operation, including determining an instruction address of a primary branch prediction path, requesting a level 0 translation lookaside buffer (TLB) to cache an address translation of the primary branch prediction path, determining one or both of an alternate control flow path instruction address and a lookahead control flow path instruction address, and requesting the level 0 TLB or an alternate level TLB to cache address translations of one or both of the alternate control flow path instruction address and the lookahead control flow path instruction address.
Owner:ADVANCED MICRO DEVICES INC

Storage instruction processing method, electronic device and storage medium

The present invention provides a data storage instruction processing method, electronic device, and storage medium, and relates to the field of computer technology. The method includes sending the block address of a target cache block corresponding to a target cache line into a data storage instruction prefetch queue, wherein all target data storage instructions cached in the target cache line do not fill the target cache line, and the data storage instruction address of the target data storage instruction includes the block address of the target cache block; performing a data cache data storage address hit check on the target data storage instruction in the data storage instruction prefetch queue, and in the event that the data storage address fails, initiating a prefetch request and executing access to a lower-layer storage system based on the block address of the target cache block, obtaining data of the target cache block and backfilling it into the data cache, and writing the data storage instruction data included in the target data storage instruction for which the prefetch request was initiated and prefetched successfully into the target cache block in the data cache. The present invention can improve the prefetch accuracy and prefetch timeliness of data storage instructions.
Owner:BEIJING VCORE TECH CO LTD

Multi-thread processor pipeline architecture system, scheduling method and equipment

The invention relates to the technical field of processor design, discloses a multi-thread processor pipeline architecture system, a scheduling method and equipment, and designs an instruction prefetching scheduling algorithm based on mixed priorities, and the algorithm realizes load balancing of thread instruction queues through a static level decision mechanism and a dynamic level decision mechanism. The static level preferentially processes an empty queue thread to avoid starvation, and the dynamic level dynamically adjusts instruction fetching priority according to each queue depth, thread validity and a branch risk state to ensure that an instruction stream is continuously and stably supplied to a subsequent stage; a coupled polling emission scheduling algorithm is provided, so that high complexity of full-permutation search is avoided, and structural conflicts and data conflicts are effectively reduced; a special hardware architecture is designed around the algorithm, a multi-program counter and an instruction queue are integrated at the front end, a double-transmitting channel, a multifunctional unit and a thread private register file are configured at the rear end, and the performance of the single-core CPU is improved by optimizing the hardware architecture.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Cooperative instruction prefetch on multicore system

Aspects of the disclosure are directed to methods, systems, and apparatuses using an instruction prefetch pipeline architecture that provides good performance without the complexity of a full cache coherent solution deployed in conventional CPUs. The architecture can include components which can be used to construct an instruction prefetch pipeline, including instruction memory (TiMem), instruction buffer (iBuf), a prefetch unit, and an instruction router.
Owner:GOOGLE LLC

Prefetching method and system based on thread branch information in graphics processor

The present application provides a prefetching method and system based on thread branch information in a graphics processor (GPU), which is applied to the field of GPU instruction prefetching technology. The method comprises determining context information and a first index of a first instruction pointer (PC) of a branch stack at a first moment and a first top instruction pointer; establishing a miss table corresponding to the branch stack based on the context information and the second index, where the second index is the previous index of the first index; and prefetching missing instruction pointers within a time period corresponding to the first instruction pointer based on the miss table each time the GPU executes the second top instruction pointer of the branch stack. The missing instruction pointers are obtained based on the context information, effectively reducing instruction cache misses caused by locality issues between branch instructions in the GPU and improving the GPU's execution speed. The method relies on the program's global context information, making it more clear and reliable, and effectively reducing design complexity.
Owner:INTERNATIONAL INNOVATION CENTER OF TSINGHUA UNIVERSITY SHANGHAI +1

FDIP-based novel instruction prefetching method and apparatus, and electronic device

The invention provides a novel instruction prefetching method and device based on FDIP and electronic device.The method comprises the steps that after a target instruction block address enters a target instruction fetching queue, the target instruction block address is matched with a label array of instruction cache for the first time; when it is determined that an instruction block corresponding to the target instruction block address is in an instruction cache, locking a target cache line corresponding to the target instruction block address, and unlocking the target cache line after an instruction fetching process of the target instruction block address is completed; when the target instruction block becomes the first entry of the target fetch queue, the fetch unit obtains the target instruction block from the instruction cache. According to the scheme, after the branch predictor generates the target instruction block address, the target instruction block address only needs to be matched with the tag array once, so that access conflicts and dynamic power consumption caused by two times of matching can be avoided, additional area overhead is avoided, the performance of a CPU can be improved, and the power consumption of the CPU can be reduced.
Owner:CHENGDU QUNXIN MICROELECTRONICS TECHNOLOGY CO LTD

Instruction prefetching method and device and artificial intelligence chip

The invention provides an instruction prefetching method and device and an artificial intelligence chip, and relates to the technical field of chip design and manufacturing, and the method comprises the steps: obtaining a thread bundle identifier of a current thread bundle; determining the current thread bundle as a memory access thread bundle based on the thread bundle identifier; and adjusting the instruction prefetch number of the current thread bundle. According to the method and the device provided by the invention, the memory access thread bundle is prevented from being excessively prefetched, and the saved instruction cache space is reserved for the calculation thread bundle which needs the instruction cache space, so that the instruction to be executed by the processor is efficiently prefetched, the overall utilization efficiency and the hit rate of instruction cache are improved, and the operation efficiency of the processor is improved. Unnecessary memory bandwidth consumption is reduced, and finally the overall processing performance of the processor is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Microprocessor with instruction prefetching

A microprocessor efficiently performs instruction prefetching, having an instruction cache, a branch predictor, a fetch target queue coupled between the branch predictor and the instruction cache, and a comparator. The instruction cache caches contents for fetching in response to a fetch address. The fetch target queue stores at least one instruction address predicted by the branch predictor to be in a branch direction, to be read out as the fetch address, or output as a plurality of prefetch candidate addresses for selection as a prefetch address to operate the instruction cache for instruction prefetching. The comparator compares the plurality of prefetch candidate addresses with a previous prefetch address to determine which prefetch candidate address is the prefetch address, to exclude instruction prefetching of consecutive instructions of the same cache line.
Owner:VIA ALLIANCE SEMICON CO LTD

Instruction prefetching system based on compensation prefetching

The invention belongs to the technical field of computers, particularly relates to an instruction prefetching system based on compensation prefetching, and aims at solving the problem that the hit rate of an instruction prefetching technology is low. The method comprises the following steps: writing a value address hit by a target address back to a cache through a prefetch storage module, so that the cache sends the value address to a value queue module, and stores a compensation address and a sequence address from a lower-level cache; whether compensation prefetching is carried out or not is judged through a value queue module on the basis of the number of repetitions of the value address, and a control signal used for starting compensation prefetching is generated and sent to a compensation prefetching module; decoding the jump instruction in the target address through a data processing module, and sending decoding information to a compensation prefetching module; and generating a compensation address and sending the compensation address to a lower-level cache through a compensation prefetching module based on the control signal and the decoding information. According to the method and the device, the hit rate of sequential prefetching to the jump instruction is improved, and the cache prefetching efficiency is improved.
Owner:SHANDONG UNIV +1

Microprocessor with instruction prefetching

A microprocessor efficiently performs instruction prefetching with an instruction cache, a branch predictor, a fetch target queue coupled between the branch predictor and the instruction cache, and a prefetch read pointer control circuit. The instruction cache caches content for fetching in response to a fetch address. The fetch target queue stores instruction addresses predicted by the branch predictor to be taken in a branch direction to be fetched as the fetch address or selected as a prefetch address to operate the instruction cache to perform instruction prefetching. The prefetch read pointer control circuit generates a prefetch read pointer to the fetch target queue such that the instruction prefetching performed by the fetch target queue in response to the prefetch address supplied by the prefetch read pointer does not fall behind the fetching performed in response to the fetch address supplied by a fetch read pointer.
Owner:VIA ALLIANCE SEMICON CO LTD

Instruction prefetching method, instruction prefetching device, processor and electronic equipment

An instruction prefetch method, an instruction prefetch device, a processor, and an electronic device. The instruction prefetch method includes: in response to a target instruction missing a target cache, writing a target access request for the target instruction into a loss status processing queue, the loss status processing queue including multiple access requests, the target access request being one of the multiple access requests, the loss status processing queue being configured to sequentially send the multiple access requests to a next-level cache of the target cache; in response to a target instruction prediction error, sending a cancel request for the target instruction to the loss status processing queue; and in response to the cancel request, releasing the queue space occupied by the access request following the target access request in the loss status processing queue. This instruction prefetch method can improve prefetch accuracy, increase the utilization of the loss status processing queue, and contribute to improving overall performance.
Owner:HYGON INFORMATION TECH CO LTD

Using retired pages history for instruction translation lookaside buffer (TLB) prefetching in processor-based devices

Using retired pages history for instruction translation lookaside buffer (TLB) prefetching in processor-based devices is disclosed herein. In some exemplary aspects, a processor-based device is provided. The processor-based device comprises a history-based instruction TLB prefetcher (HTP) circuit configured to determine that a first instruction of a first page has been retired. The HTP circuit is further configured to determine a first page virtual address (VA) of the first page. The HTP circuit is also configured to determine that the first page VA differs from a value of a last retired page VA indicator of the HTP circuit. The HTP circuit is additionally configured to, responsive to determining that the first page VA differs from the value of the last retired page VA indicator of the HTP circuit, store the first page VA as the value of the last retired page VA indicator.
Owner:QUALCOMM INC

Network processor and chip

The invention relates to a network processor and a chip, and the method comprises the steps: an instruction prefetching module generates a prefetching instruction request according to a prefetching instruction address and a thread identifier of a target thread, and transmits the prefetching instruction request to an instruction memory; the instruction preprocessing module obtains a target instruction block returned by the instruction memory, stores a first special instruction in the target instruction block and a plurality of previous instructions in the instruction cache module, obtains a next prefetch instruction address according to the first special instruction, and if prediction succeeds, sends the next prefetch instruction address to the instruction cache module; if not, sending the predicted next prefetch instruction address, the thread state information and the corresponding thread identifier to an instruction prefetch module; the instruction pre-fetching module generates a new pre-fetching instruction request according to the next pre-fetching instruction address, the thread pre-fetching control signal and the corresponding thread identifier output by the instruction preprocessing module, and sends the new pre-fetching instruction request to the instruction memory, through the application, the instruction pre-fetching time delay of the processor can be reduced.
Owner:CHENGDU JAGUAR MICROSYSTEMS CO LTD

Instruction prefetch throttling

An apparatus is provided for limiting the effective utilisation of an instruction fetch queue. The instruction fetch entries are used to control the prefetching of instructions from memory, such that those instructions are stored in an instruction cache prior to being required by execution circuitry while executing a program. By limiting the effective utilisation of the instruction fetch queue, fewer instructions will be prefetched and fewer instructions will be allocated to the instruction cache, thus causing fewer evictions from the instruction cache. In the event that the instruction fetch entries are for instructions that are unnecessary to the program, the pollution of the instruction cache with these unnecessary instructions can be mitigated.
Owner:ARM LTD

Cost-driven prefetching

Systems and techniques for cost-based prefetching are described. In one example, a processor includes a cache system having a hierarchy of one or more cache levels and an instruction prefetcher associated with a cache level of the cache system. The instruction prefetcher encodes fetch requests to a prefetcher table and determines a cost associated with each entry based on a memory level from where the instructions associated with the fetch requests were fetched. The instruction prefetcher evicts entries in the prefetcher table based on the cost associated with one or more corresponding fetch requests in response to adding new entries. The described techniques hide instruction fetch latency that a processor frontend experiences with workloads that have code working sets that do not fit in higher levels of the cache system.
Owner:ADVANCED MICRO DEVICES INC