Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

107 results about "Cache access" patented technology

Cache architecture and method of AXI interconnection module, electronic equipment and storage medium

The invention provides a cache architecture and method of an AXI interconnection module, electronic equipment and a storage medium, and relates to the technical field of storage, the cache architecture comprises an index cache region, a shared cache region and a shared cache access controller, the index cache region comprises an index entry list, and the shared cache region comprises an instruction cache region and a data cache region; each index cache region corresponds to one main device, the shared cache region corresponds to all the main devices, and the index entry list is used for recording index information of access transactions of the corresponding main devices; each index entry in the index entry list corresponds to one access transaction; and the shared cache access controller is used for storing instruction information corresponding to the access transaction in the instruction cache region based on the index information and storing data information corresponding to the access transaction in the data cache region. By applying the scheme of the invention, the utilization efficiency of on-chip resources can be improved on the basis of ensuring the transaction processing performance.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Caching method and system for access unit of superscalar processor

The invention belongs to the field of integrated circuits and computer system structures, and provides a caching method and system for a memory access unit of a superscalar processor, and the method comprises the steps: receiving a plurality of memory access instructions in the same period, and determining a corresponding Bank in a to-be-accessed cache through the memory access instructions; after the memory access instruction obtains a cache access permission, if cache line missing occurs, generating a missing request, merging all the missing requests, and performing parallel prefetching training on the merged missing requests by utilizing a mode of fusing a constant step length prefetching mode and a complex step length prefetching mode to obtain a prefetching request and a prefetching cache address corresponding to the prefetching request; requesting a missing cache line from the first-level cache to the second-level cache based on the missing queue, and writing the missing cache line back to the cache line of the corresponding data cache in the first-level cache; and storing the bus consistency request by using the sniffing queue, judging whether the data in the multi-core cache are consistent or not by using the consistency request, and performing consistency modification according to a judgment result. The cache hit rate and the bandwidth utilization rate are improved.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Database access parameter real-time cooperative processing method and device based on multi-level cache

The invention discloses a database access parameter real-time cooperative processing method and device based on multi-level cache, and belongs to the technical field of computer data caching. The method comprises the steps that a version number management table is configured; establishing a four-level cache architecture of a transaction-level cache layer, a process-level cache layer, a distributed cache layer and a database cache layer; receiving a transaction request, creating a transaction level cache layer and loading current version number information; comparing the version number in the transaction level cache layer with the version number in the process level cache layer; according to the version number comparison result, obtaining parameter data from the corresponding cache level according to a preset cache access priority strategy; and when the parameter change is detected, updating the corresponding version number in the version number management table, and updating the cache layer data as required. According to the method, real-time global effectiveness of parameter change is realized, parameter consistency of in-transit transaction is guaranteed, system processing performance and throughput are improved, system reliability and fault-tolerant capability are enhanced, and system resource utilization rate is optimized.
Owner:SHANDONG CITY COMMERCIAL BANK COOP ALLIANCE CO LTD

Data transmission method and device for interconnection between core particles, core particles and processing system

The invention relates to the technical field of integrated circuits, and discloses a data transmission method and device for interconnection between core particles, the core particles and a processing system.The method is applied to a first core particle and comprises the steps that a protocol layer of the first core particle receives to-be-transmitted data sent by a processor of the first core particle; a protocol layer carries out protocol layer overall packing on data to be transmitted to obtain a protocol layer packet, the protocol layer packet comprises a packet header field and a load area, the packet header field is used for identifying a channel type and a channel signal field of the load area, the channel type comprises at least one monitoring channel used for monitoring cache consistency between core particles, and the monitoring channel is used for monitoring cache consistency between the core particles. The packet header field is used for identifying a channel type and a channel signal field of the load area; and the protocol layer sends the to-be-transmitted data to the second core grain through the link layer and the physical layer of the first core grain based on the protocol layer packet. The problem of cache access consistency of processors among core particles is effectively solved.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

Transaction and Request Buffers for Cache Miss Handling

Methods and systems for cache miss monitoring and fulfillment are disclosed. A disclosed method comprises receiving, from a processing unit, a cache access request, determining, based on a cache failing to fulfill the cache access request, that a cache miss has occurred, populating a request buffer based on the cache miss and the cache access request, populating a transaction buffer with a transaction entry for the request buffer based on the cache miss request and the cache access request, determining by the request buffer, that information requested in the cache access request should be retrieved from main memory, determining, by the transaction buffer if information was not retrieved, that the cache access request satisfies criteria for creating a cache miss tag, creating the cache miss tag based on the cache access request, and storing the cache miss tag in a portion of the cache.
Owner:TENSTORRENT USA INC

Data cache access method and device of multi-core processor

The invention provides a data cache access method and device of a multi-core processor, and relates to the technical field of computer processors. The method comprises the following steps: accessing a data cache according to an access instruction, and responding to the access instruction to hit access data and write back the access data, or responding to the access instruction to not hit the access data, fetching back-filling data from a memory access failure queue and forwarding the back-filling data to the access queue. Accessing the data cache according to the data storage instruction, writing the stored data into the data cache in response to the access hit of the data cache, or calling a stored data retransmission unit in response to the access miss of the data cache; and sending a data storage failure request to the memory access failure queue, and taking back the backfill data and / or authority by the memory access failure queue. According to the method, the flexibility of the processing mode of the data cache access pipeline on the memory access instruction and the external consistency request can be improved, the access parallelism degree is improved, and the hardware resource utilization rate of the processor is increased.
Owner:BEIJING VCORE TECH CO LTD

Method and system for optimizing direct I / O read performance under Linux system

The invention discloses an optimization method and system for direct I / O read performance under a Linux system. The method comprises the steps that a cache control mark used for controlling cache access is introduced; in the file opening stage, whether a cache mechanism is started or not is judged, if yes, whether a cache is hit or not is detected when direct I / O reading operation is executed, data are directly read from the cache and returned when the cache is hit, standard direct I / O reading and data returning are called when the cache is not hit, meanwhile, asynchronous caching is conducted on the data, and a cache mark is set; when the direct I / O write operation is executed, whether a cache mechanism is started or not is judged, if yes, whether the cache is hit or not is detected firstly, when the cache is hit, original cache data is set to be invalid, asynchronous caching is conducted on the data, meanwhile, standard direct I / O write-in data is called, and after asynchronous caching and data write-in operation are completed, the cache data is set to be valid. According to the invention, the reading performance of direct I / O can be improved.
Owner:KYLIN CORP

Multi-level cache access method, system and equipment and storage medium

The embodiment of the invention provides a multi-level cache access method, system and device and a storage medium, and relates to the technical field of data cache access, and the method comprises the following steps: obtaining a key field in a user auditing request, and converting the key field into a feature vector; searching whether a reasoning result corresponding to the feature vector exists in a first preset cache or not; if yes, a reasoning result is returned to the user; if not, a corresponding semantic partition is found in a second preset cache, and a historical cache vector with the highest similarity and a corresponding reasoning result are determined from the semantic partition; and determining a reasoning result corresponding to the feature vector according to the feature vector, a historical cache vector and a reasoning result corresponding to the historical cache vector. In this way, a two-stage cache mechanism based on the feature vector is constructed, a corresponding reasoning result is found in combination with semantic partition, the cache hit rate is increased, repeated reasoning of a large model is reduced, response delay is reduced, and therefore the requirements for high redundancy and high real-time performance of cache data reasoning are met.
Owner:SHANGHAI XULU INFORMATION TECHNOLOGY CO LTD

System memory peak bandwidth measurement method and electronic equipment

The invention discloses a system memory peak bandwidth measurement method and electronic equipment, and relates to the technical field of bandwidth measurement, the memory bandwidth measurement process is optimized through cooperation of dynamic selection of a maximum width vector instruction, forced non-cache access and NUMA perception binding, and the bandwidth measurement efficiency is improved by utilizing the vector processing capacity of a CPU (Central Processing Unit). The data throughput of a single operation is improved to 512 bits or even higher, through cooperation of a non-temporary instruction and a memory barrier instruction, a cache level is bypassed, interference of cache hit or jitter on a measurement result is eliminated, it is ensured that the measurement result truly reflects the performance of a memory system, and the measurement accuracy is improved. Through an automatic NUMA binding mechanism, delay and congestion caused by cross-node access are avoided, and the accuracy, stability and repeatability of a test result are improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Resource-constrained persistent memory-oriented server, virtual machine monitor construction method and system and electronic equipment

The invention discloses a virtual machine monitor system for a resource-limited persistent memory server, which comprises a configuration file module, a CPU (Central Processing Unit) management module, a memory management module, a cache management module, an interrupt management module, a cross-partition communication module and a virtual machine management module. The CPU management module is used for binding a virtual CPU to a physical CPU core, so that computing resources of the virtual machines are dedicated, and meanwhile, the virtual machines are prevented from competing for the CPU; the memory management module is used for dividing a physical memory region according to a configuration file and realizing address space isolation between virtual machines by utilizing two-stage address translation; the cache management module is used for performing isolation control on cache access of a plurality of virtual machines through a coloring mechanism of the last-level cache; and the interrupt management module is used for simulating the universal interrupt controller, and the virtual machine monitor re-injects the interrupt into the target virtual machine according to the configuration file after capturing the hardware interrupt.
Owner:SHANGHAI JIAOTONG UNIV

A method for intelligent resource allocation for cognitive subzone energy

This invention discloses an intelligent resource allocation method for energy-aware partitions, belonging to the field of computer architecture and operating system task scheduling technology. The method divides computing resources into multiple independent energy-aware partitions based on the processor cores. It then monitors the energy status parameters of each partition in real time. By parsing the instruction code stream of the tasks to be assigned, the proportion of arithmetic and logical instructions is extracted as instruction mixing features, and the address span of memory access instructions is analyzed to obtain cache access density features. Based on the energy status of the partitions and the computation and memory access characteristics of the tasks, a suitability score for the task relative to each partition is generated, comprehensively reflecting the energy matching degree and hardware resource matching degree. Based on the suitability score and the remaining task acceptance capacity of the partitions, a two-dimensional assignment strategy is used to dynamically map tasks to the core execution in the optimal energy-aware partition. This invention achieves fine-grained matching of computing resources and energy status, optimizing system energy utilization efficiency.
Owner:SHENZHEN YIXING MEDICAL BEAUTY HOSPITAL

A method for dynamic voltage and frequency regulation of a RISC-V processor core

This invention relates to a dynamic voltage and frequency adjustment method for a RISC-V processor core. The method includes: monitoring and recording the critical path time of memory access requests for each missing state processing register to update the global critical path counter; monitoring prefetch instructions and extracting their delay parameters; collecting cache access events based on a time window to calculate bandwidth utilization and dynamically determining a bandwidth threshold according to preset performance parameters; and correcting the global critical path counter value by combining the prefetch delay and the bandwidth threshold, thereby triggering a dynamic voltage and frequency adjustment instruction. This method achieves fine-grained adaptive DVFS for the RISC-V processor core under complex loads, significantly improving energy efficiency and performance stability.
Owner:CHAORUI TECH (CHANGSHA) CO LTD

Dynamic sharding method and device for cache, storage medium and electronic equipment

The application relates to a dynamic cache sharding method and device, a storage medium and an electronic device. The method comprises the following steps: when it is detected that a cache system adds or removes a cache shard for capacity expansion or contraction, a cache SDK running in an application program continues to perform a cache access operation on cache data according to an original cache shard rule; a cache management component independently set performs a rehashing calculation based on a target cache shard rule after the capacity expansion or contraction, so as to perform asynchronous migration of the cache data to the target cache shard after the capacity expansion or contraction; and after the asynchronous migration is completed, the cache management component sends a shard rule switching instruction to the cache SDK, so that the cache SDK switches to the target cache shard rule to perform the cache access operation. The application solves the technical problems that cache shards are static and are prone to cause cache break-in and service fluctuation during capacity expansion or contraction.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Cache storage device and operating method thereof, system including cache storage device

A cache storage device is disclosed, including a cache circuit and a path prediction circuit. The cache circuit generates a cache hit signal indicating whether target data corresponding to an accessed address is stored in a cache line, and in path prediction mode, performs the current cache access operation primarily on candidate paths based on candidate path signals indicating candidate paths. The path prediction circuit stores accumulated information based on cache hit signals provided during previous cache access operations, by accumulating cache hit results indicating whether target data is stored in a path and path prediction hit results indicating whether target data is stored in a candidate path. In path prediction mode, the path prediction circuit generates candidate path signals based on the accumulated information by determining candidate paths for the current cache access operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Last level cache access during non-cstate self refresh

A data processor includes a data fabric, a memory controller, a last level cache, and a traffic monitor. The data fabric is for routing requests between a plurality of requestors and a plurality of responders. The memory controller is for accessing a volatile memory. The last level cache is coupled between the memory controller and the data fabric. The traffic monitor is coupled to the last level cache and operable to monitor traffic between the last level cache and the memory controller, and based on detecting an idle condition in the monitored traffic, to cause the memory controller to command the volatile memory to enter self-refresh mode while the last level cache maintains an operational power state and responds to cache hits over the data fabric.
Owner:ADVANCED MICRO DEVICES INC

Hot and cold page statistics method and apparatus

The embodiment of the present specification provides a cold and hot page statistics method and device, which is applied to a cold and hot page identification module of a memory controller, wherein the cold and hot page statistics method comprises: in response to a memory access instruction sent by a processor, generating an access record for an initial memory page according to the memory access instruction, updating the access times of the initial memory page in a data list according to the access record, and determining a cold memory page or a hot memory page according to the access times of the initial memory page in the data list. By obtaining the memory access instruction sent by the processor, the real memory access can be detected, the cache access is avoided from being regarded as the memory access to perform false statistics, and since each memory access can be recorded by the access record, the probabilistic error in the sampling statistics is avoided, and the accuracy of the cold and hot page identification is improved.
Owner:ALIBABA (CHINA) CO LTD

Cache access fabric

Examples described herein relate to a cache fabric that includes a set of routers of a first tier and a plurality of cache controller clusters of a second tier. A router of the set of routers is accessible via an interface to receive a memory access request from a processor and select from a set of cache controllers based on a cluster identifier and a memory address, and provide the memory access request to the selected set of cache controllers. The selected set of cache controllers may receive memory access requests and serve memory access requests from the cache device, or forward memory access requests to a second cache controller or second cache device associated with the cache device.
Owner:INTEL CORP

System for context-aware distributed processing and memory optimization in multi-core computer architectures

A system for context-sensitive distributed processing and memory optimization in a multi-core computer architecture, wherein the system comprises: a plurality of processor cores arranged within a processor structure and configured to perform computational tasks in parallel; a linking network that operationally connects the plurality of processor cores for data communication; a context capture unit connected to each of the plurality of processing cores and configured to monitor execution parameters such as instruction throughput, cache access behavior, memory latency, and data exchange characteristics between the cores, wherein the context capture unit generates context descriptors that are representative of the runtime execution conditions;a scheduling processor operationally connected to the context capture unit and configured to distribute computational tasks across the multitude of processor cores based on the generated context descriptors; and a memory management unit operationally connected to a hierarchical memory arrangement comprising multiple cache levels and main memory, the memory management unit being configured to dynamically allocate and migrate data segments across the hierarchical memory arrangement according to the context descriptors, thus achieving alignment between task execution and data locality.
Owner:EASWARI ENGINEERING COLLEGE TAMIL NADU +3

A cache access method, a DRAM cache, and a cache management system

The application provides a cache access method, a DRAM cache, a cache management system and a computing device. The cache access method comprises receiving a memory access request, the request specifying a data size to be accessed; determining a corresponding granularity cache area in the DRAM cache according to the data size; and the granularity cache area is obtained by dividing the DRAM cache area according to different granularity sizes. Thus, the corresponding granularity cache area is selected according to the data size of the request, the appropriate cache block can be positioned more accurately, the opportunity of DRAM cache hit is increased, and the memory access delay is reduced. The proportion of useful data in the DRAM cache is improved, and the cache space utilization is improved.
Owner:HUAWEI TECH CO LTD

Boot method, device and electronic equipment of basic input output system (BIOS)

The embodiment of the present application provides a BIOS starting method, device and electronic equipment, in which, in response to a system starting instruction, an execution area of a BIOS moving code file is determined according to a storage structure type of a CPU; the storage structure type is used to represent whether the CPU supports cache access to a bus space of a bus. The BIOS moving code file is executed in the execution area, so as to move a BIOS starting file stored in a memory to the memory of the CPU through the bus. The BIOS of the electronic equipment is started based on the BIOS starting file in the memory of the CPU. In the embodiment of the present application, the BIOS starting file in the memory is moved to the memory of the CPU based on the bus between the CPU and the memory and the BIOS moving code file, and the BIOS of the electronic equipment is started in the memory of the CPU. In this way, the method for starting the BIOS based on the bus is realized, so that the BIOS is not limited to being started only in a stored position, and the limitation of the starting mode can be reduced to a certain extent.
Owner:LOONGSON ZHONGKE (XIAN) TECH CO LTD

Servicing file restorations in a deduplication filesystem using multiple read-ahead caches

Access object (AOB) and deduplication object (DOB) services of a deduplication filesystem are provisioned across a cluster. A client-side library receives a request to restore a file, the file being divided into chunks and the chunks being assigned to similarity groups. A first table is created that maps offset ranges in the file to AOBs. Prefetch requests are issued to the AOBs for chunks of the file corresponding to the offset ranges. Upon the AOBs receiving prefetches, a second table is consulted. The second table maps similarity groups to the DOBs, each DOB being responsible for reading a chunk of an assigned similarity group from a storage layer of the filesystem. Multiple internal read-ahead streams are opened from the AOBs to the DOBs. The internal read-ahead streams prefetch the chunks read by the DOBs to populate read-ahead caches maintained at the AOBs. The request is serviced using the read-ahead caches.
Owner:DELL PROD LP

Cache management method and device, equipment, storage medium and product

PendingCN122045093ARealize negative feedback closed-loop controlAvoid access latency issuesResource allocationMemory systemsCache accessLoop control
The invention discloses a cache management method and device, equipment, a storage medium and a product, and the method comprises the steps: responding to a cache access, and obtaining a hit condition of the cache access; adjusting global leapfrogging parameters according to the hit condition; and according to the hit condition and the global leapfrogging parameter, adjusting the item level of the data item. According to the embodiment of the invention, the adjustment step length of the item level of the data item is reduced during hit, it is ensured that the hot data partition stores the high-frequency accessed data item, the occasionally accessed data item is prevented from mistakenly entering the hot area, the hot data partition entering speed of the sudden hot data item can be reduced, and the hot data partition entering efficiency is improved. The concerned data items are prevented from being quickly deleted; when the new inserted data is not hit, a large entry level is configured for the new inserted data, the new inserted data is placed in a cold data area, storage of high-frequency access data is prevented from being affected, negative feedback closed-loop control over a cache area is achieved, cold and hot data are effectively distinguished, the problem of access delay caused by the fact that the cold and hot data are difficult to distinguish is solved, and the access efficiency is improved. And the response speed and the cache stability are dynamically balanced.
Owner:深圳开鸿数字产业发展有限公司

Cache access method, cache, chip, storage medium and program product

The invention relates to a cache access method and device, a chip, equipment, a storage medium and a program product. The method comprises the following steps: receiving a cached first access request, wherein the first access request comprises a first access address; according to the first access address, tag information stored in a tag memory is inquired, each piece of tag information comprises address information and intermediate state information, and the intermediate state information is used for representing the state of a corresponding cache line in a data memory corresponding to the address information in the tag information; and under the condition that address information in mark information stored in the mark memory fails to be matched with the first access address, determining a target cache line according to intermediate state information in the mark memory, processing the first access request, and updating the intermediate state information of the target cache line, the target cache line is used for caching target data of the first access request. By adopting the method, the cache access efficiency can be improved.
Owner:SHANGHAI JAGUAR MICROSYSTEMS CO LTD +1

Instruction processing method and apparatus, processor, electronic device, and storage medium

Embodiments of this disclosure provide an instruction processing method and apparatus, a processor, an electronic device, and a storage medium. The instruction processing method includes, in response to identifying a first conditional branch instruction with a back jump from an instruction stream, recording instruction information of the first conditional branch instruction; in response to identifying the first conditional branch instruction with a back jump at least once more, determining that the first conditional branch instruction is an end branch instruction in a loop body instruction and performing a backfilling operation on the loop body instruction, the loop body instruction including at least one second branch instruction other than the first conditional branch instruction; performing the backfilling operation on the loop body instruction includes: backfilling the instruction information of the second branch instruction into an instruction information cache; invoking a branch predictor to obtain first prediction information of the second branch instruction, and backfilling the first prediction information into a branch instruction information cache. This instruction processing method expands the scope of loop body instruction recognition and reduces cache access power consumption.
Owner:BEIJING ESWIN COMPUTING TECH CO LTD

A cache access method, system, medium and product

The application discloses a cache access method and system, a medium and a product, and applies to the technical field of processors, and comprises the following steps: monitoring the memory access behavior of each processor core to a shared cache, each cache line in the shared cache is divided into a preset number of data subsegments, for each processor core, a unique corresponding state identifier is arranged for each data subsegment; after any processor core performs data writing on a target data subsegment of a target cache line in the shared cache, the target data subsegment is written back to a memory, and is loaded from the memory to a private cache of the any processor core, the state identifier of the target data subsegment of the any processor core is determined as a shared state, and the state identifier of the target data subsegment of a first processor core is determined as an invalid state. In this way, unnecessary memory data loading can be reduced, and the overall performance of a multi-core processor is improved.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Tagged-data prediction

An apparatus has cache data storage, and tagged-data prediction circuitry to generate a tagged-data prediction in response to a streaming-write request requesting that write data corresponding to a target address which missed in a previous level of cache is written to the cache data storage for the given level of cache without being allocated into the previous level of cache. The tagged-data prediction is indicative of whether a target cache data entry corresponding to the target address of the streaming-write request is predicted to be a tagged cache data entry that stores cached data associated with the target address and a valid memory safety check tag corresponding to the target address, or an untagged cache data entry that stores the cached data but does not store a valid memory safety check tag. Cache access scheduling circuitry selects, based on the tagged-data prediction generated by the tagged-data prediction circuitry for the streaming-write request, how to schedule access to the cache data storage in response to the streaming-write request.
Owner:ARM LTD

Cache access control method and device

The invention relates to the technical field of computers, and provides a cache access control method and device, and the method comprises the steps: determining hit information during cache access; and when the hit information is that the cache access competition fails and the tag is not hit, directly sending a cache access request of the target data to the next-level cache of the current cache. According to the embodiment of the invention, when the cache access competition fails and the tag is missed, the tag unit is directly accessed, the cache miss condition can be found in time, so that the request is sent to the next-level cache in time, and the pre-judgment processing remarkably shortens the response time of the miss condition.
Owner:CHENGDU QUNXIN MICROELECTRONICS TECHNOLOGY CO LTD

Virtualized caches

Systems and methods are disclosed for virtualized caches. For example, an integrated circuit (e.g., a processor) for executing instructions includes a virtually indexed physically tagged first-level (L1) cache configured to output to an outer memory system one or more bits of a virtual index of a cache access as one or more bits of a requestor identifier. For example, the L1 cache may be configured to operate as multiple logical L1 caches with a cache way of a size less than or equal to a virtual memory page size. For example, the integrated circuit may include an L2 cache of the outer memory system that is configured to receive the requestor identifier and implement a cache coherency protocol to disambiguate an L1 synonym occurring in multiple portions of the virtually indexed physically tagged L1 cache associated with different requestor identifier values.
Owner:SIFIVE INC

A cache access system supporting out-of-order processor data prefetching

The application belongs to the technical field of integrated circuit design, and particularly relates to a cache access system supporting out-of-order processor data prefetching. The system specifically comprises a LOAD memory access information tracking and sequencing module, a LOAD memory access address history buffer, a prefetcher and a target prefetch address buffer. The LOAD memory access information tracking and sequencing module changes out-of-order LOAD memory access information into in-order LOAD memory access information, which is then input into the prefetcher; the prefetcher uses the in-order memory access information to realize more accurate training and target prefetch address prediction; valid target prefetch addresses output by the prefetcher are stored in the target prefetch address buffer to wait for subsequent sending; and the target prefetch address buffer is updated in real time to invalidate untimely addresses, so as to avoid sending useless prefetch addresses. The application can improve the learning efficiency of memory access rules and the accuracy of address prediction, and reduce the resource occupation of prefetch requests on the cache system.
Owner:FUDAN UNIVERSITY

Cache line reconstruction-based gpu l2 partition parallel memory access method and device

The application discloses a GPU L2 partition parallel memory access method and device based on cache line reconstruction, and the method comprises the following steps: setting the L2 cache cache line size to 32 bytes, so that the 128-byte cache line data block is mapped to four different sub-partitions, and the mapped sub-partitions are logically grouped; after each L2 cache sub-partition receives a corresponding sub-request, it is judged whether the sub-request is hit in the local cache; if yes, the corresponding data segment is read, and is returned to the SM initiating the request through the data port and NoC; if the sub-request is not hit in the L2 cache, the sub-request is forwarded to the designated merging sub-partition, the other sub-requests belonging to the same original request are merged in the MSHR of the merging sub-partition, the memory access request of the complete cache line is formed, and the data acquisition is sent to the DRAM. The device comprises a processor and a memory. The application improves the utilization rate of the on-chip data path of the GPU, improves the cache access efficiency and the data transmission parallelism, thereby reduces the memory access delay and improves the overall performance of the GPU.
Owner:CIVIL AVIATION UNIV OF CHINA