Data prefetching method and device, electronic equipment and storage medium

By storing the memory flow information separately in the internal and external parts of the chip, and optimizing the prefetching strategy using deep learning models, the problems of insufficient prediction accuracy and large storage overhead in large-scale data sets and complex application scenarios are solved, and efficient and flexible data prefetching is achieved.

CN120353728APending Publication Date: 2025-07-22SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510208864.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When traditional data prefetching technology deals with complex data access patterns and changing data characteristics, the prediction accuracy is insufficient, poor adaptability is poor, and the storage overhead is too large, so it cannot be effectively applied to large-scale data sets and complex application scenarios.

Method used

The deep learning model is used to combine the table structure of the internal and external parts of the chip, and the prefetch address is determined using the low-latency table of the internal parts of the chip, and the tables of the external parts of the chip are stored in combination with the tables of the external parts of the chip, and the results of the deep learning model are stored on-chip and the prefetch strategy is optimized to reduce memory access delay and power consumption.

Benefits of technology

It improves the accuracy and flexibility of data prefetching, reduces memory access delay and power consumption, adapts to the needs of large-scale data sets and complex application scenarios, and optimizes system performance and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353728A_ABST
    Figure CN120353728A_ABST
Patent Text Reader

Abstract

The invention provides a data prefetching method and device, electronic equipment and a storage medium, the method is applied to the electronic equipment, the electronic equipment comprises a chip internal part and a chip external part, and the method comprises the steps that a memory access request message is received, and the memory access request message comprises a memory access address and a process identifier; querying a first cache of the internal part of the chip according to the memory access request message; and if the cache of the memory access address in the first cache is missing, determining a prefetching address according to the memory access address and a process identifier, and executing a data prefetching operation according to the prefetching address. Wherein the step of determining the prefetch address according to the memory access address and the process identifier comprises the following steps: according to the memory access address and the process identifier, querying a table based on the internal part of the chip and / or a table based on the external part of the chip, and determining the prefetch address.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of chip technology, and in particular, to a data prefetching method, apparatus, electronic device, and storage medium. Background Art

[0002] Data prefetching technology refers to predicting future data requirements by learning historical data access patterns and loading data from slow storage devices to fast storage devices in advance, thereby reducing data access latency and improving system performance. It is regarded as one of the effective ways to alleviate the "memory wall" problem of multi-core processors and is currently widely used in high-performance mainstream processors.

[0003] Traditional data prefetching technology usually makes predictions based on fixed algorithms and rules, and it is difficult to handle complex data access patterns and changing data characteristics. On the one hand, with the explosive growth of data volume and the complexity of data access patterns, the prediction accuracy of traditional methods is limited; on the other hand, due to the lack of intelligence and learning ability in traditional technologies, their adaptability is relatively poor. When the data access pattern changes, traditional technologies may not be able to adjust the prefetching strategy in time, resulting in poor prefetching effects; furthermore, the strategies and parameters of traditional technologies are usually fixed or have a limited adjustable range, which limits their flexibility and scalability in practical applications. In large-scale data sets and complex application scenarios, the limitations of traditional technologies are more obvious. There is a need to provide an efficient data prefetching method. Summary of the Invention

[0004] The present disclosure provides a data prefetching method, apparatus, electronic device, and storage medium to at least solve the above technical problems existing in the prior art.

[0005] In a first aspect, an embodiment of the present disclosure provides a data prefetching method, which is applied to an electronic device, and the electronic device includes: an internal part of the chip and an external part of the chip; the method includes:

[0006] Receiving a memory access request message, where the memory access request message includes: a memory access address and a process identifier;

[0007] Querying a first cache in the internal part of the chip according to the memory access request message; if there is a cache miss for the memory access address in the first cache, determining a prefetch address according to the memory access address and the process identifier, and performing a data prefetching operation according to the prefetch address;

[0008] Wherein, the determining the prefetch address according to the memory access address and the process identifier includes: querying a table based on the internal part of the chip and / or a table of the external part of the chip according to the memory access address and the process identifier to determine the prefetch address.

[0009] Second aspect, embodiments of the present disclosure provide a data prefetching apparatus, which is applied to an electronic device. The electronic device includes an internal part of the chip and an external part of the chip. The apparatus includes:

[0010] A receiving module, configured to receive a memory access request message, where the memory access request message includes a memory access address and a process identifier;

[0011] A processing module, configured to query a first cache in the internal part of the chip according to the memory access request message; if there is a cache miss for the memory access address in the first cache, determine a prefetch address according to the memory access address and the process identifier, and perform a data prefetch operation according to the prefetch address;

[0012] Wherein, determining the prefetch address according to the memory access address and the process identifier includes: querying a table based on the internal part of the chip and / or a table of the external part of the chip according to the memory access address and the process identifier to determine the prefetch address.

[0013] In the above solution, determining the prefetch address according to the memory access address and the process identifier includes:

[0014] Querying a first touch table according to the memory access address and the process identifier;

[0015] If the first touch table contains a flow label corresponding to the memory access address and the process identifier, updating a trigger counter corresponding to the flow label in the first touch table, and querying a first mode table according to the flow label to determine a memory access mode corresponding to the memory access address and the process identifier;

[0016] If the first touch table does not contain a flow label corresponding to the memory access address and the process identifier, querying a second touch table and a second mode table according to the flow label to determine a memory access mode corresponding to the memory access address and the process identifier;

[0017] Determining the prefetch address according to the memory access mode;

[0018] Wherein, the first touch table is used to record at least one of the following: flow label, process identifier, trigger address, flow status, trigger counter;

[0019] The first mode table is used to record at least one memory access mode, and each memory access mode includes: flow label, offset flow, time flow;

[0020] The first touch table has the same format as the second touch table; the first mode table has the same format as the second mode table;

[0021] The first touch table and the first mode table are tables stored in the internal part of the chip; the second touch table and the second mode table are tables stored in the external part of the chip.

[0022] In the above solution, determining the prefetch address according to the memory access mode includes:

[0023] Determining the offset value in the offset stream according to the memory access mode, and determining the prefetch address according to the base address and the offset value.

[0024] In the above solution, performing a data prefetch operation according to the prefetch address includes:

[0025] Performing a filtering operation according to the prefetch address, generating a prefetch request according to the prefetch address after the filtering operation; performing a data prefetch operation according to the prefetch request.

[0026] In the above solution, during the prefetch operation according to the prefetch request, the method further includes:

[0027] Tracking the execution process of the prefetch operation to obtain a tracking result;

[0028] Updating a deep learning model according to the tracking result, where the deep learning model is used to identify a memory access request message to obtain memory access mode features, and the memory access mode features at least include an offset stream and a time stream.

[0029] In the above solution, the method further includes:

[0030] Updating the first touch table and / or the first mode table according to the memory access request message.

[0031] In the above solution, updating the first touch table and / or the first mode table according to the memory access request message includes:

[0032] Using a deep learning model to identify the memory access request message to obtain memory access mode features and a confidence level, where the memory access mode features at least include: an offset stream and a time stream;

[0033] Updating the first mode table and the first touch table according to the memory access mode features and the confidence level.

[0034] In the above solution, updating the first mode table and the first touch table according to the memory access mode features and the confidence level includes:

[0035] Judging whether the confidence level exceeds a target threshold. If the confidence level exceeds the target threshold, storing the memory access mode features and the corresponding information in the second mode table and the second touch table;

[0036] Determine the memory access pattern features that meet the update conditions according to the second pattern table and the second touch table, and update the memory access pattern features that meet the update conditions and the corresponding information to the first pattern table and the first touch table;

[0037] Among them, the satisfaction of the update conditions includes at least one of the following:

[0038] The confidence level exceeds the first threshold;

[0039] The number of trigger times exceeds the second threshold;

[0040] The timeliness exceeds the third threshold;

[0041] The process state meets the state requirements.

[0042] In a third aspect, an embodiment of the present disclosure provides an electronic device of a network on chip, including:

[0043] At least one processor; and

[0044] A memory communicatively connected to the at least one processor; wherein,

[0045] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data prefetching method.

[0046] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the data prefetching method.

[0047] The embodiments of the present disclosure provide a data prefetching method, apparatus, electronic device and storage medium. The method is applied to an electronic device, and the electronic device includes an internal chip part and an external chip part; the method includes: receiving a memory access request message, where the memory access request message includes a memory access address and a process identifier; querying a first cache in the internal chip part according to the memory access request message; if a cache miss of the memory access address occurs in the first cache, determining a prefetch address according to the memory access address and the process identifier, and performing a data prefetching operation according to the prefetch address; wherein, determining the prefetch address according to the memory access address and the process identifier includes: querying a table based on the internal chip part and / or a table of the external chip part according to the memory access address and the process identifier to determine the prefetch address. In this way, the table stored in the internal chip part has a lower access latency, and the table of the internal chip part is preferentially used to determine the prefetch address, improving the prefetch efficiency; at the same time, combining the table of the external chip part ensures that more content can be stored, meeting the requirements of large-scale data sets and complex application scenarios.

[0048] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0049] Figure 1 It is a schematic flowchart of a data prefetching method provided by an embodiment of the present disclosure;

[0050] Figure 2 It is a schematic diagram of a trigger table provided by an embodiment of the present disclosure;

[0051] Figure 3 It is a schematic diagram of a confidence level provided by an embodiment of the present disclosure;

[0052] Figure 4 It is a schematic diagram of a pattern table provided by an embodiment of the present disclosure;

[0053] Figure 5 It is a schematic structural diagram of a data prefetching device provided by an application embodiment of the present disclosure;

[0054] Figure 6 It is a schematic flowchart of a data prefetching method provided by an application embodiment of the present disclosure;

[0055] Figure 7 It is a schematic structural diagram of a data prefetching device provided by an embodiment of the present disclosure;

[0056] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed Embodiments

[0057] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present disclosure.

[0058] Before describing the solutions of the embodiments of the present disclosure with reference to the accompanying drawings, the related technologies will be described first.

[0059] As mentioned above, with the advent of the big data era, the scale of data and the complexity of data access patterns have been increasing continuously, and traditional data prefetching technologies have been difficult to meet the requirements in the modern big data environment. Deep learning models, with their advantages such as high prediction accuracy, intelligence and automation, strong adaptability, and strong ability to process large-scale data, have become a new direction for the development of data prefetching technologies.

[0060] Neural network deep learning simulates the learning process of the human brain by constructing a multi-layer neural network, automatically learns and extracts features from data, and then realizes complex data processing and prediction tasks. By constructing a multi-layer neural network, complex memory access features are automatically extracted from data and predictions are made based on this. Its powerful learning and generalization capabilities enable deep learning models to far exceed traditional methods in terms of prediction accuracy. Especially when dealing with large-scale data sets and complex data access patterns, it can more accurately predict data requirements. In addition, the neural network-based data prefetching technology can learn data access patterns more intelligently and automatically, and automatically adjust the prefetching strategy according to the prediction results, which reduces the need for manual intervention and tuning, improves the flexibility and response speed of the system, shows strong adaptability for different application scenarios and data characteristics, and is suitable for large-scale data storage and access scenarios.

[0061] Although the above-mentioned neural network-based data prefetching technology has many advantages, it is not yet widely used in the market on a large scale. One of the main reasons is that its storage overhead is relatively large and it cannot be integrated on-chip, which is caused by the complexity of the neural network and the characteristics of data processing.

[0062] Neural networks, especially deep neural networks, usually contain a large number of parameters, such as weights and biases. These parameters need to be stored and updated during the training process, so they will occupy a large amount of storage space. The more layers and nodes a neural network has, the stronger its representation ability, but at the same time it also means that the number of parameters to be stored is larger. In addition, as the number of layers increases, gradient information, intermediate calculation results, etc. also need to be cached, further increasing the storage overhead. In order to train an accurate and efficient neural network model, a large amount of training data is usually required. These data need to be loaded into memory during the training process and processed through multiple iterations, thus increasing the storage overhead. In data prefetching technology, it may be necessary to process various types of data, such as text, images, audio, etc. These data need to occupy different spaces during storage and processing and may require specific preprocessing operations, thus increasing the complexity of storage and calculation. Therefore, it is necessary to make adjustments to address the problem of large storage overhead to meet the requirements of on-chip prefetching and ensure the feasibility and practicality of the technology.

[0063] In practical applications, to accelerate the inference process, reduce power consumption, or meet real-time requirements, it can be considered to store the results learned by the neural network model inside the chip. This way can significantly reduce the latency of data access and improve the overall performance. Moreover, since the need for data movement between off-chip and on-chip is reduced, the power consumption of on-chip storage is usually lower than that of off-chip DRAM (Dynamic Random Access Memory), which can reduce the overall power consumption of the system. Additionally, on-chip storage usually has higher reliability and stability and is not as susceptible to external interference and signal integrity problems as off-chip DRAM.

[0064] Deep learning models can extract complex memory access features from large-scale datasets and complex application scenarios and make predictions based on them. Their accuracy, flexibility, and intelligence are higher than those of traditional prefetching algorithms. However, due to their excessive storage overhead, they cannot be stored inside the chip. If all prefetching requires interaction with off-chip, it will cause significant memory access latency and be accompanied by problems such as excessive power consumption and poor stability. Therefore, the method of storing results inside the chip is considered. However, since the capacity of on-chip storage is usually much smaller than that of off-chip DRAM, it may not be able to store the entire neural network model.

[0065] Based on this, the embodiments of the present disclosure provide a data prefetching method, which is applied to an electronic device. The electronic device includes: an internal part of the chip and an external part of the chip. The method includes: receiving a memory access request message, where the memory access request message includes: a memory access address and a process identifier; querying a first cache in the internal part of the chip according to the memory access request message; if there is a cache miss for the memory access address in the first cache, determining a prefetch address according to the memory access address and the process identifier, and performing a data prefetching operation according to the prefetch address; where determining the prefetch address according to the memory access address and the process identifier includes: querying a table based on the internal part of the chip and / or a table of the external part of the chip according to the memory access address and the process identifier to determine the prefetch address. In this way, using the table in the internal part of the chip has a lower access latency, preferentially using the table in the internal part of the chip to determine the prefetch address to improve the prefetch efficiency; at the same time, combining the table of the external part of the chip to ensure that more content can be stored to meet the requirements of large-scale datasets and complex application scenarios.

[0066] Figure 1 It is a schematic flowchart of a data prefetching method provided by the embodiments of the present disclosure, which is applied to an electronic device. The electronic device includes: an internal part of the chip and an external part of the chip; as Figure 1 shown, the method includes:

[0067] Step 101, receiving a memory access request message, where the memory access request message includes: a memory access address and a process identifier;

[0068] Step 102: Query the first cache in the internal part of the chip according to the memory access request message; if there is a cache miss for the memory access address in the first cache, determine a prefetch address according to the memory access address and the process identifier, and perform a data prefetch operation according to the prefetch address.

[0069] Among them, determining the prefetch address according to the memory access address and the process identifier includes: querying a table based on the internal part of the chip and / or a table of the external part of the chip according to the memory access address and the process identifier, and determining the prefetch address.

[0070] Here, by storing memory access stream information with different values and different targets separately inside and outside the chip, data prefetch based on a deep learning model inside the chip is realized. Among them, the tables (the first touch table and the first mode table) in the internal part of the chip have a lower access latency. Storing information such as some commonly used and frequently accessed stream tags and prefetch addresses inside the chip can improve the access speed and reduce the overhead of reading external storage. The tables (the second touch table and the second mode table) in the external part of the chip can store more extensive historical data or less frequently used data. Since they are stored externally, the latency of accessing these tables is relatively high compared to inside the chip, but they can store more content, which is suitable for storing large-scale and less frequently accessed stream tags and mode information to meet the needs of large-scale data sets and complex application scenarios.

[0071] In some embodiments, when a process issues a memory access request, this memory access request message may include:

[0072] A memory access address for indicating the memory location to be accessed;

[0073] A process identifier for identifying the process that issues the request so that the system can correctly process the request.

[0074] Among them, the memory access address may be a virtual address (the address accessed by the process). For the virtual address, the virtual address can be converted into a physical address in combination with the page table base address. The physical address is the result after the virtual address to physical address conversion, and the page table base address is used to store the page table position of the mapping relationship between the virtual address and the physical address.

[0075] Here, the first cache in the internal part of the chip may be a first-level cache (L1 Cache) or a multi-level cache (such as including a first-level cache, a second-level cache, a third-level cache, etc.), which is not limited here.

[0076] The first cache may be a high-speed memory for storing commonly used data, and accessing the first cache is faster than directly accessing the main memory. The first cache in the internal part of the chip may be a first-level cache (L1 Cache) or a multi-level cache.

[0077] If the data corresponding to the memory access address is not found in the first cache, this means a cache miss. Here, when a cache miss occurs, it is proposed to use the memory access address and the process identifier to speculate on the data addresses that may be needed in the future and perform data prefetching.

[0078] Data prefetching is a method of reducing the latency of subsequent accesses by predicting the data that may be accessed next and preloading it into the first cache in advance. The process identifier is used to help determine which data may be used by the current process and avoid interference between the data of different processes. After determining the prefetch address, the predicted data is pre-loaded from the main memory into the first cache in advance. In this way, if these data need to be accessed next time, they are already in the cache, avoiding another cache miss. By the above method, the latency of memory access can be reduced and the system performance can be improved.

[0079] In some embodiments, determining the prefetch address according to the memory access address and the process identifier includes:

[0080] Query the first touch table according to the memory access address and the process identifier;

[0081] If the first touch table contains the flow label corresponding to the memory access address and the process identifier, update the trigger counter corresponding to the flow label in the first touch table, and query the first mode table according to the flow label to determine the memory access mode corresponding to the memory access address and the process identifier;

[0082] If the first touch table does not contain the flow label corresponding to the memory access address and the process identifier, query the second touch table and the second mode table according to the flow label to determine the memory access mode corresponding to the memory access address and the process identifier;

[0083] Determine the prefetch address according to the memory access mode;

[0084] Wherein, the first touch table is used to record at least one of the following: flow label, process identifier, trigger address, flow state, trigger counter;

[0085] The first mode table is used to record at least one memory access mode, and each memory access mode includes: flow label, offset flow, time flow;

[0086] The first touch table has the same format as the second touch table; the first mode table has the same format as the second mode table;

[0087] The first touch table and the first mode table are tables stored in part inside the chip; the second touch table and the second mode table are tables stored in part outside the chip.

[0088] Here, when a cache miss occurs (i.e., a cache miss for the memory access address in the first cache), check whether there is an existing flow label in the first touch table according to the missing memory access address and the corresponding process identifier.

[0089] If a corresponding flow label is found in the first touch table, perform the following operations according to the determined flow label:

[0090] I. Update the trigger counter: Each time this flow label is accessed, the trigger counter is incremented. The role of the trigger counter is to track the frequency or historical record of a certain flow label, so as to better predict future access patterns.

[0091] II. Query the first mode table: Search the first mode table for the memory access stream corresponding to the flow label to determine the memory access mode corresponding to this flow label (i.e., the combination of the memory access address and the process identifier). This memory access mode can be regarded as the rule or trend of accessing data, which is reflected by the offset stream and the time stream.

[0092] If no flow label is found in the first touch table, perform the following operations:

[0093] Query the second touch table and the second mode table: Query the second touch table to find the flow label of this memory access address and the process identifier. If a flow label is found, further query the second mode table according to this label to determine the corresponding memory access mode. Specifically, it can be sent to a processing unit outside the chip (assumed to be denoted as the space-time flow recording unit, which is used to maintain the second touch table and the second mode table).

[0094] If corresponding records are found through the second touch table and the second mode table, find the corresponding memory access mode after updating the trigger counter. If none are found, no prefetching is performed this time.

[0095] In this way, the tables inside the chip (the first touch table and the first mode table) have a lower access latency. Saving information such as some commonly used and frequently accessed flow labels and prefetch addresses inside the chip can improve the access speed and reduce the overhead of reading external storage. The tables outside the chip (the second touch table and the second mode table) can store more extensive historical data or less frequently used data. Since they are stored externally, the latency for accessing these tables is higher, but they can store more content and are suitable for storing large-scale and less frequently accessed flow labels and mode information.

[0096] Moreover, storing the first touch table and the first mode table inside the chip, and the second touch table and the second mode table outside the chip, while ensuring the same table format, this design brings about flexible data structure support, optimized prefetch mechanism, simplified data migration, enhanced scalability, reduced design complexity, and potential performance optimization. This enables the system to maintain consistent processing logic when switching between different storage media, while improving the overall performance and reliability of the system.

[0097] In some embodiments, determining the prefetch address according to the memory access mode includes:

[0098] Determining the offset value in the offset stream according to the memory access mode, and determining the prefetch address according to the base address and the offset value.

[0099] Specifically, after finding the corresponding memory access mode, the stream prefetching unit sequentially fetches the offset values recorded in the memory access mode and calculates the prefetch address according to the base address, and sends the prefetch address to the controller for the controller to perform selection and filtering operations, and perform prefetch operations according to the prefetch address.

[0100] Here, the stream prefetching unit can be used to predict in advance which memory locations will be accessed in the future according to the memory access mode, and calculate the specific prefetch address through the base address (starting address) according to the offset of each memory access (such as the distance between data).

[0101] In some embodiments, performing data prefetch operations according to the prefetch address includes:

[0102] Performing a filtering operation according to the prefetch address, and generating a prefetch request according to the prefetch address after the filtering operation;

[0103] Performing a prefetch operation according to the prefetch request.

[0104] Here, the prefetch request may include the physical address of the predicted data to be accessed; it will be provided to the first cache, and the first cache prefetches data from the memory of the electronic device according to the prefetch request.

[0105] In some embodiments, during the process of performing the prefetch operation according to the prefetch request, the method further includes:

[0106] Tracking the execution process of the prefetch operation to obtain a tracking result;

[0107] Updating a deep learning model according to the tracking result, where the deep learning model is used to identify memory access request messages to obtain memory access mode features, and the memory access mode features at least include an offset stream and a time stream.

[0108] Specifically, after obtaining the prefetch address, it is possible to judge the prefetch address to determine which ones need to be executed and which ones need to be fetched. Specifically, the latency of issuing a prefetch request for the prefetch address (i.e., the timing of issuing the request) can be calculated and compared with the previously recorded time interval range to ensure that the timing of the request sending is appropriate. If the request is sent too late, these requests can be temporarily stored and issued when the timing is appropriate; if the request is too early, it may affect the system stability or cause unnecessary cache pollution (i.e., the data loaded in advance may no longer be needed soon), so these requests can be filtered out and fed back.

[0109] After issuing the prefetch request, the execution of these prefetch requests will also be continuously monitored to observe whether they are effective. In this way, it can help evaluate the effect of the prefetch mechanism. By tracking the results, the prediction accuracy rate can be calculated. If the accuracy rate of the prefetch request is low (i.e., there are many cases of prediction errors), the prefetch requests for this memory access stream can be suspended or stopped from being issued continuously (i.e., truncating the emission of the stream) to avoid wasting resources or degrading performance. And based on these tracking results (such as prediction accuracy rate, latency, etc.), it can be fed back to the deep learning model for auxiliary training. Through continuous training, the deep learning module can improve future prefetch predictions and continuously improve the performance of the system.

[0110] Specifically, by tracking the execution process of the prefetch operation to obtain the tracking results, including: monitoring or recording the execution process of the prefetch operation, which can include tracking which data is prefetched, when it is prefetched, the timeliness of the prefetch, the success or failure of the prefetch, etc.

[0111] The memory access pattern features refer to the specific patterns or rules of data access when accessing memory. These features can be certain preferences. For example, the offset stream refers to the offset of data access or the rule of access location (such as when accessing data, some data may always be accessed continuously or have a certain regularity), and the time stream refers to the time series features of data access, that is, the access timestamp corresponding to each offset value.

[0112] In this way, by tracking the accuracy rate information of the prefetch request and cutting off the memory access stream when the accuracy rate is lower than the threshold, the accuracy of the prefetch is further ensured. By precisely controlling the timing of the prefetch request, filtering out unnecessary requests, and adjusting the prefetch strategy according to the feedback information. Finally, these control data will be fed back to the deep learning module to optimize the accuracy of the prefetch prediction, thereby improving the system performance.

[0113] The following combines Figure 2 、 Figure 3 、 Figure 4 To explain the above-mentioned first mode table, second mode table, first touch table, and second touch table.

[0114] For example, a schematic diagram of a second touch table is provided, as Figure 2 shown. The second touch table includes the following: flow label, process identifier, trigger address, flow status, trigger counter. Whenever a new entry is created in the second touch table, the flow label, process identifier, trigger address, and flow status are passed to and recorded in the second mode table, and an initial value is set for the trigger counter according to the result of the deep learning model.

[0115] Among them, the flow label is used to distinguish different data streams, processes, or threads, facilitating management and scheduling.

[0116] The process identifier is a number used by the operating system to uniquely identify a process.

[0117] The trigger address is the address for triggering an entry in the second mode table, that is, the memory access address required when certain events occur. In application, the memory access address can be matched with the trigger address in the second mode table.

[0118] The flow status refers to the current status of a certain data stream, process stream, or task stream. Different numbers have different meanings, reflecting different confidence levels. The specific meanings are as Figure 3 shown.

[0119] The trigger counter is a counter used to track or count the number of occurrences of trigger events.

[0120] During the management process of the second mode table and the second touch table, if an offset stream identical to the current memory access stream is found in the second mode table, then it is checked whether there is the same process identifier in the second touch table according to the flow label. If there are those from the same process and with the same flow status, then the trigger addresses in the second touch table are merged; here, merging the trigger addresses can simplify management or improve performance. Through merging, duplicate operations can be reduced or processing can be accelerated. Otherwise, if there are no those from the same process and with the same flow status, then it is checked whether there are memory access streams with the same or higher flow status numbers among those with the same label, and according to the least recently used (LRU) principle, the least recently unused entry is selected for replacement.

[0121] If there is no identical flow label and all entries in the second touch table are full, then the one with the smallest total value of the trigger counter is found for replacement, and feedback is given to the second mode table for corresponding replacement. Here, a smaller total value of the trigger counter means that these entries may not have been accessed frequently, and generally these entries can be selected for replacement. The result of the replacement operation is fed back to the second mode table to maintain data consistency and synchronization between the two tables.

[0122] Thus, various strategies (such as flow tags, flow states, least recently used, trigger counters, etc.) are used to manage the entries in the flow pattern and trigger table. For example, if it is found that the table is full, entries are replaced according to certain rules (such as the smallest trigger counter, matching flow tags, identical flow states, etc.), and finally, the data in the second pattern table and the second trigger table are updated. This process helps optimize the management of flows and ensures that the system efficiently processes memory and cache flow accesses.

[0123] Here, the entries in the second pattern table and the second trigger table are in a multiple relationship, and this multiple is the maximum number of entries that can be stored for each flow tag.

[0124] For example, a schematic diagram of the second pattern table is provided, as Figure 4 shown, which shows that at least one memory access pattern is stored in the second pattern table, and each of the memory access patterns includes: an offset flow, a time flow, and a separate tag set for each memory access pattern. For example, the flow tag is sig1, the offset flow is (2, 2, 5, -6...), and the time flow is (3, 1, 6, 3...).

[0125] Among them, the time flow refers to the time interval corresponding to each offset value in the offset flow, which is used to judge timeliness in the controller. When there are multiple identical offset flow patterns, the time flow maintains the shortest memory access interval and the longest memory access interval for each offset flow.

[0126] The offset flow refers to the offset of data access or the pattern of access positions (such as when accessing data, some data may always be accessed continuously or have a certain pattern). For example, after accessing an address, the next address accessed is an offset relative to the previous address. The purpose of the offset flow is to capture the variation pattern of relative addresses in memory accesses.

[0127] Taking Figure 4 as an example, assume that the memory access flow shown in the figure is a valid memory access pattern that can be stored. Addresses starting with the same letter indicate memory access requests from the same process. When the first cache receives a memory access request, it will synchronously record the timestamp and send it to the deep learning model.

[0128] The memory access stream records multiple offset streams and time streams. For example, for the access processes of address A, address B, and address C, the offset stream and time of address A, the offset stream and time of address B, and the offset stream and time of address C can be obtained. Taking address A as an example, for address A, the offset value from the next issued memory access address is 2, that is, the memory access address is (A + 2), and the timestamp difference is 3, that is, the access time is (t1 + 3). Then the offset value from the next one is 2, that is, the memory access address is (A + 4), and the timestamp difference is 1, that is, the access time is (t1 + 4). And so on, a stream of a specified length is obtained and stored in the second mode table, and a separate stream label is set for the offset stream (for address A, it corresponds to the record of sig1). For address B, the record corresponding to sig2 can be obtained; for address C, the record corresponding to sig3 can be obtained, which will not be elaborated here one by one.

[0129] When the deep learning model sends the obtained offset stream and time stream information to the second mode table, it will first check whether there is an existing offset stream in the second mode table with the same entry. If so, the time stream information is updated according to the relevant information, and an update operation is performed in the second trigger table with this stream label; if not, a new entry is created to save the information and a new stream label is assigned. If the entries in the table are full, a signal is sent to the mode merge table to wait for feedback on how to replace.

[0130] It should be noted that the structures of the first trigger table and the first mode table are the same as those of the second trigger table and the second mode table respectively (that is, also like the Figure 2 trigger table, Figure 4 the structure of the mode table), so the structures of the first trigger table and the first mode table will not be elaborated here. To reduce the overhead cost, operations such as reducing the number of entries, limiting the length of the offset stream, and compressing the trigger address can be performed.

[0131] Outside the chip, there can be a unit for maintaining and managing the second trigger table and the second mode table, which is assumed to be denoted as the space-time stream recording unit. This space-time stream recording unit can send the entries in the second trigger table and the second mode table that meet the conditions to the first trigger table and the first mode table after compressing the trigger address. The compression of the trigger address can be achieved by common methods such as taking the low bits of the trigger address and hash operation.

[0132] In the embodiments of the present disclosure, the trigger addresses from the same process and with the same stream characteristics and stream states in the trigger table are merged, and the storage overhead cost is reduced by merging the memory access modes, which is convenient for finding high-value memory access stream information. The trigger table is grouped according to the stream label to avoid a certain stream occupying too much space and preventing other processes from obtaining prefetch requests.

[0133] Meanwhile, two tables, namely the touch table and the pattern table, are maintained, and stream tags are used to represent stream relationships, avoiding the impact on power consumption, etc. caused by using a single large storage unit, and facilitating operations such as searching and updating the tables.

[0134] It should be noted that the touch table may not store the trigger address, but only store other stream information. When several consecutive offset values in the same process correspond one by one to the first few offset values in the stream feature pattern, the remaining offset values are taken out and the prefetch address is calculated based on this.

[0135] In some embodiments, the method further includes:

[0136] Updating the first touch table and / or the first pattern table according to the memory access request message.

[0137] Specifically, the updating the first touch table and / or the first pattern table according to the memory access request message includes:

[0138] Using a deep learning model to identify the memory access request message to obtain a memory access pattern feature and a confidence level, where the memory access pattern feature includes: an offset stream and a time stream;

[0139] Updating the first pattern table and the first touch table according to the memory access pattern feature and the confidence level.

[0140] Specifically, updating the first pattern table and the first touch table according to the memory access pattern feature and the confidence level includes:

[0141] Judging whether the confidence level exceeds a target threshold. If the confidence level exceeds the target threshold, storing the memory access pattern feature and the corresponding information into the second pattern table and the second touch table;

[0142] Determining a memory access pattern feature that meets the update condition according to the second pattern table and the second touch table, and updating the memory access pattern feature that meets the update condition and the corresponding information to the first pattern table and the first touch table;

[0143] Wherein, the meeting the update condition includes at least one of the following:

[0144] The confidence level exceeds a first threshold;

[0145] The trigger times exceed a second threshold;

[0146] The timeliness exceeds a third threshold;

[0147] The process state meets the state requirement.

[0148] Specifically, after receiving memory access request information (including memory access address, process identifier, etc.), the deep learning model will identify the memory access pattern, and when the obtained memory access pattern determines memory access pattern features (i.e., offset stream and time stream) that meet the confidence requirement and have higher value than the stored entries, they will be sent to the second pattern table and the second trigger table for storage.

[0149] According to the updated second pattern table and second trigger table, when higher-value memory access pattern features can be determined by comprehensively considering information such as confidence, trigger times, timeliness, and process status, they will be sent to the first pattern table and the first trigger table inside the chip for storage.

[0150] Here, whether to send to the first trigger table and the first pattern table needs to be considered comprehensively in terms of stream status, the trigger times recorded by the trigger counter, process status, etc.

[0151] That is to say, the judgment of whether the value is high can be determined by combining confidence, trigger times, timeliness, process status, etc.

[0152] For example, if the deep learning model gives memory access pattern features and a high confidence level when identifying a certain memory access request message, it indicates that it is more certain about the identification of the memory access pattern features and is very effective in the current process or workload. Generally, different confidence levels are represented by 0 - 3, and entries with a status of 0 are usually considered to have a higher confidence level and are more prioritized.

[0153] The trigger counter is used to track the number of times a certain memory access pattern is triggered. If the trigger times of a certain stream label are higher than a set threshold, the memory access pattern features (i.e., offset stream and time stream) corresponding to this stream label are usually more important and have higher usage value.

[0154] Timeliness can be obtained based on the feedback of specific operations. The deep learning model can adjust confidence, etc. in combination with timeliness during the training process. The process status can be determined based on the specific running status of the process. Generally, running processes are more important.

[0155] It should be noted that whether the update condition is met can also be scored by combining at least one of the confidence, trigger times, timeliness, and process status. The scoring method can be set according to requirements, and after obtaining the score, it can be determined according to whether the score exceeds the score threshold.

[0156] For example, if the confidence recorded by the stream status needs to be greater than the second threshold (i.e., the stream status is 0 or 1), then there is a need for in-chip prefetching, and entries with a stream status of 0 have higher priority than those with a status of 1;

[0157] The higher the trigger times recorded by the trigger counter, the more valuable this entry can be proven to be compared with the existing entries.

[0158] Here, the specific values of the first threshold, the second threshold, and the third threshold are not limited, and the status requirement may be the running state.

[0159] In addition, for the process state, for a process that has ended, the entries stored in the on-chip table will be sent to the off-chip table for replacement operations. Through the above operations, it is ensured that the relevant information of the most valuable memory access patterns is retained in the first trigger table and the first mode table.

[0160] In this way, considering the problem that the capacity of on-chip storage is usually much smaller than that of off-chip DRAM and may not be able to store the training results of the entire deep learning model, an effective data management and migration strategy is provided for the problem that the training results of the deep learning model cannot be fully stored inside the chip, so as to ensure that data can be quickly migrated from off-chip to on-chip when needed (that is, the information in the second mode table and the second trigger table is selectively updated to the first mode table and the first trigger table), and the usage efficiency of the cache is optimized.

[0161] Furthermore, by adopting the method of storing results on-chip, it not only utilizes the fact that the deep learning model can extract complex memory access features in large-scale data sets and complex application scenarios and make predictions based on this, with higher accuracy, flexibility, and intelligence compared to traditional prefetching algorithms, but also avoids the problems that it cannot be stored on-chip due to its excessive storage overhead, and that if all prefetching requires interaction with off-chip, it will cause a large memory access delay, as well as problems such as excessive power consumption and poor stability.

[0162] Moreover, in the embodiments of the present disclosure, operations such as updating and replacing the table comprehensively consider the trigger times, confidence (flow state), accuracy and timeliness of prefetching, process state, etc., and can effectively ensure that higher-value data is retained in the table compared to a single replacement strategy.

[0163] In some embodiments, the method further includes: training a deep learning model, specifically including:

[0164] Collecting training data and corresponding labels, where the training data includes access request data containing information such as memory access addresses and process identifiers. The labels include memory access pattern features and confidence, and the memory access pattern features specifically include: offset stream, time stream, and the confidence can comprehensively reflect the accuracy, timeliness, etc. of this memory access pattern;

[0165] Using the training data and corresponding labels to train a neural network model, and obtaining the trained neural network model as the deep learning model in the embodiments of the present disclosure.

[0166] And, after subsequent use, information such as each access request information and the finally obtained prefetch results can be continuously collected to optimize the deep learning model.

[0167] Figure 5 Schematic diagram of a data prefetching device provided by an application embodiment of the present disclosure; as Figure 2 shown, the device includes any chip, and the device includes:

[0168] A deep learning unit, a peripheral memory, and a spatio-temporal flow recording unit outside the chip;

[0169] A memory management unit, a multi-level cache, a processor (CPU), a trigger table, a mode table, a flow prefetch unit, a controller, and a CSR (Control and Status Register) register inside the chip.

[0170] In some embodiments, the deep learning unit is used to track the memory access flow information of the same process, learn the memory access pattern according to the memory access flow information, and transmit the memory access flow information that meets the conditions to the spatio-temporal flow recording unit for storage.

[0171] In some embodiments, the spatio-temporal flow recording unit maintains a second mode table and a second trigger table; among them, the second mode table can receive the memory access flow information sent by the deep learning unit, set flow tags for each type of offset flow and time flow, and send the corresponding memory access flow information to the mode merge table for storage; in the second trigger table, the trigger addresses for the same process and flow state can be merged, and a trigger counter is maintained for the entries, and when the conditions are met, it is sent to the first trigger table and the first mode table inside the chip for replacement. When the trigger condition is met, the flow prefetch unit can sequentially retrieve and calculate the prefetch address from the first trigger table, and issue a prefetch request after being selected and filtered by the controller.

[0172] Software can regulate the overall prefetch structure by writing to the CSR register, including clearing and modifying the entries in the mode table, adjusting the controller parameters, setting the prefetch degree, etc. In this way, the prefetch degree is regulated by software and the simulation delay operation is used to improve flexibility, and the way of writing to the CSR register is used to reduce the software regulation cost, realizing the data prefetch technology of software and hardware cooperation.

[0173] In some embodiments, the deep learning unit is used to receive information (such as page table base address, address space identifier, etc.) from the memory management unit and the multi-level cache, combine the information (such as page table base address, address space identifier, etc.) that identifies the process space in the memory management unit to distinguish the process where the instruction is located, generate a process identifier, and obtain the memory access pattern under this process by identifying the memory access flow information under the same process. Here, in addition to using the combination of the page table base address and the address space identifier to distinguish processes, identification information such as process identifier, process owner ID, user identifier, group identifier, external identifier (user-defined) can also be used.

[0174] When the confidence level of the memory access pattern obtained by the deep learning model in the deep learning unit is higher than the fourth threshold, the deep learning unit sends the memory access pattern features and corresponding information (such as memory access stream information, confidence level, etc.) to the spatio-temporal stream recording unit. The deep learning model can adopt common types, such as convolutional neural network, recurrent neural network and other models.

[0175] Figure 6 It is a schematic flowchart of a data prefetching method provided by an application embodiment of the present disclosure; as Figure 6 shown, the method includes:

[0176] Step 601, determine the memory access address and process identifier according to the memory access request message;

[0177] Step 602, when a cache miss occurs, query whether there is an existing entry in the first touch table according to the memory access address and process identifier; if a corresponding existing entry is found in the first touch table, go to step 603; if not, go to step 604;

[0178] Step 603, query the corresponding memory access pattern in the first pattern table according to the stream label, and go to step 605;

[0179] Step 604, send it to the spatio-temporal stream recording unit outside the chip to find the records in the second touch table and the second pattern table. If found, after maintaining the trigger counter (incrementing by 1), find the corresponding memory access pattern and go to step 605; if not found, no prefetching is performed this time.

[0180] Step 605, according to the determined memory access pattern, use the stream prefetcher to sequentially fetch the offset values and calculate the prefetch address according to the base address, and send the obtained prefetch address to the controller for selection and filtering operations.

[0181] Step 606, the controller issues a prefetch request and tracks it, and feeds back the tracking result to the deep learning unit.

[0182] Here, the controller can calculate the latency of the issued prefetch request and compare it with the recorded time interval range. For requests that are too late, save them for a while and then send them. For requests that are too early, filter them out and feedback to avoid contaminating the cache. Then, after issuing the prefetch request, the controller will track the issued prefetch request situation and decide whether to truncate the emission of the stream according to the accuracy information. Finally, the tracking information maintained in the controller will be fed back to the deep learning module for auxiliary training.

[0183] It should be noted that the above solution supports software control to improve the accuracy information of prefetching. The software can configure the parameters of the stream prefetching through writing operations to the CSR register. When the stream prefetching obtains the offset value in the stream memory access mode, the number of obtained offset values, that is, the prefetch degree, is achieved through software control. In the controller, the memory access latency simulated by each prefetch request can be adjusted by software to achieve accurate latency calculation.

[0184] The solution provided by the embodiments of the present disclosure realizes in-chip data prefetching by storing the results of deep learning models. For the problem of large storage overhead, methods such as combining memory access mode merging, ending process eviction, and memory access stream tags are used to ensure efficient and high-value in-chip storage space management, and the off-chip space is used to record more effective memory access stream information. For the situation where the accuracy, coverage, etc. information of different processes or different memory access modes are inconsistent, adaptive multi-level prefetching and prefetching termination operations are performed to reduce the pollution of the cache by prefetching and improve the overall data integrity. For the blindness problem of hardware prefetching, the software control method is used to further improve flexibility, allowing developers to make accurate predictions according to the memory access mode and data dependency of the program, and further improving the accuracy of prefetching. For the timeliness problem, the memory access latency of the data prefetch request is calculated in the controller and compared with the historical information to determine the prefetch timing and filter some requests.

[0185] Figure 7 It is a schematic structural diagram of a data prefetching device provided by an embodiment of the present disclosure; as Figure 7 shown, the device is applied to an electronic device, and the electronic device includes: an in-chip part and an out-of-chip part; the device includes:

[0186] A receiving module, configured to receive a memory access request message, where the memory access request message includes: a memory access address and a process identifier;

[0187] A processing module, configured to query a first cache in the in-chip part according to the memory access request message; if a cache miss of the memory access address occurs in the first cache, determine a prefetch address according to the memory access address and the process identifier, and perform a data prefetching operation according to the prefetch address;

[0188] Among them, the determining the prefetch address according to the memory access address and the process identifier includes: querying a table based on the in-chip part and / or a table of the out-of-chip part according to the memory access address and the process identifier to determine the prefetch address.

[0189] In some embodiments, the processing module is configured to query a first touch table according to the memory access address and the process identifier;

[0190] If the stream tag corresponding to the memory access address and the process identifier is included in the first touch table, update the trigger counter corresponding to the stream tag in the first touch table, and query the first mode table according to the stream tag to determine the memory access mode corresponding to the memory access address and the process identifier;

[0191] If the stream tag corresponding to the memory access address and the process identifier is not included in the first touch table, query the second touch table and the second mode table according to the stream tag to determine the memory access mode corresponding to the memory access address and the process identifier;

[0192] Determine the prefetch address according to the memory access mode;

[0193] Wherein, the first touch table is used to record at least one of the following: stream tag, process identifier, trigger address, stream status, trigger counter;

[0194] The first mode table is used to record at least one memory access mode, and each memory access mode includes: stream tag, offset stream, time stream;

[0195] The first touch table has the same format as the second touch table; the first mode table has the same format as the second mode table;

[0196] The first touch table and the first mode table are tables stored in part inside the chip; the second touch table and the second mode table are tables stored in part outside the chip.

[0197] In some embodiments, the processing module is configured to determine the offset value in the offset stream according to the memory access mode, and determine the prefetch address according to the base address and the offset value.

[0198] In some embodiments, the processing module is configured to perform a filtering operation according to the prefetch address, generate a prefetch request according to the prefetch address after the filtering operation; perform a data prefetch operation according to the prefetch request.

[0199] In some embodiments, the processing module is configured to track the execution process of the prefetch operation during the prefetch operation according to the prefetch request, and obtain a tracking result;

[0200] Update the deep learning model according to the tracking result, where the deep learning model is used to identify the memory access request message to obtain memory access mode features, and the memory access mode features at least include an offset stream and a time stream.

[0201] In some embodiments, the processing module is further configured to update the first touch table and / or the first mode table according to the memory access request message.

[0202] In some embodiments, the processing module is configured to identify the memory access request message by using a deep learning model, and obtain a memory access pattern feature and a confidence level. The memory access pattern feature at least includes: an offset stream and a time stream;

[0203] Update the first pattern table and the first touch table according to the memory access pattern feature and the confidence level.

[0204] In some embodiments, the processing module is configured to determine whether the confidence level exceeds a target threshold. If the confidence level exceeds the target threshold, store the memory access pattern feature and the corresponding information in the second pattern table and the second touch table;

[0205] Determine a memory access pattern feature that meets the update condition according to the second pattern table and the second touch table, and update the memory access pattern feature that meets the update condition and the corresponding information to the first pattern table and the first touch table;

[0206] Wherein, the condition for meeting the update condition includes at least one of the following:

[0207] The confidence level exceeds a first threshold;

[0208] The number of trigger times exceeds a second threshold;

[0209] The timeliness exceeds a third threshold;

[0210] The process state meets the state requirement.

[0211] It can be understood that when implementing the corresponding data prefetching method, the data prefetching device provided in the above embodiments can, as needed, allocate the above processing to different program modules to complete all or part of the processing described above. In addition, the device provided in the above embodiments and the embodiments of the corresponding method belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be elaborated here.

[0212] The embodiments of the present disclosure provide a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be triggered to execute the data prefetching method provided by the embodiments of the present disclosure.

[0213] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0214] Figure 8 It is a schematic structural diagram of an electronic device of a network-on-chip provided by the embodiments of the present disclosure; as Figure 8As shown, the electronic device 80 includes: a processor 801 and a memory 802 communicatively connected to the processor 801; the memory 802 stores instructions executable by the processor 801; when the instructions are executed by the processor 801, the processor 801 is enabled to execute a data prefetching method.

[0215] Of course, the electronic device provided in the above embodiment and the embodiment of the corresponding method belong to the same concept. The processor 801 can also execute the data prefetching method of any one of the above. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.

[0216] In practical applications, the electronic device 80 may further include: at least one network interface 803. Each component in the electronic device 80 is coupled together through a bus system 804. It can be understood that the bus system 804 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 804 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 5 all kinds of buses are labeled as the bus system 804. Among them, the number of the processors 801 can be at least one, and the number of the memories 802 can be at least one. The network interface 803 is used for the communication between the electronic device 80 and other devices in a wired or wireless manner.

[0217] The memory 802 in the embodiments of the present disclosure is used to store various types of data to support the operation of the electronic device 80.

[0218] The method disclosed in the above embodiments of the present disclosure can be applied to the processor 801 or implemented by the processor 801. The processor 801 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 801 or the instructions in software form. The above-mentioned processor 801 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 801 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present disclosure, it can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory 802. The processor 801 reads the information in the memory 802 and combines its hardware to complete the steps of the foregoing method.

[0219] In some embodiments, the electronic device 80 may be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components, and is used to execute the foregoing method.

[0220] It should be understood that the various forms of the processes shown above may be used, and steps may be reordered, added, or deleted. For example, the steps described in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. No limitation is imposed herein.

[0221] In the above description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.

[0222] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by those of ordinary skill in the technical field to which this disclosure belongs. The terms used in this disclosure are only for the purpose of describing the embodiments of this disclosure and are not intended to limit this disclosure.

[0223] It should be understood that in the various embodiments of this disclosure, the magnitude of the serial numbers of the respective implementation processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this disclosure.

[0224] In addition, the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this disclosure, "a plurality" means two or more, unless otherwise specifically defined.

[0225] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claimed rights.

Claims

1. A data prefetching method, characterized in that The method is applied to an electronic device, and the electronic device includes: an internal part of a chip and an external part of the chip; the method includes: Receiving a memory access request message, where the memory access request message includes: a memory access address and a process identifier; Querying a first cache in the internal part of the chip according to the memory access request message; if a cache miss of the memory access address occurs in the first cache, determining a prefetch address according to the memory access address and the process identifier, and performing a data prefetch operation according to the prefetch address; Wherein, determining the prefetch address according to the memory access address and the process identifier includes: querying a table in the internal part of the chip and / or a table in the external part of the chip according to the memory access address and the process identifier to determine the prefetch address.

2. The method according to claim 1, wherein Determining the prefetch address according to the memory access address and the process identifier includes: Querying a first touch table according to the memory access address and the process identifier; If the first touch table contains a flow label corresponding to the memory access address and the process identifier, updating a trigger counter corresponding to the flow label in the first touch table, and querying a first mode table according to the flow label to determine a memory access mode corresponding to the memory access address and the process identifier; If the first touch table does not contain a flow label corresponding to the memory access address and the process identifier, querying a second touch table and a second mode table according to the flow label to determine a memory access mode corresponding to the memory access address and the process identifier; Determining a prefetch address according to the memory access mode; Wherein, the first touch table is used to record at least one of the following: flow label, process identifier, trigger address, flow status, trigger counter; The first mode table is used to record at least one memory access mode, and each memory access mode includes: flow label, offset flow, time flow; The first touch table has the same format as the second touch table; the first mode table has the same format as the second mode table; The first touch table and the first mode table are tables stored in the internal part of the chip; the second touch table and the second mode table are tables stored in the external part of the chip.

3. The method according to claim 2, characterized in that, Determining the prefetch address according to the memory access mode includes: Determining an offset value in the offset flow according to the memory access mode, and determining a prefetch address according to a base address and the offset value.

4. The method according to claim 1, wherein Performing a data prefetch operation according to the prefetch address includes: Performing a filtering operation according to the prefetch address, and generating a prefetch request according to the prefetch address after the filtering operation; Performing a data prefetch operation according to the prefetch request.

5. The method according to claim 4, wherein During the process of performing a prefetch operation according to the prefetch request, the method further includes: Tracking the execution process of the prefetch operation to obtain a tracking result; Updating a deep learning model according to the tracking result, where the deep learning model is used to identify a memory access request message to obtain a memory access mode feature, and the memory access mode feature at least includes an offset flow and a time flow.

6. The method according to claim 2, characterized in that, The method further includes: Updating the first touch table and / or the first mode table according to the memory access request message.

7. The method according to claim 6, wherein Updating the first touch table and / or the first mode table according to the memory access request message includes: Identify the memory access request message using a deep learning model to obtain a memory access pattern feature and a confidence level. The memory access pattern feature at least includes: an offset stream and a time stream. Update the first pattern table and the first trigger table according to the memory access pattern feature and the confidence level.

8. The method according to claim 7, wherein Updating the first pattern table and the first trigger table according to the memory access pattern feature and the confidence level includes: Determine whether the confidence level exceeds a target threshold. If the confidence level exceeds the target threshold, store the memory access pattern feature and the corresponding information in the second pattern table and the second trigger table. Determine the memory access pattern feature that meets the update condition according to the second pattern table and the second trigger table, and update the memory access pattern feature that meets the update condition and the corresponding information to the first pattern table and the first trigger table. Wherein, the meeting the update condition includes at least one of the following: The confidence level exceeds a first threshold. The number of trigger times exceeds a second threshold. The timeliness exceeds a third threshold. The process state meets the state requirement.

9. A data prefetching device, characterized in that, The device is applied to an electronic device, and the electronic device includes: an internal part of the chip and an external part of the chip; the device includes: A receiving module, configured to receive a memory access request message, where the memory access request message includes: a memory access address and a process identifier. A processing module, configured to query a first cache in the internal part of the chip according to the memory access request message; if there is a cache miss for the memory access address in the first cache, determine a prefetch address according to the memory access address and the process identifier, and perform a data prefetch operation according to the prefetch address. Wherein, determining the prefetch address according to the memory access address and the process identifier includes: querying a table based on the internal part of the chip and / or a table of the external part of the chip according to the memory access address and the process identifier to determine the prefetch address.

10. An electronic device, characterized in that, Includes: At least one processor; And, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • High-precision memory access prediction and cache prefetching method and system based on machine learning

    CN121996576A