Caching Processing Method for Heterogeneous Devices, Heterogeneous System, Product, Device and Medium

By obtaining the number of accesses to each device in a heterogeneous system, determining and verifying the file page, and then moving it to the processor memory, the memory delay of CPU accessing heterogeneous devices in a heterogeneous system is solved, and data transmission rate and system efficiency are improved.

CN119473168BActive Publication Date: 2025-07-11LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510046220.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-07-11
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

In heterogeneous systems, the CPU has a long delay in accessing the memory of heterogeneous devices, resulting in a reduced data transmission rate bandwidth and affecting system efficiency.

Method used

By obtaining the number of accesses to target virtual address of each device in a heterogeneous system, determining the file page of the target heterogeneous device, and moving it to the processor memory after verification is passed, establishing a global address space allocation relationship, realizing cache management of heterogeneous devices, and optimizing data transmission paths.

Benefits of technology

It improves the global access efficiency on the host side, reduces the latency of CPU accessing heterogeneous device memory, and improves data transmission rate and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119473168B_ABST
    Figure CN119473168B_ABST
Patent Text Reader

Abstract

The present invention discloses a cache processing method, a heterogeneous system, a product, a device and a medium for heterogeneous devices, relating to the technical field of data storage. The access times of the target virtual addresses corresponding to the devices in the heterogeneous system are used to implement a global data caching mechanism. The file pages corresponding to the target heterogeneous devices are determined based on the access times of the target virtual addresses, enabling the file system layer to have the ability to manage the memory cache of the heterogeneous devices. The predicted file pages of the target heterogeneous devices are subjected to a verification process, and are moved in the case of passing the verification, improving the allocation accuracy of the target heterogeneous devices. The file pages of the target heterogeneous devices are mapped and moved to the memory on the host side, eliminating the need to access the target heterogeneous devices and directly accessing the file pages stored in the processor memory, avoiding the increased access rate caused by the processor accessing the target heterogeneous devices, reducing the latency of accessing the memory of the target heterogeneous devices, and improving the data transfer rate and the system efficiency of the heterogeneous system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and in particular, to a cache processing method for heterogeneous devices, a heterogeneous system, a product, a device, and a medium. Background Art

[0002] When a central processing unit (CPU) in a heterogeneous system accesses a heterogeneous device, since the rate at which the CPU accesses the memory of the heterogeneous device (such as a field-programmable gate array (FPGA)) is significantly lower than the rate at which the CPU accesses its own memory, the latency for accessing the memory of the heterogeneous device is long, and the bandwidth corresponding to the data transfer rate is reduced. If the memory of the heterogeneous device is frequently accessed, the system efficiency of the heterogeneous system will be reduced.

[0003] Therefore, how to improve the system efficiency of a heterogeneous system is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0004] The object of the present invention is to provide a cache processing method for heterogeneous devices, a heterogeneous system, a product, a device, and a medium, so as to solve the problem that the latency for accessing the memory of a heterogeneous device is long, the bandwidth corresponding to the data transfer rate is reduced, and thus the system efficiency of the heterogeneous system is reduced.

[0005] To solve the above technical problem, the present invention provides a cache processing method for heterogeneous devices, including:

[0006] Obtaining the target virtual address access times corresponding to each device in the heterogeneous system;

[0007] Determining the file page corresponding to the target heterogeneous device based on the target virtual address access times; and performing a verification process on the file page;

[0008] When the verification is passed, moving the file page to the memory of the processor so that the processor can access the file page of the target heterogeneous device.

[0009] On the one hand, obtaining the target virtual address access times corresponding to each device in the heterogeneous system includes:

[0010] Pre-establishing a global address space allocation relationship corresponding to each device in the heterogeneous system; wherein the global address space allocation relationship stores the address space access information corresponding to the processor and the heterogeneous device; the virtual address access times of the address space access information include the first virtual address access times corresponding to external devices accessing local devices and / or the second virtual address access times corresponding to local devices accessing themselves;

[0011] Determine the access times of the target virtual address according to the global address space allocation relationship.

[0012] On the other hand, determining the access times of the target virtual address according to the global address space allocation relationship includes:

[0013] Obtain the preset virtual address access times and / or the preset local access ratio;

[0014] Determine the local access ratio according to the first virtual address access times and the second virtual address access times;

[0015] In the global address space allocation relationship, select the virtual address access times that exceed the preset virtual address access times and / or the local access ratio is less than the preset local access ratio, and then use the virtual address access times corresponding to the virtual address access times that exceed the preset virtual address access times and / or the local access ratio is less than the preset local access ratio as the access times of the target virtual address.

[0016] On the other hand, when the virtual address access times include the first virtual address access times or the second virtual address access times, and there are multiple virtual addresses corresponding to the target virtual address access times, determining the file page corresponding to the target heterogeneous device based on the target virtual address access times includes:

[0017] Use the virtual address corresponding to the target virtual address access times as the second target virtual address;

[0018] When the number of the second target virtual addresses is multiple, determine whether the file pages to which the second target virtual addresses belong are the same;

[0019] If they are the same, use the heterogeneous device to which the data block corresponding to the second target virtual address belongs as the target heterogeneous device to obtain the file page corresponding to the data block of the target heterogeneous device;

[0020] If they are not the same, perform an addition process on the access times of each virtual address stored in the file pages to which the second target virtual addresses belong to obtain the sum of the access times of each;

[0021] Use the file page to which the access times of the virtual address with the largest sum of access times belong as the file page corresponding to the target heterogeneous device.

[0022] On the other hand, performing a verification process on the file page includes:

[0023] Obtain the data block to which the file page belongs and the corresponding page information;

[0024] Perform a verification process on each flag bit of the data block and the page information to obtain a verification result.

[0025] On the other hand, performing a verification process on each flag bit of the data block and the page information to obtain a verification result includes:

[0026] Perform a verification process on each flag bit of the data block to obtain a first verification result;

[0027] In the case where the first verification result is that all verifications pass, perform a verification process on the page information according to the file page hash value to obtain a second verification result;

[0028] If the second verification result is that the verification passes, then the final verification result is that the verification passes.

[0029] On the other hand, moving the file page to the memory of the processor includes:

[0030] Obtain the cache mechanism corresponding to the processor in the heterogeneous system; wherein, the cache mechanism is established by the file system layer of the file system;

[0031] Store the page attribute information corresponding to the file page into the target data structure corresponding to the cache mechanism;

[0032] Move the page information corresponding to the file page to the memory of the processor.

[0033] On the other hand, storing the page attribute information corresponding to the file page into the target data structure corresponding to the cache mechanism includes:

[0034] Obtain the target virtual address, target physical address, and the number of accesses to the target virtual address corresponding to the file page;

[0035] Save the target virtual address and the target physical address in the form of a key-value pair into the target data structure; wherein, the target data structure at least includes a key, a value, a child node pointer, a flag bit, a file page hash value, and count information;

[0036] Save the page access count to which the number of accesses to the target virtual address belongs into the count information of the target data structure;

[0037] Save the verification results of each of the flag bits of the data block corresponding to the file page into the flag bits of the target data structure;

[0038] Save the verification result of the page information corresponding to the file page into the file page hash value of the target data structure.

[0039] On the other hand, migrating the page information corresponding to the file page to the memory of the processor includes:

[0040] Storing the page information of the file page into a target data structure corresponding to the cache mechanism according to the updated memory cycle of the processor and a preset window.

[0041] On the other hand, the process of determining the preset window includes:

[0042] Obtaining the transmission system level corresponding to the file page;

[0043] Obtaining the data transmission bandwidth corresponding to the access between the processor and the target heterogeneous device;

[0044] Determining the preset window according to the data transmission bandwidth, the memory space data of the heterogeneous device, the file page information corresponding to the preset window, and the transmission system level.

[0045] On the other hand, the process of determining the updated memory cycle includes:

[0046] Obtaining an initial updated memory cycle, the number of accesses corresponding to the file page, and the transmission system level;

[0047] Determining the updated memory cycle according to the initial updated memory cycle, the number of accesses corresponding to the file page, and the transmission system level.

[0048] To solve the above technical problems, the present invention also provides a heterogeneous system, including a storage device, a heterogeneous device, a switch, and a processor; wherein, at least one heterogeneous device forms a heterogeneous device cluster;

[0049] The storage device, the heterogeneous device cluster, and the processor all achieve heterogeneous consistency through the protocol of the switch; wherein, the processor is used to implement the steps of the cache processing method of the heterogeneous device as described when executing a computer program.

[0050] To solve the above technical problems, the present invention also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the cache processing method of the heterogeneous device are implemented.

[0051] To solve the above technical problems, the present invention also provides a cache processing device for a heterogeneous device, including:

[0052] A memory for storing a computer program;

[0053] A processor for implementing the steps of the cache processing method of the heterogeneous device as described when executing the computer program.

[0054] To solve the above technical problems, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the cache processing method for heterogeneous devices as described above are implemented.

[0055] The beneficial effects of the present invention are as follows. First, the access times of the target virtual addresses corresponding to the devices in the heterogeneous system are obtained. Considering the cache management of the heterogeneous devices in the host side in the heterogeneous system, a global data cache mechanism is implemented, which is convenient for subsequent cache processing of the data of the heterogeneous devices on the host side. Second, the file pages corresponding to the target heterogeneous devices are determined based on the access times of the target virtual addresses, enabling the file system layer to have the ability to manage the memory cache of the heterogeneous devices. The determined target heterogeneous device is a target heterogeneous device predicted to be frequently used based on the access times of the target virtual addresses. Before moving the file pages of the target heterogeneous device to the memory of the processor, the file pages of the predicted target heterogeneous device are subjected to verification processing, and are moved only when the verification passes, improving the allocation accuracy of the target heterogeneous device. Finally, the file pages of the target heterogeneous device are mapped and moved to the memory on the host side to improve the global access efficiency on the host side. When accessing the target heterogeneous device on the host side, there is no need to access the target heterogeneous device directly, but to access the file pages stored in the processor memory, avoiding the increased access rate caused by the processor accessing the target heterogeneous device, reducing the latency of accessing the memory of the target heterogeneous device, and improving the data transmission rate and the system efficiency of the heterogeneous system.

[0056] Secondly, the access times of the target virtual address are determined based on the pre-established global address space allocation relationship. Through the address mapping mechanism, the global address space allocation of the heterogeneous device is added to allocate the address space and the corresponding cache allocation of the heterogeneous device, laying a foundation for the subsequent allocation process and improving the accuracy of allocation. There is a copy of the global address space allocation relationship on each device in the heterogeneous system to ensure that each device can understand the address allocation situation and relevant control information in real time, and to transmit the heterogeneous cache coherence protocol to ensure data consistency. The access times of the virtual address corresponding to the virtual address access times exceeding the preset virtual address access times and / or the local access ratio less than the preset local access ratio are selected as the target virtual address access times from the global address space allocation relationship, improving the diversity and flexibility of the determination to facilitate the subsequent determination of the accuracy of the target heterogeneous device. In the presence of multiple file pages, the access times of all virtual addresses corresponding to each of the multiple file pages need to be added up, and the file page corresponding to the maximum access times is moved for subsequent processing, improving the efficiency of cross-device memory access on the basis of making full use of the memory space of the processor memory. The check processing of each tag bit corresponding to the data block is performed to ensure the integrity of the data block. The check processing of the page information is performed to ensure whether the data in the data block has been updated or modified, improving the security of the data. First, one factor check is completed, and in the case of passing the check, another factor check is performed to improve the check accuracy and check efficiency of the file page. The page attribute information corresponding to the file page is stored in the target data structure corresponding to the cache mechanism by improving the target data structure to solve the cache management problem of the heterogeneous device in the file system layer and improve the access efficiency of the processor. The file page is moved to the memory of the processor, the page attribute information is stored in the target data structure, and the page information is moved to the memory of the processor to realize the improvement and update of the device memory host-side cache and the data structure, improving the global access efficiency of the host side. The window cache mechanism performs effective prefetching to improve the cache performance and prevent the receiver cache from overflowing. The fixed-duration cache mechanism performs effective prefetching to save memory space and improve the effectiveness of data caching.

[0057] In addition, the present invention also provides a heterogeneous system, a computer program product, a cache processing device for a heterogeneous device, and a computer-readable storage medium, which have the same beneficial effects as the above-mentioned cache processing method for a heterogeneous device. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0059] Figure 1 It is a schematic diagram of the principle of a conventional file system cache;

[0060] Figure 2 It is a flowchart of a cache processing method for heterogeneous devices provided by an embodiment of the present invention;

[0061] Figure 3 It is a schematic structural diagram of a heterogeneous system provided by an embodiment of the present invention;

[0062] Figure 4 It is a schematic diagram of the principle of a cache mechanism provided by an embodiment of the present invention;

[0063] Figure 5 It is a schematic diagram of the principle of another cache mechanism provided by an embodiment of the present invention;

[0064] Figure 6 It is a schematic diagram of a cache window provided by an embodiment of the present invention;

[0065] Figure 7 It is a structural diagram of a cache processing device for heterogeneous devices provided by an embodiment of the present invention;

[0066] Figure 8 It is a structural diagram of a cache processing apparatus for heterogeneous devices provided by an embodiment of the present invention. Detailed implementation manners

[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0068] The core of the present invention is to provide a cache processing method, a heterogeneous system, a product, a device and a medium for heterogeneous devices, so as to solve the problems of long delay in accessing the memory of heterogeneous devices and reduced bandwidth corresponding to the data transmission rate, thereby reducing the system efficiency of the heterogeneous system.

[0069] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0070] With the continuous development of technologies such as artificial intelligence, big data analysis, and scientific computing, the global data volume has shown an exponential growth trend, which has led to an increasing demand for data processing and analysis. Against this background, traditional single-architecture database systems are facing huge challenges, and their capabilities in processing massive data, ensuring data real-time performance, and coping with diverse workloads are relatively insufficient.

[0071] FPGA heterogeneous systems usually combine FPGAs with other processors (such as CPUs, Graphics Processing Units (GPUs), etc.) to make full use of the advantages of different processors to achieve high-performance computing. In this architecture, FPGAs are widely used to accelerate specific computing tasks due to their programmability, low power consumption, and high parallelism.

[0072] In FPGA heterogeneous systems, the data transfer speed is often a key factor affecting system performance. By using data caching technology, frequently accessed data can be stored in the cache close to the computing unit, thereby reducing the data access latency and improving the data access speed. At the same time, the data transfer bandwidth between the FPGA and the external memory is usually limited, and the use of data caching technology can reduce the number of accesses to the external memory, thus reducing the data transfer bandwidth requirements. In addition, storing frequently accessed data in the cache can also reduce the number of accesses to the external memory, thereby reducing the power consumption of the system. Data caching technology can further reduce the power consumption of the FPGA.

[0073] In the data transfer path of the FPGA heterogeneous system, file data is first obtained by a process, transferred to the file system for caching, then copied to the application layer Pin page, and then transferred to the heterogeneous device through DMA. Among them, the Pin page is a page locking technology that allows a process to lock one or more memory pages in physical memory instead of having the operating system's Memory Management Unit (MMU) swap them out to the swap space on the disk. This is usually used for pages that need to be retained in physical memory to ensure that the target page does not swap to the disk during the transmission process of the Direct Memory Access (DMA) hardware layer. In the whole process of the traditional data transfer path, obtaining file data, transferring it from the storage device to the file system, caching it in the file system of the CPU, copying it multiple times, accessing it to the heterogeneous device, and filtering the data by the heterogeneous device, and then transferring it to the CPU will result in more transmission delays and higher CPU loads. Figure 1 Schematic diagram of the principle of the conventional file system cache, as Figure 1As shown, in the file system caching technology, the application layer program performs read and write operations on files through the file system using the standard Portable Operating System Interface (Posix) interfaces (such as open, read, write, close, etc.). Whenever a process opens a file, a file descriptor file data structure is created in the file system. This data structure maintains the read and write offset values for the current process with respect to the file. Multiple processes can open the same file simultaneously, and a file data structure is generated in each process. The file entity operated by the process corresponds to a data structure of a unique index node (inode) in the file system. The inode data structure contains a file cache managed by a radix tree and a metadata cache. The file content is divided into pages of the same size for caching. The radix tree uses a file offset value as an index and can quickly find the file page cache block corresponding to the index. The schematic diagram of the file read and write operation is as follows, and the specific steps are as follows: 1. Allocate a file data page cache on the host side according to the read and write offset value and size of the logical file in the file data structure, and mount it into the radix tree. 2. Obtain the meta information of the file from the disk to determine the position of the file data page on the disk. 3. Obtain the corresponding file page from the disk and store it in the file system cache.

[0074] Corresponding to the Peer-to-Peer (P2P) transmission mode, data is directly transmitted from the storage device to heterogeneous devices. Although optimizing the transmission architecture improves the transmission efficiency, it bypasses the existing caching mechanism. However, the final result of data filtering through heterogeneous devices still needs to be obtained by the CPU accessing the heterogeneous devices. In either of the above ways, there will be CPU access to heterogeneous devices. Since the memory access performance of the CPU on the host side to the heterogeneous devices is different from that to its own memory on the host side, CPU access to heterogeneous devices results in relatively high latency and low bandwidth. The cache processing method for heterogeneous devices provided by the present invention can solve the above technical problems.

[0075] Figure 2 It is a flowchart of a cache processing method for a heterogeneous device provided by an embodiment of the present invention. As Figure 2 shown, the method includes:

[0076] S11: Obtain the access times of the target virtual addresses corresponding to each device in the heterogeneous system;

[0077] S12: Determine the file pages corresponding to the target heterogeneous device based on the access times of the target virtual addresses; and perform verification processing on the file pages;

[0078] S13: If the verification is passed, move the file page to the memory of the processor to facilitate the processor to access the file page of the target heterogeneous device.

[0079] Specifically, the heterogeneous system in this embodiment can be a conventional FPGA heterogeneous system. The storage device, heterogeneous device, and processor are connected through a switch corresponding to the Peripheral Component Interconnect Express (PCIE) bus, or in a point-to-point P2P transmission mode, the storage device, heterogeneous device, and processor are connected through a switch corresponding to the Compute Express Link (CXL) protocol. This is not limited here and can be set according to the actual situation. Each device in step S11 is a global device corresponding to the heterogeneous system, including a storage device, a heterogeneous device, and a processor on the host side. The heterogeneous device in this embodiment can be a single heterogeneous device or a heterogeneous device cluster composed of multiple heterogeneous devices, which is not limited here.

[0080] The acquisition basis here can be the access times of the target virtual addresses of each device based on the established database, or the access times of the target virtual addresses can be obtained from the virtual addresses of each device stored in the form of a list, table, or other forms. Either is acceptable and is not limited here.

[0081] The access times of the target virtual address can be the total access times corresponding to the target virtual address, specifically including the access times of external devices accessing local devices and / or the access times of local devices accessing themselves. The number of target virtual addresses can be one or multiple.

[0082] The acquisition here can be the access times with the highest access times of virtual addresses among all devices in the heterogeneous system, or the access times corresponding to the virtual addresses of all devices exceeding a certain threshold, etc. This is not limited here.

[0083] Determining the file page corresponding to the target heterogeneous device based on the access times of the target virtual address in step S12 can be determined based on the access times of the target virtual address itself, or can also be determined by combining other parameters on the basis of the access times of the target virtual address. This is not limited.

[0084] If it is determined based on the access count of the target virtual address itself, since the access count of the target virtual address can include the access count of an external device accessing a local device and / or the access count of the local device accessing itself, when only based on the access count of the external device accessing the local device, the virtual address corresponding to the access count exceeding a certain threshold can be used as the target virtual address, and the file page corresponding to the target virtual address can be directly obtained. When only based on the access count of the local device accessing itself, through the local access ratio or a certain threshold of the access count, if the access count is lower than the local access ratio or a certain threshold, it indicates that the access count of the external device accessing the local device is relatively large, and the corresponding virtual address needs to be used as the target virtual address.

[0085] Based on the access count of the target virtual address, if the number of target virtual addresses is multiple and corresponds to different file pages, considering the limited memory space of the processor, the access counts of all virtual addresses in different file pages can be summed up to obtain the sum of the access counts of each file page, and the file page corresponding to the highest sum of the access counts is selected as the final file page, etc.

[0086] Before moving the file page to the memory of the processor, considering the accuracy of the file page, it is necessary to perform a verification process on the file page. Here, the verification process can be performed on the data of the file page, or on the flag bit of the file, or a combined verification process of the two. The specific verification process is not limited here and can be set according to the actual situation. If the verification process fails, it can return to step S11 to re-obtain the access count of other target virtual addresses, or perform a correction process on the corresponding file page, etc., which is not limited here.

[0087] During the moving process in step S13, based on the cache mechanism of the original file system or a new file system cache mechanism can be established, and the page attribute information of the file page is updated to the cache mechanism, and the page information of the file page is stored in the memory of the processor. The cache mechanism corresponds to the storage of a data structure, and the data structure is not limited, which can be a tree structure or other data storage structures, etc.

[0088] It should be noted that based on frequent use, the file pages of the target heterogeneous device determined by the access times of the target virtual address are moved to the memory of the processor, and the cache mechanism of the corresponding processor adds the target heterogeneous device, so as to minimize the access of the processor to the heterogeneous device. The file pages corresponding to the heterogeneous device are stored in the memory of the processor. When the processor needs to access the file pages of the heterogeneous device, first check whether the file pages of the heterogeneous device are stored in the memory of the processor itself. If so, directly access the file pages of the heterogeneous device stored in its own memory without accessing the heterogeneous device, that is, move the file pages of the heterogeneous device to the memory of the processor in advance to reduce the access latency.

[0089] A cache processing method for a heterogeneous device provided by an embodiment of the present invention includes: obtaining the access times of the target virtual address corresponding to each device in a heterogeneous system; determining the file pages corresponding to the target heterogeneous device based on the access times of the target virtual address; and performing a verification process on the file pages; and in the case of passing the verification, moving the file pages to the memory of the processor to facilitate the processor to access the file pages of the target heterogeneous device. First, obtain the access times of the target virtual address corresponding to each device in the heterogeneous system. Here, considering the cache management of the heterogeneous device on the host side in the heterogeneous system, a global data cache mechanism is implemented to facilitate subsequent cache processing of the data of the heterogeneous device on the host side. Secondly, determine the file pages corresponding to the target heterogeneous device based on the access times of the target virtual address, so that the file system layer has the ability to manage the memory cache of the heterogeneous device. The determined target heterogeneous device is a target heterogeneous device predicted to be frequently used based on the access times of the target virtual address. Before moving the file pages of the target heterogeneous device to the memory of the processor, perform a verification process on the predicted file pages of the target heterogeneous device, and move them in the case of passing the verification to improve the allocation accuracy of the target heterogeneous device. Finally, map and move the file pages of the target heterogeneous device to the memory on the host side to improve the global access efficiency on the host side. When accessing the target heterogeneous device on the host side, there is no need to access the target heterogeneous device, and directly access the file pages stored in the processor memory, avoiding the increased access rate of the processor accessing the target heterogeneous device, reducing the latency of accessing the memory of the target heterogeneous device, and improving the data transmission rate and the system efficiency of the heterogeneous system.

[0090] In some embodiments, obtaining the access times of the target virtual address corresponding to each device in the heterogeneous system includes:

[0091] Pre - establish the global address space allocation relationship corresponding to each device in the heterogeneous system; among them, the global address space allocation relationship stores the address space access information corresponding to the processor and heterogeneous devices; the virtual address access times of the address space access information include the first virtual address access times when an external device accesses the local device and / or the second virtual address access times when the local device accesses itself.

[0092] Determine the target virtual address access times according to the global address space allocation relationship.

[0093] It should be noted that the global address space allocation relationship realizes unified addressing of the system memory through the heterogeneous cache coherence protocol and establishes an address mapping mechanism. The global address space allocation relationship can be in the form of a table, or a database, a set, etc., which is not limited here. It mainly stores the address space access information corresponding to the processor and heterogeneous devices. The virtual address access times of the address space access information here can be the first virtual address access times when an external device accesses the local device and the second virtual address access times when the local device accesses itself, or only the first virtual address access times when an external device accesses the local device, or only the second virtual address access times when the local device accesses itself, which is not limited here.

[0094] Determine the target virtual address access times according to the global address space allocation relationship. Since the global address space allocation relationship stores the virtual address access times corresponding to all devices in the current heterogeneous system, the virtual address with the most frequent access times here can be used as the target virtual address, and its corresponding access times as the target virtual address access times. For the screening of the most frequent access times, it can be to select the virtual address corresponding to the highest access times, or to select the virtual address corresponding to relatively high access times, which can be set according to the actual situation.

[0095] The method provided in this embodiment for determining the target virtual address access times based on the pre - established global address space allocation relationship, through the address mapping mechanism, allocates the address space and the corresponding cache of heterogeneous devices by allocating the global address space of heterogeneous devices, laying a foundation for the subsequent allocation process and improving the accuracy of allocation.

[0096] In some embodiments, copies of the global address space allocation relationship are set in each device in the heterogeneous system to facilitate the consistency of the address space access information of each device when the address of any device changes.

[0097] Specifically, there is a copy of the global address space allocation relationship on each device in the heterogeneous system to ensure that each device can understand the address allocation situation and relevant control information in real - time, and to transmit the heterogeneous cache coherence protocol to ensure data consistency.

[0098] In some embodiments, if any device has an address change, the method further includes:

[0099] receiving broadcast message delivery instructions;

[0100] The global address space allocation relationship corresponding to the address change is updated according to the broadcast message transmission instruction.

[0101] Specifically, if the address of any device changes, such as the file page corresponding to the virtual address of the heterogeneous device has been moved to the processor's memory, it is necessary to immediately initiate a broadcast message transmission instruction, such as a message corresponding to the write update mechanism. Coordination is performed between the devices in the heterogeneous system to keep the corresponding copies in the cache consistent. The hardware circuit in the heterogeneous device is used for decoding to accurately map the address in the unified address space to the corresponding actual storage module location, so that when a processing unit initiates an access request to a unified address, it can find the specific cache or main memory area where the corresponding storage location is located according to the mapping relationship.

[0102] This embodiment provides that when an address change occurs in any device, other devices are notified by broadcasting a message to pass instructions so that the copy information in the cache remains consistent, thereby ensuring data consistency.

[0103] In some embodiments, the address space access information of the global address space allocation relationship further includes device type, cluster structure type, on-chip memory channel, channel access times, current access ratio, virtual address, physical address, and read / write permission flag;

[0104] The process of determining the read and write permission mark includes:

[0105] Determine, according to a preset process, a first read-write permission mark of an access control register of a device corresponding to the heterogeneous system and a second read-write permission mark at a software level corresponding to the preset process;

[0106] If the first read-write permission mark and the second read-write permission mark are different, the final read-write permission mark is determined according to the first read-write permission mark.

[0107] Specifically, the main memory in each device is divided into different address segments, which correspond to a unified virtual address space. In addition to the virtual address and the record of the access times of the virtual address, it also includes the physical address mapped by the virtual address, device type, cluster structure type, on-chip memory channel, channel access times, this access ratio, and read / write permission flag. Table 1 is the global address space allocation relationship table. As shown in Table 1, the device type (represented by 2 bits, 2’b00 represents the CPU host device, 2’b01 represents the heterogeneous device, 2’10 represents the memory expansion device, 2’b11 is reserved), FPGA / CPU cluster identification number (Identification, ID) (represented by 4 bits, 4’h0 represents cluster 0, 4’h1 represents cluster 1, 4’h2 represents cluster 2, and so on), FPGA / CPU ID (represented by 4 bits, 4’h0 represents device 0, 4’h1 represents device 1, 4’h2 represents device 2, and so on), on-chip memory channel ID (represented by 4 bits, 4’h0 represents channel 0, 4’h1 represents channel 1, 4’h2 represents channel 2, and so on), starting address represented by 64 bits, ending address represented by 64 bits, channel access times (represented by 10 bits, unit MB / min), local access ratio (represented by 8 bits, unit %), update flag bit (represented by 1 bit, 1’b0 represents no update required, 1’b1 represents update required), permission flag bit (represented by 2 bits, 2’b00 represents no access permission, 2’b01 represents read-only, 2’b10 represents write-only, 2’b11 represents read-write).

[0108] Table 1

[0109]

[0110] The this access ratio is the number of accesses corresponding to the virtual address of the local device divided by the total number of accesses. Counting the access times of the virtual address can understand the different hot degrees of the virtual address, provide support for optimization operations such as prefetching, and improve the overall file access efficiency. In a system with the same addressing, the access permissions in different areas may be different, and overall consideration is needed to ensure the normal operation of the system and data security. For the unified address range where some shared data is located, multiple devices can dynamically coordinate read / write permissions according to the cache coherence protocol. It is also necessary to set the corresponding access control registers at the hardware level. The hardware priority is higher than the software dynamic adjustment priority. When the permissions do not match, access is refused and corresponding exception prompts may be generated, so as to ensure the orderly data access and operation of the entire heterogeneous system under the unified memory addressing. That is, when the read / write permission flag corresponding to the update has different read / write permission flags at the hardware level and the software level, the read / write permission flag at the hardware level should be the main one.

[0111] The virtual address for memory read and write access in the application is converted into a physical address (corresponding page number and offset within the page) by looking up Table 1. If the data accessed at the physical address is in the processor's memory (Cache), it is directly read from the processor's memory. If it is not in the processor's memory, the device memory space (DeviceMemory) is accessed based on the physical address through the host bridge in the root complex (Root Complex).

[0112] The multiple address access information of the global address space allocation relationship provided in this embodiment is used for subsequent auxiliary functions during the file page update process to ensure the normal operation of the system and data security.

[0113] In some embodiments, determining the target virtual address access count according to the global address space allocation relationship includes:

[0114] Obtaining a preset virtual address access count and / or a preset local access ratio;

[0115] Determining the local access ratio according to the first virtual address access count and the second virtual address access count;

[0116] In the global address space allocation relationship, select the virtual address access count that exceeds the preset virtual address access count and / or the local access ratio is less than the preset local access ratio, and use the virtual address access count corresponding to the virtual address access count that exceeds the preset virtual address access count and / or the local access ratio is less than the preset local access ratio as the target virtual address access count.

[0117] Specifically, when determining the target virtual address access count according to the global address space allocation relationship, considering that the virtual address access count can be the first virtual address access count and / or the second virtual address access count, a preset virtual address access count and / or a preset local access ratio can be set. Corresponding to the preset virtual address access count, it is compared with the first virtual address access count. If the first virtual address access count exceeds the preset virtual address access count, it is used as the target virtual address. Corresponding to the preset local access ratio, the local access ratio corresponding to each virtual address needs to be calculated (the second virtual address access count / the first virtual address access count). If the actual local access ratio is less than the preset local access ratio, it means that the local access is less and the external device access is more, which indirectly proves that the external device access count is more.

[0118] In this embodiment, through the relationship that the virtual address access count is the first virtual address access count and / or the second virtual address access count, the corresponding comparison situation also uses the relationship of and / or to determine the final target virtual address access count.

[0119] The number of corresponding target virtual addresses here can be one or multiple, which is not limited herein.

[0120] In this embodiment, the virtual address access times corresponding to the virtual addresses with the virtual address access times exceeding the preset virtual address access times and / or the local access ratio less than the preset local access ratio are selected as the target virtual address access times from the global address space allocation relationship, improving the diversity and flexibility of determination to facilitate the subsequent accuracy of determining the target heterogeneous device.

[0121] In some embodiments, when the virtual address access times include the first virtual address access times or the second virtual address access times, and the number of virtual addresses corresponding to the target virtual address access times is one, determining the file page corresponding to the target heterogeneous device based on the target virtual address access times includes:

[0122] Taking the virtual address corresponding to the target virtual address access times as the first target virtual address;

[0123] Taking the heterogeneous device to which the data block corresponding to the target virtual address belongs as the target heterogeneous device to obtain the file page of the data block corresponding to the target heterogeneous device.

[0124] Specifically, through the previous embodiment, in the case where the virtual address access times include the first virtual address access times or the second virtual address access times, and the number of virtual addresses corresponding to the target virtual address access times is one, it takes this virtual address as the first target virtual address, and the heterogeneous device to which the corresponding data block belongs is the target heterogeneous device, so that its file page can be obtained. A data block usually refers to the basic unit for storing data in a file system. It may correspond to a file page or a smaller storage unit. In a file system, a data block is the logical unit for allocating files to store data. It may contain part or all of a file page. A file page is read from the data block in the file system.

[0125] The method for determining the file page corresponding to the target heterogeneous device provided in this embodiment is determined only by the target virtual address corresponding to the target virtual address access times, improving the rapidity and simplicity of determining the target heterogeneous device.

[0126] In other embodiments, when the virtual address access times include the first virtual address access times or the second virtual address access times, and the number of virtual addresses corresponding to the target virtual address access times is multiple, determining the file page corresponding to the target heterogeneous device based on the target virtual address access times includes:

[0127] Taking the virtual addresses corresponding to the target virtual address access times as the second target virtual addresses;

[0128] When the number of second target virtual addresses is multiple, determine whether the file pages to which the second target virtual addresses belong are the same;

[0129] If they are the same, use the heterogeneous device to which the data block corresponding to the second target virtual address belongs as the target heterogeneous device to obtain the file page of the data block corresponding to the target heterogeneous device;

[0130] If they are not the same, add up the access times of the virtual addresses stored in the file pages to which the second target virtual addresses belong to obtain the sum of their respective access times;

[0131] Use the file page to which the access time of the virtual address with the largest sum of access times belongs as the file page corresponding to the target heterogeneous device.

[0132] Specifically, when there are multiple virtual addresses corresponding to the access time of the target virtual address, since what is ultimately moved is the file page of the target heterogeneous device and a file page contains multiple virtual addresses, considering the limited memory space of the processor memory, it is necessary to make full use of the memory space of the processor memory. In the case of multiple file pages, further screening is required.

[0133] Use the virtual addresses corresponding to the access time of the target virtual address as the second target virtual addresses. When the number of second target virtual addresses is multiple, it is necessary to determine whether the file pages to which the second target virtual addresses belong are the same. If so, it is the same as the method for determining the file page in the above embodiment, and the above embodiment can be referred to. If they are not the same file pages, it is necessary to add up the access times of all the virtual addresses of each file page to obtain the sum of their respective access times, and select the file page corresponding to the largest sum of the access times of the virtual addresses as the final file page.

[0134] In this embodiment, in the case where multiple file pages exist, it is necessary to add up the access times of all the virtual addresses corresponding to each of the multiple file pages respectively, obtain the file page corresponding to the largest access time for subsequent relocation, and improve the efficiency of cross-device memory access on the basis of making full use of the memory space of the processor memory.

[0135] In some embodiments, performing a verification process on the file page includes:

[0136] Obtain the data block to which the file page belongs and the corresponding page information;

[0137] Perform a verification process on each flag bit and page information of the data block to obtain a verification result.

[0138] It is understandable that before actually moving the file pages of the target heterogeneous device, it is necessary to check whether the file pages are correct, which corresponds to the verification process of each flag bit of the data block to ensure the integrity of the data block. The verification process of the page information is to ensure whether the data in the data block has been updated or modified, improving the security of the data.

[0139] The specific process of the verification process corresponding to each flag block and page data here is not limited and can be set according to the actual situation.

[0140] The verification process corresponding to each flag bit of the data block provided in this embodiment is to ensure the integrity of the data block. The verification process of the page information is to ensure whether the data in the data block has been updated or modified, improving the security of the data.

[0141] In some embodiments, the verification process of each flag bit of the data block and the page information to obtain a verification result includes:

[0142] The verification process of each flag bit of the data block to obtain a first verification result;

[0143] In the case where the first verification result is that all verifications pass, the verification process of the page information is performed according to the file page hash value to obtain a second verification result;

[0144] If the second verification result is that the verification passes, then the final verification result is that the verification passes.

[0145] It should be noted that the verification process of each flag bit and the page information here can be that after one factor (each flag bit or page information) is verified and the verification passes, the other factor (page information or each flag bit) is verified. It can also be that the two factors are verified separately and the verification results are finally summarized. Only when the verification results of both factors are passed, the final verification result is passed, so as to facilitate the moving work of the file pages.

[0146] In the verification process of the first method, if the first factor fails to pass the verification, there is no need to perform the subsequent verification of the second factor.

[0147] In this embodiment, first, one factor is verified, and in the case where the verification passes, the other factor is verified to improve the verification accuracy and efficiency of the file pages.

[0148] In some embodiments, moving the file page to the memory of the processor includes:

[0149] Obtain the cache mechanism corresponding to the processor in the heterogeneous system; among them, the cache mechanism is established by the file system layer of the file system;

[0150] Store the page attribute information corresponding to the file page into the target data structure corresponding to the cache mechanism;

[0151] Move the page information corresponding to the file page to the memory of the processor.

[0152] It can be understood that during the process of moving the file page, it is necessary to store the page attribute information corresponding to the file page into the target data structure of the cache mechanism to complete the update to the cache mechanism of the processor, so as to facilitate subsequent access and call. The page attribute information mainly includes the integrity of the data block of the file page, whether the data in the data block has been modified, and information corresponding to virtual addresses and physical addresses, etc., so as to manage and schedule the file cache.

[0153] Figure 3 A schematic structural diagram of a heterogeneous system provided by an embodiment of the present invention is shown in Figure 3 As shown, a heterogeneous consistency view is established among the hard disk storage device, the heterogeneous device cluster, and the storage device through a switch. Corresponding to the inside of the processor, it is managed through the global data cache management module and the device memory host-side cache management module in the cache mechanism of the file system layer. The global data cache management module realizes the global address space allocation relationship corresponding to the same addressing through address mapping. The device memory host-side cache management module realizes the process of moving the file page.

[0154] In some embodiments, the process of establishing the cache mechanism includes:

[0155] Obtain the first file system layer of the file system; wherein, the first file system layer includes a file system standard interface and metadata information;

[0156] On the basis of the first file system layer, add a second file system layer; wherein, the second file system layer follows the file system standard interface and metadata information of the first file system layer, and the second file system layer contains a target data structure to maintain the read / write offset value of the target process for the file;

[0157] Based on the first file system layer and the second file system layer, determine the cache mechanism.

[0158] Specifically, as shown in Figure 3As shown in the figure, a custom stacked file system layer in the kernel state, i.e., the second file system layer, is created on the original Virtual File System (VFS) layer, i.e., the original file system layer, and read and write operations on files are performed in accordance with the standard Posix interface. In the VFS, each file corresponds to an inode node with a unique number, and each inode node corresponds to a file cache managed by a tree structure. When data read and write operations are performed, a file cache page of a fixed size is allocated from the host-side memory. The VFS starts the DMA to move data blocks from the disk to the file cache and mounts it on the radix tree for maintenance. According to the offset value of file reading, the corresponding cache block is found in the radix tree for reading and writing. When each process opens a file, a File data structure is created in the VFS, and the read and write offset values of the current process for this file are maintained in this data structure. In the second file system layer, by overloading the file page cache policy managed by the tree structure for inode allocation of the file, the memory allocation and transmission between the host side and the FPGA side are added. In the conventional technology, the file system only manages the host-side memory, and all contents of the file need to be cached on the host side. Through the heterogeneous cache coherence protocol, the system memory is uniformly addressed (global address space allocation relationship), and information about the FPGA-side memory is added to the file cache on the host side and managed using a radix tree.

[0159] It should be noted that the creation of its cache mechanism can also be based on the setting of a brand-new file system layer, which is not limited here.

[0160] In this embodiment, a second file system layer is created on the original file system layer, and its cache mechanism is created here. The memory allocation and transmission of heterogeneous devices are loaded into the memory of the processor, which is convenient for the processor to access and improves the access efficiency.

[0161] Figure 4 This is a schematic diagram of the principle of a cache mechanism provided by an embodiment of the present invention. As Figure 4 shown, corresponding to storing page attribute information into the target data structure in the cache mechanism within the processor, the corresponding target data structure can be a tree data structure or other data structures, which is not limited here. It should be noted that the cache mapping relationship of the memory of heterogeneous devices on the host side is essentially to move the file pages of the target heterogeneous device to the memory of the processor, reducing the access efficiency of the processor to heterogeneous devices.

[0162] In some embodiments, storing the page attribute information corresponding to the file page into the target data structure corresponding to the cache mechanism includes:

[0163] Obtaining the target virtual address, target physical address, and target virtual address access count corresponding to the file page;

[0164] Save the target virtual address and the target physical address in the form of key-value pairs into the target data structure; wherein, the target data structure at least includes a key, a value, a child node pointer, a flag bit, a file page hash value, and count information;

[0165] Save the page access count to which the target virtual address access count belongs into the count information of the target data structure;

[0166] Save the verification results of the flag bits of each data block corresponding to the file page into the flag bits of the target data structure;

[0167] Save the verification result of the page information corresponding to the file page into the file page hash value of the target data structure.

[0168] Specifically, Table 2 is a composition table of the target data structure, including a key, a value, a child node pointer, a flag bit, a file page hash value, and count information. At the same time, it also includes attached control information.

[0169] Table 2

[0170]

[0171] As shown in Table 2, its key is the logical address (virtual address) of the file used, which contains the file page number and the page offset. The method uses the corresponding file page number as the key. In this way, no matter how the specific storage location of the file in the system changes, as long as the logical address remains unchanged, the corresponding cache node (the page corresponding to the virtual address) can be found through the page number. The value is the location information of the specific data block storing the file page in the system, recorded in the form of a pointer. Each pointer points to the specific node storing the file page (which can be nodes of different devices such as CPU memory nodes, FPGA memory nodes, etc.), and at the same time contains information such as the size and offset of the valid data of each page, which is convenient for assembling and restoring the data. Child Pointers: A tree structure is constructed by dividing branches according to different bits of the page number. Each 4-bit binary bit corresponds to a branch direction, and there can be a total of 16 branch choices. The tree structure is constructed by continuously subdividing the path according to 4-bit binary bits, enabling 16 branches to be processed at a time, which is convenient for quick searching and comparison. Flags: "Integrity Flag", which is used to verify whether the cache is complete and undamaged. This flag is updated by periodically performing checks such as Cyclic Redundancy Check (CRC). Of course, other verification processing methods can also be used, and no specific limitation is made here. If the flag indicates that the data block is incomplete, the system will trigger the re-acquisition of the corresponding data block from other backup nodes to repair the cache content. File Page Hash Value: A unique hash value generated by applying a hash algorithm (such as SHA256, etc.) to the file page content. This is convenient for quickly determining whether different files have been modified. As long as the page is modified, the hash value will change. Count Information: "Page Access Frequency", which counts the number of times each file page is accessed. Based on this statistic, the access popularity can be understood. For hot data blocks, optimization operations such as migrating to device nodes with better performance or prefetching can be considered to improve the overall file access efficiency.

[0172] It should be noted that in the present invention, a file page cache of the target data structure is allocated for the inode of each file, and the cache includes both the host-side and heterogeneous device-side memories. Figure 5 This is a schematic diagram of the principle of another cache mechanism provided by an embodiment of the present invention. As Figure 5 shown, for this child pointer, each 4-bit binary corresponds to a branch direction, and 16 branch choices can be realized, so as to speed up the traversal and search for data reading corresponding to different heterogeneous devices.

[0173] In addition, regarding the attached control information, it can be the total size of the file, the creation time, the last access time, the valid flag bit, etc., which is convenient for the system to manage and schedule the file cache as a whole.

[0174] It is understandable that the target virtual address and the target physical address are saved in the form of key-value pairs in the target data structure, and the target physical address can be obtained through Table 1 in the above embodiments. The page access count to which the target virtual address access count belongs is saved in the count information, and the verification results of each flag bit are saved in the flag bits of the target data structure. The verification result of the page information is saved in the file page hash value.

[0175] In this embodiment, storing the page attribute information corresponding to the file page in the target data structure corresponding to the cache mechanism solves the cache management problem of heterogeneous devices in the file system layer by improving the target data structure, thereby improving the access efficiency of the processor.

[0176] Corresponding to the page information migration process of the file page, the page information can be migrated based on the processor's updated memory cycle and the preset window, improving the efficiency of data migration.

[0177] In some embodiments, migrating the page information corresponding to the file page to the memory of the processor includes:

[0178] Storing the page information of the file page in the target data structure corresponding to the cache mechanism according to the processor's updated memory cycle and the preset window.

[0179] Specifically, different from the host-side memory, when the CPU host side accesses the memory of the heterogeneous device, the latency is high and the bandwidth is low. Frequent access to the FPGA heterogeneous device memory will cause the system efficiency to decrease. This embodiment enables the FPGA memory to have partial cache on the host side to improve the efficiency of cross-device memory access, thereby enhancing the global memory access efficiency of the host side. The updated memory cycle performs subsequent update processing based on the existing stored data in the processor memory, and deletes the data with a long storage cycle and infrequent use at regular intervals to save more memory space for storing the page information of the file page. The preset window is a cache mechanism that realizes data cache management based on a serial port of a specific size.

[0180] In this embodiment, migrating the file page to the memory of the processor, storing the page attribute information in the target data structure, and migrating the page information to the memory of the processor realize the cache of the device memory host side and the improvement and update of the data structure, improving the global access efficiency of the host side.

[0181] In some embodiments, the process of determining the preset window includes:

[0182] Obtaining the transmission system level corresponding to the file page;

[0183] Obtaining the data transmission bandwidth corresponding to the access between the processor and the target heterogeneous device;

[0184] Determine a preset window according to the data transmission bandwidth, the memory space data of heterogeneous devices, the file page information corresponding to the preset window, and the transmission system level.

[0185] Specifically, judge the access heat by the number of "page access frequencies" in the cache management tree node, and decide whether to map the FPGA memory to the host side for caching based on the access history of the data. Since the data that has been frequently accessed in the recent period has a high probability of being accessed in the future, it is necessary to map the FPGA cache to the host side, and the default access heat is greater than 10,000 times / minute.

[0186] Perform effective prefetching through the window cache mechanism to improve cache performance. For the acceleration computing unit, the window cache mechanism is adopted. The window cache realizes data cache management based on a "window" of a specific size. The default window size is 64 file pages (default 4KB). At the same time, the window size can usually be set according to characteristics such as the data transmission bandwidth and the transmission system level (real-time level, such as level 1 for normal, level 2 for general real-time, and level 3 for high real-time). The transmission system level here can be set in real time based on the application scenario.

[0187] Window size (number of pages) = (data transmission bandwidth / 8GB / s) * 64 * system real-time level.

[0188] If the transmission bandwidth is high and the data real-time requirement is high, the window is set larger to allow more data to wait for processing in the cache. Otherwise, it should be set smaller to avoid problems such as cache overflow at the receiving end. When new data is continuously generated and transmitted, it will enter the tree-shaped cache management structure designed above in sequence. Figure 6 A schematic diagram of a cache window provided for an embodiment of the present invention is shown as Figure 6 shown. Based on request a, set corresponding cache windows for different acceleration computing units, and each cache window enters in sequence.

[0189] The window cache mechanism provided in this embodiment performs effective prefetching, improves cache performance, and prevents cache overflow at the receiving end.

[0190] In some embodiments, the process of determining the updated memory cycle includes:

[0191] Obtain the initial updated memory cycle, the number of accesses corresponding to the file page, and the transmission system level;

[0192] Determine the updated memory cycle according to the initial updated memory cycle, the number of accesses corresponding to the file page, and the transmission system level.

[0193] Specifically, the cache data update adopts a fixed-duration caching mechanism, and manages the validity period of data by setting a timer or a timestamp mechanism. When data is first mapped to the host side and stored in the cache, a timestamp (representing the storage time) will be synchronously recorded, and a fixed duration is set as the validity period, with a default duration of 30 seconds. The setting of the fixed duration is determined based on the data access frequency, the transmission system level corresponding to the target process (real-time level, such as level 1 for normal, level 2 for general real-time, and level 3 for high real-time), and whether there is a data consistency requirement.

[0194] Fixed duration (s) = ((data access frequency / 10000) * system real-time level * 30).

[0195] For high data access frequency, high data update rate, and high data real-time level requirements, the cache has a long valid time. Conversely, the cache valid time should be shorter. For each access request to the cache data, the system checks whether the difference between the current time and the stored timestamp exceeds the set fixed duration. If not, the data is considered still valid. By calculating the SHA256 hash value of the file page, if it is equal to the file page hash value, the data can be directly obtained and returned from the cache. If it times out or does not match the file page hash value, it is determined that the cache data has expired, and the expired data processing triggers a write-back and invalidation operation to release the cache space. Specifically, the system invalidates the mapped cache through the CXL heterogeneous cache coherence protocol message to ensure data consistency.

[0196] The fixed-duration caching mechanism provided in this embodiment performs effective prefetching, saves memory space, and improves the effectiveness of data caching.

[0197] Furthermore, the present invention also provides a heterogeneous system, including a storage device, heterogeneous devices, a switch, and a processor; wherein, at least one heterogeneous device constitutes a heterogeneous device cluster;

[0198] The storage device, the heterogeneous device cluster, and the processor all achieve heterogeneous consistency through the protocol of the switch; wherein, the processor is used to implement the steps of the cache processing method of the heterogeneous device when executing a computer program.

[0199] It should be noted that based on the heterogeneous system in this embodiment, it can be a conventional heterogeneous system as shown in Figure 1 , or a heterogeneous system under the P2P transmission mode. If it is the latter, the heterogeneous consistency of the corresponding switch is implemented based on the CXL protocol.

[0200] For the introduction of the heterogeneous system provided by the present invention, please refer to the above method embodiment, and the present invention will not be elaborated here. It has the same beneficial effects as the cache processing method of the above heterogeneous device.

[0201] The above has described in detail each embodiment corresponding to the cache processing method of heterogeneous devices. On this basis, the present invention also discloses a cache processing device for heterogeneous devices corresponding to the above method. Figure 7 It is a structural diagram of a cache processing device for heterogeneous devices provided by an embodiment of the present invention. As Figure 7 shown, the device includes:

[0202] An acquisition module 11, configured to acquire the access times of target virtual addresses corresponding to each device in a heterogeneous system;

[0203] A determination module 12, configured to determine the file page corresponding to the target heterogeneous device based on the access times of the target virtual address; and perform verification processing on the file page;

[0204] A relocation module 13, configured to relocate the file page to the memory of the processor when the verification passes, so as to facilitate the processor to access the file page of the target heterogeneous device.

[0205] Since the embodiments of the device part correspond to the above embodiments, the embodiments of the device part are described with reference to the embodiments of the above method part and will not be elaborated here.

[0206] For the introduction of a cache processing device for heterogeneous devices provided by the present invention, please refer to the above method embodiments. The present invention will not elaborate here, and it has the same beneficial effects as the above cache processing method for heterogeneous devices.

[0207] Figure 8 It is a structural diagram of a cache processing device for heterogeneous devices provided by an embodiment of the present invention. As Figure 8 shown, the device includes:

[0208] A memory 21, configured to store a computer program;

[0209] A processor 22, configured to implement the steps of the cache processing method for heterogeneous devices when executing the computer program.

[0210] The cache processing device for heterogeneous devices provided in this embodiment may include, but is not limited to, a smart phone, a tablet computer, a notebook computer, or a desktop computer, etc.

[0211] Among them, the processor 22 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 22 may be implemented in at least one hardware form of a digital signal processor (DSP), an FPGA, or a programmable logic array. The processor 22 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 22 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 22 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.

[0212] The memory 21 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 21 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 21 is at least used to store the following computer program 211. After the computer program is loaded and executed by the processor 22, it can implement the relevant steps of the cache processing method of the heterogeneous device disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 21 may also include an operating system 212 and data 213, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 212 may include Windows, Unix, Linux, etc. The data 213 may include, but is not limited to, the data involved in the cache processing method of the heterogeneous device, etc.

[0213] In some embodiments, the cache processing device of the heterogeneous device may further include a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27.

[0214] Those skilled in the art can understand that Figure 8 the structure shown in

[0215] does not constitute a limitation on the cache processing device of the heterogeneous device, and may include more or fewer components than those shown in the figure.

[0216] For the introduction of a cache processing device for heterogeneous devices provided by the present invention, please refer to the above method embodiments. The present invention will not be elaborated herein again, and it has the same beneficial effects as the above cache processing method for heterogeneous devices.

[0217] Furthermore, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor 22, the steps of the cache processing method for heterogeneous devices as described above are implemented.

[0218] It can be understood that if the method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0219] For the introduction of a computer-readable storage medium provided by the present invention, please refer to the above method embodiments. The present invention will not be elaborated herein again, and it has the same beneficial effects as the above cache processing method for heterogeneous devices.

[0220] Further, the present invention also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the cache processing method for heterogeneous devices are implemented.

[0221] For the introduction of a computer program product provided by the present invention, please refer to the above method embodiments. The present invention will not be elaborated herein again, and it has the same beneficial effects as the above cache processing method for heterogeneous devices.

[0222] The above has introduced in detail a cache processing method, a heterogeneous system, a product, a device and a medium for a heterogeneous device provided by the present invention. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description of the method part for the relevant parts. It should be noted that for those of ordinary skill in the art in the technical field of the present invention, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

[0223] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including an..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

Claims

1. A cache processing method for heterogeneous devices, characterized in that Including: Obtaining the access times of the target virtual addresses corresponding to each device in the heterogeneous system; Determining the file pages corresponding to the target heterogeneous device based on the access times of the target virtual addresses; And performing verification processing on the file pages; When the verification is passed, moving the file pages to the memory of the processor so that the processor can access the file pages of the target heterogeneous device stored in its own memory; wherein, the moving process is set based on the file system layer; Correspondingly, obtaining the access times of the target virtual addresses corresponding to each device in the heterogeneous system includes: Pre - establishing the global address space allocation relationship corresponding to each device in the heterogeneous system; wherein, the global address space allocation relationship stores the address space access information corresponding to the processor and the heterogeneous device; the virtual address access times of the address space access information include the first virtual address access times of the external device accessing the local device and / or the second virtual address access times of the local device accessing itself; Determining the access times of the target virtual addresses according to the global address space allocation relationship; Correspondingly, determining the access times of the target virtual addresses according to the global address space allocation relationship includes: Obtaining the preset virtual address access times and / or the preset local access ratio; Determining the local access ratio according to the first virtual address access times and the second virtual address access times; Selecting in the global address space allocation relationship the virtual address access times that exceed the preset virtual address access times and / or the local access ratio that is less than the preset local access ratio, and taking the virtual address access times that exceed the preset virtual address access times and / or the local access ratio that is less than the preset local access ratio as the access times of the target virtual addresses.

2. The cache processing method for heterogeneous devices according to claim 1, wherein When the virtual address access times include the first virtual address access times or the second virtual address access times, and there are multiple virtual addresses corresponding to the access times of the target virtual addresses, determining the file pages corresponding to the target heterogeneous device based on the access times of the target virtual addresses includes: Taking the virtual addresses corresponding to the access times of the target virtual addresses as the second target virtual addresses; When the number of the second target virtual addresses is multiple, determining whether the file pages to which each of the second target virtual addresses belongs are the same; If they are the same, taking the heterogeneous device to which the data block corresponding to the second target virtual address belongs as the target heterogeneous device to obtain the file pages of the target heterogeneous device corresponding to the data block; If they are not the same, adding up the access times of the virtual addresses stored in the file pages to which each of the second target virtual addresses belongs to obtain the sum of the respective access times; Taking the file page to which the access times of the virtual address with the largest sum of access times belong as the file page corresponding to the target heterogeneous device.

3. The cache processing method for heterogeneous devices according to claim 1, wherein Performing verification processing on the file pages includes: Obtaining the data block to which the file page belongs and the corresponding page information; Performing verification processing on each flag bit of the data block and the page information to obtain a verification result.

4. The cache processing method for heterogeneous devices according to claim 3, wherein Performing a verification process on each flag bit of the data block and the page information to obtain a verification result, including: Performing a verification process on each flag bit of the data block to obtain a first verification result; When the first verification result is that all verifications pass, performing a verification process on the page information according to the file page hash value to obtain a second verification result; If the second verification result is that the verification passes, the final verification result is that the verification passes.

5. The cache processing method for heterogeneous devices according to claim 4, wherein Moving the file page to the memory of the processor, including: Obtaining the cache mechanism corresponding to the processor in the heterogeneous system; wherein, the cache mechanism is established by the file system layer of the file system; Storing the page attribute information corresponding to the file page into the target data structure corresponding to the cache mechanism; Moving the page information corresponding to the file page to the memory of the processor.

6. The cache processing method for heterogeneous devices according to claim 5, wherein Storing the page attribute information corresponding to the file page into the target data structure corresponding to the cache mechanism, including: Obtaining the target virtual address, target physical address, and the number of accesses to the target virtual address corresponding to the file page; Saving the target virtual address and the target physical address in the form of a key-value pair into the target data structure; wherein, the target data structure at least includes a key, a value, a child node pointer, a flag bit, a file page hash value, and count information; Saving the page access count to which the number of accesses to the target virtual address belongs into the count information of the target data structure; Saving the verification results of each of the flag bits of the data block corresponding to the file page into the flag bit of the target data structure; Saving the verification result of the page information corresponding to the file page into the file page hash value of the target data structure.

7. The caching processing method for heterogeneous devices according to claim 5, characterized in that Moving the page information corresponding to the file page to the memory of the processor, including: Storing the page information of the file page into the target data structure corresponding to the cache mechanism according to the updated memory cycle of the processor and a preset window.

8. The caching processing method for heterogeneous devices according to claim 7, wherein The determining process of the preset window includes: Obtaining the transmission system level corresponding to the file page; Obtaining the data transmission bandwidth corresponding to the access between the processor and the target heterogeneous device; Determining the preset window according to the data transmission bandwidth, the memory space data of the heterogeneous device, the file page information corresponding to the preset window, and the transmission system level.

9. The cache processing method for heterogeneous devices according to claim 7, wherein The determining process of the updated memory cycle includes: Obtaining an initial updated memory cycle, the number of accesses corresponding to the file page, and the transmission system level; Determining the updated memory cycle according to the initial updated memory cycle, the number of accesses corresponding to the file page, and the transmission system level.

10. A heterogeneous system, characterized in that, Including a storage device, a heterogeneous device, a switch, and a processor; wherein, at least one heterogeneous device forms a heterogeneous device cluster; The storage device, the heterogeneous device cluster, and the processor all achieve heterogeneous consistency through the protocol of the switch; wherein, the processor is used to implement the steps of the cache processing method for the heterogeneous device as described in any one of claims 1 to 9 when executing a computer program.

11. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the cache processing method for the heterogeneous device according to any one of claims 1 to 9 are implemented.

12. A cache processing device for heterogeneous devices, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the cache processing method for the heterogeneous device according to any one of claims 1 to 9 when executing the computer program.

13. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the cache processing method for the heterogeneous device according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Big data processing server system

    CN111651489A

  • Memory access optimization method and device, equipment, medium and program product

    CN118051189A