Page migration method and device in hierarchical memory, storage medium and program product

By using the concurrent replication function of the hardware accelerator in the processor for page migration, the problems of large CPU overhead and low migration efficiency in the prior art are solved, and more efficient memory resource management is achieved.

CN120010789AInactive Publication Date: 2025-05-16ALIBABA CLOUD COMPUTING CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510462273.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the existing hierarchical memory system is migrated on pages, it is completed by the central processor, resulting in large CPU overhead and low data migration efficiency.

Method used

By setting up the page concurrent copy function implemented by the driver based on the hardware accelerator in the processor, the access popularity of multiple pages is determined, and the hardware accelerator is triggered to migrate the page to be migrated to the corresponding target memory layer by calling this function.

Benefits of technology

Reduces processor overhead, improves page migration efficiency, and optimizes the configuration of memory resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010789A_ABST
    Figure CN120010789A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a page migration method and device in a hierarchical memory, a storage medium and a program product, an electronic device is provided with a page concurrent copying function achieved based on a drive program of a hardware accelerator, and the method comprises the steps that a processor determines the access popularity of multiple pages, the pages to be migrated in the multiple pages are determined according to the access heat, different pages to be migrated correspond to different memory layers, a hardware accelerator is triggered by calling a page concurrent copying function to migrate the pages to be migrated to a corresponding target memory layer, and a target cold page is migrated to a slow memory layer. Through the scheme, the page migration efficiency can be improved, and the overhead of a CPU (Central Processing Unit) is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, storage medium and program product for migrating pages in a hierarchical memory. Background Art

[0002] With the rapid development of data-intensive applications (such as machine learning, graph processing, and big data analysis), memory requirements have increased dramatically. Dynamic Random Access Memory (DRAM) is expensive and has limited capacity, but has fast access speeds. While emerging non-volatile memory (NVM) and compute express link memory (CXL) provide larger capacity and lower costs, they have higher latency, which may lead to performance degradation if data is not placed properly. Therefore, memory tiering technology has emerged. Memory tiering technology is a memory management technology that divides memory into multiple levels to form a hierarchical memory system. Each memory layer has different speeds and access rights, which can improve memory performance and efficiency.

[0003] The main purpose of designing a hierarchical memory system is to place appropriate data (memory pages) in the appropriate memory layer to fully utilize the advantages of different memory layers. This involves the migration of pages between different memory layers.

[0004] In existing hierarchical memory systems, page migration is performed by the central processing unit (CPU), which faces problems such as high CPU overhead and low data migration efficiency. Summary of the invention

[0005] The embodiments of the present application provide a method, device, storage medium and program product for migrating pages in a hierarchical memory, so as to improve page migration efficiency and reduce processor overhead.

[0006] In a first aspect, an embodiment of the present application provides a method for migrating pages in a hierarchical memory, which is applied to a processor in an electronic device, wherein the electronic device is provided with a concurrent page copy function implemented by a driver based on a hardware accelerator, and the method includes: Determine the popularity of multiple pages; Determine a page to be migrated among the multiple pages according to the access popularity, where different pages to be migrated correspond to different memory layers; The hardware accelerator is triggered by calling the page concurrent copy function to migrate the to-be-migrated page to the corresponding target memory layer.

[0007] In a second aspect, an embodiment of the present application provides a device for migrating pages in a hierarchical memory, which is applied to a processor in an electronic device, wherein the electronic device is provided with a concurrent page copy function implemented by a driver based on a hardware accelerator, and the device includes: A heat determination module is used to determine the access heat of multiple pages; A page classification module, used for determining a page to be migrated among the multiple pages according to the access popularity, where different pages to be migrated correspond to different memory layers; The migration processing module is used to trigger the hardware accelerator to migrate the to-be-migrated page to the corresponding target memory layer by calling the concurrent page replication function.

[0008] In a third aspect, an embodiment of the present application provides a method for migrating pages in a hierarchical memory, which is applied to a hardware accelerator in an electronic device, wherein the electronic device also includes a processor and a concurrent page copy function implemented by a driver based on the hardware accelerator, and the method includes: Obtaining a page to be migrated determined by the processor from a plurality of pages, where different pages to be migrated correspond to different memory layers; wherein the processor determines the page to be migrated according to the access heat of the plurality of pages; Based on the page concurrent replication function, the to-be-migrated page is migrated to the corresponding target memory layer.

[0009] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the page migration method in the hierarchical memory as described in the first aspect or the third aspect.

[0010] In a fifth aspect, an embodiment of the present application provides a non-temporary machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the page migration method in the hierarchical memory as described in the first aspect or the third aspect.

[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it can implement the page migration method in the hierarchical memory as described in the first aspect or the third aspect.

[0012] In the page migration scheme in the hierarchical memory provided in the embodiment of the present application, the hierarchical memory includes memory layers of different levels such as a fast memory layer and a slow memory layer, and the allocated pages (page) need to be migrated between different memory layers according to their access heat. In order to reduce the overhead of the processor (CPU), a driver based on a hardware accelerator implements a page concurrent copy function, so that when performing a page migration operation, the CPU can hand over the task to the hardware accelerator to improve the efficiency of page migration and reduce the CPU overhead. Specifically, the CPU determines the access heat of multiple pages, determines the pages to be migrated among the multiple pages according to the access heat, and then triggers the hardware accelerator to migrate different pages to be migrated to the corresponding memory layer by calling the page concurrent copy function. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0014] Figure 1 A schematic diagram of a page migration system provided in an embodiment of the present application; Figure 2 A flowchart of a method for migrating pages in a hierarchical memory provided in an embodiment of the present application; Figure 3 A schematic diagram of a page concurrent copying process provided in an embodiment of the present application; Figure 4 A flowchart of a method for migrating pages in a hierarchical memory provided in an embodiment of the present application; Figure 5 A flowchart of a method for migrating pages in a hierarchical memory provided in an embodiment of the present application; Figure 6 A flowchart of a method for migrating pages in a hierarchical memory provided in an embodiment of the present application; Figure 7 A flowchart of a method for migrating pages in a hierarchical memory provided in an embodiment of the present application; Figure 8 A flowchart of a method for migrating pages in a hierarchical memory provided in an embodiment of the present application; Fig. 9 A schematic diagram of the structure of a hierarchical memory page migration device provided in an embodiment of the present application; Fig.10 A schematic diagram of the structure of an electronic device provided in this embodiment. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. In addition, the step timing in the following method embodiments is only an example, not a strict limitation.

[0016] It should be noted that, in the case of user information involved in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to large language models or other models) are in compliance with relevant laws and standards.

[0017] First, the terms or concepts involved in the embodiments of the present application are explained: Tiered Memory: A memory hierarchy that improves memory performance and efficiency by dividing memory into multiple levels, each with different speeds and access permissions.

[0018] Data Streaming Accelerator (DSA) is an on-chip accelerator that mainly accelerates operations such as memory copy, comparison, and checksum generation.

[0019] Asynchronous replication: A memory replication technology that asynchronously copies data from one memory location to another, which can improve memory performance and efficiency.

[0020] Batch Size: refers to the number of operations performed in a batch, usually used to optimize the performance of data processing and transmission.

[0021] A tiered memory system is a technology for expanding the capacity of a memory with limited main memory capacity by using persistent memory, CXL and other media to expand the memory capacity. In the embodiment of the present application, the tiered memory system is divided into two levels: a fast memory layer and a slow memory layer. The fast memory layer corresponds to a storage medium with limited capacity but fast access speed, such as DRAM, and the slow memory layer corresponds to a storage medium with large capacity but slow access speed, such as CXL and NVW.

[0022] Each memory layer is used to store allocated pages. In the embodiment of the present application, pages are divided into hot pages and cold pages according to their access popularity. The main purpose of the layered memory is to migrate hot pages and cold pages between the fast memory layer and the slow memory layer to store the appropriate pages in the appropriate memory layer. In general, hot pages are migrated to the fast memory layer, and cold pages are migrated to the slow memory layer, because hot pages are often accessed more frequently.

[0023] If all the work of page migration is handled by the CPU in the device, it will inevitably cause a large CPU overhead and affect the CPU's processing efficiency for other tasks. In view of this, in the embodiment of the present application, in order to reduce CPU overhead and improve page migration efficiency, a hardware accelerator (such as DSA) can be used to implement part of the page migration work: mainly the copy operation (copy page) of the page to be migrated from one memory layer to another memory layer, and the parallel copy capability of the hardware accelerator can be used to provide the processing efficiency of the copy operation. In this way, the CPU can be freed up as much as possible and the CPU overhead during the page migration process can be reduced. In addition, in view of the situation where there is a mixed migration of large pages (huge pages, such as 2MB pages) and basic pages (4KB pages) in actual applications, the embodiment of the present application optimizes the mixed migration path of 4KB and 2MB pages by adopting a concurrent migration algorithm, thereby improving the efficiency of page migration.

[0024] The page migration scheme provided in the embodiment of the present application is a technical scheme for optimizing memory management, which aims to improve system performance through efficient page migration technology. The scheme achieves dynamic optimization configuration of memory resources by migrating hot pages to a fast memory layer (such as DRAM) and migrating cold pages to a slow memory layer (such as persistent memory). The following is an introduction to the page migration scheme in the layered memory provided in the embodiment of the present application.

[0025] Figure 1 A schematic diagram of a page migration system provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the page migration system located in an electronic device includes a processor (such as the CPU shown in the figure), a hardware accelerator (such as the DSA shown in the figure), a fast memory layer and a slow memory layer. Figure 1 Only these two memory layers are used as examples for illustration. In fact, the memory system can be divided into multiple memory layers with different access speeds and storage capacities, without limitation.

[0026] In practical applications, the electronic device may be various types of terminal devices or a server. The processor and hardware accelerator are not limited to CPU and DSA. The number of hardware accelerators may be multiple, such as multiple DSAs shown in the figure.

[0027] The CPU and DSA work together to complete the migration of pages between the fast memory layer and the slow memory layer.

[0028] For example, suppose the pages currently stored in the fast memory layer and the slow memory layer are as follows Figure 1 As shown in , the CPU can periodically calculate the access heat for each allocated page, such as by generating Figure 1 This is achieved using the page histogram shown in the figure.

[0029] Specifically, the CPU can generate a page histogram with a set duration (such as 100 milliseconds, 200 milliseconds) as a cycle, such as generating a page histogram corresponding to each page in the current statistical time period, so as to determine the access popularity of each page according to the page histogram, wherein the horizontal axis of the page histogram represents different access number ranges, and the vertical axis represents the corresponding page identifiers within different access number ranges. Assuming that the access number is sorted from high to low along the horizontal axis direction, the page within the access number range on the far left of the page histogram has a higher access number, that is, a higher access popularity.

[0030] Afterwards, the CPU distinguishes hot pages and cold pages based on the access popularity of these pages. Figure 1 As shown in , it is assumed that the access heat of page a originally in the fast memory layer is found to be relatively low, and it is determined as a cold page; and the access heat of page b originally in the slow memory layer is found to be relatively high, and it is determined as a hot page.

[0031] Afterwards, the CPU determines the pages to be migrated based on the available space in the fast memory layer, for example, it determines that page a needs to be migrated from the fast memory layer to the slow memory layer, and page b needs to be migrated from the slow memory layer to the fast memory layer. This triggers the DSA hardware accelerator to complete the above migration operation, that is, copying page a from the fast memory layer to the slow memory layer, and copying page b from the slow memory layer to the fast memory layer.

[0032] During the copying process, DSA can implement the above copying process based on the page concurrent copying function provided by the embodiment of the present application, and also improve the processing efficiency. Figure 1 The work queues (WQs) corresponding to the multiple DSAs shown in the figure may contain different pages that need to be migrated and are allocated by the CPU. Through concurrent replication of multiple DSAs, the delay of page migration can be reduced and the migration efficiency can be improved.

[0033] Figure 2 A flowchart of a method for migrating pages in a hierarchical memory provided by an embodiment of the present application. The method is applied to a processor in an electronic device, wherein the electronic device is provided with a page concurrent copy function implemented by a driver based on a hardware accelerator. Figure 2 As shown, the method comprises the following steps: 201. Determine the visit popularity of multiple pages.

[0034] 202. Determine hot pages and cold pages among the multiple pages according to the access popularity.

[0035] 203. According to the available space of the fast memory layer, determine the target hot page to be migrated from the hot pages, and determine the target cold page to be migrated from the cold pages.

[0036] 204. Trigger the hardware accelerator to migrate the target hot page to the fast memory layer and migrate the target cold page to the slow memory layer by calling the concurrent page copy function.

[0037] In the embodiment of the present application, in order to realize efficient migration of pages based on multiple hardware accelerators set in the electronic device, a modified driver of the hardware accelerator can be integrated in the operating system kernel of the electronic device. If the driver that can be applied to various types of operating system kernels without modification of the hardware accelerator is called a basic driver, then the driver that adds the page concurrent copy function on the basis of the basic driver can be called a new driver. Then the new driver is set in the operating system kernel of the above-mentioned electronic device, that is, the concurrent copy function of the page provided by the new driver is used to trigger the hardware accelerator to complete the concurrent copy operation of the page that needs to be migrated.

[0038] In terms of specific implementation, it can be simply considered that the page concurrent copy function is modularized and then packaged together with the basic driver to form a new driver, which can be adapted to different versions of operating system kernels.

[0039] In addition, in actual applications, the new driver program may be added to a setting directory in the operating system kernel, such as the misc / exp directory, but is not limited thereto.

[0040] In actual applications, if the hardware accelerator set in the electronic device originally copies / transfers some data in a "non-concurrent manner", then the above-mentioned page concurrent copy function provided in the embodiment of the present application is equivalent to adding an option for the hardware accelerator to copy / transfer data in a "concurrent manner". Based on this, optionally, a system interface (sysAPI) can be provided in the operating system kernel, and the above-mentioned new driver can switch between the above-mentioned two methods by calling the system interface, for example, in the page migration scenario, choose to use the "concurrent mode" corresponding to the page concurrent copy function.

[0041] In an optional embodiment, the above-mentioned page concurrent copy function can be implemented in the form of one or more application programming interfaces (Application Programming Interface, referred to as API). For example, the page concurrent copy function is provided by a first API and a second API, the first API is an interface for triggering page migration, such as dsa_copy_page_lists, and the second API is an interface for concurrently copying pages, such as dsa_multi_copy_pages. The second API is nested by the first API. Specifically, when it is determined that certain pages need to be migrated, the migration operation process for these pages can be triggered by calling the first API, and the second API will be called in the first API to actually complete the concurrent copy process for these pages: copy these pages from a certain memory layer to another memory layer. By implementing the above-mentioned page concurrent copy function in the form of two APIs, the implementation method and usage method are simple and convenient.

[0042] Based on the above page concurrent copy function implemented in the operating system kernel, the process of the processor and the hardware accelerator cooperating to complete the page migration is as follows: First, the processor determines the access heat of multiple pages. These pages are allocated pages, which may include large pages and basic pages. The access heat may be the number of accesses within a set statistical time period. In practical applications, the processor may regularly count the number of accesses within a corresponding statistical time period for each allocated page. Thus, optionally, the processor may create a migration thread and periodically wake up the migration thread to perform page migration tasks.

[0043] Afterwards, the processor determines the hot pages and cold pages contained therein according to the access heat of each page. In an optional embodiment, a threshold for identifying hot pages and cold pages can be preset, and the number of the thresholds can be one or two. When there is one threshold, pages with access heat greater than the threshold can be considered hot pages, and pages with access heat less than or equal to the threshold can be considered cold pages. When there are two thresholds, such as including a first threshold and a second threshold (the first threshold is greater than the second threshold), pages with access heat greater than or equal to the first threshold can be considered hot pages, and pages with access heat less than or equal to the second threshold can be considered cold pages.

[0044] Afterwards, the processor determines the target hot pages to be migrated from the hot pages and the target cold pages to be migrated from the cold pages based on the available space of the fast memory layer. In an optional embodiment, some of the above-mentioned multiple pages may have been previously stored in the fast memory layer, but are identified as cold pages based on their access heat in the current statistical time period, then the target cold pages to be migrated include these pages, based on which, the available space of the fast memory layer can be the sum of the remaining storage space of the fast memory layer before this migration process and the storage space that can be freed after these cold pages are assumed to be moved out. Thus, if the available space is sufficient to accommodate all of the hot pages that were not originally in the fast memory layer, then the target hot pages to be migrated include all the hot pages that are not in the fast memory layer, otherwise, some hot pages can be selected from the available space in descending order of access heat as target hot pages.

[0045] Afterwards, the processor triggers the hardware accelerator to migrate the target hot page to the fast memory layer and the target cold page to the slow memory layer by calling the concurrent page copy function implemented by the hardware accelerator-based driver in the operating system kernel.

[0046] There are often multiple target hot pages and target cold pages to be migrated, and they may include large pages and basic pages. In order to improve migration efficiency, a concurrent page replication method is adopted. Specifically, as mentioned above, multiple hardware accelerators can be set in the electronic device, and each hardware accelerator can correspond to a work queue. The processor can assign the target hot pages and target cold pages that need to be migrated to different work queues, so that the hardware accelerators corresponding to different work queues read the pages in the corresponding work queues to complete the corresponding migration operation: that is, copy the target hot pages to the fast memory layer, and copy the target cold pages to the slow memory layer.

[0047] In an optional embodiment, migration processing can be performed for target hot pages and target cold pages respectively. For example, for target cold pages, the processor first allocates different target cold pages to different work queues, so that the hardware accelerator copies the target cold pages in the respective work queues to the slow memory layer. Afterwards, after the processor senses that the migration process of the target cold page has been completed, it allocates different target hot pages to different work queues, so that the hardware accelerator copies the target hot pages in the respective work queues to the fast memory layer.

[0048] In an embodiment of the present application, through the cooperation of the above processor and hardware accelerator, the processor can only complete the task of determining the target hot page and the target cold page, and the subsequent task of copying the target hot page and the target cold page to the corresponding memory layer is offloaded to the hardware accelerator for concurrent execution, which can reduce the processor overhead and improve the page migration efficiency.

[0049] In an optional embodiment, the processor triggers the hardware accelerator to migrate the target hot page to the fast memory layer and the target cold page to the slow memory layer by calling the above-mentioned concurrent page copy function, which can be specifically implemented as follows: By calling the page concurrent replication function, the first type page and the second type page included in the target hot page and / or the target cold page are determined, and the capacity of the first type page is greater than the capacity of the second type page; and the first type page is split into multiple sub-pages, and the multiple sub-pages are allocated to multiple work queues that have been created, and different second type pages are allocated to different work queues, triggering the hardware accelerators corresponding to each of the multiple work queues to migrate the pages in the corresponding work queue to the corresponding target memory layer in an asynchronous replication manner, wherein the hot pages in a work queue are migrated to a fast memory layer, and the cold pages are migrated to a slow memory layer, and the number of the multiple sub-pages matches the number of the multiple work queues.

[0050] Wherein, optionally, when the page concurrent replication function is implemented by the first API and the second API mentioned above, the processor can call the first API to trigger the migration process for the target hot page and / or the target cold page after determining the target hot page and the target cold page. Wherein, "and / or" here means: all target hot pages that need to be migrated can be added to a list A, and all target cold pages that need to be migrated can be added to a list B, so as to perform migration processing for these two lists respectively; or, all target hot pages and target cold pages that need to be migrated can be added to a list C, and migration processing can be performed for this list C.

[0051] As described above, the second API is called during the execution of the first API, that is, after the above list is formed, the second API is called to allocate the pages in the corresponding list before the concurrent copy operation, and after the allocation processing, the hardware accelerator is triggered to perform the concurrent copy operation.

[0052] Specifically, no matter which of the above lists is used, the pages that need to be migrated (target cold pages, target hot pages) contained therein may include the first type of pages: large pages, and the second type of pages: basic pages, and these two types of pages need to be concurrently allocated separately.

[0053] In actual applications, the processor can first initialize the number of work queues corresponding to the hardware accelerator as needed. Optionally, the hardware accelerator and its corresponding work queue can be in a one-to-one correspondence, but is not limited to this. For ease of description, it is assumed here that the processor sets N work queues corresponding to N hardware accelerators, and N>1.

[0054] Afterwards, the processor can determine the first type of pages, i.e., large pages, contained in the target hot page and / or the target cold page during the first traversal, and then split each first type page into multiple sub-pages, and allocate the multiple sub-pages to N work queues, where the number of sub-pages split from each first type page is equal to the number of N work queues. The N sub-pages split from a first type page can be allocated to the N work queues using a round-robin method. Depending on the value of N, the size of the sub-page may be larger or smaller than the size of the base page (4KB). By splitting and concurrently copying large pages, the migration efficiency of large pages can be improved.

[0055] Afterwards, during the second traversal, the processor allocates the remaining second type pages, i.e., the basic pages, to different work queues. Optionally, the processor may use a round-robin method to allocate one second type page to each work queue each time. Moreover, for the second type pages finally allocated in each work queue, the processor may set a batch size, i.e., how many second type pages are copied at a time, to improve the copying efficiency. Alternatively, the processor calculates the number of second type pages that need to be evenly allocated to each work queue based on the total number of remaining second type pages and the total number of work queues N, and sets the number of second type pages allocated to each work queue in one or several batch lists, such as pre-setting the number of second type pages that need to be copied during a batch process to determine how many batch processes are required.

[0056] After completing the above allocation process, the processor can trigger the N hardware accelerators to copy the allocated pages in their respective work queues to the corresponding memory layer. In actual applications, the processor can create a copy thread (expressed as idxd_wq_thread) for the work queue corresponding to each hardware accelerator in advance.

[0057] In addition, in the embodiment of the present application, during the copy operation, the hardware accelerator uses an asynchronous copy mode to complete the copy operation of the pages in the corresponding work queue, that is, the migration operation.

[0058] Specifically, after the processor completes the allocation processing of the pages in each work queue, it triggers the copy request corresponding to each work queue to the corresponding hardware accelerator. Taking a hardware accelerator as an example, the hardware accelerator can set a status flag indicating whether the page in the corresponding work queue has been copied. Initially, the value of the status flag is a first value indicating that it is not completed. By setting the status flag to the first value, an interrupt signal (such as an MSI-X interrupt) can be triggered to wake up the corresponding copy thread (idxd_wq_thread), and the copy thread scans all pages in the corresponding work queue that have not completed the copy operation, and performs corresponding copy operations on the pages that have not completed the copy operation: for example, for hot pages, copy them to the fast memory layer, and for cold pages, copy them to the slow memory layer. After all pages in the work queue are copied, the copy thread calls the set callback function and sets the value of the status flag to the second value indicating that it has been completed. Based on this, the migration thread in the processor can determine that the current page migration process is completed based on the second value of the status flag corresponding to each of the N hardware accelerators, and can unblock and continue subsequent operations, such as entering the next round of page migration processing, and so on.

[0059] To facilitate understanding of the above page concurrent replication function, combined with Figure 3 Example description.

[0060] Figure 3 A schematic diagram of a page concurrent copy process provided by an embodiment of the present application, such as Figure 3 As shown, it is assumed that the target hot pages include multiple hot pages indicated in the "List of Pages to be Migrated" in the figure, some of which are large pages (2MB) and some are basic pages (4KB). The arrows indicate the order of access heat. Assuming that 8 work queues are set as shown in the figure, the first large page traversed is split into 8 sub-pages, each of which is 256KB in size and allocated to 8 work queues, corresponding to the first page in the 8 work queues.

[0061] Similarly, for the second large page traversed, it is split into 8 sub-pages, each sub-page has a size of 256KB, and is allocated to 8 work queues, corresponding to the second page in the 8 work queues.

[0062] After that, the remaining basic pages are allocated to the work queues in a round-robin manner. The batch size can then be set to determine which basic pages in each work queue are copied as a batch, for example Figure 3 It is shown in FIG. 2 that all basic pages allocated to each work queue form a batch list.

[0063] Based on the above allocation processing results, the hardware accelerators corresponding to the eight work queues can simultaneously and concurrently execute the copy operation tasks of the pages in their respective corresponding work queues.

[0064] In an optional embodiment, the number of work queues set by the processor can be set on demand or by default, such as a default setting of 8 or other set numbers. The on-demand setting is, for example: setting on demand according to the number of large pages contained in the target hot page and the target cold page to be migrated. In actual applications, the correspondence between the number of different large pages and the number of work queues can be preset. The more large pages there are, the more work queues there are, so as to determine how many work queues need to be set according to the correspondence. Through concurrent replication based on multiple work queues, the efficiency of page migration can be significantly improved.

[0065] Figure 4 A flowchart of a method for migrating pages in a hierarchical memory provided by an embodiment of the present application. The method is applied to a processor in an electronic device, wherein the electronic device is provided with a page concurrent copy function implemented by a driver based on a hardware accelerator. Figure 4 As shown, the method comprises the following steps: 401. Determine the visit popularity of multiple pages.

[0066] 402. Determine a first proportion corresponding to hot pages and a second proportion corresponding to cold pages according to the total number of pages of the plurality of pages, determine a hot page heat threshold and a cold page heat threshold according to the access heat, the first proportion, and the second proportion of the plurality of pages, and determine hot pages and cold pages among the plurality of pages according to the hot page heat threshold and the cold page heat threshold.

[0067] 403. According to the available space of the fast memory layer, determine the target hot page to be migrated from the hot pages, and determine the target cold page to be migrated from the cold pages.

[0068] 404. Trigger the hardware accelerator to migrate the target hot page to the fast memory layer and migrate the target cold page to the slow memory layer by calling the concurrent page copy function.

[0069] In this embodiment, the thresholds for identifying hot pages and cold pages can be dynamically determined to adapt to the total number of pages allocated in different statistical time periods, so as to control the pages in different memory layers to be at a reasonable number.

[0070] Specifically, in the current statistical time period, assuming that the total number of multiple pages allocated is n, the first proportion corresponding to hot pages can be preset, such as 20%, and the second proportion corresponding to cold pages can be preset, such as 60% and 80%. Based on the total number of pages n, the first proportion, the second proportion, and the access heat of each page, the hot page heat threshold and the cold page heat threshold can be determined.

[0071] For example, n=100, the first proportion=20%, the second proportion=60%, then optionally, the access heat values ​​can be sorted from high to low, and the access heat corresponding to the page ranked 20th is determined as the hot page heat threshold corresponding to the current statistical time period, and the access heat corresponding to the page ranked 60th from the bottom is determined as the cold page heat threshold corresponding to the current statistical time period.

[0072] Or, optionally, different non-overlapping access heat value ranges arranged from high to low can be pre-set, and the value range to which each page belongs can be determined according to its access heat. The number of hot pages n1 and the number of cold pages n2 are determined according to the total number of pages n, the first proportion, and the second proportion. Then, based on the order of the above different access heat value ranges and the number of pages contained in each value range, the lower limit of the first access heat value range is determined as the hot page heat threshold and the upper limit of the second access heat value range is determined as the cold page heat threshold, wherein the access heat of the above n1 pages is greater than or equal to the lower limit of the first access heat value range, and the access heat of the above n2 pages is less than or equal to the upper limit of the second access heat value range.

[0073] Based on the determined hot page heat threshold and cold page heat threshold in the current statistical time period, it can be determined that pages with access heat higher than the hot page heat threshold among multiple pages are hot pages, and pages with access heat lower than the cold page heat threshold are cold pages.

[0074] The implementation process of other steps in this embodiment can refer to the relevant description in the aforementioned embodiment and will not be repeated here.

[0075] By dynamically adjusting the threshold for distinguishing hot pages and cold pages according to the total number of allocated pages in different statistical time periods, it is possible to prevent the determination of results based on an excessive number of hot pages and cold pages, and avoid frequent and unnecessary page migration.

[0076] Figure 5A flowchart of a method for migrating pages in a hierarchical memory provided by an embodiment of the present application. The method is applied to a processor in an electronic device, wherein the electronic device is provided with a page concurrent copy function implemented by a driver based on a hardware accelerator. Figure 5 As shown, the method comprises the following steps: 501. Determine access heats corresponding to a first number of pages stored in a slow memory layer, and access heats corresponding to a second number of pages stored in a fast memory layer.

[0077] 502. Determine a first proportion corresponding to hot pages and a second proportion corresponding to cold pages according to the total number of pages of the plurality of pages, and determine a hot page heat threshold and a cold page heat threshold according to the access heat, the first proportion, and the second proportion of the plurality of pages.

[0078] 503. Determine hot pages from the first number of pages according to the access heats corresponding to the first number of pages and the hot page heat thresholds, and add the hot pages to a promotion list, where the promotion list corresponds to the slow memory layer.

[0079] 504. Determine cold pages from the second number of pages according to the access heat and cold page heat thresholds corresponding to each of the second number of pages, and add the cold pages to a demotion list, where the demotion list corresponds to the fast memory layer.

[0080] 505. According to the available space of the fast memory layer, determine the target hot page to be migrated from the hot pages included in the promotion list, and determine the target cold page to be migrated from the cold pages included in the demotion list.

[0081] 506. Trigger the hardware accelerator to migrate the target hot page to the fast memory layer and migrate the target cold page to the slow memory layer by calling the concurrent page copy function.

[0082] In actual applications, in the initial stage, the initially allocated pages can be optionally stored in the fast memory layer or the slow memory layer, and after a set time or after the number of pages stored in the two memory layers reaches a set value, a periodically triggered page migration process is triggered. During this period, the newly allocated pages can be stored according to the default strategy (such as allocation to the slow memory layer or allocation to the fast memory layer). Based on this, it can be considered that the page migration process is a process of migrating pages that have been stored in the fast memory layer and the slow memory layer. Specifically, the cold pages originally stored in the fast memory layer are migrated to the slow memory layer, and the hot pages originally stored in the slow memory layer are migrated to the fast memory layer.

[0083] In view of this, it can be understood that within the current statistical time period, the processor can determine the access heat corresponding to each of the first number of pages stored in the slow memory layer, and the access heat corresponding to each of the second number of pages stored in the fast memory layer, and then determine the hot pages from the first number of pages and determine the cold pages from the second number of pages.

[0084] As described above, optionally, a first proportion corresponding to hot pages and a second proportion corresponding to cold pages can be determined based on the total number of pages of a plurality of pages (the plurality of pages include a first number of pages and a second number of pages), and a hot page heat threshold and a cold page heat threshold can be determined based on the access heat, the first proportion, and the second proportion of the plurality of pages. Afterwards, based on the access heat and the hot page heat threshold corresponding to each of the first number of pages, hot pages are determined from the first number of pages, and the hot pages are added to a promotion list, and the promotion list corresponds to a slow memory layer. Based on the access heat and the cold page heat threshold corresponding to each of the second number of pages, cold pages are determined from the second number of pages, and the cold pages are added to a demotion list, and the demotion list corresponds to a fast memory layer.

[0085] In this embodiment, since the slow memory layer mainly stores cold pages, the migration direction is mainly to promote the pages that have become hot pages in the current statistical time period to the fast memory layer, therefore, a promotion list can be maintained for the slow memory layer, and the hot pages determined therein can be added to the promotion list. On the contrary, the fast memory layer mainly stores hot pages, and the migration direction is mainly to demote the pages that have become cold pages in the current statistical time period to the slow memory layer, therefore, a demote list can be maintained for the fast memory layer, and the cold pages determined therein can be added to the demote list.

[0086] In addition, when determining the access popularity of each page within the current statistical time period, a page histogram corresponding to multiple pages within the current statistical time period can be generated to determine the access popularity of multiple pages based on the page histogram, wherein the horizontal axis of the page histogram represents different preset access number ranges, and the vertical axis represents the corresponding page identifiers within different access number ranges.

[0087] In actual applications, the processor can create a migration thread in the kernel for the slow memory layer and the fast memory layer respectively, so that the corresponding migration threads can determine the access heat of each page stored therein, the page classification identification (i.e., the identification of hot pages and cold pages), the maintenance of the corresponding promotion list / demotion list, and the triggering of migration operations for the pages in the promotion list / demotion list. In other words, at this time, the migration operation can be triggered for the promotion list and the demotion list respectively.

[0088] Specifically, according to the available space of the fast memory layer, the target cold page to be migrated can be determined from the cold pages included in the demoted list, and the target hot page to be migrated can be determined from the hot pages included in the promoted list. Then, by calling the concurrent page replication function, the hardware accelerator is triggered to migrate the target hot page to the fast memory layer and the target cold page to the slow memory layer. That is, at this time, Figure 3 There are two lists of pages to be migrated, one storing the target hot pages and the other storing the target cold pages, and the migration operation of concurrent copying is performed on the two lists of pages to be migrated. The specific migration process refers to the relevant description in the above embodiment, which will not be repeated here.

[0089] Figure 6 A flowchart of a method for migrating pages in a hierarchical memory provided by an embodiment of the present application. The method is applied to a processor in an electronic device, wherein the electronic device is provided with a page concurrent copy function implemented by a driver based on a hardware accelerator. Figure 6 As shown, the method comprises the following steps: 601. Determine the visit popularity of multiple pages in the current statistical time period.

[0090] 602. If the current statistical time period meets the set heat cooling condition, the access heat of multiple pages is reduced by a set range.

[0091] 603. Determine hot pages and cold pages among the multiple pages according to the access heats of the multiple pages after the access heats are reduced.

[0092] 604. According to the available space of the fast memory layer, determine the target hot page to be migrated from the hot pages, and determine the target cold page to be migrated from the cold pages.

[0093] 605. Trigger the hardware accelerator to migrate the target hot page to the fast memory layer and migrate the target cold page to the slow memory layer by calling the concurrent page copy function.

[0094] In this embodiment, since the capacity of the fast memory layer is often limited, in order to prevent frequent overflow of the fast memory layer (that is, too many hot pages exceed its storage capacity), the access heat of the pages can be cooled regularly, so that more hot pages originally stored in the fast memory layer can be turned into cold pages, and then migrated to the slower memory layer with larger capacity to free up available space in the fast memory layer.

[0095] Specifically, it is possible to set the page cooling process to be performed every m statistical time periods, where m>1. If a current statistical time period meets this condition, then after determining the access heat corresponding to each of the multiple pages stored in the slow memory layer and the fast memory layer at this time in the manner provided in the aforementioned embodiment, a set reduction process is performed on each access heat, thereby ensuring that the relative proportion of the access heat of all pages remains unchanged. Alternatively, the access heat reduction process may be performed only on the pages in the fast memory layer. The reduction range can be customized, such as reducing by half or another proportion.

[0096] Based on this, it can be understood that if the original access heat of a page is H, it is reduced by half to H / 2, and the threshold used to distinguish between cold and hot pages remains unchanged. Assuming that H is greater than the threshold, then H / 2 is likely to become less than the threshold, so that the page may change from a hot page to a cold page and be migrated out of the fast memory layer.

[0097] Figure 7 A flowchart of a method for migrating pages in a hierarchical memory provided by an embodiment of the present application. The method is applied to a processor in an electronic device, wherein the electronic device is provided with a page concurrent copy function implemented by a driver based on a hardware accelerator. Figure 7 As shown, the method comprises the following steps: 701. Determine the visit popularity of multiple pages.

[0098] 702. Determine a page to be migrated among the multiple pages according to the access popularity, and different pages to be migrated correspond to different memory layers.

[0099] 703. Trigger the hardware accelerator to migrate the to-be-migrated page to the corresponding target memory layer by calling the concurrent page copy function.

[0100] In the above embodiments, a memory system having two levels of memory layers, namely a fast memory layer and a slow memory layer, is used as an example for explanation. In actual applications, the memory system may be divided into more levels of memory layers as needed, but the principle of page migration is generally applicable.

[0101] Specifically, the CPU can still periodically count the access heat corresponding to each of the allocated pages, and determine the pages to be migrated among multiple pages based on the access heat. Different pages to be migrated correspond to different memory layers, that is, some pages to be migrated need to be migrated to memory layer A, some pages to be migrated need to be migrated to memory layer B, and some pages to be migrated need to be migrated to memory layer C, etc. After that, the CPU triggers the hardware accelerator to complete the migration operation of these pages to be migrated to the corresponding memory layer based on the concurrent page replication function.

[0102] For the contents not described in detail in this embodiment, please refer to the relevant descriptions in other embodiments and will not be repeated here.

[0103] Figure 8 A flowchart of a method for migrating pages in a hierarchical memory provided in an embodiment of the present application. The method is applied to a hardware accelerator in an electronic device, the electronic device also includes a processor, and the electronic device is provided with a page concurrent copy function implemented by a driver of the hardware accelerator, and the page concurrent copy function is executed by the hardware accelerator. Figure 8 As shown, the method comprises the following steps: 801. Obtain a page to be migrated determined by a processor from a plurality of pages, where different pages to be migrated correspond to different memory layers; wherein the processor determines the page to be migrated according to access heats of the plurality of pages.

[0104] 802. Based on the concurrent page replication function, migrate the page to be migrated to the corresponding target memory layer.

[0105] In this embodiment, the process of the hardware accelerator executing the concurrent page copy function can refer to the relevant descriptions in other embodiments and will not be repeated here.

[0106] The page migration scheme provided in the above embodiments of the present application can be applied to various data-intensive applications, such as machine learning, deep learning, graph processing, big data analysis, database applications and other application scenarios. In summary, in these application scenarios, there are often some frequently accessed data, and the pages storing these data can be migrated to a memory layer with faster access speed (but possibly smaller capacity), which can speed up the data calculation process, and migrate the pages corresponding to some data that is no longer frequently accessed to a memory layer with larger capacity (but possibly slower access speed) for subsequent use.

[0107] For example, in a deep learning scenario, a deep neural network model often includes a backbone network and other branch networks (such as convolutional network layers, etc.), and different networks correspond to different weight parameters, where the weight parameters corresponding to the backbone network are used for each input data. Based on this, based on the access popularity of the pages corresponding to the weight parameters corresponding to the backbone network, it can be determined that these pages are hotter pages and can be migrated to the fast memory layer, while the pages corresponding to some other branch networks (which may have different functions) will be migrated to the slow memory layer.

[0108] Fig. 9 The present invention provides a schematic diagram of a device for migrating pages in a hierarchical memory, which is applied to a processor in an electronic device, wherein the electronic device is provided with a concurrent page copy function implemented by a driver based on a hardware accelerator. Fig. 9As shown, the device includes: a heat determination module 11, a page classification module 12, and a migration processing module 13.

[0109] The popularity determination module 11 is used to determine the access popularity of multiple pages.

[0110] The page classification module 12 is used to determine the pages to be migrated among the multiple pages according to the access popularity, and different pages to be migrated correspond to different memory layers.

[0111] The migration processing module 13 is used to trigger the hardware accelerator to migrate the to-be-migrated page to the corresponding target memory layer by calling the concurrent page replication function.

[0112] Optionally, the operating system kernel of the electronic device is provided with a page concurrent copy function implemented by the hardware accelerator-based driver.

[0113] Optionally, the page classification module 12 is specifically used to: determine the hot pages and cold pages among the multiple pages according to the access heat; determine the target hot pages to be migrated from the hot pages, and determine the target cold pages to be migrated from the cold pages according to the available space of the fast memory layer. The migration processing module 13 is specifically used to: trigger the hardware accelerator to migrate the target hot pages to the fast memory layer and migrate the target cold pages to the slow memory layer by calling the page concurrent replication function.

[0114] Optionally, the migration processing module 13 is specifically used to: determine the first type of page and the second type of page contained in the target hot page and / or the target cold page by calling the page concurrent replication function, the capacity of the first type of page is greater than the capacity of the second type of page; split the first type of page into multiple sub-pages, and allocate the multiple sub-pages to multiple work queues that have been created, and allocate different second type pages to different work queues, and the number of the multiple sub-pages matches the number of the multiple work queues; trigger the hardware accelerators corresponding to each of the multiple work queues to migrate the pages in the corresponding work queue to the corresponding target memory layer in an asynchronous replication manner, wherein the hot pages in a work queue are migrated to the fast memory layer, and the cold pages are migrated to the slow memory layer.

[0115] Optionally, the page concurrent copying function is provided through a first API and a second API, the first API is an interface for triggering page migration, the second API is an interface for concurrently copying pages, and the second API is nested by the first API.

[0116] Optionally, the page classification module 12 is specifically used to: determine a first proportion corresponding to hot pages and a second proportion corresponding to cold pages according to the total number of pages of the multiple pages; determine a hot page heat threshold and a cold page heat threshold according to the access heat of the multiple pages and the first proportion and the second proportion; determine hot pages and cold pages among the multiple pages according to the hot page heat threshold and the cold page heat threshold.

[0117] Optionally, the heat determination module 11 is specifically used to: determine the access heat corresponding to each of the first number of pages stored in the slow memory layer, and the access heat corresponding to each of the second number of pages stored in the fast memory layer, the multiple pages including the first number of pages and the second number of pages. The page classification module 12 is specifically used to: determine hot pages from the first number of pages according to the access heat corresponding to each of the first number of pages, and add the hot pages to a promotion list, the promotion list corresponding to the slow memory layer; determine cold pages from the second number of pages according to the access heat corresponding to each of the second number of pages, and add the cold pages to a demotion list, the demotion list corresponding to the fast memory layer. The migration processing module 13 is specifically used to: determine the target hot pages to be migrated from the hot pages included in the promotion list according to the available space of the fast memory layer, and determine the target cold pages to be migrated from the cold pages included in the demotion list.

[0118] Optionally, the heat determination module 11 is specifically used to: generate a page histogram corresponding to the multiple pages in the current statistical time period, so as to determine the access heat of the multiple pages according to the page histogram, the horizontal axis of the page histogram represents different access number ranges, and the vertical axis represents the corresponding page identifiers within different access number ranges.

[0119] Optionally, the heat determination module 11 is specifically used to: determine the access heat of the multiple pages in the current statistical time period; if the current statistical time period meets the set heat cooling condition, reduce the access heat of the multiple pages by a set range.

[0120] Fig. 9 The device shown can execute the steps in the page migration method in the hierarchical memory in the aforementioned embodiment. The detailed execution process and technical effects can be found in the description of the aforementioned embodiment, which will not be repeated here.

[0121] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Fig.10 As shown, in practice, the electronic device includes: a memory 21 and a processor 22.

[0122] The memory 21 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0123] In the embodiment of the present application, the operating system kernel of the electronic device is provided with a concurrent page copy function implemented by a driver based on a hardware accelerator. The driver and the program code corresponding to the concurrent page copy function can be stored in the memory 21.

[0124] The processor 22 is coupled to the memory 21 and is used to execute the computer program in the memory 21 to implement the following steps: Determine the popularity of multiple pages; Determine a page to be migrated among the multiple pages according to the access popularity, where different pages to be migrated correspond to different memory layers; The hardware accelerator is triggered by calling the page concurrent copy function to migrate the to-be-migrated page to the corresponding target memory layer.

[0125] Optionally, the operating system kernel of the electronic device is provided with a page concurrent copy function implemented by the hardware accelerator-based driver.

[0126] Wherein, optionally, determining the pages to be migrated among the multiple pages according to the access heat includes: determining the hot pages and cold pages among the multiple pages according to the access heat of the multiple pages; determining the target hot pages to be migrated from the hot pages and determining the target cold pages to be migrated from the cold pages according to the available space of the fast memory layer. Thus, triggering the hardware accelerator to migrate the pages to be migrated to the corresponding target memory layer by calling the page concurrent replication function includes: triggering the hardware accelerator to migrate the target hot pages to the fast memory layer and the target cold pages to the slow memory layer by calling the page concurrent replication function.

[0127] In an optional embodiment, the above-mentioned page concurrent copying function is provided through a first API and a second API, wherein the first API is an interface for triggering page migration, and the second API is an interface for concurrently copying pages, and the second API is nested by the first API.

[0128] In an optional embodiment, the hardware accelerator is triggered by calling the concurrent page copy function to migrate the target hot page to the fast memory layer and the target cold page to the slow memory layer, which can be specifically implemented as follows: By calling a page concurrent replication function, determining a first type of page and a second type of page included in a target hot page and / or a target cold page, the capacity of the first type of page is greater than the capacity of the second type of page; Splitting the first type page into multiple sub-pages, and assigning the multiple sub-pages to the multiple work queues that have been created, assigning different second type pages to different work queues, and the number of the multiple sub-pages matches the number of the multiple work queues; The hardware accelerators corresponding to the multiple work queues are triggered to migrate the pages in the corresponding work queue to the corresponding target memory layer in an asynchronous replication manner, wherein the hot pages in a work queue are migrated to the fast memory layer, and the cold pages are migrated to the slow memory layer.

[0129] In an optional embodiment, determining hot pages and cold pages among a plurality of pages according to the access heat may be specifically implemented as follows: Determine, according to the total number of pages of the plurality of pages, a first proportion corresponding to the hot pages and a second proportion corresponding to the cold pages; Determine a hot page heat threshold and a cold page heat threshold according to the access heat, the first proportion, and the second proportion of the multiple pages; According to the hot page temperature threshold and the cold page temperature threshold, hot pages and cold pages among the multiple pages are determined.

[0130] In an optional embodiment, determining the access heat of multiple pages can be implemented as follows: determining the access heat corresponding to each of the first number of pages stored in the slow memory layer, and the access heat corresponding to each of the second number of pages stored in the fast memory layer, wherein the multiple pages include the first number of pages and the second number of pages. Thus, determining the hot pages and cold pages among the multiple pages according to the access heat can be implemented as follows: determining the hot pages from the first number of pages according to the access heat corresponding to each of the first number of pages, and adding the hot pages to the promotion list, wherein the promotion list corresponds to the slow memory layer; determining the cold pages from the second number of pages according to the access heat corresponding to each of the second number of pages, and adding the cold pages to the demotion list, wherein the demotion list corresponds to the fast memory layer. Thus, determining the target hot pages to be migrated from the hot pages and determining the target cold pages to be migrated from the cold pages according to the available space of the fast memory layer can be implemented as follows: determining the target hot pages to be migrated from the hot pages included in the promotion list, and determining the target cold pages to be migrated from the cold pages included in the demotion list according to the available space of the fast memory layer.

[0131] In an optional embodiment, the access popularity of multiple pages is determined, specifically by generating a page histogram corresponding to the multiple pages in the current statistical time period, so as to determine the access popularity of the multiple pages based on the page histogram, wherein the horizontal axis of the page histogram represents different ranges of access times, and the vertical axis represents the corresponding page identifiers within different ranges of access times.

[0132] In an optional embodiment, determining the access heat of multiple pages may also be: determining the access heat of the multiple pages in a current statistical time period; if the current statistical time period meets a set heat cooling condition, reducing the access heat of the multiple pages by a set amount.

[0133] Further, if Fig.10 As shown, the electronic device also includes: a communication component 23, a display 24, a power component 25, an audio component 26 and other components. Fig.10 Only some components are shown schematically, which does not mean that the electronic device only includes Fig.10 The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT device, or a server device such as a conventional server, a cloud server or a server array.

[0134] The above-mentioned memory can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0135] The communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile communication network such as 2G, 3G, 4G / LTE, 5G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0136] The display includes a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0137] The power supply assembly provides power to various components of the device where the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply assembly is located.

[0138] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal.

[0139] Accordingly, the embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is enabled to implement each step in the above method embodiment. Among them, the computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (Phase-change Random Access Memory, PRAM), static random access memory (SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (Random-Access Memory, RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (Digital Video Disc, DVD) or other optical storage, magnetic cassette, tape disk storage or other magnetic storage device or any other non-transmission medium Accordingly, the present application embodiment also provides a computer program product, the computer program product includes a computer program or an instruction, when the computer program or the instruction is executed by the processor, the processor is enabled to implement each step in the above method embodiment. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by a computer program or an instruction. In addition, these computer programs or instructions can be applied to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor or other programmable data processing device can be implemented as a device for implementing the corresponding functions in the above method embodiment.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for migrating pages in a hierarchical memory, characterized in that: A processor applied to an electronic device, wherein the electronic device is provided with a page concurrent copy function implemented by a driver based on a hardware accelerator, and the method comprises: Determine the popularity of multiple pages; Determine a page to be migrated among the multiple pages according to the access popularity, where different pages to be migrated correspond to different memory layers; The hardware accelerator is triggered by calling the page concurrent copy function to migrate the to-be-migrated page to the corresponding target memory layer.

2. The method according to claim 1, characterized in that: The operating system kernel of the electronic device is provided with a page concurrent copying function implemented by the driver based on the hardware accelerator.

3. The method according to claim 1, characterized in that The step of determining the page to be migrated among the plurality of pages according to the access popularity includes: Determine hot pages and cold pages among the plurality of pages according to the access heat; According to the available space of the fast memory layer, determining a target hot page to be migrated from the hot pages, and determining a target cold page to be migrated from the cold pages; The triggering the hardware accelerator to migrate the to-be-migrated page to the corresponding target memory layer by calling the concurrent page copy function includes: The hardware accelerator is triggered by calling the concurrent page copy function to migrate the target hot page to the fast memory layer and migrate the target cold page to the slow memory layer.

4. The method according to claim 3, characterized in that: The triggering of the hardware accelerator to migrate the target hot page to the fast memory layer and the target cold page to the slow memory layer by calling the concurrent page copy function includes: By calling the page concurrent replication function, determining a first type of page and a second type of page included in the target hot page and / or the target cold page, the capacity of the first type of page is greater than the capacity of the second type of page; Splitting the first type page into a plurality of sub-pages, and assigning the plurality of sub-pages to a plurality of work queues that have been created, assigning different second type pages to different work queues, and the number of the plurality of sub-pages matches the number of the plurality of work queues; The hardware accelerators corresponding to each of the multiple work queues are triggered to migrate the pages in the corresponding work queue to the corresponding target memory layer in an asynchronous replication manner, wherein the hot pages in a work queue are migrated to the fast memory layer, and the cold pages are migrated to the slow memory layer.

5. The method according to claim 1, characterized in that The page concurrent copying function is provided by a first API and a second API, wherein the first API is an interface for triggering page migration, and the second API is an interface for concurrently copying pages, and the second API is nested by the first API.

6. The method according to claim 3, characterized in that The step of determining hot pages and cold pages among the plurality of pages according to the access heat comprises: Determine, according to the total number of pages of the plurality of pages, a first proportion corresponding to hot pages and a second proportion corresponding to cold pages; Determining a hot page heat threshold and a cold page heat threshold according to the access heat of the multiple pages and the first proportion and the second proportion; Hot pages and cold pages among the multiple pages are determined according to the hot page heat threshold and the cold page heat threshold.

7. The method according to claim 3, characterized in that Determining the access popularity of multiple pages includes: Determine access heats corresponding to a first number of pages stored in the slow memory layer, and access heats corresponding to a second number of pages stored in the fast memory layer, wherein the plurality of pages include the first number of pages and the second number of pages; The step of determining hot pages and cold pages among the plurality of pages according to the access heat comprises: determining hot pages from the first number of pages according to access heats corresponding to the first number of pages, and adding the hot pages to a promotion list, the promotion list corresponding to a slow memory layer; Determine cold pages from the second number of pages according to access heats corresponding to the second number of pages, and add the cold pages to a demotion list, where the demotion list corresponds to a fast memory layer; The step of determining a target hot page to be migrated from the hot pages and a target cold page to be migrated from the cold pages according to the available space of the fast memory layer includes: According to the available space of the fast memory layer, the target hot page to be migrated is determined from the hot pages included in the promotion list, and the target cold page to be migrated is determined from the cold pages included in the demotion list.

8. The method according to claim 1, characterized in that Determining the access popularity of multiple pages includes: Determine the visit popularity of the multiple pages in the current statistical time period; If the current statistical time period meets the set heat cooling condition, the access heat of the multiple pages is reduced by a set range.

9. A method for migrating pages in a hierarchical memory, characterized in that: A hardware accelerator applied to an electronic device, wherein the electronic device further comprises a processor and a concurrent page copy function implemented by a driver based on the hardware accelerator, and the method comprises: Obtaining a page to be migrated determined by the processor from a plurality of pages, where different pages to be migrated correspond to different memory layers; wherein the processor determines the page to be migrated according to the access heat of the plurality of pages; Based on the page concurrent replication function, the to-be-migrated page is migrated to the corresponding target memory layer.

10. An electronic device, characterized in that: include: A memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the page migration method in the hierarchical memory as described in any one of claims 1 to 9.

11. A non-transitory machine-readable storage medium, characterized in that: The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method for migrating pages in a hierarchical memory as claimed in any one of claims 1 to 9.

12. A computer program product, characterized in that include: A computer program, when executed by a processor of an electronic device, causes the processor to execute the method for migrating pages in a hierarchical memory as claimed in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Memory data migration method and device

    CN109582223A

  • Data acceleration calculation method and device, equipment and storage medium

    CN116225997A

  • Data migration method and device, electronic equipment and readable storage medium

    CN116319758A

  • Data migration method and device, chip and computer readable storage medium

    CN117806526A

  • Memory page migration method and device, memory equipment and program product

    CN118093197A