A virtualization method for memory, a deep learning system and its task switching method

By configuring the address translation table, the virtual address and physical address conversion can be realized, and memory resources are delayed, which solves the problems of low resource utilization and poor real-time performance of deep learning systems during task switching, and achieves more efficient resource utilization and real-time performance.

CN119248423BActive Publication Date: 2025-06-10奕行智能科技(广州)有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411322632.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-06-10
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

During task switching, existing deep learning systems have high overhead for on-site storage and recovery, resulting in low resource utilization, poor real-time performance, and inability to effectively virtualize computing power.

Method used

By configuring the address translation table, the conversion between virtual addresses and physical addresses is realized, memory resources are delayed, and memory transfers are reduced on-site.

Benefits of technology

It reduces the amount of memory transfer during task switching, improves resource utilization, reduces the overhead of on-site storage and recovery, and improves the real-time and flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119248423B_ABST
    Figure CN119248423B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for virtualizing memory. First, an address translation table is configured, and then virtual addresses are converted to physical addresses of the memory through this address translation table. Through memory virtualization, it is possible to delay the preservation of memory resources when the deep learning system switches tasks, so as to reduce the amount of memory transfer for context preservation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly to a method for virtualizing memory, a deep learning system, and a task switching method thereof. Background Art

[0002] Deep learning is increasingly widely used, and AI / DL tasks often require the computing power of GPUs / NPUs. In relatively complex AI task scenarios, such as scenarios with high real-time requirements like autonomous driving, strong scheduling capabilities for computing resources are needed to reasonably plan the execution of computing cores and efficiently complete complex computing tasks.

[0003] The main directions of existing task / kernel scheduling algorithms for DL models include two categories: non-preemptive scheduling of NPUs and preemptive scheduling of GPUs. GPUs / NPUs need to undertake computing tasks with high computing power density, and the context of computing cores is large. Therefore, the overhead of task scheduling / task switching is often large, much greater than that of computing cores designed for concurrent scenarios such as CPUs.

[0004] NPU processors usually contain multiple NPU computing cores, which are independent of each other. According to the execution priority and resource usage of computing tasks, the execution order of computing tasks is planned. When an NPU computing core becomes idle, according to the allocation scheduling algorithm, the computing task is dispatched. Therefore, the next computing task must wait for the currently executing task to complete, and high-priority tasks cannot preempt existing computing resources, resulting in poor real-time performance. Moreover, when multiple NPU computing tasks are parallel, multiple independent NPU computing cores are required, and some hardware resources are difficult to isolate, which will increase the hardware burden and result in poor flexibility. In addition, NPUs cannot implement a time-slice rotation-based task scheduling method, cannot virtualize computing power from the time dimension, and reduce resource utilization.

[0005] GPUs can support time-slice rotation of computing tasks. Time-slice rotation will save the context of the currently running computing task after the current time slice is completed and switch to the next computing task that needs to rotate. The next rotating task needs to restore the previously saved context to continue execution. In GPUs, this scheduling method involves the saving and restoration of the context. GPU scheduling is based on the SIMT architecture and has many SM units. Therefore, when saving the context, the registers and caches of each SM unit need to be saved to an outer storage unit, and when restoring the context, the context information needs to be read from the outer storage unit, resulting in a heavy overhead for context switching. Taking A100 as an example, the number of SM units is 108. Each SM unit has 64KB of regs and 384KB of cache, for a total of 448KB. Estimated at the highest bandwidth of 1.5TB / s, it still takes 30us to complete the saving of the context. Summary of the Invention

[0006] In view of some or all of the problems in the prior art, a first aspect of the present invention provides a method for virtualizing memory, including:

[0007] Configuring an address translation table; and

[0008] Implementing the conversion between the virtual address and the physical address of the memory through the address translation table.

[0009] Further, configuring the address translation table includes:

[0010] Setting at least one group of address mapping groups, including setting the starting virtual address, the physical address corresponding to the starting virtual address, and the address length included in the address mapping group.

[0011] Further, in each group of address mapping groups, the mapping relationship between the virtual address and the physical address is unique.

[0012] Further, the address range in each group of address mapping groups overlaps with the address ranges in the remaining N groups of address mapping groups, where N is an integer greater than or equal to zero.

[0013] Further, configuring the address translation table further includes:

[0014] Configuring the priorities of each group of address mapping groups, where the priorities of each group of address mapping groups are all different.

[0015] Further, configuring the address translation table further includes:

[0016] Setting the enable status of each address mapping group.

[0017] Further, implementing the conversion between the virtual address and the physical address of the memory through the address translation table includes:

[0018] Determining the address mapping group corresponding to the input address; and

[0019] Performing address conversion according to the information of the corresponding address mapping group.

[0020] Further, determining the address translation group corresponding to the input virtual address includes:

[0021] Comparing the input address with the address ranges of each address mapping group one by one according to the priority, and taking the first address mapping group that matches successfully as the corresponding address mapping group, where a successful match means that the input address is within the address range of the address mapping group.

[0022] Based on the virtualization method described above, a second aspect of the present invention provides a deep learning system, including:

[0023] A task management unit for scheduling and managing computing tasks;

[0024] A computing unit communicatively connected to the task management unit, adopting a DSA architecture, and including at least one computing core, a first address translation module, and a shared cache module, where each computing core includes an independent on-chip cache area, and the first address translation module is used to convert a virtual address into a physical address according to the virtualization method described above, so that the physical addresses of multiple available memory segments of the shared cache module are mapped to consecutive memory space virtual addresses; and

[0025] An external storage unit communicatively connected to the computing unit through a second address translation module, where the second address translation module is used to convert the address of the virtual space into a physical address according to the virtualization method described above, so that the program code of the computing task is mapped to a specified virtual address space.

[0026] Based on the deep learning system described above, a third aspect of the present invention provides a task switching method for a deep learning system, including:

[0027] When the task management unit receives a computing task with a higher priority, it notifies the computing unit to perform a context switch;

[0028] After the current calculations of each core of the computing unit are completed, the computing results of each core are saved to the shared cache module, and the memory occupancy of the shared cache module is sent to the task management unit;

[0029] The task management unit determines whether the remaining memory of the shared cache module meets the memory requirements of the computing task:

[0030] If it meets the requirements, the first address translation module is configured to map the physical address of the remaining memory to consecutive virtualized addresses; and

[0031] If it does not meet the requirements, at least part of the data in the shared cache module is first transferred to the external storage module so that the remaining memory of the shared cache module is not less than the memory requirements of the computing task, and then the first address translation module is configured to map the physical address of the remaining memory to consecutive virtualized addresses;

[0032] The computing unit saves the context of each core and sends a save completion signal to the task management unit;

[0033] The task management unit configures the second address translation module to map the computing task to a specified virtual address space;

[0034] The task management unit notifies the computing unit to execute a computing task; and

[0035] After the execution of the computing task is completed, the computing unit and the task management unit perform context restoration.

[0036] A memory virtualization method, a deep learning system, and a task switching method provided by the present invention enable the memory resources to be saved with a delay during task switching through memory virtualization, reducing the amount of memory transfer for context saving. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To further clarify the above and other advantages and features of the embodiments of the present invention, a more specific description of the embodiments of the present invention will be presented with reference to the accompanying drawings. It can be understood that these drawings only depict typical embodiments of the present invention and thus will not be considered as limiting its scope. In the drawings, for clarity, the same or corresponding components will be denoted by the same or similar reference numerals.

[0038] Figure 1 A schematic structural diagram of an address translation table showing an embodiment of the present invention;

[0039] Figure 2 A schematic structural diagram of a deep learning system showing an embodiment of the present invention; and

[0040] Figure 3 A schematic flowchart of a task switching method of a deep learning system showing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In the following description, the present invention is described with reference to the embodiments. However, those skilled in the art will recognize that the embodiments can be implemented without one or more specific details or in combination with other alternative and / or additional methods, materials, or components. In other cases, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring the inventive points of the present invention. Similarly, for purposes of explanation, specific quantities, materials, and configurations are set forth to provide a thorough understanding of the embodiments of the present invention. However, the present invention is not limited to these specific details. In addition, it should be understood that the embodiments shown in the drawings are illustrative representations and not necessarily drawn to scale.

[0042] In this specification, the reference to "an embodiment" or "the embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. The phrase "in an embodiment" appearing throughout this specification does not necessarily refer to the same embodiment.

[0043] It should be noted that the embodiments of the present invention describe the process steps in a specific order. However, this is only for explaining the specific embodiment and does not limit the order of the steps. On the contrary, in different embodiments of the present invention, the order of the steps can be adjusted according to the adjustment of the process.

[0044] During the calculation process of multiple computing cores of the NPU, when a task needs to preempt the current computing task, it is necessary to save the context. The NPU usually adopts a Domain-Specific Architecture (DSA). Its characteristic is that the context of the control flow register is very small, which is of the same order of magnitude as that of the CPU processor. However, the context of the cache is usually large, which is the main overhead of context saving. Context restoration is the reverse process of context saving, and the core overhead is also the transfer of the cache context. Based on this, if it is necessary to reduce the context switching overhead of the NPU, it can be considered to reduce the transfer of the cache context. After research, the inventor found that when users write computing tasks, the address spaces they consider are all fixed and continuous address spaces. This memory space can be regarded as the memory of virtual addresses, and users do not actually need to perceive the actual physical memory space. Therefore, it can be considered to map the discontinuous actual memory space into a continuous virtual address space through the method of memory virtualization. Then, when saving the context, the existing data in the cache does not need to be moved, thereby reducing the cache transfer overhead. Taking the computing task A[0:1024]=B[0:1024]+C[0:1024] as an example, the memory virtualization is briefly described. Here, A, B, and C are all arrays with a size of 1024B, and the unit is int8. From the user's perspective, the address spaces are all continuously available. Therefore, the address space of A can be allocated as 0-1024B, the address space of B can be allocated as 1024-2048B, and the address space of C can be allocated as 2048-3072B. However, the actually available memory space is 1024-4096B. Then, a memory virtualization method can be provided. For example, through an Address Translation Unit (ATU), a mapping relationship of VA[0:3072]->PA[1024:4096] is established, so that users only see the virtual addresses. By virtualizing the memory context and based on this, the process of task context saving and restoration can be realized, so as to achieve the delayed saving and switching of the memory and reduce the context switching overhead of the NPU.

[0045] The solution of the present invention will be further described below with reference to the accompanying drawings of the embodiments.

[0046] Figure 1 A schematic structural diagram of an address translation table showing an embodiment of the present invention is as follows Figure 1As shown, an address translation table ATU includes at least one set of address mapping groups Entry, where each Entry includes a starting virtual address VitrualAddr, the physical address Physical Addr corresponding to the starting virtual address, and the address length Size included in the Entry. The mapping relationship between the virtual address and the physical address is unique. In an embodiment of the present invention, the virtual address ranges in different Entries may overlap. Based on this, in an embodiment of the present invention, each Entry has a different priority. When the mapping ranges of addresses overlap, the unique mapping relationship is determined by the priority of the Entry. In yet another embodiment of the present invention, each Entry can be set to an enabled or disabled state. Specifically, each Entry includes an enable flag bit. If the enable flag bit is low, the Entry is invalid and does not participate in the ATU address translation query.

[0047] Based on this, in an embodiment of the present invention, the method for virtualizing memory includes: configuring an address translation table and implementing the conversion between the virtual address and the physical address of the memory through the address translation table. Configuring the address translation table means setting at least one set of address mapping groups. The setting of each set of address mapping groups includes setting the starting virtual address, the physical address corresponding to the starting virtual address, and the address length included in the address mapping group. The address range in any one set of address mapping groups may overlap with the address ranges in the remaining N sets of address mapping groups, where N is an integer greater than or equal to zero. In an embodiment of the present invention, configuring the address translation table further includes configuring the priorities of each set of address mapping groups, where the priorities of each set of address mapping groups are all different. In an embodiment of the present invention, configuring the address translation table further includes setting the enable state of each address mapping group, etc.

[0048] In an embodiment of the present invention, to implement the conversion between the virtual address and the physical address of the memory through the address translation table, it is first necessary to determine the address mapping group corresponding to the input address, and then perform address translation according to the information of the corresponding address mapping group. Specifically, in an embodiment of the present invention, the address translation process is as follows:

[0049] First, query in sequence according to the priority order of each Entry to determine the Entry corresponding to the input address. If the enable flag bit of an Entry is low, the Entry is invalid and does not participate in the address translation query. If no enabled Entry is found for the input address, it is default not to perform address translation. In an embodiment of the present invention, the address matching rule of the Entry is as follows:

[0050] If the input address is within the range [VA, VA + Size), that is, when the virtual address VA (Virtual Addr) <= input address (Input Addr) < (virtual address VA (Virtual Addr) + memory size (Size)), it is considered that the input address matches this Entry;

[0051] Next, address translation is performed based on this Entry. When the Entry matches, the output address (Output Addr) = VA – Input Addr + physical address PA (Physical Addr).

[0052] Based on the virtualization method described above, Figure 2 shows a schematic structural diagram of a deep learning system according to an embodiment of the present invention. As Figure 2 shown, a deep learning system includes a task management unit 201, a computing unit 202, and an external storage unit 203.

[0053] The task management unit 201 is used to schedule and manage computing tasks, and it can be, for example, an MCU.

[0054] The computing unit 202 is responsible for computing tasks. It is communicatively connected to the task management unit and follows a common DSA architecture, without a large amount of control logic, and a single instruction can correspond to a large number of computing tasks. As shown in the figure, the computing unit includes at least one computing core Core 221, a first address translation module 222, and a shared cache module L2 share memory 223. Each computing core 221 includes an independent on-chip cache area L1 memory, which can execute computations. The first address translation module 222 is used to map the physical addresses of multiple available memory segments of the shared cache module 223 into consecutive memory space virtual addresses according to the virtualization method described above, so that during context switching, the latency overflow of the system computing cache can be achieved, the amount of data transfer can be reduced, and the context switching speed can be accelerated.

[0055] The external storage unit 203, the shared cache module L2 share memory 223, and the on-chip cache area L1 memory constitute the three-level hierarchical memory of the deep learning system. At the same time, as shown in the figure, the external storage unit 203 is provided with a corresponding second address conversion module 231. The second address conversion module 231 is used to map the program code of the computing task to the specified virtual address space according to the virtualization method described above, so that each computing task uses the same virtual address space, so that the computing task does not need to be moved to a unified physical space, reducing the overhead of address relocation and accelerating the loading speed of the computing task. In an embodiment of the present invention, the external storage unit 203 may include storage modules such as DDR (HBM).

[0056] Based on the virtualization method described above, the deep learning system can virtualize the memory context, realize the delayed preservation and switching of the memory, and reduce the context switching overhead. In an embodiment of the present invention, through an external unit, such as the task management unit 201 described above, etc., the computing unit is notified that a context switch is required. After receiving the switch request, the computing unit will wait for the current instruction to complete and perform context preservation after completion. Then the computing unit notifies the external unit of the memory occupancy information at the current task switch moment, including the address / length of the memory occupancy, etc. After receiving the memory occupancy information, the external unit judges whether the remaining memory size meets the memory requirements of the task to be issued. If the remaining memory meets the requirements of the task to be issued, there is no need to save the current context memory back to the external storage. Instead, through the virtualization method described above, the current multiple available memory segments are remapped into a continuous memory space for the task to be issued to use. If the remaining memory does not meet the requirements of the issued task, it is necessary to overflow the excess memory, save it to the external storage, and then perform the virtualization operation, that is, calculate the address mapping of the ATU unit required by the task to be issued. At the same time, the computing unit actively saves the required resources such as registers, realizes the parallelism between the ATU switch and the register preservation, and notifies the external unit after the register preservation is completed. After the external unit processes the ATU address mapping and the memory overflows, and the computing unit completes the register preservation, a new computing task can be issued to the computing unit. After the computing unit finishes processing the current task, it notifies the external unit that the task is completed. At this time, the external unit judges whether to restore the context. When the context needs to be restored, the external unit notifies the computing unit to restore the context from the specified location. The computing unit restores the register and other configurations and waits for the ATU / memory context restoration message. If there is memory overflow, the external unit needs to copy the context memory back to the corresponding location. The external unit restores the address mapping of the ATU and notifies the computing unit that the ATU memory context has been restored. After receiving the restoration message, the computing unit jumps the PC pointer and continues to execute the switched computing task.

[0057] To better describe the task switching method,Figure 3 A flowchart showing a task switching method of a deep learning system according to an embodiment of the present invention. As Figure 3 shown, a task switching method of a deep learning system includes:

[0058] First, in step 301, confirm on-site switching. During the execution of a computing task by a computing unit, if the task management unit receives a computing task with a higher priority, at this time, the task management unit needs to interrupt the existing computing task and align and issue a higher-priority task, that is, the task management unit notifies the computing unit that on-site switching is required;

[0059] Next, in step 302, wait for the switching point. When the computing unit receives on-site switching, it needs to wait for the current multiple computing cores to complete their calculations. At this time, the calculation results of the computing cores are all retained in the L2-level memory, that is, the shared cache module, and the on-chip L1 memory of the computing cores is idle. At this time, it can be considered that a switching point has been reached;

[0060] Next, in step 303, confirm the memory occupancy. After reaching the switching point, the computing unit notifies the task management unit of the memory occupancy of the shared cache area. The shared cache module usually occupies multiple discontinuous memory segments. Therefore, it is necessary to notify the task management unit of the address and length of each memory segment;

[0061] Next, in step 304, determine whether the memory requirement is met. After receiving the memory occupancy information, the task management unit determines whether the total remaining memory is greater than the memory requirement of the task to be issued according to the resource situation of the task to be issued. If it is met, it directly enters step 306, memory virtualization. If it is not met, it first enters step 305, memory overflow;

[0062] In step 305, memory overflow. When the remaining memory is less than the memory requirement of the task to be issued, the task management unit needs to move away the redundant memory of the current on-site shared cache module. At this time, it is not necessary to move all the on-site memory of the current shared cache module, but only the redundant part needs to be moved to meet the memory requirement of the task to be issued. The task management unit can configure modules such as DMA to move the memory of the shared cache module to an external storage unit, and then enter step 306, memory virtualization;

[0063] In step 306, memory virtualization. After the remaining memory can meet the memory requirement of the task to be issued, the task management unit configures the first address conversion module on the side of the shared cache module to remap the multiple physically discontinuous memory segments of the current memory of the shared cache module into a memory segment with continuous virtual addresses. Thus, for the task to be issued, it will have a continuous memory segment that meets the memory requirement of the shared cache module;

[0064] Meanwhile, at step 307, the register context is saved. While memory virtualization is in progress, the computing unit concurrently saves the context within the computing core. Since the on-chip cache area within the computing core is idle at this time and does not need to be saved, mainly the register context of the instruction stream needs to be saved. The register context usually includes various general-purpose registers such as the PC, which is used for the control of the instruction stream. The size of the register context is typically below the order of 1KB, which is very lightweight compared to the GPU. It is saved by the computing unit to the external storage unit, and after the register context is saved, the task management unit is notified that the context within the computing core has been saved;

[0065] Next, at step 308, the computing task is executed. When the task management unit completes memory virtualization and receives the message from the computing unit notifying that the context within the core has been saved, it indicates that the current task context has been saved. At this time, the second address translation module is configured to map the computing task program code to be issued to a unified address space, and the computing unit is notified to execute the next computing task; and

[0066] Finally, at step 309, context restoration is performed. When the computing task of the computing unit is completed, the task management unit is notified that the current task has ended. If there are computing tasks to be restored, the task management unit restores them in sequence according to the preempted nesting order. In an embodiment of the present invention, context restoration mainly includes: the task management unit notifies the computing unit of the save address of the current context in the external storage unit, and the computing unit restores the context from the specified address. At this time, the computing unit and the task management unit work in parallel. The computing unit is responsible for restoring the lightweight context within the computing core, and the task management unit is responsible for restoring other contexts. That is, the computing unit restores context information such as registers and waits for the notification from the task management unit. If there was memory that overflowed to the external storage unit side previously, the task management unit needs to move the memory on the external storage unit side to the shared cache module, reconfigure the first and second address translation modules, restore the virtual address mapping between the external storage unit and the shared cache module of the previous context. After completion, the task management unit notifies the computing unit that the external context has been restored. After receiving the message, the computing unit jumps to the corresponding PC pointer and continues to execute the previous computing task.

[0067] In an embodiment of the present invention, during the execution of any computing task by the computing unit, as long as a higher-priority task is received, task switching can be performed according to the method described above. It should be noted that in scenarios where task switching is required, the mapping relationship between virtual addresses and physical addresses is different between different tasks. Therefore, it is necessary to establish the address mapping relationship in real time and dynamically, that is, the physical memory space is changing, while the virtual space seen by the user is continuous and fixed.

[0068] Although the embodiments of the present invention have been described above, it should be understood that they are presented by way of example only and not as a limitation. It will be apparent to those skilled in the relevant art that various combinations, variations and changes can be made thereto without departing from the spirit and scope of the present invention. Therefore, the breadth and scope of the present invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined only in accordance with the appended claims and their equivalents.

Claims

1. A deep learning system, characterized in that: include: A task management unit, which is configured to schedule and manage computing tasks; A computing unit, which is communicatively connected to the task management unit, adopts a DSA architecture, and includes at least one computing core, a first address translation module, and a shared cache module, wherein each computing core includes an independent on-chip cache area, and the first address translation module is configured to map the physical addresses of multiple available memory segments of the shared cache module to continuous memory space virtual addresses according to a memory virtualization method, wherein the memory virtualization method includes the steps of: configuring an address translation table, and performing conversion between virtual addresses and physical addresses of the memory through the address translation table; as well as An external storage unit is communicatively connected to the computing unit via a second address translation module, wherein the second address translation module is configured to map the program code of the computing task to a specified virtual address space according to a memory virtualization method, wherein the memory virtualization method comprises the steps of configuring an address translation table, and performing conversion between the virtual address and the physical address of the memory via the address translation table.

2. The deep learning system according to claim 1, characterized in that Configuring the address translation table includes the following steps: At least one address mapping group is set, wherein the setting of each address mapping group includes: setting a starting virtual address, a physical address corresponding to the starting virtual address, and an address length included in the address mapping group.

3. The deep learning system according to claim 2, characterized in that In each address mapping group, the mapping relationship between the virtual address and the physical address is unique.

4. The deep learning system according to claim 2, characterized in that The range of addresses in each address mapping group overlaps with the range of addresses in the remaining N address mapping groups, where N is an integer greater than or equal to zero.

5. The deep learning system according to claim 2, wherein: Configuring the address translation table also includes the following steps: Configure the priority of each address mapping group. The priority of each address mapping group is different.

6. The deep learning system according to claim 2, characterized in that Configuring the address translation table also includes the following steps: Set the enable status of each address mapping group.

7. The deep learning system according to claim 2, characterized in that The conversion between the virtual address and the physical address of the memory is realized by the address conversion table, comprising the steps of: Determine the address mapping group corresponding to the input address; and Address translation is performed according to the information of the corresponding address mapping group.

8. The deep learning system according to claim 7, characterized in that: Determining the address mapping group corresponding to the input address includes the steps of: The input address is compared with the address range of each address mapping group one by one according to the priority, and the first address mapping group that successfully matches is used as the corresponding address mapping group, wherein a successful match means that the input address is within the address range of the address mapping group.

9. A task switching method for a deep learning system according to any one of claims 1 to 8, characterized in that: Includes steps: After receiving a computing task with a higher priority, the task management unit notifies the computing unit to perform on-site switching; After the current calculation of each core is completed, the calculation unit saves the calculation result of each core to the shared cache module, and sends the memory occupancy of the shared cache module to the task management unit; The task management unit determines whether the remaining memory of the shared cache module meets the memory requirement of the computing task: If the conditions are met, configuring the first address translation module to map the physical address of the remaining memory to a continuous virtualized address; as well as If not, first move at least part of the data in the shared cache module to the external storage module so that the remaining memory of the shared cache module is not less than the memory requirement of the computing task, and then configure the first address translation module to map the physical address of the remaining memory to a continuous virtualized address; The computing unit saves the scene of each core and sends a saving completion signal to the task management unit; The task management unit configures the second address translation module to map the computing task to a specified virtual address space; The task management unit notifies the computing unit to execute a computing task; as well as After the computing task is completed, the computing unit and the task management unit perform on-site recovery.

Citation Information

Patent Citations

  • Virtualization supporting guest operating systems using memory protection units

    CN104956342A