Methods, apparatuses, electronic devices, storage media and products for driving computing devices

By combining power management and memory status management in the driver of computing devices, the compatibility and reliability issues of dynamic GPU resource management are resolved, enabling seamless recovery of computing tasks and dynamic takeover of resources, thereby improving service reliability and resource utilization.

CN121092227BActive Publication Date: 2026-03-06ALIBABA CLOUD COMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511650805.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-03-06
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Traditional CPU container saving and restoration techniques cannot be directly applied to GPU resources, affecting service reliability and resource utilization. Furthermore, existing virtualization methods have poor applicability to ordinary hardware environments.

Method used

By combining power management and video memory status management mechanisms in the driver of the computing device, the dynamic saving and restoration of the video memory status of the computing device can be realized, supporting the suspension, resource reclamation and resumption of computing tasks, without relying on dedicated virtualization hardware.

Benefits of technology

While ensuring data consistency, it enables seamless recovery of computing tasks and dynamic takeover of resources, improving service reliability and resource utilization, and supporting elastic scheduling and fault migration of computing devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121092227B_ABST
    Figure CN121092227B_ABST
Patent Text Reader

Abstract

This application provides a driving method, apparatus, electronic device, storage medium, and product for a computing device, relating to the field of computer computing. By combining a power sleep process with a video memory state saving mechanism, dynamic saving of the video memory state of a computing device is achieved without relying on dedicated virtualization hardware. This supports the suspension of computing tasks, resource reclamation, and subsequent recovery while ensuring data consistency. By combining a power wake-up process with a video memory state recovery mechanism, complete reconstruction of the video memory state of a computing device is achieved without relying on dedicated virtualization hardware. This supports seamless recovery of computing tasks, dynamic takeover of resources, and cross-device recovery while ensuring execution context consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a driving method, apparatus, electronic device, storage medium, and product for a computing device. Background Technology

[0002] With the rapid development of applications such as artificial intelligence, big data analytics, and high-performance computing, computing devices such as Graphics Processing Units (GPUs), as key parallel computing acceleration units, are being deployed on a continuously expanding scale in cloud computing platforms, data centers, and virtualization environments. Through efficient resource utilization and flexible scheduling, costs can be optimized, resource utilization rates improved, and service reliability enhanced.

[0003] Currently, containerization and virtualization technologies are widely used for resource management. However, traditional CPU container saving and restoration techniques cannot be directly applied to GPU resources, affecting service reliability and resource utilization. Virtualization resource management methods rely on specific hardware virtualization capabilities, making them difficult to apply to ordinary hardware environments. Summary of the Invention

[0004] This application provides a driving method for a computing device, a driving device for a computing device, an electronic device, a computer-readable storage medium, and a computer program product to alleviate or solve one or more technical problems existing in the prior art.

[0005] In a first aspect, embodiments of this application provide a method for driving a computing device, comprising: in response to a power wake-up request for a target computing device, sending a power wake-up command to the target computing device, the power wake-up command being used to wake up the target computing device; in response to the target computing device being woken up, sending a restore video memory command to the target computing device, the restore video memory command being used to instruct the target computing device to read target video memory data from a storage device, the target video memory data being obtained by writing video memory content generated by the source computing device during the execution of the target computing task queue into the storage device before deleting the target computing task queue; and, if the target computing device has read the target video memory data, reconstructing the target computing task queue so that the target computing device can continue to execute unfinished tasks in the target computing task queue based on the target video memory data.

[0006] Secondly, embodiments of this application provide a method for driving a computing device, comprising: in response to a power sleep request for a source computing device, sending a save video memory command to the source computing device, the save video memory command being used to instruct the video memory content generated by the source computing device during the execution of a target computing task queue to be written to a storage device; if the video memory content has been written to the storage device, deleting the target computing task queue; and in response to the deletion of the target computing task queue, sending a power sleep command to the source computing device, the power sleep command being used to instruct the source computing device to enter a sleep state.

[0007] Thirdly, embodiments of this application provide a driving device for a computing device, including a power management module. The power management module is configured to: in response to a power sleep request for a source computing device, send a save video memory command to the source computing device, the save video memory command instructing the video memory content generated by the source computing device during the execution of a target computing task queue to be written to a storage device to obtain target video memory data; if the video memory content has been written to the storage device, delete the target computing task queue; in response to the deletion of the target computing task queue, send a power sleep command to the source computing device, the power sleep command instructing the source computing device to enter a sleep state; in response to a power wake-up request for a target computing device, send a restore video memory command to the target computing device, the restore video memory command instructing the target computing device to read the target video memory data from the storage device; if the target computing device has read the target video memory data, rebuild the target computing task queue so that the target computing device can continue to execute unfinished tasks in the target computing task queue based on the target video memory data.

[0008] Fourthly, embodiments of this application provide a method for driving a computing device, comprising: sending a save video memory command to a source computing device, the save video memory command being used to instruct the video memory content generated by the source computing device during the execution of a target computing task queue to be written to a storage device to obtain target video memory data; and deleting the target computing task queue if the video memory content has been written to the storage device; and / or sending a restore video memory command to a target computing device, the restore video memory command being used to instruct the target computing device to read the target video memory data from the storage device; and rebuilding the target computing task queue if the target computing device has read the target video memory data, so that the target computing device can continue to execute unfinished tasks in the target computing task queue based on the target video memory data.

[0009] Fifthly, embodiments of this application provide a method for driving a computing device, applied to a management and control end. The method includes: in response to a source computing device executing a target computing task queue being in an abnormal state, sending a power sleep request to the source computing device, wherein the driving device is configured to put the source computing device into a sleep state based on the method described in the second aspect; in response to the source computing device entering the sleep state, sending a power wake-up request to the target computing device, wherein the driving device is further configured to wake up the target computing device based on the method described in the first aspect, so that the target computing device continues to execute unfinished tasks in the target computing task queue.

[0010] In a sixth aspect, embodiments of this application provide a method for driving a computing device, applied to a management and control end. The method includes: in response to a source computing device that needs to reclaim a queue of execution target computing tasks, sending a power sleep request to a driving device for the source computing device, wherein the driving device is used to cause the source computing device to enter a sleep state based on the method described in the second aspect.

[0011] In a seventh aspect, embodiments of this application provide a method for driving a computing device, applied to a management and control end. The method includes: in response to the need to pause a target computing task queue, sending a power sleep request to a driving device for a source computing device executing the target computing task queue, wherein the driving device is configured to put the source computing device into a sleep state based on the method described in the second aspect; and / or, in response to the need to rebuild the target computing task queue, sending a power wake-up request to the driving device for the target computing device, wherein the driving device is further configured to wake up the target computing device based on the method described in the first aspect, so that the target computing device continues to execute unfinished tasks in the target computing task queue.

[0012] Eighthly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the methods of embodiments of this application when executing the computer program.

[0013] Ninthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of any one of the embodiments of this application.

[0014] In a tenth aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method of any one of the embodiments of this application.

[0015] According to one technical solution of this application, by combining the power sleep process with the video memory state saving mechanism, dynamic saving of the video memory state of the computing device is achieved without relying on dedicated virtualization hardware. This supports the suspension of computing tasks, resource reclamation, and subsequent recovery while ensuring data consistency. According to another technical solution of this application, by combining the power wake-up process with the video memory state recovery mechanism, complete reconstruction of the video memory state of the computing device is achieved without relying on dedicated virtualization hardware. This supports seamless recovery of computing tasks, dynamic resource takeover, and cross-device recovery while ensuring execution context consistency, thereby improving service reliability and resource utilization. Thus, this application embodiment achieves complete lifecycle management of computing device resources, covering key operations such as suspension, power failure, wake-up, and cross-card recovery. The entire state saving and recovery process is completed in the kernel-mode driver device, without interfering with the user-mode application space or relying on dedicated virtualization hardware, possessing good system transparency and compatibility.

[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0017] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this application and should not be construed as limiting the scope of this application.

[0018] Figure 1 This diagram illustrates the architecture of the host system 100 provided in an embodiment of this application.

[0019] Figure 2 A schematic diagram illustrating the distributed reasoning principle of the hybrid expert model is shown.

[0020] Figure 3 A flowchart illustrating a driving method 300 for a computing device provided in an embodiment of this application is shown.

[0021] Figure 4 This is a structural block diagram of a driving device 400 for a computing device provided in an embodiment of this application.

[0022] Figure 5A , Figure 5B and Figure 5C This diagram illustrates an application example of dynamic management of computing resources provided in an embodiment of this application.

[0023] Figure 6A block diagram of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0024] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the concept or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0025] First, the terms used below will be explained.

[0026] GPU: A processor specifically designed for graphics rendering and highly parallel computing tasks. It features a massively parallel architecture and high-speed memory access capabilities, and is widely used in fields such as artificial intelligence, deep learning, scientific computing, and image processing.

[0027] Checkpoint-Restore (CR) refers to the technology of persistently saving the complete running state (including memory data, register state, storage content, device context, etc.) of a running process, virtual machine, or container, and restoring it to the same or compatible environment at a later point in time to continue execution. It is often used in fault tolerance, migration, and elastic scheduling scenarios.

[0028] GPU driver: refers to the low-level software component used to operate and manage GPU hardware devices. It is responsible for realizing communication and control between the operating system and GPU hardware devices, covering functions such as resource allocation, task scheduling, device initialization, error handling, performance monitoring and function expansion. GPU drivers are usually developed and maintained by hardware manufacturers and are the foundation for realizing the full performance of the GPU and supporting advanced features such as virtualization and power management.

[0029] Power management refers to the technology of intelligently controlling and scheduling the power status of hardware devices such as GPUs, including operations such as power-on, power-off, and power consumption management, aiming to balance computing performance and energy consumption, and achieve energy saving and efficient utilization.

[0030] Virtual GPU (vGPU): A technology that uses a hardware and software collaboration mechanism to divide a single physical GPU hardware resource into multiple independent virtual GPU instances. Each virtual GPU can be independently allocated to different virtual machines or containers, supporting multi-tenant sharing of GPU acceleration capabilities and achieving resource isolation and on-demand allocation.

[0031] Single Root I / O Virtualization (SR-IOV): A hardware virtualization technology based on the high-speed serial computer expansion bus standard (Peripheral Component Interconnect Express, PCIe) bus. It allows a single physical I / O device (such as a GPU or network card) to be divided into one physical function (PF) and multiple virtual functions (VF) at the hardware level. Each virtual function can be independently assigned to different virtual machines or containers, thereby achieving high-performance device passthrough and resource isolation.

[0032] Virtual Function I / O (VFIO): A user-space device access framework provided by the operating system kernel, which supports the direct allocation of secure access permissions for physical I / O devices to virtual machines or containers, enabling device passthrough and low-latency access to hardware resources. It is often used in conjunction with SR-IOV technology to support high-performance virtualization and device migration.

[0033] Direct Device Assignment (VFIO Passthrough): This refers to the technology of directly assigning access permissions of physical devices (such as GPUs) to virtual machines or containers through the Virtualization I / O Framework (VFIO), enabling them to bypass the traditional virtualization layer and directly manipulate hardware resources with near-native performance. It is widely used in high-performance computing and cloud platform scenarios.

[0034] Hot migration refers to the technology of migrating the running state, memory data and resources of a computing instance (such as a virtual machine or container) to another physical host without stopping the machine or with only a brief interruption while the instance is running continuously. It is used to achieve load balancing, hardware maintenance and fault avoidance.

[0035] Borrowing and returning cards: refers to a mechanism that temporarily allocates (borrows) hardware resources such as GPUs to computing instances that need them based on actual computing load requirements, and actively reclaims (returns) them to the resource pool when the task is completed or the resources are idle. It supports on-demand allocation and elastic reuse of resources, effectively improving the overall resource utilization rate.

[0036] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.

[0037] In real-world applications, user demands for GPU resources are highly dynamic and diverse. For example, developers may need to suspend and resume their development environments at any time, operations and maintenance personnel often need to perform fault migration or elastic scaling, some low-frequency computing tasks involve long waiting times and extremely low GPU utilization, while cloud platforms aim to optimize resource allocation and flexibly respond to resource peaks and troughs and multi-tenant requirements. All of these necessitate reliable dynamic GPU resource allocation solutions.

[0038] Traditional dynamic allocation of GPU resources requires the flexible allocation, reclamation, and switching of physical GPUs or their computing resources (such as computing cores, video memory, and encoding / decoding units) among different computing instances (such as containers and virtual machines) based on system load and application requirements. However, because GPUs have independent memory spaces, dedicated command queues, and complex hardware execution contexts, their state management involves the collaborative saving and reconstruction of multi-dimensional information such as video memory data, task queues, and hardware register configurations. Therefore, traditional CPU container saving and recovery techniques are difficult to directly apply to GPU resource management.

[0039] One solution utilizes virtualization extension technology to achieve online migration of GPU resources. However, this relies on hardware and protocol support such as SR-IOV and VFIO migration interfaces, capturing the GPU device state and performing hot migration operations during virtual machine or container runtime. This solution can migrate the GPU context from one host environment to another during runtime, enabling continuous execution of computing tasks. However, its drawbacks include strong dependence on the underlying hardware architecture, limiting its applicability to specific GPU models that support SR-IOV and VFIO protocols, resulting in poor versatility. Furthermore, this solution typically only supports online hot migration and cannot effectively support scenarios such as offline state preservation, cold recovery, and resource reclamation and reallocation, thus limiting its application scope in elastic resource management.

[0040] Another solution involves saving and restoring GPU task states through user-mode runtime library interception. This approach hijacks the runtime application programming interface (API) of the runtime library of the parallel computing platform and programming model to save and restore the state of GPU computing tasks. However, this solution relies on replacing or injecting user-mode library files and requires the deployment of additional middleware frameworks, significantly increasing system integration complexity. Furthermore, this solution is sensitive to the application's runtime environment, exhibits poor compatibility and stability, and struggles to achieve general support for services and applications.

[0041] In summary, the technologies for achieving dynamic management of GPU resources are limited by hardware virtualization capabilities or face issues of software compatibility and deployment complexity, making it difficult to provide efficient, reliable, and universal solutions in a common hardware environment.

[0042] In view of this, embodiments of this application provide a driving method for a computing device, a driving device for a computing device, an electronic device, a computer-readable storage medium, and a computer program product. By combining the power management mechanism and state management (save and restore) mechanism of the computing device (including the GPU) in the driver of the computing device, dynamic management of computing resources can be achieved without relying on dedicated virtualization hardware capabilities. A detailed description follows.

[0043] Figure 1 This diagram illustrates the architecture of a host system 100 provided in an embodiment of this application. Figure 1 As shown, the host system 100 includes a target application, a drive device, a storage device, and a computing device, which includes a source computing device and / or a target computing device.

[0044] The target application is an application running in the user space of the operating system, responsible for submitting computational tasks and receiving the execution results. The target application creates and manages a queue of computational tasks by calling the driver. For example, in scenarios such as deep learning inference, scientific computing, or graphics rendering, the target application can distribute a large number of parallel computational tasks to the computing device for execution.

[0045] The driver is a driver module located in the operating system kernel mode, used to drive computing devices. In the embodiments of this application, the driver combines the functions of state management (saving and restoring) and power management of the computing device.

[0046] Storage devices can be non-volatile storage media used for persistent data storage, including but not limited to solid-state drives (SSDs), non-volatile memory express (NVMe) hard drives, hard disk drives (HDDs), etc. Storage devices can also be network storage modules, such as storage volumes in distributed file systems, object storage, or cloud storage services. Network storage modules can connect to the host system 100 via a high-speed network interface to achieve data storage across nodes, racks, or data centers. In this embodiment, the storage device is used to store the video memory content of the computing device.

[0047] Computing devices are the basic units for performing computing tasks, including but not limited to GPUs, general-purpose graphics processing units (GPGPUs), neural network processing units (NPUs), tensor processing units (TPUs), or dedicated accelerators. In this embodiment, the computing device is used to execute computing tasks sent by the target application and return the execution results of the computing tasks to the target application. Depending on the application scenario, the computing device bound to the target application can be dynamically switched. For example, the computing device corresponding to the target application can be switched from the source computing device to the target computing device. Therefore, the source computing device can be understood as the computing device that previously executed the computing task, while the target computing device can be understood as the computing device that replaces the source computing device to continue executing the computing task.

[0048] For example, in the task initialization process, the target application calls the driver to create a target computing task queue; through the target computing task queue, the target application sends computing tasks to the source computing device, the source computing device receives the computing tasks, performs the computing, and feeds back the execution results to the target application.

[0049] During the save process, the control unit triggers a power sleep request for the source computing device. The driver responds to the power sleep request by sending a save video memory command to the source computing device. Upon receiving the save video memory command, the source computing device writes the video memory content to the storage device, thus saving the video memory state. After confirming that the video memory content has been written to the storage device, the driver deletes the target computing task queue and sends a power sleep command to the source computing device, driving the source computing device into sleep mode, stopping the execution of the target application's computing tasks, and releasing its computing resources.

[0050] During the recovery process, the control unit triggers a power wake-up request for the target computing device. The driver responds to the power wake-up request and sends a restore video memory command to the target computing device. Upon receiving the restore video memory command, the target computing device reads the target video memory data (the video memory content written by the source computing device) from the storage device, completing the restoration of the video memory state. With the target computing device having read the target video memory data, the driver rebuilds the target computing task queue. Through the target computing task queue, the target application sends subsequent computing tasks to the target computing device. The target computing device continues to execute computing tasks based on the target video memory data and feeds back the execution results of the computing tasks to the target application.

[0051] In an application scenario requiring elastic suspension / wake-up of computing device resources, when the system load decreases or other high-priority tasks exist, the control terminal can trigger a power sleep request for the source computing device to the driver, initiating a save process and enabling on-demand release of computing resources and energy saving. When the load increases or the computing task needs to be restarted, the control terminal can trigger a power wake-up request for the target computing device to the driver, initiating a recovery process. The target video memory data is loaded from the storage device, and the computing task is resumed, thus achieving rapid reuse and dynamic allocation of computing resources. The source and target computing devices can be the same device or different devices.

[0052] In a specific application scenario, this solution not only supports local suspension and wake-up on the same computing device, but also enables the migration and restoration of video memory state to different computing devices. It breaks through the functional limitations of traditional power management solutions (such as some manufacturers only supporting device-level sleep and wake-up), and is especially suitable for task migration between computing devices of the same model or with compatible functions.

[0053] In a fault migration application scenario, if the source computing device is in an abnormal state, such as experiencing a hardware failure or communication anomaly, a power sleep request can be triggered through the management terminal to enter the save process. Subsequently, a power wake-up request can be triggered for the target computing device through the management terminal to enter the recovery process. This rebuilds the computing environment by restoring the video memory data, automatically migrating the computing tasks to the target computing device and improving the system's fault tolerance and availability.

[0054] In a card-based computing scenario, the system can dynamically allocate computing resources based on actual computing load requirements. For example, when a computing instance (such as a virtual machine or container) needs to perform a high-load computing task, it can borrow (i.e., allocate) an idle GPU as its computing device. Once the computing task is completed or the current computing instance's demand for the GPU decreases, the system can automatically trigger a resource reclamation process. In this process, the management terminal first issues a power sleep request for the GPU (the source computing device), enters the save process, and releases the computing resource. Then, the GPU is marked as idle and returned to the resource pool, waiting to be borrowed by the next computing instance with demand.

[0055] In a scenario involving rapid batch startup, multiple computing tasks can be submitted to computing devices for execution in the initial stage. When the system needs to pause these computing tasks in batches, a unified power sleep request can be used to put all source computing devices into sleep mode and save their respective video memory contents to storage devices. When a restart is required later, simply triggering a power wake-up request on these devices as the target computing devices enables rapid startup of a large number of computing tasks.

[0056] In a resource pooling application scenario, multiple computing devices are abstracted into a shared resource pool, which is uniformly managed and scheduled by the control terminal. By triggering power sleep and power wake-up requests to computing devices, dynamic allocation and reclamation of computing resources can be realized, and the migration of computing tasks between different computing devices can be supported, thereby improving resource utilization and scheduling flexibility.

[0057] From an application perspective, the technical solution of this application provides a novel, general, flexible, and easy-to-deploy computing resource management method that enables on-demand allocation, state persistence, rapid recovery, and elastic scheduling of computing resources. In scenarios such as multi-tenant cloud environments, mixed AI training and inference workloads, and edge computing resource constraints, it can effectively improve GPU resource utilization, ensure the continuity of critical services, enhance system resource elasticity and fault tolerance, and help promote the development of cloud computing, artificial intelligence platforms, and data center infrastructure towards higher efficiency and greater adaptability.

[0058] For example, the control terminal can be deployed as a standalone management system, cloud platform controller, cluster scheduler, or job scheduling system.

[0059] For example, the source computing device and the target computing device can be deployed simultaneously in the host system 100, or they can be deployed on different physical hosts or server nodes, depending on the application scenario requirements. For instance, in a cross-data center disaster recovery application scenario, the source computing device is located in the main data center, and the target computing device is located in an off-site backup center. When the main center fails, a power wake-up request can be remotely triggered to migrate the computing tasks to the computing devices in the backup center for continued execution, thus ensuring service continuity.

[0060] This application embodiment combines the power management mechanism and state management mechanism (save and restore the video memory state) of the computing device (such as the GPU) in the driver device, which can realize the suspension, power-off, wake-up and cross-card recovery of the computing device without relying on dedicated virtualization hardware, and has good compatibility, scalability and practicality.

[0061] It is understood that other components of the host system 100 described above can employ various technical solutions now and in the future known to those skilled in the art, and will not be described in detail here. Furthermore, the deployment form or other configurations of the host system 100, computing devices, and storage devices may differ in different application scenarios, and can be flexibly adjusted according to the application scenario.

[0062] It should be noted that the application scenarios or examples provided in this application are for ease of understanding, and this application does not specifically limit the application of the technical solutions. Furthermore, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0063] The technical solution of this application and how it solves the aforementioned technical problems are described in detail below with specific embodiments. The listed specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0064] Figure 2 A flowchart illustrating a driving method 200 for a computing device according to an embodiment of this application is shown. This method 200 can be applied to a driving device for a computing device, such as a device driven by… Figure 1 The drive unit in the host system 100 shown executes the operation. For example... Figure 2 As shown, the method 200 includes steps S201, S202 and S203.

[0065] Step S201: In response to a power sleep request for the source computing device, a save video memory command is sent to the source computing device. The save video memory command is used to instruct the video memory content generated by the source computing device during the execution of the target computing task queue to be written to the storage device.

[0066] A power sleep request is a request used to trigger a computing device to enter a sleep state. Power sleep requests can be initiated by the management system. For example, a power sleep request can be triggered by the management system based on the actual needs of the application scenario (refer to scenarios such as elastic suspension, fault migration, and card borrowing / returning mentioned earlier).

[0067] In response to a power sleep request, the drive device sends a save video memory command to the source computing device, thereby writing the video memory content generated by the source computing device during the execution of the target computing task queue (as described above) to the storage device. For example, this video memory content reflects the video memory state of the source computing device, including but not limited to: the context of the target application, the execution state of the target computing task queue, and the running state of the source computing device.

[0068] As an example, the driver invokes the power management module to initiate a command to save the video memory. For instance, the power management module uses the hardware copy engine of the source computing device to write the video memory contents to a storage device, obtaining the target video memory data. The target video memory data is the stored data of the video memory state of the source computing device, which can be used for subsequent recovery operations.

[0069] Step S202: If the video memory content has been written to the storage device, delete the target computing task queue.

[0070] For example, the driver can confirm that the video memory content has been successfully written to the storage device by receiving a write completion interrupt signal from the source computing device; alternatively, the driver can detect whether the write operation is complete by polling the write completion status register of the storage device. Furthermore, the driver can obtain information about whether the video memory content has been successfully written to the storage device through a callback function or an event notification mechanism.

[0071] After confirming that the video memory content has been written to the storage device, the driver performs a deletion operation on the target computing task queue. For example, the power management module completes the deletion by clearing or freezing the target computing task queue and marking it as invalid. This ensures that the target video memory data will not be overwritten or corrupted by new computing tasks before the source computing device enters sleep mode, avoiding inconsistencies between the target video memory data and the actual operating state.

[0072] Step S203: In response to the deletion of the target computing task queue, a power sleep command is sent to the source computing device. The power sleep command is used to instruct the source computing device to enter a sleep state.

[0073] Sleep mode is an energy-saving mode in which computing devices enter a low-power state. In this mode, the computing device shuts down or reduces unnecessary functions to save power consumption. Sleep mode typically includes different levels, such as hibernation (almost completely powered off, but can still be woken up by software commands) and power-off mode (requiring a physical signal to resume).

[0074] For example, after confirming that the task queue has been successfully deleted, the drive device sends a power sleep command to the source computing device, triggering it to enter a sleep state.

[0075] According to the technical solution of the embodiments of this application, by combining the power management process with the memory state saving mechanism, dynamic saving of the memory state of computing devices is realized without relying on dedicated virtualization hardware. It can support the suspension of computing tasks, resource reclamation and subsequent recovery on the basis of ensuring data consistency, and provides reliable technical support for the elastic scheduling, fault migration and resource pooling of computing devices such as GPUs.

[0076] In one implementation, before sending the save video memory command to the source computing device in step S201, method 200 may further include: in response to a task queue creation request from the target application, sending a task command to the source computing device and recording the task command, the task command being used to create a target computing task queue.

[0077] The task queue creation request instructs the target application to establish an execution channel on the computing device for submitting and scheduling computing tasks. For example, when the target application loads a deep learning model and is ready to start an inference or training task, it can trigger this request by calling an API interface.

[0078] The driver responds to the task queue creation request by sending the corresponding task command to the source computing device to complete the initialization of the target computing task queue. The task command refers to the control instructions generated by the driver and sent to the source computing device to create the target computing task queue, such as stream creation commands and command queue allocation commands.

[0079] Meanwhile, the driver maintains a metadata structure associated with the task queue in kernel mode and records key parameter information of the task command, including but not limited to: task queue identifier (ID), execution context (Context), etc.

[0080] By recording task commands, the driver can accurately reconstruct the structure and attributes of the target computing task queue during the subsequent recovery phase without relying on the target application to re-initiate the original API calls. This provides fundamental support for achieving transparent state recovery without application involvement or awareness, improving the system's automation capabilities and response efficiency in scenarios such as resource migration, suspension, and wake-up.

[0081] In one implementation, before sending a task command to the source computing device, method 200 may further include: sending an initialization command to the source computing device and recording the initialization command, the initialization command being used to initialize the source computing device.

[0082] Specifically, when a target application accesses the source computing device for the first time, or when the system first enables the source computing device, the driver needs to execute the source computing device's initialization process. This includes, for example, detecting device hardware information, configuring registers, establishing memory management unit mapping, and initializing interrupt service mechanisms. The initialization process is completed through a series of initialization commands.

[0083] While executing these initialization commands, the driver also records them, for example, by recording them in the driver's state management module, forming an initialization context snapshot of the source computing device.

[0084] By recording initialization commands, the drive unit can quickly initialize the target computing device during the subsequent recovery phase and automatically complete the environment reconstruction based on the recorded initialization commands, thereby shortening the recovery delay and improving the compatibility and reliability of cross-device migration.

[0085] Figure 3 A flowchart illustrating a driving method 300 for a computing device according to an embodiment of this application is shown. This method 300 can be applied to a driving device for a computing device, such as a device driven by… Figure 1 The drive unit in the host system 100 shown executes the operation. For example... Figure 3 As shown, the method 300 includes steps S301, S302 and S303.

[0086] Step S301: In response to the power wake-up request for the target computing device, a power wake-up command is sent to the target computing device to wake up the target computing device.

[0087] A power wake-up request is a control command used to trigger a target computing device to resume from sleep mode to normal operation. It can be initiated by the management and control terminal based on task recovery needs, resource scheduling strategies, or user operations. For example, in application scenarios such as elastic wake-up, fault migration, card borrowing and returning, or batch startup, when it is necessary to resume previously suspended computing tasks, the management and control terminal sends a power wake-up request to the driver device.

[0088] In response to the request, the driver sends a power-on command to the target computing device, switching it from sleep mode to normal operation. During the wake-up process, the target computing device is powered on again, the clock signal is restored, and basic hardware functional modules are initialized to prepare for subsequent memory recovery and task reconstruction.

[0089] Step S302: In response to the target computing device being woken up, a restore video memory command is sent to the target computing device. The restore video memory command is used to instruct the target computing device to read the target video memory data from the storage device. The target video memory data is obtained by writing the video memory content generated by the source computing device during the execution of the target computing task queue to the storage device before deleting the target computing task queue.

[0090] For example, the driver device confirms that the target computing device has completed power wake-up and entered an operable state by detecting the device ready interrupt signal returned by the target computing device or polling its status register. Subsequently, the driver device calls the power management module to issue a restore video memory command. This command instructs the target computing device to read the target video memory data previously saved by the source computing device from a preset storage location on the storage device, and write it into its own video memory space through a hardware copy engine, thus restoring the video memory state.

[0091] Step S303: If the target computing device has read the target video memory data, rebuild the target computing task queue so that the target computing device can continue to execute the unfinished tasks in the target computing task queue based on the target video memory data.

[0092] For example, the driver confirms that the target video memory data has been successfully loaded into the video memory by receiving a read completion interrupt from the target computing device or by an event notification mechanism. Subsequently, the driver initiates the task queue reconstruction process.

[0093] Reconstructing the target computing task queue refers to recreating a computing task scheduling channel on the target computing device with the same structure and attributes as the source computing device, enabling it to receive and execute computing tasks. After the target computing task queue is reconstructed, the target application can continue to submit subsequent computing tasks through the original interface, and the target computing device will continue the computing process based on the restored target memory data, achieving seamless connection of computing tasks.

[0094] For example, the target computing task queue can be reconstructed by directly configuring the metadata information of the task queue on the target computing device. For instance, a new task queue can be built on the target device based on saved queue attributes (such as stream ID, context handle, memory space range, scheduling priority, etc.) and bound to the restored video memory data.

[0095] According to the technical solution of the embodiments of this application, by combining the power wake-up process with the memory state recovery mechanism, the complete reconstruction of the memory state of the computing device is realized without relying on dedicated virtualization hardware. It can support seamless recovery of computing tasks, dynamic takeover of resources and cross-device recovery while ensuring the consistency of execution context, and provides reliable technical support for the elastic scheduling, fault migration and resource pooling of computing devices such as GPUs.

[0096] In one implementation, rebuilding the target computing task queue in step S303 may include: replaying task commands to the target computing device to rebuild the target computing task queue, wherein the task commands are commands sent to the source computing device and recorded for creating the target computing task queue.

[0097] For example, the driver can read previously recorded task commands from the kernel-mode metadata structure, replay the task commands, and send them to the target computing device. For instance, the driver can replay commands such as "stream creation command" and "command queue allocation instruction" to reconstruct a task scheduling environment on the target computing device that is consistent with the source computing device.

[0098] By replaying task commands, the driver can accurately reproduce the topology and scheduling strategy of the target computing task queue on the target computing device, ensuring the consistency of execution behavior after the computing task is restored, and supporting application scenarios such as cross-device migration and resource reuse. This process does not require the target application to call the API interface again, achieving seamless task queue reconstruction.

[0099] In one embodiment, before sending the restore video memory command to the target computing device in step S302, method 300 may further include: replaying an initialization command to the target computing device, the initialization command being a command sent to the source computing device and recorded for initializing the source computing device; and sending the restore video memory command to the target computing device when the target computing device has completed initialization.

[0100] When the target computing device is a newly connected device or a device that has not been used for a long time, device-level initialization must be completed first. For example, the driver reads a series of previously recorded initialization commands from local storage or a configuration database to obtain an initialization context snapshot of the source computing device. By replaying these initialization commands, the target computing device is made to have an operating environment consistent with that of the source computing device.

[0101] After confirming that the target computing device has completed initialization (e.g., receiving an initialization completion interrupt or a status register returning a ready flag), the driver then executes step S302 to send a command to restore video memory.

[0102] By replaying the initialization command, the driver can ensure that the target computing device has the hardware support and operating environment required to execute the original task, avoiding memory recovery failure or task execution abnormalities due to differences in device configuration, and significantly improving the robustness and scalability of cross-device and cross-model migration.

[0103] Figure 4 This diagram illustrates the structure of a drive unit 400 for a computing device according to an embodiment of this application. Figure 4 As shown, the drive device 400 includes a power management module 401.

[0104] The power management module 401 is configured to: in response to a power sleep request from the source computing device, send a save video memory command to the source computing device, the save video memory command instructing the source computing device to write the video memory content generated during the execution of the target computing task queue to a storage device to obtain the target video memory data; if the video memory content has been written to the storage device, delete the target computing task queue; in response to the deletion of the target computing task queue, send a power sleep command to the source computing device, the power sleep command instructing the source computing device to enter a sleep state; in response to a power wake-up request from the target computing device, send a restore video memory command to the target computing device, the restore video memory command instructing the target computing device to read the target video memory data from the storage device; if the target computing device has read the target video memory data, rebuild the target computing task queue so that the target computing device can continue to execute the unfinished tasks in the target computing task queue based on the target video memory data.

[0105] According to the driver device 400 provided in the embodiments of this application, the power management module 401 is enabled to trigger the state saving and restoration. By combining the power management mechanism and the state management mechanism (saving and restoring the video memory state) of the computing device in the driver device, the suspending, power-off, waking up and cross-card restoration of the computing device can be realized without relying on dedicated virtualization hardware, and it has good compatibility, scalability and practicality.

[0106] In one implementation, such as Figure 4 As shown, the drive device 400 also includes a status management module 402.

[0107] The power management module 401 is also configured to: in response to a task queue creation request from a target application, send a task command to the state management module, the task command being used to create a target computing task queue. The state management module 402 is configured to: record the task command and forward the task command to the source computing device; and / or, replay the task command to the target computing device to rebuild the target computing task queue.

[0108] A status management module 402 is set in the drive device 400 to realize the recording (saving) and playback (restoration) of task commands.

[0109] In one embodiment, the power management module 401 is further configured to: send an initialization command to the state management module 402, the initialization command being used to initialize the source computing device; the state management module 402 is further configured to: record the initialization command and forward the initialization command to the source computing device. And / or, the power management module 401 is further configured to: send a power wake-up command to the state management module 402, the power wake-up command being used to wake up the target computing device; the state management module 402 is further configured to: forward the power wake-up command to the target computing device, and in response to the target computing device being woken up, replay the initialization command to the target computing device; the power management module 401 is further configured to: send a restore video memory command to the target computing device when the target computing device has completed initialization.

[0110] A status management module 402 is set in the drive device 400 to realize the recording (saving) and playback (restoration) of initialization commands.

[0111] According to the technical solution provided in the embodiments of this application, by setting a power management module and a state management module in the driver device, dynamic saving and restoration of computing resources can be achieved based on the enhanced power management mechanism. This can significantly improve the utilization rate of computing resources and realize state saving and restoration across computing device hardware (i.e., "card replacement"). Unlike the power management module of traditional computing devices, which only supports basic power consumption control such as sleep and wake-up, this solution customizes and enhances the power management mechanism of the power management module, enabling it to work in conjunction with the state management module to achieve complete state persistence and cross-device migration.

[0112] The state management module, as the core module of the driver device, is responsible for managing the video memory state (video memory content, task commands, initialization commands, etc.) of the computing device, realizing functions such as persistence of computing task queues, data consistency verification, and cross-device data adaptation. Specifically, the state management module provides interfaces for capturing, serializing, and persisting data such as video memory content, target computing task queues, task commands, and initialization commands. By ensuring the consistency and integrity of this data during saving and restoring, it ensures that the saved data can be accurately restored on the target computing device. Through collaboration with the power management module, it triggers the saving process when the computing device is in hibernation or powered off, and triggers the restoration process when the computing device is woken up / restarted, thereby realizing state migration across computing devices or within the same computing device.

[0113] The following is combined with Figure 5A , Figure 5B and Figure 5CThis document introduces an application example of dynamic computing resource management based on a driver device 400. In this example, the driver device is a GPU driver, the source computing device is GPU1 hardware, and the target computing device is GPU2 hardware. The application example includes three processes: a task initialization process (such as...) Figure 5A As shown), the saving process (such as...) Figure 5B (as shown) and recovery process (as shown) Figure 5C (As shown).

[0114] like Figure 5A As shown, the task initialization process includes: GPU initialization phase, task queue creation phase, and task execution phase.

[0115] During the GPU initialization phase, the target application or management terminal initiates an initialization request for the GPU1 hardware through the GPU driver. First, the GPU driver calls the power management module to pass the GPU initialization command (a specific example of the initialization command described above) to the state management module (step S501). The state management module records this GPU initialization command for use in the subsequent recovery process (step S502) and forwards the GPU initialization command to the GPU1 hardware (step S503). After completing initialization, the GPU1 hardware returns the initialization execution result to the power management module (i.e., step S504). This completes the management of computing tasks and resource allocation.

[0116] During the task queue creation phase, after the target application starts, it creates a GPU task queue by calling the GPU driver (step S505). Specifically, the target application sends a request to the GPU driver to create the target computing task queue (the GPU task queue is a specific example of the target computing task queue described above). The GPU driver generates the corresponding GPU creation task command (step S506), and the state management module records this command (step S507). The GPU driver then forwards the command to GPU hardware 1 for execution (step S508). After GPU 1 hardware completes processing, it returns the execution result of the task command to the power management module (step S509). The GPU driver then sends the execution result back to the target application (step S510), completing the creation of the GPU task queue.

[0117] During the task execution phase, the target application sends a specific computation task to the GPU (i.e., step S511). The GPU1 hardware receives and executes the computation task, and returns the execution result of the computation task to the target application after completion (i.e., step S512).

[0118] like Figure 5B As shown, the save process includes a save triggering phase and a state capture and persistence phase. The save process is usually triggered before the GPU1 hardware is suspended or powered off.

[0119] During the save trigger phase, when the control terminal needs to reclaim resources or suspend the GPU1 hardware, it triggers a power sleep request for the GPU1 hardware (step S513). The GPU driver responds to the request and calls the power management module to start the save process.

[0120] During the state capture and persistence phase, the power management module sends a save memory command to the GPU1 hardware (step S514). Using the GPU hardware copy engine, the GPU1 hardware's memory contents (including application data, task queues, running status, etc.) are written to the storage device (step S515), obtaining the target memory data. After this operation, the GPU1 hardware returns the execution result (step S516), allowing the power management module to know that the GPU1 hardware has read the target memory data. Then, the power management module deletes the user-mode task queue (target computation task queue) by clearing or freezing it, and notifies the GPU1 hardware (step S517) to prevent interference from new computation tasks and ensure data consistency. The GPU1 hardware unbinds itself from the queue and returns the execution result to the power management module (step S518). The power management module issues a power sleep command to the GPU1 hardware (step S519), causing the GPU1 hardware to enter a low-power or power-off sleep state. After completing the save process, the GPU driver returns the execution result to the management terminal (step S520).

[0121] like Figure 5C As shown, the recovery process includes: wake-up and device preparation stage, status write-back and recovery stage, and task restart stage. The recovery process can be triggered after card replacement, wake-up, or migration.

[0122] During the wake-up and device preparation phase, the control unit triggers a power wake-up request for the GPU2 hardware (step S521), and the GPU driver notifies the power management module to start the recovery process. The power management module sends a power wake-up command to the status management module (step S522), which is forwarded by the status management module to the GPU2 hardware (step S523). After executing the command, the GPU2 hardware returns an execution confirmation to the status management module (i.e., step S524). Subsequently, the status management module replays the previously recorded GPU initialization command (i.e., step S525) to initialize the GPU2 hardware, ensuring that it has the same operating environment as the GPU1 hardware to adapt to and prepare for executing the computing tasks sent by the target application. The GPU2 hardware returns the hardware execution result to the status management module (i.e., step S526), ​​and the status management module returns an execution confirmation to the power management module (i.e., step S527).

[0123] During the state write-back and recovery phase, after receiving the execution return from the state management module, the power management module sends a restore memory command to the GPU2 hardware (step S528). The GPU2 hardware reads the previously saved target memory data from the storage device through the GPU hardware copy engine and writes the data back to its memory space (step S529). After completion, the GPU2 hardware returns the execution result to the power management module (step S530). The state management module replays the task commands (step S531), and these commands are resent to the GPU2 hardware for execution, thereby reconstructing the original user-mode task queue and its execution context, restoring the task execution interface with the target application, and then the GPU2 hardware returns the hardware execution result to the state management module (step S532). The state management module returns an execution confirmation to the power management module (step S533).

[0124] During the task restart phase, after confirming that the GPU2 hardware has read the target video memory data, the GPU driver returns the execution result to the management terminal (i.e., step S533). Subsequently, the target application sends a new computing task to the GPU2 hardware (step S535), which is executed by the GPU2 hardware. The GPU2 hardware continues to execute the computing task based on the target video memory data and returns the execution result (i.e., step S536), and the service process is seamlessly continued.

[0125] The entire recovery process requires no intervention from the target application, nor does it require reloading the model or resetting the context, achieving seamless recovery of user space.

[0126] In summary, the technical solution of this application effectively overcomes the limitations of related technologies. Firstly, compared to the SR-IOV hot migration solution, this solution, through enhanced power management mechanisms combined with save-and-restore mechanisms, enables computing device resources to perform suspend, power-off, wake-up, and cross-card recovery operations without relying on dedicated virtualization hardware capabilities. This not only expands the scope of application and improves deployment flexibility but also supports a wide range of computing device models; as long as a basic power management interface is available, this solution can be used without relying on SR-IOV or VFIO protocols. Furthermore, this solution is not limited to online hot migration but also enables offline state saving and restoration, meeting the needs of service scenarios such as elastic suspension, cold migration, and card borrowing / returning. This feature provides greater flexibility for resource scheduling in cloud computing environments, especially in multi-tenant environments, improving resource utilization and service quality.

[0127] Compared to saving and restoring GPU task states through user-space runtime library interception technology, this solution implements saving and restoring directly at the driver layer, avoiding interference or replacement of user-space applications or native libraries. This greatly reduces the complexity of deployment and maintenance, enhances compatibility and stability, and ensures native support for all services and applications without additional modification work.

[0128] This application embodiment also provides a method for driving a computing device, including: sending a save video memory command to a source computing device, the save video memory command instructing the source computing device to write the video memory content generated during the execution of a target computing task queue to a storage device to obtain target video memory data; deleting the target computing task queue after the video memory content has been written to the storage device; and / or sending a restore video memory command to a target computing device, the restore video memory command instructing the target computing device to read the target video memory data from the storage device; and rebuilding the target computing task queue after the target computing device has read the target video memory data, so that the target computing device can continue to execute unfinished tasks in the target computing task queue based on the target video memory data.

[0129] The trigger for sending a save video memory command to the source computing device can be either receiving a power sleep request for the source computing device or being triggered by other forms of requests (such as a state save request for the source computing device). The trigger for sending a restore video memory command to the target computing device can be either receiving a power wake-up request for the target computing device or being triggered by other forms of requests (such as a state restore request for the target computing device).

[0130] For specific implementation methods and technical effects, please refer to the relevant or similar descriptions above, which will not be repeated here.

[0131] This application also provides a method for driving a computing device, applied to a management and control end. The method includes: in response to a source computing device executing a target computing task queue being in an abnormal state, sending a power sleep request to a driving device for the source computing device, wherein the driving device is configured to put the source computing device into a sleep state based on method 300; in response to the source computing device entering the sleep state, sending a power wake-up request to a target computing device for the driving device, wherein the driving device is further configured to wake up the target computing device based on method 200, so that the target computing device can continue to execute unfinished tasks in the target computing task queue.

[0132] For example, in a fault migration application scenario, if the source computing device is in an abnormal state, the control terminal can send a power sleep request to the drive device for the source computing device, and send a power wake-up request to the drive device for the target computing device after the source computing device enters a sleep state, thereby migrating the target computing task queue from the source computing device to the target computing device.

[0133] This application embodiment also provides a driving method for a computing device, applied to a management and control end. The method includes: in response to a source computing device that needs to reclaim the execution target computing task queue, sending a power sleep request for the source computing device to a driving device, wherein the driving device is used to put the source computing device into a sleep state based on method 200.

[0134] For example, in a scenario where the computing resources of the source computing device need to be flexibly suspended, the control terminal can send a power sleep request to the driver for the source computing device.

[0135] For example, in a scenario involving card borrowing and returning or resource pooling, if it is necessary to return or reclaim the computing resources of the source computing device, the control terminal can send a power sleep request to the driver for the source computing device.

[0136] This application also provides a method for driving a computing device, applied to a management and control end. The method includes: in response to the need to pause a target computing task queue, sending a power sleep request to a driving device for a source computing device executing the target computing task queue, wherein the driving device is configured to put the source computing device into a sleep state based on method 300; and / or, in response to the need to rebuild the target computing task queue, sending a power wake-up request to a target computing device, wherein the driving device is further configured to wake up the target computing device based on method 200, so that the target computing device continues to execute unfinished tasks in the target computing task queue.

[0137] For example, in a batch rapid startup application scenario, multiple computing tasks can be submitted to the computing device for execution in the initial stage, forming a target computing task queue. When the system needs to pause these computing tasks in batches, the control terminal can send a power sleep request to the driver for the source computing device; when it is necessary to restart the computing tasks in batches later, the control terminal can send a power wake-up request to the driver for the target computing device.

[0138] Figure 6 This is a block diagram of an electronic device used to implement embodiments of this application. For example... Figure 6 As shown, the electronic device includes a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the computer program, it implements the method described in the above embodiments. The number of memories 601 and processors 602 can be one or more. In a specific implementation, the electronic device may also include a communication interface 603 for communicating with external devices and exchanging data.

[0139] In practical implementation, if the memory 601, processor 602, and communication interface 603 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0140] Optionally, in a specific implementation, if the memory 601, processor 602 and communication interface 603 are integrated on a single chip, the memory 601, processor 602 and communication interface 603 can communicate with each other through an internal interface.

[0141] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in this application.

[0142] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method provided in this application.

[0143] This application also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device with the chip installed to perform the method provided in this application.

[0144] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.

[0145] It should be understood that the aforementioned processor can be a CPU, or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.

[0146] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0147] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0148] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0149] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0150] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.

[0151] The logic and / or steps described in the flowchart or otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0152] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.

[0153] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0154] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A driving method of a computing device, applied to an operating system kernel state of a host system, the computing device comprising a target computing device and a source computing device in the host system, the target computing device being used to replace the source computing device, the method comprising: in response to a power wake-up request for the target computing device, sending a power wake-up command to the target computing device, the power wake-up command being used to wake up the target computing device; in response to the target computing device having been woken up, sending a restore GPU memory command to the target computing device, the restore GPU memory command being used to instruct the target computing device to read target GPU memory data from a storage device, the target GPU memory data being obtained by writing GPU memory content generated by the source computing device in the process of executing a target computing task queue to the storage device before the target computing task queue is deleted; in a case where the target computing device has read the target GPU memory data, reconstructing the target computing task queue for the target computing device to continue executing an unfinished task in the target computing task queue based on the target GPU memory data.

2. The method of claim 1, wherein, The reconstructing the target computing task queue comprises: playing back a task command to the target computing device to reconstruct the target computing task queue, the task command being a command sent to the source computing device for creating the target computing task queue and recorded.

3. The method of claim 1 or 2, wherein, Before the sending the restore GPU memory command to the target computing device, the method further comprises: playing back an initialization command to the target computing device, the initialization command being a command sent to the source computing device for initializing the source computing device and recorded; in a case where the target computing device completes the initialization, sending the restore GPU memory command to the target computing device. 4.A driving method of a computing device, applied to an operating system kernel state of a host system, the computing device comprising a target computing device and a source computing device in the host system, the target computing device being used to replace the source computing device, the method comprising: in response to a power sleep request for the source computing device, sending a save GPU memory command to the source computing device, the save GPU memory command being used to instruct writing GPU memory content generated by the source computing device in the process of executing a target computing task queue to a storage device; in a case where the GPU memory content has been written to the storage device, deleting the target computing task queue; in response to the target computing task queue having been deleted, sending a power sleep command to the source computing device, the power sleep command being used to instruct the source computing device to enter a sleep state.

5. The method of claim 4, wherein, Before the sending the save GPU memory command to the source computing device, the method further comprises: in response to a task queue creation request from a target application, sending a task command to the source computing device and recording the task command, the task command being used to create the target computing task queue.

6. The method of claim 5, wherein, Before the sending the task command to the source computing device, the method further comprises: sending an initialization command to the source computing device, the initialization command being used to initialize the source computing device.

7. A driving apparatus of a computing device, applied in an operating system kernel state of a host system, the computing device comprising a target computing device and a source computing device in the host system, the target computing device being used to replace the source computing device, the apparatus comprising a power management module, the power management module being configured to: in response to a power sleep request for the source computing device, sending a save display memory command to the source computing device, the save display memory command being used to instruct the source computing device to write display memory content generated in a process of executing a target computing task queue into a storage device to obtain target display memory data; deleting the target computing task queue in a case that the display memory content has been written into the storage device; in response to the target computing task queue having been deleted, sending a power sleep command to the source computing device, the power sleep command being used to instruct the source computing device to enter a sleep state; in response to a power wake-up request for the target computing device, sending a restore display memory command to the target computing device, the restore display memory command being used to instruct the target computing device to read the target display memory data from the storage device; reconstructing the target computing task queue for the target computing device to continue executing an unfinished task in the target computing task queue based on the target display memory data in a case that the target computing device has read the target display memory data.

8. The apparatus of claim 7, wherein, the apparatus further comprising a state management module, the power management module being further configured to, in response to a task queue creation request from a target application, send a task command to the state management module, the task command being used to create the target computing task queue; the state management module being configured to record the task command and forward the task command to the source computing device; and / or, playing back the task command to the target computing device to reconstruct the target computing task queue.

9. The apparatus of claim 8, wherein, the power management module being further configured to send an initialization command to the state management module, the initialization command being used to initialize the source computing device; the state management module being further configured to record the initialization command and forward the initialization command to the source computing device; and / or, the power management module being further configured to send a power wake-up command to the state management module, the power wake-up command being used to wake up the target computing device; the state management module being further configured to forward the power wake-up command to the target computing device and, in response to the target computing device having been woken up, play back the initialization command to the target computing device; the power management module being further configured to, in a case that the target computing device completes initialization, send the restore display memory command to the target computing device.

10. A method for driving a computing device, applied in the operating system kernel mode of a host system, wherein the computing device includes a target computing device and a source computing device in the host system, the target computing device being used to replace the source computing device, comprising: Send a save video memory command to the source computing device. The save video memory command is used to instruct the video memory content generated by the source computing device during the execution of the target computing task queue to be written to the storage device to obtain the target video memory data. If the video memory content has been written to the storage device, delete the target computing task queue. And / or, A restore video memory command is sent to the target computing device. The restore video memory command is used to instruct the target computing device to read the target video memory data from the storage device. If the target computing device has read the target video memory data, the target computing task queue is rebuilt so that the target computing device can continue to execute the unfinished tasks in the target computing task queue based on the target video memory data.

11. A method for driving a computing device, applied at a control terminal, the method comprising: In response to a source computing device executing a target computing task queue being in an abnormal state, a power sleep request is sent to the driving device for the source computing device, the driving device being configured to put the source computing device into a sleep state based on the method of any one of claims 4 to 6; In response to the source computing device entering a sleep state, a power wake-up request for the target computing device is sent to the driving device. The driving device is further configured to wake up the target computing device based on the method of any one of claims 1 to 3, so that the target computing device can continue to execute unfinished tasks in the target computing task queue.

12. A method for driving a computing device, applied at a control terminal, the method comprising: In response to the need to reclaim the queue of target computing tasks, a power sleep request is sent to the driving device for the source computing device, the driving device being configured to put the source computing device into a sleep state based on the method of any one of claims 4 to 6.

13. A method for driving a computing device, applied at a control terminal, the method comprising: In response to the need to pause the target computing task queue, a power sleep request is sent to the driving device for the source computing device executing the target computing task queue, the driving device being configured to put the source computing device into a sleep state based on the method of any one of claims 4 to 6; and / or, In response to the need to rebuild the target computing task queue, a power wake-up request for the target computing device is sent to the driving device. The driving device is also configured to wake up the target computing device according to the method of any one of claims 1 to 3, so that the target computing device can continue to execute the unfinished tasks in the target computing task queue.

14. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 6, 10 to 13.

15. A computer-readable storage medium having stored therein a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 6, 10 to 13.

16. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6, 10 to 13.

Citation Information

Patent Citations

  • GPU resource management method, device and system and readable storage medium

    CN114816741A

  • GPU application migration method, device and system and storage medium

    CN116149818A