Task playback method, apparatus, device, medium, and product
Patent Information
- Application Number
- CN202611364463.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-04
- Publication Date
- 2026-10-02
AI Technical Summary
[0004]然而上述方法在动态的用户队列场景下,对异常事件的处理方式难以应对复杂的任务调度,任务回放效率较低
[0011]本申请实施例提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN122861901A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a task playback method, apparatus, device, medium, and product. Background Technology
[0002] Task replay refers to the process of resubmitting previously recorded instruction replay execution information to the processor to restore or reproduce the task execution process, thereby achieving processor behavior verification.
[0003] In related technologies, during task replay, for a single or fixed number of logical queues, the resources corresponding to the task are released after the replay is completed; in the event of abnormal events such as queue blocking or task switching during the replay, the replay is re-executed or an execution failure message is returned.
[0004] However, in dynamic user queue scenarios, the above methods are inadequate for handling complex task scheduling and have low task replay efficiency. Summary of the Invention
[0005] This application provides a task playback method, apparatus, device, medium, and product. The technical solution includes at least one of the following aspects.
[0006] On one hand, embodiments of this application provide a task replay method, the method comprising: If the playback of the first task data is interrupted when replaying the first task data through at least one hardware queue, a context snapshot of the first task data is obtained based on the target queue and the context snapshot is stored. The context snapshot is used to indicate the playback progress of the first task data. The target queue is the queue in the hardware queue that caused the playback interruption when replaying the first task data. In response to the first task data being triggered to resume playback, load a context snapshot of the first task data; The first task data is replayed based on the context snapshot.
[0007] On the other hand, a task playback device is provided, the device comprising: The storage module is configured to, when replaying first task data through at least one hardware queue, if the replay of the first task data is interrupted, obtain a context snapshot of the first task data based on a target queue and store the context snapshot, wherein the context snapshot is used to indicate the replay progress of the first task data, and the target queue is the queue in the hardware queue that caused the replay interruption when replaying the first task data. The recovery module is configured to load a context snapshot of the first task data in response to the first task data being triggered for recovery playback; The recovery module is also configured to recover and replay the first task data based on the context snapshot.
[0008] On the other hand, embodiments of this application provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, which is loaded and executed by the processor to implement the task playback method provided in the embodiments of this application above.
[0009] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the task playback method provided in the embodiments of this application as described above.
[0010] On the other hand, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the task playback method provided in the embodiments of this application described above.
[0011] The beneficial effects of the technical solutions provided in this application include at least the following: In a dynamic user queue scenario, at least one hardware queue is used to replay the first task data. When the target queue causing the replay interruption is determined, a context snapshot of the first task data is obtained and saved based on the target queue. This context snapshot indicates the execution state when the first task data replay was interrupted. Based on the context snapshot, the execution state can be restored during replay recovery. That is, in response to the first task data being triggered to resume replay, the context progress is reloaded, restoring the replay progress at the time of the interruption, thereby enabling the switching of the first task data replay. The handling of abnormal events during replay is implemented by saving the aforementioned context snapshot. In the event of an abnormality, the continuity of the logical queue is maintained by switching the replay task, addressing different task scheduling scenarios, ensuring the feasibility and stability of task replay switching, and improving task replay efficiency. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of the structure of a task playback system provided in an exemplary embodiment of this application; Figure 2 This is a flowchart of a task replay method provided in an exemplary embodiment of this application; Figure 3 This is a flowchart of a task replay method provided in another exemplary embodiment of this application; Figure 4 This is a schematic diagram of a task replay process provided in an exemplary embodiment of this application; Figure 5 This is a flowchart of a task replay method provided in yet another exemplary embodiment of this application; Figure 6 This is a schematic diagram of a first logical queue indicating replay task data provided in an exemplary embodiment of this application; Figure 7 This is a structural block diagram of a task playback device provided in an exemplary embodiment of this application; Figure 8 This is a structural block diagram of a task playback device provided in another exemplary embodiment of this application; Figure 9 This is a schematic block diagram of a processor provided in an exemplary embodiment of this application; Figure 10 This is a schematic block diagram of a board provided in an exemplary embodiment of this application; Figure 11 This is a schematic block diagram of a computer device provided in an exemplary embodiment of this application; Figure 12 This is a schematic block diagram of a computer device provided in another exemplary embodiment of this application; Figure 13 This is a schematic block diagram of a computer device provided in yet another exemplary embodiment of this application; Figure 14 This is a schematic block diagram of a computer device provided in another exemplary embodiment of this application; Figure 15 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation
[0014] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0015] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0017] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user, processor, and computer device data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0019] First, the terms used in the embodiments of this application will be introduced.
[0020] Task replay: Task replay refers to the process of resubmitting previously recorded task data to the processor as input to reproduce task execution. The task data includes instruction replay execution information, and the processor is the same execution environment or a simulated execution environment as the previous instructions, thereby verifying the processor's behavior. A logical queue is used to store the task data to be replayed, maintaining the execution order of instructions within the task data. During replay, the replay engine corresponding to the logical queue pulls task data from the logical queue and drives the hardware queue to replay the task data.
[0021] Task replay involves replaying different task data on different processors. For example, a Graphics Processing Unit (GPU) performs trace replay using a GPU execution sequence (trace) toolchain. This toolchain captures, records, and analyzes timing and processing status information during GPU operation, such as kernel startup, memory operations, and synchronization events. This provides a basis for GPU performance testing and functional debugging, such as reproducing faults and comparing whether anomalies were introduced before and after version updates. The task data is implemented as an execution sequence, including but not limited to image processing commands, parameter configuration commands, and computational commands received from the Graphics Application Programming Interface (API).
[0022] User Queue Scenario: In a user queue scenario, users corresponding to the processor who require replay submit task data to a logical queue via user submission requests. A correspondence exists between the logical queue and the user. Upon receiving the submitted task data, the logical queue triggers the replay of the task data, replaying it according to the order in which the data was enqueued. The same user corresponds to the same logical queue. The user submission request writes the task data as payload to the tail of the logical queue, without directly triggering task replay. The logical queue triggers the replay of task data according to a preset triggering strategy, which instructs the replay scheduling of task data in the logical queue.
[0023] Secondly, in related technologies, during task replay, for a single or fixed number of queue instances, the logical queue releases the resources corresponding to the task after replay is complete. In the event of abnormal events such as queue blocking or task switching during replay, replay is either attempted again or an execution failure message is returned. For example, detecting a buffer stop or an execution queue idle exception indicates that the current replay of task data has failed. However, in dynamic user queue scenarios, multiple rounds of user submission requests target the same logical queue. During the replay of task data corresponding to multiple rounds of user submission requests, the semantics of this logical queue are shared; that is, the global attributes maintained by the logical queue during the replay process remain continuous and are not reset across multiple rounds of task data.
[0024] The methods described above provide a one-time replay failure handling for abnormal events, which is insufficient for complex task scheduling. When replaying task data corresponding to user-submitted requests in subsequent rounds, the original execution context cannot be resumed, resulting in low task replay efficiency. Furthermore, these methods lack a mechanism to determine whether the hardware queue used for replay can be used for recovery, and they also lack a management mechanism for context snapshots used to recover replay progress. After an abnormal event occurs, only at least one of the following handling methods can be applied: replay failure or loss of the original execution context. This results in low utilization of the logical queue's continuity, further contributing to low task replay efficiency.
[0025] In this embodiment, when replaying first task data through at least one hardware queue, if the replay of the first task data is interrupted, a context snapshot of the first task data is obtained based on the target queue and stored. The context snapshot is used to indicate the replay progress of the first task data. The target queue is the hardware queue that caused the replay interruption when replaying the first task data was executed. In response to the first task data being triggered to resume replay, the context snapshot of the first task data is loaded. The first task data is replayed based on the context snapshot, thereby realizing the replay task switching to deal with replay interruption and abnormal events, maintaining the continuity of the logical queue, ensuring the feasibility and stability of task replay switching, and improving task replay efficiency. The replay task switching includes a saving phase and a recovery phase.
[0026] The task playback method provided in this application can be executed by a terminal, by a server, or by a combination of both. Taking the method executed by a combination of a terminal and a server as an example... Figure 1 This is a schematic diagram of the structure of a task playback system provided in an exemplary embodiment of this application. For example... Figure 1 As shown, the task playback system may include: terminal 120 and server 140.
[0027] Terminal 120 can be an electronic device such as a mobile phone, tablet computer, vehicle terminal (vehicle system), or personal computer (PC). Terminal 120 may deploy a processor for playing back task data. Optionally, the processor includes a device for executing task playback. Exemplarily, the processor may be at least one of a central processing unit (CPU), GPU, general-purpose graphics processor (GPGPU), application-specific integrated circuit (ASIC), neural network processing unit (NPU), tensor processing unit (TPU), vision processing unit (VPU), digital signal processing (DSP), or field-programmable gate array (FPGA). Terminal 120 may also deploy a playback engine for instructing the processor to play back the task data.
[0028] The task playback method provided in this application embodiment can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. Figure 1 Taking the task replay system shown as an example, the task replay method can be executed by terminal 120 (such as by the client of the performance prediction program installed and running in terminal 120), or by server 140, or by interaction between terminal 120 and server 140. This application does not limit this.
[0029] Terminal 120 is connected to server 140 via a wireless network or a wired network.
[0030] Those skilled in the art will understand that the number of the aforementioned devices can be more or less. For example, there may be only one device, or there may be dozens or hundreds of devices, or even more. This application does not limit the number or type of devices.
[0031] Server 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 140 provides functional services for implementing model training. Optionally, server 140 undertakes the main computing work, and terminal 120 undertakes the secondary computing work; or, server 140 undertakes the secondary computing work, and terminal 120 undertakes the main computing work; or, server 140 and terminal 120 collaborate on computing using a distributed computing architecture.
[0032] In one example, taking the interaction between terminal 120 and server 140 to perform a task playback method as an example, server 140, in response to receiving a first user submission request for a first logical queue, sends first task data to terminal 120. When terminal 120 is playing back the first task data through at least one hardware queue, if the playback of the first task data is interrupted, it obtains a context snapshot of the first task data based on the target queue and stores the context snapshot. The context snapshot is used to indicate the playback progress of the first task data. The target queue is the queue in the hardware queue of terminal 120's processor that caused the playback interruption during the playback of the first task data. In response to the first task data being triggered to resume playback, the context snapshot of the first task data is loaded; the first task data is resumed for playback based on the context snapshot. When terminal 120 completes the playback of the first task data, the playback result is obtained and transmitted back to server 140.
[0033] It is worth noting that the aforementioned terminal 120 refers to an electronic device with output display capabilities and input control capabilities. The aforementioned server 140 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud security, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. This application embodiment does not limit this.
[0034] Cloud technology refers to a managed technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.
[0035] In some embodiments, the server 140 described above can also be implemented as a node in a blockchain system.
[0036] Figure 2 This is a flowchart of a task playback method provided in an exemplary embodiment of this application, the method being implemented by a computer device (which can be configured as follows). Figure 1The method can be executed by the terminal 120 or server 140 shown, or it can be executed by the terminal, the server, or both. This embodiment takes the method being executed by the terminal as an example, and the method includes the following steps 210 to 230.
[0037] Step 210: If the playback of the first task data is interrupted when the first task data is replayed through at least one hardware queue, then obtain a context snapshot of the first task data based on the target queue and store the context snapshot.
[0038] Task replay refers to reproducing the execution of task data corresponding to a business task in a processor. By reproducing the historical operation sequence, the correctness, performance, and stability of the business system corresponding to the business task under specific inputs can be verified, thereby discovering regression defects, locating root causes of failures, or assessing the impact of changes. The replay system for the processor-corresponding execution replay task is the same execution environment or a simulation execution environment as the business system. For example, task data includes instruction replay execution information, such as historical request call chains.
[0039] In some embodiments, the application scenarios corresponding to replay tasks include, but are not limited to: 1. Software regression testing scenarios, such as replaying task data captured by the old version after a code version update to ensure that the new version does not break the existing and unchanged functions of the old version. 2. Performance benchmarking scenarios, such as replaying the same load to compare throughput, latency, and resource consumption under different hardware and / or software configurations. 3. Fault reproduction and diagnosis scenarios, such as replaying task data at the time of the anomaly when a rare anomaly occurs in a business scenario, reproducing the problem in an isolated environment and analyzing the root cause. 4. Driver and / or firmware verification: after the processor driver or firmware is upgraded, replaying the historical command stream to verify compatibility and execution correctness. 5. Artificial Intelligence (AI) model inference verification scenarios, such as replaying historical inference requests to ensure that the output results after the model update meet expectations. 6. Image processing verification scenarios, such as replaying task data corresponding to historical API call sequences, re-executing image processing after hardware or driver changes, and verifying the consistency of rendering results or performance changes. Optionally, the processor performing the replay includes, but is not limited to: a GPU performing trace replay, a CPU performing multi-threaded program replay, and a distributed system performing historical message consumption replay.
[0040] The playback system includes, but is not limited to, different logical queues for managing playback data sources and playback progress, a playback engine that receives scheduling from the logical queues and controls the hardware queues during playback, and a hardware queue corresponding to the processor as the physical execution unit. These three components are decoupled and work collaboratively through playback offsets, driver interfaces, and status feedback. The logical queue is the core of the playback system, storing task data to be played back and playback offsets. It triggers playback by pulling task data from the playback engine through polling or event-triggered strategies. The playback engine reads task data from the logical queue, generates a command stream based on the playback control signals corresponding to the logical queue, and submits it to the processor's hardware queue through the driver. The hardware queue receives the command stream corresponding to the playback control signals issued by the playback engine, performs calculations, and feeds back to the playback engine via interrupts. The playback engine updates the playback offset of the logical queue based on the feedback.
[0041] Task replay is automatically managed and triggered via logical queues, avoiding wasted human resources. Task data to be replayed is prioritized according to its delivery time, facilitating recording the replay process and generating replay results. The logical queue acts as an abstract buffer structure, maintaining the order of task data to be replayed, replay offsets, and metadata. Task data is arranged in the logical queue in the order it was enqueued. Illustratively, subsequent batches of task data are appended to the tail of the logical queue to avoid disrupting the order of previous batches. The continuously accumulating replay offset provides the replay system with a basis for resuming interrupted downloads. Task data from all submission rounds within the same logical queue share the same replay triggering strategy and context snapshot binding relationship.
[0042] Optionally, the logical queue schedules the processor's computing core to execute replay computing tasks through the hardware queue scheduling processor. The two are bound based on a mapping relationship, and after unbinding, the mapping relationship is maintained and associated by external metadata.
[0043] The system decouples the triggering of replay task data from the acquisition of task data through a preset triggering strategy. The logical queue indirectly triggers replay by ensuring that changes in the state of the task data within the queue meet the triggering conditions indicated by the triggering strategy. Triggering strategies include, but are not limited to: 1. A polling strategy, where the replay system periodically scans the logical queue for non-empty states. 2. An event-driven strategy, where the logical queue sends a replay control signal when it changes from empty to non-empty, waking up the replay engine to retrieve task data. It is important to note that the logical queue itself does not actively initiate replay; rather, it serves as the data source and trigger condition carrier for the replay operation.
[0044] Each logical queue typically corresponds to an independent replay task, which is determined by the user requesting replay and the business scenario corresponding to the task data to be replayed. Each logical queue has a unique identifier; for example, the identifier can be implemented as an identity (ID). Different logical queues carry different replay tasks, and each logical queue independently manages the order and replay offset of a set of task data corresponding to its replay task. The replay offset indicates the replay progress of the task data. A set of task data includes task data from at least one submission batch. Different logical queues may have the same or different replay priorities, triggering strategies, or resource allocations. For example, when a logical queue with a higher replay priority triggers replay, it can preempt the hardware resources of a lower-priority logical queue that is currently replaying. The replay process of the first logical queue, including but not limited to replay interruption, completion, or failure, does not affect the replay execution of other logical queues. In other words, logical queues are independent data pipelines divided according to replay tasks, and the differences between different logical queues are reflected in the content of the replay task and the data to be replayed.
[0045] For different batches of task data within the same logical queue that correspond to the same replay task, the data sources or generation methods of the different batches of task data may differ. For example, subsequent non-first task data may be incremental data newly generated during the execution of previous tasks, meaning there is a data generation order dependency between the two batches of task data. Alternatively, due to the capacity limitations of the first logical queue, users may submit task data to be replayed in batches. Optionally, the hardware queues corresponding to different batches of task data within the same logical queue for replay can be the same.
[0046] In a user queue scenario, the replay scheduling of task data in the logical queue includes three common scheduling logics: initial submission, subsequent submission, and final submission. The scheduling logic is implemented through user submission requests. Illustratively, when the logical queue receives a user submission request for the first time, the user submission request is considered an initial submission; when the logical queue receives a user submission request after the first time, the user submission request is considered a subsequent submission; and when the logical queue receives a request indicating that the task should be reset or the queue should be deregistered, the request is considered a final submission.
[0047] To illustrate, the scheduling logic described above is related to the task replay process. This includes, but is not limited to, the task data corresponding to the first submission being interrupted during replay, triggering context saving; the task data corresponding to subsequent submissions being interrupted, or triggering the restoration of the task data corresponding to the previous submission; and the termination of submissions being used to trigger data cleanup of the logical queue. For example, when a subsequent non-first submission to the same logical queue arrives, if a context snapshot corresponding to that logical queue is detected in the persistent storage space, the context snapshot is written back to the hardware queue through the original target queue or a new recovery queue to perform context loading and restore the replay.
[0048] In some embodiments, users of the replay system include, but are not limited to, testers, operations and maintenance personnel, and automated testing systems, thereby enabling at least one of the following: regression testing of the execution behavior of business tasks, building a continuous integration / continuous deployment (CI / CD) pipeline, and reproducing faults during the execution of business tasks. Users of the replay system are data providers who indicate replay requirements through task data, thereby verifying the processor execution behavior or performance of business tasks. They are not the users who generated the original business task requests. The users of the original business task requests can instruct the generation of task data, but there is no necessary connection between them and the task replay logic; they do not directly participate in task replay.
[0049] The context snapshot is used to indicate the playback progress of the first task data, and the target queue is the queue in the hardware queue that generates a playback interrupt when the first task data is played back.
[0050] In some embodiments, playback interruption includes at least one of the following situations: 1. A higher-priority playback task forcibly preempts the execution right of the current first logical queue through the driver or hardware layer. 2. The processor's computing core times out or insufficient storage resources cause the driver to terminate the queue. 3. The hardware queue corresponding to the logical queue experiences hardware anomalies such as link jitter or error-correcting code (ECC) errors, causing the hardware queue to be suspended.
[0051] If the execution progress of the hardware queue stalls abnormally during playback interruption, it can be obtained by checking the command completion status and timestamp timeout attribute of the hardware queue.
[0052] The target queue is the hardware queue in which the playback task corresponding to the first logical queue was executed before the first task data was interrupted and played back. It is also a hardware queue that can be reused during the playback recovery process. The target queue is the specific hardware queue selected by the playback system from the hardware queue for playing back the first task data during the context saving or context recovery process. As the execution anchor point of the current interrupted playback task in the first logical queue at the hardware layer, the target queue retains the execution context of the playback task corresponding to the first logical queue, such as register snapshots and cache information, to indicate the execution context when the playback terminal is obtained.
[0053] Among them, reusable is used to indicate that the target queue is in a healthy and available state and is not bound to other logical queues, including but not limited to at least one of the following: the driver response is normal, there is no freezing or timeout; the internal commands of the hardware queue have been cleared or can be safely reset; it is not occupied by other logical queues, or the correspondence between it and other logical queues has been released through abnormal reclamation.
[0054] In some embodiments, the playback system obtains context information from the registers, cache controls, and incomplete command list of the target queue through the driver interface of the target queue, obtains a structured context snapshot by serializing the context information, and stores the context snapshot corresponding to the first task data to a preset storage space, such as a memory cache, a distributed shared storage space, or a temporary file system.
[0055] Step 220: In response to the first task data being triggered to resume playback, load the context snapshot of the first task data.
[0056] In some embodiments, the methods for triggering the recovery of the first task data playback include, but are not limited to: 1. The user of the playback system actively issues a recovery command manually through API, command line, or control interface. The recovery command is used to instruct the playback system to recover the first task data playback. 2. The playback system automatically triggers the recovery of the first task data playback according to a preset strategy. Illustratively, when subsequent batches of task data arrive in the first logical queue, the playback system detects the existence of a context snapshot corresponding to the first logical queue, automatically loads and continues to play back the first task data; after the interrupted hardware resources are restored to availability, the playback system automatically resumes the suspended playback task; the background daemon of the playback system periodically checks the hardware queue status, and automatically starts the context recovery process for the first task data playback when the preset recovery conditions are met.
[0057] A context snapshot is a record of key state information captured from the hardware queue when the playback task corresponding to the first logical queue is paused due to an interruption, enabling the reconstruction of the execution context. The context preservation achieved through context snapshots only saves the playback progress of the first task data, not the first task data itself. Unplayed task data remains in the first logical queue. When the first task data is triggered to resume playback and the context snapshot is loaded, playback continues from the position indicated by the playback progress. Illustratively, in response to the first task data being triggered to resume playback, the playback system loads the context snapshot, continues reading the first logical queue from the recorded playback offset, and reconstructs the hardware state of the hardware queue, thereby achieving precise continuation of the first task data playback without depending on the state of other hardware queues.
[0058] In some embodiments, the playback progress of the playback task corresponding to the first logical queue includes, but is not limited to, the completed playback offset, such as the read position of the first task data in the first logical queue; hardware local state, such as submitted but incomplete commands, the dependencies between the first task data and the hardware queue, and the dependencies between the hardware queue and the corresponding computing or storage resources. There is a correspondence between the playback task, the identification information of the first logical queue, and the context snapshot corresponding to the first task data.
[0059] Step 230: Restore the first task data based on the context snapshot.
[0060] In some embodiments, the playback result is obtained after the playback of the first task data or the task data in the first logical queue has ended. Optionally, playback execution information during the playback process can be generated and output as a record to obtain the playback result.
[0061] For example, replay results include, but are not limited to, replay execution reports, such as performance metrics (time consumption, throughput, etc.), functional correctness comparison results (consistency with and / or difference from expected output, etc.); replay anomaly diagnostic information, such as outputting error codes, crash logs, or anomaly information in the event of failure during replay, for fault location.
[0062] In summary, the method of this embodiment determines the target queue for replaying the first task data in a dynamic user queue scenario, and obtains and saves a context snapshot of the first task data based on the target queue. The context snapshot is used to indicate the execution state when the replay of the first task data is interrupted. Based on the context snapshot, the execution state can be restored when the replay is resumed. That is, in response to the first task data being triggered to resume replay, the context progress is reloaded to restore the replay progress when the replay was interrupted, thereby realizing the switching of the replay of the first task data. The method for handling abnormal events during the replay process is to save the above-mentioned context snapshot and maintain the continuity of the logical queue in the event of an abnormality by switching the replay task. This addresses different task scheduling scenarios, ensures the feasibility and stability of task replay switching, and improves task replay efficiency.
[0063] Figure 3 This is a flowchart of a task playback method provided in another embodiment of this application. The method is implemented by a computer device (which can be configured as follows). Figure 1 The method can be executed by the terminal 120 or server 140 shown, or it can be executed by the terminal, the server, or both. This embodiment takes the method being executed by the terminal as an example, and the method further includes steps 310 to 320.
[0064] Step 310: During the process of replaying the first task data through at least one hardware queue, obtain the hardware status corresponding to at least one hardware queue.
[0065] The hardware queue used for replaying the first task data is the actual execution carrier. It receives and schedules the processor command stream issued by the playback engine, driving the processor hardware to complete the computational tasks during the playback process. Illustratively, taking GPU as the processor for trace playback as an example, the hardware queue replaying the first task data includes at least one of the following steps: the playback engine constructs a GPU command buffer based on the task data, the command buffer including the computational core parameters and resource binding information corresponding to the playback command to be executed; the playback command is sent to the hardware queue through the driver interface; the hardware queue calls the GPU computing unit to execute the playback command; the GPU computing unit reports the progress of the playback process to the playback engine and records the playback execution information.
[0066] Hardware status is indicator data that indicates the execution status of the hardware queue, including but not limited to at least one of the following: command buffer pointer, resource binding information, register snapshot, etc. The command buffer pointer is used to indicate the list of incomplete commands during the first task data replay process, and the resource binding information is used to indicate the dependency relationship between the first task data and the hardware queue, as well as between the hardware queue and the corresponding computing or storage resources.
[0067] To illustrate, taking GPU as the processor for trace playback as an example, the hardware status also includes: the hardware queue doorbell value, which is the sequence number of the currently submitted command buffer, used to synchronize the hardware scheduling progress when restoring the first task data; the semaphore used to synchronize the execution order of hardware queue commands, which is the count of currently waiting synchronization objects; the page table base address register, which is the root address for GPU memory address translation; the exception status register, which records the reason for playback interruption or failure, used to assist in determining whether the hardware is in a normal state when restoring the first task data; the current shader program handle, used to indicate the computing core or shader loaded when the first task data playback was interrupted; and the temporary register group, used to calculate temporary data or intermediate results in the playback task.
[0068] Step 320: Determine the recovery queue from at least one hardware queue whose hardware status meets the preset status conditions.
[0069] The status condition is used to indicate that the queue to be recovered has the ability to recover the data of the first task in the event of a playback interruption.
[0070] In some embodiments, during the replay of the first task data triggered by the logical queue, the hardware status of each hardware queue corresponding to the logical queue is continuously detected, and the existence of a recoverable hardware queue is identified based on the combination of hardware statuses.
[0071] It is worth noting that steps 310 to 320 can be implemented before step 220, i.e., before the playback of the first task data is interrupted, or they can be implemented in response to the playback interruption of the first task data. Hardware queue status detection does not necessarily require continuous polling and can also be implemented as interrupt-triggered detection; interrupt-triggered detection includes, but is not limited to, when the playback task is interrupted due to preemption, timeout, or fault, the detection logic is passively triggered by the driver or system event to identify the status of the interrupted hardware queue used for playing back the first task data and mark the hardware queue as a queue to be recovered; among them, interrupt-triggered detection reduces system overhead, while continuous polling provides more comprehensive coverage of abnormal situations, and this application does not limit this. Step 210 can include the following steps 212 to 214.
[0072] Step 212: Obtain the target queue based on the queue to be recovered.
[0073] The status condition is used to indicate that the hardware queue is ready to resume playback of the first task data.
[0074] To illustrate, the hardware queue is paused due to an interrupt, and the execution progress of the replay task for the first task data has not yet been persistently saved. The internal hardware of the hardware queue retains the execution context of the replay task corresponding to the first logical queue, such as register snapshots and cache information. The replay system detects through status checks that the hardware queue meets the status conditions and is in a recoverable state. In response to the interruption of the first task data replay, when the hardware queue corresponding to the first logical queue needs to relinquish its hardware, the replay system saves its execution context and generates a context snapshot based on the residual hardware state of the hardware queue corresponding to the first logical queue. For example, when the hardware queue is detected to be in a recoverable state, a "to be recovered" mark is added to the hardware queue, instructing the organization to directly reclaim the hardware queue as a regular free queue.
[0075] To illustrate, before the context saving process is started, the playback system selects the target queue from the hardware queues whose hardware states meet the state conditions, and obtains the context information from the registers, cache space and command buffer corresponding to the target queue through the driver interface of the target queue. The context information is serialized to obtain a context snapshot, and the context snapshot corresponding to the first task data is persistently stored.
[0076] In some embodiments, a recovery tag is marked on the recovery queues whose hardware states meet the state conditions. The recovery tag includes the queue number. In response to the playback interruption of the first task data, the recovery queues are sorted according to the queue number of the recovery tag to obtain a recovery queue sequence. The first recovery queue is determined as the target queue based on the recovery queue sequence.
[0077] Indicatively, when the logical queue is not bound to a target queue, the target queue is selected from the hardware queues marked with a pending recovery flag. Optionally, the determination of the target queue is implemented through a traversal query: the playback system sequentially queries the execution status of the hardware queues, including busy or idle states, and selects the first hardware queue in the idle state as the target queue according to the order of queue identifiers (e.g., queue IDs) (e.g., 0, 1, 2, 3). The pending recovery flag decouples the hardware status query of the hardware queue from the determination of the target queue, simplifies the selection strategy of the target queue through the queue identifier of the hardware queue, and improves the efficiency of task playback.
[0078] In some embodiments, determining a queue to be recovered from at least one hardware queue whose hardware state meets preset state conditions includes at least one of the following: determining a queue to be recovered from at least one hardware queue whose hardware state indicator queue buffer is in a stagnant state; determining a queue to be recovered from at least one hardware queue whose hardware state indicator queue depth is in an empty state.
[0079] The target queue includes at least one of the following: a first queue in the hardware queue that causes playback blocking when replaying the first task data; a second queue in the hardware queue that causes playback exception when replaying the first task data, wherein the driver interface of the second queue is in a normal response state.
[0080] To illustrate, playback blocking includes, but is not limited to, buffer stagnation, meaning that a command has been submitted to the hardware queue but has not received a completion interrupt or synchronization semaphore for a long time, and the driver returns a "busy" or "suspended" state when querying the hardware status; an empty task is an abnormal idle state of the hardware queue, which can be obtained by detecting the queue depth and command quantity attributes, that is, after the playback engine pulls task data from the logical queue, the command corresponding to the generated playback control signal does not contain a valid calculation command, such as at least one of the following: a zero-length buffer, an unbound resource, etc., resulting in no actual load to execute in the hardware queue.
[0081] By analyzing the combination of buffer stagnation and empty tasks in the hardware state, it is determined whether the hardware queue meets the state conditions; hardware queues whose hardware states meet the state conditions are identified as recovery queues. Conversely, if the hardware state of a hardware queue does not meet the state conditions, it indicates that the hardware queue is in an unrecoverable error state, such as queue suspension, timeout error, or the associated resources of the hardware queue being released or reallocated. Consequently, the driver or hardware state corresponding to the hardware queue has been corrupted and cannot be reset to a state where task data can be replayed. By clearly defining the state conditions of the queue to be recovered, and identifying the target queue for obtaining context snapshots among at least one hardware queue, task replay efficiency is improved.
[0082] Schematic representation: State conditions include, but are not limited to: the hardware queue being able to normally receive and respond to operation requests issued by the driver; the hardware queue being able to correctly return the current execution state of the queue after reset processing; the control registers and / or status registers corresponding to the hardware queue being able to normally respond to read and write operations of the playback engine without timeouts or bus errors; the video memory or host memory region corresponding to the hardware queue being able to be correctly mapped to the address space without memory management unit (MMU) failures; and the queue operation functions of the playback system driver layer, such as playback commands corresponding to task data and query status functions, being able to be called normally and return valid results without freezing or crashing. In other words, from the perspective of the playback system driver, the driver can communicate normally with the hardware queue and execute playback operations, and the hardware queue meets the state conditions.
[0083] The queue to be recovered must meet certain conditions to be readable, such as a normal driver interface response and accessible registers. It does not need to meet the condition of being able to execute new commands, thus allowing the acquisition of a context snapshot of the first task's data. In other words, the queue to be recovered is a recoverable hardware queue temporarily in an abnormal state. State recovery processing, such as resetting or driver reset, can restore the execution of playback operations.
[0084] In some embodiments, the target queue is the hardware queue that executes the replay task of the first logical queue. To avoid errors in selecting the target queue, interrupted recovery queues corresponding to other logical queues, such as hardware queues corresponding to previously executed tasks, tasks corresponding to other logical queues, or interrupted replay tasks, are not included in the at least one hardware queue mentioned in step 212. That is, there is an association between at least one hardware queue and the first logical queue. For example, the association between the queue identifier of at least one hardware queue and the identifier information of the first logical queue is maintained through a mapping table.
[0085] Schematic, in response to a playback task interruption, the binding relationship between the first logical queue and at least one hardware queue is unbound. The execution context of the playback task corresponding to the first logical queue, such as register snapshots and cache information, remains on the at least one hardware queue. The playback system records the association between the queue IDs of the hardware queues and the logical queue IDs in the playback system's metadata. That is, at least one hardware queue itself is in an unbound state, but the mapping relationship between hardware queues and logical queues is maintained externally. When a target queue is determined, the playback system queries this mapping relationship to select the target queue corresponding to the first logical queue from the queue IDs of the at least one hardware queue.
[0086] With the execution context preserved, the hardware queue for replaying the first task data that is compatible with the replay task corresponding to the first logical queue is selected as the target queue. The residual cache, page table, translation lookaside buffer (TLB), register mapping and other local states are used to reduce the overhead of subsequent replay recovery.
[0087] Step 214: Obtain a context snapshot based on the target queue and write the context snapshot to the preset persistent storage space.
[0088] In some embodiments, the context snapshot does not simply save a single state bit, but records a set of key context information that can reconstruct the execution context. Using the identifier of the first logical queue as an index key, it ensures that the context snapshot corresponds to the playback task of the first logical queue. This prevents interference when different logical queues maintain their own context snapshots corresponding to their respective task data, facilitating playback system management. Furthermore, when restoring the first task data for playback, the context snapshot can be directly located using the identifier of the first logical queue, eliminating the need to traverse the stored content, reducing query overhead, and improving task playback switching efficiency. The context snapshot includes at least one of the following.
[0089] 1. Identification information of the first logical queue; The identification information of the first logical queue corresponds to the replay task, and the identification information can be implemented as the logical queue ID.
[0090] 2. The submission batch identifier corresponding to the first task data is used to distinguish different task data units in the logical queue; the submission batch identifier corresponding to the first task data is used to indicate the context metadata of the replay task, including but not limited to the replay order, the timestamp of submission to the logical queue, etc.
[0091] Optionally, the generation of the submission batch identifier includes, but is not limited to, explicitly specifying the logical queue ID and the submission batch identifier when the user initiates a user submission request; or, through context inheritance, if a subsequent submission is identified by the replay system as a continuation of the replay of previous task data, the submission batch identifier of the task data is consistent with that of the previous task data; or, if two different submission batches of task data belong to the same workflow at the business level, such as continuous operation records of the same user session, their order, submission batch identifier, and progress are uniformly managed through a logical queue.
[0092] 3. The hardware state corresponding to the first task data, which includes at least one of the following: command buffer pointer, resource binding information, register snapshot, etc. The command buffer pointer is used to indicate the list of incomplete commands during the playback of the first task data, and the resource binding information is used to indicate the dependency relationship between the first task data and the hardware queue, as well as between the hardware queue and the corresponding computing or storage resources.
[0093] Optionally, for the hardware queue used to replay the first task data, it is not necessary to save all hardware states; illustratively, the hardware program counter does not need to be saved separately, but can be reconfigured when replaying the first task data.
[0094] 4. Playback offset of the first logical queue; The playback offset is used to indicate the playback progress of the first task data. The playback offset is used to indicate the offset or sequence number corresponding to the position in the logical queue where the first task data has been completed for playback, and corresponds to the read position in the logical queue when resuming the playback of the first task data.
[0095] In the persistent storage space, the context snapshot is stored using the identification information corresponding to the first logical queue, such as the logical queue ID, as the index key. Illustratively, there is a mapping relationship between the context snapshot corresponding to the task data of the playback interruption and the logical queue of the task data. The context snapshot and the logical queue can be stored in a separate persistent storage space in the form of a mapping table.
[0096] The persistent storage space corresponds to the playback system; for example, the persistent storage space includes, but is not limited to, at least one of a local file system, a database, or a distributed key-value store. Optionally, the persistent storage space used to store context snapshots, task data, and playback results may be the same or different.
[0097] By storing context snapshots and logical queues independently in persistent storage, the playback progress corresponding to the first task data that was interrupted during playback is maintained, enabling context switching for task playback and improving the stability of task playback switching.
[0098] In some embodiments, after step 230 above, in response to the playback end triggered by the first logical queue, a playback end signal is obtained, which is used to indicate at least one of playback execution completion and playback execution exception; data cleanup is performed on the recovery queue based on the playback end signal.
[0099] In some embodiments, the daemon process corresponding to the playback engine is used to periodically check the hardware queue that has ended playback and perform data cleanup on the hardware queue based on the playback end signal. Optionally, the playback end signal obtained in response to the playback end triggered by the first logical queue includes at least one of indicating that the playback of the interrupted first task data has ended and indicating that the playback of all task data in the first logical queue has ended.
[0100] Schematic, data cleanup of the recovery queue based on the playback end signal includes, but is not limited to: recording playback execution information during the playback process to indicate the playback result, wherein the playback result includes playback exception information; deleting the context snapshot and releasing the physical resources of at least one hardware queue, including the recovery queue, when playback execution is complete, such as resetting the hardware state, handling incomplete commands, unbinding memory or video memory addresses, etc.; marking the playback task as abnormal and terminating the current playback execution chain when playback execution is abnormal and causing playback failure, and releasing the physical resources of the hardware queue.
[0101] Optionally, when performing data cleanup on the hardware queue, data cleanup is also performed on the first logical queue, including but not limited to cleaning up metadata, such as resetting the replay offset of task data, removing the binding relationship between the first logical queue and the target queue and the recovery queue respectively, and updating the logical queue status, etc., at least one of the following.
[0102] Based on the playback end signal, a single recycling entry point is shared between the completed and failed states. The playback engine responds to the playback end signal by triggering a pre-defined cleanup function to perform data cleanup; or it retrieves data cleanup instructions based on the playback end signal. This is achieved by uniformly recycling the hardware queue state upon recovery failure, playback failure, or playback completion. In the case of playback recovery, exception recycling is uniformly performed for both successful and failed playback paths, improving the stability of playback control and thus increasing the efficiency of task playback.
[0103] In some embodiments, a replay execution exception is used to indicate that an unrecoverable exception has occurred in the hardware queue, such as the recovery queue, during the replay of the first task data. The replay execution exception includes at least one of the following: a failure in the recovery queue; a corrupted context snapshot of the first task data; the recovery queue being unresponsive; or an abnormality in the storage address corresponding to the first task data.
[0104] Indicatively, unrecoverable abnormal situations include, but are not limited to: 1. The time spent waiting for the semaphore used to synchronize the hardware queue during the recovery queue replay process reaches a preset waiting limit, such as a timeout caused by at least one reason, such as the previous replay task instruction not being completed due to hardware failure, hardware scheduler blocking, or loss of context snapshot; where waiting timeout refers to unrecoverable waiting failures due to hardware state failure. For example, taking the GPU as the processor, waiting failure includes GPU hardware unresponsiveness (GPU Hang), and the corresponding abnormal event scenarios include, but are not limited to, at least one of the following: shader infinite loop, MMU failure, power management abnormality, etc. 2. The processor attempts to access a memory address for which no mapping relationship has been established, such as an address access abnormality caused by at least one reason, such as command buffer pointer out of bounds or resource binding failure.
[0105] Optionally, if a playback end signal is obtained after the playback of the first logical queue is completed, the playback execution exception can be used to indicate that during the playback of the second task data and other task data corresponding to the first logical queue, there is an unrecoverable exception in the hardware queue used to play back the second task data. This application does not limit this.
[0106] During the continued execution after replay, if the target queue has completed replaying the first task data, the hardware state of the recovery queue is cleared and the context snapshot corresponding to the first task data is removed. If an unrecoverable exception occurs during the replay process, the first task data is marked as failed and exception recycling is triggered. The recycling process extends to different event scenarios corresponding to execution exceptions, with different event scenarios corresponding to different exception causes. Based on the replay end signal obtained under different exception event scenarios, the recovery queue is uniformly recycled. Specifically, by directly marking unrecoverable exceptions as failures and recycling them, unlike retryable normal wait timeouts, the task replay and recycling process is clearly defined, improving task replay efficiency.
[0107] For illustrative purposes, please refer to the following: Figure 4 , Figure 4 This is a schematic diagram of a task playback process provided in an exemplary embodiment of this application, such as... Figure 4As shown, when the playback of the first task data is triggered, step 410 is executed to obtain the hardware status corresponding to at least one hardware queue. Step 420 is executed to mark the queues to be recovered in the at least one hardware queue whose hardware status meets the preset status conditions as pending recovery. Step 430 is executed to determine the target queue from the queues to be recovered in response to the playback interruption of the first task data. Step 440 is executed to obtain and store the context snapshot of the first task data based on the target queue. Step 450 is executed to load the context snapshot of the first task data in response to the triggering of the recovery playback of the first task data, and to recover and play back the first task data based on the context snapshot. Step 460 is executed to perform data cleanup on the recovery queue in response to the end of the first task data playback.
[0108] In summary, the method of this embodiment determines the target queue for replaying the first task data in a dynamic user queue scenario, and obtains and saves a context snapshot of the first task data based on the target queue. The context snapshot is used to indicate the execution state when the replay of the first task data is interrupted. Based on the context snapshot, the execution state can be restored when the replay is resumed. That is, in response to the first task data being triggered to resume replay, the context progress is reloaded to restore the replay progress when the replay was interrupted, thereby realizing the switching of the replay of the first task data. The method for handling abnormal events during the replay process is to save the above-mentioned context snapshot and maintain the continuity of the logical queue in the event of an abnormality by switching the replay task. This addresses different task scheduling scenarios, ensures the feasibility and stability of task replay switching, and improves task replay efficiency.
[0109] The method provided in this embodiment detects the hardware status of at least one hardware queue in a user queue scenario, distinguishes between recoverable anomalies and ordinary unrecoverable playback failures of the playback task of the first logical queue from the perspective of hardware execution, and obtains a context snapshot of the first task data playback from the recoverable target queue. The context snapshot is used to realize the context switching of task playback, including determining the playback progress of the first task data when playback is interrupted based on the context snapshot when the playback task of the first logical queue continues to be executed, thereby reducing unnecessary overall interruption of the playback process.
[0110] Figure 5 This is a flowchart of a task playback method provided in another embodiment of this application. The method is implemented by a computer device (which can be configured as follows). Figure 1 The method can be executed by the terminal 120 or server 140 shown, or it can be executed by the terminal, the server, or both. This embodiment uses the method executed by the terminal as an example, and the method further includes the following steps.
[0111] Step 510: In response to the first task data being triggered to resume playback, determine the recovery queue from the queue to be recovered.
[0112] In some embodiments, before step 210 above, at least one of the following steps is included: receiving a first user submission request for a first logical queue corresponding to the first task data, the first user submission request being used to indicate that the first task data is provided for the task playback of the first logical queue, the first logical queue being used to manage the first task data and trigger the playback of the first task data; obtaining the first task data based on the first user submission request; and, in the case of obtaining the first task data, triggering the playback of the first task data through the first logical queue in response to the first logical queue being non-empty.
[0113] The first user submission request includes the identification information of the first logical queue and the data of the first task or the storage address of the first task data. The first task data is obtained through the first user submission request, and the first task data corresponding to the first logical queue is triggered to play back under the preset triggering condition, that is, when the first logical queue is not empty. The data preparation and playback execution of the playback task are decoupled through the user submission request, supporting asynchronous scheduling and realizing flexible playback task orchestration.
[0114] The user-submitted request is used to store the task data to be replayed into the logical queue corresponding to the replay task. The logical relationship between the user-submitted request and the task data includes, but is not limited to, the user-submitted request providing the original input to the replay system by delivering the task data; when the task data is stored in the logical queue for testing, it changes the logical queue from empty to non-empty, indirectly activating the triggering strategy, thereby triggering the logical queue to send replay control instructions to the hardware queue, resulting in a change in the state of the hardware queue.
[0115] Optionally, the user submission request is implemented as a write operation. Users in the replay system store the replay task load, i.e., task data, into a logical queue through the user submission request, where a correspondence exists between the user and the logical queue. The user submission request is implemented as a submission action, and the user submitting the corresponding task data is implemented as the submission content. The purpose of the user submission request is to prepare the basis for replay data, not to directly start replay. The initiation of replay for the first task data is triggered by the first logical queue based on the first task data in the queue.
[0116] In some embodiments, the execution timing relationship between the replay task and the business task corresponding to the task data source includes: 1. The replay task and the business task are in real-time or near real-time relationship, that is, the replay system acquires the task data corresponding to the business task and replays the task data at the same time as the business task is executed or within a time period below a threshold. The replay task and the business task run in parallel, and there may be a short delay between their execution times. Accordingly, the task replay is implemented as online replay, which can be used for at least one of real-time detection, dynamic verification, etc. 2. The replay task and the business task are in asynchronous or delayed processing relationship, that is, after the business task is completed, the task data is stored in the storage space corresponding to the historical task data. The user in the replay system sends the task data from the above storage space to the logical queue and then triggers the replay of the task data. The replay task and the business task run serially. Accordingly, the task replay is implemented as offline replay, which can be used for at least one of regression testing, performance analysis, or fault reproduction.
[0117] Optionally, in the different playback scenarios mentioned above, users of the playback system can pre-set playback strategies. Playback strategies are used to indicate the objects, timing, etc., for online or offline playback. Alternatively, users can submit user submission requests in real time to realize data delivery to the logical queue in offline playback.
[0118] In some embodiments, before the playback of the first task data is interrupted, at least one of the following steps is included: during the playback of the first task data, acquiring the third task data corresponding to the second logical queue, wherein the playback tasks corresponding to the first logical queue and the second logical queue are different; in response to the second logical queue triggering the playback of the third task data, acquiring the playback interruption signal for the first logical queue, wherein the playback interruption signal is used to indicate the interruption of the playback of the first task data.
[0119] The first and second logical queues correspond to different replay tasks, and the users corresponding to the first and second logical queues are also different. Users in the logical queues submit task data to the logical queues via user submission requests. If the replay task corresponding to the second logical queue interrupts the replay task corresponding to the first logical queue, the replay priority of the second logical queue can be indicated to be higher than that of the first logical queue; this application does not limit this. By clarifying the rules for obtaining the replay interruption signal, scheduling management of multi-round submissions in user queue scenarios is achieved. The switching rules during task replay are clarified. Upon obtaining the third task data, an interruption is triggered to replay the first task data in the first logical queue. In dynamic user queue scenarios, switchover processing is performed on recoverable abnormal events, improving task recovery efficiency.
[0120] The recovery queue is used to recover and replay the first task data. In some embodiments, in response to the first task data being triggered to recover and replay, determining the recovery queue from at least one hardware queue includes at least one of the following: 1. If the hardware state of the target queue meets preset recovery conditions, the target queue is determined as the recovery queue, and the recovery conditions include at least one of the following: the driver interface of the target queue is in a normal response state, the hardware of the target queue is in a normal operating state, and the resources of the target queue are valid; 2. If the hardware state of the target queue does not meet the recovery conditions, a recovery queue whose hardware state meets the recovery conditions is determined from the queue to be recovered.
[0121] The recovery conditions are used to indicate the conditions that the hardware state of the hardware queue used to recover the data of the first task must meet, including but not limited to at least one condition such as the hardware queue being in an idle state, there being no residual replay commands to be executed, and the storage page table mapping being normal.
[0122] In some embodiments, the recovery condition indicates the conditions that the hardware state of the hardware queue must meet during context recovery, and the state condition indicates the conditions that the hardware state of the hardware queue must meet during context saving. The state condition only requires the hardware queue to be readable, while the recovery condition requires the hardware queue to be writable and executable. Optionally, since the hardware state of the target queue meets the state condition, in response to the triggering of recovery playback of the first task data, the target queue is reprocessed to make it idle and writable, that is, the target queue is determined as the recovery queue when the hardware state of the target queue meets the recovery condition.
[0123] In response to the triggering of the first task data recovery playback, the system first checks whether the hardware status of the target queue meets the recovery conditions. If the target queue's hardware status meets the recovery conditions, it is prioritized for recovery playback of the first task data. If the target queue's hardware status does not meet the recovery conditions, other hardware queues used to execute the playback task corresponding to the first logical queue are checked, and the recovery queue is determined. Illustratively, the queue identifiers corresponding to the queues to be recovered are sorted according to queue number size to obtain a sequence of queues to be recovered, and the recovery queue is determined based on this sequence.
[0124] By prioritizing the selection of a target queue to restore the replay of the first task data, the migration overhead of the context snapshot is reduced by leveraging the local state affinity of the target. The local state affinity is used to indicate that the local state of the hardware queue after executing the replay task corresponding to the first task data has a high degree of matching with the address space and resource binding of the current replay task. This makes it easier to continue replaying the first task data on the hardware queue than to migrate the context snapshot to another hardware queue and restore the replay.
[0125] To illustrate, taking the GPU as the processor for trace playback as an example, determining the target queue as the recovery queue eliminates the need to re-establish page table mapping or reload shader programs. The GPU cache may still contain intermediate data from the first task, avoiding re-initializing the queue state or reallocating video memory resources, thus reducing scheduling costs.
[0126] In some embodiments, after determining the recovery queue, a binding relationship is established between the first logical queue and the recovery queue, and the context data corresponding to the first task data is loaded through the recovery queue and the first task data is restored and replayed.
[0127] Step 520: Load the context snapshot via the recovery queue.
[0128] In some embodiments, before step 220 above, at least one of the following steps is included: when the second task data corresponding to the first logical queue is obtained, the recovery playback of the first task data is triggered through the first logical queue; in response to the end of the playback triggered by the first logical queue, the playback of the second task data is triggered through the first logical queue.
[0129] Wherein, the second task data may be the same as or different from the first task data; if the second task data is the same as the first task data, the second user submits a request to instruct the restoration of the first task data in the first logical queue; if the first task data is different from the second task data, the first logical queue triggers the first task data based on the enqueue order of the task data when the second task data is obtained; wherein, the second user submits a request to provide the second task data to the first logical queue.
[0130] The logical queues corresponding to the first task data and the second task data are the same, but the submission batches are different. To illustrate, the first task data that was interrupted during playback is a batch that has been queued but not yet fully played back, while the second task data is a batch that was newly submitted after the first task data was interrupted during playback. Both are appended to the end of the same logical queue. When resuming playback of the first task data, the playback engine continues to read the first logical queue from the playback offset recorded in the context snapshot. It first completes the playback of the remaining part of the first task data of the interrupted batch, and then plays back the second task data of the subsequent new batch. The first task data and the second task data are arranged consecutively in the first logical queue in the order of their enqueueing.
[0131] Schematic illustration: When the second task data is the same as the first task data, the submission batch identifiers of the first and second task data are the same; in response to the task data obtained in the preceding and following rounds having the same submission batch identifier and this submission batch identifier having a context snapshot in the persistent storage space. When the second task data is different from the first task data, in response to the task data obtained in the preceding and following rounds corresponding to the same first logical queue and the identifier information of the first logical queue having a context snapshot in the persistent storage space, the first logical queue triggers the replay and recovery of the first task data based on the context snapshot.
[0132] Optionally, the determination of whether the task data or its submission batch identifier obtained in previous and subsequent rounds is the same during the maintenance of the logical queue can be implemented in different ways. For example, the determination can be made in response to the acquisition of task data or in response to the acquisition of task data after triggering replay. This application does not limit this. Illustratively, the determination in response to the acquisition of task data after triggering replay is implemented as follows: In response to the acquisition of new task data, check whether there is a context snapshot corresponding to the current logical queue in the persistent storage space. If a context snapshot exists, the current logical queue triggers the restoration and replay of the task data corresponding to the context snapshot. If the replay of the task data corresponding to the context snapshot ends and the replay of new task data continues, determine whether the submission batch identifier of the new task data has been replayed in the current logical queue. If the new task data is the same as the task data of the previous round, skip the repeated replay of the new task data. If the new task data is different from the task data of the previous round, replay the new task data.
[0133] By managing logical queues, scheduling and management of multi-round submissions in user queue scenarios are achieved. The switching rules during task replay are clarified. When the second task data is obtained, the first task data that was previously saved in the context in the first logical queue is triggered to be restored and replayed first, and then the second task data is triggered to be replayed, maintaining the consistency of logical queue management. In dynamic user queue scenarios, recoverable abnormal events are switched and unrecoverable abnormal events are recycled, improving task recycling efficiency.
[0134] For illustrative purposes, please refer to the following: Figure 6 , Figure 6 This is a schematic diagram of a first logical queue indicating playback task data provided in an exemplary embodiment of this application, such as... Figure 6 As shown, the first logical queue 610 manages task data submitted in different batches, such as first task data and second task data. According to the timing of user submission requests, the second task data is enqueued after the first task data.
[0135] like Figure 6 As shown, the first logical queue 610 receives a user submission request 620 and obtains the task data corresponding to the user submission request 620, including but not limited to obtaining first task data based on the first user submission request and obtaining second task data based on the second user submission request; wherein the user submission request 620 is used to submit task data to be replayed. Upon obtaining the task data, the first logical queue 610 triggers the replay of the task data and sends a replay control command to at least one hardware queue, wherein the at least one hardware queue includes a target queue 630.
[0136] like Figure 6 As shown, in response to the interruption of the first task data playback, when the target queue 630 corresponding to the first logical queue 610 is determined, a context snapshot 640 is obtained based on the target queue 630 to save the context of the first task data playback; in response to the triggering of the playback recovery of the first task data (for example, triggering playback recovery based on the enqueueing of the second task data), the context snapshot 640 is read, and the first task data is resumed at the position of the interrupted playback based on the context snapshot 640, and the second task data is played back in sequence after the first task data playback ends.
[0137] In summary, the method of this embodiment determines the target queue for replaying the first task data in a dynamic user queue scenario, and obtains and saves a context snapshot of the first task data based on the target queue. The context snapshot is used to indicate the execution state when the replay of the first task data is interrupted. Based on the context snapshot, the execution state can be restored when the replay is resumed. That is, in response to the first task data being triggered to resume replay, the context progress is reloaded to restore the replay progress when the replay was interrupted, thereby realizing the switching of the replay of the first task data. The method for handling abnormal events during the replay process is to save the above-mentioned context snapshot and maintain the continuity of the logical queue in the event of an abnormality by switching the replay task. This addresses different task scheduling scenarios, ensures the feasibility and stability of task replay switching, and improves task replay efficiency.
[0138] The method provided in this embodiment, when the first task data is triggered to be restored and replayed, determines a recovery queue for restoring and replaying the first task data in at least one hardware queue of the replay task corresponding to the first logical queue, and restores and replays the first task data based on the hardware queue state and context snapshot retained in the recovery queue. By establishing a context saving and restoration mechanism around the logical queue, the continuity of replay in user queue scenarios with multiple submission rounds is enhanced, thereby realizing task replay switching and improving the efficiency of task replay.
[0139] Figure 7 This is a structural block diagram of a task playback device provided in an exemplary embodiment of this application, such as... Figure 7 As shown, the device includes: The storage module 710 is configured to, when replaying the first task data through at least one hardware queue, if the replay of the first task data is interrupted, obtain a context snapshot of the first task data based on the target queue and store the context snapshot. The context snapshot is used to indicate the replay progress of the first task data. The target queue is the queue in the hardware queue that generated the replay interruption when replaying the first task data. Recovery module 720 is configured to load a context snapshot of the first task data in response to the first task data being triggered for recovery playback. Recovery module 720 is also configured to recover playback of the first task data based on context snapshot.
[0140] In an optional embodiment, Figure 8 This is a structural block diagram of a task playback device provided in another exemplary embodiment of this application, such as... Figure 8 As shown, the device also includes an acquisition module 730; the acquisition module 730 is configured to acquire the hardware status corresponding to at least one hardware queue during the process of replaying the first task data through at least one hardware queue. The storage module 710 is also configured to, in response to a playback interruption of the first task data, determine from at least one hardware queue a target queue whose hardware state meets a preset state condition, the state condition being used to indicate that the hardware queue is ready to resume playback of the first task data. The storage module 710 is also configured to obtain a context snapshot based on the target queue; The save module 710 is also configured to write context snapshots to a preset persistent storage space.
[0141] In an optional embodiment, such as Figure 8 As shown, the device also includes a control module 740; the control module 740 is configured to mark a queue to be recovered in at least one hardware queue with a recovery mark, the hardware state of the queue to be recovered meeting the state conditions. The storage module 710 is also configured to, in response to a playback interrupt of the first task data, sort the queue identifiers corresponding to at least one hardware queue according to the identifier number to obtain a queue identifier sequence; The storage module 710 is also configured to determine the target queue marked with a recovery tag based on the queue identifier sequence.
[0142] In an optional embodiment, the hardware state of the queue to be recovered meets a preset state condition, including at least one of the following: the hardware state indicates that the buffer of the queue to be recovered is in a stagnant state; the hardware state indicates that the queue depth of the queue to be recovered is in an empty state.
[0143] In an optional embodiment, the recovery module 720 is further configured to determine a recovery queue from at least one hardware queue in response to a first task data being triggered to resume playback, the recovery queue being used to resume playback of the first task data; Recovery module 720 is also configured to load context snapshots via the recovery queue.
[0144] In an optional embodiment, the recovery module 720, in response to the first task data being triggered to resume playback, determines a recovery queue from at least one hardware queue, including at least one of the following: The recovery module 720 is also configured to identify the target queue as a recovery queue if the hardware status of the target queue meets the preset recovery conditions. The recovery module 720 is also configured to determine a recovery queue from at least one hardware queue whose hardware state meets the recovery conditions if the hardware state of the target queue does not meet the recovery conditions.
[0145] In an optional embodiment, the target queue includes at least one of the following: a first queue in the hardware queue that causes playback blocking when replaying the first task data; and a second queue in the hardware queue that causes playback exception when replaying the first task data, wherein the driver interface of the second queue is in a normal response state.
[0146] In an optional embodiment, the context snapshot includes at least one of the following: identification information of the first logical queue; submission batch identifier corresponding to the first task data; hardware status corresponding to the first task data, the hardware status including at least one of command buffer pointer, resource binding information, register snapshot, etc.; and replay offset in the first logical queue, the replay offset being used to indicate the replay progress of the first task data.
[0147] In an optional embodiment, the acquisition module 730 is further configured to receive a first user submission request for a first logical queue, the first user submission request being used to indicate that first task data is provided for task replay of the first logical queue; The acquisition module 730 is also configured to acquire first task data based on a request submitted by the first user, and the first logical queue is used to manage the first task data; The control module 740 is also configured to, upon acquiring the first task data, trigger the replay of the first task data via the first logical queue in response to the first logical queue being non-empty.
[0148] In an optional embodiment, the recovery module 720 is further configured to trigger the recovery playback of the first task data through the first logical queue when the second task data corresponding to the first logical queue is obtained. The control module 740 is also configured to trigger the playback of the second task data through the first logical queue in response to the end of the first task data playback.
[0149] In an optional embodiment, the acquisition module 730 is further configured to acquire the third task data corresponding to the second logical queue during the process of replaying the first task data, wherein the replay tasks corresponding to the first logical queue and the second logical queue are different. The storage module 710 is also configured to, in response to the second logical queue triggering the playback of the third task data, acquire a playback interrupt signal for the first logical queue, the playback interrupt signal being used to indicate an interruption in the playback of the first task data.
[0150] In an optional embodiment, the acquisition module 730 is further configured to acquire a playback end signal in response to the end of the first task data playback, the playback end signal being used to indicate at least one of playback execution completion and playback execution exception; The control module 740 is also configured to perform data cleanup on the recovery queue based on the playback end signal.
[0151] In an optional embodiment, the replay execution exception includes at least one of the following: the recovery queue malfunctions; the context snapshot of the first task data is corrupted; the recovery queue is in an unresponsive state; or the storage address corresponding to the first task data is abnormal.
[0152] In summary, in the device of this embodiment, a target queue for replaying the first task data is determined in a dynamic user queue scenario. A context snapshot of the first task data is obtained and saved based on the target queue. The context snapshot is used to indicate the execution state when the replay of the first task data is interrupted. Based on the context snapshot, the execution state can be restored when the replay is resumed. That is, in response to the first task data being triggered to resume replay, the context progress is reloaded to restore the replay progress when the replay was interrupted, thereby realizing the switching of the replay of the first task data. The method for handling abnormal events during the replay process is to save the above-mentioned context snapshot and maintain the continuity of the logical queue in the event of an abnormality by switching the replay task. This addresses different task scheduling scenarios, ensures the feasibility and stability of task replay switching, and improves task replay efficiency.
[0153] It should be noted that the task playback device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the task playback device provided in the above embodiments belongs to the same concept as the task playback method embodiments, and its specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0154] This application provides a processor. Figure 9 This is a schematic block diagram of a processor provided in an exemplary embodiment of this application. The processor 900 includes the task playback device provided in the above embodiments. Alternatively, an embodiment of this application provides a processor including programmable logic circuits and / or program instructions, which, when the processor is run on a computer device, are used to implement the task playback method provided in the above method embodiments.
[0155] This application provides a circuit board. Figure 10 This is a schematic block diagram of a board provided in an exemplary embodiment of this application. The board 1000 includes the task playback device provided in the above embodiment. Alternatively, this application embodiment provides a board that, when running on a computer device, is used to implement the task playback method provided in the above method embodiment. Specifically, the board can also be called a server board. A board is a type of printed circuit board (PCB), which is manufactured with a socket and can be inserted into a slot on the motherboard of a server to control the operation of hardware, such as controlling the operation of hardware devices like displays and acquisition cards. After installing a driver or computer program on the board, the board can realize the corresponding function. The driver or computer program can be installed in the processor, which controls the execution of the driver or computer program, and in conjunction with the task playback device, realizes the corresponding function of the board 1000.
[0156] This application provides a computer device. Figure 11 This is a schematic block diagram of a computer device provided in an exemplary embodiment of this application. The first computer device 1100 includes the task playback device provided in the above embodiment. Alternatively, an embodiment of this application provides a computer device. Figure 12 This is a schematic block diagram of a computer device provided in another exemplary embodiment of this application. The second computer device 1200 includes the processor provided in the above embodiments, and the computer device can be implemented as a terminal. Alternatively, an embodiment of this application provides a computer device. Figure 13 This is a schematic block diagram of a computer device provided in yet another exemplary embodiment of this application. The third computer device 1300 includes the board provided in the above embodiments, and the third computer device 1300 can be implemented as a server.
[0157] Optionally, embodiments of this application also provide a computer device, which includes: a processor and a memory, wherein the memory stores a computer program; the processor is used to execute the computer program in the memory to implement the task playback method provided in the above-described method embodiments.
[0158] Figure 14This is a schematic block diagram of a computer device provided in another exemplary embodiment of this application. The computer device is a server 140. Typically, the server 140 includes a first processor 1401 and a first memory 1402.
[0159] The first processor 1401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The first processor 1401 may be implemented using at least one hardware form selected from DSP, FPGA, and Programmable Logic Array (PLA). The first processor 1401 may also include a main processor and a coprocessor. The main processor, also known as a CPU, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the first processor 1401 may integrate a GPU, which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the first processor 1401 may also include an AI processor, which is used to handle computational operations related to machine learning.
[0160] The first memory 1402 may include one or more computer-readable storage media, which may be non-transitory. The first memory 1402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the first memory 1402 are used to store at least one instruction, which is executed by the first processor 1401 to implement the task playback method provided in the method embodiments of this application.
[0161] In some embodiments, the server 140 may optionally include an input interface 1403 and an output interface 1404. The first processor 1401, the first memory 1402, and the input interface 1403 and output interface 1404 can be connected via a bus or signal lines. Various peripheral devices can be connected to the input interface 1403 and output interface 1404 via a bus, signal lines, or a circuit board. The input interface 1403 and output interface 1404 can be used to connect at least one I / O-related peripheral device to the first processor 1401 and the first memory 1402. In some embodiments, the first processor 1401, the first memory 1402, and the input interface 1403 and output interface 1404 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the first processor 1401, the first memory 1402, and the input interface 1403 and output interface 1404 can be implemented on separate chips or circuit boards, and this application embodiment does not limit this.
[0162] Figure 15 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Optionally, the computer device 1500 is a terminal.
[0163] The computer device 1500 may be a portable mobile terminal, also referred to as a mobile terminal in this embodiment. Examples include smartphones, tablets, MP3 players, and MP4 players. The computer device 1500 may also be referred to as user equipment, portable terminal, or other names.
[0164] Typically, computer device 1500 includes: a second processor 1501 and a second memory 1502.
[0165] The second processor 1501 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The second processor 1501 may be implemented using at least one hardware form selected from DSP, FPGA, and PLA. The second processor 1501 may also include a main processor and a coprocessor. The main processor, also known as a CPU, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the second processor 1501 may integrate a GPU, which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, the second processor 1501 may also include an AI processor, which is used to handle computational operations related to machine learning.
[0166] The second memory 1502 may include one or more computer-readable storage media, which may be tangible and non-transitory. The second memory 1502 may also include high-speed random access memory devices and non-volatile storage devices, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the second memory 1502 are used to store at least one instruction, which is executed by the second processor 1501 to implement the task playback method provided in the various method embodiments of this application.
[0167] In some embodiments, the computer device 1500 may also optionally include: a peripheral device interface 1503 and at least one peripheral device. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1504, a touch display screen 1505, a camera assembly 1506, an audio circuit 1507, and a power supply 1508. The computer device 1500 also includes one or more sensors 1509. The one or more sensors 1509 include, but are not limited to: an accelerometer 1510, a gyroscope 1511, a pressure sensor 1512, an optical sensor 1513, and a proximity sensor 1514.
[0168] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on the processor. Figure 10 The structure shown does not constitute a limitation on the board. Figure 11 , Figure 12 , Figure 13 , Figure 14 and Figure 15 The structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0169] On the other hand, embodiments of this application provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, which is loaded and executed by the processor to implement the task playback method provided in the embodiments of this application above.
[0170] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the task playback method provided in the embodiments of this application as described above.
[0171] Without loss of generality, computer-readable storage media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage media are not limited to the above-mentioned types. On the other hand, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the task playback method provided in the embodiments of this application described above.
[0172] On the other hand, embodiments of this application provide a computer device including the processor described above. Optionally, the processor is a GPU. The computer device can be at least one of a portable computer, a desktop computer, a server, a server cluster, an AI computing cluster, and a cloud computing cluster. The AI computing cluster can also be simply referred to as an intelligent computing cluster or a smart computing cluster.
[0173] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A task replay method, characterized in that, The method includes: If the playback of the first task data is interrupted when replaying the first task data through at least one hardware queue, a context snapshot of the first task data is obtained based on the target queue and the context snapshot is stored. The context snapshot is used to indicate the playback progress of the first task data. The target queue is the queue in the hardware queue that caused the playback interruption when replaying the first task data. In response to the first task data being triggered to resume playback, load a context snapshot of the first task data; The first task data is replayed based on the context snapshot.
2. The method according to claim 1, characterized in that, The method further includes: During the process of replaying the first task data through the at least one hardware queue, the hardware status corresponding to the at least one hardware queue is obtained; From the at least one hardware queue, determine the queue to be recovered whose hardware status meets the preset status conditions. The status conditions are used to indicate that the queue to be recovered has the ability to recover and replay the first task data in the event of a playback interruption. In the case of replaying first task data through at least one hardware queue, if the replay of the first task data is interrupted, a context snapshot of the first task data is obtained based on the target queue, and the context snapshot is stored, including: In response to the interruption of the playback of the first task data, the target queue is obtained based on the queue to be recovered; The context snapshot is obtained based on the target queue, and the context snapshot is written to a preset persistent storage space.
3. The method according to claim 2, characterized in that, The step of determining the recovery queue from the at least one hardware queue that meets the preset state conditions includes: A recovery tag is marked on the recovery queue whose hardware state meets the state conditions, and the recovery tag includes the queue number; The step of responding to a playback interruption of the first task data by obtaining the target queue based on the queue to be recovered includes: In response to the interruption of the playback of the first task data, the queues to be recovered are sorted according to the size of the queue numbers of the markers to be recovered, to obtain a sequence of queues to be recovered; The first queue to be recovered is determined as the target queue based on the sequence of queues to be recovered.
4. The method according to claim 2, characterized in that, The step of determining the recovery queue from the at least one hardware queue that meets the preset state conditions includes at least one of the following: Identify, from the at least one hardware queue, the queue whose hardware status indicates that the queue buffer is in a stagnant state and needs to be recovered. From the at least one hardware queue, determine the queue to be recovered that has an empty hardware status indicator queue depth.
5. The method according to claim 2, characterized in that, The step of loading a context snapshot of the first task data in response to the triggering of playback recovery includes: In response to the first task data being triggered to resume playback, a recovery queue is determined from the queue to be recovered, the recovery queue being used to resume playback of the first task data; The context snapshot is loaded through the recovery queue.
6. The method according to claim 5, characterized in that, The step of determining a recovery queue from the queue to be recovered in response to the first task data being triggered for playback includes at least one of the following: If the hardware status of the target queue meets the preset recovery conditions, the target queue is determined to be the recovery queue. The recovery conditions include at least one of the following: the driver interface of the target queue is in a normal response state, the hardware of the target queue is in a normal operation state, and the resources of the target queue are valid. If the hardware status of the target queue does not meet the recovery conditions, a recovery queue whose hardware status meets the recovery conditions is determined from the queue to be recovered.
7. The method according to claim 1, characterized in that, The method further includes: Receive a first user submission request for a first logical queue corresponding to the first task data. The first user submission request is used to indicate that task data is provided for the first task playback of the first logical queue. The first logical queue is used to manage the first task data and trigger the playback of the first task data. The first task data is obtained based on the request submitted by the first user; If the first task data is obtained, in response to the first logical queue being non-empty, the first task data is replayed through the first logical queue.
8. The method according to claim 7, characterized in that, The context snapshot includes at least one of the following: The identification information of the first logical queue; The submission batch identifier corresponding to the first task data; The hardware status corresponding to the first task data includes at least one of the following: command buffer pointer, resource binding information, register snapshot, etc. The replay offset of the first logical queue, which is used to indicate the replay progress of the first task data.
9. The method according to claim 7, characterized in that, Before loading the context snapshot of the first task data in response to the triggering of playback recovery of the first task data, the method further includes: When the second task data corresponding to the first logical queue is obtained, the recovery and replay of the first task data is triggered through the first logical queue. The method further includes: In response to the end of the first task data playback, the second task data is triggered to play back through the first logical queue.
10. The method according to claim 7, characterized in that, Before the step of obtaining and storing a context snapshot of the first task data based on the target queue and then replaying the first task data via at least one hardware queue, if the replay of the first task data is interrupted, the method further includes: During the playback of the first task data, the third task data corresponding to the second logical queue is obtained, and the playback tasks corresponding to the first logical queue and the second logical queue are different. In response to the second logical queue triggering the playback of the third task data, a playback interrupt signal for the first logical queue is obtained, the playback interrupt signal being used to indicate an interruption in the playback of the first task data.
11. The method according to claim 7, characterized in that, The method further includes: In response to the playback end triggered by the first logical queue, a playback end signal is obtained, the playback end signal being used to indicate at least one of playback execution completion and playback execution exception; Data cleanup is performed on the recovery queue based on the playback end signal.
12. The method according to claim 11, characterized in that, The replay execution exception includes at least one of the following: The recovery queue malfunctioned; The context snapshot of the first task data is corrupted; The recovery queue is in an unresponsive state; The storage address corresponding to the first task data is abnormal.
13. A task playback device, characterized in that, The device includes: The storage module is configured to, when replaying first task data through at least one hardware queue, if the replay of the first task data is interrupted, obtain a context snapshot of the first task data based on a target queue and store the context snapshot, wherein the context snapshot is used to indicate the replay progress of the first task data, and the target queue is the queue in the hardware queue that caused the replay interruption when replaying the first task data. The recovery module is configured to load a context snapshot of the first task data in response to the first task data being triggered for recovery playback; The recovery module is also configured to recover and replay the first task data based on the context snapshot.
14. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction, the at least one instruction being loaded and executed by the processor to implement the task playback method as described in any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement the task playback method as described in any one of claims 1 to 12.
16. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the task playback method as described in any one of claims 1 to 12.