Coding task processing method and device, electronic equipment and readable storage medium

CN122340277BActive Publication Date: 2026-09-25MOORE THREADS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610803312.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-25
Estimated Expiration
2046-06-04

AI Technical Summary

Technical Problem

[0002]在相关多核视频编解码系统中,对硬件核心工作状态的监控通常较为粗放,往往仅能感知核心是否“存活”或通过轮询寄存器获取笼统的忙闲指示,而无法精细区分核心内部任务流所处的具体处理阶段

Benefits of technology

[0010]本实施例提供的编解码任务的处理方案,对待处理的编解码任务流,通过状态机线程来记录相应编解码核心在执行该编解码任务时的处理状态,且该处理状态包括生命周期状态和编解码工作流状态,以及该编解码工作流状态至少包含有单帧编解码执行状态、完成状态及暂时空闲等待状态,通过此种状态监测机制能够对编解码任务流的处理进度进行精细化建模与追踪。尤其通过将“暂时空闲等待状态”明确定义为完成针对非尾帧编解码操作后的合法状态,使得在完成编解码任务流的过程中能够精准区分正常的间歇性暂停与真正的执行卡死异常。在此基础上,通过将各状态的实际持续时长与对应的预设阈值进行比对,得以可靠地识别出发生在特定处理阶段内的超时异常事件,进而得以通过恢复线程依据异常所对应的实际处理状态触发与之匹配的修复操作,实现了对异常的精准定位与差异化恢复。通过应用该方案可有效避免对正常业务流程的误干扰,显著减少了因单一恢复策略导致的无效开销与业务中断范围,从而提升了编解码任务处理的可靠性、资源利用效率、对异常处置的针对性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122340277B_ABST
    Figure CN122340277B_ABST
Patent Text Reader

Abstract

The present disclosure provides a processing method and device of a coding task, electronic equipment and a readable storage medium, and relates to the technical field of video coding. The method comprises: receiving a to-be-processed coding task stream; recording, by a state machine thread, a processing state of the coding task stream by a target coding core, the processing state comprising a life cycle state of a hardware unit and a coding workflow state, the coding workflow state comprising a single-frame coding execution state, a single-frame coding completion state and a temporary idle waiting state, the temporary idle waiting state being a next processing state to which the single-frame coding completion state should be in after completing a coding operation on a current non-tail frame; generating a corresponding abnormal event for a target processing state with a processing timeout condition in the processing state; and performing abnormal repair on the abnormal event by a recovery thread in a repair mode corresponding to the target processing state. The method can improve the overall processing efficiency of the coding task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video encoding and decoding technology, specifically to the fields of codecs, scheduling, finite state machines, etc., and particularly to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing encoding and decoding tasks. Background Technology

[0002] In multi-core video codec systems, monitoring the operating status of hardware cores is often rather rudimentary. It typically only detects whether a core is "alive" or obtains a general busy / idle indication by polling registers, failing to precisely distinguish the specific processing stage of the core's internal task flow. This makes it difficult for the system to accurately identify fundamentally different situations, such as "being stuck while processing frame data" versus "normally waiting for the next frame data."

[0003] Meanwhile, when an anomaly is detected, related technologies often adopt a single, crude recovery strategy such as resetting the entire core or restarting the driver. This fails to address the specific context of the anomaly, resulting in high recovery overhead, a wide range of impacts, and a tendency to misjudge and disrupt normal business processes (for example, misjudging idle time caused by user-initiated pauses as a fault and making unnecessary interventions). This restricts the overall reliability, resource utilization efficiency, and user experience of the system. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing encoding and decoding tasks.

[0005] In a first aspect, embodiments of this disclosure propose a method for processing encoding and decoding tasks, comprising: receiving a stream of encoding and decoding tasks to be processed; recording the processing state of the target encoding and decoding core on the stream of encoding and decoding tasks through a state machine thread; wherein the processing state includes the lifecycle state of the hardware unit and the encoding and decoding workflow state, the encoding and decoding workflow state including: single-frame encoding and decoding execution state, single-frame encoding and decoding completion state, and temporary idle waiting state, the temporary idle waiting state being the next processing state that should be in after the single-frame encoding and decoding completion state has completed the encoding and decoding operation on the current non-tail frame; generating corresponding abnormal events for the target processing state where there is a processing timeout in the processing state recorded by the state machine thread; and performing abnormal repair on the abnormal events according to the repair method corresponding to the target processing state through a recovery thread.

[0006] Secondly, embodiments of this disclosure propose a processing apparatus for encoding and decoding tasks, comprising: a task stream receiving unit configured to receive a task stream to be processed; a processing status monitoring unit configured to record the processing status of the target encoding and decoding core on the task stream through a state machine thread; wherein the processing status includes the lifecycle status of the hardware unit and the encoding and decoding workflow status, the encoding and decoding workflow status including: single-frame encoding and decoding execution status, single-frame encoding and decoding completion status, and temporary idle waiting status, the temporary idle waiting status being the next processing status that the single-frame encoding and decoding completion status should be in after completing the encoding and decoding operation on the current non-tail frame; an exception event generation unit configured to generate corresponding exception events for the target processing status where processing timeout occurs in the processing status recorded by the state machine thread; and a targeted repair unit configured to perform exception repair on the exception events through a recovery thread according to the repair method corresponding to the target processing status.

[0007] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the processing method for the encoding / decoding task as described in the first aspect.

[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to perform a processing method for encoding / decoding tasks as described in the first aspect.

[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the processing method for the encoding / decoding task as described in the first aspect.

[0010] The encoding / decoding task processing scheme provided in this embodiment records the processing state of the corresponding encoding / decoding core when executing the encoding / decoding task flow through a state machine thread. This processing state includes a lifecycle state and an encoding / decoding workflow state. The encoding / decoding workflow state includes at least a single-frame encoding / decoding execution state, a completion state, and a temporary idle waiting state. This state monitoring mechanism enables refined modeling and tracking of the processing progress of the encoding / decoding task flow. In particular, by explicitly defining the "temporary idle waiting state" as a legitimate state after completing encoding / decoding operations on non-tail frames, it is possible to accurately distinguish between normal intermittent pauses and genuine execution freezes during the completion of the encoding / decoding task flow. Based on this, by comparing the actual duration of each state with corresponding preset thresholds, timeout exceptions occurring in specific processing stages can be reliably identified. Then, a recovery thread can trigger a matching repair operation based on the actual processing state corresponding to the exception, achieving precise anomaly localization and differentiated recovery. By applying this solution, erroneous interference with normal business processes can be effectively avoided, and the ineffective overhead and business interruption caused by a single recovery strategy can be significantly reduced, thereby improving the reliability of encoding and decoding task processing, resource utilization efficiency, and the targeted handling of anomalies.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart illustrating a method for processing encoding and decoding tasks provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating a method for generating and repairing a first abnormal event, as provided in an embodiment of this disclosure; Figure 4 A flowchart illustrating a method for generating and repairing a second abnormal event, as provided in an embodiment of this disclosure; Figure 5 A flowchart illustrating a method for generating and repairing a third abnormal event, as provided in an embodiment of this disclosure; Figure 6 A flowchart illustrating a method for generating and repairing a fourth abnormal event, as provided in an embodiment of this disclosure; Figure 7A flowchart illustrating a method for generating and repairing a fifth abnormal event, as provided in an embodiment of this disclosure; Figure 8 A flowchart illustrating a method for generating and repairing a sixth abnormal event, as provided in an embodiment of this disclosure; Figure 9 A flowchart illustrating a method for generating and repairing a seventh abnormal event, as provided in an embodiment of this disclosure; Figure 10 A flowchart illustrating a method for generating and repairing an eighth abnormal event, as provided in an embodiment of this disclosure; Figure 11a This is a schematic diagram illustrating the flow relationship between eight different processing states provided in an embodiment of the present disclosure; Figure 11b This is a schematic diagram illustrating the process by which multiple encoder cores and decoder cores contained within a single GPU core, provided in an embodiment of the present disclosure, jointly process different task flows. Figure 12 A structural block diagram of a processing apparatus for encoding and decoding tasks provided in an embodiment of this disclosure; Figure 13 This is a schematic diagram of the structure of an electronic device suitable for performing a processing method for encoding and decoding tasks, as provided in an embodiment of this disclosure. Detailed Implementation

[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0014] The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.

[0015] Figure 1 An exemplary system architecture 100 is shown, illustrating embodiments of processing methods, apparatuses, electronic devices, and computer-readable storage media for encoding and decoding tasks to which the present disclosure may be applied.

[0016] like Figure 1As shown, the exemplary system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0017] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include video encoding / decoding processing applications, encoding / decoding resource scheduling applications, and instant messaging applications.

[0018] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0019] Server 105 can provide various services through its built-in applications. Taking a codec resource scheduling application that provides reasonable codec resource scheduling for video codec tasks as an example, server 105 can achieve the following effects when running this application: First, terminal devices 101, 102, and 103 send the codec task stream to be processed to server 105 via network 104 according to the user's instructions. Next, server 105 uses a state machine thread to record the processing status of the target codec core on the codec task stream. The processing status includes the lifecycle status of the hardware unit and the codec workflow status, which includes: single-frame codec execution status, single-frame codec completion status, and temporary idle waiting status. The temporary idle waiting status is the next processing status that the single-frame codec completion status should be in after completing the codec operation on the current non-tail frame. Next, in the processing status recorded by the state machine thread, corresponding abnormal events are generated for the target processing status where there is a processing timeout. Finally, the abnormal events are repaired by the recovery thread according to the repair method corresponding to the target processing status.

[0020] Furthermore, the server 105 can return the result data obtained after the encoding and decoding process to the terminal devices 101, 102, and 103 via the network 104.

[0021] It should be noted that the stream of codec tasks to be processed can be obtained from terminal devices 101, 102, and 103 via network 104, or it can be pre-stored in the task queue on the local machine of server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when it starts processing previously stored codec tasks), it can choose to obtain this data directly from the local machine. In this case, the exemplary system architecture 100 may not include terminal devices 101, 102, and 103 and network 104.

[0022] Since encoding and decoding a task stream containing a large number of video frames requires significant computing resources and power, the encoding and decoding task processing methods provided in the subsequent embodiments of this disclosure are generally executed by a server 105 or a server cluster with strong computing power and abundant computing resources. Correspondingly, the encoding and decoding task processing device is also generally located in the server 105 or server cluster. However, it should also be noted that when the terminal devices 101, 102, and 103 also have sufficient computing power and resources, the terminal devices 101, 102, and 103 can also complete the aforementioned calculations performed by the server 105 through the encoding and decoding resource scheduling applications installed on them, and thus output the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the encoding / decoding resource scheduling application determines that the terminal device has strong computing power and sufficient remaining computing resources, the terminal device can be allowed to perform the aforementioned operations, thereby appropriately reducing the computing pressure on server 105. Correspondingly, the processing device for encoding / decoding tasks can also be located in terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude server 105 and network 104.

[0023] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0024] Please refer to Figure 2 , Figure 2 A flowchart of a method for processing encoding and decoding tasks provided in this embodiment of the disclosure, wherein process 200 includes the following steps: Step 201: Receive the stream of encoding / decoding tasks to be processed; This step is intended for the execution body of the encoding / decoding task processing method (e.g., Figure 1 The server 105, server cluster, and (specifically, codec resource scheduling applications running on servers or high-performance terminal devices) shown receive the pending codec task streams.

[0025] The encoding / decoding task flow refers to a collection of video frame data and their processing requirements that are temporally or logically related. It should be a work description entity encapsulating the processing context. Its core content typically includes, but is not limited to: an operation type identifier indicating whether it is an encoding or decoding task; source and destination address information of the video data; the encoding / decoding standard used (e.g., H.264 / AVC, H.265 / HEVC, AV1, etc.); key technical parameters (e.g., resolution, frame rate, bitrate, etc.); and the task's own attributes (e.g., priority, session number, etc.). It should be noted that the boundary determination criteria for a so-called complete 'encoding / decoding task' described in the embodiments of this application all depend on the logical integrity of the processing unit carried by the task flow. As one possible implementation, this task can correspond to all consecutive video frames contained in a task package encapsulated by upper-layer services. Within this unit, all frames belong to the same processing target and share the same context. As another possible implementation, in continuous video streaming scenarios, task boundaries can also be defined by 'start frames' and 'end frames' actively labeled by the upper-level business layer. This means that the start and end of an independent processing cycle are announced through explicit flow control protocols or frame marking signals. The typical representation of a task in software data structures is that all unprocessed data blocks are associated with a globally unique session ID or task ID, ensuring that the driver can accurately classify data packets during reading and assign them a unique hardware core and state machine context.

[0026] From a system interaction perspective, this sending action implies the existence of a higher-level scheduler or management system. This scheduler, based on strategies such as global load, core idle status, and task priority, decides to assign a newly arrived or queued task to the current target codec core and its associated management module. Essentially, this receiving process is the initialization operation where the executing entity responds to the scheduling instruction and parses and carries the task flow description information. In practical terms, this receiving behavior can be implemented through various inter-process communication or system call mechanisms. For example, in a message queue-based architecture, the upstream scheduler places the encapsulated task flow descriptor into a specific message queue, and the resource scheduling application can then act as a consumer to asynchronously retrieve tasks from this queue.

[0027] Furthermore, in cloud transcoding services, "receiving" might correspond to retrieving a video segment to be transcoded from a distributed task queue; while in real-time communication terminals, it might correspond to receiving a media stream handle that needs to be encoded or decoded in real time from a camera acquisition module or network receiving module. In addition, this step typically includes preliminary validity checks, such as checking whether the encoding / decoding format is supported by the current hardware core and whether the parameters are within the allowed range. If the check fails, an error might be immediately returned upstream without entering the subsequent status monitoring loop, thus reducing the ineffective waste of resources caused by handling low-level errors.

[0028] Step 202: Record the processing status of the target codec core on the codec task flow through the state machine thread; Building upon step 201, this step aims to have the aforementioned execution entity record the processing state of the target codec core for the codec task flow via a state machine thread. This processing state can include the lifecycle state of the hardware unit (i.e., the physical hardware manifestation of the target codec core) and the codec workflow state. The lifecycle state describes different life stages of the codec core before and after supporting codec task processing, such as power-on, initialization, and power-off states. The codec workflow state describes different working stages of the codec core during codec task processing, and may include at least: a single-frame codec execution state, a single-frame codec completion state, and a temporary idle waiting state. The temporary idle waiting state indicates the next processing state after completing the codec operation on the current frame (which is not the last frame). The last frame is the last frame in the codec task flow that requires codec operation, and a non-last frame indicates that the current frame is not the last frame.

[0029] The technical solution described in this step aims to actively and continuously observe and divide the internal stages of the hardware codec core's task execution process through software-defined monitoring logic. This transforms what was originally a black-box, continuous processing process into a series of discrete processing states with clear business meanings. The state machine thread can refer to an independent monitoring thread, daemon process, or finite state machine (FSM) running in the codec resource scheduling application on a server or terminal device. It can obtain real-time activity information of the target core by periodically querying (e.g., through driver interfaces, hardware registers, or kernel events) or passively receiving event notifications from the codec core driver layer. Its direct monitoring and recording object is the stage or mode the core is in when processing the task flow, and it abstracts this stage information into a predefined, finite set of states.

[0030] Of the at least three core encoding / decoding workflow states described in this step, the "single-frame encoding / decoding execution state" refers to the period when the hardware circuitry or dedicated logic of the target encoding / decoding core is performing actual encoding or decoding operations on a specific frame of data in the task flow. From a monitoring perspective, this state may correspond to the core's busy flag being set, a specific working register continuously changing, or power consumption being in an active high range. The "single-frame encoding / decoding completion state" is a crucial instantaneous or briefly held state, indicating that the hardware operation for processing a frame of data has ended, the result has been written to the designated buffer, and the core is ready to receive the next instruction. In practice, this state can be indicated by a hardware interrupt, driver callback, or task completion signal. The "temporarily idle waiting state" is explicitly defined as the legitimate waiting state that the codec should enter when the "single-frame encoding / decoding completion state" confirms that the currently processed frame is not the last frame of the task flow (i.e., not the tail frame). This precise definition strictly distinguishes, from a technical principle perspective, normal process interruptions caused by business logic from abnormal pauses caused by hardware or software failures. To more clearly illustrate the correspondence between states and individual frames, and the specific method of state reporting, the following example uses a typical implementation scenario: Assume a certain encoding / decoding task stream contains multiple consecutive video frames (frame numbers A, B, C, etc.). The state machine thread can maintain a frame-level state tracking variable for each task stream. For the currently being processed frame (e.g., frame A), the state machine thread records a triplet: the current frame number, the current processing state, and the duration of that state. When the hardware decoding operation for frame A is complete, the driver reports an event via an interrupt or callback message. This event carries the following data structure: {frame number: A, state identifier: decode_done, timestamp: T1}. Upon receiving this report, the state machine thread updates the state of frame A in the task stream to 'single frame encoding / decoding complete state'. Subsequently, since frame A is not the last frame (frame B is still pending processing), and the data for frame B is not yet ready, the task stream immediately switches to 'temporarily idle waiting state'. At this point, the state machine thread can update the current state of the task flow to {frame number: A (keeping it pointing to the previous processed frame), state identifier: stream_idle, timestamp: T2}, and start timing for this idle state. Until the upper-layer business sends data for frame B, the driver triggers a state switch again, and the state machine thread can record {frame number: B, state identifier: decoding / encoding, timestamp: T3}, thus starting to monitor the execution status of frame B. Therefore, each 'single-frame encoding / decoding complete state' strictly corresponds to a specific frame and is uniquely identified by its frame number; while the 'temporarily idle waiting state' is a transitional state of the task flow between two frames, and its associated frame number is fixed to the previous completed frame until the next frame begins processing.This reporting mechanism, based on a combination of frame identifiers and state identifiers, enables the state machine thread to perform fine-grained monitoring at the frame-by-frame level without causing confusion in state attribution due to the characteristics of streaming processing.

[0031] In continuous streaming video encoding and decoding, it is perfectly normal for the core to briefly refrain from actual encoding and decoding operations after the completion of non-tail frames, due to data pipeline incompleteness, downstream module processing delays, or simple inter-frame scheduling intervals. However, monitoring solutions provided by related technologies often fail to recognize this legitimacy and misjudge it as a freeze, triggering unnecessary recovery operations and disrupting normal business processes. This step introduces and correctly defines this state, providing an accurate benchmark for subsequent anomaly detection based on duration. Furthermore, in addition to the three core states described in this step, more processing states can be included according to actual needs, such as power-on state, power-off state, encoding / decoding core initialization, and encoding / decoding parameter initialization.

[0032] Furthermore, in addition to identifying the current processing state, this state machine thread can also timestamp each state, recording its entry time, thereby enabling the calculation of its "actual duration." Monitoring behavior can typically be periodic polling or event recording triggered during state transitions. For the "temporarily idle waiting state," its starting point is the moment the "single-frame encoding / decoding completion state" is confirmed (and it is known not to be the last frame), and its ending point is the moment the next frame of data begins processing (i.e., entering the next "single-frame encoding / decoding execution state"), or the moment a timeout triggers an exception. This monitoring mechanism constructs a time-series map that accurately reflects the actual working state changes of the encoding / decoding core, providing a primary basis for subsequent time-based anomaly diagnosis.

[0033] Step 203: Generate corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread. Building upon step 202, this step aims to have the aforementioned execution entity generate corresponding exception events for any target processing state (i.e., any processing state with a timeout) within any processing state recorded by the state machine thread. For example, this can be done by comparing the actual duration of any processing state with a corresponding preset duration upper limit, and then confirming which processing states have timeouts. In other words, a reasonable time expectation boundary can be pre-established for each clearly defined processing state. By continuously measuring the state dwell time and comparing it with this boundary, abstract exception possibilities are transformed into concrete, operable timeout events. The actual duration refers to the time interval elapsed from when the state machine thread confirms the encoder / decoder core enters a specific processing state until the current moment or until it leaves that state. The corresponding preset duration upper limit is the maximum allowable time threshold configured individually for each state based on prior knowledge, historical statistics, or performance benchmarks. It represents the reasonable upper limit of time the system believes should be spent completing its expected operation in that state. For example, the preset duration of the "single frame encoding / decoding execution state" may be dynamically set based on frame type, resolution, and core computing power; while the preset duration of the "temporary idle waiting state" may be set longer based on data pipeline latency and business tolerance to distinguish between normal waiting and abnormal deadlock.

[0034] In practical terms, this comparison and judgment logic is typically implemented using timed polling or an event-driven mechanism within a state machine thread. This state machine thread maintains the state sequence experienced by each task flow on the current core and the entry timestamp of each state. When a state is continuously monitored, the system periodically, or at a possible transition moment, calculates the difference between the current time and the state entry timestamp to obtain the "actual duration". Subsequently, it reads the "preset duration upper limit" threshold corresponding to the state type from the preset configuration. The comparison operation itself is a simple numerical judgment: if the actual duration is greater than the preset duration upper limit, it is determined that "a processing timeout has occurred", and a corresponding exception event is generated. This event is usually a structured data object, which can at least contain the following information: exception type identifier (associated with a specific timeout state, such as single-frame encoding / decoding execution timeout), associated target encoding / decoding core identifier, associated encoding / decoding task flow identifier, the time point of the timeout, and the duration of the state that triggered the exception. Generating this event signifies that the exception has been formally identified and recorded, and placed in a pending queue or notification channel. This decouples the "monitoring and judgment" stage from the subsequent "repair execution" stage, enabling the aforementioned execution entities to handle exceptions asynchronously and in an orderly manner.

[0035] Furthermore, the preset duration limits corresponding to different processing states do not necessarily have to be static values; they can also be variables that are dynamically adjusted based on system load, historical average processing time, or encoding / decoding task complexity. The process of generating abnormal events can also include preliminary filtering or aggregation logic. For example, to avoid false alarms caused by momentary jitter, an event is only ultimately generated if timeouts are detected in multiple consecutive monitoring periods.

[0036] Step 204: Perform exception repair on the exception event using the recovery thread according to the repair method corresponding to the target processing state.

[0037] Based on step 203, this step aims to improve the overall encoding and decoding efficiency by having the aforementioned execution entity perform abnormal repair on abnormal events in accordance with the repair method corresponding to the target processing state through a recovery thread (which can be set independently in the state machine thread to avoid being affected by the state machine thread).

[0038] This solution predefines a mapping relationship between each potentially timed-out processing state and its most matching and likely effective recovery strategy. Therefore, when an abnormal event carrying the "target processing state" identifier is generated, the aforementioned execution entity can, based on this mapping relationship, not take a uniform "restart" or "retry" operation, but trigger a customized repair process for that specific state's failure mode. This idea stems from an in-depth analysis of the different working stages represented by different states and their typical root causes of failure. For example, the root causes and required recovery actions for a pause occurring during the data computation stage are completely different from those for a pause occurring during the synchronization signal waiting stage. Therefore, the essence of this step is to achieve "targeted treatment" for fault repair, resolving the fault with minimal intervention and maximizing the protection of normal business data and the overall system stability.

[0039] In practice, this step can be handled by the "repair executor" module or sub-thread within the application responsible for resource scheduling and management. This module listens to an exception event queue. When an exception event is detected, it first parses out the key field, namely the "target processing state" that triggered the exception. Then, a "state-repair strategy mapping table" maintained internally by the module is queried. This mapping table is indexed by state type, linking to a specific repair operation function or a series of repair instructions. For example, the mapping table might associate "single-frame codec execution state" with the instruction sequence "attempt to reset the core context and retry the current frame," and "temporarily idle waiting state" with the instruction sequence "check and release potentially deadlocked synchronization resources." Further, based on the query results, the repair executor calls the corresponding function or executes instructions sequentially to perform repair operations on the target codec core and its associated task flow context. The entire repair process is usually automated, and upon completion, it reports the repair result (success or failure) to the monitoring system. It can also update the state machine thread (or finite state machine) corresponding to the codec core, returning it to a safe initial or ready state.

[0040] Furthermore, the anomaly repair mechanism provided in this step can be designed as a layered or rollback-enabled mechanism. For example, for a given state, a primary repair plan and one or more alternative repair plans can be configured. When the primary repair plan fails, the aforementioned execution entity can automatically try alternative plans or escalate the anomaly report. In addition, the repair operation itself can be atomic or contain multiple sub-steps. To enhance system robustness, before executing critical repair operations, a snapshot backup of the current task flow context (such as processed frame data and parameters) can be taken, allowing for partial rollback in case of repair failure and preventing complete state corruption. This differentiated repair strategy based on precise state mapping significantly reduces the secondary risks and performance overhead introduced by the recovery operation itself, thereby improving overall task execution efficiency.

[0041] The encoding / decoding task processing method provided in this disclosure uses a state machine thread to record the processing state of the corresponding encoding / decoding core when executing the encoding / decoding task flow. This processing state includes a lifecycle state and an encoding / decoding workflow state. The encoding / decoding workflow state includes at least a single-frame encoding / decoding execution state, a completion state, and a temporary idle waiting state. This state monitoring mechanism enables refined modeling and tracking of the processing progress of the encoding / decoding task flow. In particular, by explicitly defining the "temporary idle waiting state" as a legitimate state after completing encoding / decoding operations on non-tail frames, it is possible to accurately distinguish between normal intermittent pauses and genuine execution deadlock anomalies during the completion of the encoding / decoding task flow. Based on this, by comparing the actual duration of each state with the corresponding preset threshold, timeout anomalies occurring in specific processing stages can be reliably identified. Then, a recovery thread can trigger a matching repair operation based on the actual processing state corresponding to the anomaly, achieving precise anomaly localization and differentiated recovery. By applying this solution, erroneous interference with normal business processes can be effectively avoided, and the ineffective overhead and business interruption caused by a single recovery strategy can be significantly reduced, thereby improving the reliability of encoding and decoding task processing, resource utilization efficiency, and the targeted handling of anomalies.

[0042] Based on the above embodiments illustrating the overall solution, to further deepen the understanding of how to determine timeouts, generate corresponding abnormal events, and handle abnormal events in different processing states, the following will provide specific explanations for different actual processing states through several different embodiments: Please refer to the following first. Figure 3 , Figure 3 The flowchart of a method for generating and repairing a first abnormal event provided in this embodiment of the disclosure targets a single-frame encoding / decoding execution state included in the processing state, wherein process 300 includes the following steps: Step 301: In response to the fact that the actual duration of the single frame codec execution state is greater than the preset duration limit corresponding to the single frame codec execution state, determine that there is a processing timeout in the single frame codec execution state, and generate a first abnormal event for the single frame codec execution state with a processing timeout. When the state machine thread continuously tracks the target codec core to a single-frame encoding / decoding execution state, it initiates a timing logic to calculate the actual duration of this state since its inception. This state signifies that the core's hardware circuitry is fully engaged in mathematical transformations and data processing for encoding or decoding a specific frame in the task flow. The execution entity compares this actual duration with a preset upper limit for this state. This upper limit is typically a reasonable time boundary set based on factors such as the frame's encoding complexity, resolution, the core's nominal computing power, and historical average processing time. Once the comparison confirms that the actual duration exceeds this upper limit, the execution entity determines that a processing timeout has occurred in the current state. This usually means that the codec core encountered an unexpected obstacle while processing the frame, such as getting stuck in a complex computation loop, accessing an abnormal memory address causing a suspension, or a transient hardware unit failure. Therefore, the execution entity generates a structured first exception event. The event is a data object containing a specific identifier. It records at least the exception type (i.e., single frame execution timeout), the core number where the exception occurred, the associated task flow ID, and the sequence number of the frame that malfunctioned, so as to provide as complete a context as possible for subsequent accurate repair.

[0043] Step 302: Repair the first abnormal event by using at least one of the following first repair methods through the recovery thread: initialize the target codec core, replace the memory used, and re-encode and decode the current frame to be encoded and decoded.

[0044] Based on preset fault diagnosis logic, the aforementioned execution entity selects at least one of several repair strategies to attempt recovery: The first strategy is to initialize the target codec core, which typically means sending a reset or reinitialization command to the hardware core via the driver, restoring all its internal registers, pipelines, and state machines to a known clean initial state, thereby clearing deadlocks that may be caused by software instruction sequence errors or transient hardware state disturbances; The second strategy is to replace the memory used. Considering that the encoding and decoding process requires frequent reading and writing of frame data and intermediate results, timeouts may originate from failed access to a specific physical memory page or buffer. This repair method attempts to switch the input, output, or intermediate buffers used by the current task flow to another pre-allocated spare memory area to avoid possible memory hardware defects or software mapping errors; The third strategy is to re-encode and decode the current frame to be encoded and decoded. After ruling out core state and memory problems, timeouts may also originate from occasional bit errors or transient interference in the frame data stream. This method retrieves the original frame data from the data cache, along with the necessary parameters, and resubmits it to the recovered core for processing.

[0045] Furthermore, the three repair methods mentioned above can be combined or executed sequentially based on configuration or a preliminary judgment of the error type. For example, the execution entity can first attempt the lightest reprocessing of the current frame operation; if it fails rapidly and consecutively, it can then perform memory replacement; if the problem persists, it can perform the most thorough kernel initialization. After each repair operation, the state machine thread can reassess the kernel state; if it returns to normal, the task continues; otherwise, it may try the next strategy or report a higher-level fault. This mechanism enables the execution entity to "self-heal" with minimal impact and maximum speed.

[0046] Please refer to the following: Figure 4 , Figure 4 The flowchart of a method for generating and repairing a second abnormal event provided in this embodiment of the disclosure is for a single-frame encoding / decoding completion state included in the processing state, wherein process 400 includes the following steps: Step 401: In response to the fact that the actual duration of the single frame encoding and decoding completion state is greater than the preset duration limit corresponding to the single frame encoding and decoding completion state, determine that there is a processing timeout in the single frame encoding and decoding completion state, and generate a second abnormal event for the single frame encoding and decoding completion state with a processing timeout. The single-frame encoding / decoding completion state is a critical control point. It signifies that the hardware core has finished its substantive computational processing of a frame of data, and the result has been output. However, the system software layer has not yet fully completed the subsequent cleanup and state transition logic. This is typically a transient state that should be transitioned to quickly. But if the state machine thread finds that the actual duration of the encoding / decoding core in this state exceeds the preset very short upper limit threshold for that state, it determines that a timeout has occurred. This timeout is often not due to insufficient hardware computing power, but more likely to stem from software-level signal transmission failures or resource coordination blockages. For example, a hardware completion interrupt may have occurred, but the driver failed to respond in time and notify the upper-layer application; or the software semaphore or callback function used to mark the completion of frame processing may have failed to execute in time due to thread scheduling issues. The "second exception event" generated at this time is used to mark the stagnation of this completion confirmation process.

[0047] Step 402: The second abnormal event is repaired by the recovery thread using at least one of the following second repair methods: resend the encoding / decoding completion notification corresponding to the current frame to the state machine thread, and re-encode / decode the current frame.

[0048] Building upon step 401, this step provides a precise repair solution for this communication or coordination layer failure. The first repair method involves resending the encoding / decoding completion notification corresponding to the current frame to the state machine thread. This works by attempting to resend or re-trigger the potentially lost completion signal. This can be achieved by resimulating a hardware interrupt callback, resetting the completion flag, or resubmitting a completion message to the message queue, aiming to wake up monitoring or scheduling logic that may be waiting due to event loss. The second repair method, re-encoding / decoding the current frame, is a more conservative but thorough recovery strategy. It is based on the judgment that since the completion confirmation process has experienced unrecoverable chaos, the safest approach is to have the hardware core re-execute from the processing start point of the frame. The execution entity will resubmit the frame data to the encoding / decoding core, which is equivalent to using a complete recalculation to overwrite previously completed operations that may be in an uncertain state, naturally generating a new, complete processing completion event, thus bypassing the previous state machine bottleneck.

[0049] The two anomaly repair methods mentioned above can be regarded as a progressive relationship from lightweight retries to heavy reconstruction. The aforementioned execution entity can prioritize trying the low-cost resending notification operation; if the status is still not progressing after a short time, the repair of reprocessing the current frame will be initiated.

[0050] By specifically addressing this state, the finite state machine is ensured to reliably continue even with brief communication failures in the processed encoding / decoding task stream, maintaining the continuity of the processing flow.

[0051] Please continue to refer to this. Figure 5 , Figure 5 The flowchart of a method for generating and repairing a third abnormal event provided in this embodiment of the disclosure targets a temporary idle waiting state included in the processing state, wherein process 500 includes the following steps: Step 501: In response to the actual duration of the temporary idle waiting state being greater than the preset duration limit corresponding to the temporary idle waiting state, determine that there is a processing timeout in the temporary idle waiting state, and generate a third abnormal event for the temporary idle waiting state with a processing timeout. As described in step 202 regarding the temporary idle waiting state, this state is defined as a legitimate waiting phase after completing non-tail frame encoding and decoding. Its preset duration is typically set longer than that of the execution state to accommodate reasonable pipeline delays or scheduling intervals. When the state machine thread detects that the actual time the core spends in this state exceeds this relaxed but still limited threshold, it determines that a timeout has occurred. Technically, this timeout strongly suggests that the task flow is not waiting healthily, but rather is stuck in some kind of synchronization or resource coordination deadlock, commonly known as deadlock or livelock. For example, the producer-consumer queue may be full and unable to be consumed, or the mutex lock may fail to release as expected, causing the next frame processing start signal to never be triggered. The "third exception event" generated at this time is specifically used to identify this type of synchronization failure.

[0052] Step 502: For the third abnormal event, perform deadlock detection on the encoding and decoding task flow through the recovery thread, and perform unlocking and reset processing on the detected deadlock.

[0053] Building upon step 501, this step provides specific repair operations for this type of synchronization failure. Its core is to perform deadlock detection on the encoding / decoding task flow and then unlock and reset any detected deadlocks. In practice, this is typically performed by the diagnostic module in a resource scheduling application. Deadlock detection can be accomplished by analyzing the lock resources held by the task flow and its associated cores, the semaphores they are waiting for, and the dependency graph between them. For example, the execution entity can check whether the worker threads of the task flow are permanently waiting for a mutex that cannot be released, or whether the associated buffer queue is in a state of being both full (producer blocked) and unable to be consumed (consumer blocked). Once a specific resource (such as a lock object or queue) involved in the deadlock is detected, the repair module will implement unlock and reset procedures. This may include forcibly releasing the held locks, clearing and resetting the blocked buffers, or terminating and restarting the relevant threads or subtasks in the deadlock loop, thereby breaking the deadlock. After a reset, the aforementioned execution entity may attempt to restore the task flow from an abnormal deadlock-like idle state to a node where it can safely continue, such as rolling back to the checkpoint where the previous frame was completed, or reinitializing the relevant data channels.

[0054] Similarly, deadlock detection and remediation can be designed as a tiered strategy. Initial detection might only quickly identify obvious circular waits and attempt to force an unlock. If the problem is complex or rapid remediation fails, the aforementioned execution entity can initiate more in-depth analysis, such as generating resource dependency snapshots for offline diagnostics. Furthermore, after successful unlocking and reset, the execution entity can record the context information that led to the deadlock, which can be used to optimize subsequent task scheduling strategies and prevent the recurrence of the same pattern, thereby improving the system's preventative capabilities while resolving the fault.

[0055] The solution provided in this embodiment can ensure that stagnation caused by software logic complexity can be effectively identified and eliminated, thus guaranteeing the overall robustness of long-running multi-task flow processing.

[0056] In the above passage Figures 3 to 5 Based on the detailed descriptions of the single-frame encoding / decoding execution state, single-frame encoding / decoding completion state, and temporary idle waiting state in the corresponding embodiments, further references can be made. Figure 6 , Figure 6 The flowchart of a method for generating and repairing a fourth abnormal event provided in this embodiment of the disclosure is intended to address the case where the encoding / decoding workflow state may also include a task flow end state. This task flow end state is a single-frame encoding / decoding completion state indicating the next processing state after the encoding / decoding operation of the last frame in the encoding / decoding task flow is completed. The process 600 includes the following steps: Step 601: In response to the actual duration of the task flow end state being greater than the preset duration limit corresponding to the task flow end state, determine that there is a processing timeout in the task flow end state, and generate a fourth abnormal event for the task flow end state with a processing timeout. Compared to the previous embodiment, this embodiment adds a new state type to the processing state: the task flow end state. Technically, this state signifies that the lifecycle of a single encoding / decoding task flow has reached its end. That is, the hardware core has successfully processed the last frame (tail frame) of data in the task flow, and all core computational work has been completed. The executing entity then needs to perform subsequent cleanup operations such as resource cleanup and context release. This state is similar to... Figure 3 The difference between the single-frame encoding / decoding completion states described in the embodiments is that the former indicates the end of the entire task, while the latter only indicates the completion of processing a single frame. Defining this state allows the aforementioned execution entity to further distinguish between the two different stages of "processing a frame sequence" and "ending the entire task," thereby enabling independent health monitoring of the task completion process itself.

[0057] This step specifically focuses on timeout detection for this state. Once the state machine thread confirms that the core has entered the task flow completion state, it starts timing for this state. The preset maximum duration of this state is usually set to a reasonable time window that allows for normal cleanup operations. If the actual time exceeds this threshold, it is considered a timeout. This timeout usually indicates an obstacle encountered during task completion, such as: delays when releasing large amounts of memory or shutting down specific hardware function channels; conflicts when returning resources to the resource manager; or software errors causing the state machine to fail to correctly respond to the task completion confirmation signal. The fourth exception event generated in this case is specifically used to identify the specific failure mode of "task completion process stall."

[0058] Step 602: Repair the fourth abnormal event by using at least one of the following third repair methods through the recovery thread: initialize the task parameters of the encoding / decoding task flow, and call the preset hardware decoding repair tool.

[0059] This step provides a dedicated repair strategy for this type of end-of-phase failure, known as the "Third Repair Method": The first method initializes the task parameters of the encoding / decoding task flow. Its principle is to attempt to reset the task control block at the software level. This includes clearing or reinitializing all configuration parameters, handles, and status flags associated with the task flow, aiming to resolve cleanup deadlocks that may be caused by parameter corruption or resource reference counting errors, allowing the system to retry ending the process from a clean software starting point. The second method invokes a preset hardware decoding repair tool, which is a hardware-level intervention. This repair tool refers to a set of dedicated command sequences provided in the driver or firmware, used to force the hardware core to exit the current task mode, clear the internal pipeline cache, and reset the hardware state machine associated with a specific task. This is equivalent to forcibly performing a hardware reset at the hardware level to unbind the task and ensure that hardware resources are released.

[0060] Furthermore, the two repair methods provided above can be viewed as progressive interventions from the software to the hardware level. The executing entity can first attempt to initialize task parameters. If no progress is observed in the state (such as resources being successfully released) within a shorter time, it escalates to calling a lower-level "hardware repair tool." Additionally, after performing the repair, the executing entity can also record detailed context information that caused the task to time out (such as the type of resource that failed to be released) for subsequent analysis and optimization of resource management strategies, thereby preventing similar problems from recurring in future tasks. This embodiment aims to maintain controllability and recoverability even at the very end of the task lifecycle through the above mechanism.

[0061] Please continue to refer to further information. Figure 7 , Figure 7 The flowchart of a method for generating and repairing a fifth abnormal event provided in this embodiment of the disclosure is intended for situations where the lifecycle state includes the codec core initialization state. The codec core initialization state is a task flow end state indicating the next processing state after the processing of the codec task flow has ended. The process 700 includes the following steps: Step 701: In response to the actual duration of the codec core initialization state being greater than the preset duration limit corresponding to the codec core initialization state, determine that there is a processing timeout in the codec core initialization state, and generate a fifth exception event for the codec core initialization state with a processing timeout. Building upon the above embodiments, this embodiment further introduces the codec core initialization state as a new processing state in the finite state machine model. Technically, this state represents a transitional preparation phase: after the core successfully completes a codec task flow (entering the task flow completion state), it does not immediately enter the active state to process the next frame. Instead, it needs to return to a clean, ready initial state to prepare for potential subsequent tasks. The core activities of this state typically include: clearing the residual context of the previous task, resetting relevant hardware configuration registers, reclaiming dedicated memory, and resetting the internal state machine to standard standby mode. By defining this independent state, this embodiment enables the aforementioned execution entity to independently monitor the crucial process of core reset and preparation between tasks, ensuring that the process of releasing old task resources and preparing for new tasks is healthy and controllable.

[0062] This step involves timeout monitoring for this preparation phase. The state machine thread starts timing when the core enters the encoding / decoding core initialization state. The preset maximum duration of this state is based on the normal time consumption of initialization operations (such as register reset and memory reclamation). If the actual time consumption exceeds the limit, it is determined to be an initialization timeout. This type of timeout is often not a computationally intensive problem, but is more likely to indicate resource release conflicts or internal coordination deadlocks. For example, it may be due to a submodule failing to confirm the reset completion in time, or the driver software being blocked while waiting for an internal resource to become available. The fifth exception event generated at this time identifies this specific failure mode of the core reset process stalling.

[0063] Step 702: For the fifth abnormal event, the recovery thread performs lock detection on the worker threads of the encoding and decoding task flow, and unlocks the detected locks.

[0064] This step provides a dedicated remediation strategy for such coordination failures: lock detection is performed on the worker threads of the encoding / decoding task flow, and any detected locks are unlocked. In practice, this operation targets the software worker threads managing the encoding / decoding core. During initialization, these threads may need to acquire and release a series of software locks (such as mutexes and spinlocks) to safely access shared driver data structures or hardware registers. The remediation module diagnoses the state of these worker threads, checking whether they are waiting for a lock already held by another thread and unable to be released, or whether their own locks were not released at the correct node in the initialization sequence due to a logical error. Once such deadlocks or lock leaks are detected, the remediation operation forcibly releases these illegally held locks, possibly by sending a forced release command to the lock manager or safely terminating and restarting the worker thread. Its purpose is to break down software synchronization barriers that prevent the initialization process from progressing.

[0065] Furthermore, the repair mechanism provided in this embodiment can also be combined with resource auditing. That is, after the unlocking process, the aforementioned execution entity can further verify and release the memory or hardware resources that were not properly reclaimed due to the lock problem, ensuring that the core returns to a truly clean state. In addition, this process can also record the call stack or resource identifier that caused the lock contention, which can be used for subsequent optimization of the driver layer code, reducing the probability of similar conflicts occurring in future initialization processes, thereby improving the long-term stability of the system while resolving the current fault.

[0066] Please continue to refer to further information. Figure 8 , Figure 8 The flowchart of a method for generating and repairing a sixth abnormal event provided in this embodiment of the present disclosure is intended to address the case where the codec workflow state may also include a codec parameter initialization state. This codec parameter initialization state is the next processing state that the target codec core of a codec task flow still awaiting processing should be in after successfully completing core initialization through the codec core initialization state. The next processing state after the codec parameter initialization state is the single-frame codec execution state. Flow 800 includes the following steps: Step 801: In response to the actual duration of the codec parameter initialization state being greater than the preset duration limit corresponding to the codec parameter initialization state, determine that there is a processing timeout in the codec parameter initialization state, and generate a sixth exception event for the codec parameter initialization state with a processing timeout. Building upon the above embodiments, this embodiment further introduces a codec parameter initialization state as a new processing state in the finite state machine model. This state follows immediately after the codec core initialization state. Its technical principle is that when a core that has completed general initialization is about to serve a specific codec task flow, the hardware codec must be precisely and individually configured according to the specific requirements of that task flow. This process involves parsing, verifying, and ultimately loading a series of setting parameters (e.g., video encoding standard, resolution, frame rate, bitrate control mode, GOP structure, number of reference frames, etc., where GOP is an abbreviation for Group of Pictures) from the task flow descriptor into the corresponding configuration registers or internal memory of the hardware core. Defining this independent state allows the execution entity to separate the general core readiness check from the parameter adaptation process for a specific task, thereby implementing independent health monitoring and protection for this critical and error-prone configuration step.

[0067] This step involves continuous monitoring of the parameter loading and configuration phase. Once the codec core enters the codec parameter initialization state, the state machine thread begins timing. The preset time limit for this state is typically set based on parameter complexity and hardware configuration speed. If the actual time exceeds the threshold, it is considered a timeout. Technically, this timeout strongly suggests that the problem is not due to core hardware failure or insufficient general resources, but rather is likely related to the input parameters themselves or their compatibility with the hardware. For example, the parameters may contain values ​​beyond the hardware's supported range (such as unsupported encoding levels), have internal contradictions (such as a severe mismatch between resolution and bitrate), or have been corrupted during transmission or parsing. The generated "sixth exception event" is specifically used to identify specific pre-startup failures such as "parameter configuration failure."

[0068] Step 802: For the sixth exception event, the incoming setting parameters are validated by the recovery thread, and the reason for the validation failure is returned.

[0069] This step provides a precise and efficient repair and diagnosis strategy for such parameter-related faults: validating the incoming setting parameters and returning the reason for the validity verification failure. In practice, this operation initiates a diagnostic verification process. The repair module (or a dedicated parameter verification submodule) reacquires or retrieves the task flow parameter set that caused the timeout from the context, and then performs a comprehensive validity check based on known hardware capability specifications, codec standard syntax rules, and system policies. This may include syntax checks (such as parameter structure integrity), semantic checks (such as whether the values ​​are within the valid range), and compatibility checks (such as whether a specific parameter combination is supported by the current hardware core). After verification, the module generates a structured failure reason report, such as "Unsupported H.265 Main10 Profile," "Frame width exceeds the maximum supported 4096 pixels," or "Bandwidth control parameter CBR (Constant Bitrate) conflicts with GOP structure." This report is returned to the upper-layer scheduler or logging system along with the abnormal event.

[0070] This design allows the execution entity to not only mark the current task flow as failed and potentially attempt other cores based on the specific failure reasons returned, but also provides direct error location information to the upstream task submitter or system administrator, significantly reducing troubleshooting time. Furthermore, the execution entity can learn from the types of failure reasons returned; for example, if a certain type of parameter error occurs frequently, it can provide parameter constraint hints to the upstream system, thereby reducing misconfiguration at the source and improving overall availability and maintainability.

[0071] Please continue to refer to further information. Figure 9 , Figure 9 The flowchart of a method for generating and repairing a seventh abnormal event provided in this embodiment of the disclosure is intended to address the case where the lifecycle state may also include a power-off state. This state is the next processing state that the target codec core should be in after not receiving any new codec task streams to be processed within a preset time period after successful initialization. The process 900 includes the following steps: Step 901: In response to the actual duration of the power-down state being greater than the preset duration limit corresponding to the power-down state, determine that there is a processing timeout in the power-down state, and generate a seventh abnormal event for the power-down state with a processing timeout. Building upon the aforementioned embodiments, this embodiment further introduces a power-down state as a new processing state within the finite machine model. Technically, this state represents the low-power sleep phase that the target codec core should enter according to energy-saving strategies after completing its general initialization (being in the codec core initialization state and confirmed as ready) and, within a continuously preset idle watch period, has not received any newly issued codec task streams. The core objective of this state is to safely and orderly shut down the core's hardware power or put it into deep sleep mode during idle periods to reduce overall system energy consumption. Managing power-down as a defined and monitorable state allows the system to independently and controllably supervise this potentially risky and sensitive operation of hardware power-down, preventing core freezes or resource unrecoverability due to abnormal power-down processes.

[0072] This step involves monitoring the time consumption of the power-down process itself. Once the core enters the power-down state, the state machine thread begins timing. The preset maximum duration of this state is set based on the maximum reasonable time required for normal hardware power-down or entering low-power mode. If the actual time exceeds this limit, it is considered a power-down timeout. This timeout usually indicates a blockage in the hardware power management process. Possible causes include: the hardware core has unfinished minor tasks before responding to the power-down command, a slow response from the Power Management Unit (PMU), or a timeout in the driver software while waiting for a power-down confirmation signal. The "seventh exception event" generated at this time is specifically used to identify the specific fault of "hardware power-down process stall."

[0073] Step 902: For the seventh abnormal event, the core hardware decoding power-off condition of the target codec is detected by the recovery thread, and the hardware decoding power-off condition is modified when it is not met.

[0074] This step provides a specific intervention strategy for this type of power management fault: detecting the hardware power-down conditions of the target codec core and modifying them when they are not met. In practice, this involves diagnostic checks of the preconditions required for hardware power-down. These conditions are typically defined by the driver and hardware specifications and may include: whether the core's clock gating is ready, whether all internal caches have been flushed, whether the power domains are isolated, and whether there are any pending interrupts or DMA (Direct Memory Access) operations. The repair module checks these conditions by reading a series of hardware status registers. If a condition is not met (e.g., a flag indicates that data transfer is still incomplete), the repair operation attempts to modify it. This might be done by sending a forced completion or reset command to the relevant submodule, clearing an erroneous pending status bit, or adjusting a power state switching threshold parameter. The goal is to manually satisfy or bypass the obstructing condition, allowing the standard power-down sequence to continue.

[0075] This repair mechanism demonstrates the ability to safely intervene in the low-level hardware state. After the execution conditions are modified, the aforementioned execution entity will retry the standard power-down procedure. Furthermore, the entire process can be recorded in detail. If a particular power-down condition repeatedly becomes an obstacle, this information can be fed back to hardware driver development or power management policy configuration to optimize the default power-down sequence or hardware design, thereby improving the system's power management reliability in the long term.

[0076] Please refer to Figure 10 , Figure 10 The flowchart of a method for generating and repairing an eighth abnormal event provided in this embodiment of the disclosure is intended to address the case where the lifecycle state may also include a power-on state. The power-on state is the next processing state that the target codec core, which is in the power-off state, should be in after receiving the codec task stream to be processed. The next processing state after the power-on state is the codec core initialization state. The process 1000 includes the following steps: Step 1001: In response to the actual duration of the power-on state being greater than the preset duration limit corresponding to the power-on state, determine that there is a processing timeout in the power-on state, and generate an eighth abnormal event for the power-on state with a processing timeout. Building upon the above embodiments, this embodiment further introduces a power-on state as a new processing state within the finite state machine model, representing the wake-up and activation phase of the hardware core from sleep to ready state. Technically, this state characterizes the hardware power-on, basic resource preparation, and operationalization process initiated by the target codec core in a low-power power-off state after receiving a new codec task flow instruction. This process typically includes, but is not limited to: issuing a power-on instruction to the core's power management unit, restoring clock signal supply, loading necessary boot firmware or microcode, and performing minimal hardware self-tests and register initializations to prepare the core for its subsequent general codec core initialization state. By explicitly defining this power-on process as an independent and monitorable state, the system can independently supervise the critical and potentially risky wake-up of the hardware module, ensuring that the core can reliably recover from deep sleep to a working state.

[0077] This step involves monitoring the duration of the hardware power-on process. When the scheduling system decides to wake up a core that is currently powered down to execute a new task, the core state switches to the power-on state, and the state machine thread starts timing. The preset maximum duration of this state is based on the maximum time required for power-on reset, stable power supply, and basic initialization as specified in the hardware specifications. If the actual time exceeds this threshold, it is considered a power-on timeout. This timeout usually indicates that the hardware-level wake-up process is blocked. Possible causes include: power rails failing to power on normally, clock source failure, corrupted or failed boot firmware, or hardware self-test program errors and stalling. The eighth exception event generated at this time is specifically used to identify this serious hardware power-on failure, indicating a potential problem with the basic hardware functionality.

[0078] Step 1002: For the eighth exception event, restart the hardware module of the target codec core by restoring the thread.

[0079] This step provides the most fundamental recovery strategy for this type of low-level hardware failure: restarting the target codec core's hardware module. In practice, this represents a complete, power-on reboot of the hardware module. This is typically not a simple software restart, but rather a complete power-down / power-on cycle executed across the entire power domain or functional module containing the target codec core, or triggering its hard reset pin, via a system-level power management controller or hardware reset circuitry. The goal is to clear any hardware latches, temporary electrical faults, or firmware execution malfunctions that could have stalled the power-on process by completely cutting off and resupplying power, or by forcibly resetting all internal logic, creating an opportunity for a clean, new power-on sequence.

[0080] Given that hardware reboot is a high-impact operation, it is typically used only after other software-level recovery attempts have failed. Before triggering the reboot, the execution entity can attempt to record the states of relevant power and reset control registers for subsequent fault analysis. Furthermore, the reboot operation itself should have timeout protection. If the core still cannot enter the encoding / decoding core initialization state within a reasonable time after reboot, it may be marked as permanently faulty and reported, while the scheduling system isolates the core from the available resource pool. The aim of this mechanism is to ensure that even if a failure occurs at the lowest level of the hardware wake-up process, strong intervention measures can be used to attempt recovery and make clear fault isolation decisions, thereby maintaining the availability of the remaining parts.

[0081] Based on the above embodiments, a self-learning and adaptive optimization step can also be added to dynamically calibrate the time base used for anomaly detection (i.e., the preset duration upper limit mentioned in the above embodiments). This step can include two consecutive operations: first, calculating the corresponding average processing time for each processing state within a preset period; second, updating the corresponding preset duration upper limit using the average processing time corresponding to each processing state.

[0082] From a technical perspective, this solution aims to address the inherent limitations of statically preset thresholds. In actual system operation, the complexity of different task flows, system load, and even the performance status of the hardware itself (such as temperature and aging) can cause the actual completion time of the same processing state (e.g., "single-frame encoding / decoding execution state") to fluctuate within a reasonable range. By periodically collecting historical data and calculating the average processing time, the aforementioned execution entity can capture this "baseline" of actual performance, allowing the monitoring standard to dynamically evolve with the operating environment. This avoids missed anomalies due to overly lenient fixed thresholds, or frequent false alarms due to overly strict thresholds. In practical terms, this preset period can be a fixed time window (e.g., the past 24 hours) or a window based on the number of processed tasks (e.g., the most recently completed 1000 frames). The statistical process can be executed by the monitoring module or a separate performance analysis thread. For each defined processing state, the aforementioned execution entity collects samples of the actual processing time spent each time entering and successfully leaving that state within that period. Then, it calculates a representative "average processing time" using algorithms such as arithmetic mean, moving average, or the mean after removing outliers. After obtaining the new average, the execution entity can update the corresponding preset duration limit. The update strategy is not a simple replacement; it typically involves a weighted calculation between the original preset value and the newly calculated average, or adding a safety margin of a standard deviation multiple to the average to form a new threshold that better reflects the current actual performance. For example, the new preset duration limit might be set as: average processing time × (1 + dynamic coefficient) + fixed margin. This coefficient can be configured according to the tolerance for false positives and false negatives in the actual application scenario.

[0083] Furthermore, independent statistical models and threshold sets can be pre-established and maintained for tasks of different complexity levels (such as 1080p encoding and 4K encoding). In addition, the update process can incorporate gradual adjustment constraints to ensure that the magnitude of a single update is not too large, avoiding drastic threshold fluctuations due to individual extreme tasks. Simultaneously, the aforementioned execution entity can record the historical adjustment trajectory of the thresholds. When the preset duration of a certain state continues to abnormally increase, it may indicate hardware performance degradation or excessive software load, thus triggering a preventative maintenance warning. This self-optimizing mechanism transforms the entire monitoring system from a static set of rules into an intelligent system capable of adapting to actual workloads and possessing continuous learning capabilities, significantly improving the accuracy of anomaly detection and the robustness and efficiency of the entire resource management system.

[0084] Building upon the core state monitoring and anomaly repair solutions provided in the above embodiments, more refined predictive resource management and power consumption control strategies can be introduced. For example, in addition to the existing timed monitoring and recovery threads, a machine learning module can be added. This module, in principle, performs in-depth analysis of the historical state machine data accumulated over long-term operation of each hardware codec core to learn its workload patterns, task arrival intervals, and statistical patterns of state transitions. Specifically, this module can utilize time series analysis models or regression models to predict the probability of task suspension or the arrival time of new tasks in the near future based on historical data (such as task flow startup frequency, duration distribution of stream idle state, processing time of decoding / encoding state, etc.). In practical terms, this machine learning module can run as an independent background service, periodically obtaining state machine log data from all cores from the global monitoring module and performing model inference. When it is predicted that a core about to be woken up from idle is highly likely to receive a new task soon, this module can suggest to the scheduler that the scheduler perform a "pre-wake-up" operation in advance, that is, initialize the core in the coreidle or power-off state to the ready state in advance. This predictive pre-wake-up effectively hides the latency of hardware power-on and initialization, thereby improving the system's response speed to sudden tasks and achieving an optimal balance between resource availability and response latency. From a logical expansion perspective, this module can also learn online based on prediction accuracy, dynamically adjust prediction strategies, and combine with load balancing strategies to achieve more precise resource scheduling.

[0085] Furthermore, based on the above embodiment's solution of directly transitioning to a power-down state when a timeout occurs during the temporary idle waiting state, a more gradual power consumption control method can be attempted: when the duration of the temporary idle waiting state exceeds the first-level threshold (this threshold may be shorter than the threshold for triggering power-down), the power-down process is not immediately triggered. Instead, the target codec core's operating mode is switched to a low-frequency mode using hardware-supported dynamic voltage and frequency adjustment technology. Technically, this operation significantly reduces the core's static and dynamic power consumption by lowering its operating voltage and clock frequency, while maintaining the core's basic circuitry in a powered and responsive state. In practice, the driver can call the power management interface to switch the core's clock source to a low-frequency clock domain and adjust the supply voltage accordingly. When the task flow associated with the core resumes data transmission, the execution entity can quickly raise the core frequency and voltage to normal operating levels in a time much shorter than required to fully power on and wake up from the power-off state, thus continuing task processing. This solution achieves a better trade-off between power saving and response speed, and is more suitable for application scenarios where task pause times are uncertain but may be short. From a logically feasible expansion perspective, the system can be configured with multiple power consumption states, dynamically switching between different performance-power consumption levels based on the duration of the temporary idle waiting state. It can even incorporate historical pause patterns of the task flow to adaptively adjust thresholds at each level, achieving personalized and optimal energy efficiency management.

[0086] To enhance understanding, this disclosure also provides a specific implementation scheme based on a particular application scenario. Please refer to the example below. Figure 11a and Figure 11b : This embodiment aims to establish a robust and powerful video codec driver that can self-monitor its driver, software, and hardware status during operation. When an anomaly occurs, the system can automatically check, reset, and restart the relevant hardware and software units to ensure the long-term, stable operation of the encoding and decoding services.

[0087] The complete technical solution is described in detail below with reference to the accompanying drawings: like Figure 11aAs shown, for hardware systems (such as GPUs) containing multiple codec cores (e.g., multiple decoding cores and multiple encoding cores), this embodiment constructs a hierarchical management and monitoring architecture. The core driver of the entire system creates and maintains a dedicated worker thread for each independent hardware codec core (including decoding thread 1 corresponding to decoder core 1, decoding thread N corresponding to decoder core N, encoding thread 1 corresponding to encoder core 1, and encoding thread M corresponding to encoder core M). Decoding thread 1 receives task flow 1 and task flow 2 for video playback, decoding thread 2 receives task flow for video decoding, encoding thread 1 receives task flow 1 and task flow 2 for video encoding, and encoding thread M receives task flow similarly for video encoding. This thread is not only responsible for driving the hardware unit and handling communication with upper-layer services, but also maintains a finite state machine (i.e., each codec thread has its own dedicated finite state machine) to accurately characterize the lifecycle and workflow of the core. Simultaneously, each core has its own independent task queue for receiving and scheduling multiple codec task flows assigned to it. At the system level, a global timed monitoring and recovery thread is set up. This thread polls the state machine information maintained by all core threads at fixed intervals to achieve centralized monitoring of the overall system health and anomaly recovery.

[0088] The workflow of a single codec core can be abstracted as a finite state machine containing eight distinct states, with the state transition relationships as follows: Figure 11b As shown. The initial state is power off, indicating that the hardware is in a power-off sleep state. When the driver receives a codec task request, it triggers the hardware power-on process, and the state changes to power on. This process includes loading firmware, enabling the clock, and other operations. After the hardware completes basic self-tests and initialization, it enters coreidle (codec core initialization state), indicating that the core is ready and can receive tasks at any time. When there is a specific codec task flow to process, the core enters decode / encode init (codec parameter initialization state). This stage is responsible for initializing the codec according to the task flow parameters (such as encoding format and resolution). After the codec initialization is successful, the codec core immediately begins to process the first frame of data, and the state changes to decoding / encoding (single frame codec execution state). This state indicates that the core is performing actual encoding or decoding operations on the frame data. When a frame of data is processed, the core state changes to decode / encode done (single frame codec completion state), and the driver thread returns the processing result to the upper layer application.

[0089] Subsequently, different paths exist depending on the business logic. If the current frame is the last frame (tail frame) of the task stream, the state will transition to close decode / encode (task stream end process). In this state, the driver releases all resources occupied by the task stream, and then the core returns to the core idle state. If the current frame is not the last frame (non-tail frame), the core enters stream idle (temporarily idle waiting state). This is a legal waiting state used to handle normal intervals between data streams (such as playback buffering) or active pauses by the business. When the next frame of data arrives, the state will switch directly from stream idle back to decode / encode to continue processing. In addition, if the core is idle in the core idle state for more than a preset time, or if all task streams are in the stream idle state for more than a set time, the core will automatically power down to save power, and the state will transition from coreidle to power off. Figure 11b The dashed arrows clearly indicate the path from stream idle to core idle when a timeout occurs, distinguishing between service interruption and abnormal deadlock.

[0090] Throughout its lifecycle, each core state machine thread meticulously records the timestamps, context information (such as task flow ID, frame sequence number, and encoding / decoding parameters), and performance data (such as single-frame processing time and average frame rate) for each state transition. Figure 11a As shown, the global timed monitoring and recovery thread achieves fine-grained monitoring by periodically querying this rich state machine data. It can not only detect whether the core is "alive," but also determine whether it is stuck in a specific state (such as decoding / encoding), whether task processing has timed out, or whether the queue is blocked.

[0091] When the monitoring thread detects any anomaly, it triggers a differentiated recovery strategy based on the specific state of the core at the time of the anomaly, rather than performing a uniform "restart" operation. For example: for a power-on anomaly, it attempts to restart the hardware module; for a decode / encode init anomaly, it rolls back the faulty configuration and retryes initialization; for a decoding / encoding anomaly, it may only reset the current abnormal task flow rather than the entire core; for a stream idle anomaly, it performs deadlock detection and unlocking reset. This highly targeted recovery mechanism minimizes the impact on normal business operations and achieves fast and accurate fault self-healing.

[0092] In summary, this embodiment constructs a complete closed-loop system from state awareness and anomaly diagnosis to differentiated recovery by establishing a refined finite state machine model for each hardware codec core and combining it with independent task queues and global intelligent monitoring. This solution can significantly improve the utilization rate of multi-core hardware resources and system throughput, shorten anomaly recovery time, achieve intelligent power consumption management, and has good scalability, thereby greatly enhancing the reliability and operational stability of the system in complex high-concurrency video processing scenarios.

[0093] Further reference Figure 12 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a processing apparatus for encoding and decoding tasks, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0094] like Figure 12 As shown, the encoding and decoding task processing device 1200 of this embodiment may include: a task stream receiving unit 1201, a processing status monitoring unit 1202, an abnormal event generation unit 1203, and a targeted repair unit 1204. The system includes a task flow receiving unit 1201, configured to receive a stream of encoding / decoding tasks to be processed; a processing status monitoring unit 1202, configured to record the processing status of the target encoding / decoding core on the encoding / decoding task flow through a state machine thread; wherein the processing status includes the lifecycle status of the hardware unit and the encoding / decoding workflow status, the encoding / decoding workflow status includes: single-frame encoding / decoding execution status, single-frame encoding / decoding completion status, and temporary idle waiting status, the temporary idle waiting status being the next processing status after the single-frame encoding / decoding completion status has completed the encoding / decoding operation on the current non-tail frame; an exception event generation unit 1203, configured to generate corresponding exception events for the target processing status where there is a processing timeout, based on the processing status recorded by the state machine thread; and a targeted repair unit 1204, configured to perform exception repair on the exception events through a recovery thread according to the repair method corresponding to the target processing status.

[0095] In this embodiment, the specific processing and technical effects of the following components in the encoding / decoding task processing device 1200—namely, the task stream receiving unit 1201, the processing status monitoring unit 1202, the abnormal event generation unit 1203, and the targeted repair unit 1204—can be found by referring to [reference needed]. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.

[0096] In some other optional implementations of this embodiment, the exception event generation unit 1203 is further configured to: In response to the actual duration of the single-frame codec execution state being greater than the preset duration limit corresponding to the single-frame codec execution state, it is determined that there is a processing timeout in the single-frame codec execution state, and a first abnormal event is generated for the single-frame codec execution state with a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: The first abnormal event is repaired by employing at least one of the following first repair methods through the recovery thread: Initialize the target codec core, replace the memory used, and re-encode and decode the current frame to be encoded and decoded.

[0097] In some other optional implementations of this embodiment, the exception event generation unit 1203 is further configured to: In response to the fact that the actual duration of the single frame encoding and decoding completion state is greater than the preset duration limit corresponding to the single frame encoding and decoding completion state, it is determined that there is a processing timeout in the single frame encoding and decoding completion state, and a second abnormal event is generated for the single frame encoding and decoding completion state with a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: The second abnormal event is repaired by employing at least one of the following second repair methods through the recovery thread: Resend the encoding / decoding completion notification corresponding to the current frame to the state machine thread, and re-encode / decode the current frame.

[0098] In some other optional implementations of this embodiment, the exception event generation unit 1203 is further configured to: In response to the actual duration of the temporary idle waiting state being greater than the preset duration limit corresponding to the temporary idle waiting state, it is determined that there is a processing timeout in the temporary idle waiting state, and a third abnormal event is generated for the temporary idle waiting state that has a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: In response to the third abnormal event, a deadlock detection is performed on the encoding and decoding task flow through a recovery thread, and the detected deadlock is unlocked and reset.

[0099] In some other optional implementations of this embodiment, the encoding / decoding workflow state further includes: a task flow end state, which is the next processing state after the encoding / decoding operation of the tail frame is completed. The exception event generation unit 1203 is further configured to: In response to the actual duration of the task flow end state being greater than the preset duration limit corresponding to the task flow end state, it is determined that there is a processing timeout in the task flow end state, and a fourth abnormal event is generated for the task flow end state with a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: The fourth abnormal event is repaired by employing at least one of the following third repair methods through the recovery thread: Initialize the task parameters of the encoding / decoding task flow and call the preset hardware decoding repair tool.

[0100] In some other optional implementations of this embodiment, the lifecycle states include: the codec core initialization state, which is a task flow end state indicating the next processing state after the processing of the codec task flow has ended. The exception event generation unit 1203 is further configured to: In response to the fact that the actual duration of the codec core initialization state is greater than the preset duration limit corresponding to the codec core initialization state, it is determined that there is a processing timeout in the codec core initialization state, and a fifth exception event is generated for the codec core initialization state with a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: In response to the fifth abnormal event, the recovery thread performs lock detection on the worker threads of the encoding and decoding task flow and unlocks the detected locks.

[0101] In some other optional implementations of this embodiment, the codec workflow state further includes: a codec parameter initialization state, which is the next processing state that the target codec core of the codec task flow that still has pending processing should be in after successfully completing core initialization through the codec core initialization state. The next processing state after the codec parameter initialization state is the single-frame codec execution state. The exception event generation unit 1203 is further configured to: In response to the actual duration of the codec parameter initialization state being greater than the preset duration limit corresponding to the codec parameter initialization state, it is determined that there is a processing timeout in the codec parameter initialization state, and a sixth exception event is generated for the codec parameter initialization state with a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: For the sixth exception event, the validity of the input setting parameters is validated by the recovery thread, and the reason for the validity validation failure is returned.

[0102] In some other optional implementations of this embodiment, the lifecycle state also includes: a power-off state, which is the next processing state that the target codec core should be in after receiving a new codec task stream within a preset time period after successful initialization. The exception event generation unit 1203 is further configured to: If the actual duration of the power-down state is greater than the preset duration limit corresponding to the power-down state, it is determined that there is a processing timeout in the power-down state, and a seventh abnormal event is generated for the power-down state with a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: For the seventh exception event, the core hardware decoding power-off condition of the target codec is detected by the recovery thread, and the hardware decoding power-off condition is modified when it is not met.

[0103] In some other optional implementations of this embodiment, the lifecycle state further includes: a power-on state, which is the next processing state that the target codec core, which is in the power-off state, should be in after receiving the codec task stream to be processed. The next processing state after the power-on state is the codec core initialization state. The exception event generation unit 1203 is further configured to: If the actual duration of the power-on state is greater than the preset duration limit corresponding to the power-on state, it is determined that there is a processing timeout in the power-on state, and an eighth abnormal event is generated for the power-on state with a processing timeout. Correspondingly, the targeted repair unit 1204 is further configured as follows: In response to the eighth abnormal event, the hardware module of the target codec core is restarted by the recovery thread.

[0104] In some other optional implementations of this embodiment, the encoding / decoding task processing device 1200 may further include: The average processing time statistics unit is configured to calculate the corresponding average processing time for each processing state within a preset period. The upper limit update unit is configured to update the corresponding preset duration upper limit using the average processing time corresponding to each processing state.

[0105] This embodiment is a device embodiment corresponding to the above method embodiment. The encoding and decoding task processing device provided in this embodiment records the processing state of the corresponding encoding and decoding core when executing the encoding and decoding task flow through a state machine thread. The processing state includes a lifecycle state and an encoding and decoding workflow state. The encoding and decoding workflow state includes at least a single frame encoding and decoding execution state, a completion state, and a temporary idle waiting state. This state monitoring mechanism enables fine-grained modeling and tracking of the processing progress of the encoding and decoding task flow. In particular, by explicitly defining the "temporary idle waiting state" as a legal state after completing the encoding and decoding operation for non-tail frames, it is possible to accurately distinguish between normal intermittent pauses and true execution deadlock anomalies during the completion of the encoding and decoding task flow. Based on this, by comparing the actual duration of each state with the corresponding preset threshold, timeout anomalies occurring in specific processing stages can be reliably identified. Then, the recovery thread can trigger a matching repair operation based on the actual processing state corresponding to the anomaly, achieving accurate anomaly location and differentiated recovery. By applying this solution, erroneous interference with normal business processes can be effectively avoided, and the ineffective overhead and business interruption caused by a single recovery strategy can be significantly reduced, thereby improving the reliability of encoding and decoding task processing, resource utilization efficiency, and the targeted handling of anomalies.

[0106] According to embodiments of this disclosure, this disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to implement the encoding / decoding task processing method described in any of the above embodiments.

[0107] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to execute the processing method for the encoding / decoding task described in any of the above embodiments.

[0108] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the encoding / decoding task processing method described in any of the above embodiments.

[0109] Figure 13A schematic block diagram of an electronic device 1300, exemplified by embodiments of the present disclosure, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0110] like Figure 13 As shown, the electronic device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. Various programs and data can also be stored in the RAM 1303 for on-demand operation of the electronic device 1300. The computing unit 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0111] Multiple components in electronic device 1300 are connected to I / O interface 1305, including: input unit 1306, such as keyboard, mouse, etc.; output unit 1307, such as various types of displays, speakers, etc.; storage unit 1308, such as disk, optical disk, etc.; and communication unit 1309, such as network card, modem, wireless transceiver, etc. Communication unit 1309 allows device 1300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0112] The computing unit 1301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 performs the various methods and processes described above, such as methods for processing encoding and decoding tasks. For example, in some embodiments, methods for processing encoding and decoding tasks may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by the computing unit 1301, one or more steps of the methods for processing encoding and decoding tasks described above may be performed. Alternatively, in other embodiments, computing unit 1301 may be configured by any other suitable means (e.g., by means of firmware) to perform processing methods for encoding and decoding tasks.

[0113] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0114] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0115] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0117] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0118] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0119] According to the technical solution of this disclosure, a state monitoring mechanism that presets at least a single-frame encoding / decoding execution state, a completion state, and a temporary idle waiting state enables refined modeling and tracking of the processing progress of the encoding / decoding task flow. In particular, by explicitly defining the "temporary idle waiting state" as a legitimate state after completing non-tail frame encoding / decoding operations, it is possible to accurately distinguish between normal intermittent pauses and genuine execution deadlock anomalies during the completion of the encoding / decoding task flow. Based on this, by comparing the actual duration of each state with the corresponding preset threshold, timeout anomalies occurring in specific processing stages can be reliably identified. Then, based on the actual processing state corresponding to the anomaly, a matching repair operation is triggered, achieving precise anomaly localization and differentiated recovery. Applying this solution effectively avoids erroneous interference with normal business processes, significantly reduces the ineffective overhead and business interruption scope caused by a single recovery strategy, thereby improving the reliability of encoding / decoding task processing, resource utilization efficiency, and the targeted nature of anomaly handling.

[0120] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0121] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for processing encoding and decoding tasks, characterized in that, include: Receive the stream of encoding / decoding tasks to be processed; The state machine thread records the processing state of the target codec core on the codec task flow; wherein, the processing state includes the life cycle state of the hardware unit and the codec workflow state, the codec workflow state includes: single frame codec execution state, single frame codec completion state and temporary idle waiting state, the temporary idle waiting state is the next processing state that should be entered after the single frame codec completion state has completed the codec operation on the current non-tail frame; In the processing states recorded by the state machine thread, corresponding exception events are generated for target processing states where processing timeouts occur; The abnormal event is repaired by restoring the thread according to the repair method corresponding to the target processing state.

2. The method according to claim 1, characterized in that, The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the fact that the actual duration of the single-frame codec execution state is greater than the preset duration upper limit corresponding to the single-frame codec execution state, it is determined that the single-frame codec execution state has the processing timeout situation, and a first abnormal event is generated for the single-frame codec execution state with the processing timeout situation; Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: The recovery thread performs anomaly repair on the first abnormal event using at least one of the following first repair methods: The target codec core is initialized, the memory used is replaced, and the current frame to be encoded and decoded is re-encoded and decoded.

3. The method according to claim 1, characterized in that, The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the fact that the actual duration of the single-frame encoding and decoding completion state is greater than the preset duration upper limit corresponding to the single-frame encoding and decoding completion state, it is determined that the single-frame encoding and decoding completion state has a processing timeout, and a second abnormal event is generated for the single-frame encoding and decoding completion state with the processing timeout. Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: The recovery thread performs anomaly repair on the second abnormal event using at least one of the following second repair methods: The encoding / decoding completion notification corresponding to the current frame is resent to the state machine thread, and the current frame is re-encoded and decoded.

4. The method according to claim 1, characterized in that, The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the actual duration of the temporary idle waiting state being greater than the preset duration limit corresponding to the temporary idle waiting state, it is determined that the temporary idle waiting state has a processing timeout, and a third abnormal event is generated for the temporary idle waiting state with the processing timeout. Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: In response to the third abnormal event, the recovery thread performs deadlock detection on the encoding / decoding task stream and performs unlocking and reset processing on the detected deadlock.

5. The method according to claim 1, characterized in that, The encoding / decoding workflow state also includes a task flow end state, which is the next processing state that should be entered after the encoding / decoding operation of the tail frame is completed, as indicated by the single-frame encoding / decoding completion state. The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the actual duration of the task flow end state being greater than the preset duration upper limit corresponding to the task flow end state, it is determined that the task flow end state has a processing timeout, and a fourth abnormal event is generated for the task flow end state with the processing timeout. Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: The recovery thread performs anomaly repair on the fourth abnormal event using at least one of the following third repair methods: Initialize the task parameters of the encoding / decoding task stream and call the preset hardware decoding repair tool.

6. The method according to claim 5, characterized in that, The lifecycle states include: the codec core initialization state, which is the next processing state that should be entered after the task flow end state indicates the end of processing the codec task flow. The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the fact that the actual duration of the codec core initialization state is greater than the preset duration upper limit corresponding to the codec core initialization state, it is determined that the codec core initialization state has a processing timeout, and a fifth abnormal event is generated for the codec core initialization state with the processing timeout. Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: In response to the fifth abnormal event, the recovery thread performs lock detection on the worker thread of the encoding / decoding task flow and unlocks the detected lock.

7. The method according to claim 6, characterized in that, The encoding / decoding workflow state also includes: a codec parameter initialization state, which is the next processing state that the target codec core of the encoding / decoding task flow that still has pending processing should be in after successfully completing core initialization through the codec core initialization state. The next processing state after the codec parameter initialization state is the single-frame encoding / decoding execution state. The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the fact that the actual duration of the codec parameter initialization state is greater than the preset duration upper limit corresponding to the codec parameter initialization state, it is determined that there is a processing timeout in the codec parameter initialization state, and a sixth abnormal event is generated for the codec parameter initialization state with the processing timeout. Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: For the sixth abnormal event, the recovery thread verifies the validity of the input setting parameters and returns the reason for the failure of the validity verification.

8. The method according to claim 7, characterized in that, The lifecycle state also includes a power-down state, which is the next processing state that the target codec core should be in after it has not received any new codec task streams to be processed within a preset time period after successful initialization. The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the actual duration of the power-down state being greater than the preset duration limit corresponding to the power-down state, it is determined that the power-down state has a processing timeout, and a seventh abnormal event is generated for the power-down state with the processing timeout. Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: In response to the seventh abnormal event, the recovery thread detects the hardware decoding power-off condition of the target codec core, and modifies the hardware decoding power-off condition if it is not met.

9. The method according to claim 8, characterized in that, The lifecycle states also include: a power-on state, which is the next processing state that the target codec core, which is in the power-off state, should be in after receiving the codec task stream to be processed. The next processing state after the power-on state is the codec core initialization state. The step of generating corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread, includes: In response to the actual duration of the power-on state being greater than the preset duration limit corresponding to the power-on state, it is determined that the power-on state has a processing timeout, and an eighth abnormal event is generated for the power-on state with the processing timeout. Correspondingly, the abnormal event is repaired by the recovery thread according to the repair method corresponding to the target processing state, including: In response to the eighth abnormal event, the hardware module of the target codec core is restarted through the recovery thread.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: The average processing time for each of the aforementioned processing states is statistically calculated within a preset period. The corresponding preset duration upper limit is updated using the average processing time corresponding to each of the processing states.

11. A processing apparatus for encoding and decoding tasks, characterized in that, include: The task stream receiving unit is configured to receive the codec task stream to be processed; The processing status monitoring unit is configured to record the processing status of the target codec core on the codec task flow through a state machine thread; wherein, the processing status includes the life cycle status of the hardware unit and the codec workflow status, the codec workflow status includes: single frame codec execution status, single frame codec completion status and temporary idle waiting status, the temporary idle waiting status is the next processing status that should be in after the single frame codec completion status has completed the codec operation on the current non-tail frame; The exception event generation unit is configured to generate corresponding exception events for target processing states where processing timeouts occur, based on the processing states recorded by the state machine thread. The targeted repair unit is configured to perform abnormal repair on the abnormal event through a recovery thread in a repair method corresponding to the target processing state.

12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the encoding / decoding task processing method according to any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the processing method for the encoding and decoding task according to any one of claims 1-10.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the processing method for the encoding / decoding task according to any one of claims 1-10.

Citation Information

Patent Citations

  • Image decoding heterogeneous acceleration system architecture based on CPU + FPGA

    CN121691713A

  • Permitting unaborted processing of transaction after exception mask update instruction

    EP3462312A1