Artificial intelligence based in-loop enhancement method for ultra-low bitrate video coding

By employing an in-loop enhancement method for video encoding and decoding based on spiking neural networks, the problems of visual distortion and structural integrity in video encoding and decoding technology at extremely low bit rates are solved. This method restores video structural boundaries and texture details, thereby improving video quality and temporal consistency.

CN121509669BActive Publication Date: 2026-07-24GUOSEN RUIAN (BEIJING) INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUOSEN RUIAN (BEIJING) INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-12-09
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to maintain the subjective quality and structural integrity of videos at extremely low bitrates. Traditional methods result in visual distortion, blockiness, texture blurring, and loss of detail. Furthermore, convolutional neural network-based enhancement methods are difficult to deploy in mobile terminals and real-time encoding environments.

Method used

An in-loop enhancement method based on spiking neural networks is adopted for video encoding and decoding. The reconstructed frame is transformed into edge spiking layer, texture spiking layer and block boundary spiking layer through spiking coding layer. Spiking events are issued on short time scales within the frame and long time scales across frames. Combined with interleaved phase review and asynchronous residual back-injection mechanism, an enhanced spiking event stream is generated to replace the reference frame.

Benefits of technology

It preserves the restoration of video structural boundaries and texture details at ultra-low bitrates, achieves temporal consistency and dynamic bitrate matching, reduces the risk of cross-frame distortion accumulation, and improves subjective viewing quality and the reliability of reference frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509669B_ABST
    Figure CN121509669B_ABST
Patent Text Reader

Abstract

The application discloses an in-loop enhancement method for super-low code rate video coding based on artificial intelligence, and relates to the technical field of artificial intelligence, which comprises the following steps: step 1: obtaining a reconstructed frame generated by a current decoder as a visual input and a current code rate indicator as a control input; generating an initial pulse event stream containing event timestamps and polarity and a pulse timing consistency graph thereof; step 2: under the constraint of the pulse timing consistency graph, performing pulse consistency re-estimation and asynchronous residual back-annotation based on interleaved phase cross-checking; step 3: under the control of in-loop closed-loop and code rate collaborative regulation, generating a pulse reconstruction enhancement frame by using the enhanced pulse event stream, and replacing a reference frame with the pulse reconstruction enhancement frame. The application can significantly reduce the cross-frame distortion accumulation risk, improve the subjective viewing quality and the reliability of the reference frame, and has practical application value in low-bandwidth transmission and limited storage scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an in-loop enhancement method for ultra-low bitrate video encoding and decoding based on artificial intelligence. Background Technology

[0002] Video encoding and decoding technology is a core component of modern multimedia communication and storage systems, and its development directly determines the availability and quality of video content under different bandwidth and storage conditions. With the rapid development of ultra-high-definition video, virtual reality, immersive interaction, and streaming media services, maintaining the subjective quality and structural integrity of video at extremely low bitrates has become a major challenge for the industry. Traditional video compression standards such as H.264 / AVC, H.265 / HEVC, and the next-generation H.266 / VVC all rely on classic frameworks such as prediction, transform, quantization, and entropy coding. They reduce redundant information by utilizing pixel correlation in the spatial and temporal domains, thereby reducing the bitstream size. However, when the bitrate drops to extremely low levels, these traditional methods often lead to severe visual distortion, manifesting as blockiness, ringing artifacts, texture blurring, and loss of detail, making it difficult to meet the quality requirements of practical applications.

[0003] In recent years, academia and industry have begun to introduce AI-based loop enhancement techniques, deploying deep learning models in the codec loop as filtering or enhancement modules to compensate for distortion and restore details under low bitrate conditions. For example, some techniques utilize convolutional neural networks (CNNs) to achieve deblocking and super-resolution on the decoder side, while others employ generative adversarial networks (GANs) to restore texture within the decoder loop. These methods do significantly improve subjective quality compared to traditional filters, especially in terms of texture detail and edge structure restoration. However, because CNNs generally rely on floating-point computation and have high hardware resource requirements, these methods face significant challenges in deployment on mobile terminals, embedded systems, and real-time coding environments. Furthermore, traditional neural network enhancement methods typically use frames as the basic processing unit, lacking strict constraints on temporal consistency, often resulting in inter-frame flickering and detail drift. In existing research, spiking neural networks (SNNs) are gradually gaining attention as a bio-inspired computational model. Their core idea is to use sparse temporal representations of pulse events to transmit information in a manner similar to the firing mechanism of neurons. Compared to traditional CNNs, spiking neural networks have advantages in event-driven operation and energy consumption, and can run efficiently on low-power hardware platforms. Meanwhile, spiking neural networks inherently possess time series modeling capabilities, enabling them to encode temporal information simultaneously on both short and long timescales. However, existing publicly available research on video enhancement based on spiking neural networks is still in the exploratory stage, with most results focusing on areas such as event camera data processing, low-power visual recognition, and sparse pattern recognition. Practical applications within video encoding / decoding loops are not yet mature. Summary of the Invention

[0004] The purpose of this invention is to provide an AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding. This method can prioritize the preservation of video structural boundaries and controllably restore texture details under ultra-low bitrate conditions, while achieving temporal consistency and dynamic bitrate matching. This significantly reduces the risk of cross-frame distortion accumulation, improves subjective viewing quality and the reliability of reference frames, and has practical application value in low-bandwidth transmission and limited storage scenarios.

[0005] To address the aforementioned technical problems, this invention provides an artificial intelligence-based in-loop enhancement method for ultra-low bitrate video encoding and decoding, the method comprising: Step 1: Obtain the reconstructed frame generated by the current decoder as visual input, and obtain the current bitrate indicator as control input; transform the reconstructed frame into three types of pulse event streams through the pulse coding layer of the spiking neural network, namely edge pulse layer, texture pulse layer and block boundary pulse layer; the pulse coding layer adopts leakage integration and trigger firing mechanism to fire pulse events simultaneously on two time axes: short time scale within the frame and long time scale across the frame, and sets the upper limit of firing density and reset hysteresis range according to the bitrate indicator, thereby generating an initial pulse event stream containing event timestamps and polarities and its pulse timing consistency map; Step 2: Under the constraints of the pulse timing consistency map, pulse consistency re-estimation and asynchronous residual back-injection based on interleaved phase kernel are performed. This process specifically includes: interleaving the current initial pulse event stream and the reference pulse event stream on short time scales within the frame and long time scales across frames to form an interleaved phase window; performing threshold voting and polarity determination on pulse clusters in each window to suppress isolated pulses and confirm the dominant firing timing; performing peak offset compensation on pulse residuals formed by unmatched pulses according to adjacent time slices, and splitting them through structure priority branches and texture priority branches, and then merging them into an enhanced pulse event stream according to pulse timing consistency weights; at the same time, performing threshold drift self-correction based on the inter-frame dominant firing timing difference to maintain the consistency and stability of long time scales, thereby forming an enhanced pulse event stream that can enhance structure and refine texture at ultra-low bit rates. Step 3: Under the control of closed-loop and code rate coordinated adjustment within the loop, the enhanced pulse event stream is used to generate a pulse reconstruction enhanced frame, and the pulse reconstruction enhanced frame is used to replace the reference frame.

[0006] Furthermore, in step 1, the reconstructed frames are first uniformly converted into single-channel images of the luminance component, and linearly stretched to a fixed integer range according to the pixel value range while maintaining the spatial resolution; the image width and height are recorded to determine the processing grid; two time axes are established: a short time scale within the frame and a long time scale across the frame. The short time scale within the frame is divided into several equally spaced discrete scales within a frame period, and the number of these discrete scales is fixed at 1000; the long time scale across the frame spans the cumulative window of consecutive frames, and its window length is fixed at 8 frames and a sliding update strategy is adopted; the bitrate indicator is normalized to a discrete level from 0 to 100.

[0007] Furthermore, in step 1, the rules for setting the upper limit of the release density and the reset hysteresis range based on the bitrate indicator are as follows: if the bitrate indicator is in the range of 0 to 33, the upper limit of the release density is set to no more than 100,000 pulse events per megapixel per frame, and the reset hysteresis range is 6 to 10 discrete scales; if the bitrate indicator is in the range of 34 to 66, the upper limit of the release density is set to no more than 250,000 pulse events per megapixel per frame, and the reset hysteresis range is 4 to 6 discrete scales; if the bitrate indicator is in the range of 67 to 100, the upper limit of the release density is set to no more than 500,000 pulse events per megapixel per frame, and the reset hysteresis range is 2 to 4 discrete scales.

[0008] Furthermore, in step 1, within a long time-stamped window spanning multiple frames, the consistency of the pulse time series for each pixel is evaluated as follows: the pixel is clustered based on the major firing scales of the last 8 frames. When a cluster appears in at least 5 frames and the difference between each firing scale does not exceed 3 discrete scales, the pixel is determined to have high consistency, and the major firing scale corresponding to the cluster is recorded as the dominant firing time sequence of the pixel. If a cluster appears in 3 or 4 frames and the difference between each firing scale does not exceed 5 discrete scales, it is determined to have medium consistency. All other cases are determined to have low consistency. The aforementioned consistency levels are encoded as integers in a raster with the same resolution as the image to form a consistency level map, and the dominant firing time sequence is encoded as integer values ​​of short time-stamped scales to form a time sequence index map. Finally, the consistency level map and the time sequence index map together constitute a pulse time sequence consistency map.

[0009] Furthermore, in step 2, the process of forming the interleaved phase window includes: interleaving the current initial pulse event stream and the reference pulse event stream pixel-by-pixel on intra-frame short timescales and cross-frame long timescales; using the timing index map in the pulse timing consistency map as a reference, taking three short timescales before and after the dominant transmission timing of the current frame and three short timescales before and after the dominant transmission timing of the reference frame at each pixel, arranging them in an interleaved order of the current scale and the reference scale, forming an interleaved phase window containing 12 short timescales; when When any end crosses the boundary, the boundary scale is repeatedly filled to keep the window length at 12 short timescales. Within each staggered phase window, pulse events of the edge pulse layer, texture pulse layer, and block boundary pulse layer are merged according to the following rules: if the short timescale interval between two adjacent pulse events in the same pulse layer does not exceed 2 short timescales and the polarity is the same, they are merged into the same pulse cluster; otherwise, a new cluster is created. Pulse clusters containing only 1 pulse event are marked as isolated pulse clusters, and pulse clusters containing 2 or more pulse events are marked as valid pulse clusters.

[0010] Furthermore, in step 2, the process of threshold voting and polarity determination for pulse clusters in each window includes: within each interleaved phase window, threshold voting and polarity determination are performed on all pulse clusters of each pulse layer, according to the following rules: the base number of votes for each pulse cluster is equal to the number of pulse events within that cluster; when the difference between the center short timescale of the pulse cluster and the dominant emission timing of that pixel does not exceed 3 short timescales and the consistency level is high, 2 votes are added to that cluster; when it is medium, 1 vote is added; when it is low, no votes are added; the pulse cluster with the highest number of votes is recorded as the winning cluster, and its polarity is used as the polarity determination result of that pulse layer within the interleaved phase window, and its center short timescale is recorded as the candidate dominant emission timing; if there is a tie in the number of votes, the pulse cluster with the earlier center short timescale is selected; isolated pulse clusters are directly eliminated when the consistency level is low; when the consistency level is medium, they are retained only when isolated pulse clusters of the same polarity also appear in adjacent interleaved phase windows; when the consistency level is high, they are retained but their votes do not receive additional votes.

[0011] Furthermore, in step 2, the process of performing peak-shifting compensation on the pulse residuals formed by unmatched pulses according to adjacent time slices, and splitting them through the structure-priority branch and the texture-priority branch, includes: performing peak-shifting compensation on each pulse residual based on its distance from the candidate dominant release timing, wherein: when the difference between the center short time scale and the candidate dominant release timing does not exceed 2 short time scales, the residual as a whole is moved 1 to 2 short time scales away from the candidate dominant release timing, and the minimum movement amplitude that does not overlap with the confirmed winning cluster is preferred; when both sides are idle, 2 short time scales are selected for movement; the default movement amplitude for edge pulse layers and block boundary pulse layers is 1 short time scale, and texture pulse layers are allowed to move up to 2 short time scales; when the intersection When there are no idle short-timescales within the phase-shifted window, the residual is allowed to be moved to the edge short-timescale of the adjacent phase-shifted window, but it must not overlap with the paired clusters on the reference side. The pulse residuals after peak-shifting compensation are split into pulse layers, specifically: edge pulse layers and block boundary pulse layers enter the structure priority branch, and texture pulse layers enter the texture priority branch. The structure priority branch retains at most one pulse residual in each phase-shifted window of each pixel, and prioritizes retaining the pulse residual with the smallest timing difference from the candidate dominant release. The texture priority branch retains at most two pulse residuals in each phase-shifted window of each pixel, and sets an upper limit on the number of pulse events in each pulse residual. When the upper limit is exceeded, the last pulse event is discarded evenly in chronological order until the upper limit is met.

[0012] Furthermore, in step 2, the process of merging pulse events into an enhanced pulse event stream according to pulse timing consistency weights includes: merging the confirmed winning clusters of each pulse layer with the pulse residuals output from the two branches to form a candidate set for the enhanced pulse event stream; when multiple pulse events conflict at the same short timescale of the same pixel, conflict resolution is performed, which is achieved by mapping consistency levels to merging weights, where high consistency corresponds to the third weight, medium consistency corresponds to the second weight, and low consistency corresponds to the first weight, and pulse events with larger merging weights are preferentially retained; when the merging weights are the same, selection is based on pulse layer priority, which is edge pulse layers > block boundary pulse layers > texture pulse layers; if the priorities are still the same, pulse events with earlier generation times are retained; the pulse events after the aforementioned resolution are written into the enhanced pulse event stream in ascending order of short timescale.

[0013] Furthermore, in step 2, the process of performing threshold drift self-correction based on the inter-frame dominant release timing difference includes: within the cross-frame long time-stamp window, the dominant release timing of the current frame and the reference frame are statistically analyzed for each pixel, and the inter-frame dominant release timing difference is calculated; when the dominant release timing difference of the same pixel is greater than or equal to 5 short time-stamps in 3 consecutive frames and the consistency level is high, the retention threshold of the threshold voting in this step is reduced by 1 level to improve the tolerance for delay changes; when the dominant release timing difference of the same pixel is less than or equal to 2 short time-stamps in 3 consecutive frames and the pixel no longer generates pulse residuals in the interleaved phase window, the retention threshold of the threshold voting in this step is increased by 1 level to reduce unnecessary retention; the single-frame adjustment range of threshold drift self-correction does not exceed 1 level; when the consistency level is low or the pulse residual density per unit area exceeds 50,000 pulse events per megapixel per frame, the threshold drift self-correction is paused.

[0014] The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding of this invention has the following advantages: it can simultaneously maintain the structural integrity and visual quality of video under ultra-low bitrate conditions, and solves the problems of structural blurring, texture drift, temporal instability, and insufficient bitrate coordination in existing technologies. By introducing an event-driven mechanism of a spiking neural network within the loop, reconstructed frames are uniformly encoded into three types of pulse event streams: edge pulse layer, texture pulse layer, and block boundary pulse layer. Pulse events are issued on two time axes: short timescale within the frame and long timescale across frames. This dual-timescale processing method ensures both detail enhancement and global consistency. Combined with the constraints of the pulse temporal consistency graph, an interleaved phase review kernel and threshold voting decision mechanism are adopted to effectively suppress isolated pulses and pseudo-events. At the same time, through asynchronous residual back-injection and peak offset compensation, unmatched events can be differentiated in the structure-priority and texture-priority branches, and form an enhanced pulse event stream in the final coherent merging, thereby ensuring that structural boundaries are preferentially preserved and texture details are restored within a controllable range. Furthermore, in the closed-loop and rate-coordinated adjustment within the loop, the pulse reconstruction enhancement frames generated by the enhanced pulse event stream are used to replace the reference frames, and the pulse firing statistics are fed back to the rate control, ensuring that the quantization adjustment of subsequent frames remains consistent with the structural target. When the bitrate drops sharply, the system automatically enters a density reduction protection mode to suppress excessive firing of texture pulses and increase edge weights, thereby preventing structural collapse. When the bitrate recovers, a gradual release strategy is adopted to gradually restore the event firing of the texture layer, maintaining a stable transition of the enhancement effect. Through the above design, this invention, while ensuring coding efficiency, can achieve a comprehensive enhancement effect of structure priority, texture control, and temporal stability at ultra-low bitrates, significantly improving the subjective quality of the video and the reliability of the reference frame, reducing the risk of cross-frame cumulative distortion, and thus having outstanding practical application value under low-bandwidth transmission and limited storage conditions. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 A schematic diagram illustrating the principle of the pulse coding layer leakage integral and triggered emission (LIF) mechanism provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the construction principle of an interleaved phase window based on an interleaved phase core, provided in an embodiment of the present invention. Figure 3 This is a statistical diagram illustrating the threshold voting and polarity decision of a pulse cluster provided in an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] An AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding, comprising: Step 1: Obtain the reconstructed frame generated by the current decoder as visual input, and obtain the current bitrate indicator as control input; transform the reconstructed frame into three types of pulse event streams through the pulse coding layer of the spiking neural network, namely edge pulse layer, texture pulse layer and block boundary pulse layer; the pulse coding layer adopts leakage integration and trigger firing mechanism to fire pulse events simultaneously on two time axes: short time scale within the frame and long time scale across the frame, and sets the upper limit of firing density and reset hysteresis range according to the bitrate indicator, thereby generating an initial pulse event stream containing event timestamps and polarities and its pulse timing consistency map; Step 2: Under the constraints of the pulse timing consistency map, pulse consistency re-estimation and asynchronous residual back-injection based on interleaved phase kernel are performed. This process specifically includes: interleaving the current initial pulse event stream and the reference pulse event stream on short time scales within the frame and long time scales across frames to form an interleaved phase window; performing threshold voting and polarity determination on pulse clusters in each window to suppress isolated pulses and confirm the dominant firing timing; performing peak offset compensation on pulse residuals formed by unmatched pulses according to adjacent time slices, and splitting them through structure priority branches and texture priority branches, and then merging them into an enhanced pulse event stream according to pulse timing consistency weights; at the same time, performing threshold drift self-correction based on the inter-frame dominant firing timing difference to maintain the consistency and stability of long time scales, thereby forming an enhanced pulse event stream that can enhance structure and refine texture at ultra-low bit rates. Step 3: Under the control of closed-loop and code rate coordinated adjustment within the loop, the enhanced pulse event stream is used to generate a pulse reconstruction enhanced frame, and the pulse reconstruction enhanced frame is used to replace the reference frame.

[0019] In one implementation, after the decoder completes decoding of the current video frame, the reconstructed frame generated by the decoder is directly used as the visual input. Simultaneously, the current bitrate indicator corresponding to that frame is obtained from the encoding side or the bitstream parsing side as the control input. The bitrate indicator can be calculated comprehensively based on the frame's bit consumption, quantization step size, target bitrate, or buffer state, and mapped to discrete levels from 0 to 100 in an offline pre-defined lookup table. Through this normalization process, the bitrate states under different resolutions, scenes, and encoding configurations are compressed into a unified integer range, enabling the subsequent spiking neural network to adaptively adjust the pulse event firing behavior on a unified control scale, without requiring separate design for each encoding configuration.

[0020] Before pulse coding, the reconstructed frames undergo uniform image preprocessing. Specifically, the reconstructed frames are converted into single-channel images of the luma component, for example, retaining only the luma component according to the separation method of luma and chroma components in common video coding formats. The reason for using a single-channel luma component image is that in ultra-low bitrate videos, the human eye's perception of structural contours and texture details is mainly affected by changes in luma. The chroma component is already significantly simplified under strong compression conditions, so directly enhancing the luma component can bring a more significant improvement in subjective quality with the same complexity. Subsequently, through a linear stretching operation, the luma pixel values ​​are mapped to a fixed integer range, such as 0 to 255, while maintaining the spatial resolution. This linear stretching can normalize scenes that are too dark or too bright in the input video, giving the intermediate grayscale areas more dynamic utilization space. Therefore, in the subsequent leakage integration and trigger firing processes, luma differences can be stably mapped to differences in pulse firing frequency. During this process, the system records the image width and height to determine the spatial grid used in subsequent processing, ensuring that each pixel position can be uniquely identified throughout the entire processing flow using a fixed row and column index.

[0021] After spatial preprocessing, to simultaneously characterize the fine temporal sequence within a single frame and the long-term stability across multiple frames in the temporal domain, two time axes are established: intra-frame short timescales and cross-frame long timescales. The intra-frame short timescale is divided into a fixed number of discrete scales within a frame period. In the preferred implementation, this number is fixed at 1000, dividing a frame into 1000 evenly spaced time slices. Using 1000 short timescale scales allows for sufficient subdivision of a frame's time to distinguish the triggering order of different structures and textures under common frame rate conditions, while also preventing the time grid from becoming too dense, thus avoiding a sharp increase in the amount of pulse event timestamp data. The cross-frame long timescale spans a cumulative window across consecutive frames. In the preferred implementation, this window length is fixed at 8 frames, and a sliding update strategy is employed. By using a long time-stamped window spanning eight frames, a time range from hundreds of milliseconds to tens of milliseconds can be covered. This allows stable edges and block boundaries to be observed multiple times within the window, while random noise and transient compression artifacts are unlikely to reappear in most frames, thus providing a reliable observational basis for subsequent consistency assessments.

[0022] After establishing the timeline, the upper limit of the firing density and the range of the reset hysteresis of the spiking neural network are controlled using the current bitrate indicator. Specifically, the normalized bitrate indicator is divided into three levels. When the bitrate indicator is in the range of 0 to 33, it indicates that the current bitrate is extremely low, and the bit resources allowed for detailed description are very limited. In this case, the upper limit of the firing density is set to no more than 100,000 pulse events per megapixel per frame, while the reset hysteresis range is set to 6 to 10 discrete increments. A lower upper limit of the firing density can significantly suppress excessive pulse events in noisy regions and regions with large quantization errors, while a longer reset hysteresis range allows each pulse to remain silent for a period of time after firing, which is beneficial for highlighting truly structurally significant edges and block boundaries and avoiding fine artifacts that are incorrectly amplified at ultra-low bitrates. When the bitrate indicator is between 34 and 66, the video has moderate bitrate resources. At this point, a moderate increase in detail enhancement is allowed, raising the upper limit of the release density to no more than 250,000 pulse events per megapixel per frame, while shortening the reset hysteresis range to 4 to 6 discrete increments. This allows more texture details to be encoded in pulse form. When the bitrate indicator is between 67 and 100, the bitrate is relatively ample, and the encoder has already retained more original information. Therefore, the expressive power of pulse events can be further unleashed during the enhancement stage. The upper limit of the release density is raised to no more than 500,000 pulse events per megapixel per frame, and the reset hysteresis range is further shortened to 2 to 4 discrete increments. This reduces the time interval between adjacent pulse releases, thus finely depicting high-frequency textures and fine structures. Under this hierarchical control, the upper limit of the release density and the reset hysteresis range are always coordinated with the bitrate indicator, ensuring that the number and temporal distribution of pulse events are neither too sparse, resulting in insufficient enhancement capabilities, nor too dense, amplifying compressed noise under low bitrate conditions.

[0023] In the specific implementation of the spiking neural network, three types of pulse event streams can be constructed for each pixel location, corresponding to the edge pulse layer, texture pulse layer, and block boundary pulse layer, respectively. To obtain the inputs for these three pulse layers, various spatial filtering operations are performed on the single-channel image of the luminance component. For example, the directional gradient operator is used to extract luminance changes in the horizontal and vertical directions to emphasize structurally significant regions such as object contours as input for the edge pulse layer; high-pass or band-pass filtering operators are used to extract local high-frequency changes to characterize surface texture and fine details as input for the texture pulse layer; and block contrast operations matching the coding block division are used to highlight inter-block boundaries and locations with large differences in mean within blocks, thus constructing the input for the block boundary pulse layer. Through this feature splitting method, the same luminance information is decomposed into three different feature channels: structural edges, detailed textures, and block boundary transitions. Subsequent pulse coding can apply different triggering strategies to different types of information, making the enhancement results more in line with the human eye's preferences in ultra-low bitrate video scenarios: first ensuring the stability of the overall contour and block structure, and then gradually adding texture details as resources allow.

[0024] For each type of feature channel, the pulse coding layer employs a leakage integration and trigger firing mechanism to simultaneously fire pulse events on both the intra-frame short timescale and the cross-frame long timescale. Specifically, each pixel position can be considered a pulse unit with an internal state. This pulse unit accumulates potential based on the current feature intensity at each intra-frame short timescale, superimposing a portion of the potential left over from the previous short timescale, while simultaneously leaking potential according to a preset ratio. This allows the potential to gradually decay towards its initial state when there is no significant feature input for an extended period. When the accumulated potential exceeds the pulse trigger threshold, the pulse unit generates a pulse event at the current short timescale, determining the polarity of the pulse event based on the direction of feature change. For example, in the edge pulse layer, a transition from dark to bright can correspond to a positive polarity pulse event, and a transition from bright to dark can correspond to a negative polarity pulse event. After generating a pulse event, the pulse unit's potential is partially or completely reset and remains in a reset lag state for several subsequent short timescales. During this period, even if a new feature input is strong, it will not immediately trigger a new pulse firing. The duration of the reset lag is determined by the aforementioned reset lag range and a specific value is selected based on the current bit rate indication, thereby increasing the time interval between adjacent pulses under low bit rate conditions to ensure that only significant structural changes in both space and time can continuously trigger pulse events.

[0025] refer to Figure 1 , Figure 1This diagram illustrates the principle of the leakage integration and trigger firing mechanism employed in the pulse coding layer of this invention. The figure details the dynamic behavior of a pulse unit at a single pixel location along the intra-frame short-timescale dimension using a two-dimensional coordinate system. The horizontal axis represents the intra-frame short-timescale scale, and the vertical axis represents the accumulated potential value within the pulse unit. In the ultra-low bitrate video encoding / decoding in-loop enhancement method, this mechanism is the core step in converting discrete digital image signals into a temporally sparsity-based pulse event stream. As shown, the system first acquires the reconstructed frame generated by the current decoder as visual input. After luminance component extraction and linear stretching preprocessing, these pixel values ​​are fed into the input of the spiking neural network. The black solid line in the figure represents the trajectory of the pulse unit's membrane potential changing over time. At each intra-frame short-timescale scale, the pulse unit receives an input current from the current feature intensity, which is proportional to the pixel's luminance value or its spatial gradient. Simultaneously, the system introduces a leakage mechanism, causing the potential left over from the previous moment to leak according to a preset attenuation ratio. This process is represented in the figure as the natural downward trend of the potential curve when there is no strong input. The existence of the leakage mechanism is crucial. It simulates the characteristics of biological neurons, ensuring that only recent and significant feature inputs can maintain a high potential level, thereby effectively filtering out random high-frequency noise and instantaneous quantization errors common in ultra-low bitrate videos.

[0026] The diagram shows a horizontal dashed line as the trigger threshold. When the accumulated potential, represented by the black solid line, rises above this threshold at a certain short timescale, the pulse unit immediately enters the trigger state and emits a pulse event. In the diagram, this event is visually represented by an arrow emanating from above the threshold line. The position of the arrow corresponds to the timestamp of the emission, and the direction or color of the arrow indicates the polarity of the pulse, i.e., whether the brightness changes towards brighter or darker. Once a trigger emission occurs, the internal potential of the pulse unit is immediately reset, typically to zero or a low resting potential, in preparation for the next integration. Figure 1 A particularly noteworthy design feature is the reset hysteresis region after triggering, which is the gray shaded area immediately following the output in the diagram. Within this region, regardless of the intensity of the input feature, the pulse unit is in either an absolute or relative refractory period, temporarily ceasing its response to new inputs. This reset hysteresis range is not fixed but serves as a key control variable directly governed by the current bit rate indicator.

[0027] When explaining the control logic of this mechanism in detail, the context of ultra-low bitrates must be considered. When the bitrate indicator shows that the current bitrate is extremely low, the system automatically lengthens the time span of the gray shaded area in the image. This means that after a pulse is emitted, the pulse unit will remain silent for a longer period of time. This design physically forces a reduction in the maximum frequency of pulse emission, thus limiting the upper limit of emission density. From a visual enhancement perspective, this is a protective strategy: when bit resources are extremely scarce, reconstructed frames are often filled with block artifacts and glitches caused by large quantization steps. If pulse units are allowed to respond frequently, these artifacts will be converted into high-frequency pulses and incorrectly amplified. By expanding the reset hysteresis range, the system forces the pulse coding layer to respond only to those structural features (such as the main outline of an object) with the highest intensity and longest duration, while ignoring rapidly flickering noise. Conversely, when the bitrate indicator is high, the system shortens the reset hysteresis range, making the shaded area in the image narrower, allowing the pulse unit to reset at a faster rate and respond to new detail changes, thereby capturing texture information in the image. Therefore, Figure 1 This not only demonstrates a static integral distribution process, but also dynamically reflects how the present invention achieves adaptive control of enhancement intensity at the microscale of short intra-frame timescales by adjusting potential dynamics parameters, ensuring that the generated initial pulse event stream matches the current coding quality state in terms of temporal distribution, thus laying a precise foundation for subsequent consistency processing.

[0028] In the cross-frame long-timescale dimension, the spiking neural network retains the firing history of each pixel location within a window of eight consecutive frames, including the intra-frame short-timescale scale of each firing and the corresponding pulse polarity. In this design, the intra-frame short-timescale reflects the fine temporal distribution within a single frame, while the cross-frame long-timescale records the repetition and stability across multiple frames. For example, for a stable object edge, the edge pulse layer will repeatedly trigger pulse events in adjacent frames, and the firing scale will not differ significantly across frames. For random textures caused by quantization noise, pulse events are often triggered only occasionally at specific short-timescale scales in individual frames, and the firing scale lacks temporal concentration. Utilizing this temporal distribution characteristic, a pulse temporal consistency map can be constructed subsequently to distinguish stable structures from transient noise from a temporal perspective.

[0029] When constructing the initial pulse event stream, the system organizes each triggered pulse event in the edge pulse layer, texture pulse layer, and block boundary pulse layer into a time-ordered event sequence. Each pulse event can include the row and column indices of the pixel, the pulse layer category, the intra-frame short timestamp scale, and the pulse polarity. To control data volume, the number of pulse events per megapi per frame is constantly monitored during pulse event generation. When the number of pulse events in a frame approaches the upper limit of the release density corresponding to the current bitrate indicator, the pulse coding layer can increase the trigger threshold or correspondingly extend the reset hysteresis, reducing the trigger probability in the remaining short timestamp scales and ensuring that the final number of generated pulse events does not exceed the design limit. This proactively reduces secondary events without disrupting the main structure and significant texture, maintaining the overall match between the pulse event stream and the available bitrate.

[0030] After generating the initial pulse event stream, to provide prior constraints for subsequent pulse consistency reassessment and asynchronous residual back-injection, a pulse timing consistency map needs to be constructed based on the firing history within a cross-frame long time-stamped window. Specifically, for each pixel location within the cross-frame long time-stamped window, the dominant firing scale for that pixel in the last 8 frames is calculated. That is, within each frame, the short time-stamped scale with the highest firing intensity or closest to the trigger threshold is selected as the dominant firing scale for that frame. These 8 dominant firing scales are then clustered. When a cluster appears in at least 5 frames, and the difference between firing scales within that cluster does not exceed 3 short time-stamped scales, the pixel is considered highly consistent, and the dominant firing scale corresponding to that cluster is recorded as the dominant firing timing for that pixel. Highly consistent pixels typically correspond to static object outlines, stable block boundaries, or structural boundaries that remain aligned after camera motion compensation. These areas are most likely to expose compression artifacts subjectively under ultra-low bitrate conditions, and therefore require higher weighting in subsequent enhancement. When a cluster appears in 3 or 4 frames, and the difference between each firing scale does not exceed 5 short timescales, the pixel is judged as having medium consistency. These pixels generally correspond to slowly moving structures or periodic textures, and their temporal distribution has a certain stability but is not as concentrated as static edges. The rest are judged as having low consistency, which usually represents fast-moving areas, complex textures, or pure noise areas. These areas are difficult to reliably recover at ultra-low bitrates, and excessive enhancement may introduce flicker.

[0031] The aforementioned consistency levels are encoded as integers in a raster with the same resolution as the image to form a consistency level map. For example, 0 can represent low consistency, 1 represents medium consistency, and 2 represents high consistency. Simultaneously, the dominant release timing determined for high or medium consistency pixels is written as an integer value with a short intra-frame timestamp to the corresponding raster position to form a timing index map. Low consistency pixels can be filled with a preset default scale or according to the scale of the most recent valid cluster. Finally, the consistency level map and the timing index map are combined into a pulse timing consistency map, which is output along with the initial pulse event stream. The existence of the pulse timing consistency map allows subsequent processing to determine whether a pulse event should be retained, offset, or merged, no longer relying solely on local information of the current frame, but instead referencing the temporal consistency of the last 8 frames. In ultra-low bitrate scenarios, this temporal consistency-based weighted guidance can robustly distinguish between occasional pulses perturbed by compressed noise and structural pulses that truly need enhancement without significantly increasing computational complexity, helping to maintain the stability of video structure and the naturalness of texture even with extremely low bit budgets.

[0032] In optional implementations, the linear stretching range of the luminance component can be adjusted according to the bit depth of the source video. For example, for source videos with higher bit depths, pixel values ​​can be stretched to an integer range of 0 to 1023, and then mapped back to a potential range suitable for leakage integration through a lookup table within the spiking neural network, thus ensuring compatibility with video sources of different bit depths. While maintaining a fixed number of short timescales within a frame (1000) and a fixed cross-frame long timescale window length of 8 frames, the frame number threshold and scale difference threshold used in consistency determination can be fine-tuned during offline training based on different content types. For example, for applications with many motion scenes, appropriately increasing the scale difference threshold can improve tolerance for smooth motion, allowing the edges of moving objects to be classified as having medium or high consistency, thereby preventing motion regions from remaining in a low-consistency state for extended periods and lacking sufficient enhancement during the enhancement process. Through these methods, the initial pulse event stream and its pulse timing consistency map can maintain stable and flexible performance under different bitrates, content, and shooting conditions, providing a balanced and reliable input foundation for subsequent pulse consistency re-estimation and asynchronous residual back-injection based on interleaved phase kernels.

[0033] After constructing the current initial pulse event stream and the pulse timing consistency map, pulse consistency re-estimation and asynchronous residual back-injection based on interleaved phase verification can be performed under the constraints of the pulse timing consistency map. This process utilizes the temporal relationship between the current initial pulse event stream and the reference pulse event stream to reconfirm and correct the dominant delivery timing. On the other hand, it performs controlled peak-shifting compensation and back-injection on pulse residuals that are not directly matched, making the enhanced pulse event stream more stable in time while taking into account structure and texture in space, thereby recovering as many clear edges and natural details as possible under ultra-low bitrate conditions.

[0034] In one implementation, the current initial pulse event stream and the reference pulse event stream are first interleaved and aligned pixel-by-pixel on intra-frame short timescales and cross-frame long timescales. Specifically, for each pixel position, the dominant firing timing scale of the current frame and the reference frame is read using the timing index map in the aforementioned pulse timing consistency map. In a preferred implementation, the reference frame can be the frame immediately preceding the current frame, or the frame with the smallest residual after motion compensation from the current frame can be selected as the reference frame from among multiple frames. To form an interleaved phase window, at each pixel position, based on the timing index map in the pulse timing consistency map, three intra-frame short timescale scales are selected before and after the dominant firing timing of the current frame, and three intra-frame short timescale scales are also selected before and after the dominant firing timing of the reference frame. These are then arranged in an interleaved order of the current scale and the reference scale, so that the scales of the current frame and the reference frame alternate on the time axis, resulting in an interleaved phase window with a length of 12 intra-frame short timescale scales. When the current frame's dominant delivery timing or the reference frame's dominant delivery timing is close to the beginning or end of a frame, one or both ends of the window may go out of bounds. In this case, the out-of-bounds position is filled by repeating the scale at the boundary, so that the interleaved phase window always keeps the short time scale within 12 frames unchanged.

[0035] refer to Figure 2 , Figure 2 This diagram illustrates the construction principle of an interleaved phase window based on an interleaved phase kernel. The diagram details how the system aligns and reassembles the temporal information of the current frame and the reference frame under the constraints of a pulse timing consistency map, enabling precise consistency evaluation. In ultra-low bitrate video enhancement, distinguishing between real structure and compressed noise is a significant challenge, and accurate judgment is often difficult based solely on single-frame information. Therefore, this invention introduces the concept of a long timescale across frames, specifically implemented as the interleaved phase window shown in the diagram. The horizontal axis represents the index sequence within the interleaved phase window, typically set to a fixed length, such as an integer sequence from 1 to 12; the vertical axis represents the absolute intra-frame short timescale. The core logic of this diagram lies in demonstrating how to extract segments from a continuous timeline and interweave time points from different sources.

[0036] During the construction process, the system first determines the dominant firing timing (denoted as C) of the current pixel in the current frame and the dominant firing timing (denoted as R) in the reference frame based on the timing index map in the pulse timing consistency map. These two timing points constitute the reference anchor points for window construction. As shown in the figure, in order to capture the firing behavior near these two time points, the system selects several short time scales before and after the dominant firing timing C in the current frame, and selects the same number of short time scales before and after the dominant firing timing R in the reference frame. In the preferred embodiment shown in the figure, three scales are selected before and after, covering the small fluctuations within the local time range. Subsequently, these selected scales are not simply spliced ​​together, but arranged on the horizontal axis in a specific staggered order. The figure clearly depicts this staggered arrangement: the sampling points of the current frame and the sampling points of the reference frame appear alternately on the horizontal axis. For example, position 1 corresponds to a scale in the current frame, position 2 corresponds to a scale in the reference frame, and so on. Solid dots represent samples from the current frame, while hollow dots represent samples from a reference frame. They are connected by dashed lines to form a unified viewing window.

[0037] The physical significance of this staggered phase window construction is profound. It essentially establishes a virtual local time domain where the temporal behavior of the current frame and the reference frame are placed side-by-side for direct comparison. If a structure is stable, then in the graph, the positions of the solid and hollow dots on the vertical axis should be very close, appearing as a relatively gentle broken line; this means that the pulse firing time of the current frame highly coincides with the historical firing time of the reference frame, indicating temporal continuity. Conversely, if it is a random pulse caused by noise, the position of its solid dot may be far from that of the hollow dot on the vertical axis, appearing as a sharp jump in the broken line. In this way, the staggered phase window transforms the complex cross-frame consistency problem into an analysis of the window's internal geometry. Furthermore, Figure 2 It also implicitly includes instructions on window boundary handling. When the dominant firing timing is near the beginning or end of a frame, the selected range may exceed the limits. In this case, the system uses boundary duplication or mirror padding to fill the window, ensuring that the generated interleaved phase window always maintains a fixed length (e.g., 12 units) regardless of when the pulse occurs within the frame. This standardized window construction provides a normalized operating platform for subsequent pulse cluster merging and threshold voting, eliminating the need for separate processing logic for each time point and greatly improving the efficiency of parallel processing. Figure 2 As shown in the construction principle, this invention successfully introduces a macroscopic historical reference into a microscopic time slice, providing solid logical support for accurately capturing and restoring video structure in ultra-low bitrate environments.

[0038] This staggered phase window approach allows for simultaneous focusing of the dominant firing timing of both the current and reference frames within a local timeframe, providing alignment references for pulse events around both frames. The current frame provides the target information to be enhanced, while the reference frame reflects the temporal history of the same pixel. By staggering these two frames, it's possible to directly compare the consistency of the current and reference pulse distributions within a single window. If a window is only established within the current frame, without considering the reference frame, it becomes difficult to distinguish between temporal drift caused by genuine content changes and pseudo-pulses due to compressed noise. By introducing the dominant firing timing of the reference frame and staggering it, verification can be performed from the same temporal perspective, reducing the probability of misjudgment.

[0039] After forming staggered phase windows, pulse events from the edge pulse layer, texture pulse layer, and block boundary pulse layer are merged within each window to generate pulse clusters. Specifically, within the same pulse layer, two pulse events with adjacent short timescales are compared. If the intra-frame short timescale interval between the two events is no more than two short timescales and the pulse polarity is the same, then these two pulse events are merged into the same pulse cluster. If the interval is greater than two intra-frame short timescales or the polarities are different, then a new pulse cluster is created. After the above merging, pulse clusters containing only one pulse event are marked as isolated pulse clusters, and pulse clusters containing two or more pulse events are marked as valid pulse clusters. This merging method is equivalent to treating close pulse events with the same polarity as the same continuous activity on the time axis. It can reflect the concentrated firing behavior of the same structure or texture within a local time range, while distinguishing occasional one-off firings for subsequent suppression and filtering. For quantization noise and residual block effects commonly found in ultra-low bitrate videos, their triggering often occurs in isolation on a single time scale. By marking pulse clusters containing only one pulse event as isolated pulse clusters, conditions can be prepared for subsequent compression.

[0040] After constructing the pulse clusters within each interleaved phase window, the dominant firing sequence needs to be confirmed through threshold voting and polarity determination, suppressing unreliable isolated pulse clusters in the process. In practice, threshold voting and polarity determination can be performed on all pulse clusters of each pulse layer within each interleaved phase window. The base number of votes for each pulse cluster can be set to the number of pulse events within that cluster; the more events, the more frequent the firing in that time period, and the higher the base number of votes. Simultaneously, the consistency level in the pulse timing consistency map is incorporated into the voting process. When the difference between the center short timescale of a pulse cluster and the dominant firing sequence of that pixel does not exceed three intra-frame short timescales, and the consistency level is high, two additional votes are added to that cluster; one additional vote is added when the consistency level is medium; and no additional votes are added when the consistency level is low. This approach considers both the number of pulse events within a given time period and the stability of that time period across long timescales, ensuring that pulse clusters that repeatedly occur across multiple frames and maintain consistency with the dominant firing sequence are significantly favored.

[0041] After counting the votes, the pulse cluster with the highest number of votes is selected as the winning cluster. Its pulse polarity is used as the polarity determination result of that pulse layer within the interleaved phase window, and its center short timescale is recorded as the candidate dominant emission sequence. If multiple pulse clusters have the same number of votes, the pulse cluster with the earlier center short timescale is selected as the winning cluster. This prioritizes the retention of earlier structural changes and reduces unstable flickering caused by delay fluctuations. When processing isolated pulse clusters, to further suppress noise, in cases of low consistency, all isolated pulse clusters can be directly eliminated to prevent the retention of a single pulse event in obviously unstable regions. In cases of medium consistency, isolated pulse clusters with the same polarity are only retained if they also appear in adjacent interleaved phase windows. This is equivalent to requiring that the isolated emission spans at least two adjacent windows in time to prove that it has a certain degree of persistence. In cases of high consistency, isolated pulse clusters can be retained, but their votes do not receive additional votes; they are still compared based on the basic vote count, thus ensuring that truly recurring effective pulse clusters maintain their advantage in high consistency regions. The reason for this design is that high consistency regions usually correspond to stable structural edges or block boundaries. Even if a pulse appears only once in a frame, it may represent subtle and realistic structural changes. In contrast, low consistency regions are mostly noise and irregular textures, and retaining isolated pulses can easily cause flickering.

[0042] After threshold voting and polarity determination, some pulse events may still fail to be included in the winning cluster or be determined to be isolated pulse clusters that need to be retained. These pulses constitute the pulse residuals of the current pixel within the interleaved phase window. Simply discarding all these pulse residuals would result in over-smoothing of texture regions and excessive loss of detail; while directly superimposing them with the winning cluster would cause overly dense firing in time, increasing noise. Therefore, it is necessary to perform peak offset compensation on these pulse residuals and inject them asynchronously into the enhanced pulse event stream through structure-first and texture-first branches.

[0043] In practical implementation, the peak-shifting compensation strategy can be determined based on the distance between the center short-timescale mark of each pulse residual and the candidate dominant firing sequence. When the difference does not exceed two intra-frame short-timescale marks, it indicates that the pulse residual is very close to the dominant firing sequence on the time axis. Direct superposition would result in a concentrated burst with the winning cluster, which visually appears as overly sharp or flickering edges. In this case, the pulse residual is moved as a whole away from the candidate dominant firing sequence by one to two intra-frame short-timescale marks, prioritizing the smallest movement amplitude that does not overlap with the confirmed winning cluster after the movement, to ensure a peak-shifting effect in time. If there are idle short-timescale marks on both sides of the candidate dominant firing sequence, it is preferred to move two intra-frame short-timescale marks, so that the residual maintains a moderate interval with the backbone structure in time, similar to arranging soft auxiliary pulses outside the main outline, thus preserving texture energy without stacking too many pulses on the same scale. By default, the staggered movement of the edge pulse layer and the block boundary pulse layer can be limited to a maximum of one intra-frame short time scale to avoid the structural information being stretched too far in time, affecting the clarity of the main outline; the texture pulse layer is allowed to move to two intra-frame short time scales to more fully unfold local details, allowing the texture information to spread smoothly along the time axis and improve the fineness.

[0044] When all short timescales within an interleaved phase window are occupied by the winning cluster or other reserved pulses, and no free timescale can be found within the window, the pulse residual can be moved to the edge short timescale of an adjacent interleaved phase window. This cross-window peak-shifting compensation allows the residual to be released over a wider time range, but it must not overlap with paired clusters on the reference side to avoid disrupting the established temporal correspondence between the current frame and the reference frame. After peak-shifting compensation, the pulse residuals are split by pulse layer: pulse residuals from edge pulse layers and block boundary pulse layers enter the structure-priority branch, and pulse residuals from texture pulse layers enter the texture-priority branch. The structure-priority branch retains at most one pulse residual in each interleaved phase window per pixel, prioritizing the one with the smallest difference in timing from the candidate dominant release, thus ensuring that only one auxiliary pulse trajectory is retained near the structural information to enhance the stability of the boundary without generating multiple interfering structural lines. The texture-first branch retains a maximum of two pulse residuals within each interleaved phase window of each pixel, and sets an upper limit on the number of pulse events within each pulse residual, such as limiting it to no more than four or eight. When the upper limit is exceeded, the last pulse events are discarded evenly in chronological order until the upper limit is met. This keeps the texture information within a limited temporal density range, preserving sufficient detail sampling while avoiding graininess caused by excessively dense pulse events in the texture area.

[0045] After the structure-first and texture-first branches filter the pulse residuals, the confirmed winning clusters of each pulse layer and the pulse residuals output by the two branches need to be merged to form a candidate set for enhancing the pulse event stream. Since pulse events from multiple sources may conflict at the same short-timescale of the same pixel within the same frame, conflict resolution is required to avoid repeated issuance at the same time. During the resolution process, priority is achieved by mapping the consistency level in the pulse temporal consistency map to a merging weight. For example, high consistency can be assigned to the third weight, medium consistency to the second weight, and low consistency to the first weight, with the third weight greater than the second weight, and the second weight greater than the first weight. When multiple pulse events exist at the same pixel and the same short-timescale, the pulse event with the larger merging weight is retained first. This ensures that, in the case of competition among multiple candidate events, events that are more temporally stable and have more consistent cross-frame performance always gain priority. If multiple events have the same merging weight, the selection is based on the pulse layer priority, with the priority order being edge pulse layer > block boundary pulse layer > texture pulse layer. This sorting method reflects the consideration of ensuring the clarity of structural edges and block boundaries first in ultra-low bitrate scenarios, and then supplementing texture details when resources allow. When the merging weight and pulse layer priority are still the same, for example, when multiple events within the same layer come from different branches, the pulse event with the earlier generation time can be retained as the representative event at that scale, thereby further reducing redundancy.

[0046] After the aforementioned conflict resolution, all retained impulse events are written into the enhanced impulse event stream in ascending order of intra-frame short timescales. In this configuration, the enhanced impulse event stream follows the dominant firing sequence in the temporal dimension, while also interspersing a certain amount of auxiliary impulses around it through structure-priority and texture-priority branches. This creates a stable main firing sequence near structural boundaries, and the addition of several staggered supplementary firings makes edges appear more continuous and block boundaries smoother. Simultaneously, it maintains a certain temporal diversity in texture areas, preventing excessive smoothing of the image.

[0047] To maintain the consistency and stability of pulse timing over long periods and adapt to timing changes caused by content movement or encoding delays, threshold drift self-correction can be performed based on the inter-frame dominant release timing difference within a long-time-stamped window spanning multiple frames. Specifically, for each pixel, the dominant release timing of the current frame and the reference frame are statistically analyzed within the long-time-stamped window, and the inter-frame dominant release timing difference is calculated. When the dominant release timing difference of the same pixel is greater than or equal to 5 short-time-stamped intervals within three consecutive frames, and the consistency level is high, it indicates that the main structure of the region where the pixel is located has experienced a stable and continuous temporal shift, possibly due to slow movement, camera panning, or accumulated delay compensation. If a high threshold voting retention threshold is maintained, these shifts will be considered anomalous, thus incorrectly weakening the true structure. Therefore, in this case, the threshold voting retention threshold in this process can be lowered by one level, allowing pulse clusters with slightly lower vote counts but consecutive occurrences to be retained, thereby increasing the tolerance for delay changes and enabling the dominant release timing to gradually migrate with the actual changes in content.

[0048] Conversely, when the dominant firing timing difference of the same pixel is less than or equal to two intra-frame short timescales in three consecutive frames, and the pixel no longer generates pulse residuals within the interleaved phase window, it indicates that the current dominant firing timing is highly stable, and there are almost no redundant residuals requiring peak compensation. In this case, if the threshold voting retention threshold is too low, too many major events may be retained in highly stable regions, causing slight noise in the image. Therefore, in this situation, the threshold voting retention threshold can be increased by one level, allowing only pulse clusters with higher votes to pass, thereby reducing unnecessary retention and further smoothing continuous regions. To avoid rapid and repeated changes in the threshold between frames, the adjustment range of the threshold drift self-correction in each frame does not exceed one level. Through this gradual strategy, the system slowly adapts to long-term trends without over-responding to short-term fluctuations.

[0049] refer to Figure 3 , Figure 3 This diagram illustrates the threshold voting and polarity decision statistics for pulse clusters, providing an in-depth analysis of how a competition mechanism filters out effective pulses and determines the enhancement direction within the interleaved phase window. Continuing the logic from the previous diagram, once the interleaved phase window is constructed, it contains raw pulse events from the current frame and the reference frame. Figure 3The decision-making process for processing these pulse events is visually illustrated using histograms or stacked bar charts. The horizontal axis represents the different pulse clusters identified within the window, labeled as cluster A, cluster B, cluster C, etc.; the vertical axis represents the weighted vote count received by each pulse cluster. In this step, the system first performs a merge operation, treating pulse events with extremely short time intervals (e.g., differing by no more than two short time scales) and the same polarity within the window as the same pulse cluster. The purpose of this operation is to integrate discrete pulse signals into "event packets" with a certain energy density, because in video signals, a real physical structure often triggers a series of dense pulses, while noise usually manifests as isolated single points.

[0050] Figure 3 The core of the demonstration lies in showcasing the scoring mechanism of "threshold voting." The height of each bar chart consists of two parts: the base vote count and the consistency-weighted vote count. The base vote count corresponds to the physical number of pulse events contained within that pulse cluster; the denser the pulses, the higher the base vote count. This reflects the direct contribution of signal strength. Above the base vote count, a consistency-weighted vote count, distinguished by different texture fills or colors, is superimposed. This portion of the votes originates from prior information in the pulse temporal consistency map. If a pixel is judged as "high consistency" (i.e., consistently appearing in the past 8 frames) over a long timescale, the system awards an additional 2 votes; "medium consistency" awards 1 vote; and "low consistency" awards no points. Figure 3 The diagram clearly illustrates the Matthew effect resulting from this weighting: cluster A, representing a stable structure, may have a similar number of base pulses as cluster B, representing texture, but due to its "high consistency" weighting, its total votes increase significantly, allowing it to stand out in the competition. A horizontal dashed line is also shown in the diagram, representing the "retention threshold." Only pulse clusters with a total vote count exceeding this threshold are considered the "winning cluster" and retained as the dominant signal.

[0051] also, Figure 3 The algorithm specifically demonstrates the decision-making logic for "isolated pulse clusters." For clusters containing only one pulse event and with a low consistency level (such as cluster C), their total votes are extremely low, far below the retention threshold, and are therefore marked as "suppressed / discarded" in the diagram. This vividly illustrates how this method filters out quantization noise in ultra-low bitrate videos. For some isolated clusters located in medium-to-high consistency regions (such as cluster D), they may still be retained as long as their vote count barely exceeds the threshold or meets specific cross-window continuity conditions. The polarity of the winning cluster will determine the enhancement direction (brightening or darkening) of that pixel at the current moment, while its center moment will be determined as the dominant firing sequence. Figure 3What is demonstrated is not merely a simple counting process, but an intelligent decision-making model that integrates temporal density analysis and historical reliability assessment. Through a quantitative voting mechanism, it resolves the perennial contradiction between "preserving detail" and "suppressing noise" under ultra-low bitrate conditions. By assigning higher weight to historical consistency, the system can accurately "recall" the necessary structural pulses based on historical memory, even when the current frame quality is extremely poor and the signal is weak, while ruthlessly eliminating random artifacts lacking historical basis. This mechanism ensures that the enhanced pulse event stream possesses both rich detail and extremely high temporal stability, providing crucial decision outputs for ultimately generating clear and flicker-free enhanced video frames.

[0052] When the consistency level of a pixel is low, or the pulse residual density per unit area exceeds 50,000 pulse events per megapixel per frame, threshold drift self-correction can be temporarily paused. Low-consistency regions inherently lack temporal stability; frequent threshold adjustments in such regions can cause the system to repeatedly change its retention strategy under noise-driven conditions, leading to highly unstable pulse event distribution. When the pulse residual density is excessively high, it indicates drastic scene changes, numerous compression artifacts, or overly aggressive initial parameter settings. Continuing to adjust the threshold under these conditions will cause the entire system to fall into a non-convergent state. By pausing threshold drift self-correction in these situations, the relative stability of the threshold voting strategy can be maintained. Self-correction can then be reactivated as the content gradually stabilizes or the residual density decreases, ensuring the system maintains controllable behavior even in complex scenes.

[0053] Through the aforementioned pulse consistency re-estimation and asynchronous residual back-injection process based on interleaved phase kernels, the enhanced pulse event stream forms a structural skeleton centered on the dominant firing sequence on the time axis. Pulse residuals, after peak-shifting compensation, are then strategically arranged around this skeleton. This strengthens structural edges and block boundaries under ultra-low bitrate conditions, while simultaneously dispersing texture details temporally, avoiding hard edges and flickering that are common in traditional enhancement schemes against high-compression artifacts. Overall, in ultra-low bitrate scenarios, the encoder has already discarded a significant amount of detail. Simply using static spatial filtering or single-frame enhancement would make it difficult to balance noise and real details. This method, however, utilizes temporal information from both intra-frame short-timescale and cross-frame long-timescale dimensions, combined with pulse temporal consistency maps, threshold voting, peak-shifting compensation, and threshold drift self-correction. This approach maximizes the extraction of remaining recoverable information while maintaining long-timescale consistency, making the resulting enhanced pulse event stream more suitable as the basis for subsequent pulse reconstruction enhancement frames.

[0054] In the optional implementation, the specific interleaving order of the current frame and reference frame ticks in the interleaved phase window can be fixed offline according to different content types. For example, it can be an order that expands outwards from the current frame's dominant release sequence, or it can be an order that expands outwards from the reference frame's dominant release sequence, as long as the window length is fixed at 12 short timescale ticks within the frame. The numerical interval of the consistency level mapping to the merging weights can also be adjusted during training. For example, in some surveillance videos with extremely high stability requirements, the weight gap between high consistency and medium consistency can be significantly widened, so that the high consistency region almost completely dominates the conflict resolution; while in sports videos with a lot of movement, the gap between high consistency and medium consistency can be appropriately narrowed to improve the tolerance for motion edges. In addition, the upper limit of the number of pulse events within each pulse residual in the texture-first branch can be reset for different resolutions and different target bitrates to ensure that the ability to preserve texture details can be correspondingly enhanced when the resolution increases or the target bitrate is slightly higher. Under these optional methods, pulse consistency reestimation and asynchronous residual backinjection based on staggered phase verification still operate under the constraints of the pulse timing consistency graph. The core process remains unchanged, only the specific values ​​and strategies are adapted to meet the visual requirements of different application scenarios.

[0055] After constructing the enhanced impulse event stream, it can be used to generate impulse reconstruction enhancement frames under the control of closed-loop and bitrate coordinated adjustment within the loop. These enhanced impulse event streams then replace the reference frames, ensuring stable operation of the entire video encoding and decoding process within the closed loop. The core idea here is to reconstruct the existing reconstructed luminance image at the decoding end using an enhanced impulse event stream that has been calibrated and strengthened with structural and texture information in time. This allows truly stable edges, block boundaries, and textures to be highlighted, while minimizing the amplification of quantization noise and prediction errors under ultra-low bitrate conditions.

[0056] In one implementation, the generation of the pulse reconstruction enhancement frame uses the reconstructed luminance image of the current frame from the decoder as a base estimate. Specifically, for each pixel location, the initial luminance value of that pixel in the reconstructed luminance image is first read. This luminance value has typically been linearly stretched to an integer range of 0 to 255 and is consistent with the luminance scale used in the enhancement pulse event stream. Then, at the spatial location of that pixel, the corresponding pulse event sequence is extracted from the enhancement pulse event stream, including three types of events: edge pulse layer, texture pulse layer, and block boundary pulse layer, arranged in the order of intra-frame short timestamps. Each event in the enhancement pulse event stream carries information such as the intra-frame short timestamp timestamp, pulse polarity, and the pulse layer to which it belongs. Simultaneously, the stability and dominant firing timing of the pixel across long timestamps can be obtained through the consistency level map and timing index map in the pulse timing consistency map. By combining this information, a temporally ordered "enhancement instruction sequence" can be constructed for each pixel to guide fine-tuning of the luminance value.

[0057] When reconstructing a single pixel, the edge pulse layer, texture pulse layer, and block boundary pulse layer can be considered as three sub-paths operating at different frequency bands. The edge pulse layer corresponds to structural contours with steep brightness changes, and therefore mainly controls local contrast and edge sharpness during reconstruction. The texture pulse layer corresponds to local details with higher frequencies but smaller amplitudes, used to restore subtle textures and complex details on the surface. The block boundary pulse layer corresponds to block boundaries and brightness steps within blocks generated by coded block division, mainly used to reduce block artifacts and smooth block boundary transitions. Whenever a pulse event is encountered, a brightness gain or brightness attenuation for the current pixel is calculated based on the pulse layer and polarity to which the event belongs, combined with the pixel's level in the consistency level map and the distance between the event timestamp and the dominant release sequence. In the preferred implementation, edge pulse layer events at high-consistency pixels can be assigned a higher base impact level. For example, within the range of 0 to 255, a single high-consistency edge pulse event can adjust pixel brightness by 2 to 4 gray levels, while the adjustment range in medium-consistency regions can be limited to 1 to 2 gray levels, and the adjustment range in low-consistency regions can be controlled within 1 gray level or even directly prohibited from structural enhancement. The reason for this is that high-consistency regions exhibit a stable dominant firing sequence over the past 8 frames, indicating that these edges or block boundaries are persistent real structures, and the enhancement intensity can be increased with confidence; while low-consistency regions show significant temporal oscillations, and excessive enhancement can easily mistake noise and transient artifacts for real details.

[0058] The processing of texture pulse layers differs from that of edge pulse layers. Since textures inherently allow for a degree of randomness, and texture details are often sacrificed at the encoder in ultra-low bitrate scenes, when restoring textures in pulse reconstruction enhancement frames, the emphasis is more on overall density and uniformity of distribution, rather than the absolute intensity of a single event. In practice, the texture pulse layer can be configured with smaller single-event brightness adjustments; for example, each event can adjust pixel brightness by no more than 1 to 2 gray levels, but multiple texture events can be accumulated near short time markers within a frame. Through this "small-step, multiple-step" stacking method, the textured area gradually exhibits subtle brightness fluctuations without producing excessively strong false edges. To prevent noticeable graininess from texture enhancement, the local contrast of the pixel can be calculated based on a 3×3 or 5×5 local neighborhood before each texture event. When the local contrast is already high, the enhancement amplitude of subsequent texture events is appropriately reduced; when the local contrast is low, the texture event is allowed to function fully. This allows texture enhancement to focus more on areas that were originally smooth but slightly blurry, while avoiding excessive amplification of noise in areas with already significant details.

[0059] The reconstruction of the block boundary pulse layer focuses on eliminating block artifacts and ensuring smooth transitions between blocks. For pixels located near block boundaries, when an enhanced pulse event aligned with the dominant firing sequence is detected in the block boundary pulse layer, several pixels on both sides of the block boundary can be adjusted simultaneously. For example, within one pixel of the block boundary, a slight brightness leveling operation can be performed on pixels on the inner side of the block, moderately lowering pixels above the local mean and moderately raising pixels below the local mean, while only increasing the contrast of pixels on the outer side of the block. This avoids creating harsh brightness steps while preserving the general outline of the edges between blocks. The polarity of the block boundary pulse layer can be used to indicate which side is brighter and which side is darker, thus guiding the leveling direction. Simply applying a uniform blur filter to the entire region can reduce block artifacts but also damage the true object boundaries. However, using directional adjustments based on block boundary pulse events can reduce block artifacts while maintaining the clarity of the main structure.

[0060] After accumulating all enhancement pulse events for each pixel, a pulse-reconstructed enhanced luminance image is obtained. To avoid introducing global luminance shift or contrast drift during the enhancement process, a lightweight correction can be performed on the entire frame image after generating the pulse-reconstructed enhanced frame. For example, in one implementation, the average luminance of the original reconstructed luminance image and the pulse-reconstructed enhanced frame can be statistically analyzed. When the difference between the two exceeds one gray level, a uniform shift is performed on all pixels in the pulse-reconstructed enhanced frame to bring the average luminance back to the level of the original reconstructed luminance image, reducing the perceived luminance drift during long playback. Furthermore, at the block or macroblock level, the local average luminance before and after enhancement can be compared. When the average luminance of a certain block shifts beyond a preset range relative to its original value (e.g., more than four gray levels), the enhancement result of that block is proportionally pulled back, making the local luminance distribution more closely match the prediction result at the encoding end, thereby suppressing loop accumulation errors.

[0061] Regarding the closed-loop and bitrate-coordinated adjustment within the loop, the generation of pulse reconstruction enhancement frames is not at a fixed intensity but is directly controlled by the current bitrate indicator. Combining the aforementioned bitrate indicator gradation from 0 to 100, different global enhancement coefficients and constraint strategies can be configured for each type of pulse layer when generating pulse reconstruction enhancement frames. When the bitrate indicator is in the range of 0 to 33, it indicates that the video is in an extremely low bitrate state. At this point, the encoder has already abandoned a large amount of detail information, and quantization noise and prediction errors are relatively large. If a significant enhancement is still applied to the texture pulse layer under such conditions, residual quantization errors will be mistaken for real texture, resulting in flickering and noise in large areas. Therefore, at low bitrates, the enhancement coefficient of the texture pulse layer can be significantly reduced, or even only texture enhancement can be retained in high-consistency and medium-consistency regions, while completely disabling texture enhancement in low-consistency regions. Simultaneously, the weights of the edge pulse layer and block boundary pulse layer can be increased, concentrating the limited enhancement capabilities on contours and block boundaries, thereby ensuring that the main structure of the image is as clear as possible, and that even if the texture is slightly simple, the subjective visual experience remains stable. At this point, the maximum allowable enhancement amplitude for a single pixel can be reduced, for example, limiting the total enhancement amplitude of a single pixel to no more than 8 gray levels, in order to prevent excessive ringing in areas with more compression artifacts.

[0062] When the bitrate indicator is between 34 and 66, it indicates that the video bitrate is at a medium level. At this point, the encoder retains some texture information, but there is still significant loss. In this situation, the enhancement coefficient of the texture pulse layer can be moderately increased, allowing texture events to play a role on more pixels, while keeping the weights of the edge pulse layer and block boundary pulse layer roughly the same as at low bitrates. Specifically, the upper limit of the single-pixel cumulative amplitude of texture enhancement can be increased to about 12 gray levels, allowing more texture events to be superimposed in medium consistency regions, while still strictly limiting them in low consistency regions. This allows for richer detail in areas without significantly amplifying noise.

[0063] When the bitrate indicator is between 67 and 100, the video bitrate is relatively ample, and the encoder has already preserved a significant amount of high-frequency information. In this scenario, the role of the pulse reconstruction enhancement frame is more focused on fine-tuning details and unifying the style, rather than significantly restoring missing information. Therefore, the enhancement coefficient of the texture pulse layer can be further increased, allowing texture events to play an almost complete role in high and medium consistency regions, and the total enhancement amplitude per pixel can be widened to about 16 gray levels. At the same time, the smoothing intensity of the block boundary pulse layer can be appropriately reduced to avoid over-processing the already relatively smooth block boundaries, which would affect detail contrast. Through this adjustment method in conjunction with the bitrate indicator, the pulse reconstruction enhancement frame can achieve a balance of "structure priority and moderate texture" at different bitrates, without the situation where a fixed enhancement strategy has excessively different effects at different bitrates.

[0064] In the closed-loop implementation, the pulse reconstruction enhancement frame is used not only to output the enhanced image to the display side, but also as a reference frame for subsequent frames. Specifically, after encoding a frame, the decoder first obtains the regular reconstructed frame, then generates a pulse reconstruction enhancement frame based on the enhanced pulse event stream, and replaces the original reference frame with this pulse reconstruction enhancement frame. This ensures that subsequent motion estimation, motion compensation, and prediction processes are all performed based on the enhanced reference. During playback, the decoder performs pulse reconstruction enhancement operations on each frame and updates the reference frame in the same completely consistent order, thus ensuring that the encoder and decoder always share the same set of enhanced reference images within the loop. This closed-loop approach is adopted to avoid the "enhancement drift" problem caused by only enhancing the decoder without updating the encoding side's reference frame. If the displayed image is only enhanced at the decoder, while the encoder still uses the unenhanced reconstructed frame as a reference, the image predicted by the encoder will gradually become disconnected from the content actually displayed by the decoder over time. Especially in motion compensation and multi-step prediction scenarios, this disconnect manifests as increasing blurriness with compression and the inability to accumulate enhancement effects. By updating the reference frame within the loop, the enhanced structural information of each frame becomes the starting point for the prediction of the next frame, allowing the enhancement effect to accumulate continuously over time.

[0065] To prevent errors from accumulating and amplifying over time due to closed-loop enhancement within the loop, a consistency check can be performed on the pulse reconstruction enhancement frame before replacing the reference frame. In one implementation, the difference image between the pulse reconstruction enhancement frame and the original reconstructed brightness image can be statistically analyzed. If a certain percentage of pixels (e.g., more than 10%) across the entire frame exhibit significant brightness changes, and these pixels are concentrated in low-consistency regions or regions with abnormally high pulse residual density, it indicates that the current parameter configuration may be too aggressive. In this case, a complete rollback operation can be performed on the pulse reconstruction enhancement frame, for example, reducing the overall enhancement amplitude to 50% to 70% of the original, and then using the rolled-back result as the final reference frame. This closed-loop strategy of "enhancement followed by verification and re-entry into the loop" can prevent misjudged enhancement results from continuously propagating to subsequent frames in extreme scenarios, ensuring the system remains stable during long-term operation.

[0066] In optional implementations, the generation of pulse reconstruction enhancement frames can also employ a layered reconstruction strategy. For example, a structural reconstruction map can be generated first using an edge pulse layer and a block boundary pulse layer to restore the main contours and block boundaries to a high degree of clarity. Then, a texture enhancement map reconstructed by a texture pulse layer is superimposed on the structural reconstruction map. The structural reconstruction map can be generated by redistributing edge pulse lines in a one-dimensional direction perpendicular to the edge, making the brightness gradient on both sides of the edge position more obvious but not excessively steep. The texture enhancement map can be generated by locally randomizing texture pulse events, making the texture exhibit slight randomness in space, close to the detail distribution of a natural image. Finally, the structural reconstruction map and the texture enhancement map are linearly synthesized according to weights related to the bitrate indicator to obtain the pulse reconstruction enhancement frame. In this layered approach, structural information and texture information can be more clearly separated, making it easier to fine-tune their respective enhancement strategies for different application scenarios and bitrate conditions.

[0067] Through the above steps, under the control of closed-loop and bitrate coordinated adjustment within the loop, the enhanced frame is reconstructed using the pulse generated by the enhanced pulse event stream. This not only strengthens the structure and refines the texture within a single frame, but also injects enhanced information into the entire prediction and compensation link by replacing the reference frame. This allows subsequent frames to be based on a clearer and more stable reference during encoding and decoding, thereby achieving a more stable and natural overall viewing effect under ultra-low bitrate conditions.

[0068] The present invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. An AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding, characterized in that, The method includes: Step 1: Obtain the reconstructed frame generated by the current decoder as visual input, and obtain the current bitrate indicator as control input; transform the reconstructed frame into three types of pulse event streams through the pulse coding layer of the spiking neural network, namely edge pulse layer, texture pulse layer and block boundary pulse layer; the pulse coding layer adopts leakage integration and trigger firing mechanism to fire pulse events simultaneously on two time axes: short time scale within the frame and long time scale across the frame, and sets the upper limit of firing density and reset hysteresis range according to the bitrate indicator, thereby generating an initial pulse event stream containing event timestamps and polarities and its pulse timing consistency map; Step 2: Under the constraints of the pulse timing consistency map, pulse consistency re-estimation and asynchronous residual back-injection based on interleaved phase kernel are performed. This process specifically includes: interleaving the current initial pulse event stream and the reference pulse event stream on short time scales within the frame and long time scales across frames to form an interleaved phase window; performing threshold voting and polarity determination on pulse clusters in each window to suppress isolated pulses and confirm the dominant firing timing; performing peak offset compensation on pulse residuals formed by unmatched pulses according to adjacent time slices, and splitting them through structure priority branches and texture priority branches, and then merging them into an enhanced pulse event stream according to pulse timing consistency weights; at the same time, performing threshold drift self-correction based on the inter-frame dominant firing timing difference to maintain the consistency and stability of long time scales, thereby forming an enhanced pulse event stream that can enhance structure and refine texture at ultra-low bit rates. Step 3: Under the control of closed-loop and rate-coordinated adjustment within the loop, a pulse reconstruction enhancement frame is generated using the enhanced pulse event stream, and the reference frame is replaced by this pulse reconstruction enhancement frame. The generation of the pulse reconstruction enhancement frame is based on the reconstructed luminance image of the current frame from the decoder. For each pixel location, the initial luminance value of that pixel in the reconstructed luminance image is read, and a pulse event sequence corresponding to that pixel location is extracted from the enhanced pulse event stream. The pulse event sequence includes three types of events: edge pulse layer, texture pulse layer, and block boundary pulse layer, arranged in the order of short timestamps within the frame. Whenever a pulse event is encountered, the luminance gain or luminance attenuation for the current pixel is calculated based on the pulse layer and polarity to which the event belongs, combined with the pixel's level in the consistency level map and the distance between the event timestamp and the dominant release sequence. After accumulating all enhanced pulse events for each pixel, the luminance image after pulse reconstruction enhancement is obtained.

2. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 1, characterized in that, In step 1, the reconstructed frames are first uniformly converted into single-channel images of the luminance component, and linearly stretched to a fixed integer range according to the pixel value range while maintaining the spatial resolution. The image width and height are recorded to determine the processing grid. Two time axes are established: a short time scale within the frame and a long time scale across the frame. The short time scale within the frame is divided into several equally spaced discrete scales within a frame period, and the number of these discrete scales is fixed at 1000. The long time scale across the frame spans the cumulative window of consecutive frames, and its window length is fixed at 8 frames and a sliding update strategy is adopted. The bitrate indicator is normalized to a discrete level from 0 to 100.

3. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 2, characterized in that, In step 1, the rules for setting the upper limit of the release density and the reset hysteresis range based on the bitrate indicator are as follows: if the bitrate indicator is in the range of 0 to 33, the upper limit of the release density is set to no more than 100,000 pulse events per megapixel per frame, and the reset hysteresis range is 6 to 10 discrete scales; if the bitrate indicator is in the range of 34 to 66, the upper limit of the release density is set to no more than 250,000 pulse events per megapixel per frame, and the reset hysteresis range is 4 to 6 discrete scales; if the bitrate indicator is in the range of 67 to 100, the upper limit of the release density is set to no more than 500,000 pulse events per megapixel per frame, and the reset hysteresis range is 2 to 4 discrete scales.

4. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 3, characterized in that, In step 1, within a long time-stamped window spanning multiple frames, the consistency of the pulse time series for each pixel is evaluated as follows: The pixel is clustered based on the dominant firing scales of the last 8 frames. If a cluster appears in at least 5 frames and the difference between firing scales does not exceed 3 discrete scales, the pixel is classified as having high consistency, and the dominant firing scale corresponding to that cluster is recorded as the dominant firing time sequence for that pixel. If a cluster appears in 3 or 4 frames and the difference between firing scales does not exceed 5 discrete scales, it is classified as having medium consistency. All other cases are classified as having low consistency. The aforementioned consistency levels are encoded as integers in a raster with the same resolution as the image to form a consistency level map. Simultaneously, the dominant firing time sequence is encoded as integer values ​​of short time-stamped scales to form a time-series index map. Finally, the consistency level map and the time-series index map together constitute a pulse time-series consistency map.

5. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 4, characterized in that, In step 2, the process of forming an interleaved phase window includes: interleaving the current initial pulse event stream and the reference pulse event stream pixel-by-pixel on the intra-frame short timescale and cross-frame long timescale; using the timing index map in the pulse timing consistency map as a reference, taking the three short timescale scales before and after the dominant release timing of the current frame and the three short timescale scales before and after the dominant release timing of the reference frame at each pixel, arranging them in an interleaved order of the current scale and the reference scale to form an interleaved phase window containing 12 short timescale scales; when any end exceeds the boundary, it is repeatedly filled with boundary scales to keep the window length at 12 short timescale scales; within each interleaved phase window, the pulse events of the edge pulse layer, texture pulse layer, and block boundary pulse layer are merged according to the following rule: when the short timescale scale interval of two adjacent pulse events in the same pulse layer does not exceed 2 short timescale scales and the polarity is the same, they are merged into the same pulse cluster; otherwise, a new cluster is created; only containing 1 A pulse cluster containing two or more pulse events is marked as an isolated pulse cluster, and a pulse cluster containing two or more pulse events is marked as a valid pulse cluster.

6. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 5, characterized in that, In step 2, the process of threshold voting and polarity determination for pulse clusters in each window includes: within each interleaved phase window, threshold voting and polarity determination are performed on all pulse clusters of each pulse layer, according to the following rules: the base number of votes for each pulse cluster is equal to the number of pulse events within that cluster; when the difference between the center short timescale of the pulse cluster and the dominant emission timing of that pixel does not exceed 3 short timescales and the consistency level is high, 2 votes are added to that cluster; when it is medium, 1 vote is added; when it is low, no votes are added; the pulse cluster with the highest number of votes is recorded as the winning cluster, and its polarity is used as the polarity determination result of that pulse layer within the interleaved phase window, and its center short timescale is recorded as the candidate dominant emission timing; if there is a tie in the number of votes, the pulse cluster with the earlier center short timescale is selected; isolated pulse clusters are directly eliminated when the consistency level is low; when the consistency level is medium, they are retained only when isolated pulse clusters of the same polarity also appear in adjacent interleaved phase windows; when the consistency level is high, they are retained but their votes do not receive additional votes.

7. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 6, characterized in that, In step 2, the process of performing peak offset compensation on the pulse residuals formed by unmatched pulses according to adjacent time slices, and then splitting them through the structure-priority branch and the texture-priority branch, includes: performing peak offset compensation on each pulse residual based on its distance from the candidate dominant release timing, wherein: when the difference between the center short time scale and the candidate dominant release timing does not exceed 2 short time scales, the residual is moved as a whole to the side away from the candidate dominant release timing by 1 to 2 short time scales, and the minimum movement amplitude that does not overlap with the confirmed winning cluster is preferred; when both sides are idle, a movement of 2 short time scales is selected; the default movement amplitude for edge pulse layers and block boundary pulse layers is 1 short time scale, and texture pulse layers are allowed to move up to 2 short time scales. A short timescale; when there is no free short timescale within the interleaved phase window, the residual is allowed to be moved to the edge short timescale of the adjacent interleaved phase window, but it must not overlap with the paired cluster on the reference side; the pulse residual after peak compensation is split into pulse layers, specifically: the edge pulse layer and the block boundary pulse layer enter the structure priority branch, and the texture pulse layer enters the texture priority branch; the structure priority branch retains at most 1 pulse residual in each interleaved phase window of each pixel, and prioritizes retaining the pulse residual with the smallest timing difference with the candidate dominant release; the texture priority branch retains at most 2 pulse residuals in each interleaved phase window of each pixel, and sets an upper limit on the number of pulse events in each pulse residual. When the upper limit is exceeded, the last pulse event is discarded evenly in chronological order until the upper limit is met.

8. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 7, characterized in that, In step 2, the process of merging pulse events into an enhanced pulse event stream according to pulse timing consistency weights includes: merging the confirmed winning clusters of each pulse layer with the pulse residuals output from the two branches to form a candidate set for the enhanced pulse event stream; when multiple pulse events conflict at the same short timescale of the same pixel, conflict resolution is performed. This resolution is achieved by mapping consistency levels to merging weights, where high consistency corresponds to the third weight, medium consistency corresponds to the second weight, and low consistency corresponds to the first weight, with priority given to retaining pulse events with larger merging weights; when the merging weights are the same, selection is based on pulse layer priority, which is edge pulse layers > block boundary pulse layers > texture pulse layers; if the priorities are still the same, pulse events with earlier generation times are retained; the pulse events after the aforementioned resolution are written into the enhanced pulse event stream in ascending order of short timescale.

9. The AI-based in-loop enhancement method for ultra-low bitrate video encoding and decoding as described in claim 8, characterized in that, In step 2, the process of performing threshold drift self-correction based on the inter-frame dominant transmission timing difference includes: within the cross-frame long time-stamp window, the dominant transmission timing of the current frame and the reference frame are statistically analyzed for each pixel, and the inter-frame dominant transmission timing difference is calculated; when the dominant transmission timing difference of the same pixel is greater than or equal to 5 short time-stamps in 3 consecutive frames and the consistency level is high, the retention threshold of the threshold voting in this step is reduced by 1 level to improve the tolerance for delay changes; when the dominant transmission timing difference of the same pixel is less than or equal to 2 short time-stamps in 3 consecutive frames and the pixel no longer generates pulse residuals in the interleaved phase window, the retention threshold of the threshold voting in this step is increased by 1 level to reduce unnecessary retention; the single-frame adjustment range of threshold drift self-correction does not exceed 1 level; when the consistency level is low or the pulse residual density per unit area exceeds 50,000 pulse events per megapixel per frame, the threshold drift self-correction is paused.