Super-resolution reconstruction system for ultra-high definition video based on heterogeneous computing architecture
Patent Information
- Application Number
- CN202611008570.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-08
AI Technical Summary
前者会导致图像细节模糊,且无法逆转已发生的相位错位;后者则会带来高达三倍以上的显存占用与内部总线带宽消耗,破坏了异构计算流水线的流式处理效能,难以满足高帧率超高清视频实时输出的时延指标
[0007]本发明的有益效果在于:本发明针对高分辨率与宽色域条件下的超高清多媒体流处理需求,通过将亮度平面与色度平面的处理管线在底层缓存硬件中解耦,并利用高频亮度图像重建后产生的边缘方向梯度来同步引导色度平面的空间插值计算,消除了由异构体系中分层存储驻留、独立总线搬运以及显示栅格时序差异所引发的半像素级色度相位错位缺陷;进而改善了高饱和度细线及运动边缘区域的色度溢出与彩晕现象,保证了重构图像色彩边界与亮度边缘的对齐,提升了最终视频序列的显示逼真度,满足了流式数据高通量实时处理的延迟要求。
Smart Images

Figure CN122510094B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of super-resolution reconstruction technology, and more specifically, to an ultra-high-definition video super-resolution reconstruction system based on a heterogeneous computing architecture. Background Technology
[0002] In modern multimedia processing systems, real-time super-resolution reconstruction of ultra-high-definition video (such as 8K) typically relies on a heterogeneous computing architecture with multiple processing units working together. Limited by bus bandwidth and chip power consumption, the video stream output after front-end decoding is often stored and transmitted in a YCbCr4:2:0 half-sampling format. In conventional super-resolution pipelines, to maximize the utilization of the specific computing power of each hardware accelerator, a plane-based processing strategy is usually adopted: the luminance plane, containing the main high-frequency details, is preferentially assigned to a neural network processor, which excels at dense matrix operations, for local block-level residual reconstruction; while the chrominance plane, with half the data volume, is handled by a graphics processor, which excels at spatial resampling, for subsequent scaling and transport.
[0003] This heterogeneous allocation mechanism exposes significant image quality shortcomings under ultra-high pixel density and wide color gamut (such as BT.2020) display conditions. Due to objective differences between luminance and chrominance samples in memory residency paths, processor cache capacity allocation, direct memory access (DMA) transfer cycles, and the final output raster scan timing, sharp luminance edges that have undergone super-resolution reconstruction in the early stages and chrominance edges that have been amplified by bilinear interpolation by the graphics processor in the later stages will experience spatial phase deviation in the target high-resolution grid. This deviation is particularly noticeable in the edge areas of thin lines, text, or fast-moving objects with low brightness and high saturation, and visually manifests as visible color smudges, chrominance trailing, and edge color breaks.
[0004] Most existing conventional avoidance solutions attempt to perform global color smoothing through post-processing filters, or force the YCbCr format to be pre-converted to full-resolution RGB format before being fed into the neural network for joint reconstruction. The former leads to blurred image details and cannot reverse the phase misalignment that has already occurred; the latter results in more than three times the memory usage and internal bus bandwidth consumption, which disrupts the streaming processing performance of heterogeneous computing pipelines and makes it difficult to meet the latency requirements for real-time output of high frame rate ultra-high-definition video. Summary of the Invention
[0005] This invention provides an ultra-high-definition video super-resolution reconstruction system based on a heterogeneous computing architecture, which solves the technical problems mentioned in the background art.
[0006] This invention provides an ultra-high-definition video super-resolution reconstruction system based on a heterogeneous computing architecture. It is applied to a heterogeneous computing architecture including a main processor, graphics processor, neural network processor, reconfigurable line cache accelerator, and output interface, and is configured to execute: The system captures an initial image frame sequence carrying a timestamp through the input interface, parses the metadata, divides the input samples of the initial image frame sequence into a luminance plane and a chrominance plane, and writes them into a fixed-point video buffer. The luminance plane is fed into the neural network processor and the reconfigurable line buffer accelerator cache, and the chroma plane is fed into the video frame cache of the graphics processor. The neural network processor is used to perform residual reconstruction on the brightness plane to generate a brightness plane with target resolution, and the brightness edge direction feature map is extracted and generated. Using the graphics processor, and combining the brightness edge direction feature map, edge phase alignment calculation is performed on the chroma plane to obtain the target resolution chroma plane; The target resolution luminance plane and the target resolution chrominance plane are fused and written into a unified frame buffer to form an intermediate image frame containing complete luminance and chrominance samples; Perform color space conversion and output timing synchronization on the intermediate image frames to generate a target frame rate output frame buffer with an output timestamp; The target frame rate output frame buffer is encapsulated in the order of the output timestamps and output as target image sequence data through the output interface.
[0007] The beneficial effects of this invention are as follows: Addressing the ultra-high-definition multimedia streaming processing requirements under high resolution and wide color gamut conditions, this invention decouples the processing pipelines of the luminance plane and chrominance plane in the underlying cache hardware. It utilizes the edge direction gradient generated after high-frequency luminance image reconstruction to synchronously guide the spatial interpolation calculation of the chrominance plane, eliminating the half-pixel-level chrominance phase misalignment defect caused by hierarchical storage, independent bus transport, and display raster timing differences in heterogeneous systems. Furthermore, it improves chrominance overflow and halo phenomena in high-saturation fine lines and moving edge areas, ensuring the alignment of the reconstructed image's color boundaries with luminance edges, enhancing the display realism of the final video sequence, and meeting the latency requirements of high-throughput real-time streaming data processing. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the chromaticity edge phase alignment calculation mechanism of the present invention. Detailed Implementation
[0009] The technical solution of the present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. The following embodiments are used to illustrate the implementation of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can make equivalent substitutions for processor type, cache structure, network structure, interpolation kernel function, parameter values, output interface type, or video format without departing from the concept of the present invention, and all such substitutions should be understood as falling within the scope covered by the disclosure of the present invention.
[0010] Example 1: like Figure 1 As shown, in one embodiment, the ultra-high-definition video super-resolution reconstruction system based on a heterogeneous computing architecture includes an input interface 100, a main processor 200, a graphics processor 300, a neural network processor 400, a reconfigurable line buffer accelerator 500, a fixed-point video buffer 600, a circular buffer 700, a unified frame buffer 800, a synchronization arbitration unit 900, and an output interface 1000. The input interface 100 receives an initial image frame sequence carrying timestamps; the main processor 200 parses metadata, generates working descriptors, establishes a direct memory access list, and schedules each hardware unit; the neural network processor 400 performs super-resolution residual reconstruction on the luminance plane and generates a luminance edge direction feature map; the graphics processor 300 performs edge phase alignment calculations on the chrominance plane based on the luminance edge direction feature map; the reconfigurable line buffer accelerator 500 performs luminance stripe caching, sideband splicing, on-chip tensor organization, and luminance output handling; and the synchronization arbitration unit 900 matches the luminance branch and chrominance branch based on the input frame timestamp and stripe number.
[0011] Input interface 100 can be an HDMI receiver interface, a DisplayPort receiver interface, a MIPICSI interface, a PCIe video capture interface, an on-chip video decoder output interface, or other interfaces capable of outputting continuous image frame sequences. Output interface 1000 can be an HDMI transmitter interface, a DisplayPort transmitter interface, a MIPICSI interface, a PCIe display output interface, a network video output interface, or an on-chip display controller interface. The above interface types do not affect the core processing logic of this invention, which utilizes the luminance edge direction to guide chroma phase alignment.
[0012] In this embodiment, the luminance plane and chrominance plane are processed using separate paths. The luminance plane forms a luminance reconstruction path via the reconfigurable line buffer accelerator 500 and the neural network processor 400; the chrominance plane forms a chrominance phase alignment path via the graphics processor 300; the two converge at a unified frame buffer 800. Compared to the approach of first converting the YCbCr format image to full-resolution RGB and then uniformly sending it to the neural network processor, this embodiment maintains the separate-plane storage format of luminance and chrominance data, which can reduce video memory usage and internal bus traffic, while compensating for the chrominance phase deviation caused by separate-plane processing through luminance edge direction feature maps.
[0013] The main processor 200 is connected to the graphics processor 300, neural network processor 400, and reconfigurable line cache accelerator 500 via an on-chip bus, shared memory, direct memory access controller, or on-chip network. The fixed-point video buffer 600, circular buffer 700, and unified frame buffer 800 can reside in the same physical memory, or they can be located in main memory, on-chip static random access memory, graphics memory, or on-chip shared cache, respectively. For those skilled in the art, as long as the addressable relationships of the luminance plane, chroma plane, working descriptor, timestamp, and stripe number can be maintained during processing, the data flow of this embodiment can be achieved.
[0014] The synchronous arbitration unit 900 is integrated into the front end of the on-chip interconnect memory controller. It uses a pure hardware state machine to implement the matching logic. Internally, it maintains a queue of luminance entries to be matched and a queue of chrominance entries to be matched with a depth of 8. Each entry stores the frame timestamp, stripe number, and physical base address of the corresponding data. When the luminance branch or chrominance branch completes data writing and sends a write completion signal, the corresponding entry is written to the tail of the corresponding queue. The state machine successively retrieves the first entries of the two queues and performs a full match comparison between the frame timestamp and the stripe number. If a match is successful, the merge completion flag of the unified frame buffer 800 is triggered, and the corresponding data merge descriptor is moved into the output merge queue. If the first entry of the luminance queue waits for more than 16 bus clock cycles without matching the corresponding chrominance entry, the timeout discard logic is triggered, the luminance entry is cleared, and the abnormal status register is reported. If the chrominance entry arrives first, it waits for the luminance entry. The waiting threshold is the same as that of the luminance branch. In non-matching scenarios, data write submission is not performed.
[0015] Example 2: The main processor 200 receives the initial image frame sequence in real time through the input interface 100. Each initial image frame carries at least an input frame timestamp and spatial metadata. The input frame timestamp can be a microsecond-level time recording marker, a hardware counter marker, a frame sequence number marker, or a synchronization marker derived from the display clock; under continuous input conditions, the input frame timestamp of the subsequent frame monotonically increases relative to the previous frame. The spatial metadata includes at least the input image resolution, color conversion matrix parameters, chroma sampling position parameters, transfer function characteristics, pixel aspect ratio, and image cropping window coordinates. In another embodiment, the spatial metadata may also include a source color gamut identifier, dynamic range identifier, interlaced or progressive scan identifier, and an input frame valid area identifier.
[0016] The main processor 200 divides the input samples in the initial image frame into luminance samples and chrominance samples. Luminance samples are written to the luminance plane in the fixed-point video buffer 600 in a two-dimensional raster order; chrominance samples are written to the interleaved chrominance plane in an interleaved arrangement of blue and red components. For YCbCr4:2:0 input, the chrominance plane is half-sampled relative to the luminance plane in both the horizontal and vertical directions; for YCbCr4:2:2 input, the chrominance plane is half-sampled relative to the luminance plane in the horizontal direction; for YCbCr4:4:4 input, the chrominance plane can maintain the same sampling density as the luminance plane. All of the above sampling formats can be identified through the sampling ratio field in the working descriptor.
[0017] To facilitate fixed-point hardware computation, the main processor 200 performs bit-width normalization on the input samples. Let the original bit width of the input samples be... The target effective bit width is In a preferred embodiment, Take 10 bits; in other embodiments, It can be 8 bits, 12 bits, or 16 bits. When When the main processor 200 performs a left shift and zero-padding on the high-order bits of the input sample; when At that time, the main processor 200 retains the high-order bits. The lower-order bits are discarded. The normalized luminance and chrominance samples can be stored left-aligned in a 16-bit unsigned fixed-point storage container, or in a fixed-point container that matches the target bit width.
[0018] The main processor 200 generates a working descriptor for each frame. The working descriptor includes at least the frame identifier, luma plane base address, chroma plane base address, luma plane row step size, chroma plane row step size, input width, input height, target width, target height, chroma sampling format, chroma sampling center offset, input frame timestamp, circular buffer node number, and effective bit width. The luma plane row step size and chroma plane row step size represent the number of consecutive bytes occupied in memory for a row of samples and its end-padding bytes in the corresponding plane. The chroma sampling center offset represents the horizontal and vertical offset of the chroma sampling point relative to the geometric center of the corresponding luma sampling unit.
[0019] In one embodiment, the chroma sampling center offset is expressed as: ,in Indicates horizontal offset. This indicates the vertical offset. The main processor 200 can configure the chroma sampling position (center sampling, left-aligned sampling, or other chroma sampling positions defined by video standards) via settings. and This enables the graphics processor 300 to obtain the correct chroma sampling geometry center in subsequent spatial mapping.
[0020] The main processor 200 writes the luma plane, interleaved chroma plane, and corresponding working descriptors into a circular buffer 700. The circular buffer 700 is managed using a circular queue and can hold at least the luma plane, chroma plane, and working descriptors of the current frame, the previous frame, and the frame before that. In a preferred embodiment, the circular buffer 700 stores 3 to 8 frames of data; after the current frame is written, the corresponding nodes of adjacent historical frames remain valid and unoverwritten. When processing the current luma stripe, the reconfigurable line buffer accelerator 500 can read the luma stripe at the same spatial location in adjacent historical frames based on the frame identifier, stripe number, and node offset, thereby providing historical luma information for temporal super-resolution reconstruction.
[0021] The system uses the video pixel clock received by input interface 100 as the timestamp reference clock domain. All timestamps are generated based on a free-running counter in this clock domain, with a counter width of 32 bits and a monotonically increasing value. The main processor 200, reconfigurable line buffer accelerator 500, and neural network processor 400 are located in the system bus clock domain, while the graphics processor 300 is located in the graphics engine clock domain. When transferring timestamps across clock domains, a two-level register timing synchronization process is used. The first-level register acquires the input signal, and the second-level register outputs a stable signal to eliminate metastability risks. The timestamp values in different clock domains maintain the original count value of the reference clock domain without frequency conversion; only timing ensures the correctness of value transmission. When the counter overflows, it automatically wraps back to 0. Since the frame timestamps monotonically increase and the queue depth does not exceed 8 frames, the order can still be correctly identified by comparing the values after wrapping, and no matching errors will occur.
[0022] Example 3: The main processor 200 reads the working descriptors from the circular buffer 700 and establishes a luminance direct memory access list and a chrominance direct memory access list based on the working descriptors. Each direct memory access list consists of multiple transfer descriptors. Each transfer descriptor includes at least the source address, destination address, transfer length, source row step size, destination row step size, next descriptor address, stripe number, frame timestamp, and completion flag. Multiple transfer descriptors are connected by the next descriptor address according to physical address order or stripe scan order, enabling the direct memory access controller to complete continuous stripe transfer without the main processor 200's line-by-line intervention.
[0023] A direct memory access list for luminance is used to move the luminance plane to the reconfigurable line cache accelerator 500. The main processor 200 divides the luminance plane vertically into multiple luminance stripes according to a set number of rows processed per cycle. Each luminance strip includes multiple rows of continuous luminance samples, and the strip height can be determined based on the input tensor height of the neural network processor 400, the on-chip cache capacity, the maximum receptive field per side, and the bus burst transfer length. In one embodiment, the strip height is 16, 32, 64, or 128 rows; in another embodiment, the strip height can be dynamically adjusted based on the input resolution and the target magnification.
[0024] The chroma direct memory access list is used to move the chroma plane to the video frame buffer or texture sampling buffer of the graphics processor 300. During the movement, the blue and red components are not de-interleaved, nor is the chroma plane pre-interpolated and magnified to full resolution. The graphics processor 300 performs fractional coordinate addressing and interpolation sampling on the chroma plane based on the texture descriptor. The texture descriptor includes at least the chroma plane base address, chroma row step size, chroma plane width, chroma plane height, sampling format, blue and red component interleaving order, horizontal sampling ratio, vertical sampling ratio, chroma sample center offset, effective bit width, and input frame timestamp.
[0025] When the main processor 200 dispatches a luminance reconstruction task to the neural network processor 400, it writes the input frame timestamp and stripe number into the luminance task control block; when it dispatches a chroma phase alignment task to the graphics processor 300, it writes the same input frame timestamp and corresponding stripe number into the chroma task control block. The synchronization arbitration unit 900 compares the input frame timestamps and stripe numbers at the luminance branch output and the chroma branch output to determine whether they belong to the same input frame and the same target output stripe. Only when they match will the system allow subsequent fusion writing to the unified frame buffer 800.
[0026] The main processor 200 dispatches tasks to the neural network processor 400 and the graphics processor 300 through a task control block queue in shared memory. Each task control block occupies 64 bytes of contiguous physical address space, and its fields include frame timestamp, stripe number, input data base address, output data base address, working descriptor pointer, and completion flag. The main processor 200 writes the task control block to the tail of the corresponding processor's task queue and sends an interrupt signal to the corresponding processor to notify that the task has arrived by writing to the trigger register. After the neural network processor 400 and the graphics processor 300 complete their current tasks, they set the completion flag in the corresponding task control block to 1 and send a completion interrupt to the main processor 200. Task scheduling adopts a stripe-level pipelined strategy, with a fixed priority: luminance residual reconstruction tasks are higher than chrominance edge phase alignment tasks. Within the same frame, tasks are dispatched sequentially according to stripe number from smallest to largest. Processors establish a one-to-one correspondence between tasks through frame timestamps and stripe numbers, without the need for additional interaction signals.
[0027] Direct memory access (DMI) transfers use physical address addressing. The main processor 200 is responsible for translating virtual addresses into physical addresses and writing them to the descriptor. The DMI controller does not have address translation capabilities; all transfer addresses are physically contiguous. The row step size refers to the total number of bytes corresponding to a row of data, including the number of valid pixel data bytes and the number of row end padding bytes. The row end padding bytes are automatically calculated based on address alignment requirements, ensuring that the starting address of each row meets the 32-byte alignment requirement. The padding data is not involved in the actual calculation and is only used for address alignment. The transfer triggering method uses a linked list triggering mode. The main processor 200 writes the first address of the linked list to the linked list start address register of the DMI controller, and then writes to the start register to trigger the transfer. The controller automatically executes the transfer of each descriptor item sequentially according to the linked list order, without requiring the main processor 200 to intervene row by row. After each descriptor item is transferred, the controller automatically sets the corresponding completion flag to 1. After all descriptor items in the entire linked list have been transferred, the controller sends a completion interrupt to the main processor 200. If a bus error, address out-of-bounds error, or other exception occurs during the transfer, the controller automatically terminates the transfer, sets the exception flag, and sends an exception interrupt, which is then handled by the main processor 200.
[0028] Example 4: The reconfigurable line buffer accelerator 500 includes a line buffer array, an address mapping unit, a sideband splicing unit, a direct memory access interface, and a status register. The line buffer array is used to store the current luminance stripe and adjacent line data; the address mapping unit is used to calculate the source address and destination address based on the working descriptor and stripe number; the sideband splicing unit is used to organize the current luminance stripe, adjacent historical frame luminance stripes at the same position, and stripe sidebands into an input tensor that can be read by the neural network processor 400; the status register is used to record the current stripe number, frame timestamp, transfer completion status, and abnormal status.
[0029] The reconfigurable line cache accelerator 500's reconfigurability is achieved through register parameter configuration, requiring no modification to the hardware circuit logic. The configuration interface uses the standard APB slave interface, and the configurable parameters include four categories: line cache depth, sideband width, tensor channel count, and stripe height. Each level corresponds to a set of preset hardware address mappings and data concatenation rules. During system initialization, the main processor 200 calculates the optimal configuration values based on the current input resolution, magnification, and residual reconstruction network structure parameters, and writes them to the corresponding configuration register. The hardware automatically adjusts the line cache read / write address mapping logic, the concatenation range of the sideband concatenation units, and the tensor organization format based on the register values, completing the static reconstruction. Inter-frame reconfiguration and reconstruction can be performed based on changes in input parameters. During single-frame processing, the configuration parameters remain fixed and do not dynamically adjust. In the 2x super-resolution scenario, the default configuration is a strip height of 32 lines, a sideband width of 17 input pixels, and 2 channels. In the 4x super-resolution scenario, the default configuration is a strip height of 16 lines, a sideband width of 17 input pixels, and 2 channels. All configuration levels ensure that hardware resources match processing requirements and that there will be no buffer overflow.
[0030] The reconfigurable line cache accelerator 500 processes the first... When reading a luminance stripe, the current luminance stripe of the current frame, the luminance stripes at the same spatial position in adjacent historical frames, and the stripe sidebands are read. The stripe sidebands include several outward extensions above and below the current luminance stripe, as well as the corresponding outward extensions above and below the stripes at the same position in adjacent historical frames. The sideband width is determined based on the maximum receptive field of the network structure being run by the neural network processor 400. For example, when the maximum receptive field of the network structure is... When the input pixel is 1, the strip sideband can include at least the upper part. Rows and below Row brightness sample.
[0031] The reconfigurable line buffer accelerator 500 aligns the current frame luminance strip, historical frame luminance strips, and stripe sidebands in memory according to temporal and spatial coordinate order to generate an input tensor. The input tensor may include the current frame luminance channel, the previous frame luminance channel, the frame before that luminance channel, and the corresponding sideband channel; or it may only include the current frame and the previous frame luminance channels. This invention does not limit the specific number of historical frames; as long as the input tensor can provide the spatial and temporal neighborhood information of the current stripe, it can meet the luminance reconstruction requirements of this invention.
[0032] Input tensor A four-dimensional arrangement is adopted, with the dimensions in the order of batch, channel, height, and width. The batch dimension is fixed at 1. The channel dimension equals the sum of the number of historical frames participating in residual reconstruction and the number of current frames. In this embodiment, two frames of data, the current frame and the previous frame, are used, corresponding to 2 channels. The height dimension equals the sum of the height of the input luminance strip and the width of the upper and lower sidebands. The width dimension equals the number of horizontal pixels in the input luminance plane. The memory arrangement adopts row-major contiguous storage. Each pixel data is stored in 16-bit unsigned fixed-point format, with the row start address aligned to 32 bytes. Multi-channel data is arranged contiguously in channel order, and data from different channels at the same spatial location are adjacent in memory. During stitching, the luminance strip and upper and lower sideband data of the current frame are first written into the memory area corresponding to the first channel in row order. Then, the strip and sideband data of the same position of the previous frame are written into the corresponding area of the second channel in the same order. After stitching, the starting address of the tensor data is mapped to the base address of the input tensor buffer of the neural network processor 400 for direct reading by the neural network processor 400.
[0033] In one embodiment, the neural network processor 400 internally includes an input tensor reading unit, a convolution matrix multiply-accumulate unit, a nonlinear activation unit, a residual block unit, an upsampling unit, and an output reconstruction unit. The input tensor reading unit reads input tensors from the on-chip shared address space of the reconfigurable line cache accelerator 500; the convolution matrix multiply-accumulate unit performs convolution calculations on the input tensors; the nonlinear activation unit performs activation mapping on the convolution results; the residual block unit extracts high-frequency detail residuals; the upsampling unit maps low-resolution brightness features to a target resolution grid; and the output reconstruction unit fuses the base magnified brightness features and residual features point-by-point.
[0034] The network weights of the neural network processor 400 can be fixed weights that are stored in read-only memory, on-chip static random access memory, or external memory after offline training, or they can be a set of parameters loaded during system initialization. The network structure can be a residual convolutional network, a lightweight convolutional network, a depthwise separable convolutional network, a subpixel rearrangement upsampling network, or other super-resolution networks suitable for execution on the neural network processor 400. As long as the network can output a target resolution brightness plane and preserve local brightness edge direction information, it can be used in this embodiment.
[0035] The residual convolutional network used in this system consists of eight cascaded residual blocks and one subpixel upsampling layer. The input is a single-channel luminance tensor, and the output is a single-channel luminance plane at the target resolution. Each residual block contains two 3×3 convolutional layers, each with 64 channels, a stride of 1, and zero-padding. The ReLU activation function is used. The residual connections directly add the block input and output pixel by pixel. The subpixel upsampling layer is located at the end of the network. At a magnification of 2, the number of channels expands to 4; at a magnification of 4, the number of channels expands to 16. Pixel rearrangement maps the channel dimension to the spatial dimension, achieving size magnification. The network's maximum single-sided receptive field is [not specified]. The calculation formula is: in In this embodiment, the number of residual blocks is... Setting the value to 8 corresponds to a maximum single-sided receptive field of 17 input pixels. The width of the strip edge is set to be equal to this maximum single-sided receptive field value to ensure complete coverage of the receptive field in the strip edge area.
[0036] Let the target resolution brightness plane be The basic amplified brightness characteristics are The residual characteristics are In one embodiment, the neural network processor 400 generates the target resolution brightness plane in the following manner: in, and These represent the horizontal and vertical coordinates in the target resolution grid, respectively. The neural network processor 400 can output in stripes and bind the output brightness stripes to the corresponding frame timestamp, strip number, and target coordinate origin.
[0037] The neural network processor 400 extracts the horizontal and vertical gradients on the target resolution brightness plane. In one embodiment, central difference calculation is used: At the boundary of the target image, the neural network processor 400 can obtain the values of adjacent pixels using boundary copying, mirror expansion, zero padding, or single-sided difference. This boundary processing method only affects a small number of pixels at the boundary and does not change the overall mechanism of this invention, which guides chromatic phase alignment through the luminance edge direction.
[0038] The neural network processor 400 generates a brightness edge direction feature map based on the horizontal and vertical gradients. Let the minimum stability constant be... Then the normalized denominator is: Horizontal component in brightness edge direction feature map and vertical components They are respectively: in, This is used to avoid division by zero anomalies in flat regions. In a preferred embodiment, In other embodiments, Can be taken to A constant within a certain range. The neural network processor 400 will... and Together they form a brightness edge direction feature map, and make it share the same target coordinate origin, the same frame timestamp, and the same strip number with the target resolution brightness plane.
[0039] The horizontal and vertical gradients of all pixels within a strip are calculated using the center difference method. For the boundary pixels of the top and bottom rows of the current strip, the gradient calculation reuses the adjacent row data from the strip's sidebands, requiring no additional padding. The brightness edge direction feature map output by each strip includes the strip's main region and the overlapping regions of the top and bottom rows. The width of the overlapping region is equal to the range of neighboring pixels required for gradient calculation, and the overlapping region data of adjacent strips are completely identical. When stitching together the entire frame feature map, the last row of the main region of the previous strip is arranged adjacently to the first row of the main region of the next strip, discarding overlapping duplicate rows to ensure the spatial continuity of the entire frame feature map. Because the sideband data ensures the integrity of the boundary pixel gradient calculation, the feature values at the strip boundaries are completely consistent with the continuous calculation results of the entire frame, and there will be no abrupt changes or artifacts in the strip boundary features.
[0040] The brightness edge direction feature map contains two channels: a horizontal component and a vertical component. Each component is stored in an 8-bit signed fixed-point format, with a value range of -128 to 127, corresponding to the normalized interval of -1 to 1, and a quantization scaling factor of 128. After generating the feature map, the neural network processor 400 writes it to the feature output buffer of the reconfigurable line cache accelerator 500 in stripe units. The buffer adopts a double-buffered structure. When the current stripe feature is written to the first buffer, the previous stripe feature can be moved from the second buffer to the unified feature cache area in main memory via direct memory access. The graphics processor 300 reads the feature cache area data in main memory through the texture sampling unit. The texture format is configured as a dual-channel 8-bit signed integer, and the texture coordinates correspond one-to-one with the target resolution space. The graphics processor 300 can directly address and read the edge direction component at the corresponding position through the target pixel coordinates without additional address translation.
[0041] Example 5: like Figure 2 As shown, the graphics processor 300 reads the target resolution pixel position, texture descriptor, chroma sampling center offset, and luminance edge direction feature map, and performs edge phase alignment calculation on the chroma plane. The graphics processor 300 can implement this calculation through a programmable shader, computation shader, texture sampling unit, vector operation unit, or dedicated image processing core.
[0042] For each pixel in the target resolution grid, the graphics processor 300 maps it to the input luminance space and the input chrominance space according to the scaling factor. Let the horizontal linear magnification factor be... The vertical linear magnification factor is The input chromaticity sampling point coordinates are The chroma sampling center offset is The coordinate mapping of the target resolution space, input luminance space, and input chrominance space follows the pixel center alignment principle. The general mapping relationship between the coordinates of the input chrominance sampling points and the coordinates of the target pixels is as follows: in , These represent the chroma downsampling ratios in the horizontal and vertical directions, respectively, for the 4:2:0 format. , For the 4:2:2 format , For 4:4:4 format , For the 4:2:2 format, half-sample mapping can be performed only in the horizontal direction; for the 4:4:4 format, half-sample mapping can be omitted. The mapping relationship between different sampling formats is determined by the horizontal and vertical sampling ratios in the texture descriptor.
[0043] The graphics processor 300 calculates the center-mapped coordinates of the input chroma sampling point in the target resolution grid based on the input chroma sampling point coordinates and the chroma sampling center offset. In one embodiment, for the 4:2:0 format, the center-mapped coordinates are calculated as follows: The graphics processor 300 reads the edge phase direction corresponding to the current target pixel position from the brightness edge direction feature map. ,in: The position deviation vector of the target pixel position relative to the center mapping coordinates is The phase offset along the normal to the brightness edge is obtained through the vector dot product: This phase offset is used to characterize the geometric deviation of the current target pixel relative to its chroma sampling center in the normal direction of the luminance edge. When the target pixel is located near a strong luminance edge and its normal phase offset is small, it indicates that this location is more suitable for using chroma features extracted along the edge tangentially to avoid chroma samples spreading across the luminance boundary.
[0044] The graphics processor 300 calculates the local gradient magnitude. In one embodiment: Before calculating the spatial phase alignment weights, first calculate the horizontal gradient. with vertical gradient Normalize the gradient magnitude based on the effective bit width of the brightness target. The calculation formula is: in To determine the target effective bit width, this embodiment uses 10, which corresponds to a denominator of 1024, normalizing the gradient magnitude to the range of 0 to 2; video quantization accuracy threshold. For a constant within the normalization domain, the preferred value is... To align with the normalized gradient, the weight calculation formula is revised as follows: in Used to control the extent of chromaticity correction along the edge normal. The value selection rule is as follows: the default value is the value corresponding to 0.5 target resolution pixels. When the magnification is greater than or equal to 4, The value is adjusted to correspond to the value of one target resolution pixel when the chroma sampling format is 4:4:4. The value is adjusted to the value corresponding to 0.25 target resolution pixels; The parameters are dynamically adjusted according to the target effective bit width. The general calculation formula is: This calculation ensures that the weight suppression effect in flat regions remains consistent across different bit widths; the weight calculation employs an exponential decay form to ensure that the alignment weight decreases faster as the phase offset increases, guaranteeing that it only affects the near-neighbor region in the normal direction at the luminance edge and avoiding large-scale chromatic distortion. In a preferred embodiment, , Take the spatial span corresponding to half the target resolution pixel; in other embodiments... Based on the output bit width to Values within the range, The spatial span can be taken as 0.25 to 1.5 pixels corresponding to the target resolution.
[0045] The graphics processor 300 calculates the brightness edge tangency based on the edge phase direction. In one embodiment: Edge tangential vector Defined in the target resolution space, when used for input chroma plane sampling offset, it needs to be scaled according to the chroma downsampling ratio and magnification. The scaled input space tangential offset vector The calculation formula is: in , These are the horizontal and vertical components of the tangent vector in the target space, respectively. , These are unit vectors in the horizontal and vertical directions, respectively, ensuring that the sampling offset is consistent with the geometric scale of the input chromaticity space and the edge direction of the target space; phase offset. and The parameters maintain the same numerical scale, requiring no additional conversion.
[0046] The graphics processor 300 extracts multiple chromaticity samples tangentially along the luminance edge on the input chromaticity plane. In one embodiment, respectively in... , and Perform bilinear texture sampling at three locations to obtain , and The edge-direction chromaticity features are calculated as follows: The edge-direction chromaticity features employ a 3-point tangential sampling weighted calculation design, based on the visual characteristic that chromaticity changes at brightness edges are continuous along the tangential direction but abruptly change along the normal direction. 3-point sampling achieves an optimal balance between computational complexity and edge smoothing effect. The sampling step size is set to 1 / 2 unit of the tangential vector to ensure the sampling range covers the chromaticity information of half a pixel on each side of the edge. This ensures that the sampling area does not cross the edge and cause color mixing, while fully utilizing neighboring chromaticity samples to improve interpolation accuracy. The weighting coefficients use 1 / 4, 1 / 2, and 1 / 4 triangular window weighting, corresponding to the first-order approximation of the Gaussian kernel. This low-pass filtering of the tangential chromaticity signal suppresses sampling noise while maintaining the naturalness of chromaticity transitions at the edges. For different magnification scenarios, the number of sampling points and the step size remain fixed. Only the actual sampling step size within the input chromaticity space needs to be adjusted according to the aforementioned spatial vector scaling rules to ensure consistent edge alignment at different magnifications.
[0047] At the same time, the graphics processor 300 performs conventional spatial interpolation on the input chromaticity plane to obtain the reference chromaticity features. The conventional spatial interpolation can be bilinear interpolation, bicubic interpolation, Lanczos interpolation, polynomial interpolation, or other texture interpolation methods supported by the graphics processor 300. The baseline chromaticity feature is used to ensure chromaticity continuity in flat and weakly edged regions, while the edge-direction chromaticity feature is used to constrain the chromaticity diffusion direction near bright edges.
[0048] The graphics processor 300 fuses edge-direction chromaticity features and reference chromaticity features based on spatial phase alignment weights to obtain the target resolution chromaticity plane. In this embodiment, It can be a two-dimensional vector containing blue and red components. The fusion process can be performed on the blue and red components separately, or it can be performed all at once as a two-dimensional vector. The graphics processor 300 traverses the pixels within the target resolution grid in parallel to generate a target resolution chroma plane with the same spatial grid as the target resolution luminance plane.
[0049] The interleaved chroma plane uses a storage format that alternates between Cb and Cr components. In each row, two adjacent bytes correspond to the Cb and Cr components at the same position, respectively. The graphics processor 300 configures this plane as a dual-channel 8-bit or 10-bit unsigned integer texture, with the texture format corresponding to the dual-channel interleaved format. The hardware texture sampling unit can automatically identify the two chroma components and simultaneously return the Cb and Cr values at the corresponding positions during sampling. Chromaticity edge phase alignment calculations are performed independently for the Cb and Cr components. The two components share the same set of luminance edge direction features, spatial phase alignment weights, and sampling coordinates. Only the sampled chroma values participate in the calculation of the reference chroma features and edge direction chroma features of their respective components, ultimately outputting the target resolution chroma planes for each component. The sampling boundaries of the chroma plane adopt an edge copying and filling strategy, consistent with the boundary processing method of the luminance plane. When the sampling coordinates exceed the range of the input chroma plane, the nearest boundary pixel value is taken as the sampling result to avoid chroma anomalies in the boundary region.
[0050] Example 6: The input luminance bands are divided evenly along the vertical direction, with each row as a unit, and the band height... Two constraints must be met: first, the value must be an integer power of 2; second, the value must be a power of 2. The maximum line cache depth must be less than or equal to 500 for the reconfigurable line cache accelerator. This represents the maximum one-sided receptive field of the network. In this embodiment, the maximum depth of the down-buffer is 128 lines, which is suitable for a 2x super-resolution scenario. Taking 17 corresponds to a strip height of 32 rows; in a 4x super-resolution scenario, the corresponding strip height is 16 rows. The ratio of the input chroma strip height to the input luminance strip height is equal to the chroma vertical downsampling ratio. In the 4:2:0 format, the chroma strip height is half the luminance strip height; in the 4:2:2 format, they are equal. The target output strip height is equal to the input luminance strip height multiplied by the vertical magnification factor. There is a one-to-one correspondence between the input and output stripes. After residual reconstruction, a single input strip generates a single target output strip at the corresponding position. The remaining lines at the end of the frame that are less than a complete strip are padded to the height of a complete strip by copying and filling the bottom edge. After processing, the output lines corresponding to the padded areas are discarded to ensure the accuracy of the output frame size.
[0051] When reconstructing the luminance residual for the current strip, it is necessary to read the values above and below the current strip. When processing the first strip of the row's sideband data, if there is no valid data above it, the top edge is copied to fill the gap. The top edge band data is filled by copying the pixel values of the first row to fill all the top edge band rows; when processing the last strip, since there is no valid data below, the bottom edge is copied to fill it. The bottom sideband data is copied to fill all bottom sideband rows. When processing the current strip for chroma edge phase alignment, only the luminance edge direction feature map strip data at the corresponding position is needed, without the feature data of adjacent strips, thus eliminating inter-strip dependencies. The system supports two modes: single-strip serial processing and dual-strip parallel processing. In serial mode, each strip is processed sequentially, while in parallel mode, two adjacent strips can be processed simultaneously. During parallel processing, the overlapping sideband data of the two strips share the same row buffer area, eliminating the need for repeated data transfer and ensuring bandwidth efficiency. The dependencies in strip processing only exist between adjacent strips within the same frame; strip processing in different frames is completely independent with no cross-frame dependencies.
[0052] The main processor 200 divides the target resolution luminance plane and the target resolution chrominance plane into multiple target output stripes. Each target output strip corresponds to luminance data and chrominance data within the same vertical range. The main processor 200 generates a data merging descriptor for each target output strip. The data merging descriptor includes at least the frame timestamp, stripe number, luminance target storage address, chrominance target storage address, luminance data length, chrominance data length, luminance write completion flag, chrominance write completion flag, exception flag, and next descriptor address.
[0053] The unified frame buffer 800 includes a luma region and a chroma region. The luma region stores the target resolution luma plane, and the chroma region stores the target resolution chroma plane. In one embodiment, the luma region and the chroma region are located in the same contiguous physical address space and are distinguished by a fixed offset; in another embodiment, the luma region and the chroma region may be located in different storage partitions and logically associated through the target storage address in the data merge descriptor.
[0054] The reconfigurable line buffer accelerator 500 writes the target resolution luminance data corresponding to the target output stripe into the luminance region of the unified frame buffer 800. The graphics processor 300 writes the target resolution chrominance data corresponding to the same target output stripe into the chrominance region of the unified frame buffer 800. Both the luminance output direct memory access transaction and the chrominance output direct memory access transaction carry the same frame timestamp and stripe number. The synchronization arbitration unit 900 identifies whether the two write transactions correspond to the same spatial segment based on the frame timestamp and stripe number.
[0055] In one embodiment, the main processor 200 writes the stripe number to the luminance write control register of the reconfigurable line cache accelerator 500 and the same stripe number to the chroma write control register of the graphics processor 300. The reconfigurable line cache accelerator 500 carries the stripe number as a transaction tag in luminance write transactions; the graphics processor 300 carries the same stripe number as a transaction tag in chroma write transactions. The storage controller or synchronization arbitration unit 900 avoids spatial mismatches of luminance and chroma data between different stripes by matching the transaction tags.
[0056] When the unified frame buffer 800 receives a luminance write complete signal, it sets the luminance write complete flag in the corresponding data merge descriptor; when it receives a chrominance write complete signal, it sets the chrominance write complete flag in the corresponding data merge descriptor. Let the luminance write complete flag be... The chromaticity writing is complete. Then the completion status of the target output strip merging can be represented as: when When true, the main processor 200 moves the data merging descriptor corresponding to the target output stripe into the output merging queue. When all target output stripes corresponding to the same input frame enter the output merging queue, and their frame timestamps and stripe numbers are sequential, a complete intermediate image frame is formed in the unified frame buffer 800. This intermediate image frame contains target resolution luminance samples and target resolution chrominance samples aligned with their spatial phase.
[0057] The data merging descriptor is pre-generated by the main processor 200 during the frame initialization phase. Each target output stripe corresponds to one descriptor. All descriptors are stored in the descriptor table in main memory in order of stripe number. Each descriptor occupies 32 bytes and includes the frame timestamp, stripe number, luminance target storage address, chrominance target storage address, luminance data length, chrominance data length, luminance write completion flag, chrominance write completion flag, exception flag, and next descriptor address pointer. The synchronization matching logic is automatically executed by the hardware state machine in the storage controller. When the reconfigurable line cache accelerator 500 completes the luminance data writing and sends a write completion signal, the hardware automatically indexes the corresponding descriptor according to the stripe number and sets the luminance write completion flag to 1. Similarly, when the graphics processor 300 completes the chrominance data writing and sends a write completion signal, it sets the chrominance write completion flag to 1. After each completion flag is set, the hardware automatically determines whether the two completion flags in the same descriptor are both 1. If they are both 1, the descriptor is automatically moved to the end of the output merging queue, and a merging completion interrupt is sent to the main processor 200. The completion flag is cleared by the main processor 200 after reading the output merging queue. After clearing, the descriptor can be recycled for strip merging in the next frame.
[0058] Example 7: The graphics processor 300 or display interface controller reads intermediate image frames from the unified frame buffer 800 and performs normalization, quantization range limiting, and wide color gamut format writing on the luminance and chrominance samples. In one embodiment, both luminance and chrominance samples are stored as fixed-point integers, and normalization is used to map the sample values to a continuous range between 0 and 1. Let the input sample value be... The minimum theoretical value of the sample is The maximum theoretical value of the sample is The normalized value is ,but: In an embodiment where the output is a 10-bit video with a valid quantization range, the quantization range for the luminance component can be 64 to 940, and the quantization range for the chrominance component can be 64 to 960. Let the lower limit of the quantization range be... The upper limit of the quantization interval is The output quantization value after limitation is ,but: The quantization interval automatically adapts to the output bit width and quantization range type. Within a limited range, the general calculation formula for the brightness quantization interval is: The general formula for calculating the colorimetric quantization interval is: in For output bit width, for 8-bit output The corresponding brightness range is 16-235, and the chroma range is 16-240, for 10-bit output. This corresponds to a brightness range of 64-940 and a chroma range of 64-960 for 12-bit output. This corresponds to a brightness range of 256-3760 and a chromaticity range of 256-3840; across the entire range, the upper and lower limits of the quantization range are 0 and 1, respectively. The quantization process uses a rounding method that takes the nearest even number. When the calculated value is between two integers, the even number is taken as the final quantization result to reduce quantization error. The normalization benchmark uses the upper and lower limits of the quantization interval. That is, the input sample value is first mapped to the normalization interval of 0 to 1, and then multiplied by the quantization interval span and the lower and lower limits are added to obtain the quantization result, ensuring the consistency of color mapping under different bit widths and different ranges.
[0059] After quantization and constraint, the samples are written into a wide color gamut format container according to the set output line step size, forming a working frame buffer. The wide color gamut format container is compatible with the BT.2020 color space and other wide color gamut display standards. The output line step size is determined based on the target resolution, output bit width, chroma sampling format, and alignment byte count.
[0060] Color space conversion follows the BT.2020 standard and is performed in two steps. The first step converts the limited range of YCbCr values to linear RGB values. The second step performs photoelectric conversion according to output requirements. The YCbCr to RGB conversion matrix uses BT.2020 standard coefficients, and the conversion formula is as follows: in , , To normalize luminance and chromaticity values to the 0-1 range, the normalization formula for converting a finite range to the full range is: in , These represent the upper and lower limits of the brightness quantization range. , This represents the upper and lower limits of the colorimetric measurement range. The photoelectric conversion uses the BT.1886 gamma transfer function by default, with a gamma value of 2.4. It can also be switched to the PQ or HLG transfer function according to the output interface configuration. When switching, simply replace the corresponding conversion formula. The output sampling format is consistent with the input by default. If you need to output 4:4:4 or RGB format, you can directly output the converted full sample data without additional interpolation.
[0061] After the working frame buffer is written, the storage controller generates a write completion interrupt or a completion status flag. Upon responding to the write completion interrupt or polling for the completion status flag, the main processor 200 writes the output timestamp to the header control field of the working frame buffer. The output timestamp can inherit the corresponding input frame timestamp or be regenerated by the display clock according to the target output frame rate. Regardless of the method used, the output timestamp is used to ensure that the continuous image frame sequence is encapsulated and output in the correct time order.
[0062] The main processor 200 sends the working frame buffer address, frame control field, and status field (bound with the output timestamp) to the display timing synchronization queue. The display timing synchronization queue outputs the working frame buffer according to the target refresh cycle. The target refresh cycle can be 50Hz, 60Hz, 90Hz, 120Hz, 144Hz, or other refresh frequencies supported by the display device. In one embodiment, the display interface controller extracts the working frame buffer at the head of the queue from the display timing synchronization queue at the arrival of each synchronization cycle and delivers it to the output interface 1000 as the target frame rate output frame buffer.
[0063] The display timing synchronization queue adopts a circular queue structure with a maximum depth of 4 frames. Each node in the queue stores the base address of the working frame buffer and the corresponding output timestamp. When the input frame rate is equal to the output frame rate, the output timestamp directly inherits the timestamp of the corresponding input frame. The queue outputs frames one by one in a first-in-first-out order. In each display synchronization cycle, the first frame of the queue is taken out and delivered to output interface 1000. When the input frame rate is lower than the output frame rate, a frame repetition strategy is used for frame rate conversion. That is, the same input frame corresponds to multiple output cycles, and the output timestamp increases at equal intervals according to the output frame rate. When there is only one frame of valid data in the queue and the next output cycle arrives, the data of that frame is repeatedly output until a new input frame is processed and enqueued. When the input frame rate is higher than the output frame rate, a frame dropping strategy is used. That is, when the queue is full, the newly processed frame is directly discarded, and the frame data with the latest timestamp in the queue is retained to ensure the real-time performance of the output screen. The timestamp mapping rule is as follows: the output frame timestamp is equal to the corresponding input frame timestamp plus the difference between the output frame sequence number and the input frame sequence number multiplied by the output period, ensuring that the output timestamp increases monotonically and the intervals are uniform; when the queue is full, a backpressure signal is triggered to send a pause command to the front-end processing pipeline to avoid data overflow.
[0064] The output interface 1000 encapsulates the target frame rate output frames into a continuous image frame sequence according to the output timestamp order. During the encapsulation process, each frame can be appended with display color metadata, data bit width information, target resolution information, transfer function information, and synchronization identifier. The display color metadata may include target color space matrix parameters, photoelectric conversion function parameters, white point information, and effective display area information. The data bit width information can be 8 bits, 10 bits, 12 bits, or 16 bits.
[0065] Example 8: In the above embodiments, the target resolution can be 3840*2160, 7680*4320, 8192*4320, or other ultra-high-definition video resolutions; the input frame rate and output frame rate can be the same or different; the input chroma sampling format can be 4:2:0, 4:2:2, or 4:4:4; the input bit width and output bit width can be 8 bits, 10 bits, 12 bits, or 16 bits. The neural network processor 400 can be implemented by a dedicated NPU, DSP, GPU computing core, FPGA matrix operation array, or other hardware units capable of performing matrix multiplication and convolution operations. The graphics processor 300 can be implemented by a discrete GPU, on-chip GPU, programmable shader array, texture sampling array, or dedicated image processor.
[0066] In another embodiment, the luminance edge direction feature map does not need to be generated all at once across the entire frame, but can be generated, buffered, and sent to the graphics processor 300 strip by strip. When processing chroma stripes, the graphics processor 300 only needs to read the luminance edge direction feature map within the same spatial range as the current target output strip. This reduces on-chip cache usage and allows the luminance reconstruction path and chroma phase alignment path to be executed in parallel in a pipelined manner.
[0067] In another embodiment, the spatial phase alignment weights can also be determined based on edge strength, phase offset, motion confidence, chroma noise estimate, or texture complexity. For example, in high-speed motion regions, the contribution of historical frames to the luminance tensor can be reduced; in regions with high chroma noise, the weight of the reference chroma features can be increased; and in high-saturation thin lines or text edge regions, the weight of edge-direction chroma features can be increased. These changes do not alter the core idea of this invention: constraining the chroma interpolation direction using the luminance edge direction.
[0068] The system employs a tiered handling mechanism for different types of anomalies. Input anomalies include frame loss, timestamp discontinuity, and format mutation. When frame loss or timestamp jumps exceeding a threshold are detected, the main processor 200 automatically clears the circular buffer 700 and all processing queues, resets the pipeline state, and re-establishes the processing flow starting from the next valid input frame. When a format mutation is detected, the main processor 200 reconfigures the format parameters of all hardware modules, adjusts the striping rules and buffer size, and ensures correct processing of subsequent frames. Computational anomalies include timeouts by the neural network processor 400 and errors by the graphics processor 300. When a computational anomaly is detected, the corresponding processor automatically terminates its current task, reports the anomaly status, and the main processor 200 discards all intermediate data of the current frame, clears the corresponding task queue, resets the processor state, and continues processing the next frame. Transmission anomalies include direct memory access errors and bus congestion. When a transmission error is detected, the direct memory access controller automatically terminates the current transmission, reports the anomaly, and the main processor 200 re-initiates the transmission of the corresponding data. If three consecutive transmissions fail, the current frame is discarded, and the bus state is reset. When any exception occurs, the exception type and the corresponding frame timestamp are recorded in the status register for upper-layer software to read and analyze. After the exception is resolved, the system automatically resumes normal streaming processing without the need for a global reset.
[0069] The technical advantages of this embodiment are as follows: First, the luminance plane and chrominance plane still maintain separate heterogeneous processing, avoiding premature expansion of half-sampled YCbCr data into full-resolution RGB data, thereby reducing memory usage and bus bandwidth pressure; Second, the high-frequency edge direction information generated by the luminance plane after reconstruction by the neural network processor 400 is used for spatial interpolation of the chrominance plane, so that the chrominance interpolation direction extends tangentially along the edge at the strong luminance edge, reducing cross-edge mixing; Third, a complete synchronization chain is established by input frame timestamp, stripe number, write completion flag and output timestamp, avoiding spatial or temporal mismatch between the luminance branch and the chrominance branch due to different heterogeneous processing delays.
[0070] Compared to schemes that only perform bilinear or bicubic interpolation on the chroma plane, this embodiment can reduce the spatial offset of the chroma boundary relative to the luminance edge in high-saturation text edges, dark background colored lines, moving object edges, and wide color gamut display scenarios, thus mitigating color smudges, chroma ghosting, and edge color banding. Compared to schemes that send luminance, chroma, or all RGB channels to the neural network processor 400 for joint reconstruction, this embodiment does not require processing complete three-channel full-resolution data in the neural network processor 400, making it more suitable for real-time ultra-high-definition video streaming reconstruction.
[0071] The embodiments of the present invention have been described above, but the present invention is not limited to the above implementations. The above implementations are merely illustrative and not restrictive. Those skilled in the art, guided by the present invention, can make various equivalent transformations to the specific hardware form, cache capacity, network structure, interpolation kernel function, parameter range, data format, and output interface. As long as they still utilize the edge direction information in the luminance reconstruction result to guide the chromaticity plane phase alignment, they should all fall within the protection scope of the present invention.
Claims
1. A super-high-definition video super-resolution reconstruction system based on a heterogeneous computing architecture, applied to a heterogeneous computing architecture including a main processor, a graphics processor, a neural network processor, a reconfigurable line cache accelerator, and an output interface, characterized in that, Configured for execution: The system captures an initial image frame sequence carrying a timestamp through the input interface, parses the metadata, divides the input samples of the initial image frame sequence into a luminance plane and a chrominance plane, and writes them into a fixed-point video buffer. The luminance plane is fed into the neural network processor and the reconfigurable line buffer accelerator cache, and the chroma plane is fed into the video frame cache of the graphics processor. The neural network processor is used to perform residual reconstruction on the brightness plane to generate a target resolution brightness plane, and a brightness edge direction feature map is extracted and generated, including: extracting the horizontal and vertical gradients of the target resolution brightness plane, and generating the brightness edge direction feature map based on the horizontal and vertical gradients; Using the graphics processor, and combining the luminance edge direction feature map, edge phase alignment calculation is performed on the chroma plane to obtain the target resolution chroma plane. This includes: extracting local gradient vectors based on the horizontal and vertical gradients, and normalizing the local gradient vectors to determine the edge phase direction; calculating the phase offset along the edge normal, where the phase offset characterizes the geometric deviation of the current target pixel relative to its chroma sampling center along the luminance edge normal; calculating spatial phase alignment weights based on the local gradient vectors and the phase offset along the edge normal; determining the luminance edge tangent based on the edge phase direction, performing interpolation along the luminance edge tangent on the chroma plane to obtain edge direction chroma features, and performing spatial interpolation on the chroma plane to obtain reference chroma features; and fusing the edge direction chroma features and the reference chroma features based on the spatial phase alignment weights to output the target resolution chroma plane. The target resolution luminance plane and the target resolution chrominance plane are fused and written into a unified frame buffer to form an intermediate image frame containing complete luminance and chrominance samples; Perform color space conversion and output timing synchronization on the intermediate image frames to generate a target frame rate output frame buffer with an output timestamp; The target frame rate output frame buffer is encapsulated in the order of the output timestamps and output as target image sequence data through the output interface.
2. The ultra-high-definition video super-resolution reconstruction system based on heterogeneous computing architecture according to claim 1, characterized in that, The process of capturing an initial image frame sequence carrying a timestamp through an input interface, parsing metadata, dividing the input samples of the initial image frame sequence into a luminance plane and a chrominance plane, and writing them into a fixed-point video buffer includes: Receive the initial image frame sequence, and extract spatial metadata and the input frame timestamp as the timestamp; Write the luminance sample from the input sample into the luminance plane, and write the chrominance sample from the input sample into the interleaved chrominance plane as the chrominance plane, and truncate it according to the set bit width; Generate a working descriptor that includes the luminance plane base address, the interleaved chrominance plane base address, the plane row step size, the chrominance sampling center offset, and the input frame timestamp; The working descriptor, along with the corresponding luminance plane and the interleaved chrominance plane, is fed into the circular buffer.
3. The ultra-high-definition video super-resolution reconstruction system based on heterogeneous computing architecture according to claim 2, characterized in that, The step of sending the luminance plane to the neural network processor and the reconfigurable line buffer accelerator cache, and sending the chroma plane to the video frame buffer of the graphics processor, includes: Establish a luminance direct memory access list and a chrominance direct memory access list based on the working descriptor; The brightness plane is divided into brightness stripes according to the grid stripes, and the brightness stripes are stored in the reconfigurable row cache accelerator through the brightness direct memory access linked list and mapped to the input tensor cache area of the neural network processor. Maintain the original sampling format of the chroma plane, store it in the texture sampling buffer of the graphics processor through the chroma direct memory access linked list, and record the chroma sampling center offset in the texture descriptor; The frame timestamps are bound to the luminance plane and the chrominance plane.
4. The ultra-high-definition video super-resolution reconstruction system based on heterogeneous computing architecture according to claim 3, characterized in that, The step of using the neural network processor to perform residual reconstruction on the brightness plane to generate a target resolution brightness plane and extracting and generating a brightness edge direction feature map includes: The reconfigurable line buffer accelerator is used to stitch together the current brightness stripe of the brightness plane, the brightness stripe at the same position in adjacent frames, and the stripe sidebands to generate tensor data and store it in the input tensor buffer area. The neural network processor reads the tensor data to generate amplified brightness features and residual features, and then fuses the amplified brightness features and the residual features to reconstruct the target resolution brightness plane. Extract the horizontal and vertical gradients of the target resolution brightness plane within the input tensor buffer; The brightness edge direction feature map is generated based on the horizontal gradient and the vertical gradient, and the brightness edge direction feature map enters the strip buffer area along with the target resolution brightness plane.
5. The ultra-high-definition video super-resolution reconstruction system based on heterogeneous computing architecture according to claim 4, characterized in that, The step of using the graphics processor to perform edge phase alignment calculation on the chroma plane in conjunction with the luminance edge direction feature map to obtain the target resolution chroma plane includes: Obtain the grid coordinates of the target resolution pixel and its corresponding input chroma sampling point coordinates. Map the input chroma sampling point coordinates to obtain the center mapping coordinates based on the chroma sampling center offset. Calculate the phase offset along the edge normal by combining the grid coordinates and the edge phase direction.
6. The ultra-high-definition video super-resolution reconstruction system based on heterogeneous computing architecture according to claim 1, characterized in that, The step of fusing the target resolution luminance plane and the target resolution chrominance plane and writing them into a unified frame buffer to form an intermediate image frame containing complete luminance and chrominance samples includes: The target resolution luminance plane and the target resolution chrominance plane are divided into multiple target output stripes, and a data merging descriptor is generated for each target output stripe in the main processor; The unified frame buffer is configured to include a luma region and a chroma region; The reconfigurable line buffer accelerator is used to write the target resolution luminance plane corresponding to the target output stripe into the luminance region; The graphics processor is used to write the target resolution chroma plane corresponding to the same target output strip into the chroma region; The writing process of the luminance region is bound to the writing process of the chrominance region according to the strip number; Once the luminance and chrominance samples within the target output strip have been written, the data merging descriptor is moved into the output merging queue to generate the intermediate image frame.
7. The ultra-high-definition video super-resolution reconstruction system based on heterogeneous computing architecture according to claim 1, characterized in that, The step of performing color space conversion and output timing synchronization on the intermediate image frames to generate a target frame rate output frame buffer with an output timestamp includes: Read the intermediate image frame and perform normalization processing on the luminance and chrominance samples in the intermediate image frame; The normalized sample values are restricted to the set video quantization range and written into a wide color gamut format container according to the set output line step size to generate a working frame buffer. After the working frame buffer has been written, the output timestamp is bound and the working frame buffer is sent to the display timing synchronization queue. The display timing synchronization queue is controlled to output the working frame buffer according to the set synchronization period, thereby generating the timing-aligned target frame rate output frame buffer.
8. The ultra-high-definition video super-resolution reconstruction system based on heterogeneous computing architecture according to claim 7, characterized in that, The step of encapsulating the target frame rate output frame buffer according to the output timestamp order and outputting it as target image sequence data through the output interface includes: The target frame rate output frame buffer is extracted sequentially from the display timing synchronization queue; The extracted target frame rate output frame buffer is encapsulated into a continuous image frame sequence according to the output timestamp order, and the target frame rate output frame buffer is appended with display color metadata and data bit width information; The continuous image frame sequence is output through the output interface to generate the target image sequence data.
Citation Information
Patent Citations
An image super-resolution reconstruction method based on sparse representation and deep learning
CN109741256A
Autonomous controllable architecture software and hardware collaborative audio and video signal transmission method and system
CN120751133A