Display devices, display methods based on content-adaptive metadata, and storage media
By combining a metadata scheduler and a conditional super-resolution reconstructor, the optimal subset of metadata is dynamically selected, solving the problem that existing super-resolution methods cannot adapt to diverse video content. This achieves a balance between quality and overhead under resource constraints, improving display quality and system efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TCL CHINA STAR OPTOELECTRONICS TECHNOLOGY CO LTD
- Filing Date
- 2026-03-20
- Publication Date
- 2026-07-03
AI Technical Summary
Existing super-resolution methods based on metadata conditions have failed to effectively adapt to the diverse needs of varied video content, resulting in poor reconstruction quality and wasted resources, and failing to achieve a balance between quality and cost under resource budget constraints.
A metadata scheduler is used to perform feature analysis on video content, dynamically select the optimal subset of metadata, and process it through a conditional super-resolution reconstructor to achieve adaptive super-resolution processing.
Under resource constraints, the display effect and system efficiency are improved, the reconstruction fidelity of key areas is ensured, and the device resource overhead caused by unnecessary metadata is reduced, achieving the optimal balance between super-resolution quality and overhead.
Smart Images

Figure CN122335545A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of display technology, specifically to display devices, display methods based on content-adaptive metadata, and storage media. Background Technology
[0002] Super-resolution technology aims to recover high-resolution images or video content from low-resolution observed signals, and is one of the main technologies for improving the output image quality of display devices. With the development of generative models, super-resolution technology has evolved from traditional deterministic regression to conditional generation, capable of generating fine visual details that surpass traditional pixel-level restoration. In practical application scenarios of display devices, the video content to be displayed often has complex and diverse characteristics. Different segments may alternately contain elements such as text overlay, fast motion, smooth animation, and low-light faces. The metadata requirements for auxiliary reconstruction vary significantly among different types of content. For example, dense text areas require sharp edge guidance, while high-speed motion scenes rely on temporal motion cues.
[0003] Existing metadata-based super-resolution methods typically employ fixed metadata configuration schemes, meaning that a single or multiple preset metadata types are always enabled during the inference phase, without considering the differentiated metadata requirements of different content segments. This fixed configuration mode has significant limitations in practical applications. When optimal auxiliary information is highly dependent on content features, fixed metadata input cannot adapt to diverse video segments, not only making it difficult to guarantee the reconstruction quality of various content types but also potentially introducing unnecessary computational, bandwidth, and latency overhead due to indiscriminately enabling metadata. Furthermore, traditional super-resolution systems have high coupling between metadata and the reconstructor, lack a unified interface design, and cannot flexibly adapt to different types of super-resolution backbone networks, further limiting their widespread application in display devices.
[0004] In the actual operation of display devices, resource budgets are strictly limited. It is neither realistic nor necessary to enable all available metadata indiscriminately. In some cases, some metadata may even introduce reconstruction artifacts in unsuitable content scenarios, thereby reducing the display effect. Summary of the Invention
[0005] This application provides a display device, a display method based on content-adaptive metadata, and a storage medium, which can dynamically select the optimal subset of metadata according to the video content, achieve the optimal balance between quality and overhead under resource constraints, and improve display effect and system efficiency.
[0006] In a first aspect, a display device is provided, including a display system based on content-adaptive metadata, the display system including a metadata scheduler and a conditional super-resolution reconstructor; The metadata scheduler is used to perform feature analysis on the video content to be displayed at a first resolution, extract the content features of the video content, and select a target metadata subset from multiple preset metadata categories based on the content features; wherein the content features include attribute information used to describe the dynamic characteristics and content composition of the video content; The conditional super-resolution reconstructor is used to perform super-resolution processing on the video content based on each metadata category in the target metadata subset, and output display content at a second resolution, wherein the second resolution is greater than the first resolution.
[0007] Secondly, a display method based on content-adaptive metadata is also provided, which runs on a display device, the method comprising: Feature analysis is performed on the video content to be displayed at a first resolution to extract the content features of the video content, and a target metadata subset is selected from multiple preset metadata categories based on the content features; wherein, the content features include attribute information used to describe the dynamic characteristics and content composition of the video content; Super-resolution processing is performed on the video content based on each metadata category in the target metadata subset, and the display content at a second resolution is output, wherein the second resolution is greater than the first resolution.
[0008] Thirdly, a computer-readable storage medium is also provided, the computer-readable storage medium storing a computer program configured to be executed by a processor to implement the above-described content-adaptive metadata-based display method.
[0009] In summary, this invention, through the implementation of a metadata scheduler, performs feature analysis on the first-resolution video content to be displayed and extracts corresponding attribute information. This allows for the collection of differences in the dynamic characteristics and content composition of video segments, and then the targeted selection of a subset of metadata, avoiding the limitations of traditional fixed metadata configuration schemes. The metadata scheduler's adaptive selection strategy based on content features can match optimal metadata auxiliary information for different types of video segments under resource constraints. This ensures the reconstruction fidelity of key areas (such as text, faces, and high-speed motion areas) while reducing the device resource overhead caused by unnecessary metadata. It achieves a balance between super-resolution quality and device resource overhead, improving the system versatility and scalability of the display device. Ultimately, it outputs high-definition, highly stable second-resolution (super-resolution) display content, optimizing the output image quality of the display device. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of a display device provided in an embodiment of this application.
[0012] Figure 2 This is a schematic diagram of the metadata scheduler provided in an embodiment of this application.
[0013] Figure 3 This is a schematic diagram of a conditional super-resolution reconstructor provided in an embodiment of this application.
[0014] Figure 4 This is a flowchart illustrating the content-adaptive metadata-based display method provided in this application embodiment.
[0015] Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of this application.
[0016] Figure 6 An example diagram of the system architecture provided for an exemplary implementation of this application.
[0017] Figure 7 This is a flowchart illustrating a video processing method provided in an exemplary embodiment of this application.
[0018] Figure 8 This is one of the structural example diagrams of the sending end and receiving end provided in an exemplary embodiment of this application.
[0019] Figure 9 This is the second example diagram showing the structure of the sending end and receiving end provided in an exemplary embodiment of this application.
[0020] Figures 10A-10B This is one of the structural example diagrams of the sending end and receiving end provided in an exemplary embodiment of this application.
[0021] Figures 11-12 This is the second example diagram showing the structure of the sending end and receiving end provided in an exemplary embodiment of this application. Detailed Implementation
[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0024] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0025] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration of devices to perform additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more conditions or values may, as examples, be based on additional conditions or values beyond those stated.
[0026] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0027] Based on the technical issues described in the background, and considering that generative super-resolution technology, as a core technology for improving image and video display quality, can recover high-resolution content with rich details from low-resolution original input, it has been widely used in various real-world scenarios such as smart display devices, video playback terminals, and streaming media transmission. In actual super-resolution processing, the input content to be processed exhibits a high degree of diversity, with significant differences in its domain, content type, and degradation patterns between different segments. This results in differentiated requirements for metadata (i.e., super-resolution prompts) to assist in super-resolution reconstruction for various types of content.
[0028] As an example, video content in real-world applications is often composed of alternating segments of various features. These include dense text segments containing subtitles and UI overlays, fast-moving segments such as sports events and action sequences, smooth animation clips like those in anime and cartoons, and low-light photo segments of faces and scenery taken at night or in low-light environments. The types of auxiliary information required for high-quality super-resolution reconstruction vary significantly depending on the content segments' features: dense text areas require metadata guided by edge sharpness to ensure clear text; high-speed motion scenes rely on metadata based on temporal motion cues to mitigate motion blur and artifacts; smooth animation clips are better suited to metadata based on texture completion to enhance image detail; and low-light photo segments require metadata based on noise suppression and brightness compensation to optimize image purity and visual effects. All types of content can achieve significant super-resolution quality improvements by using specific forms of auxiliary information that match their features.
[0029] In existing technologies, generative super-resolution methods based on metadata conditions mostly adopt fixed metadata condition design schemes. That is, during the inference phase, a preset single metadata type is always used, or a fixed combination of multiple metadata types is used indiscriminately. This does not consider the personalized and differentiated needs of different content and fragments for metadata. This fixed-configuration processing mode has significant performance limitations in practical applications: First, when the optimal prompt information for super-resolution reconstruction is highly dependent on the characteristics of the input content itself, fixed metadata input cannot flexibly adapt to diverse content and fragments, making it difficult to guarantee the reconstruction quality of various feature contents. It may even introduce reconstruction artifacts due to the mismatch between metadata and content features, reducing the display effect. Second, super-resolution processing in real-world scenarios faces strict resource budget constraints, including limitations on computational load, transmission bandwidth, and processing latency. Indiscriminate use of metadata will result in a large waste of resources, while a single metadata type cannot meet the reconstruction needs of different content. Neither can achieve a balance between resource consumption and super-resolution quality, ultimately leading to poor implementation results in real-world scenarios.
[0030] In summary, how to dynamically match the optimal metadata auxiliary information for diverse input content features, while adapting to resource budget constraints, and achieving the optimal balance between quality and overhead in generative super-resolution processing has become a pressing technical problem that needs to be solved in the application of generative super-resolution technology in real-world scenarios.
[0031] In view of this, embodiments of this application provide a display device, a display method based on content-adaptive metadata, and a storage medium. By using a metadata scheduler and a conditional super-resolution reconstructor, adaptive metadata selection based on content features is achieved, which solves the problems of low efficiency and quality caused by fixed metadata configuration in the prior art. It can dynamically select the optimal metadata subset according to the video content, achieve the optimal balance between quality and overhead under resource constraints, and improve display effect and system efficiency.
[0032] The technical terms used in this application will be explained below: Display device: This can refer to a hardware device capable of receiving video signals and converting them into visual images, such as a television, monitor, projector, or mobile terminal. The device is configured to integrate a display system for optimizing display quality.
[0033] Content-adaptive metadata refers to auxiliary information that can be dynamically adjusted and selected based on the characteristics of the video content itself. Metadata is used to guide the super-resolution processing to improve the reconstruction quality of different types of video content.
[0034] Display system: This can refer to a collection of software or hardware modules integrated within a display device, responsible for processing and optimizing the video content to be displayed. The system is configured to perform super-resolution processing of video content to provide a higher quality visual experience.
[0035] Metadata scheduler: This can refer to a functional module in a display system whose function is to analyze video content and intelligently select the most suitable metadata for the current content based on the analysis results. This scheduler is configured to ensure that, within resource constraints, only the metadata most beneficial to improving quality is enabled.
[0036] Conditional super-resolution reconstructor: This can refer to another functional module in a display system that is configured to receive a subset of target metadata selected by a metadata scheduler and perform super-resolution processing on low-resolution video content based on the metadata. This reconstructor is configured to utilize the additional information provided by the metadata to generate display content with higher resolution and finer details.
[0037] First-resolution video content: This can refer to the original video signal input to the display system for processing, which has a relatively low spatial resolution.
[0038] Content features: These refer to attribute information extracted from video content that describes the dynamic characteristics and content composition of the video. They are used by the metadata scheduler for decision-making, such as the video's motion intensity, text ratio, and texture density.
[0039] Metadata categories: These can refer to pre-defined, selectable sets of different types of auxiliary information. Each category represents a specific source of information, such as degradation cues, structural information, temporal motion cues, semantic masks, or text region information.
[0040] Target metadata subset: This can refer to the set of metadata that the metadata scheduler selects from multiple metadata categories based on content characteristics, and that is most suitable for super-resolution processing of the current video content.
[0041] The second resolution display content can refer to the video content output after being processed by the conditional super-resolution reconstructor. Its spatial resolution is higher than the original first resolution, thus providing a clearer and more detailed visual effect.
[0042] like Figure 1 The diagram shown is a schematic of a display device provided in an embodiment of this application. The display device includes a content-adaptive metadata-based display system, which comprises a metadata scheduler and a conditional super-resolution reconstructor. The metadata scheduler performs feature analysis on the video content to be displayed at a first resolution, extracts content features of the video content, and selects a target metadata subset from multiple preset metadata categories based on the content features. The content features include attribute information describing the dynamic characteristics and content composition of the video content. The conditional super-resolution reconstructor performs super-resolution processing on the video content based on each metadata category in the target metadata subset, outputting display content at a second resolution, which is greater than the first resolution.
[0043] The metadata categories include at least a portion of degradation cues, structural and edge information, temporal motion cues, semantic masks, regions of interest (ROI) graphs, and text regions. Degradation cues include noise levels, artifact scoring, and blur proxy information; structural and edge information includes edge graphs and gradient graphs; temporal motion cues include optical flow, motion vectors, and shot transition markers. Semantic masks may include coarse-grained region labeling; ROI graphs may include importance weights for faces, salient regions, and subject regions; and text regions may include positional mask information for video subtitles and UI overlays. The conditional super-resolution reconstructor can receive video content, a subset of target metadata, and a selection mask through a unified interface.
[0044] As an example, the display device can be a standalone display terminal, such as a smart TV or computer monitor, or a display module integrated into other devices, such as a smartphone or tablet. The display system is configured to dynamically adjust the auxiliary information used in super-resolution processing based on the characteristics of the video content to be displayed, in order to optimize the display effect. For example, the display system can be implemented as a set of software programs running on the processor of the display device, or as a dedicated hardware acceleration module. Further, the display system is configured to include a metadata scheduler and a conditional super-resolution reconstructor. The metadata scheduler is configured to analyze the input video content and intelligently select appropriate metadata. The conditional super-resolution reconstructor is configured to perform super-resolution processing using the selected metadata. For example, the metadata scheduler can be implemented as a lightweight neural network model trained to predict the optimal combination of metadata based on the features of the input video. The conditional super-resolution reconstructor can be implemented as a deep learning-based super-resolution model, such as a convolutional neural network or a generative adversarial network, designed to receive and utilize external conditional information.
[0045] As an example, a metadata scheduler is used to perform feature analysis on the video content to be displayed at a first resolution and extract content features from the video content. This feature analysis process may include image processing of video frames, such as calculating the average brightness, contrast, or color histogram of the image. Statistical methods, such as sampling and statistically analyzing the pixel values of video frames, can also be used to obtain their basic attributes. For example, the video content can be analyzed frame-by-frame to calculate the average pixel value of each frame as a brightness feature, or the pixel differences between adjacent frames can be calculated as motion features. Based on this, the metadata scheduler is used to select a target metadata subset from multiple preset metadata categories based on these content features. This selection process can be based on preset rules; for example, if low brightness is detected in the video content, a metadata category related to low-light enhancement is selected. Alternatively, a simple lookup table mechanism can be used to map different content features to a predetermined combination of metadata categories. For example, when the video content is identified as "cartoon," the scheduler can be configured to select the "smooth texture" metadata category; when the video content is identified as "news broadcast," the scheduler can be configured to select the "text sharpening" metadata category. Content features include attribute information describing the dynamic characteristics and content composition of the video content. Attribute information may include basic information such as the video's frame rate, encoding format, and resolution. It may also include macroscopic descriptions obtained from preliminary analysis of the video content, such as whether the video is static or dynamic, and whether it is an indoor or outdoor scene.
[0046] For example, frame rate and encoding format can be obtained from the metadata of the video file, or the video content can be roughly classified into static or dynamic scenes. Therefore, a conditional super-resolution reconstructor is used to perform super-resolution processing on the video content based on the metadata categories in the target metadata subset. This processing can involve inputting selected metadata as additional input signals along with the low-resolution video content into the super-resolution model. For example, if the target metadata subset includes "edge information" metadata, this edge information can be directly superimposed onto the low-resolution video frame, and the superimposed image is then input into the super-resolution model for processing. Alternatively, if the target metadata subset includes "color correction parameters" metadata, these parameters can be used to adjust the color of the video content before or after super-resolution processing. Ultimately, the conditional super-resolution reconstructor outputs a second resolution display content, which is greater than the first resolution. Thus, the processed video content has a higher pixel density, resulting in a clearer and more detailed image on the display device.
[0047] For example, the display system includes a metadata orchestrator and a backbone-agnostic conditional super-resolution (SR) reconstructor. The metadata orchestrator performs feature analysis on the low-resolution input signal y, which serves as the first-resolution video content in the file. It extracts a context descriptor c, composed of high-level content identifiers and segment-level attributes. This descriptor represents the content features describing the dynamic characteristics and content composition of the video frame. Based on these features, the orchestrator selects the optimal metadata subset m(z) from a pre-defined set of K candidate metadata categories M, corresponding to the target metadata subset. The conditional super-resolution reconstructor then performs generative super-resolution processing on the low-resolution video content based on the selected optimal metadata subset, ultimately outputting a high-resolution estimate of the content to be displayed at the second resolution. Specifically, it is expressed as: z = π ( c ), Where z is an explicit mask, which supports the variability and missingness of metadata, making the display system robust under heterogeneous processing flow and runtime constraints, and the high resolution is higher than the original low resolution.
[0048] This embodiment achieves adaptive super-resolution processing of video content by introducing a metadata scheduler and a conditional super-resolution reconstructor. This scheme dynamically selects the most suitable subset of metadata based on the dynamic characteristics and content composition of the video content, effectively overcoming the limitations of traditional fixed metadata configuration schemes when processing diverse video content. Therefore, under resource budget constraints, this embodiment can optimize the quality and efficiency of super-resolution processing, improve the clarity and detail of displayed content, and enhance the flexibility and robustness of the display system.
[0049] In some embodiments, the content features include content feature identifiers and segment-level attributes; the content feature identifiers include at least one of the domain, subject matter, or channel type information of the video content; the segment-level attributes include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level.
[0050] As an example, content feature identifiers can correspond to lower-level content identifiers. Content feature identifiers can be attributes used for macro-level classification and identification of video content. They can be obtained from various sources, such as video metadata (e.g., file header information, streaming media tags), user viewing history, or content provider classifications. For example, a domain can refer to the macro-level category to which the video content belongs, such as movie, TV series, sports events, or news; a genre can be further refined into science fiction, comedy, action, or documentary; and a channel type refers to the specific channel from which the video content originates, such as "original content from channel XX." Identifiers help the metadata scheduler to initially understand the overall nature of the video and the expected viewing experience, thereby enabling coarse-grained metadata selection.
[0051] Fragment-level attributes describe the local and dynamic characteristics of video content in time or space. These attributes are calculated in real-time or offline using image processing and analysis algorithms on video frames or frame sequences. Text proportion measures the percentage of text regions in a video frame, which can be estimated using lightweight text detectors (e.g., optimized deep learning OCR models) or edge density combined with stroke heuristics (e.g., identifying text regions by analyzing image edge features and stroke structure). Object proportion measures the percentage of specific objects (e.g., faces, people, vehicles, specific products) in a video frame, which can be detected at a first resolution using compact detectors (e.g., lightweight object detection models like MobileNet and YOLO-tiny) to reduce computational complexity. Texture density measures the richness of texture details in a video frame, which can be characterized using gradient energy (e.g., calculating the average or sum of image gradient magnitudes) or local variance statistics (e.g., calculating the variance of pixel values in local image regions). Low-light noise level (e.g., low light or noise level) is used to measure the significance of noise in a video image under low light conditions. It can be estimated by using brightness distribution and noise proxy (e.g., by analyzing the brightness histogram distribution of the image and combining it with the statistical characteristics of high-frequency components to estimate the noise level).
[0052] For example, the proportion of the target object can be the proportion of text, face or person; the level of low light noise can be a low light or noise indicator; the content feature extracted by the metadata scheduler, i.e. the context descriptor c, can be decomposed into two parts g and a, which can be written as c=(g,a), where g is a high-level content identifier, which contains information such as the domain, theme, and channel type of the video content, and a is a segment-level attribute. The definition of a in the file explicitly includes dimensions such as motion intensity, text proportion, face or target object proportion, texture density, low light or noise indicator.
[0053] This application refines content features into content feature identifiers and segment-level attributes, enabling a more comprehensive and accurate understanding of the macro-type and micro-dynamic characteristics of video content. Content feature identifiers allow the metadata scheduler to perform preliminary, categorized metadata selection based on the overall attributes of the video (such as domain, subject matter, and channel type), ensuring that the selected metadata matches the general type of the video content. Simultaneously, segment-level attributes (such as motion intensity, text proportion, target object proportion, texture density, and low-light noise level) provide detailed information about the video content in both local and temporal dimensions, allowing the metadata scheduler to perform more refined and targeted metadata selection based on the specific details and dynamic changes of the video content. For example, for video segments with high motion intensity, metadata related to motion compensation can be prioritized; for segments with a high text proportion, text enhancement metadata can be emphasized; and for segments with high low-light noise levels, noise reduction metadata can be prioritized. Layered and detailed feature analysis significantly improves the accuracy and adaptability of target metadata subset selection, enabling the conditional super-resolution reconstructor to perform more effective super-resolution processing based on more accurate metadata guidance, ultimately outputting higher quality display content and optimizing the display effect of different types of video content.
[0054] In some embodiments, the metadata scheduler includes at least one of the following modules: a motion intensity calculation module for estimating the motion intensity of video content using block motion statistics, optical flow amplitude, or frame difference energy; a text detection module for estimating the text proportion of video content using a lightweight text detector or an edge density combined with a stroke heuristic; a target object detection module for estimating the target object proportion of video content using a compact detector at a first resolution; a texture density calculation module for characterizing the texture density of video content using gradient energy or local variance statistics; and a dark light noise detection module for estimating the dark light noise level of video content using brightness distribution and noise proxy.
[0055] As an example, a metadata scheduler may include a motion intensity calculation module for estimating the motion intensity of video content using block motion statistics, optical flow amplitude, or frame difference energy. For instance, block motion statistics divides video frames into macroblocks, calculates motion vectors between macroblocks in adjacent frames, and statistically analyzes the amplitude of these motion vectors to quantify the overall intensity of motion. Optical flow amplitude methods calculate the instantaneous motion vector of each pixel in the video; its magnitude directly reflects the intensity of local motion. Frame difference energy calculates the sum of the absolute or squared differences between pixel values in adjacent frames; a larger difference indicates more significant motion.
[0056] In addition, the metadata scheduler may include a text detection module for estimating the text proportion of video content using either a lightweight text detector or an edge density combined with stroke heuristics. Lightweight text detectors are typically optimized, small deep learning models capable of quickly identifying text regions in a video frame and calculating their proportion of the total frame. Edge density combined with stroke heuristics, on the other hand, utilizes image processing techniques to identify and locate text regions by analyzing the edge features of the image and the inherent stroke structure of the text characters.
[0057] To identify key content in the video, the metadata scheduler can further include a target object detection module to estimate the proportion of target objects in the video content at a first resolution using a compact detector. A compact detector refers to a target detection model with low computational resource consumption and a simplified model structure, capable of quickly and accurately identifying predefined target objects (such as faces, specific products, scene elements, etc.) in the video and calculating the area proportion of the target object in the frame without resolution upscaling.
[0058] To characterize the richness of detail in video footage, the metadata scheduler can include a texture density calculation module to represent the texture density of video content using gradient energy or local variance statistics. Gradient energy reflects the drasticness of local brightness changes by calculating the gradient magnitude of image pixels; the larger the gradient magnitude, the richer the texture. Local variance statistics calculate the variance of pixel values within local regions of the image; the larger the variance, the richer the texture detail in that region.
[0059] To address video quality issues in low-light environments, the metadata scheduler can also include a low-light noise detection module. This module estimates the level of low-light noise in the video content by analyzing brightness distribution and noise proxies. The low-light noise detection module first analyzes the brightness histogram of the video frames to identify low-brightness areas. Then, within these low-brightness areas, it quantifies the noise intensity by statistically analyzing pixel value fluctuations, local contrast, or using a specific noise model.
[0060] For example, the previously described 'a' includes the following: methods for estimating motion intensity through block motion statistics, optical flow amplitude, or frame difference energy constitute the core implementation of the motion intensity calculation module; methods for estimating text proportion through a lightweight text detector or edge density combined with stroke heuristics represent the functionality of the text detection module; methods for estimating the proportion of faces or target objects at low resolution (i.e., first resolution) using a compact detector correspond to the technical features of the target object detection module; methods for characterizing texture density through gradient energy or local variance statistics represent the implementation of the texture density calculation module; and methods for estimating the degree of dark light noise through brightness distribution and noise proxy constitute the core function of the dark light noise detection module.
[0061] In summary, the metadata scheduler can quantitatively analyze segment-level attributes of video content, such as motion intensity, text proportion, target object proportion, texture density, and low-light noise levels. This allows the metadata scheduler to obtain more comprehensive and detailed content feature information, enabling it to more accurately match the actual needs of the video content when selecting a subset of target metadata from multiple preset metadata categories. For example, for video content with high motion intensity, metadata categories related to motion compensation or temporal enhancement can be prioritized; for video content with a high text proportion, metadata categories that improve text clarity can be selected. This metadata selection mechanism based on refined content features significantly improves the targeting and effectiveness of super-resolution processing, avoiding blind or unnecessary metadata processing. This ensures the quality of displayed content while optimizing resource utilization efficiency, ultimately outputting visually superior second-resolution display content.
[0062] In some embodiments, Figure 2 This is a schematic diagram of the metadata scheduler provided in an embodiment of this application. For example... Figure 2 As shown, the metadata scheduler also includes a mask generation module, which generates a corresponding selection mask based on the target metadata subset, so that the conditional super-resolution reconstructor can determine the metadata categories included in the target metadata subset based on the selection mask. The selection mask is used to identify the metadata categories included in the target metadata subset.
[0063] As an example, the mask generation module transforms the target metadata subset selected by the metadata scheduler into a structured, easily parsed format—a selection mask. This mask generation module can be a software module that receives the identification information of the target metadata subset as input and transforms it into a standardized data structure according to preset encoding rules. For example, if multiple metadata categories are preset, the mask generation module can generate a binary bitmask, where each bit represents a predefined metadata category, and the bit's state (e.g., 0 or 1) indicates whether the category is selected. Furthermore, the selection mask can also be a list containing the index numbers or unique identifiers of the selected metadata categories, or a sparse vector where only the dimension corresponding to the selected metadata category has a non-zero value.
[0064] The selection mask explicitly identifies the various metadata categories contained in the target metadata subset. It serves as the interface between the metadata scheduler and the conditional super-resolution reconstructor, ensuring the accuracy and efficiency of information transmission. Through this explicit identification, the conditional super-resolution reconstructor no longer needs to analyze or infer the content of the target metadata subset itself; instead, it directly obtains the required information by parsing the selection mask. The conditional super-resolution reconstructor may internally contain a parser or decoder to read and interpret the selection mask. For example, when the selection mask is a binary bitmask, the parser iterates through each bit of the mask, activating or disabling the corresponding metadata processing branch or module based on the bit's state. This ensures that super-resolution processing can precisely adjust its behavior according to the metadata scheduler's decisions, thereby achieving adaptive processing of the video content.
[0065] For example, one of the core designs of the MetaSR framework is that a binary selection mask z is generated by the metadata scheduler, z=(z1,…,zK)∈{0,1}, where K represents selection. The mask is generated based on the optimal metadata subset m(z) selected by the scheduler, m(z)={mk|zk=1}, which is the target metadata subset. zk=1 indicates that the corresponding metadata category is enabled, and zk=0 indicates that it is disabled. The mask directly identifies the target metadata subset. z=π(c) is generated by the scheduler, which maps the context c to the metadata selection mask z. The reconstructor takes (y,m(z),z) as input and determines the selected metadata category based on the mask z. (See also...) Furthermore, the binary selection mask z is used to identify the selected metadata category in the candidate metadata category set. The conditional super-resolution reconstructor can know the currently selected metadata category by parsing the mask z.
[0066] This application introduces a mask generation module, enabling the metadata scheduler to generate a corresponding selection mask based on the selected target metadata subset. This selection mask identifies the metadata categories contained in the target metadata subset in a clear and standardized manner, allowing the conditional super-resolution reconstructor to efficiently and accurately determine and utilize the required metadata categories based on this selection mask. This avoids the conditional super-resolution reconstructor performing additional parsing or inference on the metadata subset during processing, significantly improving the efficiency and accuracy of super-resolution processing and ensuring that super-resolution processing accurately responds to the metadata scheduler's decisions, thereby optimizing resource utilization and improving the quality of displayed content.
[0067] In some embodiments, such as Figure 2As shown, the metadata scheduler also includes a gain prediction module and a greedy selection module; the gain prediction module is used to predict the expected quality gain of each metadata category under the current content features; the greedy selection module is used to calculate the unit cost quality gain index of each metadata category, and select the metadata category according to the unit cost quality gain index of each metadata category until the preset resource budget constraint is met, and generate a selection mask.
[0068] The gain prediction module uses a small network to predict the expected quality gain of each metadata category under the current content features. For example, it can estimate the potential improvement in display quality after video content super-resolution processing based on historical data, machine learning models, or predefined rules. For instance, predicting temporal motion cue metadata may bring a higher quality gain for specific content features (such as high-motion scenes), while text region metadata may be more valuable for scenes containing a large amount of text. The expected quality gain can be quantified as an improvement in PSNR (Peak Signal-to-Noise Ratio) or SSIM (Structural Similarity Index), or represented by predicted values of user-perceived quality (such as MOS score).
[0069] The greedy selection module chooses the optimal metadata category at each step, i.e., the category with the highest quality-per-unit-cost (MPC) gain. The MPC gain is typically calculated by dividing the expected quality gain by the resource overhead (e.g., computational complexity, memory usage, or processing latency) of that metadata category under the current content characteristics. The preset resource budget constraint refers to the maximum resource consumption limit the system can withstand when selecting metadata categories. This may include, but is not limited to, total computation (e.g., FLOPs), total memory usage, total processing latency, or power consumption. This constraint is a threshold pre-set during system design based on hardware capabilities, real-time requirements, or user experience goals. The greedy selection module continuously accumulates the resource overhead of the selected metadata categories. Once the accumulated value reaches or exceeds the preset resource budget constraint, the selection process stops, ensuring the system operates within acceptable resource limits. The MPC gain describes the ratio between the improvement in super-resolution quality brought by selecting a particular metadata category and the computational resource cost (including bandwidth, computation, or latency) incurred by using that metadata. The MPC gain can be selected in descending order. Resource budget constraints can refer to resource limitations that are dynamically adjusted based on the contextual characteristics of the video content when allocating resources for super-resolution processing.
[0070] For example, when the metadata scheduler generates the selection mask, it uses a small network hω( A learning-based gain predictor is constructed to predict the gain score Δk(c) of each metadata category under the current content feature, i.e., the context descriptor c. This predictor is the gain prediction module in the claims, realizing the prediction function of expected quality gain. Simultaneously, the scheduler calculates the normalized score sk(c) of the unit cost quality gain based on the gain score and the abstract cost function of each metadata category, and selects metadata categories in descending order of this score using a greedy selection strategy until the preset resource budget constraint B(c) in the file is met. Finally, a binary selection mask z is generated, and it is explicitly stated that the selection problem is equivalent to a 0 / 1 knapsack problem, with the core being gain maximization under budget constraints. Specifically, a small network is used... h ω ( Implement the scheduler π ( c ), which predicts gain scores for each metadata category. ( c ), Then the normalized score is calculated. , and according to s k ( c Greedily select categories in descending order until the budget constraint B(c) is satisfied, generating a binary mask z. Let... Δ k ( c The expression represents the (unknown) expected quality gain resulting from enabling category k in context c. A natural selection principle is... suchthat ≤ B ( c This problem is equivalent to the classic 0 / 1 knapsack problem.
[0071] This application effectively addresses the problem of optimizing metadata selection under limited resources by introducing a gain prediction module and a greedy selection module. The gain prediction module intelligently assesses the potential contribution of different metadata categories to display quality based on video content characteristics, providing a quantitative basis for subsequent decisions. Building upon this, the greedy selection module calculates a unit cost quality gain index, prioritizing metadata categories that deliver the greatest quality improvement with minimal resource consumption at each selection step in an efficient and pragmatic manner. This dynamic selection mechanism, based on expected quality gain and resource budget constraints, enables the metadata scheduler to adaptively generate the optimal selection mask according to the characteristics of the current video content and system resource status. Ultimately, this not only ensures that super-resolution processing achieves significant display quality improvements even in resource-constrained environments but also avoids unnecessary resource waste, thus achieving both high performance and economic efficiency in system operation.
[0072] In some embodiments, such as Figure 2 As shown, the metadata scheduler also includes a hybrid rules module, which introduces domain heuristic rules to supplement the selection logic of the greedy selection module; the heuristic rules include selecting text region metadata when the text proportion exceeds a threshold and selecting temporal motion clue metadata when the motion intensity exceeds a threshold.
[0073] The hybrid rules module is configured to introduce domain-specific heuristics into the metadata selection process. Its primary function is to provide the greedy selection module with additional decision-making support based on experience or expert knowledge, compensating for the limitations of purely quantitative indicator-based selection. Through the hybrid rules module, the system can more intelligently identify strong correlations between specific content features and specific metadata categories, ensuring that, within the resource budget, metadata that has a decisive impact on a specific content type is prioritized. Domain-specific heuristics are formulated based on a deep understanding of video content characteristics and their impact on super-resolution processing. They typically exist in the form of "if condition A is met, then perform operation B," aiming to capture decision logic that is difficult to reflect through simple cost-benefit analysis but is crucial to visual quality. Rules can be predefined and stored in the system and invoked when the metadata scheduler makes selections.
[0074] As an example, when the proportion of text in the video content (e.g., calculated by the text detection module) exceeds a preset threshold, it indicates the presence of a large amount of text information in the video. In this case, even if the unit cost quality gain of the text region metadata is not the highest, the hybrid rule module will force or prioritize the selection of text region metadata. Selecting text region metadata helps to better preserve and enhance the clarity and edge sharpness of text in super-resolution processing, avoiding text blurring or distortion, thereby significantly improving the user's reading experience. This threshold can be set according to the actual application scenario and user experience requirements; for example, it can be triggered when the text proportion reaches 5% or 10%. Simultaneously, when the motion intensity of the video content (e.g., estimated by the motion intensity calculation module) exceeds a preset threshold, it indicates the presence of intense or rapid motion in the video. In this case, the hybrid rule module will force or prioritize the selection of temporal motion cue metadata. Temporal motion cue metadata can provide the conditional super-resolution reconstructor with information about the trajectory and speed of object motion, thereby better handling motion blur, artifacts, and other issues in super-resolution processing, ensuring the smoothness and clarity of moving images. This threshold can also be adjusted according to the type of video content (e.g., sports events, action movies) and the desired visual effects.
[0075] For example, to enhance the robustness and interpretability of the greedy selection strategy of the metadata scheduler, a domain heuristic rule is introduced to initialize and supplement the greedy selection logic based on the learned gain predictor. The implementation of this rule corresponds to the hybrid rule module in the claims. When the text proportion exceeds a threshold, the metadata of the text region is forcibly selected. m text When the motion intensity exceeds a threshold, temporal motion cue metadata should be forcibly selected. m temp .
[0076] This application effectively addresses the limitations that may arise when selecting metadata solely based on quality gain per unit cost by introducing a hybrid rule module and supplementing the selection logic of the greedy selection module with domain-based heuristics. For example, when video content has significant features such as high text content or high motion intensity, even if the corresponding text region metadata or temporal motion cue metadata might not be prioritized in a conventional greedy selection, the hybrid rule module ensures that key metadata is included in the target metadata subset. This allows the conditional super-resolution reconstructor to obtain more targeted auxiliary information when processing specific types of video content, resulting in clearer and sharper text display in text regions and smoother, less blurry dynamic images in motion scenes. The hybrid selection strategy, combining quantitative indicators and domain experience, not only optimizes the intelligence of metadata selection but also significantly improves the super-resolution display quality of display devices when processing diverse video content, especially in the detail representation of specific content elements (such as text and motion), providing a superior visual experience.
[0077] In some embodiments, Figure 3 This is a schematic diagram of a conditional super-resolution reconstructor provided in an embodiment of this application. For example... Figure 3 As shown, the conditional super-resolution reconstructor includes a metadata encoder and a feature fusion module. The metadata encoder is used to extract features from heterogeneous metadata in the target metadata subset and generate metadata features in a unified format. The heterogeneous metadata includes dense spatial graphs, temporal signals, compact vectors, or scalars. The feature fusion module is used to fuse the metadata features with the features of the video content to obtain enhanced features, enabling the conditional super-resolution reconstructor to perform super-resolution processing on the video content based on the enhanced features.
[0078] The metadata encoder is a key component for processing heterogeneous metadata. Its main function is to transform heterogeneous metadata from different metadata categories, with different data structures and semantics, into a unified metadata feature representation that can be processed by the subsequent feature fusion module. For example, for dense spatial graphs (such as semantic segmentation graphs and saliency graphs), the metadata encoder can use a convolutional neural network (CNN) structure for feature extraction, converting it into a high-dimensional feature map or feature vector; for temporal signals (such as motion vectors and inter-frame differences), it can use temporal models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or Transformers to capture their temporal dependencies and generate temporal feature vectors; for compact vectors (such as scene classification vectors and style description vectors), it can perform dimensionality transformation and nonlinear mapping through fully connected layers or multilayer perceptrons (MLPs); for scalars (such as mean brightness and contrast), they can be embedded into a high-dimensional space, for example, by converting them into vector representations through linear layers or lookup tables. In this way, the metadata encoder effectively eliminates the heterogeneity between metadata, laying the foundation for subsequent feature fusion.
[0079] The feature fusion module is responsible for effectively combining the unified-format metadata features output by the metadata encoder with the features of the original video content, thereby generating enhanced features containing richer contextual information. These enhanced features serve as input to the conditional super-resolution reconstructor for super-resolution processing. Feature fusion can be implemented in various ways. For example, concatenation fusion can connect metadata features and video content features along the channel dimension to form a wider feature tensor, which is then used to learn the interactions between features through convolutional layers. Addition fusion or multiplication fusion can also be used, performing element-wise operations after dimension matching to achieve feature modulation. More complex fusion mechanisms can include gating fusion or cross-modal attention, allowing metadata features to act as gating signals or query vectors, dynamically adjusting the weights and representations of video content features to learn deeper, more complex relationships between them. Furthermore, feature fusion can be performed at multiple layers of the super-resolution network to provide more refined guidance. Enhanced features can be fused using concatenation fusion or feature-level modulation methods. For example, feature-level modulation uses the FiLM gating mechanism, which generates layer-by-layer modulation parameters through metadata embedding, and uses the modulation parameters to perform affine transformation on the intermediate feature map of the video content to achieve feature modulation.
[0080] For example, the conditional super-resolution reconstructor sets up a unified metadata encoder ε to handle heterogeneous metadata. This encoder extracts features from heterogeneous metadata, including dense spatial graphs, time-series signals, compact vectors, and scalars, in candidate metadata categories, generating metadata features ek in a unified format. e k =ε k ( m k ),for z k =1 Simultaneously, the reconstructor incorporates a feature fusion stage. Through methods such as concatenation fusion or feature-wise modulation (e.g., FiLM / gating mechanisms), it fuses metadata features in a unified format with visual features extracted from low-resolution video content, resulting in enhanced features containing richer contextual information. Based on these enhanced features, super-resolution processing is then performed on the video content. As an example, feature-wise modulation utilizes metadata embedding to generate layer-by-layer modulation parameters, which can be found in [reference needed]. F′=γ ( e,z )⊙ F + β ( e,z FiLM-style modulation is particularly effective when metadata contains a mixture of spatial and non-spatial signals, and supports variable metadata inputs through an explicit mask z.
[0081] This application effectively addresses the challenge of directly applying heterogeneous metadata to super-resolution processing. The metadata encoder unifies different types of metadata into standardized metadata features, eliminating format and semantic differences. Subsequently, the feature fusion module deeply integrates the unified metadata features with video content features, generating enhanced features rich in contextual information. These enhanced features not only include the visual information of the video itself but also incorporate high-level semantic information provided by the metadata, such as dynamic characteristics and content composition. Therefore, when performing super-resolution processing, the conditional super-resolution reconstructor can perform targeted detail reconstruction and texture enhancement of video content based on more comprehensive and accurate enhanced features. For example, it can achieve clearer text reconstruction in text areas, reduce artifacts in motion areas, or enhance detail in specific target object areas. This improves the quality and efficiency of super-resolution processing, resulting in visually clearer and more natural output at second-resolution display, better preserving the semantic information of the original video content, thus providing users with a superior viewing experience.
[0082] In some embodiments, each metadata category corresponds to an abstract cost function, which is a non-negative function that depends on content features and is used to characterize the resource overhead of metadata under the corresponding metadata category; the total cost corresponding to the selection mask is the sum of the costs of each selected metadata category under the current content features.
[0083] As an example, the abstract cost function is a mathematical model used to quantify the resource consumption required to process a specific metadata category. Resource overhead can include, but is not limited to, computation time, memory usage, bandwidth requirements, latency, or power consumption. The function is designed to depend on the content characteristics of the current video content; for example, for video content with high motion intensity, the cost of processing temporal motion cue metadata may be higher than the cost of processing static texture metadata. Furthermore, the function is set to be non-negative to ensure that the quantification of resource overhead is practically meaningful. In implementation, the abstract cost function can be predicted using a pre-trained model (e.g., a machine learning-based model) that takes content features as input and outputs an estimate of the resource overhead for the corresponding metadata category; alternatively, it can be determined using empirical formulas or lookup tables based on extensive testing and statistical analysis of the actual resource consumption of different metadata categories under different content characteristics.
[0084] When the metadata scheduler generates a selection mask to identify the metadata categories contained in the target metadata subset, it calculates a total cost. The total cost is obtained by summing the abstract cost function values of each selected metadata category under the current content features. The total cost reflects the total resource consumption required to process the entire target metadata subset, providing a quantitative basis for subsequent resource budget management and optimization. For example, if three metadata categories A, B, and C are selected, and their abstract costs under the current content features are Cost(A), Cost(B), and Cost(C), respectively, then the total cost corresponding to the selection mask is Cost(A) + Cost(B) + Cost(C).
[0085] As an example, the function of a selection mask is to use a binary identifier consisting of 0s and 1s. "1" represents that the corresponding metadata category is selected, and "0" represents that it is not selected. The cost of a single metadata category: Each metadata category corresponds to an abstract cost function. The cost value of the function is not fixed and changes dynamically based on the characteristics of the input content. For example, the cost of motion cue metadata in a high-speed motion segment is higher than in a static segment. The logic for calculating the total cost: It iterates through all metadata categories marked "1" in the selection mask, extracts the cost value of each metadata category under the current content features, and adds the values together. The result is the total cost corresponding to the selection mask. For example, suppose there are three categories of candidate metadata: degradation cues, structural and edge information, and temporal motion cues. In a static landscape segment, the mask is [1,1,0], representing the selection of the first two metadata categories. If the cost of degradation cues is 2, the cost of structural and edge information is 3, and the cost of temporal motion cues is 5, then the total cost corresponding to this selection mask = 2 + 3 = 5.
[0086] For example, to achieve budget-aware optimization of metadata selection, an abstract cost function costk(c) is defined for each metadata category k. This function is a non-negative function that depends on the content features, i.e., the context descriptor c, and represents the resource overhead in terms of bandwidth, computation, latency, etc., when using the corresponding metadata category. At the same time, the total cost corresponding to the selection mask z is defined in the file as Cost(z;c)=∑zk costk(c), where z_k is the binary selection mask, and costk(c) is the cost of category k in context c. That is, the sum of the costs of the selected metadata categories under the current content features. As an example, let costk(c) ≥ 0 represent the context-dependent cost of using metadata category k (e.g., a proxy metric of bandwidth, computation, or latency). The total cost of the selection mask z is... The goal is to maximize super-resolution quality under resource constraints. This can be written as a constrained optimization problem: , making z = π ( c ), Cost ( z ; c )≤ B ( c ), where L( B(c) is the reconstruction loss (e.g., mean squared error / perceptual / ROI weighted loss and temporal loss for video), and B(c) is the context-dependent budget. Equivalently, a learning-friendly Lagrangian (penalized) form is used. where z = π ( c), where λ>0 represents the trade-off between control quality and overhead.
[0087] This application incorporates an abstract cost function and calculates the total cost corresponding to the selection mask, thereby taking resource overhead into account when selecting metadata. This allows the metadata scheduler to not only improve display quality based on content features when selecting a subset of target metadata, but also to simultaneously assess and manage the required computing resources. For example, in resource-constrained environments, metadata categories with lower resource overhead can be prioritized while ensuring a certain quality gain, thus achieving optimal super-resolution results within a limited resource budget. This avoids the problem of resource exhaustion or performance degradation caused by blindly selecting high-cost metadata, significantly improving the adaptability and efficiency of display devices under different operating conditions, and ensuring the stability and sustainability of super-resolution processing.
[0088] In some embodiments, the conditional super-resolution reconstructor is integrated into the diffusion model to generate second-resolution display content through an iterative denoising process by using video content, a subset of target metadata, and a selection mask as conditional inputs to the denoiser.
[0089] As an example, the conditional super-resolution reconstructor is designed with a diffusion model-based architecture. A diffusion model is a generative model that learns the distribution of data by simulating the process of gradually denoising data from noise. In super-resolution tasks, a diffusion model learns how to recover a high-resolution image from a low-resolution image (or its noisy representation). This ensemble typically involves a deep neural network (e.g., U-Net) acting as a denoiser, which receives the current noisy image and conditional information at each time step and predicts the noise to be removed. By iteratively performing the denoising process, the model can progressively transform low-resolution input into high-resolution output. Applying a diffusion model to super-resolution reconstruction effectively improves the processed visual quality and realism, especially when dealing with complex textures and details, generating more natural and realistic high-resolution images.
[0090] In the above scheme, video content, a subset of target metadata, and a selection mask are used as conditional inputs to the denoiser. Conditional inputs refer to providing additional guiding information beyond the current noisy image during the diffusion model's denoising process, directing the generation process towards a specific target or satisfying specific conditions. The video content, i.e., the original first-resolution video content, can be directly used as the input feature map to the denoiser, or its features can be extracted by the encoder and then used as conditions. For example, low-resolution video frames can be encoded using convolutional layers to obtain their spatial feature representations, which can then be fused with the intermediate layer features of the diffusion model, for example, through cross-attention mechanisms or feature concatenation. The subset of target metadata may contain various heterogeneous data, such as dense spatial maps, temporal signals, compact vectors, or scalars. The metadata needs to be processed by a metadata encoder to extract features, generating metadata features in a uniform format. Metadata features can be used as global conditions, for example, injected into different layers of the denoiser through adaptive layer normalization (AdaLN) or FiLM layers, or interact locally with video features through attention mechanisms. The selection mask is a binary vector or tensor used to indicate which metadata categories are selected within the subset of target metadata. It can serve as additional conditional input, for example, by embedding it as a vector and injecting it into the denoiser along with metadata features, or by directly controlling the activation or weighting of metadata features, enabling the denoiser to perceive which metadata is valid and which is invalid. By using multimodal information as conditional input, the denoiser can gain a more comprehensive understanding of the characteristics of the video content, the key areas or attributes of interest to the user or system, and which metadata is key guiding information for the current super-resolution task. This allows super-resolution processing to be finely tuned according to the specific needs of the content and the indications of the metadata, thereby generating higher-quality second-resolution display content that better meets expectations.
[0091] For example, using the diffusion model as the main implementation scheme of the conditional super-resolution reconstructor, the reconstructor is integrated into the diffusion model, and the low-resolution input y (as video content), m(z) (as a subset of target metadata), and z (as a selection mask) are collectively used as the denoising unit εθ in the diffusion model. The model takes the joint conditional input of the diffusion model and, through an iterative denoising reverse process, gradually recovers the high-resolution video content from the noisy low-resolution content, ultimately generating a high-resolution estimate as the second-resolution display content. As an example, let x0 represent a high-resolution target image in single-image super-resolution (ISR), or a high-resolution frame in video super-resolution (VSR). The diffusion model defines a Markov forward process that progressively adds Gaussian noise to x0 over T steps: The diffusion model defines the forward process as follows: Where α_t∈(0,1) is a predefined noise schedule, and The reverse process learns through a neural network denoiser, which predicts the added noise (or equivalent parameterization), conditioned on low-resolution observations y and scheduling metadata. In MetaSR, the conditional input is a tuple (y, m(z), z). in, ε θ ( Typically, it uses a U-Net structure. The mask z explicitly informs the denoiser of the currently available metadata categories, enabling a single model to robustly handle variable and missing metadata.
[0092] This application generates second-resolution display content through an iterative denoising process. Starting with a random noise sample of the same size as the target second-resolution image, a denoiser is applied progressively across multiple time steps. At each time step t, the denoiser receives the current noisy image Xt, the encoding of time step t (e.g., by sinusoidal positional encoding), and the aforementioned video content, a subset of target metadata, and a selection mask as conditional inputs. The denoiser predicts the noise εt contained in the current noisy image. Then, based on the predicted noise εt and the sampling strategy of the diffusion model (e.g., DDPM, DDIM, etc.), the image is updated to Xt-1 to make it closer to the real image. This iterative process is repeated until a preset minimum time step (typically 0) is reached, and the final image obtained is the second-resolution display content. The iterative denoising process allows the model to progressively refine image details across multiple stages, from coarse structure to fine texture, thereby generating a high-fidelity and realistic super-resolution image. The introduction of conditional input ensures that the entire denoising process is always guided by content features and metadata, resulting in a final output of second-resolution display content that is not only higher in resolution but also significantly improved in terms of visual quality and content consistency.
[0093] This application fully leverages the powerful image generation and restoration capabilities of the diffusion model by integrating a conditional super-resolution reconstructor into the diffusion model and using video content, a subset of target metadata, and a selection mask as conditional inputs to the denoiser. The iterative denoising process of the diffusion model allows it to progressively recover high-resolution details from low-resolution content. Simultaneously, the introduction of multimodal conditional inputs ensures that the reconstruction process is always precisely guided by content features and metadata selection. This results in generated second-resolution display content that not only has higher resolution but also significantly improved visual quality, detail richness, and content consistency, effectively avoiding artifacts or distortions that may occur in traditional super-resolution methods, thus providing users with a clearer, more realistic, and content-consistent viewing experience.
[0094] In some embodiments, a conditional super-resolution reconstructor is integrated into a convolutional neural network or generative adversarial network architecture to input fused enhanced features into the convolutional neural network backbone to output display content at a second resolution.
[0095] As an example, conditional super-resolution reconstructors can be integrated into convolutional neural network (CNN) or generative adversarial network (GAN) architectures. A CNN is a feedforward neural network that automatically learns features from input data through structures such as convolutional layers, pooling layers, and fully connected layers. In super-resolution reconstruction tasks, CNNs can learn the complex mapping relationship from first-resolution video content to second-resolution display content, gradually recovering image details through multi-layer nonlinear transformations. For example, advanced CNN structures such as residual networks, densely connected networks, or attention mechanisms can be used to enhance feature extraction and information transfer capabilities. A generative adversarial network consists of a generator and a discriminator, which optimize each other through adversarial training. The generator aims to generate realistic high-resolution images to deceive the discriminator; the discriminator attempts to distinguish between real high-resolution images and those generated by the generator. This adversarial training mechanism enables GANs to excel in generating super-resolution images with realistic textures and details, effectively mitigating the problem of blurred results easily produced by traditional CNN models trained based on the mean squared error (MSE) loss function.
[0096] The fused enhanced features refer to the comprehensive feature representation obtained by fusing metadata features from a subset of target metadata with features from the video content. The enhanced features contain visual information from the video content itself, as well as semantic and structural information provided by the metadata. The fused enhanced features are input into the backbone of a convolutional neural network (CNN), serving as the primary input to the deep learning model and guiding the network's learning and reconstruction process. The CNN backbone is the part of a CNN or GAN architecture responsible for core feature extraction, transformation, and upsampling. For example, in a pure CNN architecture, it refers to the main feature extraction and upsampling network; while in a GAN architecture, it typically refers to the main body of the generator network. Input methods can be varied; for example, the enhanced features can be concatenated with the first-resolution video content at the network's input layer, or injected into intermediate layers of the network via skip connections to provide conditional information at different levels of abstraction. These input methods ensure that metadata information is integrated throughout the entire super-resolution reconstruction process, enabling fine-grained control and optimization of the reconstruction results.
[0097] The output of the second-resolution display content refers to the high-resolution image or video frame reconstructed from the first-resolution video content after processing by a convolutional neural network or generative adversarial network architecture. The second resolution is significantly larger than the first resolution, representing an improvement in image detail and sharpness. This output is the final result of the super-resolution reconstruction task, and its quality directly reflects the performance of the entire system. Through the learning capabilities of deep learning architectures, the system can generate display content with richer details, fewer artifacts, and better conformity to human visual perception, thereby enhancing the user's viewing experience.
[0098] For example, the MetaSR framework is backbone-independent, and the conditional super-resolution reconstructor can be directly integrated into a convolutional neural network (CNN) or generative adversarial network (GAN) architecture. In this integration scheme, the reconstructor injects the enhanced features (i.e., the fused conditional representation h) obtained through metadata encoding and fusion into the CNN backbone network via input concatenation or feature-level modulation. The CNN backbone network then performs the super-resolution processing and outputs a high-resolution estimate of the content to be displayed at the second resolution. As an example, given a low-resolution input signal y, a context descriptor c, and a scheduler output z=π(c), a CNN-based reconstructor generates a high-resolution output: Similar to the diffusion model, z explicitly indicates the currently available metadata category, enabling a single CNN backbone to unambiguously handle variable or missing metadata inputs. CNN-based super-resolution needs to handle heterogeneous metadata formats. A unified encoding strategy is adopted: spatial metadata (such as edge maps, text / ROI masks, semantic segmentation) is represented as an image-aligned mapping, which can be concatenated with y at the input or encoded into intermediate feature maps. Non-spatial metadata (such as degenerate scalars or compact descriptors) is embedded as feature vectors through a multilayer perceptron (MLP) and injected into the CNN through feature-level modulation. Formally, the metadata encoder E(·) is defined to generate conditional representations: h = E ( m ( z ), z ), where h can contain spatial conditional mapping and global embedding. The conditionalization mechanisms widely used in both standards are compatible with common SR backbone networks (such as EDSR): Input Concatenation, for spatial metadata, constructs enhanced input: This is the simplest integration method, providing strong baseline performance. Feature-level modulation (FiLM / gating mechanism) is used to inject conditional embeddings into the residual blocks through feature-level affine modulation for more flexible fusion of spatial and non-spatial metadata. F′=γ ( h )⊙ F + β (h ), where F is the intermediate feature map, γ ( )and β ( For a small network, h is mapped to channel-level modulation parameters, and ⊙ represents element-wise multiplication. This mechanism allows metadata to act as a "control signal," adaptively adjusting reconstruction behavior to match the content context. The CNN-based implementation uses the standard reconstruction loss, with optional enhancements to ROI sensitivity or structure-aware terms. z = π ( c To improve robustness to changes in metadata selection, a random metadata discarding strategy is employed during training: random sampling or perturbation of z exposes the CNN backbone to diverse combinations of metadata (including partially missing signals). This strategy enhances inference stability, especially when the scheduler selects different subsets of metadata across segments. This CNN implementation highlights the important characteristic of MetaSR: the scheduling layer π(c) is orthogonal to the backbone network selection.
[0099] This application integrates a conditional super-resolution reconstructor into a convolutional neural network or generative adversarial network architecture, using the fused enhanced features as input, effectively leveraging the powerful feature learning and reconstruction capabilities of deep learning. This integration allows the system to capture the complex correlation between video content and adaptive metadata more precisely, thereby generating second-resolution display content with richer details, higher visual quality, and a better match to content features during super-resolution processing. Compared to traditional super-resolution methods based on interpolation or simple filtering, this approach significantly improves the robustness and image quality of super-resolution reconstruction, especially when processing videos with complex dynamic characteristics and diverse content. It better utilizes metadata information to guide the reconstruction process, avoiding poor reconstruction results caused by insufficient metadata utilization or improper integration methods.
[0100] In some embodiments, the conditional super-resolution reconstructor is integrated into the generative adversarial network (GAN) to input a selected subset of target metadata into the generator and discriminator in the GAN, respectively. The generator generates second-resolution display content based on each metadata in the subset of target metadata, and the discriminator determines the authenticity of the display content based on the metadata. Metadata consistency loss is introduced to constrain the output details of the display content to match the metadata.
[0101] As an example, a Generative Adversarial Network (GAN) is a deep learning model consisting of a generator and a discriminator. The generator is responsible for generating data samples from random noise or given conditions, while the discriminator is responsible for distinguishing the generated data from real data. Through adversarial training between the generator and the discriminator, the generator can learn to generate highly realistic data that conforms to specific conditions. Integrating a conditional super-resolution reconstructor into a GAN leverages the powerful generative capabilities of GANs to improve the quality of super-resolution processing, enabling it to generate more realistic and detailed second-resolution display content. This integration makes the super-resolution task not just a simple pixel interpolation, but an intelligent reconstruction and enhancement of content details.
[0102] The selected subset of target metadata, serving as conditional information, is provided to both the generator and the discriminator. For the generator, the metadata provides key attribute information about the video content, such as motion intensity, text proportion, and texture density. The generator can use this conditional information to guide its generation process, ensuring that the generated second-resolution display content maintains consistency in detail with the features described by the metadata. For example, if the metadata indicates that the video content contains a large amount of text, the generator will tend to generate clear and sharp text areas. For the discriminator, the metadata is used not only to judge the overall realism of the generated content but also to evaluate whether the generated content conforms to specific metadata conditions. While judging realism, the discriminator also determines whether the generated content "looks like" real content with metadata characteristics.
[0103] The generator receives video content at a first resolution and a subset of target metadata as input. Using this metadata as a guiding signal, it upscales the low-resolution video content to a second resolution through a series of convolutional layers, upsampling layers, and other network structures. The metadata acts as a "content prior," helping the generator to perform targeted enhancements based on the specific characteristics of the video content (such as texture, edges, and motion) when reconstructing high-frequency details, rather than blindly interpolating or generalizing. For example, for regions containing fine textures, the generator will focus on reconstructing texture details based on metadata instructions; for scenes with rapid motion, it will focus on the smoothness and sharpness of motion trajectories.
[0104] The discriminator receives the second-resolution display content (or real second-resolution content) output by the generator, along with a corresponding subset of target metadata. The discriminator's task is to distinguish whether the input content is real or generated by the generator, and simultaneously evaluate whether the content matches the input metadata conditions. In this way, the discriminator learns not only the statistical realism of the image but also the semantic consistency between the image content and the metadata. For example, if the metadata indicates that the video content has high motion intensity, but the generator's output image appears blurry or lacks motion, the discriminator will identify it as "fake" or "inconsistent."
[0105] Metadata consistency loss is an additional loss function used to quantify the degree of matching between the generator's output second-resolution display content and the input target metadata subset. The loss can be implemented in various ways; for example, features (such as texture features, motion features, text features, etc.) can be extracted from the generated content, and then compared with the original target metadata. If a significant difference exists, the consistency loss increases, prompting the generator to adjust its generation strategy so that the output display content more accurately reflects the characteristics indicated by the metadata in detail. For example, an auxiliary network can be designed to predict the metadata of the generated image, and its predictions can be compared with the true metadata to calculate the loss.
[0106] For example, when a conditional super-resolution reconstructor is integrated into a Generative Adversarial Network (GAN), a selected subset of target metadata, m(z), is simultaneously input into both the generator and discriminator of the GAN as a common conditional input. The generator processes low-resolution video content based on the target metadata subset to generate high-resolution video as the second-resolution display content. The discriminator then judges the realism of the generator's output based on the target metadata subset. To constrain the matching degree between the generated content and the metadata, a metadata consistency loss is introduced during training to force the details of the output content to remain aligned with the metadata information. As an example, when a conditional super-resolution reconstructor is integrated into a GAN, the generator generates perceptibly realistic HR images, and the discriminator encourages realistic textures. The scheduler selects a context-dependent subset of metadata, m(z), using z = π(c), and conditionally conditions the generator (and the discriminator). During training, z is randomly sampled / perturbed (random metadata is discarded), allowing the generator and discriminator to observe diverse combinations of metadata. Combined with explicit masking conditionalization, a single GAN model can adapt to different scheduling decisions during inference.
[0107] This example integrates a conditional super-resolution reconstructor into a generative adversarial network (GAN), and makes both the generator and discriminator use a subset of the target metadata as conditional input. This significantly improves the quality and realism of super-resolution processing. Guided by the metadata, the generator can produce second-resolution display content that highly matches the content features, avoiding detail distortion or artifacts that may occur in traditional super-resolution methods. Simultaneously, the discriminator can not only identify the visual realism of the generated content but also ensure its consistency with the metadata conditions. Furthermore, a metadata consistency loss is introduced, directly forcing the output details of the generated content to precisely match the metadata during training, effectively solving the problem of semantic discrepancies between the super-resolution content and the original metadata. This results in the final output display content not only having higher resolution but also being highly consistent with user expectations or content characteristics in key content features such as texture, motion, and text, greatly improving the viewing experience and the accuracy of the display effect.
[0108] Figure 4 This is a flowchart illustrating the content-adaptive metadata-based display method provided in this application embodiment; as follows: Figure 4 As shown, the display method operates on a display device and includes: S401: Perform feature analysis on the video content to be displayed at the first resolution, extract the content features of the video content, and select a subset of target metadata from multiple preset metadata categories based on the content features; wherein, the content features include attribute information used to describe the dynamic characteristics and content composition of the video content.
[0109] S402: Perform super-resolution processing on the video content based on each metadata category in the target metadata subset, and output the display content at a second resolution, which is greater than the first resolution.
[0110] The core innovation of this embodiment lies in combining content feature analysis with an adaptive metadata selection mechanism. This allows for the dynamic selection of the most suitable subset of metadata based on the video content's dynamic characteristics and composition, avoiding the limitations of fixed metadata configuration schemes when handling diverse video content. This achieves the effect of optimizing super-resolution quality and efficiency under resource budget constraints. As an example, the method first performs feature analysis on the video content to be displayed at a first resolution, extracting its content features. These features include attribute information describing the dynamic characteristics and composition of the video content, such as motion intensity, text proportion, target object proportion, texture density, and low-light noise level. The attribute information is obtained through a lightweight context descriptor. Motion intensity is estimated using block motion statistics, optical flow amplitude, or frame difference energy proxy. Text proportion is estimated using a lightweight text detector or edge density combined with stroke heuristics. Target object proportion is estimated using a compact detector at the first resolution. Texture density is characterized by gradient energy or local variance statistics, and low-light noise level is estimated using brightness distribution and noise proxy.
[0111] Based on the extracted content features, this method selects a subset of target metadata from multiple predefined metadata categories. These categories include at least a portion of degradation cues, structural and edge information, temporal motion cues, semantic masks, regions of interest (ROI) maps, and text regions. Degradation cues include noise levels, artifact scores, and blur proxy information; structural and edge information includes edge maps and gradient maps; temporal motion cues include optical flow, motion vectors, and shot transition markers; semantic masks include coarse-grained region labeling; ROI maps include importance weights for faces, salient regions, and main body regions; and text regions include positional masks for video subtitles and UI overlays. The selection process is implemented through a metadata scheduler, which predicts the expected quality gain of each metadata category under the current content features, calculates a unit cost quality gain index, and selects metadata categories based on the unit cost quality gain index until a predefined resource budget constraint is met. This selection mechanism can be supplemented with domain heuristics, such as selecting text region metadata when the text proportion exceeds a threshold, and selecting temporal motion cue metadata when the motion intensity exceeds a threshold, thereby ensuring the selection of the optimal metadata subset under resource constraints.
[0112] Subsequently, the method performs super-resolution processing on the video content based on each metadata category in the target metadata subset. The conditional super-resolution reconstructor extracts features from the heterogeneous metadata in the target metadata subset using a metadata encoder, generating metadata features in a unified format. The heterogeneous metadata includes dense spatial graphs, temporal signals, compact vectors, or scalars. The feature fusion module fuses the metadata features with the video content features to obtain enhanced features. This reconstructor can be integrated into diffusion models, convolutional neural networks, or generative adversarial networks (GANs): when integrated into a diffusion model, it generates second-resolution display content by using the video content, target metadata subset, and selection mask as conditional inputs to the denoiser through an iterative denoising process; when integrated into a convolutional neural network or GAN, the fused enhanced features are input to the backbone network, outputting second-resolution display content; when integrated into a GAN, the target metadata subset is input to both the generator and discriminator. The generator generates display content based on the metadata, and the discriminator determines authenticity based on the metadata, introducing metadata consistency loss constraints to match output details with the metadata.
[0113] Ultimately, this method outputs display content at a second resolution, which is greater than the first resolution. This achieves adaptive super-resolution processing of video content. This scheme dynamically selects the most suitable subset of metadata based on the dynamic characteristics and content composition of the video content, effectively overcoming the limitations of traditional fixed metadata configuration schemes when processing diverse video content. Therefore, under resource budget constraints, this embodiment can optimize the quality and efficiency of super-resolution processing, improve the clarity and detail of the displayed content, especially significantly enhancing fidelity in key areas such as text and faces, while suppressing artifact generation and enhancing system stability.
[0114] In some embodiments, the content features include content feature identifiers and segment-level attributes; the content feature identifiers include at least one of the domain, subject matter, or channel type information of the video content; the segment-level attributes include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level.
[0115] Here, "domain" refers to the macro-level domain of the video content, such as sports events, news reports, movies, and animations; "genre" refers to the specific theme of the video content, such as science fiction, comedy, documentaries, and education; and "channel type" refers to the channel attribute from which the video content originates, such as movie channels, children's channels, and news channels. This identification information can be pre-obtained through manual annotation or by classifying and identifying the video content using machine learning models, providing the metadata scheduler with overall contextual information about the video content.
[0116] Meanwhile, content features also include segment-level attributes, used to describe the local dynamic characteristics and compositional details of video content in time or space. Segment-level attributes can include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level. Motion intensity measures the intensity of motion of objects or background in the video frame, and can be obtained by calculating inter-frame pixel changes, average amplitude of optical flow field, or motion vector statistics. Text proportion measures the proportion of text areas in the video frame, and can be obtained by text detection algorithms in image processing, optical character recognition (OCR) technology combined with area calculation, etc. Target object proportion measures the proportion of a specific target object (e.g., face, vehicle, specific item) in the video frame, and can be identified by an object detection model and the target bounding box area proportion is calculated. Texture density measures the richness of texture details in the video frame, and can be obtained by calculating image gradient information, local variance, or Gabor filter response, etc. Low-light noise level measures the significance of noise in the video frame under low-light conditions, and can be obtained by analyzing the image's brightness histogram, noise model estimation, or local pixel variance, etc. Fragment-level attributes are typically obtained through real-time or offline analysis and calculation of video frames or video fragments, providing the metadata scheduler with micro-level details of the video content.
[0117] The metadata scheduler in this application achieves a more comprehensive and refined understanding of video content. Macro-level "content feature identifiers" provide the overall context of the video content, while micro-level "fragment-level attributes" reveal image details and dynamic characteristics. This multi-layered feature description allows the metadata scheduler to more accurately determine which metadata categories are most critical and effective for the super-resolution processing of the current video content. For example, for videos with high motion intensity, temporal motion cue metadata can be prioritized to reduce motion blur; for videos with a high proportion of text, text region metadata can be prioritized to improve text clarity; and for videos with high levels of noise in low light, noise suppression metadata can be prioritized to improve image purity. Targeted metadata selection avoids blindly processing all metadata categories, thereby optimizing the allocation of computational resources, improving the efficiency and targeting of super-resolution processing, and enhancing the visual quality and user experience of the final displayed content.
[0118] In some embodiments, feature analysis is performed on the video content to be displayed at a first resolution to extract content features of the video content, including at least one of the following steps: Motion intensity in video content can be estimated using block motion statistics, optical flow amplitude, or frame difference energy. Specifically, motion intensity can be estimated in several ways. For example, block motion statistics can be used to divide video frames into multiple image blocks, calculate the motion vector of each image block between consecutive frames, and quantify the intensity of motion in the overall or local scenes by statistically analyzing the amplitude and direction of the motion vectors. Alternatively, optical flow amplitude can be used to characterize motion intensity. By calculating the optical flow vectors of pixels between video frames, the magnitude of the amplitude directly reflects the movement speed and distance of the pixels, thus aggregating the motion intensity information of the scene. Another approach is to estimate motion intensity using frame difference energy, which involves calculating the absolute or squared difference of pixel values between consecutive frames and summing or averaging these differences to reflect the intensity of changes in the scene; the greater the change, the higher the motion intensity.
[0119] The text proportion of video content can be estimated using a lightweight text detector or a combination of edge density and stroke heuristics. As an example, text proportion estimation can employ a lightweight text detector, typically a small, optimized deep learning model capable of quickly identifying text regions in video frames and calculating their pixel proportion. Another approach combines edge density and stroke heuristics. First, image edges are extracted using edge detection algorithms (such as the Canny operator). Then, leveraging the high edge density and specific stroke structures (such as stroke width consistency and connectivity) typically found in text regions, morphological operations and heuristic rules are used to identify and calculate the proportion of text regions.
[0120] The proportion of target objects in video content is estimated at a first resolution using a compact detector. As an example, the estimation of the proportion of target objects can be achieved using a compact detector, which is a computationally efficient and small-scale object detection model (e.g., based on MobileNet or YOLO-Tiny architecture). This detector can quickly and accurately identify and locate predefined target objects directly in the first-resolution video content without resolution upscaling, and calculate the area proportion of the target objects in the frame.
[0121] Texture density in video content can be characterized using gradient energy or local variance statistics. For example, texture density can be represented by gradient energy, which involves applying gradient operators (such as Sobel or Prewitt operators) to video frames to obtain the gradient magnitude of the image. A larger gradient magnitude indicates richer image detail. By accumulating or averaging the gradient magnitudes, the density of the texture can be quantified. Another method is local variance statistics, which calculates the variance of pixel intensity within local regions of the image. A larger variance generally indicates more complex texture and higher density in that region.
[0122] The level of low-light noise in video content can be estimated using brightness distribution and noise proxy. As an example, the estimation of low-light noise can be achieved by analyzing the brightness distribution of the video content. First, low-light areas in the image are identified, and then noise proxy methods are used to assess the noise level of these areas. For instance, relatively flat image patches can be selected within the low-light areas, and the variance or standard deviation of their pixel values can be calculated as a surrogate indicator of noise. Alternatively, a statistical model-based method can be used to estimate noise parameters, thereby quantifying the noise level in low-light environments.
[0123] This application's embodiments enable the metadata scheduler to obtain more comprehensive and accurate video content attribute information, thus providing a solid data foundation for the intelligent selection of subsequent metadata subsets. This not only improves the accuracy of metadata selection, ensuring a high degree of match between the selected metadata and video content characteristics, but also reduces the computational overhead of feature extraction by employing efficient calculation methods. This guarantees the performance of the entire display system in real-time or near-real-time scenarios, ultimately improving the targeting of super-resolution processing and the overall visual quality of the displayed content.
[0124] In some embodiments, the method further includes: generating a corresponding selection mask based on a target metadata subset, so as to determine each metadata category included in the target metadata subset based on the selection mask, wherein the selection mask is used to identify the metadata categories included in the target metadata subset.
[0125] As an example, after performing feature analysis on the displayed video content at a first resolution and selecting a target metadata subset from multiple preset metadata categories based on the content features, this application further proposes generating a corresponding selection mask based on the target metadata subset. This selection mask can be understood as a data structure, such as a binary vector or bitmask, whose function is to explicitly indicate which metadata categories are specifically included in the target metadata subset. For example, if N metadata categories are preset, the selection mask can be a binary vector of length N, where each position corresponds to a metadata category. When a metadata category is selected and included in the target metadata subset, its corresponding position is set to "1", otherwise it is set to "0". This generation method ensures a clear and quantifiable representation of the selected metadata categories.
[0126] After generating the selection mask, subsequent super-resolution processing modules (such as conditional super-resolution reconstructors) will receive and parse it. By parsing the selection mask, the super-resolution processing module can accurately identify which metadata categories need to be actually utilized in the current processing cycle. For example, the super-resolution processing module can iterate through all possible metadata categories and, according to the indication of the selection mask, only activate and process those metadata categories marked as "1" in the mask. This ensures that the input to super-resolution processing is accurately filtered, avoiding redundant processing of unnecessary metadata.
[0127] The function of a selection mask is to provide a clear identifier indicating the specific metadata categories contained in the target metadata subset. Regardless of the form in which the target metadata subset exists (e.g., a collection of metadata objects), the selection mask provides a standardized, easily parsed interface, enabling the super-resolution processing module to quickly and accurately understand the scope of metadata it should focus on. This identification function is crucial for achieving refined metadata-driven super-resolution processing.
[0128] This application, after selecting a subset of target metadata based on video content features, further generates and utilizes a selection mask to explicitly identify and determine the included metadata categories, thereby providing accurate and efficient guidance for subsequent super-resolution processing. This ensures that the conditional super-resolution reconstructor can accurately identify and utilize only those intelligently scheduled and selected metadata categories, avoiding redundant processing of unnecessary metadata and improving the efficiency and resource utilization of super-resolution processing. Simultaneously, because the super-resolution processing can focus more on the key features of the content, the final output second-resolution display content better reflects the dynamic characteristics and composition of the video content, thus optimizing the quality and display effect of the super-resolution.
[0129] In some embodiments, the method further includes: predicting the expected quality gain of each metadata category under the current content features; calculating the unit cost quality gain index of each metadata category, and selecting metadata categories according to the unit cost quality gain index of each metadata category until a preset resource budget constraint is met, and generating a selection mask.
[0130] As an example, expected quality gain refers to a quantitative assessment of the improvement in display quality that can be achieved by applying a specific metadata category for super-resolution processing, given the current video content features. Prediction can be achieved in several ways. For example, a pre-trained machine learning model can be used, which learns from a large amount of video content, corresponding metadata, and quality assessment data after super-resolution processing (such as objective metrics like PSNR, SSIM, or VMAF, or subjective quality scores) to establish a mapping between content features and quality gain. At runtime, the model receives the content features of the current video content as input and outputs the expected quality gain for each metadata category. Alternatively, a series of heuristic rules or lookup table mechanisms can be constructed based on domain expert experience or experimental data to estimate the quality gain. For example, when content features indicate that the video content has high motion intensity, motion-related metadata categories may be predicted to have a higher quality gain.
[0131] Building upon this, this application further calculates the unit cost quality gain metric for each metadata category. The unit cost quality gain metric is a key quantitative indicator for measuring the efficiency of a metadata category, comprehensively considering the quality gain that the metadata category can bring and the resource overhead required for processing it. Typically, this metric can be obtained by dividing the predicted quality gain by the abstract cost function (i.e., resource overhead) of the metadata category under the current content characteristics. Resource overhead may include, but is not limited to, computational complexity (such as floating-point operations (FLOPs), memory usage, processing time, or power consumption. This allows for a fair comparison of the cost-effectiveness of different metadata categories.
[0132] Subsequently, metadata categories are selected based on their unit cost quality-gain metric until a preset resource budget constraint is met, ultimately generating a selection mask. This selection process typically employs a greedy algorithm: in each iteration, the metadata category with the highest current unit cost quality-gain metric is prioritized and added to the selected subset of metadata, while the remaining resource budget is updated. This process continues until no more metadata categories can be added without exceeding the preset resource budget constraint. Finally, based on the selected set of metadata categories, a selection mask in binary vector or bitmask form is generated, explicitly indicating which metadata categories will be used for subsequent super-resolution processing.
[0133] This application enables quantitative evaluation of metadata categories and optimizes their selection based on expected quality gains and resource costs. This ensures that, with limited computing resources, the system prioritizes metadata categories that contribute most to display quality improvement while consuming relatively reasonable amounts of resources. This intelligent selection process avoids blind or inefficient use of metadata, thereby improving the efficiency of super-resolution processing and the visual quality of the final displayed content, while effectively controlling system resource consumption and achieving an optimal balance between performance and cost.
[0134] In some embodiments, super-resolution processing is performed on video content based on each metadata category in the target metadata subset to output display content at a second resolution, including: extracting features from heterogeneous metadata in the target metadata subset to generate metadata features in a unified format, wherein the heterogeneous metadata includes dense spatial graphs, temporal signals, compact vectors, or scalars; fusing the metadata features with the features of the video content to obtain enhanced features, and performing super-resolution processing on the video content based on the enhanced features.
[0135] As an example, a subset of target metadata may contain various data types. For instance, dense spatial graphs can represent regional information in video frames, such as saliency maps or semantic segmentation maps, providing pixel-level spatial context; temporal signals can represent dynamic changes in video content, such as motion vectors or scene transition information, providing cues in the temporal dimension; compact vectors can represent global attributes of video content, such as scene type or sentiment tags, providing high-level semantic information; and scalars can represent numerical attributes, such as mean brightness or noise level, providing quantified statistical information. Heterogeneous data types have different structures and semantics, and using them directly would increase processing complexity. To effectively utilize heterogeneous metadata, feature extraction is required, converting it into a unified, processable format. As an example, for dense spatial graphs, convolutional neural networks (CNNs) can be used for feature encoding to extract their spatial features; for temporal signals, models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or Transformers can be used to extract their temporal dependency features; and for compact vectors or scalars, they can be embedded into a high-dimensional feature space using multilayer perceptrons (MLPs) or linear mappings. After feature extraction, all heterogeneous metadata is converted into feature vectors or feature maps with the same dimensions and semantic representation, such as a fixed-length feature vector or a feature map with a specific number of channels. This unified format facilitates subsequent fusion with video content features, ensuring that different metadata types can participate in super-resolution processing in a consistent manner.
[0136] Before performing super-resolution processing, it is typically necessary to extract the visual features of the original first-resolution video content. This can be achieved using pre-trained deep neural networks (such as ResNet, VGG, Vision Transformer, etc.) as feature extractors, obtaining multi-scale feature representations with rich semantic information from video frames. Fusion is the process of combining metadata features in a unified format with video content features. Fusion methods can be diverse. For example, metadata features and video content features can be concatenated channel by channel, and then dimensionality reduction and information integration can be performed through convolutional layers; element-wise addition or multiplication can also be used to modulate video content features with metadata features; more complex fusion mechanisms can employ attention mechanisms, enabling the model to adaptively focus on the parts of the metadata most relevant to the current video content super-resolution, thereby generating more instructive fused features. The fused features are called enhanced features, which not only contain low-level visual information and high-level semantic information of the video content itself, but also incorporate additional contextual information provided by the metadata, such as content type, motion characteristics, and texture details. Enhanced features provide more comprehensive and accurate input for subsequent super-resolution reconstruction, helping the model to better understand video content and thus generate higher-quality second-resolution display content.
[0137] After obtaining the enhanced features, they are used as input or conditions for the super-resolution model. The super-resolution model can be an end-to-end model based on a convolutional neural network (CNN), a generative adversarial network (GAN), or a diffusion model, etc. For example, in a CNN-based model, the enhanced features can be directly used as input feature maps, and a high-resolution image is gradually reconstructed through a series of upsampling layers and convolutional layers; in a GAN, the enhanced features can be used as conditional inputs to the generator, guiding the generator to generate high-resolution images that conform to the metadata features; in a diffusion model, the enhanced features can be used as conditional inputs to a denoising network, guiding the denoising process to generate high-quality second-resolution display content.
[0138] This application effectively solves the problem of heterogeneous metadata being difficult to directly apply to super-resolution processing. First, by extracting features from heterogeneous metadata in a subset of the target metadata and generating metadata features in a unified format, it overcomes the differences in format and semantics between different metadata types, enabling various metadata information to be effectively utilized by the super-resolution model. Second, it fuses the unified-format metadata features with video content features to generate enhanced features, allowing the super-resolution model to simultaneously consider the visual information of the video content and the rich contextual information provided by the metadata during reconstruction, thus obtaining more comprehensive and instructive input. Finally, performing super-resolution processing based on the enhanced features improves the reconstruction quality of the second-resolution displayed content. For example, it performs better in preserving video content details, suppressing artifacts, and improving visual perception, ensuring a high degree of match between the super-resolution results and the dynamic characteristics and composition of the video content, thereby providing users with a superior viewing experience.
[0139] To implement the content-adaptive metadata display method of this application embodiment, Figure 5 A schematic diagram of an electronic device provided in an embodiment of this application; as shown Figure 5 As shown, the electronic device 500 may include: a memory 501 for storing a computer program; and a processor 502 for implementing the methods described above when executing the computer program. The processor 502 may also implement the steps of any of the methods described above, which will not be repeated here. Exemplarily, the electronic device 500 may be the display device itself provided in this application embodiment or an executable unit integrated within the display device, such as an embedded system; this embodiment does not limit this.
[0140] It should be noted that the electronic device provided in the above embodiments and the display method embodiment based on content adaptive metadata belong to the same concept. For details of its specific implementation process, please refer to the method embodiment, which will not be repeated here.
[0141] Of course, in practical applications, such as Figure 5 As shown, the electronic device 500 may further include at least one network interface 503. Various components in the electronic device are coupled together via a bus system 504. It is understood that the bus system 504 is used to implement communication between these components. In addition to a data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 5Various buses are labeled as bus systems 504. The number of processors 502 can be at least one. Network interface 503 is used for wired or wireless communication between electronic devices and other devices. Memory 501 in this embodiment is used to store various types of data to support the operation of the electronic device. The methods disclosed in the above embodiments can be applied to processor 502, or implemented by processor 502. Processor 502 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 502 or by instructions in software form. The processor 502 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 502 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly reflected in the combined execution of hardware and software modules in a microcontroller. The software module may reside in a storage medium located in memory 501. Processor 502 reads information from memory 501 and, in conjunction with its hardware, completes the steps of the aforementioned method. In an exemplary embodiment, electronic device 500 may be implemented using one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to execute the aforementioned method.
[0142] As an example, this application provides a computer-readable storage medium storing a computer program thereon, such as a memory 501 storing the computer program, which can be executed by a processor 502 to complete the aforementioned method steps. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0143] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0144] In some embodiments, the display device provided in this application can also serve as Figure 6 The receiving end in the video processing system shown is used to perform the functions of the receiving end in the video processing system. The following is a detailed description of the video processing system provided in this embodiment.
[0145] For example, Figure 6 An exemplary architecture of a video processing system is illustrated. The system may include a transmitting end and a receiving end connected by communication. The transmitting end includes a first processor and an encoder. The first processor is used to extract metadata of the video content. As one possible implementation, the first processor may run one or more agents (or modules) for extracting metadata. The encoder is used to encode the video content to form a video stream. The transmitting end transmits the video stream to the receiving end for display. The transmitting end and the receiving end can communicate via a transmission network for data interaction. This application embodiment does not limit the specific implementation of the transmission network.
[0146] Metadata, also known as perceptual metadata.
[0147] For example, metadata must meet the following conditions to improve its broad interoperability: (1) Lightweight: For example, semantic masks are compressed using run-length encoding (RLE) and optimized as bit fields.
[0148] (2) Frame addressable: Allows for granular optimization of applications by frame or scene segmentation.
[0149] (3) Encoder compatibility: Supports insertion into SEI, or transmission via MPEG-V or CTA-861.3 over HDMI.
[0150] (4) Scalable: Allows the addition of additional agent signals (such as gaze, attention map).
[0151] For example, the field structure of metadata is as follows: { "frame": 2345, "semantic": {"face": [[32,45,80,120]], "sky": [[0,0,1920,500]]}, "depthMap": "base64-encoded", "intentFlags": ["preserve_face_tone", "reduce_eyestrain"], "sceneLighting": {"CCT": 5100, "intensity": 0.75}, "motionLevel": 0.62, "viewerPosition": {"x": 0.3, "y": 0.1} } The data in the metadata can also take other forms without restriction. For example, the form of sceneLighting can also be scene Lighting.
[0152] As one possible implementation, the sending end performs the above metadata extraction steps offline.
[0153] For example, the agent at the sending end can be a computationally intensive agent responsible for high-complexity analysis.
[0154] For example, one or more of the aforementioned agents can be integrated into a post-processing pipeline or into a custom batch processing workflow. This integration allows content creators to embed optimized metadata early in the content creation cycle to ensure that the recipient has rich information to enhance the viewing experience.
[0155] As one possible implementation, the sender can package metadata along with a DASH or HLS manifest. This allows for the use of the sender's metadata (such as optimization intent) to form an end-to-end system that maximizes visual quality and efficiency.
[0156] The receiving end includes a second processor and a decoder. The decoder decodes the video stream of the video content. The second processor parses the metadata of the video content and, based on this metadata and the receiving end's local context information, adjusts the display parameters of the receiving end to better present the video content. For example, the local context information includes panel conditions and the context of the viewing user. A detailed explanation of the local context information will be given below.
[0157] As one possible implementation, the second processor can run one or more agents to dynamically adjust display parameters. Exemplarily, these agents can run on a display system-on-chip (SoC) or timing controller (TCON). This tight hardware-level integration ensures low latency and efficient command processing, enabling fine-grained control over display parameters such as local dimming and color calibration.
[0158] For example, the receiving agent can be a lightweight agent with real-time processing capabilities, responsible for real-time execution. The receiving agent can respond to incoming metadata in real time and, by combining metadata and local context information, perform scene-specific optimizations, such as adaptive backlight control, refresh rate adjustment, color gamut remapping, 2D-to-3D enhancement (or 2D-to-3D conversion), and personalized eye comfort adjustments. This receiving end can predictably perform low-latency and low-power intelligent rendering; in some cases, even with limited computing resources, intelligent rendering can be performed without increasing system complexity.
[0159] The receiving end can also be called the playback end or the display end.
[0160] One possible implementation is to integrate the intelligent agent as a middleware service into the receiving end's operating system. Alternatively, it can be integrated into the Hardware Abstraction Layer (HAL) to ensure broad applicability at the receiving end.
[0161] As one possible implementation, the aforementioned agents can be distributed across one or more layers to ensure seamless optimization and enhancement of visual performance. For example, at the media player or decoder layer, the software could be modified to parse key metadata (such as SEI messages, OBU, or manifest information) and expose it to the receiving agent. In this way, the receiving end can dynamically adjust display parameters based on this metadata and local context information.
[0162] As one possible implementation, multiple agents can share an access data buffer. Agents can read data from the data buffer, which contains metadata and local context information. Agents can also write data to the data buffer.
[0163] As one possible implementation, the agent can also respond to prompts from upstream agents and trigger optimizations from downstream agents. For example, in response to metadata provided by an upstream agent: if the scene contains faces, the agent can preserve the skin color of the face region to ensure image fidelity.
[0164] The system described above separates high-complexity analysis from real-time execution through modular and distributed processing pipeline operations. Table 1 shows the task characteristics executed by the sending and receiving ends, respectively.
[0165] Table 1
[0166] In this way, a single complex visual intelligence processing operation performed in the sending pipeline can be widely applied to multiple receiving ends (playback devices) to generate metadata. The receiving ends do not need to re-encode or retrain, reducing their overhead. Furthermore, the receiving end can combine this metadata with its own local context information to adaptively adjust the display parameters of the video content, ensuring the video content matches the receiving end's local context and providing a personalized video playback experience. Moreover, the task division between the sending and receiving ends balances computational efficiency and real-time adaptability, preventing dynamic rendering from overloading the display hardware.
[0167] The aforementioned hybrid pipeline system framework or architecture, known as DP Vision, is used for content-aware display optimization. DP Vision achieves intelligent rendering through collaboration among multiple agents. For example, DP Vision can be deployed in various devices, such as, but not limited to, smart TVs, mobile phones, tablets, AR / VR headsets, in-vehicle displays, and streaming media platforms.
[0168] In the embodiments of this application, the intelligent agent may also be referred to as a module, intelligent module, model, agent, or other names, without limitation, and will be uniformly described here.
[0169] Please see Figure 7 , Figure 7 A flowchart example of a video processing method provided in this application embodiment. The method may specifically include: S101. The sending end analyzes the video content to obtain the video content's metadata.
[0170] Metadata is used to describe at least one dimension of the visual characteristics of the video content. The video content may include the content of one or more video frames.
[0171] For example, the sending end is a device provided by the content provider, such as a content production device, edge server, or encoder. As one possible implementation, the sending end can be responsible for performing computationally intensive (or data-intensive) video content analysis tasks to obtain the visual features of the video content in various dimensions.
[0172] As one possible implementation, the sending end can perform the aforementioned video content analysis task through one or more modules (or agents). For example, refer to... Figure 8The one or more modules may include a semantic analysis module (or semantic segmentation module or scene interpretation module), a depth and geometry module (or depth estimation module), an edge detection module, a causal reasoning module (or causal inference module), a first noise suppression module, and an illumination analysis module.
[0173] For example, except Figure 8 The modules listed in the examples may also include a motion profiling module for motion estimation. Other modules may also be included for analyzing the video content to obtain corresponding metadata.
[0174] The specific actions performed by these modules will be described below.
[0175] For example, the sending end can analyze the video content based on its global visibility and complete temporal context to obtain metadata. The temporal context represents the temporal relationship between multiple video frames contained within the video content. For instance, the sending end analyzes 60 seconds of video frames and predicts a smooth transition in color tone from warm to cool. It then generates "lighting model" metadata: the sending end can create a temporally continuous, smoothly varying lighting parameter curve (e.g., a color temperature smoothly changing from 4500K to 6500K) and package this curve into the metadata. An optimization intent is then generated. After receiving this metadata, the receiving end renders the video based on the lighting parameter curve to present a video with a smooth color tone transition.
[0176] S102. The sending end sends metadata to the receiving end.
[0177] Correspondingly, the receiving end receives metadata.
[0178] As one possible implementation, the sender can encapsulate the metadata in supplemental enhancement information (SEI) and send the SEI to the receiver. That is, the metadata is carried within the SEI. Alternatively, the metadata can be encapsulated in an open bitstream unit (OBU) and sent to the receiver. Or, the metadata can be carried in a bypass manifest, such as sidecarJSON or XML manifests. Alternatively, the sender can transmit the metadata to the receiver via a metadata API.
[0179] As one possible implementation, the transmitter can transmit metadata via a custom HDMI or DisplayPort channel.
[0180] In this way, the sending end can efficiently and with low latency send metadata to the receiving end through the appropriate transmission path, ensuring compatibility with various device types. For example, it can support the transmission of metadata to receiving ends such as TVs, tablets, in-vehicle displays, and AR / VR displays.
[0181] Furthermore, metadata can be directly embedded into the encoded bitstream or provided as a bypass manifest to ensure broad compatibility and low transmission overhead.
[0182] S103. The receiving end adjusts the display parameters used to display video content based on metadata and the receiving end's local context information.
[0183] The local context information of the receiving end is used to indicate at least one of the following: the environmental characteristics of the receiving end, the user characteristics of the viewing user of the receiving end, and the device operation characteristics of the receiving end.
[0184] As one possible implementation, the receiver interprets the metadata in real time and combines it with local context information to perform scene-aware optimization to adjust the display parameters of the video content.
[0185] For example, if a moving object is detected in the video content, the refresh rate is adjusted. Another example is adjusting the color temperature of the video content based on the ambient lighting conditions at the receiving end. This ensures consistent rendering of video frames and intelligently optimizes for the user characteristics, device characteristics, and environmental characteristics of the current viewer at the receiving end.
[0186] As one possible implementation, the receiving end can perform the task of adjusting the display parameters described above through one or more modules. For example, refer to... Figure 8 These modules may include a super-resolution module, a stereo conversion module, a color and color gamut module, a language module, a second noise suppression module, a lighting adjustment module, an orchestration module, and a verification module. The specific actions performed by these modules will be described below.
[0187] In this way, each module is responsible for a specific rendering task, such as semantic segmentation, tone mapping, refresh rate control, or color gamut adjustment. This not only ensures high-efficiency parallelism but also supports the flexible expansion of new perception modules, achieving a plug-and-play integration effect. This modular and scalable approach ensures consistent performance across receivers, content types, and user environments.
[0188] This article primarily uses the example of a dedicated module performing a specific action. In other embodiments, the functions of some of the modules involved in the embodiments of this application can also be integrated into a single module. This application does not impose any limitations on this.
[0189] The method provided in this application determines display parameters based on the analysis and understanding of video content, the perception of the viewer and the environment, and the state of the receiving end itself. Therefore, the receiving end's display parameter adjustment decisions do not solely rely on the universal metadata provided by the sending end, but are deeply integrated with the local context. This allows the same video content to be rendered in the most suitable way for the current situation in different environments (such as living rooms and bedrooms), for different users (such as visually sensitive individuals and ordinary users), and on different devices (such as high-end OLED TVs and ordinary LCD monitors), achieving a transformation from "one-size-fits-all" to "personalized experiences for each user and each time."
[0190] Furthermore, it decouples content analysis (computation-intensive, which can be performed offline or in the cloud) from real-time rendering (latency-sensitive, performed locally). The sending end can leverage powerful computing resources to perform complex, multi-dimensional video analysis and generate high-quality metadata without considering the heterogeneity of the receiving end hardware. This solves the computational load, power consumption, and heat generation problems associated with performing deep analysis and high-performance rendering simultaneously on a single device.
[0191] In some embodiments, the metadata includes one or more of the following data: edge features of the video content, local contrast of the video content, RGB channel statistics of low grayscale areas in the video content, illumination data of the video content, depth information of displayed objects in the video content, semantic segmentation data of the video content, source color configuration data of the video content, causal relationships corresponding to visual elements in the video content, motion estimation data of the video content, key visual regions (or important visual regions) in the video content, creative intent of the video content, saliency score of the video content, sentiment index of the video content, and optimization intent for the video content; the saliency score is used to describe the content weight index of different regions in the video content; the sentiment index is used to describe the emotional tone of the video content.
[0192] The metadata includes one or more of the following data, which may be explicitly included or implicitly included. Implicitly including X means that the metadata implicitly indicates X.
[0193] The aforementioned low grayscale region can refer to an area with brightness below a threshold. This low grayscale region can also be called a dark area. In this low grayscale region, the signal strength is low, making it susceptible to noise interference and affecting image quality. The method provided in this application, considering that this low grayscale region is more susceptible to noise, specifically uses the RGB channel statistics of the low grayscale region to suppress noise, thereby enhancing the image quality of the low grayscale region. For example, it can dynamically suppress artifacts while maintaining the integrity of shadow details and color fidelity. Furthermore, it ensures a clean and stable dark background without introducing compressed textures or losing subtle gradients.
[0194] The semantic segmentation described above can be used to obtain semantic masks and object labels for each object in a video frame, such as semantic masks and object labels for faces, text, and the sky. In this way, the receiving end can understand the scene classification in the image and identify the region of interest. For example, it can identify where the face is and where the sky is, thus enabling it to perform fidelity processing on the face and color enhancement on the sky.
[0195] The content weight metric for a specific region within video content indicates the degree to which that region attracts the viewer's attention. Regions of interest have higher weight metrices. One possible implementation is to use an attention heatmap to represent the content weight metric. This attention heatmap can be a grayscale or pseudo-color image of the same size as the video frame, where the brightness or value of each pixel represents the probability that the region attracts the viewer's attention. The higher the probability that a region attracts the viewer's attention, the higher its weight metric.
[0196] Visual elements in video content can also be called perceptual variables or causal variables. The causal relationship described above can also be called a causal dependency relationship. This can be represented by a causal graph, let G = (V, E), where V is the set of visual elements and E represents the causal relationship between the causal variable and the outcome variable. Causal variables include brightness, motion, and viewer comfort. Outcome variables include, for example, the degree of eye fatigue. A causal relationship could be: brightness → eye fatigue. This causal relationship indicates that brightness can affect eye fatigue, and there is a causal relationship between brightness and eye fatigue.
[0197] For example, the joint distribution of the causal variable and the outcome variable satisfies the following formula (1): (1) in, Representative variable The set of parent nodes, that is, those directly pointed to in the cause-effect graph. All nodes are The direct cause. In other words, It is a causal variable. It is the outcome variable.
[0198] Indicates that, given all direct causes In the case of variables The conditional probability of taking a certain state.
[0199] Given observed metadata O V, and action space A, the module (or agent) selects the optimal action a in the following way. : (2) U(X) is a utility function that encodes perceived targets such as contrast and sharpness, color fidelity, eye comfort, and energy efficiency.
[0200] This represents a causal interference quantifier, which indicates taking an action, such as increasing the local backlight, and observing the corresponding results to evaluate the causal effect of the action itself.
[0201] E[ ∣do(a),O] represents the expected value of the system utility U(X) given the known observational evidence O and the implementation of action a.
[0202] Argmax represents maximizing the parameter solution. Specifically, it iterates or searches the action space A to find the action a that maximizes the expected utility E[U(X)|do(a),O], and uses it as the final decision a. .
[0203] The above method can determine the adjustment strategy of display parameters based on the expected results (causal effect) that the action will lead to, and realize the transformation from "experience-driven, black-box decision-making" to "causal-driven, explainable decision-making", thereby improving the display performance of the receiving end.
[0204] For example, the depth information mentioned above includes depth maps, 3D spatial cues, and stereo disparity data. Stereo disparity data is, for example, a disparity vector. 3D spatial cues represent information or features that can be inferred from 2D images or videos to characterize the three-dimensional spatial relationships, shapes, layouts, and depths of objects in a scene.
[0205] This depth information can provide a basis for generating stereoscopic images and adjusting parallax, so as to provide users with highly immersive video content, enhance spatial perception, and provide a richer visual experience.
[0206] For example, motion estimation data includes motion vectors and motion patterns.
[0207] In some embodiments, the local context information includes at least one of environmental information, viewing user preference information, and receiving device information; the device information includes: the status of the receiving device's display panel and hardware capabilities; the preference information includes one or more of the following: the viewing user's focus, viewing angle, viewing distance, relevant interaction data of the viewing user watching the video, or viewing mood.
[0208] The status of the display panel, such as, but not limited to, the brightness, temperature, and power consumption of the display panel.
[0209] Hardware capabilities can be understood or replaced as: system capability constraints (system constraints), or hardware capability constraints. For example, hardware capabilities include panel capabilities or power supply capabilities, such as power supply limits, receiver thermal limitations, and the display range of the panel. The method provided in this application embodiment limits the state of the display panel within the range of hardware capabilities to extend the lifespan of the display panel. For example, the temperature of the display panel is controlled within the thermal limit range, and the power consumption of the display panel is controlled within the power supply limit range.
[0210] Environmental information includes, but is not limited to, lighting conditions that affect visual viewing, ambient noise levels, electromagnetic interference levels, geographical location and time, and environmental content.
[0211] For example, while ambient noise levels do not directly affect visual perception, the receiver can adjust its subtitle rendering strategy when high noise levels are detected, automatically increasing subtitle size and contrast, or enabling speech-to-text assistance. This can meet users' video viewing needs under high-noise conditions.
[0212] For example, the level of electromagnetic interference may affect the stability of the screen signal, and the receiver can adjust the refresh synchronization strategy or enable signal anti-interference processing accordingly.
[0213] For example, the receiving end can switch the corresponding rendering mode based on geographical location and time, such as a nighttime viewing mode or a car interior daylight mode.
[0214] For example, the receiving end can use a front-facing camera or sensors to identify the content of the user's environment. When a user is conducting a video conference in a meeting room, priority can be given to ensuring shared visibility and reducing glare. When a user is in a bedroom, low blue light and flicker-free dimming can be automatically enabled.
[0215] User preference information indicates a user's liking or state when watching video content. This preference information can be user-defined or inferred by the receiver based on the user's behavior. Examples of user preference information include a preference for comfort mode, a preference for 3D mode, high sensitivity to eye fatigue, and high sensitivity to flicker.
[0216] For example, the rendering effect on the receiving end can be different when the user is at different viewing angles, in order to improve the viewing experience of the video at the corresponding viewing angles.
[0217] For example, the rendering effect on the receiving end can be different when the user is at different viewing distances, in order to improve the viewing experience of the video at the corresponding viewing distance.
[0218] For example, when watching HDR videos, users tend to lower the brightness, and the receiver can determine the user's sensitivity to peak brightness. Subsequently, when the user watches an HDR video again, the receiver can lower the brightness. Similarly, users may prefer vibrant colors in animated content. Later, when the user watches similar videos, the receiver can adjust the video's color gamut to meet the user's visual needs. In this method, the receiver can learn the user's visual preferences based on user interaction data. Over time, the learned visual preferences become increasingly accurate, helping to achieve personalized video playback effects.
[0219] For example, the receiving end can analyze the user's viewing emotions in real time based on facial expressions, tone of voice, or behavioral patterns, and adjust the display parameters of the video content to adapt to the viewer's emotional state. For instance, in a suspenseful scene, it can be rendered with higher contrast and a cooler color tone. In a relaxed mood, it can be adjusted to a soft, warm color tone. Another example is adjusting emotional lighting based on the user's facial expressions, tone of voice, or behavioral patterns. Emotional lighting can be a visual attribute of the screen itself, such as dynamically fine-tuning the overall color tendency (color grading) of the image. Alternatively, emotional lighting can be the emotional lighting of the external environment, such as using smart light strips or the room's main light near the receiving end to change the color and brightness of the surrounding environment based on the identified viewer emotions and the content of the image.
[0220] For example, if a user is colorblind or has low vision, the receiver can adaptively enhance colors based on the user's characteristics to provide a visual experience that meets the user's viewing needs.
[0221] The method provided in this application allows the receiving end to combine rich metadata of the video content with local context information to perform video content rendering. This enables the display effect to dynamically adapt to the content of each frame and the local context of the receiving end, achieving high-quality, personalized rendering effects.
[0222] It should be noted that the local context information listed above is merely an example. Any information characterizing user features, device operating characteristics, and environmental characteristics can be used as the local context information of the receiving end. This allows for personalized video effects to match different user characteristics, as well as video effects to match different environmental characteristics and different device operating characteristics.
[0223] Similarly, metadata can also include other data that characterizes the video content, such as RGB to grayscale distribution data.
[0224] In some embodiments, display parameters include one or more of the following data: resolution, contrast ratio, brightness, color temperature, color gamut, saturation, or refresh rate. These display parameters are merely examples; any parameter that can affect the final presentation of the video can be used as a display parameter to be adjusted.
[0225] In some embodiments, the receiver is specifically configured to: determine a first target area in the display area of the receiver's display panel based on metadata and the receiver's local context information, and adjust the display parameters of the receiver for the first target area.
[0226] In one example, the primary target area is the user's focus of attention. In wearable devices and AR / VR headsets, the receiver may include a gaze tracking module. This module, based on gaze-aware technology and combined with the user's real-time eye gaze data (an example of interactive data), determines the user's focus on the video content and adjusts the display parameters (such as brightness, contrast, and resolution) of that focus area accordingly to ensure its clarity.
[0227] This prioritizes rendering quality for the user's visual focus areas, preserving details in those areas and improving visual resolution. In some scenarios, it reduces computation in peripheral areas, lowering power consumption.
[0228] In one example, the first target region could also be a key visual region. Of course, the first target region could also be other local regions.
[0229] The first target region listed above is merely an example. The receiving end may determine other regions as the first target region based on metadata. Alternatively, the receiving end may determine other regions as the first target region based on local context information. Or, the receiving end may determine other regions as the first target region based on both metadata and local context information.
[0230] As one possible implementation, adjusting the display parameters of the receiving end for the first target area can also be understood or replaced as follows: For the first target area, a first rendering strategy is adopted. For areas outside the first target area, a second rendering strategy is adopted. This second rendering strategy differs from the first rendering strategy. In this way, the first target area can have a different rendering effect than other areas to meet the user's specific visual needs for the first target area. For example, refer to... Figure 9 The saturation of the first target region is A, and the saturation of other regions is B.
[0231] As one possible implementation, the aforementioned eye fixation data can also be used for fatigue estimation models and visual comfort scoring. This would allow the display system to dynamically adjust the brightness, sharpness, and motion intensity of video content to reduce visual fatigue.
[0232] Compared to related technologies where instructions are applied to the entire video frame, the method provided in this application can identify important display objects or regions within the video content, such as distinguishing faces, text, or other visually important areas. This allows for priority given to important display objects or regions. On one hand, this improves the display performance of these important objects or regions. On the other hand, when receiving end resources are limited, rendering resources can be preferentially allocated to important regions to achieve relatively better rendering results under limited rendering resource conditions.
[0233] In some embodiments, the transmitting end includes a depth and geometry module for generating depth information of objects displayed in the video content; The receiving end includes a stereo conversion module, which renders the video content into a stereo image based on depth information. The perspective of the stereo image is determined based on the position or posture information of the viewing user.
[0234] The stereoscopic conversion module, also known as the 3D projection module, refers to the perspective of a stereoscopic image, which can be understood as the perspective of the multi-view image corresponding to the stereoscopic image.
[0235] As one possible implementation, the depth and geometry module can obtain a depth map based on semantic segmentation and motion estimation of video content, combined with the contours of displayed objects, and use this depth map as depth information.
[0236] Based on this depth information, the receiving end can convert the original 2D video into a 3D video, and the viewing angle in the 3D video can be adjusted according to the user's position or posture.
[0237] For example, refer to Figure 10A In (a), if the system detects that the user is sitting on the right side of the TV, the stereoscopic conversion module in the TV can adjust the viewing angle of the 3D video to provide the user in that position with a better 3D viewing experience. For example, see reference... Figure 10A (b) If the TV detects that the user is sitting in the middle position, it can adjust the viewing angle of the 3D video to provide the user in that position with a better 3D viewing experience.
[0238] As one possible implementation, in VR / AR, glasses-free 3D, and other similar scenarios, the perspective of the stereoscopic image changes as the user's head turns, as if observing a real three-dimensional world through a window. This can enhance the sense of presence and interactivity.
[0239] For example, for systems equipped with cameras or sensors, the stereo conversion module can dynamically adjust the rendering based on the depth information of the video content to match the user's head position or eye tracking, ensuring that each user gets the best parallax.
[0240] For example, in prism or light field displays, stereo conversion modules can calibrate multiple viewpoints to optimize the naked-eye 3D reality without artifacts or ghosting.
[0241] The method provided in this application embodiment allows the receiving end to generate a stereoscopic image based on the depth information of the video content. The stereoscopic effect of the image originates from the three-dimensional structure contained in the video content itself. Therefore, a more realistic, natural, and seamless 3D effect can be obtained, providing a truly immersive 3D experience.
[0242] Furthermore, since the viewing angle of a stereoscopic image can be determined based on the user's position or posture, it can provide users in different positions or postures with corresponding viewing angles, thereby further enhancing the sense of immersion.
[0243] In some embodiments, the receiver can perform color shift compensation for wide-viewing-angle scenes. For example, LCD panels are prone to color distortion at wide viewing angles, causing hue and saturation shifts in key visual areas such as skin tones and skies. As a possible implementation, the receiver may also include a color compensation module for performing color compensation based on key visual areas in metadata to maintain color fidelity at wide viewing angles. For example, refer to... Figure 10B If the system detects that a user is watching video from a side angle on the TV, it preserves the natural tones of faces and compensates for the red and yellow components that are easily lost at side views. Furthermore, it preserves the natural tones of the sky at side views. This enhances the viewing experience with wide viewing angles. Similarly, in environments like living rooms or public displays, where different viewers are at different angles, the receiver ensures that skin tones and colors of natural landscapes appear true to life for each viewer.
[0244] In some embodiments, the receiver further includes a color and gamut module (also known as a color and gamut optimization intelligent module) for determining a tone mapping strategy based on metadata and local context information of the receiver, wherein the tone mapping strategy is used to compress the brightness range of high dynamic range (HDR) video content to the renderable range of the receiver.
[0245] As one possible implementation, the receiver can generate an adaptive lookup table (LUT) based on user preference information. For example, the LUT might record: the user's habit of increasing screen brightness in high-saturation game scenes; and the user's habit of reducing blue light and increasing screen contrast in late-night movie viewing scenarios. Subsequently, upon detecting a late-night movie viewing scenario, the color and color gamut module can adjust the display parameters of the video content to enhance screen contrast and reduce blue light.
[0246] As one possible implementation, the color and gamut module can determine the tone mapping strategy based on the color offset matrix.
[0247] The color and gamut module can also be used to adjust color depth (such as brightness), gamut, and white point (such as brightness) based on metadata and local context information from the receiver to achieve perceptual balance and visual comfort.
[0248] For example, the color and gamut module adjusts color depth, gamut, and white point based on a LUT or color offset matrix.
[0249] The white point represents the chromaticity coordinates of white as perceived by the human eye under specific lighting conditions. In some cases, the white point serves as the baseline reference point for the entire color space. In other cases, it may affect the user's perception of white in video content.
[0250] For example, the color temperature of ambient light greatly affects our perception of "white" on a screen. Under warm yellow light, a white dot may appear bluish. Under candlelight, a white dot may appear warm yellow. Therefore, embodiments of this application can adjust the display parameters of the white dot to improve the display effect. For instance, under different ambient light conditions, in order to maintain the accuracy of the "white dot" color of the image, the color temperature of the corresponding area of the display panel can be dynamically compensated.
[0251] As another example, the color and gamut module can adjust the color balance, brightness curve, and white point according to the scene-specific target of the video content, panel capabilities, and ambient lighting conditions.
[0252] The method provided in this application can match different tone mapping strategies for different viewing scenarios such as gaming, streaming media, and reading. This allows for a personalized and comfortable viewing experience based on the viewing scenario, ensuring color accuracy and achieving superior image quality and eye comfort.
[0253] In some embodiments, the receiving end further includes a language module, configured to adjust the rendering effect of the subtitles based on metadata and the local context information of the receiving end. The rendering effect of the subtitles includes, but is not limited to, font, position, or contrast.
[0254] For example, this language module can adjust the rendering effect of subtitles based on the scene semantics of the video content. For instance, in movie A, a character in a dark indoor setting contrasts sharply with the glaring desert sun outside the window. If the subtitles use a fixed white, they would blend into the background and disappear in bright areas, while being too glaring in dark areas, disrupting the atmosphere. Therefore, after receiving metadata, the language module identifies the current scene as "extremely high contrast" and obtains the brightness distribution map of the image. When the subtitles appear in a bright area outside the window, the language module can automatically adjust the subtitle color to dark gray or black and add an outline to ensure clear readability. When the subtitles move to a dark indoor area, the color is automatically adjusted back to a soft light gray or off-white, and the brightness is reduced to prevent glare. If the bright area in the center of the image is too large, the language module can also temporarily adjust the subtitles to a darker area at the bottom of the image. This provides a better subtitle rendering effect.
[0255] As another example, the language module can adjust the font, position, and contrast of the subtitles based on user-defined preference information.
[0256] For another example, refer to Figure 11 (a) If the system detects that the user is positioned on the right side of the TV, the language module can render the subtitles on the right side for the user's convenience. (See reference) Figure 11 (b) If the user is detected to be in the center of the TV, the language module can render the subtitles in the center for the user's convenience.
[0257] In some embodiments, the transmitting end further includes a first noise suppression module, configured to generate RGB channel statistics for low grayscale regions in the video content. The receiving end further includes a second noise suppression module, configured to perform noise suppression on the low grayscale regions based on the RGB channel statistics for the low grayscale regions.
[0258] For example, RGB channel statistics include: the brightness distribution, noise characteristics, and color balance relationships of the RGB channels in low grayscale regions.
[0259] The brightness distribution can refer to the distribution of brightness values for each channel in the low grayscale region. For example, a histogram can be used to represent this brightness distribution.
[0260] Noise characteristics can characterize the type and intensity of noise generated by each channel in the low grayscale region. Noise types include random noise and fixed-pattern noise.
[0261] Color balance can be characterized by the relative intensity relationship of the R, G, and B channels in low grayscale areas.
[0262] In low grayscale areas, image sensor noise and panel display noise (such as color shift in VA panels) are particularly noticeable. For example, visual artifacts such as color banding and hue shift are more apparent in dark scenes, especially on VA panels. Therefore, the method provided in this application allows the transmitting end to accurately quantify the distribution and type of noise by analyzing the RGB channel statistics of the low grayscale area. The receiving end then performs noise suppression accordingly, avoiding excessive blurring of image details (especially textures) by traditional noise reduction algorithms, achieving "noise removal without leaving a trace."
[0263] This can improve the viewing experience of dark scenes and night scenes that are prevalent in movies, games and other videos, making dark details purer and clearer, and significantly improving the overall texture and dynamic range perception of the picture.
[0264] As one possible implementation, the second noise suppression module can also perform noise suppression by combining local context information.
[0265] For example, in a bright viewing environment, the noise reduction intensity can be reduced to prioritize the preservation of details and textures, thereby improving image clarity. In a dimly lit viewing environment, where the human eye is extremely sensitive to noise, the second noise suppression module increases the noise reduction intensity in the aforementioned low grayscale areas to achieve a cleaner and more comfortable viewing experience.
[0266] For example, OLED panels themselves have no backlight and produce pure blacks, but they may exhibit "black crush" (loss of detail in dark areas) or color unevenness at low grayscale levels. Therefore, receivers employ cautious video noise reduction strategies, such as prioritizing color noise suppression over luminance noise suppression, to minimize the loss of detail in dark areas. For LCD / mini-LED panels, backlight leakage may occur, resulting in grayish dark areas. Therefore, receivers adopt more aggressive noise reduction strategies, which can be combined with local dimming to reduce backlighting while suppressing noise in low grayscale areas.
[0267] For example, for the low-grayscale area that the user is focusing on, the receiver performs mild noise reduction based on detail preservation to retain as much image detail as possible. For the surrounding visual areas, stronger noise reduction is performed because the human eye is less sensitive to peripheral details.
[0268] In some embodiments, the transmitting end further includes a lighting analysis module for analyzing lighting data of the video content; the lighting data includes one or more of the following: shadow direction, lighting direction, lighting area, or flare intensity. The lighting area can be an area illuminated by light, or an area where the light intensity is greater than a threshold. An area where the light intensity is greater than a threshold can be referred to as a highlight area.
[0269] The receiver also includes a lighting adjustment module, which is used to determine a second target area in the display area of the receiver's display panel based on lighting data, and adjust the display parameters of the receiver for the second target area.
[0270] As one possible implementation, the transmitting end indicates a lighting model to the receiving end, which may include one or more of the following data: shadow direction, lighting direction, lighting area, or flare intensity. Alternatively, the transmitting end directly indicates one or more of the following data to the receiving end: shadow direction, lighting direction, lighting area, or flare intensity.
[0271] For example, refer to Figure 12 In scene (b), the protagonist is in a bedroom with a lamp next to him. The transmitter can identify key objects in the scene: the lamp and the face of the character illuminated by the lamp. By analyzing the color tendencies of the highlights, shadows, and midtones of these objects, the transmitter can infer that the dominant light source is the lamp. This light is characterized by a low color temperature, exhibiting a strong warm yellow hue. Based on this, the transmitter can generate lighting data, for example: SceneLighting: {"CCT": 2000, "intensity": 0.6}. Where CCT: 2000 indicates that the detected scene color temperature is 2000K.
[0272] After receiving the metadata containing CCT: 2000, the receiver doesn't simply make the entire screen yellowish. Instead, it adjusts the color temperature (color temperature 1) of the area where the light is located to match the visual characteristics of a 2000K color temperature. Viewers will feel as if the screen is actually emitting light, perfectly conveying the warmth of the light and creating a highly immersive experience.
[0273] As another example, the receiving end can set the backlight level of the core area of the image (such as the area where the person is located) to 92% based on the metadata of the video content. It can also slightly boost the pixel data in that area to compensate for the brightness. Ultimately, the perceived brightness is close to 100% backlight, but the actual backlight power consumption is reduced by 8%, resulting in less heat generation.
[0274] For transitional areas in the image, the receiver smoothly transitions the backlight level from 92% to 30%, creating an optical "feathered" edge. This actively pre-compensates for light diffusion from the display panel, suppressing the formation of halos.
[0275] Compared to Figure 12 (a) By uniformly adjusting the display parameters of a frame, the method provided in this application embodiment allows the transmitting end to identify the light source (such as a lamp) and its illumination range (illuminated area) in the image. The receiving end can then precisely adjust the display parameters of the illuminated area accordingly, thereby simulating a realistic lighting effect, making the displayed color tone consistent with the original scene environment, and enhancing the sense of realism and immersion.
[0276] As one possible implementation, the receiver can brighten the illuminated areas and darken the shadow areas based on the lighting data. This results in a higher dynamic range, more transparent highlights, purer shadows, and reduced halo effects due to precise brightness control, thus creating a realistic HDR effect.
[0277] As one possible implementation, the receiver can adjust display parameters based on the direction of light within the image to enhance the sense of depth. For example, if light is detected entering from the left side of the image, the backlight in the corresponding area on the left side of the screen can be slightly enhanced. This makes the displayed objects appear more three-dimensional, creating a three-dimensional light and shadow effect from a two-dimensional display.
[0278] In some embodiments, the receiver may further include a dynamic refresh and backlight intelligence module, which can be used to adjust the refresh rate based on motion estimation data or source color configuration data.
[0279] As one possible implementation, the dynamic refresh and backlight smart module can output timing parameters to trigger an adjustment of the refresh rate.
[0280] For example, based on content awareness, and combining scene motion and brightness intensity represented by metadata from the sending end, the display refresh rate is adjusted frame by frame. In fast-moving scenes, a high refresh rate is used to maintain sharpness. Static or dimly lit scenes can be rendered at a lower rate to reduce flicker and save power.
[0281] Compared to related technologies where the video refresh rate is fixed or only responds to device settings, the method provided in this application can adaptively adjust the refresh rate based on the characteristics of the video content itself. This allows for better video display by matching the characteristics of the video content.
[0282] The dynamic refresh and backlight intelligence module can also be used to make one or more of the following adjustments based on metadata (such as screen brightness and scene semantics): adjust the local dimming area or adjust the backlight.
[0283] In some examples, the dynamic refresh and intelligent backlight module adjusts the local dimming area or backlight direction of the panel in real time based on lighting data such as shadow direction, highlight areas, or flare intensity of the video content. For instance, by analyzing the lighting cues in the video content itself, it can be determined that there is a window on the left side of the movie scene, with sunlight shining in from the left and illuminating the character's face. The receiver can then simulate these lighting directions and characteristics, such as adjusting the backlight unit to minimize light scattered to the left. Thus, when a user watches the video, they will perceive the light on the screen shining from the window on the left side of the screen, illuminating the character and casting a shadow on their right. It is evident that this method allows for dynamic adjustment of the lighting direction, creating a visual experience with a sense of three-dimensional depth and environmental consistency on a two-dimensional screen. Furthermore, because it adjusts the backlight strategy for a local area, power consumption can be reduced.
[0284] In some examples, the dynamic refresh and backlight intelligence modules can also simulate the lighting consistency between the screen environment and the viewer's perspective. For instance, if sunlight shines on the receiver's screen from the left, the system can adjust the screen lighting to match the direction of the light, enhancing the three-dimensional continuity of the scene and improving the realism, immersion, and 3D depth perception of the video content.
[0285] In some embodiments, the transmitting end further includes an edge detection module for generating edge features and local contrast in the video content. For example, this edge detection module generates a high-precision edge map and statistical information on local contrast.
[0286] In some embodiments, the receiver further includes a super-resolution module for performing upsampling based on edge features and local contrast, the upsampling being used to adjust the resolution of the video content. This process can also be referred to as content-aware super-resolution.
[0287] As one possible implementation, the super-resolution module uses edge-aware and scene-aware models to upsample video content and output enhanced frames to achieve edge enhancement, strengthen spatial details, and provide efficient rendering of high-resolution videos (such as 8K).
[0288] For example, this method can be applied to action scenes, such as sports or action movies, to preserve the visual details of the action scene.
[0289] Compared to traditional super-resolution algorithms that apply similar processing to all regions, which can easily lead to oversharpening of textured areas or ringing artifacts in smooth areas, the method provided in this application allows the receiving end to upscale video content to a higher native resolution based on content-aware super-resolution. Furthermore, the receiving end utilizes edge features to accurately locate areas requiring contour enhancement and detail preservation, and uses local contrast to determine the intensity of the enhancement. For example, it can strongly enhance text and building edges while softening skin and sky areas, thereby achieving a clearer and more natural display effect.
[0290] Furthermore, it allows the use of lower-resolution source video during transmission or storage, reducing bandwidth dependence and saving costs for the transmission and storage of video content such as 4K / 8K.
[0291] In some embodiments, the receiver further includes an orchestration module for instructing the receiver to adjust the receiver's display parameters based on metadata and the receiver's local context information.
[0292] In other words, the orchestration module is responsible for the overall orchestration. For example, the orchestration module is located in the orchestration layer.
[0293] In some embodiments, the orchestration module is specifically used for: Analyze the optimization intent for the video content; Based on the optimization intent and the local context information of the receiving end, determine the modules to be used and the adjustment strategy for the display parameters; The module to be used is instructed to adjust the display parameters of the receiver according to the adjustment strategy.
[0294] As one possible implementation, the orchestration module at the receiving end can interpret the optimization intent in the metadata and make decisions based on local context information. The decisions include one or more of the following: the module (agent) to be activated or used, the adjustment strategy for display parameters, or the conflict resolution strategy.
[0295] This orchestration module acts as an execution manager. It maps display parameter adjustment strategies to specific module commands (or agent commands) to manage resource allocation and execution order across modules. For example, in a dimly lit scene with a face, the causal inference module can use commands to enhance local contrast and adjust color temperature. The orchestration module then sends instructions to the color and gamut module, which performs the local contrast enhancement. It can also send instructions to the lighting adjustment module, which adjusts the face area to a warmer color temperature.
[0296] The conflict resolution strategy described above can be used to resolve rendering conflicts between modules. This strategy can also be applied to conflicts between display metrics, such as those between thermal limits and contrast / sharpness.
[0297] For example, to prevent the panel from overheating, the peak backlight brightness is dynamically limited.
[0298] For example, a trade-off is made between display objectives such as power consumption and color fidelity, based on the current system status (such as the status of the display panel), real-time hardware constraints, and user preference information.
[0299] As one possible implementation, the orchestration module can also instruct each module on the aforementioned conflict resolution strategy, so that each module can perform rendering tasks based on the conflict resolution strategy, within the limits of hardware capabilities.
[0300] The method provided in this application embodiment allows the orchestration module to provide unified orchestration instructions to each module, thereby minimizing potential conflicts between modules. For example, when the optimization intention is to enhance color vibrancy and save power, the orchestration module will enable the corresponding module for color enhancement, but limit the peak brightness of the illumination adjustment module and may reduce the refresh rate, thus achieving the coordination of multiple objectives.
[0301] Furthermore, this orchestration module generates corresponding strategies in real time based on optimization intent and local context information. On the one hand, this ensures operational consistency. On the other hand, it allows the generated strategies to adapt to different video scenarios. Thus, it maintains intelligent and robust display performance across various receiving ends and video content types.
[0302] In some embodiments, the sending end further includes a semantic analysis module for semantic segmentation of the video content to generate semantic segmentation data. For example, the semantic analysis module can perform object-level segmentation and region classification to obtain semantic masks such as faces, text, sky, and background, as well as object labels.
[0303] In some embodiments, the sending end further includes: a causal reasoning module, used to generate an optimization intent based on a predefined causal graph, wherein the causal graph is used to describe the causal relationships corresponding to visual elements in the video content.
[0304] As one possible implementation, the causal reasoning module uses a causal decision network (CDN) to perform reasoning, evaluate and model the impact of video content and local contextual information (such as environment and user preference information) on the rendering results, in order to obtain the optimization intent of the video content.
[0305] For example, in a scenario where a user is watching a dark-toned movie in a bright living room, strictly adhering to the creative intent (maintaining an extremely dark tone) might result in the user not being able to see details, leading to a poor experience. Conversely, completely adapting to the environment (significantly brightening) could disrupt the atmosphere and render the film's artistic expression ineffective. The causal reasoning module can identify causal intervention points to optimize visibility while minimizing disruption to the creative intent. For instance, the generated optimization intent might be: precisely and slightly increase the brightness of only key subjects in the image (such as a person's face), while keeping pure background shadows and black borders extremely dark. Simultaneously, slightly increase the contrast of midtones. In this way, the user can clearly see important content in the video, while the overall dark tone is preserved to the greatest extent possible.
[0306] As another example, the causal reasoning module can obtain an optimization intent based on the semantic understanding of the video content and the state of the audience: adjust the tone curve to adapt to a scene containing human subjects under low ambient light.
[0307] As one possible implementation, the causal reasoning module can evaluate the scene, user intent, and perceptual priority of the video content, and combine this with a causal graph to output the optimized intent of the video content. For example, the optimized intent could be: preserve skin color contrast, preserve facial tones, minimize motion blur, reduce blue light in low ambient light, and enhance text readability.
[0308] Compared to black-box neural networks or dependency-based deep learning models, which yield unclear optimization intentions, the method provided in this application uses causal graphs to explicitly describe the logical causal relationships corresponding to visual elements. For example, "strong light → shadow → shadow needs to be brightened to preserve details." Therefore, the optimization intentions generated based on causal graphs have high interpretability and trustworthiness, as well as high robustness and reliability. Furthermore, its decision-making process is transparent and verifiable. This facilitates user understanding and alignment with the user's display goals. This is crucial for safety-critical applications such as medical and automotive applications.
[0309] Furthermore, causal reasoning can anticipate conflicts before problems occur, offering a proactive advantage in conflict resolution. For example, the sending end might reason that "increasing overall brightness" could lead to "overexposed faces," thus the generated optimization intent would include a constraint to preserve skin tone. This allows the receiving end's orchestration module to formulate a more comprehensive adjustment strategy.
[0310] Furthermore, the layered collaboration of the causal reasoning module, orchestration module, and other modules enables the display system to provide perceptual optimization, contextual responsiveness, and interpretable real-time rendering. This further enhances the personalized rendering of video content.
[0311] In some embodiments, display parameters are determined based on the weights of each module. For example, when the orchestration module coordinates multiple agents (modules), the influence of each agent on the final image can be quantified as a "weight". For instance, in high dynamic range (HDR) scenes, the color and color gamut modules may have higher weights to ensure color accuracy. In motion scenes, the dynamic refresh and backlight intelligence modules have higher weights to ensure smoothness.
[0312] In some embodiments, the receiving end further includes a verification module, used to: detect whether the rendering effect after adjusting the display parameters meets the hardware capabilities of the receiving end and the user's perception; if the rendering effect after adjusting the display parameters does not meet the hardware capabilities of the receiving end, or does not meet the user's perception, then the display parameters are adjusted to spare parameters, or the weights of each module are adjusted.
[0313] The verification module, also known as the verification and security execution module, is, for example, a low-latency execution module. For example, this verification module can be deployed in the SoC firmware or display pipeline to reduce the impact on the display performance of the receiving end. For example, it can reduce illusion effects and decrease grayscale instability.
[0314] As one possible implementation, the verification module can detect whether the rendering effect after adjusting the display parameters meets the hardware capabilities of the receiving end and the user's perception, based on preset rules. For example, a statistical threshold can be set, and the display parameters can be compared with the statistical threshold to determine whether the rendering effect meets the hardware capabilities of the receiving end and the user's perception.
[0315] Here, "satisfying user perception" means that the rendering result does not exhibit any visual anomalies that the user can perceive. Examples of visual anomalies include brightness flickering, color banding, clipping, unnatural skin tone changes, or incomplete visual images.
[0316] As one possible implementation, the verification module can trigger a correction process upon detecting a visual anomaly. In one possible design, the verification module can work in conjunction with the orchestration module. If the verification module detects that the rendering effect does not meet the hardware capabilities of the receiving end and user perception, it can send a message to the orchestration module. The orchestration module can then, based on this message, either fall back to alternative parameters or re-determine the weights of other modules and, based on those weights, instruct each module to adjust its display parameters.
[0317] For example, a fallback to backup parameters can be implemented, whose rendering effect meets the hardware capabilities of the receiving end and the user's perception. For instance, based on the maximum brightness limit of the OLED panel, a fallback to backup parameters can be implemented to avoid degradation caused by the panel being set to maximum brightness for extended periods. In this way, the probability of visual anomalies in the video can be reduced without exceeding hardware capabilities.
[0318] Another example is reducing the weight of certain agents.
[0319] The method provided in this application embodiment, through the verification module, can promptly detect rendering problems caused by hardware limitations (such as color banding, severe halo, and overheating throttling). This fallback mechanism ensures that the display system will not output unacceptable or hardware-damaging images, providing system-level robustness assurance.
[0320] Furthermore, the dynamic weight adjustment of the intelligent agent enables the display system to self-optimize. This allows the display system to learn and adapt, fine-tuning during operation to approach optimal performance. For example, if verification finds that the current motion vector causes a "jelly effect," the verification module can instruct the weight of the motion vector to be reduced to avoid the jelly effect as much as possible.
[0321] Furthermore, this verification module works in conjunction with the orchestration module to ensure consistency across scenes and devices, as well as safe and high-quality rendering.
[0322] The above only lists some application scenarios of the method in the embodiments of this application. The method can also be applied to other scenarios.
[0323] For example, the receiving end dynamically adjusts the display color gamut based on the color ratio information in the metadata and the panel characteristics, balancing color richness and visual comfort, and reducing visual fatigue from prolonged viewing.
[0324] For example, the color ratio information is as follows: { "color_primaries": { "red": {"x": 0.680, "y": 0.320}, / / Slightly different from standard P3 red (0.680, 0.320) "green": {"x": 0.265, "y": 0.690}, / / Different from standard P3 green (0.265, 0.690) "blue": {"x": 0.150, "y": 0.060}, "white_point": {"x": 0.314, "y": 0.351} / / White point at D63, not D65 }, "transfer_function": "PQ", / / Electro-optical conversion function is PQ "matrix_coefficients": "BT.2020" / / Use the BT.2020 color matrix } The receiver can dynamically adjust the color gamut by combining the coordinates of the three primary colors and the white point of its own panel to match the display capabilities of the receiver.
[0325] For example, the receiving end can dynamically adjust the contrast or local dimming based on the deep scene analysis and semantic segmentation of the sending end, so as to maintain high image quality in various scenarios.
[0326] For example, the receiving end can use semantically segmented metadata to adjust saturation and hue, correcting color distortion when viewed from a wide angle.
[0327] For example, the receiver can use illumination data to dynamically calibrate white balance and color temperature to ensure that the screen color is consistent with the real-world lighting environment.
[0328] For example, in a vehicle head-up display (HUD) system, it can support real-time adaptation to changing environmental conditions (such as lighting conditions) and driver status. By monitoring cockpit lighting and driver alertness through sensors, the system adjusts display parameters such as contrast and brightness of the displayed content accordingly to ensure that it does not interfere with the driver's normal driving, thereby enhancing driving safety and comfort.
[0329] For example, for educational video content, DP Vision can enhance important elements (such as key concepts or charts) in real time based on scene semantic understanding and causal understanding of the video content to align with the user's learning goals, help the user understand these important elements, and without disrupting the overall visual effect.
[0330] For example, for game-related video content, DP Vision can identify high-priority visual elements (such as co-op members, equipment, or UI overlays) based on semantic understanding of the video content, and ensure that these visual elements are rendered with minimal latency and maximum clarity. This provides a responsive and immersive gaming experience. Furthermore, it reduces motion blur and improves responsiveness in high-FPS scenes.
[0331] For example, for streaming media platforms, intelligent agents can be embedded in the content delivery process to optimize playback on heterogeneous devices by utilizing cloud-side metadata, player parsing, and edge intelligence.
[0332] For example, in scenarios such as holography and light fields, the stereoscopic conversion module can perform viewpoint synthesis, parallax balancing, and light plane segmentation to achieve a more immersive, multi-angle projection experience.
[0333] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0334] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0335] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0336] The display device, content-adaptive metadata-based display method, and storage medium provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application. The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0337] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0338] "A and or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0339] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration of devices to perform additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more conditions or values may, as examples, be based on additional conditions or values beyond those stated.
[0340] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0341] Based on the technical issues described in the background, and considering that generative super-resolution technology, as a core technology for improving image and video display quality, can recover high-resolution content with rich details from low-resolution original input, it has been widely used in various real-world scenarios such as smart display devices, video playback terminals, and streaming media transmission. In actual super-resolution processing, the input content to be processed exhibits a high degree of diversity, with significant differences in its domain, content type, and degradation patterns between different segments. This results in differentiated requirements for metadata (i.e., super-resolution prompts) to assist in super-resolution reconstruction for various types of content.
[0342] As an example, video content in real-world applications is often composed of alternating segments of various features. These include dense text segments containing subtitles and UI overlays, fast-moving segments such as sports events and action sequences, smooth animation clips like those in anime and cartoons, and low-light photo segments of faces and scenery taken at night or in low-light environments. The types of auxiliary information required for high-quality super-resolution reconstruction vary significantly depending on the content segments' features: dense text areas require metadata guided by edge sharpness to ensure clear text; high-speed motion scenes rely on metadata based on temporal motion cues to mitigate motion blur and artifacts; smooth animation clips are better suited to metadata based on texture completion to enhance image detail; and low-light photo segments require metadata based on noise suppression and brightness compensation to optimize image purity and visual effects. All types of content can achieve significant super-resolution quality improvements by using specific forms of auxiliary information that match their features.
[0343] In existing technologies, generative super-resolution methods based on metadata conditions mostly adopt fixed metadata condition design schemes. That is, during the inference phase, a preset single metadata type is always used, or a fixed combination of multiple metadata types is used indiscriminately. This does not consider the personalized and differentiated needs of different content and fragments for metadata. This fixed-configuration processing mode has significant performance limitations in practical applications: First, when the optimal prompt information for super-resolution reconstruction is highly dependent on the characteristics of the input content itself, fixed metadata input cannot flexibly adapt to diverse content and fragments, making it difficult to guarantee the reconstruction quality of various feature contents. It may even introduce reconstruction artifacts due to the mismatch between metadata and content features, reducing the display effect. Second, super-resolution processing in real-world scenarios faces strict resource budget constraints, including limitations on computational load, transmission bandwidth, and processing latency. Indiscriminate use of metadata will result in a large waste of resources, while a single metadata type cannot meet the reconstruction needs of different content. Neither can achieve a balance between resource consumption and super-resolution quality, ultimately leading to poor implementation results in real-world scenarios.
[0344] In summary, how to dynamically match the optimal metadata auxiliary information for diverse input content features, while adapting to resource budget constraints, and achieving the optimal balance between quality and overhead in generative super-resolution processing has become a pressing technical problem that needs to be solved in the application of generative super-resolution technology in real-world scenarios.
[0345] In view of this, embodiments of this application provide a display device, a display method based on content-adaptive metadata, and a storage medium. By using a metadata scheduler and a conditional super-resolution reconstructor, adaptive metadata selection based on content features is achieved, which solves the problems of low efficiency and quality caused by fixed metadata configuration in the prior art. It can dynamically select the optimal metadata subset according to the video content, achieve the optimal balance between quality and overhead under resource constraints, and improve display effect and system efficiency.
[0346] The technical terms used in this application will be explained below: Display device: This can refer to a hardware device capable of receiving video signals and converting them into visual images, such as a television, monitor, projector, or mobile terminal. The device is configured to integrate a display system for optimizing display quality.
[0347] Content-adaptive metadata refers to auxiliary information that can be dynamically adjusted and selected based on the characteristics of the video content itself. Metadata is used to guide the super-resolution processing to improve the reconstruction quality of different types of video content.
[0348] Display system: This can refer to a collection of software or hardware modules integrated within a display device, responsible for processing and optimizing the video content to be displayed. The system is configured to perform super-resolution processing of video content to provide a higher quality visual experience.
[0349] Metadata scheduler: This can refer to a functional module in a display system whose function is to analyze video content and intelligently select the most suitable metadata for the current content based on the analysis results. This scheduler is configured to ensure that, within resource constraints, only the metadata most beneficial to improving quality is enabled.
[0350] Conditional super-resolution reconstructor: This can refer to another functional module in a display system that is configured to receive a subset of target metadata selected by a metadata scheduler and perform super-resolution processing on low-resolution video content based on the metadata. This reconstructor is configured to utilize the additional information provided by the metadata to generate display content with higher resolution and finer details.
[0351] First-resolution video content: This can refer to the original video signal input to the display system for processing, which has a relatively low spatial resolution.
[0352] Content features: These refer to attribute information extracted from video content that describes the dynamic characteristics and content composition of the video. They are used by the metadata scheduler for decision-making, such as the video's motion intensity, text ratio, and texture density.
[0353] Metadata categories: These can refer to pre-defined, selectable sets of different types of auxiliary information. Each category represents a specific source of information, such as degradation cues, structural information, temporal motion cues, semantic masks, or text region information.
[0354] Target metadata subset: This can refer to the set of metadata that the metadata scheduler selects from multiple metadata categories based on content characteristics, and that is most suitable for super-resolution processing of the current video content.
[0355] The second resolution display content can refer to the video content output after being processed by the conditional super-resolution reconstructor. Its spatial resolution is higher than the original first resolution, thus providing a clearer and more detailed visual effect.
[0356] like Figure 1 The diagram shown is a schematic of a display device provided in an embodiment of this application. The display device includes a content-adaptive metadata-based display system, which comprises a metadata scheduler and a conditional super-resolution reconstructor. The metadata scheduler performs feature analysis on the video content to be displayed at a first resolution, extracts content features of the video content, and selects a target metadata subset from multiple preset metadata categories based on the content features. The content features include attribute information describing the dynamic characteristics and content composition of the video content. The conditional super-resolution reconstructor performs super-resolution processing on the video content based on each metadata category in the target metadata subset, outputting display content at a second resolution, which is greater than the first resolution.
[0357] The metadata categories include at least a portion of degradation cues, structural and edge information, temporal motion cues, semantic masks, regions of interest (ROI) graphs, and text regions. Degradation cues include noise levels, artifact scoring, and blur proxy information; structural and edge information includes edge graphs and gradient graphs; temporal motion cues include optical flow, motion vectors, and shot transition markers. Semantic masks may include coarse-grained region labeling; ROI graphs may include importance weights for faces, salient regions, and subject regions; and text regions may include positional mask information for video subtitles and UI overlays. The conditional super-resolution reconstructor can receive video content, a subset of target metadata, and a selection mask through a unified interface.
[0358] As an example, the display device can be a standalone display terminal, such as a smart TV or computer monitor, or a display module integrated into other devices, such as a smartphone or tablet. The display system is configured to dynamically adjust the auxiliary information used in super-resolution processing based on the characteristics of the video content to be displayed, in order to optimize the display effect. For example, the display system can be implemented as a set of software programs running on the processor of the display device, or as a dedicated hardware acceleration module. Further, the display system is configured to include a metadata scheduler and a conditional super-resolution reconstructor. The metadata scheduler is configured to analyze the input video content and intelligently select appropriate metadata. The conditional super-resolution reconstructor is configured to perform super-resolution processing using the selected metadata. For example, the metadata scheduler can be implemented as a lightweight neural network model trained to predict the optimal combination of metadata based on the features of the input video. The conditional super-resolution reconstructor can be implemented as a deep learning-based super-resolution model, such as a convolutional neural network or a generative adversarial network, designed to receive and utilize external conditional information.
[0359] As an example, a metadata scheduler is used to perform feature analysis on the video content to be displayed at a first resolution and extract content features from the video content. This feature analysis process may include image processing of video frames, such as calculating the average brightness, contrast, or color histogram of the image. Statistical methods, such as sampling and statistically analyzing the pixel values of video frames, can also be used to obtain their basic attributes. For example, the video content can be analyzed frame-by-frame to calculate the average pixel value of each frame as a brightness feature, or the pixel differences between adjacent frames can be calculated as motion features. Based on this, the metadata scheduler is used to select a target metadata subset from multiple preset metadata categories based on these content features. This selection process can be based on preset rules; for example, if low brightness is detected in the video content, a metadata category related to low-light enhancement is selected. Alternatively, a simple lookup table mechanism can be used to map different content features to a predetermined combination of metadata categories. For example, when the video content is identified as "cartoon," the scheduler can be configured to select the "smooth texture" metadata category; when the video content is identified as "news broadcast," the scheduler can be configured to select the "text sharpening" metadata category. Content features include attribute information describing the dynamic characteristics and content composition of the video content. Attribute information may include basic information such as the video's frame rate, encoding format, and resolution. It may also include macroscopic descriptions obtained from preliminary analysis of the video content, such as whether the video is static or dynamic, and whether it is an indoor or outdoor scene.
[0360] For example, frame rate and encoding format can be obtained from the metadata of the video file, or the video content can be roughly classified into static or dynamic scenes. Therefore, a conditional super-resolution reconstructor is used to perform super-resolution processing on the video content based on the metadata categories in the target metadata subset. This processing can involve inputting selected metadata as additional input signals along with the low-resolution video content into the super-resolution model. For example, if the target metadata subset includes "edge information" metadata, this edge information can be directly superimposed onto the low-resolution video frame, and the superimposed image is then input into the super-resolution model for processing. Alternatively, if the target metadata subset includes "color correction parameters" metadata, these parameters can be used to adjust the color of the video content before or after super-resolution processing. Ultimately, the conditional super-resolution reconstructor outputs a second resolution display content, which is greater than the first resolution. Thus, the processed video content has a higher pixel density, resulting in a clearer and more detailed image on the display device.
[0361] For example, the display system includes a metadata orchestrator and a backbone-agnostic conditional super-resolution (SR) reconstructor. The metadata orchestrator performs feature analysis on the low-resolution input signal y, which serves as the first-resolution video content in the file. It extracts a context descriptor c, composed of high-level content identifiers and segment-level attributes. This descriptor represents the content features describing the dynamic characteristics and content composition of the video frame. Based on these features, the orchestrator selects the optimal metadata subset m(z) from a pre-defined set of K candidate metadata categories M, corresponding to the target metadata subset. The conditional super-resolution reconstructor then performs generative super-resolution processing on the low-resolution video content based on the selected optimal metadata subset, ultimately outputting a high-resolution estimate of the content to be displayed at the second resolution. Specifically, it is expressed as: z = π ( c ), Where z is an explicit mask, which supports the variability and missingness of metadata, making the display system robust under heterogeneous processing flow and runtime constraints, and the high resolution is higher than the original low resolution.
[0362] This embodiment achieves adaptive super-resolution processing of video content by introducing a metadata scheduler and a conditional super-resolution reconstructor. This scheme dynamically selects the most suitable subset of metadata based on the dynamic characteristics and content composition of the video content, effectively overcoming the limitations of traditional fixed metadata configuration schemes when processing diverse video content. Therefore, under resource budget constraints, this embodiment can optimize the quality and efficiency of super-resolution processing, improve the clarity and detail of displayed content, and enhance the flexibility and robustness of the display system.
[0363] In some embodiments, the content features include content feature identifiers and segment-level attributes; the content feature identifiers include at least one of the domain, subject matter, or channel type information of the video content; the segment-level attributes include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level.
[0364] As an example, content feature identifiers can correspond to lower-level content identifiers. Content feature identifiers can be attributes used for macro-level classification and identification of video content. They can be obtained from various sources, such as video metadata (e.g., file header information, streaming media tags), user viewing history, or content provider classifications. For example, a domain can refer to the macro-level category to which the video content belongs, such as movie, TV series, sports events, or news; a genre can be further refined into science fiction, comedy, action, or documentary; and a channel type refers to the specific channel from which the video content originates, such as "original content from channel XX." Identifiers help the metadata scheduler to initially understand the overall nature of the video and the expected viewing experience, thereby enabling coarse-grained metadata selection.
[0365] Fragment-level attributes describe the local and dynamic characteristics of video content in time or space. These attributes are calculated in real-time or offline using image processing and analysis algorithms on video frames or frame sequences. Text proportion measures the percentage of text regions in a video frame, which can be estimated using lightweight text detectors (e.g., optimized deep learning OCR models) or edge density combined with stroke heuristics (e.g., identifying text regions by analyzing image edge features and stroke structure). Object proportion measures the percentage of specific objects (e.g., faces, people, vehicles, specific products) in a video frame, which can be detected at a first resolution using compact detectors (e.g., lightweight object detection models like MobileNet and YOLO-tiny) to reduce computational complexity. Texture density measures the richness of texture details in a video frame, which can be characterized using gradient energy (e.g., calculating the average or sum of image gradient magnitudes) or local variance statistics (e.g., calculating the variance of pixel values in local image regions). Low-light noise level (e.g., low light or noise level) is used to measure the significance of noise in a video image under low light conditions. It can be estimated by using brightness distribution and noise proxy (e.g., by analyzing the brightness histogram distribution of the image and combining it with the statistical characteristics of high-frequency components to estimate the noise level).
[0366] For example, the proportion of the target object can be the proportion of text, face or person; the level of low light noise can be a low light or noise indicator; the content feature extracted by the metadata scheduler, i.e. the context descriptor c, can be decomposed into two parts g and a, which can be written as c=(g,a), where g is a high-level content identifier, which contains information such as the domain, theme, and channel type of the video content, and a is a segment-level attribute. The definition of a in the file explicitly includes dimensions such as motion intensity, text proportion, face or target object proportion, texture density, low light or noise indicator.
[0367] This application refines content features into content feature identifiers and segment-level attributes, enabling a more comprehensive and accurate understanding of the macro-type and micro-dynamic characteristics of video content. Content feature identifiers allow the metadata scheduler to perform preliminary, categorized metadata selection based on the overall attributes of the video (such as domain, subject matter, and channel type), ensuring that the selected metadata matches the general type of the video content. Simultaneously, segment-level attributes (such as motion intensity, text proportion, target object proportion, texture density, and low-light noise level) provide detailed information about the video content in both local and temporal dimensions, allowing the metadata scheduler to perform more refined and targeted metadata selection based on the specific details and dynamic changes of the video content. For example, for video segments with high motion intensity, metadata related to motion compensation can be prioritized; for segments with a high text proportion, text enhancement metadata can be emphasized; and for segments with high low-light noise levels, noise reduction metadata can be prioritized. Layered and detailed feature analysis significantly improves the accuracy and adaptability of target metadata subset selection, enabling the conditional super-resolution reconstructor to perform more effective super-resolution processing based on more accurate metadata guidance, ultimately outputting higher quality display content and optimizing the display effect of different types of video content.
[0368] In some embodiments, the metadata scheduler includes at least one of the following modules: a motion intensity calculation module for estimating the motion intensity of video content using block motion statistics, optical flow amplitude, or frame difference energy; a text detection module for estimating the text proportion of video content using a lightweight text detector or an edge density combined with a stroke heuristic; a target object detection module for estimating the target object proportion of video content using a compact detector at a first resolution; a texture density calculation module for characterizing the texture density of video content using gradient energy or local variance statistics; and a dark light noise detection module for estimating the dark light noise level of video content using brightness distribution and noise proxy.
[0369] As an example, a metadata scheduler may include a motion intensity calculation module for estimating the motion intensity of video content using block motion statistics, optical flow amplitude, or frame difference energy. For instance, block motion statistics divides video frames into macroblocks, calculates motion vectors between macroblocks in adjacent frames, and statistically analyzes the amplitude of these motion vectors to quantify the overall intensity of motion. Optical flow amplitude methods calculate the instantaneous motion vector of each pixel in the video; its magnitude directly reflects the intensity of local motion. Frame difference energy calculates the sum of the absolute or squared differences between pixel values in adjacent frames; a larger difference indicates more significant motion.
[0370] In addition, the metadata scheduler may include a text detection module for estimating the text proportion of video content using either a lightweight text detector or an edge density combined with stroke heuristics. Lightweight text detectors are typically optimized, small deep learning models capable of quickly identifying text regions in a video frame and calculating their proportion of the total frame. Edge density combined with stroke heuristics, on the other hand, utilizes image processing techniques to identify and locate text regions by analyzing the edge features of the image and the inherent stroke structure of the text characters.
[0371] To identify key content in the video, the metadata scheduler can further include a target object detection module to estimate the proportion of target objects in the video content at a first resolution using a compact detector. A compact detector refers to a target detection model with low computational resource consumption and a simplified model structure, capable of quickly and accurately identifying predefined target objects (such as faces, specific products, scene elements, etc.) in the video and calculating the area proportion of the target object in the frame without resolution upscaling.
[0372] To characterize the richness of detail in video footage, the metadata scheduler can include a texture density calculation module to represent the texture density of video content using gradient energy or local variance statistics. Gradient energy reflects the drasticness of local brightness changes by calculating the gradient magnitude of image pixels; the larger the gradient magnitude, the richer the texture. Local variance statistics calculate the variance of pixel values within local regions of the image; the larger the variance, the richer the texture detail in that region.
[0373] To address video quality issues in low-light environments, the metadata scheduler can also include a low-light noise detection module. This module estimates the level of low-light noise in the video content by analyzing brightness distribution and noise proxies. The low-light noise detection module first analyzes the brightness histogram of the video frames to identify low-brightness areas. Then, within these low-brightness areas, it quantifies the noise intensity by statistically analyzing pixel value fluctuations, local contrast, or using a specific noise model.
[0374] For example, the previously described 'a' includes the following: methods for estimating motion intensity through block motion statistics, optical flow amplitude, or frame difference energy constitute the core implementation of the motion intensity calculation module; methods for estimating text proportion through a lightweight text detector or edge density combined with stroke heuristics represent the functionality of the text detection module; methods for estimating the proportion of faces or target objects at low resolution (i.e., first resolution) using a compact detector correspond to the technical features of the target object detection module; methods for characterizing texture density through gradient energy or local variance statistics represent the implementation of the texture density calculation module; and methods for estimating the degree of dark light noise through brightness distribution and noise proxy constitute the core function of the dark light noise detection module.
[0375] In summary, the metadata scheduler can quantitatively analyze segment-level attributes of video content, such as motion intensity, text proportion, target object proportion, texture density, and low-light noise levels. This allows the metadata scheduler to obtain more comprehensive and detailed content feature information, enabling it to more accurately match the actual needs of the video content when selecting a subset of target metadata from multiple preset metadata categories. For example, for video content with high motion intensity, metadata categories related to motion compensation or temporal enhancement can be prioritized; for video content with a high text proportion, metadata categories that improve text clarity can be selected. This metadata selection mechanism based on refined content features significantly improves the targeting and effectiveness of super-resolution processing, avoiding blind or unnecessary metadata processing. This ensures the quality of displayed content while optimizing resource utilization efficiency, ultimately outputting visually superior second-resolution display content.
[0376] In some embodiments, Figure 2 This is a schematic diagram of the metadata scheduler provided in an embodiment of this application. For example... Figure 2 As shown, the metadata scheduler also includes a mask generation module, which generates a corresponding selection mask based on the target metadata subset, so that the conditional super-resolution reconstructor can determine the metadata categories included in the target metadata subset based on the selection mask. The selection mask is used to identify the metadata categories included in the target metadata subset.
[0377] As an example, the mask generation module transforms the target metadata subset selected by the metadata scheduler into a structured, easily parsed format—a selection mask. This mask generation module can be a software module that receives the identification information of the target metadata subset as input and transforms it into a standardized data structure according to preset encoding rules. For example, if multiple metadata categories are preset, the mask generation module can generate a binary bitmask, where each bit represents a predefined metadata category, and the bit's state (e.g., 0 or 1) indicates whether the category is selected. Furthermore, the selection mask can also be a list containing the index numbers or unique identifiers of the selected metadata categories, or a sparse vector where only the dimension corresponding to the selected metadata category has a non-zero value.
[0378] The selection mask explicitly identifies the various metadata categories contained in the target metadata subset. It serves as the interface between the metadata scheduler and the conditional super-resolution reconstructor, ensuring the accuracy and efficiency of information transmission. Through this explicit identification, the conditional super-resolution reconstructor no longer needs to analyze or infer the content of the target metadata subset itself; instead, it directly obtains the required information by parsing the selection mask. The conditional super-resolution reconstructor may internally contain a parser or decoder to read and interpret the selection mask. For example, when the selection mask is a binary bitmask, the parser iterates through each bit of the mask, activating or disabling the corresponding metadata processing branch or module based on the bit's state. This ensures that super-resolution processing can precisely adjust its behavior according to the metadata scheduler's decisions, thereby achieving adaptive processing of the video content.
[0379] For example, one of the core designs of the MetaSR framework is that a binary selection mask z is generated by the metadata scheduler, z=(z1,…,zK)∈{0,1}, where K represents selection. The mask is generated based on the optimal metadata subset m(z) selected by the scheduler, m(z)={mk|zk=1}, which is the target metadata subset. zk=1 indicates that the corresponding metadata category is enabled, and zk=0 indicates that it is disabled. The mask directly identifies the target metadata subset. z=π(c) is generated by the scheduler, which maps the context c to the metadata selection mask z. The reconstructor takes (y,m(z),z) as input and determines the selected metadata category based on the mask z. (See also...) Furthermore, the binary selection mask z is used to identify the selected metadata category in the candidate metadata category set. The conditional super-resolution reconstructor can know the currently selected metadata category by parsing the mask z.
[0380] This application introduces a mask generation module, enabling the metadata scheduler to generate a corresponding selection mask based on the selected target metadata subset. This selection mask identifies the metadata categories contained in the target metadata subset in a clear and standardized manner, allowing the conditional super-resolution reconstructor to efficiently and accurately determine and utilize the required metadata categories based on this selection mask. This avoids the conditional super-resolution reconstructor performing additional parsing or inference on the metadata subset during processing, significantly improving the efficiency and accuracy of super-resolution processing and ensuring that super-resolution processing accurately responds to the metadata scheduler's decisions, thereby optimizing resource utilization and improving the quality of displayed content.
[0381] In some embodiments, such as Figure 2As shown, the metadata scheduler also includes a gain prediction module and a greedy selection module; the gain prediction module is used to predict the expected quality gain of each metadata category under the current content features; the greedy selection module is used to calculate the unit cost quality gain index of each metadata category, and select the metadata category according to the unit cost quality gain index of each metadata category until the preset resource budget constraint is met, and generate a selection mask.
[0382] The gain prediction module uses a small network to predict the expected quality gain of each metadata category under the current content features. For example, it can estimate the potential improvement in display quality after video content super-resolution processing based on historical data, machine learning models, or predefined rules. For instance, predicting temporal motion cue metadata may bring a higher quality gain for specific content features (such as high-motion scenes), while text region metadata may be more valuable for scenes containing a large amount of text. The expected quality gain can be quantified as an improvement in PSNR (Peak Signal-to-Noise Ratio) or SSIM (Structural Similarity Index), or represented by predicted values of user-perceived quality (such as MOS score).
[0383] The greedy selection module chooses the optimal metadata category at each step, i.e., the category with the highest quality-per-unit-cost (MPC) gain. The MPC gain is typically calculated by dividing the expected quality gain by the resource overhead (e.g., computational complexity, memory usage, or processing latency) of that metadata category under the current content characteristics. The preset resource budget constraint refers to the maximum resource consumption limit the system can withstand when selecting metadata categories. This may include, but is not limited to, total computation (e.g., FLOPs), total memory usage, total processing latency, or power consumption. This constraint is a threshold pre-set during system design based on hardware capabilities, real-time requirements, or user experience goals. The greedy selection module continuously accumulates the resource overhead of the selected metadata categories. Once the accumulated value reaches or exceeds the preset resource budget constraint, the selection process stops, ensuring the system operates within acceptable resource limits. The MPC gain describes the ratio between the improvement in super-resolution quality brought by selecting a particular metadata category and the computational resource cost (including bandwidth, computation, or latency) incurred by using that metadata. The MPC gain can be selected in descending order. Resource budget constraints can refer to resource limitations that are dynamically adjusted based on the contextual characteristics of the video content when allocating resources for super-resolution processing.
[0384] For example, when the metadata scheduler generates the selection mask, it uses a small network hω( A learning-based gain predictor is constructed to predict the gain score Δk(c) of each metadata category under the current content feature, i.e., the context descriptor c. This predictor is the gain prediction module in the claims, realizing the prediction function of expected quality gain. Simultaneously, the scheduler calculates the normalized score sk(c) of the unit cost quality gain based on the gain score and the abstract cost function of each metadata category, and selects metadata categories in descending order of this score using a greedy selection strategy until the preset resource budget constraint B(c) in the file is met. Finally, a binary selection mask z is generated, and it is explicitly stated that the selection problem is equivalent to a 0 / 1 knapsack problem, with the core being gain maximization under budget constraints. Specifically, a small network is used... h ω ( Implement the scheduler π ( c ), which predicts gain scores for each metadata category. ( c ), Then the normalized score is calculated. , and according to s k ( c Greedily select categories in descending order until the budget constraint B(c) is satisfied, generating a binary mask z. Let... Δ k ( c The expression represents the (unknown) expected quality gain resulting from enabling category k in context c. A natural selection principle is... suchthat ≤ B ( c This problem is equivalent to the classic 0 / 1 knapsack problem.
[0385] This application effectively addresses the problem of optimizing metadata selection under limited resources by introducing a gain prediction module and a greedy selection module. The gain prediction module intelligently assesses the potential contribution of different metadata categories to display quality based on video content characteristics, providing a quantitative basis for subsequent decisions. Building upon this, the greedy selection module calculates a unit cost quality gain index, prioritizing metadata categories that deliver the greatest quality improvement with minimal resource consumption at each selection step in an efficient and pragmatic manner. This dynamic selection mechanism, based on expected quality gain and resource budget constraints, enables the metadata scheduler to adaptively generate the optimal selection mask according to the characteristics of the current video content and system resource status. Ultimately, this not only ensures that super-resolution processing achieves significant display quality improvements even in resource-constrained environments but also avoids unnecessary resource waste, thus achieving both high performance and economic efficiency in system operation.
[0386] In some embodiments, such as Figure 2 As shown, the metadata scheduler also includes a hybrid rules module, which introduces domain heuristic rules to supplement the selection logic of the greedy selection module; the heuristic rules include selecting text region metadata when the text proportion exceeds a threshold and selecting temporal motion clue metadata when the motion intensity exceeds a threshold.
[0387] The hybrid rules module is configured to introduce domain-specific heuristics into the metadata selection process. Its primary function is to provide the greedy selection module with additional decision-making support based on experience or expert knowledge, compensating for the limitations of purely quantitative indicator-based selection. Through the hybrid rules module, the system can more intelligently identify strong correlations between specific content features and specific metadata categories, ensuring that, within the resource budget, metadata that has a decisive impact on a specific content type is prioritized. Domain-specific heuristics are formulated based on a deep understanding of video content characteristics and their impact on super-resolution processing. They typically exist in the form of "if condition A is met, then perform operation B," aiming to capture decision logic that is difficult to reflect through simple cost-benefit analysis but is crucial to visual quality. Rules can be predefined and stored in the system and invoked when the metadata scheduler makes selections.
[0388] As an example, when the proportion of text in the video content (e.g., calculated by the text detection module) exceeds a preset threshold, it indicates the presence of a large amount of text information in the video. In this case, even if the unit cost quality gain of the text region metadata is not the highest, the hybrid rule module will force or prioritize the selection of text region metadata. Selecting text region metadata helps to better preserve and enhance the clarity and edge sharpness of text in super-resolution processing, avoiding text blurring or distortion, thereby significantly improving the user's reading experience. This threshold can be set according to the actual application scenario and user experience requirements; for example, it can be triggered when the text proportion reaches 5% or 10%. Simultaneously, when the motion intensity of the video content (e.g., estimated by the motion intensity calculation module) exceeds a preset threshold, it indicates the presence of intense or rapid motion in the video. In this case, the hybrid rule module will force or prioritize the selection of temporal motion cue metadata. Temporal motion cue metadata can provide the conditional super-resolution reconstructor with information about the trajectory and speed of object motion, thereby better handling motion blur, artifacts, and other issues in super-resolution processing, ensuring the smoothness and clarity of moving images. This threshold can also be adjusted according to the type of video content (e.g., sports events, action movies) and the desired visual effects.
[0389] For example, to enhance the robustness and interpretability of the greedy selection strategy of the metadata scheduler, a domain heuristic rule is introduced to initialize and supplement the greedy selection logic based on the learned gain predictor. The implementation of this rule corresponds to the hybrid rule module in the claims. When the text proportion exceeds a threshold, the metadata of the text region is forcibly selected. m text When the motion intensity exceeds a threshold, temporal motion cue metadata should be forcibly selected. m temp .
[0390] This application effectively addresses the limitations that may arise when selecting metadata solely based on quality gain per unit cost by introducing a hybrid rule module and supplementing the selection logic of the greedy selection module with domain-based heuristics. For example, when video content has significant features such as high text content or high motion intensity, even if the corresponding text region metadata or temporal motion cue metadata might not be prioritized in a conventional greedy selection, the hybrid rule module ensures that key metadata is included in the target metadata subset. This allows the conditional super-resolution reconstructor to obtain more targeted auxiliary information when processing specific types of video content, resulting in clearer and sharper text display in text regions and smoother, less blurry dynamic images in motion scenes. The hybrid selection strategy, combining quantitative indicators and domain experience, not only optimizes the intelligence of metadata selection but also significantly improves the super-resolution display quality of display devices when processing diverse video content, especially in the detail representation of specific content elements (such as text and motion), providing a superior visual experience.
[0391] In some embodiments, Figure 3 This is a schematic diagram of a conditional super-resolution reconstructor provided in an embodiment of this application. For example... Figure 3 As shown, the conditional super-resolution reconstructor includes a metadata encoder and a feature fusion module. The metadata encoder is used to extract features from heterogeneous metadata in the target metadata subset and generate metadata features in a unified format. The heterogeneous metadata includes dense spatial graphs, temporal signals, compact vectors, or scalars. The feature fusion module is used to fuse the metadata features with the features of the video content to obtain enhanced features, enabling the conditional super-resolution reconstructor to perform super-resolution processing on the video content based on the enhanced features.
[0392] The metadata encoder is a key component for processing heterogeneous metadata. Its main function is to transform heterogeneous metadata from different metadata categories, with different data structures and semantics, into a unified metadata feature representation that can be processed by the subsequent feature fusion module. For example, for dense spatial graphs (such as semantic segmentation graphs and saliency graphs), the metadata encoder can use a convolutional neural network (CNN) structure for feature extraction, converting it into a high-dimensional feature map or feature vector; for temporal signals (such as motion vectors and inter-frame differences), it can use temporal models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or Transformers to capture their temporal dependencies and generate temporal feature vectors; for compact vectors (such as scene classification vectors and style description vectors), it can perform dimensionality transformation and nonlinear mapping through fully connected layers or multilayer perceptrons (MLPs); for scalars (such as mean brightness and contrast), they can be embedded into a high-dimensional space, for example, by converting them into vector representations through linear layers or lookup tables. In this way, the metadata encoder effectively eliminates the heterogeneity between metadata, laying the foundation for subsequent feature fusion.
[0393] The feature fusion module is responsible for effectively combining the unified-format metadata features output by the metadata encoder with the features of the original video content, thereby generating enhanced features containing richer contextual information. These enhanced features serve as input to the conditional super-resolution reconstructor for super-resolution processing. Feature fusion can be implemented in various ways. For example, concatenation fusion can connect metadata features and video content features along the channel dimension to form a wider feature tensor, which is then used to learn the interactions between features through convolutional layers. Addition fusion or multiplication fusion can also be used, performing element-wise operations after dimension matching to achieve feature modulation. More complex fusion mechanisms can include gating fusion or cross-modal attention, allowing metadata features to act as gating signals or query vectors, dynamically adjusting the weights and representations of video content features to learn deeper, more complex relationships between them. Furthermore, feature fusion can be performed at multiple layers of the super-resolution network to provide more refined guidance. Enhanced features can be fused using concatenation fusion or feature-level modulation methods. For example, feature-level modulation uses the FiLM gating mechanism, which generates layer-by-layer modulation parameters through metadata embedding, and uses the modulation parameters to perform affine transformation on the intermediate feature map of the video content to achieve feature modulation.
[0394] For example, the conditional super-resolution reconstructor sets up a unified metadata encoder ε to handle heterogeneous metadata. This encoder extracts features from heterogeneous metadata, including dense spatial graphs, time-series signals, compact vectors, and scalars, in candidate metadata categories, generating metadata features ek in a unified format. e k =ε k ( m k ),for z k =1 Simultaneously, the reconstructor incorporates a feature fusion stage. Through methods such as concatenation fusion or feature-wise modulation (e.g., FiLM / gating mechanisms), it fuses metadata features in a unified format with visual features extracted from low-resolution video content, resulting in enhanced features containing richer contextual information. Based on these enhanced features, super-resolution processing is then performed on the video content. As an example, feature-wise modulation utilizes metadata embedding to generate layer-by-layer modulation parameters, which can be found in [reference needed]. F′=γ ( e,z )⊙ F + β ( e,z FiLM-style modulation is particularly effective when metadata contains a mixture of spatial and non-spatial signals, and supports variable metadata inputs through an explicit mask z.
[0395] This application effectively addresses the challenge of directly applying heterogeneous metadata to super-resolution processing. The metadata encoder unifies different types of metadata into standardized metadata features, eliminating format and semantic differences. Subsequently, the feature fusion module deeply integrates the unified metadata features with video content features, generating enhanced features rich in contextual information. These enhanced features not only include the visual information of the video itself but also incorporate high-level semantic information provided by the metadata, such as dynamic characteristics and content composition. Therefore, when performing super-resolution processing, the conditional super-resolution reconstructor can perform targeted detail reconstruction and texture enhancement of video content based on more comprehensive and accurate enhanced features. For example, it can achieve clearer text reconstruction in text areas, reduce artifacts in motion areas, or enhance detail in specific target object areas. This improves the quality and efficiency of super-resolution processing, resulting in visually clearer and more natural output at second-resolution display, better preserving the semantic information of the original video content, thus providing users with a superior viewing experience.
[0396] In some embodiments, each metadata category corresponds to an abstract cost function, which is a non-negative function that depends on content features and is used to characterize the resource overhead of metadata under the corresponding metadata category; the total cost corresponding to the selection mask is the sum of the costs of each selected metadata category under the current content features.
[0397] As an example, the abstract cost function is a mathematical model used to quantify the resource consumption required to process a specific metadata category. Resource overhead can include, but is not limited to, computation time, memory usage, bandwidth requirements, latency, or power consumption. The function is designed to depend on the content characteristics of the current video content; for example, for video content with high motion intensity, the cost of processing temporal motion cue metadata may be higher than the cost of processing static texture metadata. Furthermore, the function is set to be non-negative to ensure that the quantification of resource overhead is practically meaningful. In implementation, the abstract cost function can be predicted using a pre-trained model (e.g., a machine learning-based model) that takes content features as input and outputs an estimate of the resource overhead for the corresponding metadata category; alternatively, it can be determined using empirical formulas or lookup tables based on extensive testing and statistical analysis of the actual resource consumption of different metadata categories under different content characteristics.
[0398] When the metadata scheduler generates a selection mask to identify the metadata categories contained in the target metadata subset, it calculates a total cost. The total cost is obtained by summing the abstract cost function values of each selected metadata category under the current content features. The total cost reflects the total resource consumption required to process the entire target metadata subset, providing a quantitative basis for subsequent resource budget management and optimization. For example, if three metadata categories A, B, and C are selected, and their abstract costs under the current content features are Cost(A), Cost(B), and Cost(C), respectively, then the total cost corresponding to the selection mask is Cost(A) + Cost(B) + Cost(C).
[0399] As an example, the function of a selection mask is to use a binary identifier consisting of 0s and 1s. "1" represents that the corresponding metadata category is selected, and "0" represents that it is not selected. The cost of a single metadata category: Each metadata category corresponds to an abstract cost function. The cost value of the function is not fixed and changes dynamically based on the characteristics of the input content. For example, the cost of motion cue metadata in a high-speed motion segment is higher than in a static segment. The logic for calculating the total cost: It iterates through all metadata categories marked "1" in the selection mask, extracts the cost value of each metadata category under the current content features, and adds the values together. The result is the total cost corresponding to the selection mask. For example, suppose there are three categories of candidate metadata: degradation cues, structural and edge information, and temporal motion cues. In a static landscape segment, the mask is [1,1,0], representing the selection of the first two metadata categories. If the cost of degradation cues is 2, the cost of structural and edge information is 3, and the cost of temporal motion cues is 5, then the total cost corresponding to this selection mask = 2 + 3 = 5.
[0400] For example, to achieve budget-aware optimization of metadata selection, an abstract cost function costk(c) is defined for each metadata category k. This function is a non-negative function that depends on the content features, i.e., the context descriptor c, and represents the resource overhead in terms of bandwidth, computation, latency, etc., when using the corresponding metadata category. At the same time, the total cost corresponding to the selection mask z is defined in the file as Cost(z;c)=∑zk costk(c), where z_k is the binary selection mask, and costk(c) is the cost of category k in context c. That is, the sum of the costs of the selected metadata categories under the current content features. As an example, let costk(c) ≥ 0 represent the context-dependent cost of using metadata category k (e.g., a proxy metric of bandwidth, computation, or latency). The total cost of the selection mask z is... The goal is to maximize super-resolution quality under resource constraints. This can be written as a constrained optimization problem: , making z = π ( c ), Cost ( z ; c )≤ B ( c ), where L( B(c) is the reconstruction loss (e.g., mean squared error / perceptual / ROI weighted loss and temporal loss for video), and B(c) is the context-dependent budget. Equivalently, a learning-friendly Lagrangian (penalized) form is used. where z = π ( c), where λ>0 represents the trade-off between control quality and overhead.
[0401] This application incorporates an abstract cost function and calculates the total cost corresponding to the selection mask, thereby taking resource overhead into account when selecting metadata. This allows the metadata scheduler to not only improve display quality based on content features when selecting a subset of target metadata, but also to simultaneously assess and manage the required computing resources. For example, in resource-constrained environments, metadata categories with lower resource overhead can be prioritized while ensuring a certain quality gain, thus achieving optimal super-resolution results within a limited resource budget. This avoids the problem of resource exhaustion or performance degradation caused by blindly selecting high-cost metadata, significantly improving the adaptability and efficiency of display devices under different operating conditions, and ensuring the stability and sustainability of super-resolution processing.
[0402] In some embodiments, the conditional super-resolution reconstructor is integrated into the diffusion model to generate second-resolution display content through an iterative denoising process by using video content, a subset of target metadata, and a selection mask as conditional inputs to the denoiser.
[0403] As an example, the conditional super-resolution reconstructor is designed with a diffusion model-based architecture. A diffusion model is a generative model that learns the distribution of data by simulating the process of gradually denoising data from noise. In super-resolution tasks, a diffusion model learns how to recover a high-resolution image from a low-resolution image (or its noisy representation). This ensemble typically involves a deep neural network (e.g., U-Net) acting as a denoiser, which receives the current noisy image and conditional information at each time step and predicts the noise to be removed. By iteratively performing the denoising process, the model can progressively transform low-resolution input into high-resolution output. Applying a diffusion model to super-resolution reconstruction effectively improves the processed visual quality and realism, especially when dealing with complex textures and details, generating more natural and realistic high-resolution images.
[0404] In the above scheme, video content, a subset of target metadata, and a selection mask are used as conditional inputs to the denoiser. Conditional inputs refer to providing additional guiding information beyond the current noisy image during the diffusion model's denoising process, directing the generation process towards a specific target or satisfying specific conditions. The video content, i.e., the original first-resolution video content, can be directly used as the input feature map to the denoiser, or its features can be extracted by the encoder and then used as conditions. For example, low-resolution video frames can be encoded using convolutional layers to obtain their spatial feature representations, which can then be fused with the intermediate layer features of the diffusion model, for example, through cross-attention mechanisms or feature concatenation. The subset of target metadata may contain various heterogeneous data, such as dense spatial maps, temporal signals, compact vectors, or scalars. The metadata needs to be processed by a metadata encoder to extract features, generating metadata features in a uniform format. Metadata features can be used as global conditions, for example, injected into different layers of the denoiser through adaptive layer normalization (AdaLN) or FiLM layers, or interact locally with video features through attention mechanisms. The selection mask is a binary vector or tensor used to indicate which metadata categories are selected within the subset of target metadata. It can serve as additional conditional input, for example, by embedding it as a vector and injecting it into the denoiser along with metadata features, or by directly controlling the activation or weighting of metadata features, enabling the denoiser to perceive which metadata is valid and which is invalid. By using multimodal information as conditional input, the denoiser can gain a more comprehensive understanding of the characteristics of the video content, the key areas or attributes of interest to the user or system, and which metadata is key guiding information for the current super-resolution task. This allows super-resolution processing to be finely tuned according to the specific needs of the content and the indications of the metadata, thereby generating higher-quality second-resolution display content that better meets expectations.
[0405] For example, using the diffusion model as the main implementation scheme of the conditional super-resolution reconstructor, the reconstructor is integrated into the diffusion model, and the low-resolution input y (as video content), m(z) (as a subset of target metadata), and z (as a selection mask) are collectively used as the denoising unit εθ in the diffusion model. The model takes the joint conditional input of the diffusion model and, through an iterative denoising reverse process, gradually recovers the high-resolution video content from the noisy low-resolution content, ultimately generating a high-resolution estimate as the second-resolution display content. As an example, let x0 represent a high-resolution target image in single-image super-resolution (ISR), or a high-resolution frame in video super-resolution (VSR). The diffusion model defines a Markov forward process that progressively adds Gaussian noise to x0 over T steps: The diffusion model defines the forward process as follows: Where α_t∈(0,1) is a predefined noise schedule, and The reverse process learns through a neural network denoiser, which predicts the added noise (or equivalent parameterization), conditioned on low-resolution observations y and scheduling metadata. In MetaSR, the conditional input is a tuple (y, m(z), z). in, ε θ ( Typically, it uses a U-Net structure. The mask z explicitly informs the denoiser of the currently available metadata categories, enabling a single model to robustly handle variable and missing metadata.
[0406] This application generates second-resolution display content through an iterative denoising process. Starting with a random noise sample of the same size as the target second-resolution image, a denoiser is applied progressively across multiple time steps. At each time step t, the denoiser receives the current noisy image Xt, the encoding of time step t (e.g., by sinusoidal positional encoding), and the aforementioned video content, a subset of target metadata, and a selection mask as conditional inputs. The denoiser predicts the noise εt contained in the current noisy image. Then, based on the predicted noise εt and the sampling strategy of the diffusion model (e.g., DDPM, DDIM, etc.), the image is updated to Xt-1 to make it closer to the real image. This iterative process is repeated until a preset minimum time step (typically 0) is reached, and the final image obtained is the second-resolution display content. The iterative denoising process allows the model to progressively refine image details across multiple stages, from coarse structure to fine texture, thereby generating a high-fidelity and realistic super-resolution image. The introduction of conditional input ensures that the entire denoising process is always guided by content features and metadata, resulting in a final output of second-resolution display content that is not only higher in resolution but also significantly improved in terms of visual quality and content consistency.
[0407] This application fully leverages the powerful image generation and restoration capabilities of the diffusion model by integrating a conditional super-resolution reconstructor into the diffusion model and using video content, a subset of target metadata, and a selection mask as conditional inputs to the denoiser. The iterative denoising process of the diffusion model allows it to progressively recover high-resolution details from low-resolution content. Simultaneously, the introduction of multimodal conditional inputs ensures that the reconstruction process is always precisely guided by content features and metadata selection. This results in generated second-resolution display content that not only has higher resolution but also significantly improved visual quality, detail richness, and content consistency, effectively avoiding artifacts or distortions that may occur in traditional super-resolution methods, thus providing users with a clearer, more realistic, and content-consistent viewing experience.
[0408] In some embodiments, a conditional super-resolution reconstructor is integrated into a convolutional neural network or generative adversarial network architecture to input fused enhanced features into the convolutional neural network backbone to output display content at a second resolution.
[0409] As an example, conditional super-resolution reconstructors can be integrated into convolutional neural network (CNN) or generative adversarial network (GAN) architectures. A CNN is a feedforward neural network that automatically learns features from input data through structures such as convolutional layers, pooling layers, and fully connected layers. In super-resolution reconstruction tasks, CNNs can learn the complex mapping relationship from first-resolution video content to second-resolution display content, gradually recovering image details through multi-layer nonlinear transformations. For example, advanced CNN structures such as residual networks, densely connected networks, or attention mechanisms can be used to enhance feature extraction and information transfer capabilities. A generative adversarial network consists of a generator and a discriminator, which optimize each other through adversarial training. The generator aims to generate realistic high-resolution images to deceive the discriminator; the discriminator attempts to distinguish between real high-resolution images and those generated by the generator. This adversarial training mechanism enables GANs to excel in generating super-resolution images with realistic textures and details, effectively mitigating the problem of blurred results easily produced by traditional CNN models trained based on the mean squared error (MSE) loss function.
[0410] The fused enhanced features refer to the comprehensive feature representation obtained by fusing metadata features from a subset of target metadata with features from the video content. The enhanced features contain visual information from the video content itself, as well as semantic and structural information provided by the metadata. The fused enhanced features are input into the backbone of a convolutional neural network (CNN), serving as the primary input to the deep learning model and guiding the network's learning and reconstruction process. The CNN backbone is the part of a CNN or GAN architecture responsible for core feature extraction, transformation, and upsampling. For example, in a pure CNN architecture, it refers to the main feature extraction and upsampling network; while in a GAN architecture, it typically refers to the main body of the generator network. Input methods can be varied; for example, the enhanced features can be concatenated with the first-resolution video content at the network's input layer, or injected into intermediate layers of the network via skip connections to provide conditional information at different levels of abstraction. These input methods ensure that metadata information is integrated throughout the entire super-resolution reconstruction process, enabling fine-grained control and optimization of the reconstruction results.
[0411] The output of the second-resolution display content refers to the high-resolution image or video frame reconstructed from the first-resolution video content after processing by a convolutional neural network or generative adversarial network architecture. The second resolution is significantly larger than the first resolution, representing an improvement in image detail and sharpness. This output is the final result of the super-resolution reconstruction task, and its quality directly reflects the performance of the entire system. Through the learning capabilities of deep learning architectures, the system can generate display content with richer details, fewer artifacts, and better conformity to human visual perception, thereby enhancing the user's viewing experience.
[0412] For example, the MetaSR framework is backbone-independent, and the conditional super-resolution reconstructor can be directly integrated into a convolutional neural network (CNN) or generative adversarial network (GAN) architecture. In this integration scheme, the reconstructor injects the enhanced features (i.e., the fused conditional representation h) obtained through metadata encoding and fusion into the CNN backbone network via input concatenation or feature-level modulation. The CNN backbone network then performs the super-resolution processing and outputs a high-resolution estimate of the content to be displayed at the second resolution. As an example, given a low-resolution input signal y, a context descriptor c, and a scheduler output z=π(c), a CNN-based reconstructor generates a high-resolution output: Similar to the diffusion model, z explicitly indicates the currently available metadata category, enabling a single CNN backbone to unambiguously handle variable or missing metadata inputs. CNN-based super-resolution needs to handle heterogeneous metadata formats. A unified encoding strategy is adopted: spatial metadata (such as edge maps, text / ROI masks, semantic segmentation) is represented as an image-aligned mapping, which can be concatenated with y at the input or encoded into intermediate feature maps. Non-spatial metadata (such as degenerate scalars or compact descriptors) is embedded as feature vectors through a multilayer perceptron (MLP) and injected into the CNN through feature-level modulation. Formally, the metadata encoder E(·) is defined to generate conditional representations: h = E ( m ( z ), z ), where h can contain spatial conditional mapping and global embedding. The conditionalization mechanisms widely used in both standards are compatible with common SR backbone networks (such as EDSR): Input Concatenation, for spatial metadata, constructs enhanced input: This is the simplest integration method, providing strong baseline performance. Feature-level modulation (FiLM / gating mechanism) is used to inject conditional embeddings into the residual blocks through feature-level affine modulation for more flexible fusion of spatial and non-spatial metadata. F′=γ ( h )⊙ F + β (h ), where F is the intermediate feature map, γ ( )and β ( For a small network, h is mapped to channel-level modulation parameters, and ⊙ represents element-wise multiplication. This mechanism allows metadata to act as a "control signal," adaptively adjusting reconstruction behavior to match the content context. The CNN-based implementation uses the standard reconstruction loss, with optional enhancements to ROI sensitivity or structure-aware terms. z = π ( c To improve robustness to changes in metadata selection, a random metadata discarding strategy is employed during training: random sampling or perturbation of z exposes the CNN backbone to diverse combinations of metadata (including partially missing signals). This strategy enhances inference stability, especially when the scheduler selects different subsets of metadata across segments. This CNN implementation highlights the important characteristic of MetaSR: the scheduling layer π(c) is orthogonal to the backbone network selection.
[0413] This application integrates a conditional super-resolution reconstructor into a convolutional neural network or generative adversarial network architecture, using the fused enhanced features as input, effectively leveraging the powerful feature learning and reconstruction capabilities of deep learning. This integration allows the system to capture the complex correlation between video content and adaptive metadata more precisely, thereby generating second-resolution display content with richer details, higher visual quality, and a better match to content features during super-resolution processing. Compared to traditional super-resolution methods based on interpolation or simple filtering, this approach significantly improves the robustness and image quality of super-resolution reconstruction, especially when processing videos with complex dynamic characteristics and diverse content. It better utilizes metadata information to guide the reconstruction process, avoiding poor reconstruction results caused by insufficient metadata utilization or improper integration methods.
[0414] In some embodiments, the conditional super-resolution reconstructor is integrated into the generative adversarial network (GAN) to input a selected subset of target metadata into the generator and discriminator in the GAN, respectively. The generator generates second-resolution display content based on each metadata in the subset of target metadata, and the discriminator determines the authenticity of the display content based on the metadata. Metadata consistency loss is introduced to constrain the output details of the display content to match the metadata.
[0415] As an example, a Generative Adversarial Network (GAN) is a deep learning model consisting of a generator and a discriminator. The generator is responsible for generating data samples from random noise or given conditions, while the discriminator is responsible for distinguishing the generated data from real data. Through adversarial training between the generator and the discriminator, the generator can learn to generate highly realistic data that conforms to specific conditions. Integrating a conditional super-resolution reconstructor into a GAN leverages the powerful generative capabilities of GANs to improve the quality of super-resolution processing, enabling it to generate more realistic and detailed second-resolution display content. This integration makes the super-resolution task not just a simple pixel interpolation, but an intelligent reconstruction and enhancement of content details.
[0416] The selected subset of target metadata, serving as conditional information, is provided to both the generator and the discriminator. For the generator, the metadata provides key attribute information about the video content, such as motion intensity, text proportion, and texture density. The generator can use this conditional information to guide its generation process, ensuring that the generated second-resolution display content maintains consistency in detail with the features described by the metadata. For example, if the metadata indicates that the video content contains a large amount of text, the generator will tend to generate clear and sharp text areas. For the discriminator, the metadata is used not only to judge the overall realism of the generated content but also to evaluate whether the generated content conforms to specific metadata conditions. While judging realism, the discriminator also determines whether the generated content "looks like" real content with metadata characteristics.
[0417] The generator receives video content at a first resolution and a subset of target metadata as input. Using this metadata as a guiding signal, it upscales the low-resolution video content to a second resolution through a series of convolutional layers, upsampling layers, and other network structures. The metadata acts as a "content prior," helping the generator to perform targeted enhancements based on the specific characteristics of the video content (such as texture, edges, and motion) when reconstructing high-frequency details, rather than blindly interpolating or generalizing. For example, for regions containing fine textures, the generator will focus on reconstructing texture details based on metadata instructions; for scenes with rapid motion, it will focus on the smoothness and sharpness of motion trajectories.
[0418] The discriminator receives the second-resolution display content (or real second-resolution content) output by the generator, along with a corresponding subset of target metadata. The discriminator's task is to distinguish whether the input content is real or generated by the generator, and simultaneously evaluate whether the content matches the input metadata conditions. In this way, the discriminator learns not only the statistical realism of the image but also the semantic consistency between the image content and the metadata. For example, if the metadata indicates that the video content has high motion intensity, but the generator's output image appears blurry or lacks motion, the discriminator will identify it as "fake" or "inconsistent."
[0419] Metadata consistency loss is an additional loss function used to quantify the degree of matching between the generator's output second-resolution display content and the input target metadata subset. The loss can be implemented in various ways; for example, features (such as texture features, motion features, text features, etc.) can be extracted from the generated content, and then compared with the original target metadata. If a significant difference exists, the consistency loss increases, prompting the generator to adjust its generation strategy so that the output display content more accurately reflects the characteristics indicated by the metadata in detail. For example, an auxiliary network can be designed to predict the metadata of the generated image, and its predictions can be compared with the true metadata to calculate the loss.
[0420] For example, when a conditional super-resolution reconstructor is integrated into a Generative Adversarial Network (GAN), a selected subset of target metadata, m(z), is simultaneously input into both the generator and discriminator of the GAN as a common conditional input. The generator processes low-resolution video content based on the target metadata subset to generate high-resolution video as the second-resolution display content. The discriminator then judges the realism of the generator's output based on the target metadata subset. To constrain the matching degree between the generated content and the metadata, a metadata consistency loss is introduced during training to force the details of the output content to remain aligned with the metadata information. As an example, when a conditional super-resolution reconstructor is integrated into a GAN, the generator generates perceptibly realistic HR images, and the discriminator encourages realistic textures. The scheduler selects a context-dependent subset of metadata, m(z), using z = π(c), and conditionally conditions the generator (and the discriminator). During training, z is randomly sampled / perturbed (random metadata is discarded), allowing the generator and discriminator to observe diverse combinations of metadata. Combined with explicit masking conditionalization, a single GAN model can adapt to different scheduling decisions during inference.
[0421] This example integrates a conditional super-resolution reconstructor into a generative adversarial network (GAN), and makes both the generator and discriminator use a subset of the target metadata as conditional input. This significantly improves the quality and realism of super-resolution processing. Guided by the metadata, the generator can produce second-resolution display content that highly matches the content features, avoiding detail distortion or artifacts that may occur in traditional super-resolution methods. Simultaneously, the discriminator can not only identify the visual realism of the generated content but also ensure its consistency with the metadata conditions. Furthermore, a metadata consistency loss is introduced, directly forcing the output details of the generated content to precisely match the metadata during training, effectively solving the problem of semantic discrepancies between the super-resolution content and the original metadata. This results in the final output display content not only having higher resolution but also being highly consistent with user expectations or content characteristics in key content features such as texture, motion, and text, greatly improving the viewing experience and the accuracy of the display effect.
[0422] Figure 4 This is a flowchart illustrating the content-adaptive metadata-based display method provided in this application embodiment; as follows: Figure 4 As shown, the display method operates on a display device and includes: S401: Perform feature analysis on the video content to be displayed at the first resolution, extract the content features of the video content, and select a subset of target metadata from multiple preset metadata categories based on the content features; wherein, the content features include attribute information used to describe the dynamic characteristics and content composition of the video content.
[0423] S402: Perform super-resolution processing on the video content based on each metadata category in the target metadata subset, and output the display content at a second resolution, which is greater than the first resolution.
[0424] The core innovation of this embodiment lies in combining content feature analysis with an adaptive metadata selection mechanism. This allows for the dynamic selection of the most suitable subset of metadata based on the video content's dynamic characteristics and composition, avoiding the limitations of fixed metadata configuration schemes when handling diverse video content. This achieves the effect of optimizing super-resolution quality and efficiency under resource budget constraints. As an example, the method first performs feature analysis on the video content to be displayed at a first resolution, extracting its content features. These features include attribute information describing the dynamic characteristics and composition of the video content, such as motion intensity, text proportion, target object proportion, texture density, and low-light noise level. The attribute information is obtained through a lightweight context descriptor. Motion intensity is estimated using block motion statistics, optical flow amplitude, or frame difference energy proxy. Text proportion is estimated using a lightweight text detector or edge density combined with stroke heuristics. Target object proportion is estimated using a compact detector at the first resolution. Texture density is characterized by gradient energy or local variance statistics, and low-light noise level is estimated using brightness distribution and noise proxy.
[0425] Based on the extracted content features, this method selects a subset of target metadata from multiple predefined metadata categories. These categories include at least a portion of degradation cues, structural and edge information, temporal motion cues, semantic masks, regions of interest (ROI) maps, and text regions. Degradation cues include noise levels, artifact scores, and blur proxy information; structural and edge information includes edge maps and gradient maps; temporal motion cues include optical flow, motion vectors, and shot transition markers; semantic masks include coarse-grained region labeling; ROI maps include importance weights for faces, salient regions, and main body regions; and text regions include positional masks for video subtitles and UI overlays. The selection process is implemented through a metadata scheduler, which predicts the expected quality gain of each metadata category under the current content features, calculates a unit cost quality gain index, and selects metadata categories based on the unit cost quality gain index until a predefined resource budget constraint is met. This selection mechanism can be supplemented with domain heuristics, such as selecting text region metadata when the text proportion exceeds a threshold, and selecting temporal motion cue metadata when the motion intensity exceeds a threshold, thereby ensuring the selection of the optimal metadata subset under resource constraints.
[0426] Subsequently, the method performs super-resolution processing on the video content based on each metadata category in the target metadata subset. The conditional super-resolution reconstructor extracts features from the heterogeneous metadata in the target metadata subset using a metadata encoder, generating metadata features in a unified format. The heterogeneous metadata includes dense spatial graphs, temporal signals, compact vectors, or scalars. The feature fusion module fuses the metadata features with the video content features to obtain enhanced features. This reconstructor can be integrated into diffusion models, convolutional neural networks, or generative adversarial networks (GANs): when integrated into a diffusion model, it generates second-resolution display content by using the video content, target metadata subset, and selection mask as conditional inputs to the denoiser through an iterative denoising process; when integrated into a convolutional neural network or GAN, the fused enhanced features are input to the backbone network, outputting second-resolution display content; when integrated into a GAN, the target metadata subset is input to both the generator and discriminator. The generator generates display content based on the metadata, and the discriminator determines authenticity based on the metadata, introducing metadata consistency loss constraints to match output details with the metadata.
[0427] Ultimately, this method outputs display content at a second resolution, which is greater than the first resolution. This achieves adaptive super-resolution processing of video content. This scheme dynamically selects the most suitable subset of metadata based on the dynamic characteristics and content composition of the video content, effectively overcoming the limitations of traditional fixed metadata configuration schemes when processing diverse video content. Therefore, under resource budget constraints, this embodiment can optimize the quality and efficiency of super-resolution processing, improve the clarity and detail of the displayed content, especially significantly enhancing fidelity in key areas such as text and faces, while suppressing artifact generation and enhancing system stability.
[0428] In some embodiments, the content features include content feature identifiers and segment-level attributes; the content feature identifiers include at least one of the domain, subject matter, or channel type information of the video content; the segment-level attributes include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level.
[0429] Here, "domain" refers to the macro-level domain of the video content, such as sports events, news reports, movies, and animations; "genre" refers to the specific theme of the video content, such as science fiction, comedy, documentaries, and education; and "channel type" refers to the channel attribute from which the video content originates, such as movie channels, children's channels, and news channels. This identification information can be pre-obtained through manual annotation or by classifying and identifying the video content using machine learning models, providing the metadata scheduler with overall contextual information about the video content.
[0430] Meanwhile, content features also include segment-level attributes, used to describe the local dynamic characteristics and compositional details of video content in time or space. Segment-level attributes can include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level. Motion intensity measures the intensity of motion of objects or background in the video frame, and can be obtained by calculating inter-frame pixel changes, average amplitude of optical flow field, or motion vector statistics. Text proportion measures the proportion of text areas in the video frame, and can be obtained by text detection algorithms in image processing, optical character recognition (OCR) technology combined with area calculation, etc. Target object proportion measures the proportion of a specific target object (e.g., face, vehicle, specific item) in the video frame, and can be identified by an object detection model and the target bounding box area proportion is calculated. Texture density measures the richness of texture details in the video frame, and can be obtained by calculating image gradient information, local variance, or Gabor filter response, etc. Low-light noise level measures the significance of noise in the video frame under low-light conditions, and can be obtained by analyzing the image's brightness histogram, noise model estimation, or local pixel variance, etc. Fragment-level attributes are typically obtained through real-time or offline analysis and calculation of video frames or video fragments, providing the metadata scheduler with micro-level details of the video content.
[0431] The metadata scheduler in this application achieves a more comprehensive and refined understanding of video content. Macro-level "content feature identifiers" provide the overall context of the video content, while micro-level "fragment-level attributes" reveal image details and dynamic characteristics. This multi-layered feature description allows the metadata scheduler to more accurately determine which metadata categories are most critical and effective for the super-resolution processing of the current video content. For example, for videos with high motion intensity, temporal motion cue metadata can be prioritized to reduce motion blur; for videos with a high proportion of text, text region metadata can be prioritized to improve text clarity; and for videos with high levels of noise in low light, noise suppression metadata can be prioritized to improve image purity. Targeted metadata selection avoids blindly processing all metadata categories, thereby optimizing the allocation of computational resources, improving the efficiency and targeting of super-resolution processing, and enhancing the visual quality and user experience of the final displayed content.
[0432] In some embodiments, feature analysis is performed on the video content to be displayed at a first resolution to extract content features of the video content, including at least one of the following steps: Motion intensity in video content can be estimated using block motion statistics, optical flow amplitude, or frame difference energy. Specifically, motion intensity can be estimated in several ways. For example, block motion statistics can be used to divide video frames into multiple image blocks, calculate the motion vector of each image block between consecutive frames, and quantify the intensity of motion in the overall or local scenes by statistically analyzing the amplitude and direction of the motion vectors. Alternatively, optical flow amplitude can be used to characterize motion intensity. By calculating the optical flow vectors of pixels between video frames, the magnitude of the amplitude directly reflects the movement speed and distance of the pixels, thus aggregating the motion intensity information of the scene. Another approach is to estimate motion intensity using frame difference energy, which involves calculating the absolute or squared difference of pixel values between consecutive frames and summing or averaging these differences to reflect the intensity of changes in the scene; the greater the change, the higher the motion intensity.
[0433] The text proportion of video content can be estimated using a lightweight text detector or a combination of edge density and stroke heuristics. As an example, text proportion estimation can employ a lightweight text detector, typically a small, optimized deep learning model capable of quickly identifying text regions in video frames and calculating their pixel proportion. Another approach combines edge density and stroke heuristics. First, image edges are extracted using edge detection algorithms (such as the Canny operator). Then, leveraging the high edge density and specific stroke structures (such as stroke width consistency and connectivity) typically found in text regions, morphological operations and heuristic rules are used to identify and calculate the proportion of text regions.
[0434] The proportion of target objects in video content is estimated at a first resolution using a compact detector. As an example, the estimation of the proportion of target objects can be achieved using a compact detector, which is a computationally efficient and small-scale object detection model (e.g., based on MobileNet or YOLO-Tiny architecture). This detector can quickly and accurately identify and locate predefined target objects directly in the first-resolution video content without resolution upscaling, and calculate the area proportion of the target objects in the frame.
[0435] Texture density in video content can be characterized using gradient energy or local variance statistics. For example, texture density can be represented by gradient energy, which involves applying gradient operators (such as Sobel or Prewitt operators) to video frames to obtain the gradient magnitude of the image. A larger gradient magnitude indicates richer image detail. By accumulating or averaging the gradient magnitudes, the density of the texture can be quantified. Another method is local variance statistics, which calculates the variance of pixel intensity within local regions of the image. A larger variance generally indicates more complex texture and higher density in that region.
[0436] The level of low-light noise in video content can be estimated using brightness distribution and noise proxy. As an example, the estimation of low-light noise can be achieved by analyzing the brightness distribution of the video content. First, low-light areas in the image are identified, and then noise proxy methods are used to assess the noise level of these areas. For instance, relatively flat image patches can be selected within the low-light areas, and the variance or standard deviation of their pixel values can be calculated as a surrogate indicator of noise. Alternatively, a statistical model-based method can be used to estimate noise parameters, thereby quantifying the noise level in low-light environments.
[0437] This application's embodiments enable the metadata scheduler to obtain more comprehensive and accurate video content attribute information, thus providing a solid data foundation for the intelligent selection of subsequent metadata subsets. This not only improves the accuracy of metadata selection, ensuring a high degree of match between the selected metadata and video content characteristics, but also reduces the computational overhead of feature extraction by employing efficient calculation methods. This guarantees the performance of the entire display system in real-time or near-real-time scenarios, ultimately improving the targeting of super-resolution processing and the overall visual quality of the displayed content.
[0438] In some embodiments, the method further includes: generating a corresponding selection mask based on a target metadata subset, so as to determine each metadata category included in the target metadata subset based on the selection mask, wherein the selection mask is used to identify the metadata categories included in the target metadata subset.
[0439] As an example, after performing feature analysis on the displayed video content at a first resolution and selecting a target metadata subset from multiple preset metadata categories based on the content features, this application further proposes generating a corresponding selection mask based on the target metadata subset. This selection mask can be understood as a data structure, such as a binary vector or bitmask, whose function is to explicitly indicate which metadata categories are specifically included in the target metadata subset. For example, if N metadata categories are preset, the selection mask can be a binary vector of length N, where each position corresponds to a metadata category. When a metadata category is selected and included in the target metadata subset, its corresponding position is set to "1", otherwise it is set to "0". This generation method ensures a clear and quantifiable representation of the selected metadata categories.
[0440] After generating the selection mask, subsequent super-resolution processing modules (such as conditional super-resolution reconstructors) will receive and parse it. By parsing the selection mask, the super-resolution processing module can accurately identify which metadata categories need to be actually utilized in the current processing cycle. For example, the super-resolution processing module can iterate through all possible metadata categories and, according to the indication of the selection mask, only activate and process those metadata categories marked as "1" in the mask. This ensures that the input to super-resolution processing is accurately filtered, avoiding redundant processing of unnecessary metadata.
[0441] The function of a selection mask is to provide a clear identifier indicating the specific metadata categories contained in the target metadata subset. Regardless of the form in which the target metadata subset exists (e.g., a collection of metadata objects), the selection mask provides a standardized, easily parsed interface, enabling the super-resolution processing module to quickly and accurately understand the scope of metadata it should focus on. This identification function is crucial for achieving refined metadata-driven super-resolution processing.
[0442] This application, after selecting a subset of target metadata based on video content features, further generates and utilizes a selection mask to explicitly identify and determine the included metadata categories, thereby providing accurate and efficient guidance for subsequent super-resolution processing. This ensures that the conditional super-resolution reconstructor can accurately identify and utilize only those intelligently scheduled and selected metadata categories, avoiding redundant processing of unnecessary metadata and improving the efficiency and resource utilization of super-resolution processing. Simultaneously, because the super-resolution processing can focus more on the key features of the content, the final output second-resolution display content better reflects the dynamic characteristics and composition of the video content, thus optimizing the quality and display effect of the super-resolution.
[0443] In some embodiments, the method further includes: predicting the expected quality gain of each metadata category under the current content features; calculating the unit cost quality gain index of each metadata category, and selecting metadata categories according to the unit cost quality gain index of each metadata category until a preset resource budget constraint is met, and generating a selection mask.
[0444] As an example, expected quality gain refers to a quantitative assessment of the improvement in display quality that can be achieved by applying a specific metadata category for super-resolution processing, given the current video content features. Prediction can be achieved in several ways. For example, a pre-trained machine learning model can be used, which learns from a large amount of video content, corresponding metadata, and quality assessment data after super-resolution processing (such as objective metrics like PSNR, SSIM, or VMAF, or subjective quality scores) to establish a mapping between content features and quality gain. At runtime, the model receives the content features of the current video content as input and outputs the expected quality gain for each metadata category. Alternatively, a series of heuristic rules or lookup table mechanisms can be constructed based on domain expert experience or experimental dat...
Claims
1. A display device comprising a display system based on content adaptive metadata, characterized in that, The display system includes a metadata scheduler and a conditional super-resolution reconstructor; The metadata scheduler is used to perform feature analysis on the video content to be displayed at a first resolution, extract the content features of the video content, and select a target metadata subset from multiple preset metadata categories based on the content features; wherein the content features include attribute information used to describe the dynamic characteristics and content composition of the video content; The conditional super-resolution reconstructor is used to perform super-resolution processing on the video content based on each metadata category in the target metadata subset, and output display content at a second resolution, wherein the second resolution is greater than the first resolution.
2. The display device of claim 1, wherein, The content features include content feature identifiers and segment-level attributes; the content feature identifiers include at least one of the domain, subject matter, or channel type information of the video content; the segment-level attributes include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level.
3. The display device according to claim 2, characterized in that, The metadata scheduler includes at least one of the following modules: The motion intensity calculation module is used to estimate the motion intensity of the video content through block motion statistics, optical flow amplitude, or frame difference energy. The text detection module is used to estimate the text proportion of the video content using a lightweight text detector or an edge density combined with a stroke heuristic. The target object detection module is used to estimate the proportion of target objects in the video content at a first resolution using a compact detector; A texture density calculation module is used to characterize the texture density of the video content through gradient energy or local variance statistics; The low-light noise detection module is used to estimate the low-light noise level of the video content by means of brightness distribution and noise proxy.
4. The display device according to claim 1, characterized in that, The metadata scheduler further includes a mask generation module, which generates a corresponding selection mask based on the target metadata subset, so that the conditional super-resolution reconstructor determines each metadata category included in the target metadata subset based on the selection mask, wherein the selection mask is used to identify the metadata categories included in the target metadata subset.
5. The display device according to claim 4, characterized in that, The metadata scheduler also includes a gain prediction module and a greedy selection module; The gain prediction module is used to predict the expected quality gain of each of the metadata categories under the current content features; The greedy selection module is used to calculate the unit cost quality gain index of each metadata category, and select the metadata category according to the unit cost quality gain index of each metadata category until the preset resource budget constraint is met, and generate the selection mask.
6. The display device according to claim 5, characterized in that, The metadata scheduler also includes a hybrid rule module, which introduces domain heuristic rules to supplement the selection logic of the greedy selection module; the heuristic rules include selecting text region metadata when the text proportion exceeds a threshold and selecting temporal motion cue metadata when the motion intensity exceeds a threshold.
7. The display device according to claim 1, characterized in that, The conditional super-resolution reconstructor includes a metadata encoder and a feature fusion module; The metadata encoder is used to extract features from heterogeneous metadata in the target metadata subset and generate metadata features in a unified format. The heterogeneous metadata includes dense spatial graphs, time-series signals, compact vectors, or scalars. The feature fusion module is used to fuse the metadata features with the features of the video content to obtain enhanced features, so that the conditional super-resolution reconstructor performs super-resolution processing on the video content based on the enhanced features.
8. The display device according to claim 4, characterized in that, Each of the metadata categories corresponds to an abstract cost function, which is a non-negative function that depends on the content features and is used to characterize the resource overhead of the metadata under the corresponding metadata category; the total cost corresponding to the selection mask is the sum of the costs of the selected metadata categories under the current content features.
9. The display device according to claim 4, characterized in that, The conditional super-resolution reconstructor is integrated into the diffusion model. It generates the second resolution display content by using the video content, the target metadata subset, and the selection mask as conditional inputs to the denoiser through an iterative denoising process.
10. The display device according to claim 4, characterized in that, The conditional super-resolution reconstructor is integrated into a convolutional neural network or generative adversarial network architecture to input the fused enhanced features into the convolutional neural network backbone to output the display content at the second resolution.
11. The display device according to claim 4, characterized in that, The conditional super-resolution reconstructor is integrated into a generative adversarial network (GAN) to input a selected subset of target metadata into the generator and discriminator in the GAN. The generator generates the second resolution display content based on each metadata in the subset of target metadata, and the discriminator determines the authenticity of the display content based on the metadata. Metadata consistency loss is introduced to constrain the output details of the display content to match the metadata.
12. A display method based on content-adaptive metadata, running on a display device, characterized in that, The method includes; Feature analysis is performed on the video content to be displayed at a first resolution to extract the content features of the video content, and a target metadata subset is selected from multiple preset metadata categories based on the content features; wherein, the content features include attribute information used to describe the dynamic characteristics and content composition of the video content; Super-resolution processing is performed on the video content based on each metadata category in the target metadata subset, and the display content at a second resolution is output, wherein the second resolution is greater than the first resolution.
13. The display method according to claim 12, characterized in that, The content features include content feature identifiers and segment-level attributes; the content feature identifiers include at least one of the domain, subject matter, or channel type information of the video content; the segment-level attributes include at least one of motion intensity, text proportion, target object proportion, texture density, and low-light noise level.
14. The display method according to claim 13, characterized in that, The video content to be displayed at a first resolution is subjected to feature analysis to extract its content features. Includes at least one of the following steps: The motion intensity of the video content is estimated by block motion statistics, optical flow amplitude, or frame difference energy. The text percentage of the video content is estimated using a lightweight text detector or an edge density-stroke heuristic. The proportion of the target object in the video content is estimated at a first resolution using a compact detector; The texture density of the video content is characterized by gradient energy or local variance statistics; The low-light noise level of the video content is estimated by using brightness distribution and noise proxy.
15. The display method according to claim 12, characterized in that, The method further includes: Based on the target metadata subset, a corresponding selection mask is generated to determine each of the metadata categories included in the target metadata subset. The selection mask is used to identify the metadata categories included in the target metadata subset.
16. The display method according to claim 15, characterized in that, The method further includes: Predict the expected quality gain for each of the metadata categories under the current content features; Calculate the unit cost quality gain index for each of the metadata categories, and select metadata categories based on the unit cost quality gain index of each metadata category until a preset resource budget constraint is met, and generate the selection mask.
17. The display method according to claim 12, characterized in that, The step of performing super-resolution processing on the video content based on each metadata category in the target metadata subset, and outputting display content at a second resolution, includes: Feature extraction is performed on the heterogeneous metadata in the target metadata subset to generate metadata features in a unified format. The heterogeneous metadata includes dense spatial graphs, time-series signals, compact vectors, or scalars. The metadata features are fused with the features of the video content to obtain enhanced features, and super-resolution processing is performed on the video content based on the enhanced features.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program configured to be executed by a processor to implement the content-adaptive metadata-based display method according to any one of claims 12 to 17.