A video processing system, a sending end and a receiving end

By having the sending and receiving ends of the video processing system work together to analyze video content, generate metadata, and combine it with local context information to dynamically adjust display parameters, the system solves the problem of insufficient video presentation caused by differences in environment, user, and device in traditional video processing methods, and achieves personalized video display effects.

CN122138003APending Publication Date: 2026-06-02TCL CHINA STAR OPTOELECTRONICS TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TCL CHINA STAR OPTOELECTRONICS TECHNOLOGY CO LTD
Filing Date
2026-02-11
Publication Date
2026-06-02

Smart Images

  • Figure CN122138003A_ABST
    Figure CN122138003A_ABST
Patent Text Reader

Abstract

This application discloses a video processing system, a transmitter, and a receiver, belonging to the field of display technology. The system includes a transmitter and a receiver connected by communication. The transmitter analyzes video content to obtain metadata of the video content and sends the metadata to the receiver. The metadata describes at least one dimension of the visual features of the video content. The receiver receives the metadata and adjusts display parameters for displaying the video content based on the metadata and its local context information. The local context information of the receiver indicates at least one of the following: environmental characteristics of the receiver, user characteristics of the viewing user, and device operating characteristics of the receiver. This enables personalized video effects that match different user characteristics, environmental characteristics, and device operating characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display technology, and in particular to a video processing system, a transmitter, and a receiver. Background Technology

[0002] With the increasing prevalence of high-resolution, high dynamic range display devices, users' demands for video viewing experience are rising. Traditional video processing and display chains typically follow a linear model of "acquisition-encoding-transmission-decoding-display," where content creation and display are independent of each other. After receiving a video signal, the display device usually relies on its built-in, general-purpose post-processing algorithms (such as contrast enhancement, color optimization, and dynamic backlight control) to optimize the image.

[0003] However, this video processing method makes it difficult to guarantee the quality of the video presentation. Summary of the Invention

[0004] A video processing system, a transmitter, and a receiver are provided to solve the aforementioned technical problems.

[0005] In a first aspect, a video processing system is provided, comprising a transmitting end and a receiving end connected by communication, wherein:

[0006] The sending end is used to analyze the video content to obtain the metadata of the video content, and send the metadata to the receiving end; the metadata is used to describe the visual features of the video content in at least one dimension. The receiving end is used to receive the metadata and adjust the display parameters for displaying the video content based on the metadata and the local context information of the receiving end; the local context information of the receiving end is used to indicate at least one of the following: the environmental characteristics of the receiving end, the user characteristics of the viewing user of the receiving end, and the device operation characteristics of the receiving end.

[0007] In one possible design, the metadata includes one or more of the following: edge features of the video content, local contrast of the video content, RGB channel statistics of low grayscale areas in the video content, lighting data of the video content, depth information of displayed objects in the video content, semantic segmentation data of the video content, source color configuration data of the video content, causal relationships corresponding to visual elements in the video content, motion estimation data of the video content, key visual regions in the video content, creative intent of the video content, saliency score of the video content, sentiment index of the video content, and optimization intent for the video content; the saliency score is used to describe the content weight index of different regions in the video content; the sentiment index is used to describe the emotional tone of the video content.

[0008] In one possible design, the local context information includes at least one of environmental information, the viewing user's preference information, and the receiving end's device information; the device information includes: the status of the receiving end's display panel and hardware capabilities; the preference information includes one or more of the following: the viewing user's focus, viewing angle, viewing distance, and the viewing user's relevant interaction data or viewing emotions.

[0009] In one possible design, the display parameters include one or more of the following data: resolution, contrast ratio, brightness, color temperature, color gamut, saturation, or refresh rate.

[0010] In one possible design, the receiver is specifically used for: Based on the metadata and the local context information of the receiving end, a first target area is determined in the display area of ​​the display panel of the receiving end, and the display parameters of the receiving end for the first target area are adjusted.

[0011] In one possible design, the transmitting end includes a depth and geometry module for generating depth information of objects displayed in the video content; The receiving end includes a stereo conversion module, used to render the video content into a stereoscopic image based on the depth information, wherein the perspective of the stereoscopic image is determined based on the position or posture information of the viewing user.

[0012] In one possible design, the receiver further includes: The color and gamut module is used to determine a tone mapping strategy based on the metadata and the local context information of the receiving end. The tone mapping strategy is used to compress the brightness range of high dynamic range (HDR) video content to the renderable range of the receiving end.

[0013] In one possible design, the receiver further includes: The language module is used to adjust the rendering effect of the subtitles based on the metadata and the local context information of the receiving end.

[0014] In one possible design, the transmitting end further includes a first noise suppression module for generating RGB channel statistics for low grayscale regions in the video content; The receiving end also includes a second noise suppression module, which is used to perform noise suppression on the low grayscale region based on the RGB channel statistical information of the low grayscale region.

[0015] In one possible design, the transmitting end further includes a lighting analysis module for analyzing the lighting data of the video content; the lighting data includes one or more of the following: shadow direction, lighting direction, lighting area, or flare intensity; The receiving end also includes a light adjustment module, which is used to determine a second target area in the display area of ​​the display panel of the receiving end based on the light data, and adjust the display parameters of the receiving end for the second target area.

[0016] In one possible design, the transmitting end further includes an edge detection module for generating edge features and local contrast in the video content; The receiving end also includes a super-resolution module, used to perform upsampling based on the edge features and the local contrast, the upsampling being used to adjust the resolution of the video content.

[0017] In one possible design, the receiver further includes an orchestration module for instructing the receiver to adjust its display parameters based on the metadata and the receiver's local context information.

[0018] In one possible design, the orchestration module is specifically used for: Analyze the optimization intent for the video content; Based on the optimization intent and the local context information of the receiving end, determine the modules to be used and the adjustment strategy for the display parameters; The module to be used is instructed to adjust the display parameters of the receiving end according to the adjustment strategy.

[0019] In one possible design, the transmitter further includes: The semantic analysis module is used to perform semantic segmentation on the video content and generate the semantic segmentation data; The causal reasoning module is used to generate the optimization intent based on a predefined causal graph, which describes the causal relationships corresponding to visual elements in the video content.

[0020] In one possible design, the display parameters are determined based on the weights of each module; the receiving end further includes a verification module, used for: The rendering effect after adjusting the display parameters is checked to see if it meets the hardware capabilities of the receiving end and the user's perception. If the rendering effect after adjusting the display parameters does not meet the hardware capabilities of the receiving end, or if the rendering effect after adjusting the display parameters does not meet the user's perception, then the display parameters will be adjusted to standby parameters, or the weights of each module will be adjusted.

[0021] Secondly, a receiving end is provided, which is communicatively connected to a sending end in a video processing system. The receiving end is used to receive the metadata and adjust the display parameters for displaying the video content based on the metadata and the local context information of the receiving end. The local context information of the receiving end is used to indicate at least one of the following: the environmental characteristics of the receiving end, the user characteristics of the viewing user of the receiving end, and the device operation characteristics of the receiving end.

[0022] In one possible design, the local context information includes at least one of environmental information, the viewing user's preference information, and the receiving end's device information; the device information includes: the status of the receiving end's display panel and hardware capabilities; the preference information includes one or more of the following: the viewing user's focus, viewing angle, viewing distance, and the viewing user's relevant interaction data or viewing emotions.

[0023] In one possible design, the receiver is specifically used for: Based on the metadata and the local context information of the receiving end, a first target area is determined in the display area of ​​the display panel of the receiving end, and the display parameters of the receiving end for the first target area are adjusted.

[0024] In one possible design, the receiving end includes a stereo conversion module for rendering the video content into a stereoscopic image based on the depth information, wherein the perspective of the stereoscopic image is determined based on the position or posture information of the viewing user.

[0025] In one possible design, the receiver further includes: The color and gamut module is used to determine a tone mapping strategy based on the metadata and the local context information of the receiving end. The tone mapping strategy is used to compress the brightness range of high dynamic range (HDR) video content to the renderable range of the receiving end.

[0026] In one possible design, the receiver further includes: The language module is used to adjust the rendering effect of the subtitles based on the metadata and the local context information of the receiving end.

[0027] In one possible design, the receiver further includes a second noise suppression module for performing noise suppression on the low grayscale region based on the RGB channel statistics of the low grayscale region.

[0028] In one possible design, the receiver further includes a light adjustment module for determining a second target area in the display area of ​​the receiver's display panel based on the light data, and adjusting the display parameters of the receiver for the second target area.

[0029] In one possible design, the receiver further includes a super-resolution module for performing upsampling based on the edge features and the local contrast, the upsampling being used to adjust the resolution of the video content.

[0030] In one possible design, the receiver further includes an orchestration module for instructing the receiver to adjust its display parameters based on the metadata and the receiver's local context information.

[0031] In one possible design, the orchestration module is specifically used for: Analyze the optimization intent for the video content; Based on the optimization intent and the local context information of the receiving end, determine the modules to be used and the adjustment strategy for the display parameters; The module to be used is instructed to adjust the display parameters of the receiving end according to the adjustment strategy.

[0032] In one possible design, the display parameters are determined based on the weights of each module; the receiving end further includes a verification module, used for: The rendering effect after adjusting the display parameters is checked to see if it meets the hardware capabilities of the receiving end and the user's perception. If the rendering effect after adjusting the display parameters does not meet the hardware capabilities of the receiving end, or if the rendering effect after adjusting the display parameters does not meet the user's perception, then the display parameters will be adjusted to standby parameters, or the weights of each module will be adjusted.

[0033] Thirdly, a transmitting end is provided, which is communicatively connected to a receiving end in a video processing system. The transmitting end is used to analyze video content to obtain metadata of the video content and send the metadata to the receiving end. The metadata is used to describe at least one dimension of visual features of the video content. The metadata includes one or more of the following: edge features of the video content, local contrast of the video content, RGB channel statistics of low grayscale areas in the video content, lighting data of the video content, depth information of displayed objects in the video content, semantic segmentation data of the video content, source color configuration data of the video content, causal relationships corresponding to visual elements in the video content, motion estimation data of the video content, key visual regions in the video content, creative intent of the video content, saliency score of the video content, sentiment index of the video content, and optimization intent for the video content; the saliency score is used to describe the content weight index of different regions in the video content; the sentiment index is used to describe the emotional tone of the video content.

[0034] In one possible design, the transmitting end further includes a first noise suppression module for generating RGB channel statistics for low grayscale regions in the video content; The transmitting end also includes a lighting analysis module for analyzing the lighting data of the video content; the lighting data includes one or more of the following: shadow direction, lighting direction, lighting area, or flare intensity; The transmitting end also includes an edge detection module, used to generate edge features and local contrast in the video content; The semantic analysis module is used to perform semantic segmentation on the video content and generate the semantic segmentation data; The causal reasoning module is used to generate the optimization intent based on a predefined causal graph, wherein the causal graph is used to describe the causal relationships corresponding to visual elements in the video content. The transmitting end includes a depth and geometry module, used to generate depth information of objects displayed in the video content.

[0035] In this application, the display parameters are determined based on the analysis and understanding of the video content, the perception of the viewer and the environment, and the state of the receiving end itself. It is evident that the receiving end's display parameter adjustment decisions do not solely rely on the universal metadata provided by the sending end, but rather deeply integrate it with the local context. This allows the same video content to be rendered in the most suitable way for the current situation in different environments (such as living rooms and bedrooms), for different users (such as visually sensitive individuals and ordinary users), and on different devices (such as high-end OLED TVs and ordinary LCD monitors), achieving a transformation from "one-size-fits-all" to "personalized experiences for each individual and each time." Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.

[0038] Figure 1 An example diagram of a system architecture provided for an exemplary implementation of this disclosure.

[0039] Figure 2 This is a flowchart illustrating a video processing method provided in an exemplary embodiment of the present disclosure.

[0040] Figure 3 This is a structural example diagram of the sending end and receiving end provided in an exemplary embodiment of this disclosure.

[0041] Figure 4 This is a structural example diagram of the sending end and receiving end provided in an exemplary embodiment of this disclosure.

[0042] Figures 5A-5B This is a structural example diagram of the sending end and receiving end provided in an exemplary embodiment of this disclosure.

[0043] Figures 6-7 This is a structural example diagram of the sending end and receiving end provided in an exemplary embodiment of this disclosure.

[0044] Figures 8-10 This is a structural example diagram of a device provided as an exemplary embodiment of the present disclosure. Detailed Implementation

[0045] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0046] In the embodiments of this application, "at least one" refers to one or more; "multiple" refers to two or more. In the description of this application, the terms "first," "second," "third," etc., are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order.

[0047] References such as “one embodiment” or “some embodiments” as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the terms “comprising,” “including,” “having,” and variations thereof, as used in this specification, mean “including, but not limited to,” unless otherwise specifically emphasized.

[0048] It should be noted that in the embodiments of this application, "and / or" describes the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. In addition, the character " / ", unless otherwise specified, generally indicates that the associated objects before and after it are in an "or" relationship.

[0049] It should be noted that in the embodiments of this application, "connection" can be understood as electrical connection. The connection between two electrical components can be a direct or indirect connection between the two electrical components. For example, the connection between A and B can be a direct connection between A and B, or an indirect connection between A and B through one or more other electrical components.

[0050] This application provides a video processing method applied to a video processing system. This video processing system can also be referred to as a display system.

[0051] For example, Figure 1 An exemplary architecture of a video processing system is illustrated. The system may include a transmitting end and a receiving end connected by communication. The transmitting end includes a first processor and an encoder. The first processor is used to extract metadata of the video content. As one possible implementation, the first processor may run one or more agents (or modules) for extracting metadata. The encoder is used to encode the video content to form a video stream. The transmitting end transmits the video stream to the receiving end for display. The transmitting end and the receiving end can communicate via a transmission network for data interaction. This application embodiment does not limit the specific implementation of the transmission network.

[0052] Metadata, also known as perceptual metadata.

[0053] For example, metadata must meet the following conditions to improve its broad interoperability: (1) Lightweight: For example, semantic masks are compressed using run-length encoding (RLE) and optimized as bit fields.

[0054] (2) Frame addressable: Allows for granular optimization of applications by frame or scene segmentation.

[0055] (3) Encoder compatibility: Supports insertion into SEI, or transmission via MPEG-V or CTA-861.3 over HDMI.

[0056] (4) Scalable: Allows the addition of additional agent signals (such as gaze, attention map).

[0057] For example, the field structure of metadata is as follows: { "frame": 2345, "semantic": {"face": [[32,45,80,120]], "sky": [[0,0,1920,500]]}, "depthMap": "base64-encoded", "intentFlags": ["preserve_face_tone", "reduce_eyestrain"], "sceneLighting": {"CCT": 5100, "intensity": 0.75}, "motionLevel": 0.62, "viewerPosition": {"x": 0.3, "y": 0.1} } The data in the metadata can also take other forms without restriction. For example, the form of sceneLighting can also be scene Lighting.

[0058] As one possible implementation, the sending end performs the above metadata extraction steps offline.

[0059] For example, the agent at the sending end can be a computationally intensive agent responsible for high-complexity analysis.

[0060] For example, one or more of the aforementioned agents can be integrated into a post-processing pipeline or into a custom batch processing workflow. This integration allows content creators to embed optimized metadata early in the content creation cycle to ensure that the recipient has rich information to enhance the viewing experience.

[0061] As one possible implementation, the sender can package metadata along with a DASH or HLS manifest. This allows for the use of the sender's metadata (such as optimization intent) to form an end-to-end system that maximizes visual quality and efficiency.

[0062] The receiving end includes a second processor and a decoder. The decoder decodes the video stream of the video content. The second processor parses the metadata of the video content and, based on this metadata and the receiving end's local context information, adjusts the display parameters of the receiving end to better present the video content. For example, the local context information includes panel conditions and the context of the viewing user. A detailed explanation of the local context information will be given below.

[0063] As one possible implementation, the second processor can run one or more agents to dynamically adjust display parameters. Exemplarily, these agents can run on a display system-on-chip (SoC) or timing controller (TCON). This tight hardware-level integration ensures low latency and efficient command processing, enabling fine-grained control over display parameters such as local dimming and color calibration.

[0064] For example, the receiving agent can be a lightweight agent with real-time processing capabilities, responsible for real-time execution. The receiving agent can respond to incoming metadata in real time and, by combining metadata and local context information, perform scene-specific optimizations, such as adaptive backlight control, refresh rate adjustment, color gamut remapping, 2D-to-3D enhancement (or 2D-to-3D conversion), and personalized eye comfort adjustments. This receiving end can predictably perform low-latency and low-power intelligent rendering; in some cases, even with limited computing resources, intelligent rendering can be performed without increasing system complexity.

[0065] The receiving end can also be called the playback end or the display end.

[0066] One possible implementation is to integrate the intelligent agent as a middleware service into the receiving end's operating system. Alternatively, integration can be performed in the Hardware Abstraction Layer (HAL) to ensure broad applicability at the receiving end.

[0067] As one possible implementation, the aforementioned agents can be distributed across one or more layers to ensure seamless optimization and enhancement of visual performance. For example, at the media player or decoder layer, the software could be modified to parse key metadata (such as SEI messages, OBU, or manifest information) and expose it to the receiving agent. In this way, the receiving end can dynamically adjust display parameters based on this metadata and local context information.

[0068] As one possible implementation, multiple agents can share an access data buffer. Agents can read data from the data buffer, which contains metadata and local context information. Agents can also write data to the data buffer.

[0069] As one possible implementation, the agent can also respond to prompts from upstream agents and trigger optimizations from downstream agents. For example, in response to metadata provided by an upstream agent: if the scene contains faces, the agent can preserve the skin color of the face region to ensure image fidelity.

[0070] The system described above separates high-complexity analysis from real-time execution through modular and distributed processing pipeline operations. Table 1 shows the task characteristics executed by the sending and receiving ends, respectively.

[0071] Table 1

[0072] In this way, a single complex visual intelligence processing operation performed in the sending pipeline can be widely applied to multiple receiving ends (playback devices) to generate metadata. The receiving ends do not need to re-encode or retrain, reducing their overhead. Furthermore, the receiving end can combine this metadata with its own local context information to adaptively adjust the display parameters of the video content, ensuring the video content matches the receiving end's local context and providing a personalized video playback experience. Moreover, the task division between the sending and receiving ends balances computational efficiency and real-time adaptability, preventing dynamic rendering from overloading the display hardware.

[0073] The aforementioned hybrid pipeline system framework or architecture, known as DP Vision, is used for content-aware display optimization. DP Vision achieves intelligent rendering through collaboration among multiple agents. For example, DP Vision can be deployed in various devices, such as, but not limited to, smart TVs, mobile phones, tablets, AR / VR headsets, in-vehicle displays, and streaming media platforms.

[0074] In the embodiments of this application, the intelligent agent may also be referred to as a module, intelligent module, model, agent, or other names, without limitation, and will be uniformly described here.

[0075] Please see Figure 2 , Figure 2 A flowchart example of a video processing method provided in this application embodiment. The method may specifically include: S101. The sending end analyzes the video content to obtain the video content's metadata.

[0076] Metadata is used to describe at least one dimension of the visual characteristics of the video content. The video content may include the content of one or more video frames.

[0077] For example, the sending end is a device provided by the content provider, such as a content production device, edge server, or encoder. As one possible implementation, the sending end can be responsible for performing computationally intensive (or data-intensive) video content analysis tasks to obtain the visual features of the video content in various dimensions.

[0078] As one possible implementation, the sending end can perform the aforementioned video content analysis task through one or more modules (or agents). For example, refer to... Figure 3 The one or more modules may include a semantic analysis module (or semantic segmentation module or scene interpretation module), a depth and geometry module (or depth estimation module), an edge detection module, a causal reasoning module (or causal inference module), a first noise suppression module, and an illumination analysis module.

[0079] For example, except Figure 3 The modules listed in the examples may also include a motion profiling module for motion estimation. Other modules may also be included for analyzing the video content to obtain corresponding metadata.

[0080] The specific actions performed by these modules will be described below.

[0081] For example, the sending end can analyze the video content based on its global visibility and complete temporal context to obtain metadata. The temporal context represents the temporal relationship between multiple video frames contained within the video content. For instance, the sending end analyzes 60 seconds of video frames and predicts a smooth transition in color tone from warm to cool. It then generates "lighting model" metadata: the sending end can create a temporally continuous, smoothly varying lighting parameter curve (e.g., a color temperature smoothly changing from 4500K to 6500K) and package this curve into the metadata. An optimization intent is then generated. After receiving this metadata, the receiving end renders the video based on the lighting parameter curve to present a video with a smooth color tone transition.

[0082] S102. The sending end sends metadata to the receiving end.

[0083] Correspondingly, the receiving end receives metadata.

[0084] As one possible implementation, the sender can encapsulate the metadata in supplemental enhancement information (SEI) and send the SEI to the receiver. That is, the metadata is carried within the SEI. Alternatively, the metadata can be encapsulated in an open bitstream unit (OBU) and sent to the receiver. Or, the metadata can be carried in a bypass manifest, such as sidecarJSON or XML manifests. Alternatively, the sender can transmit the metadata to the receiver via a metadata API.

[0085] As one possible implementation, the transmitter can transmit metadata via a custom HDMI or DisplayPort channel.

[0086] In this way, the sending end can efficiently and with low latency send metadata to the receiving end through the appropriate transmission path, ensuring compatibility with various device types. For example, it can support the transmission of metadata to receiving ends such as TVs, tablets, in-vehicle displays, and AR / VR displays.

[0087] Furthermore, metadata can be directly embedded into the encoded bitstream or provided as a bypass manifest to ensure broad compatibility and low transmission overhead.

[0088] S103. The receiving end adjusts the display parameters used to display video content based on metadata and the receiving end's local context information.

[0089] The local context information of the receiving end is used to indicate at least one of the following: the environmental characteristics of the receiving end, the user characteristics of the viewing user of the receiving end, and the device operation characteristics of the receiving end.

[0090] As one possible implementation, the receiver interprets the metadata in real time and combines it with local context information to perform scene-aware optimization to adjust the display parameters of the video content.

[0091] For example, if a moving object is detected in the video content, the refresh rate is adjusted. Another example is adjusting the color temperature of the video content based on the ambient lighting conditions at the receiving end. This ensures consistent rendering of video frames and intelligently optimizes for the user characteristics, device characteristics, and environmental characteristics of the current viewer at the receiving end.

[0092] As one possible implementation, the receiving end can perform the task of adjusting the display parameters described above through one or more modules. For example, refer to... Figure 3These modules may include a super-resolution module, a stereo conversion module, a color and color gamut module, a language module, a second noise suppression module, a lighting adjustment module, an orchestration module, and a verification module. The specific actions performed by these modules will be described below.

[0093] In this way, each module is responsible for a specific rendering task, such as semantic segmentation, tone mapping, refresh rate control, or color gamut adjustment. This not only ensures high-efficiency parallelism but also supports the flexible expansion of new perception modules, achieving a plug-and-play integration effect. This modular and scalable approach ensures consistent performance across receivers, content types, and user environments.

[0094] This article primarily uses the example of a dedicated module performing a specific action. In other embodiments, the functions of some of the modules involved in the embodiments of this application can also be integrated into a single module. This application does not impose any limitations on this.

[0095] The method provided in this application determines display parameters based on the analysis and understanding of video content, the perception of the viewer and the environment, and the state of the receiving end itself. Therefore, the receiving end's display parameter adjustment decisions do not solely rely on the universal metadata provided by the sending end, but are deeply integrated with the local context. This allows the same video content to be rendered in the most suitable way for the current situation in different environments (such as living rooms and bedrooms), for different users (such as visually sensitive individuals and ordinary users), and on different devices (such as high-end OLED TVs and ordinary LCD monitors), achieving a transformation from "one-size-fits-all" to "personalized experiences for each user and each time."

[0096] Furthermore, it decouples content analysis (computation-intensive, which can be performed offline or in the cloud) from real-time rendering (latency-sensitive, performed locally). The sending end can leverage powerful computing resources to perform complex, multi-dimensional video analysis and generate high-quality metadata without considering the heterogeneity of the receiving end hardware. This solves the computational load, power consumption, and heat generation problems associated with performing deep analysis and high-performance rendering simultaneously on a single device.

[0097] In some embodiments, the metadata includes one or more of the following data: edge features of the video content, local contrast of the video content, RGB channel statistics of low grayscale areas in the video content, illumination data of the video content, depth information of displayed objects in the video content, semantic segmentation data of the video content, source color configuration data of the video content, causal relationships corresponding to visual elements in the video content, motion estimation data of the video content, key visual regions (or important visual regions) in the video content, creative intent of the video content, saliency score of the video content, sentiment index of the video content, and optimization intent for the video content; the saliency score is used to describe the content weight index of different regions in the video content; the sentiment index is used to describe the emotional tone of the video content.

[0098] The metadata includes one or more of the following data, which may be explicitly included or implicitly included. Implicitly including X means that the metadata implicitly indicates X.

[0099] The aforementioned low grayscale region can refer to an area with brightness below a threshold. This low grayscale region can also be called a dark area. In this low grayscale region, the signal strength is low, making it susceptible to noise interference and affecting image quality. The method provided in this application, considering that this low grayscale region is more susceptible to noise, specifically uses the RGB channel statistics of the low grayscale region to suppress noise, thereby enhancing the image quality of the low grayscale region. For example, it can dynamically suppress artifacts while maintaining the integrity of shadow details and color fidelity. Furthermore, it ensures a clean and stable dark background without introducing compressed textures or losing subtle gradients.

[0100] The semantic segmentation described above can be used to obtain semantic masks and object labels for each object in a video frame, such as semantic masks and object labels for faces, text, and the sky. In this way, the receiving end can understand the scene classification in the image and identify the region of interest. For example, it can identify where the face is and where the sky is, thus enabling it to perform fidelity processing on the face and color enhancement on the sky.

[0101] The content weight metric for a specific region within video content indicates the degree to which that region attracts the viewer's attention. Regions of interest have higher weight metrices. One possible implementation is to use an attention heatmap to represent the content weight metric. This attention heatmap can be a grayscale or pseudo-color image of the same size as the video frame, where the brightness or value of each pixel represents the probability that the region attracts the viewer's attention. The higher the probability that a region attracts the viewer's attention, the higher its weight metric.

[0102] Visual elements in video content can also be called perceptual variables or causal variables. The causal relationship described above can also be called a causal dependency relationship. This can be represented by a causal graph, let G = (V, E), where V is the set of visual elements and E represents the causal relationship between the causal variable and the outcome variable. Causal variables include brightness, motion, and viewer comfort. Outcome variables include, for example, the degree of eye fatigue. A causal relationship could be: brightness → eye fatigue. This causal relationship indicates that brightness can affect eye fatigue, and there is a causal relationship between brightness and eye fatigue.

[0103] For example, the joint distribution of the causal variable and the outcome variable satisfies the following formula (1): (1) in, Representative variable The set of parent nodes, that is, those directly pointed to in the cause-effect graph. All nodes are The direct cause. In other words, It is a causal variable. It is the outcome variable.

[0104] Indicates that, given all direct causes In the case of variables The conditional probability of taking a certain state.

[0105] Given observed metadata O V, and action space A, the module (or agent) selects the optimal action a in the following way. : (2) U(X) is a utility function that encodes perceived targets such as contrast and sharpness, color fidelity, eye comfort, and energy efficiency.

[0106] This represents a causal interference quantifier, which indicates taking an action, such as increasing the local backlight, and observing the corresponding results to evaluate the causal effect of the action itself.

[0107] E[ ∣do(a),O] represents the expected value of the system utility U(X) given the known observational evidence O and the implementation of action a.

[0108] Argmax represents maximizing the parameter solution. Specifically, it iterates or searches the action space A to find the action a that maximizes the expected utility E[U(X)|do(a),O], and uses it as the final decision a. .

[0109] The above method can determine the adjustment strategy of display parameters based on the expected results (causal effect) that the action will lead to, and realize the transformation from "experience-driven, black-box decision-making" to "causal-driven, explainable decision-making", thereby improving the display performance of the receiving end.

[0110] For example, the depth information mentioned above includes depth maps, 3D spatial cues, and stereo disparity data. Stereo disparity data is, for example, a disparity vector. 3D spatial cues represent information or features that can be inferred from 2D images or videos to characterize the three-dimensional spatial relationships, shapes, layouts, and depths of objects in a scene.

[0111] This depth information can provide a basis for generating stereoscopic images and adjusting parallax, so as to provide users with highly immersive video content, enhance spatial perception, and provide a richer visual experience.

[0112] For example, motion estimation data includes motion vectors and motion patterns.

[0113] In some embodiments, the local context information includes at least one of environmental information, viewing user preference information, and receiving device information; the device information includes: the status of the receiving device's display panel and hardware capabilities; the preference information includes one or more of the following: the viewing user's focus, viewing angle, viewing distance, relevant interaction data of the viewing user watching the video, or viewing mood.

[0114] The status of the display panel, such as, but not limited to, the brightness, temperature, and power consumption of the display panel.

[0115] Hardware capabilities can be understood or replaced as: system capability constraints (system constraints), or hardware capability constraints. For example, hardware capabilities include panel capabilities or power supply capabilities, such as power supply limits, receiver thermal limitations, and the display range of the panel. The method provided in this application embodiment limits the state of the display panel within the range of hardware capabilities to extend the lifespan of the display panel. For example, the temperature of the display panel is controlled within the thermal limit range, and the power consumption of the display panel is controlled within the power supply limit range.

[0116] Environmental information includes, but is not limited to, lighting conditions that affect visual viewing, ambient noise levels, electromagnetic interference levels, geographical location and time, and environmental content.

[0117] For example, while ambient noise levels do not directly affect visual perception, the receiver can adjust its subtitle rendering strategy when high noise levels are detected, automatically increasing subtitle size and contrast, or enabling speech-to-text assistance. This can meet users' video viewing needs under high-noise conditions.

[0118] For example, the level of electromagnetic interference may affect the stability of the screen signal, and the receiver can adjust the refresh synchronization strategy or enable signal anti-interference processing accordingly.

[0119] For example, the receiving end can switch the corresponding rendering mode based on geographical location and time, such as a nighttime viewing mode or a car interior daylight mode.

[0120] For example, the receiving end can use a front-facing camera or sensors to identify the content of the user's environment. When a user is conducting a video conference in a meeting room, priority can be given to ensuring shared visibility and reducing glare. When a user is in a bedroom, low blue light and flicker-free dimming can be automatically enabled.

[0121] User preference information indicates a user's liking or state when watching video content. This preference information can be user-defined or inferred by the receiver based on the user's behavior. Examples of user preference information include a preference for comfort mode, a preference for 3D mode, high sensitivity to eye fatigue, and high sensitivity to flicker.

[0122] For example, the rendering effect on the receiving end can be different when the user is at different viewing angles, in order to improve the viewing experience of the video at the corresponding viewing angles.

[0123] For example, the rendering effect on the receiving end can be different when the user is at different viewing distances, in order to improve the viewing experience of the video at the corresponding viewing distance.

[0124] For example, when watching HDR videos, users tend to lower the brightness, and the receiver can determine the user's sensitivity to peak brightness. Subsequently, when the user watches an HDR video again, the receiver can lower the brightness. Similarly, users may prefer vibrant colors in animated content. Later, when the user watches similar videos, the receiver can adjust the video's color gamut to meet the user's visual needs. In this method, the receiver can learn the user's visual preferences based on user interaction data. Over time, the learned visual preferences become increasingly accurate, helping to achieve personalized video playback effects.

[0125] For example, the receiving end can analyze the user's viewing emotions in real time based on facial expressions, tone of voice, or behavioral patterns, and adjust the display parameters of the video content to adapt to the viewer's emotional state. For instance, in a suspenseful scene, it can be rendered with higher contrast and a cooler color tone. In a relaxed mood, it can be adjusted to a soft, warm color tone. Another example is adjusting emotional lighting based on the user's facial expressions, tone of voice, or behavioral patterns. Emotional lighting can be a visual attribute of the screen itself, such as dynamically fine-tuning the overall color tendency (color grading) of the image. Alternatively, emotional lighting can be the emotional lighting of the external environment, such as using smart light strips or the room's main light near the receiving end to change the color and brightness of the surrounding environment based on the identified viewer emotions and the content of the image.

[0126] For example, if a user is colorblind or has low vision, the receiver can adaptively enhance colors based on the user's characteristics to provide a visual experience that meets the user's viewing needs.

[0127] The method provided in this application allows the receiving end to combine rich metadata of the video content with local context information to perform video content rendering. This enables the display effect to dynamically adapt to the content of each frame and the local context of the receiving end, achieving high-quality, personalized rendering effects.

[0128] It should be noted that the local context information listed above is merely an example. Any information characterizing user features, device operating characteristics, and environmental characteristics can be used as the local context information of the receiving end. This allows for personalized video effects to match different user characteristics, as well as video effects to match different environmental characteristics and different device operating characteristics.

[0129] Similarly, metadata can also include other data that characterizes the video content, such as RGB to grayscale distribution data.

[0130] In some embodiments, display parameters include one or more of the following data: resolution, contrast ratio, brightness, color temperature, color gamut, saturation, or refresh rate. These display parameters are merely examples; any parameter that can affect the final presentation of the video can be used as a display parameter to be adjusted.

[0131] In some embodiments, the receiver is specifically configured to: determine a first target area in the display area of ​​the receiver's display panel based on metadata and the receiver's local context information, and adjust the display parameters of the receiver for the first target area.

[0132] In one example, the primary target area is the user's focus of attention. In wearable devices and AR / VR headsets, the receiver may include a gaze tracking module. This module, based on gaze-aware technology and combined with the user's real-time eye gaze data (an example of interactive data), determines the user's focus on the video content and adjusts the display parameters (such as brightness, contrast, and resolution) of that focus area accordingly to ensure its clarity.

[0133] This prioritizes rendering quality for the user's visual focus areas, preserving details in those areas and improving visual resolution. In some scenarios, it reduces computation in peripheral areas, lowering power consumption.

[0134] In one example, the first target region could also be a key visual region. Of course, the first target region could also be other local regions.

[0135] The first target region listed above is merely an example. The receiving end may determine other regions as the first target region based on metadata. Alternatively, the receiving end may determine other regions as the first target region based on local context information. Or, the receiving end may determine other regions as the first target region based on both metadata and local context information.

[0136] As one possible implementation, adjusting the display parameters of the receiving end for the first target area can also be understood or replaced as follows: For the first target area, a first rendering strategy is adopted. For areas outside the first target area, a second rendering strategy is adopted. This second rendering strategy differs from the first rendering strategy. In this way, the first target area can have a different rendering effect than other areas to meet the user's specific visual needs for the first target area. For example, refer to... Figure 4 The saturation of the first target region is A, and the saturation of other regions is B.

[0137] As one possible implementation, the aforementioned eye fixation data can also be used for fatigue estimation models and visual comfort scoring. This would allow the display system to dynamically adjust the brightness, sharpness, and motion intensity of video content to reduce visual fatigue.

[0138] Compared to related technologies where instructions are applied to the entire video frame, the method provided in this application can identify important display objects or regions within the video content, such as distinguishing faces, text, or other visually important areas. This allows for priority given to important display objects or regions. On one hand, this improves the display performance of these important objects or regions. On the other hand, when receiving end resources are limited, rendering resources can be preferentially allocated to important regions to achieve relatively better rendering results under limited rendering resource conditions.

[0139] In some embodiments, the transmitting end includes a depth and geometry module for generating depth information of objects displayed in the video content; The receiving end includes a stereo conversion module, which renders the video content into a stereo image based on depth information. The perspective of the stereo image is determined based on the position or posture information of the viewing user.

[0140] The stereoscopic conversion module, also known as the 3D projection module, refers to the perspective of a stereoscopic image, which can be understood as the perspective of the multi-view image corresponding to the stereoscopic image.

[0141] As one possible implementation, the depth and geometry module can obtain a depth map based on semantic segmentation and motion estimation of video content, combined with the contours of displayed objects, and use this depth map as depth information.

[0142] Based on this depth information, the receiving end can convert the original 2D video into a 3D video, and the viewing angle in the 3D video can be adjusted according to the user's position or posture.

[0143] For example, refer to Figure 5A In (a), if the system detects that the user is sitting on the right side of the TV, the stereoscopic conversion module in the TV can adjust the viewing angle of the 3D video to provide the user in that position with a better 3D viewing experience. For example, see reference... Figure 5A (b) If the TV detects that the user is sitting in the middle position, it can adjust the viewing angle of the 3D video to provide the user in that position with a better 3D viewing experience.

[0144] As one possible implementation, in VR / AR, glasses-free 3D, and other similar scenarios, the perspective of the stereoscopic image changes as the user's head turns, as if observing a real three-dimensional world through a window. This can enhance the sense of presence and interactivity.

[0145] For example, for systems equipped with cameras or sensors, the stereo conversion module can dynamically adjust the rendering based on the depth information of the video content to match the user's head position or eye tracking, ensuring that each user gets the best parallax.

[0146] For example, in prism or light field displays, stereo conversion modules can calibrate multiple viewpoints to optimize the naked-eye 3D reality without artifacts or ghosting.

[0147] The method provided in this application embodiment allows the receiving end to generate a stereoscopic image based on the depth information of the video content. The stereoscopic effect of the image originates from the three-dimensional structure contained in the video content itself. Therefore, a more realistic, natural, and seamless 3D effect can be obtained, providing a truly immersive 3D experience.

[0148] Furthermore, since the perspective of a stereoscopic image can be determined based on the viewer's position or posture, it can provide users in different positions or postures with corresponding perspectives, thereby further enhancing the sense of immersion.

[0149] In some embodiments, the receiver can perform color shift compensation for wide-viewing-angle scenes. For example, LCD panels are prone to color distortion at wide viewing angles, causing hue and saturation shifts in key visual areas such as skin tones and skies. As a possible implementation, the receiver may also include a color compensation module for performing color compensation based on key visual areas in metadata to maintain color fidelity at wide viewing angles. For example, refer to... Figure 5B If the system detects that a user is watching video from a side angle on the TV, it preserves the natural tones of faces and compensates for the red and yellow components that are easily lost at side views. Furthermore, it preserves the natural tones of the sky at side views. This enhances the viewing experience with wide viewing angles. Similarly, in environments like living rooms or public displays, where different viewers are at different angles, the receiver ensures that skin tones and colors of natural landscapes appear true to life for each viewer.

[0150] In some embodiments, the receiver further includes a color and gamut module (also known as a color and gamut optimization intelligent module) for determining a tone mapping strategy based on metadata and local context information of the receiver, wherein the tone mapping strategy is used to compress the brightness range of high dynamic range (HDR) video content to the renderable range of the receiver.

[0151] As one possible implementation, the receiver can generate an adaptive lookup table (LUT) based on user preference information. For example, the LUT might record: the user's habit of increasing screen brightness in high-saturation game scenes; and the user's habit of reducing blue light and increasing screen contrast in late-night movie viewing scenarios. Subsequently, upon detecting a late-night movie viewing scenario, the color and color gamut module can adjust the display parameters of the video content to enhance screen contrast and reduce blue light.

[0152] As one possible implementation, the color and gamut module can determine the tone mapping strategy based on the color offset matrix.

[0153] The color and gamut module can also be used to adjust color depth (such as brightness), gamut, and white point (such as brightness) based on metadata and local context information from the receiver to achieve perceptual balance and visual comfort.

[0154] For example, the color and gamut module adjusts color depth, gamut, and white point based on a LUT or color offset matrix.

[0155] The white point represents the chromaticity coordinates of white as perceived by the human eye under specific lighting conditions. In some cases, the white point serves as the baseline reference point for the entire color space. In other cases, it may affect the user's perception of white in video content.

[0156] For example, the color temperature of ambient light greatly affects our perception of "white" on a screen. Under warm yellow light, a white dot may appear bluish. Under candlelight, a white dot may appear warm yellow. Therefore, embodiments of this application can adjust the display parameters of the white dot to improve the display effect. For instance, under different ambient light conditions, in order to maintain the accuracy of the "white dot" color of the image, the color temperature of the corresponding area of ​​the display panel can be dynamically compensated.

[0157] As another example, the color and gamut module can adjust the color balance, brightness curve, and white point according to the scene-specific target of the video content, panel capabilities, and ambient lighting conditions.

[0158] The method provided in this application can match different tone mapping strategies for different viewing scenarios such as gaming, streaming media, and reading. This allows for a personalized and comfortable viewing experience based on the viewing scenario, ensuring color accuracy and achieving superior image quality and eye comfort.

[0159] In some embodiments, the receiving end further includes a language module, configured to adjust the rendering effect of the subtitles based on metadata and the local context information of the receiving end. The rendering effect of the subtitles includes, but is not limited to, font, position, or contrast.

[0160] For example, this language module can adjust the rendering effect of subtitles based on the scene semantics of the video content. For instance, in movie A, a character in a dark indoor setting contrasts sharply with the glaring desert sun outside the window. If the subtitles use a fixed white, they would blend into the background and disappear in bright areas, while being too glaring in dark areas, disrupting the atmosphere. Therefore, after receiving metadata, the language module identifies the current scene as "extremely high contrast" and obtains the brightness distribution map of the image. When the subtitles appear in a bright area outside the window, the language module can automatically adjust the subtitle color to dark gray or black and add an outline to ensure clear readability. When the subtitles move to a dark indoor area, the color is automatically adjusted back to a soft light gray or off-white, and the brightness is reduced to prevent glare. If the bright area in the center of the image is too large, the language module can also temporarily adjust the subtitles to a darker area at the bottom of the image. This provides a better subtitle rendering effect.

[0161] As another example, the language module can adjust the font, position, and contrast of the subtitles based on user-defined preference information.

[0162] For another example, refer to Figure 6(a) If the system detects that the user is positioned on the right side of the TV, the language module can render the subtitles on the right side for the user's convenience. (See reference) Figure 6 (b) If the user is detected to be in the center of the TV, the language module can render the subtitles in the center for the user's convenience.

[0163] In some embodiments, the transmitting end further includes a first noise suppression module, configured to generate RGB channel statistics for low grayscale regions in the video content. The receiving end further includes a second noise suppression module, configured to perform noise suppression on the low grayscale regions based on the RGB channel statistics for the low grayscale regions.

[0164] For example, RGB channel statistics include: the brightness distribution, noise characteristics, and color balance relationships of the RGB channels in low grayscale regions.

[0165] The brightness distribution can refer to the distribution of brightness values ​​for each channel in the low grayscale region. For example, a histogram can be used to represent this brightness distribution.

[0166] Noise characteristics can characterize the type and intensity of noise generated by each channel in the low grayscale region. Noise types include random noise and fixed-pattern noise.

[0167] Color balance can be characterized by the relative intensity relationship of the R, G, and B channels in low grayscale areas.

[0168] In low grayscale areas, image sensor noise and panel display noise (such as color shift in VA panels) are particularly noticeable. For example, visual artifacts such as color banding and hue shift are more apparent in dark scenes, especially on VA panels. Therefore, the method provided in this application allows the transmitting end to accurately quantify the distribution and type of noise by analyzing the RGB channel statistics of the low grayscale area. The receiving end then performs noise suppression accordingly, avoiding excessive blurring of image details (especially textures) by traditional noise reduction algorithms, achieving "noise removal without leaving a trace."

[0169] This can improve the viewing experience of dark scenes and night scenes that are prevalent in movies, games and other videos, making dark details purer and clearer, and significantly improving the overall texture and dynamic range perception of the picture.

[0170] As one possible implementation, the second noise suppression module can also perform noise suppression by combining local context information.

[0171] For example, in a bright viewing environment, the noise reduction intensity can be reduced to prioritize the preservation of details and textures, thereby improving image clarity. In a dimly lit viewing environment, where the human eye is extremely sensitive to noise, the second noise suppression module increases the noise reduction intensity in the aforementioned low grayscale areas to achieve a cleaner and more comfortable viewing experience.

[0172] For example, OLED panels themselves have no backlight and produce pure blacks, but they may exhibit "black crush" (loss of detail in dark areas) or color unevenness at low grayscale levels. Therefore, receivers employ cautious video noise reduction strategies, such as prioritizing color noise suppression over luminance noise suppression, to minimize the loss of detail in dark areas. For LCD / mini-LED panels, backlight leakage may occur, resulting in grayish dark areas. Therefore, receivers adopt more aggressive noise reduction strategies, which can be combined with local dimming to reduce backlighting while suppressing noise in low grayscale areas.

[0173] For example, for the low-grayscale area that the user is focusing on, the receiver performs mild noise reduction based on detail preservation to retain as much image detail as possible. For the surrounding visual areas, stronger noise reduction is performed because the human eye is less sensitive to peripheral details.

[0174] In some embodiments, the transmitting end further includes a lighting analysis module for analyzing lighting data of the video content; the lighting data includes one or more of the following: shadow direction, lighting direction, lighting area, or flare intensity. The lighting area can be an area illuminated by light, or an area where the light intensity is greater than a threshold. An area where the light intensity is greater than a threshold can be referred to as a highlight area.

[0175] The receiver also includes a lighting adjustment module, which is used to determine a second target area in the display area of ​​the receiver's display panel based on lighting data, and adjust the display parameters of the receiver for the second target area.

[0176] As one possible implementation, the transmitting end indicates a lighting model to the receiving end, which may include one or more of the following data: shadow direction, lighting direction, lighting area, or flare intensity. Alternatively, the transmitting end directly indicates one or more of the following data to the receiving end: shadow direction, lighting direction, lighting area, or flare intensity.

[0177] For example, refer to Figure 7In scene (b), the protagonist is in a bedroom with a lamp next to him. The transmitter can identify key objects in the scene: the lamp and the face of the character illuminated by the lamp. By analyzing the color tendencies of the highlights, shadows, and midtones of these objects, the transmitter can infer that the dominant light source is the lamp. This light is characterized by a low color temperature, exhibiting a strong warm yellow hue. Based on this, the transmitter can generate lighting data, for example: SceneLighting: {"CCT": 2000, "intensity": 0.6}. Where CCT: 2000 indicates that the detected scene color temperature is 2000K.

[0178] After receiving the metadata containing CCT: 2000, the receiver doesn't simply make the entire screen yellowish. Instead, it adjusts the color temperature (color temperature 1) of the area where the light is located to match the visual characteristics of a 2000K color temperature. Viewers will feel as if the screen is actually emitting light, perfectly conveying the warmth of the light and creating a highly immersive experience.

[0179] As another example, the receiving end can set the backlight level of the core area of ​​the image (such as the area where the person is located) to 92% based on the metadata of the video content. It can also slightly boost the pixel data in that area to compensate for the brightness. Ultimately, the perceived brightness is close to 100% backlight, but the actual backlight power consumption is reduced by 8%, resulting in less heat generation.

[0180] For transitional areas in the image, the receiver smoothly transitions the backlight level from 92% to 30%, creating an optical "feathered" edge. This actively pre-compensates for light diffusion from the display panel, suppressing the formation of halos.

[0181] Compared to Figure 7 (a) By uniformly adjusting the display parameters of a frame, the method provided in this application embodiment allows the transmitting end to identify the light source (such as a lamp) and its illumination range (illuminated area) in the image. The receiving end can then precisely adjust the display parameters of the illuminated area accordingly, thereby simulating a realistic lighting effect, making the displayed color tone consistent with the original scene environment, and enhancing the sense of realism and immersion.

[0182] As one possible implementation, the receiver can brighten the illuminated areas and darken the shadow areas based on the lighting data. This results in a higher dynamic range, more transparent highlights, purer shadows, and reduced halo effects due to precise brightness control, thus creating a realistic HDR effect.

[0183] As one possible implementation, the receiver can adjust display parameters based on the direction of light within the image to enhance the sense of depth. For example, if light is detected entering from the left side of the image, the backlight in the corresponding area on the left side of the screen can be slightly enhanced. This makes the displayed objects appear more three-dimensional, creating a three-dimensional light and shadow effect from a two-dimensional display.

[0184] In some embodiments, the receiver may further include a dynamic refresh and backlight intelligence module, which can be used to adjust the refresh rate based on motion estimation data or source color configuration data.

[0185] As one possible implementation, the dynamic refresh and backlight smart module can output timing parameters to trigger an adjustment of the refresh rate.

[0186] For example, based on content awareness, and combining scene motion and brightness intensity represented by metadata from the sending end, the display refresh rate is adjusted frame by frame. In fast-moving scenes, a high refresh rate is used to maintain sharpness. Static or dimly lit scenes can be rendered at a lower rate to reduce flicker and save power.

[0187] Compared to related technologies where the video refresh rate is fixed or only responds to device settings, the method provided in this application can adaptively adjust the refresh rate based on the characteristics of the video content itself. This allows for better video display by matching the characteristics of the video content.

[0188] The dynamic refresh and backlight intelligence module can also be used to make one or more of the following adjustments based on metadata (such as screen brightness and scene semantics): adjust the local dimming area or adjust the backlight.

[0189] In some examples, the dynamic refresh and intelligent backlight module adjusts the local dimming area or backlight direction of the panel in real time based on lighting data such as shadow direction, highlight areas, or flare intensity of the video content. For instance, by analyzing the lighting cues in the video content itself, it can be determined that there is a window on the left side of the movie scene, with sunlight shining in from the left and illuminating the character's face. The receiver can then simulate these lighting directions and characteristics, such as adjusting the backlight unit to minimize light scattered to the left. Thus, when a user watches the video, they will perceive the light on the screen shining from the window on the left side of the screen, illuminating the character and casting a shadow on their right. It is evident that this method allows for dynamic adjustment of the lighting direction, creating a visual experience with a sense of three-dimensional depth and environmental consistency on a two-dimensional screen. Furthermore, because it adjusts the backlight strategy for a local area, power consumption can be reduced.

[0190] In some examples, the dynamic refresh and backlight intelligence modules can also simulate the lighting consistency between the screen environment and the viewer's perspective. For instance, if sunlight shines on the receiver's screen from the left, the system can adjust the screen lighting to match the direction of the light, enhancing the three-dimensional continuity of the scene and improving the realism, immersion, and 3D depth perception of the video content.

[0191] In some embodiments, the transmitting end further includes an edge detection module for generating edge features and local contrast in the video content. For example, this edge detection module generates a high-precision edge map and statistical information on local contrast.

[0192] In some embodiments, the receiver further includes a super-resolution module for performing upsampling based on edge features and local contrast, the upsampling being used to adjust the resolution of the video content. This process can also be referred to as content-aware super-resolution.

[0193] As one possible implementation, the super-resolution module uses edge-aware and scene-aware models to upsample video content and output enhanced frames to achieve edge enhancement, strengthen spatial details, and provide efficient rendering of high-resolution videos (such as 8K).

[0194] For example, this method can be applied to action scenes, such as sports or action movies, to preserve the visual details of the action scene.

[0195] Compared to traditional super-resolution algorithms that apply similar processing to all regions, which can easily lead to oversharpening of textured areas or ringing artifacts in smooth areas, the method provided in this application allows the receiving end to upscale video content to a higher native resolution based on content-aware super-resolution. Furthermore, the receiving end utilizes edge features to accurately locate areas requiring contour enhancement and detail preservation, and uses local contrast to determine the intensity of the enhancement. For example, it can strongly enhance text and building edges while softening skin and sky areas, thereby achieving a clearer and more natural display effect.

[0196] Furthermore, it allows the use of lower-resolution source video during transmission or storage, reducing bandwidth dependence and saving costs for the transmission and storage of video content such as 4K / 8K.

[0197] In some embodiments, the receiver further includes an orchestration module for instructing the receiver to adjust the receiver's display parameters based on metadata and the receiver's local context information.

[0198] In other words, the orchestration module is responsible for the overall orchestration. For example, the orchestration module is located in the orchestration layer.

[0199] In some embodiments, the orchestration module is specifically used for: Analyze the optimization intent for the video content; Based on the optimization intent and the local context information of the receiving end, determine the modules to be used and the adjustment strategy for the display parameters; The module to be used is instructed to adjust the display parameters of the receiver according to the adjustment strategy.

[0200] As one possible implementation, the orchestration module at the receiving end can interpret the optimization intent in the metadata and make decisions based on local context information. The decisions include one or more of the following: the module (agent) to be activated or used, the adjustment strategy for display parameters, or the conflict resolution strategy.

[0201] This orchestration module acts as an execution manager. It maps display parameter adjustment strategies to specific module commands (or agent commands) to manage resource allocation and execution order across modules. For example, in a dimly lit scene with a face, the causal inference module can use commands to enhance local contrast and adjust color temperature. The orchestration module then sends instructions to the color and gamut module, which performs the local contrast enhancement. It can also send instructions to the lighting adjustment module, which adjusts the face area to a warmer color temperature.

[0202] The conflict resolution strategy described above can be used to resolve rendering conflicts between modules. This strategy can also be applied to conflicts between display metrics, such as those between thermal limits and contrast / sharpness.

[0203] For example, to prevent the panel from overheating, the peak backlight brightness is dynamically limited.

[0204] For example, a trade-off is made between display objectives such as power consumption and color fidelity, based on the current system status (such as the status of the display panel), real-time hardware constraints, and user preference information.

[0205] As one possible implementation, the orchestration module can also instruct each module on the aforementioned conflict resolution strategy, so that each module can perform rendering tasks based on the conflict resolution strategy, within the limits of hardware capabilities.

[0206] The method provided in this application embodiment allows the orchestration module to provide unified orchestration instructions to each module, thereby minimizing potential conflicts between modules. For example, when the optimization intention is to enhance color vibrancy and save power, the orchestration module will enable the corresponding module for color enhancement, but limit the peak brightness of the illumination adjustment module and may reduce the refresh rate, thereby achieving the coordination of multiple objectives.

[0207] Furthermore, this orchestration module generates corresponding strategies in real time based on optimization intent and local context information. On the one hand, this ensures operational consistency. On the other hand, it allows the generated strategies to adapt to different video scenarios. Thus, it maintains intelligent and robust display performance across various receiving ends and video content types.

[0208] In some embodiments, the sending end further includes a semantic analysis module for semantic segmentation of the video content to generate semantic segmentation data. For example, the semantic analysis module can perform object-level segmentation and region classification to obtain semantic masks such as faces, text, sky, and background, as well as object labels.

[0209] In some embodiments, the sending end further includes: a causal reasoning module, used to generate an optimization intent based on a predefined causal graph, wherein the causal graph is used to describe the causal relationships corresponding to visual elements in the video content.

[0210] As one possible implementation, the causal reasoning module uses a causal decision network (CDN) to perform reasoning, evaluate and model the impact of video content and local contextual information (such as environment and user preference information) on the rendering results, in order to obtain the optimization intent of the video content.

[0211] For example, in a scenario where a user is watching a dark-toned movie in a bright living room, strictly adhering to the creative intent (maintaining an extremely dark tone) might result in the user not being able to see details, leading to a poor experience. Conversely, completely adapting to the environment (significantly brightening) could disrupt the atmosphere and render the film's artistic expression ineffective. The causal reasoning module can identify causal intervention points to optimize visibility while minimizing disruption to the creative intent. For instance, the generated optimization intent might be: precisely and slightly increase the brightness of only key subjects in the image (such as a person's face), while keeping pure background shadows and black borders extremely dark. Simultaneously, slightly increase the contrast of midtones. In this way, the user can clearly see important content in the video, while the overall dark tone is preserved to the greatest extent possible.

[0212] As another example, the causal reasoning module can obtain an optimization intent based on the semantic understanding of the video content and the state of the audience: adjust the tone curve to adapt to a scene containing human subjects under low ambient light.

[0213] As one possible implementation, the causal reasoning module can evaluate the scene, user intent, and perceptual priority of the video content, and combine this with a causal graph to output the optimized intent of the video content. For example, the optimized intent could be: preserve skin color contrast, preserve facial tones, minimize motion blur, reduce blue light in low ambient light, and enhance text readability.

[0214] Compared to black-box neural networks or dependency-based deep learning models, which yield unclear optimization intentions, the method provided in this application uses causal graphs to explicitly describe the logical causal relationships corresponding to visual elements. For example, "strong light → shadow → shadow needs to be brightened to preserve details." Therefore, the optimization intentions generated based on causal graphs have high interpretability and trustworthiness, as well as high robustness and reliability. Furthermore, its decision-making process is transparent and verifiable. This facilitates user understanding and alignment with the user's display goals. This is crucial for safety-critical applications such as medical and automotive applications.

[0215] Furthermore, causal reasoning can anticipate conflicts before problems occur, offering a proactive advantage in conflict resolution. For example, the sending end might reason that "increasing overall brightness" could lead to "overexposed faces," thus the generated optimization intent would include a constraint to preserve skin tone. This allows the receiving end's orchestration module to formulate a more comprehensive adjustment strategy.

[0216] Furthermore, the layered collaboration of the causal reasoning module, orchestration module, and other modules enables the display system to provide perceptual optimization, contextual responsiveness, and interpretable real-time rendering. This further enhances the personalized rendering of video content.

[0217] In some embodiments, display parameters are determined based on the weights of each module. For example, when the orchestration module coordinates multiple agents (modules), the influence of each agent on the final image can be quantified as a "weight". For instance, in high dynamic range (HDR) scenes, the color and color gamut modules may have higher weights to ensure color accuracy. In motion scenes, the dynamic refresh and backlight intelligence modules have higher weights to ensure smoothness.

[0218] In some embodiments, the receiving end further includes a verification module, used to: detect whether the rendering effect after adjusting the display parameters meets the hardware capabilities of the receiving end and the user's perception; if the rendering effect after adjusting the display parameters does not meet the hardware capabilities of the receiving end, or does not meet the user's perception, then the display parameters are adjusted to spare parameters, or the weights of each module are adjusted.

[0219] The verification module, also known as the verification and security execution module, is, for example, a low-latency execution module. For example, this verification module can be deployed in the SoC firmware or display pipeline to reduce the impact on the display performance of the receiving end. For example, it can reduce illusion effects and decrease grayscale instability.

[0220] As one possible implementation, the verification module can detect whether the rendering effect after adjusting the display parameters meets the hardware capabilities of the receiving end and the user's perception, based on preset rules. For example, a statistical threshold can be set, and the display parameters can be compared with the statistical threshold to determine whether the rendering effect meets the hardware capabilities of the receiving end and the user's perception.

[0221] Here, "satisfying user perception" means that the rendering result does not exhibit any visual anomalies that the user can perceive. Examples of visual anomalies include brightness flickering, color banding, clipping, unnatural skin tone changes, or incomplete visual images.

[0222] As one possible implementation, the verification module can trigger a correction process upon detecting a visual anomaly. In one possible design, the verification module can work in conjunction with the orchestration module. If the verification module detects that the rendering effect does not meet the hardware capabilities of the receiving end and user perception, it can send a message to the orchestration module. The orchestration module can then, based on this message, either fall back to alternative parameters or re-determine the weights of other modules and, based on those weights, instruct each module to adjust its display parameters.

[0223] For example, a fallback to backup parameters can be implemented, whose rendering effect meets the hardware capabilities of the receiving end and the user's perception. For instance, based on the maximum brightness limit of the OLED panel, a fallback to backup parameters can be implemented to avoid degradation caused by the panel being set to maximum brightness for extended periods. In this way, the probability of visual anomalies in the video can be reduced without exceeding hardware capabilities.

[0224] Another example is reducing the weight of certain agents.

[0225] The method provided in this application embodiment, through the verification module, can promptly detect rendering problems caused by hardware limitations (such as color banding, severe halo, and overheating throttling). This fallback mechanism ensures that the display system will not output unacceptable or hardware-damaging images, providing system-level robustness assurance.

[0226] Furthermore, the dynamic weight adjustment of the intelligent agent enables the display system to self-optimize. This allows the display system to learn and adapt, fine-tuning during operation to approach optimal performance. For example, if verification finds that the current motion vector causes a "jelly effect," the verification module can instruct the weight of the motion vector to be reduced to avoid the jelly effect as much as possible.

[0227] Furthermore, this verification module works in conjunction with the orchestration module to ensure consistency across scenes and devices, as well as secure and high-quality rendering.

[0228] The above only lists some application scenarios of the method in the embodiments of this application. The method can also be applied to other scenarios.

[0229] For example, the receiving end dynamically adjusts the display color gamut based on the color ratio information in the metadata and the panel characteristics, balancing color richness and visual comfort, and reducing visual fatigue from prolonged viewing.

[0230] For example, the color ratio information is as follows: { "color_primaries": { "red": {"x": 0.680, "y": 0.320}, / / Slightly different from standard P3 red (0.680, 0.320) "green": {"x": 0.265, "y": 0.690}, / / Different from standard P3 green (0.265, 0.690) "blue": {"x": 0.150, "y": 0.060}, "white_point": {"x": 0.314, "y": 0.351} / / White point at D63, not D65 }, "transfer_function": "PQ", / / Electro-optical conversion function is PQ "matrix_coefficients": "BT.2020" / / Use the BT.2020 color matrix } The receiver can dynamically adjust the color gamut by combining the coordinates of the three primary colors and the white point of its own panel to match the display capabilities of the receiver.

[0231] For example, the receiving end can dynamically adjust the contrast or local dimming based on the deep scene analysis and semantic segmentation of the sending end, so as to maintain high image quality in various scenarios.

[0232] For example, the receiving end can use semantically segmented metadata to adjust saturation and hue, correcting color distortion when viewed from a wide angle.

[0233] For example, the receiver can use illumination data to dynamically calibrate white balance and color temperature to ensure that the screen color is consistent with the real-world lighting environment.

[0234] For example, in a vehicle head-up display (HUD) system, it can support real-time adaptation to changing environmental conditions (such as lighting conditions) and driver status. By monitoring cockpit lighting and driver alertness through sensors, the system adjusts display parameters such as contrast and brightness of the displayed content accordingly to ensure that it does not interfere with the driver's normal driving, thereby enhancing driving safety and comfort.

[0235] For example, for educational video content, DP Vision can enhance important elements (such as key concepts or charts) in real time based on scene semantic understanding and causal understanding of the video content to align with the user's learning goals, help the user understand these important elements, and without disrupting the overall visual effect.

[0236] For example, for game-related video content, DP Vision can identify high-priority visual elements (such as co-op members, equipment, or UI overlays) based on semantic understanding of the video content, and ensure that these visual elements are rendered with minimal latency and maximum clarity. This provides a responsive and immersive gaming experience. Furthermore, it reduces motion blur and improves responsiveness in high-FPS scenes.

[0237] For example, for streaming media platforms, intelligent agents can be embedded in the content delivery process to optimize playback on heterogeneous devices by utilizing cloud-side metadata, player parsing, and edge intelligence.

[0238] For example, in scenarios such as holography and light fields, the stereoscopic conversion module can perform viewpoint synthesis, parallax balancing, and light plane segmentation to achieve a more immersive, multi-angle projection experience.

[0239] These scenarios are merely examples, and the methods in this application embodiment can also be applied to other scenarios without limitation.

[0240] The above primarily describes the solutions provided by the embodiments of this application from a methodological perspective. It is understood that, in order to achieve the above functions, the electronic device includes hardware structures and / or software modules corresponding to the execution of each function. Based on the units and algorithm steps of the various examples described in the embodiments disclosed in this application, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by a computer driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solutions of the embodiments of this application.

[0241] This application embodiment can divide the electronic device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional module. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0242] This application also provides a video processing system, including a transmitter and a receiver. The specific structure and implementation steps of the transmitter and receiver can be found in the above embodiments and will not be repeated here.

[0243] For example, Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown.

[0244] like Figure 8 As shown, the electronic device 500 may include a processor 510. Optionally, the electronic device may also include a memory 520 and a display screen 530, etc. For example, the electronic device 500 is the aforementioned transmitting end or receiving end, or a chip in the transmitting end or a chip in the receiving end.

[0245] Processor 510 may include one or more processing units, such as application processors (APs), modem processors, tensor processing units (TPUs), graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0246] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0247] The processor 510 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory can store instructions or data that the processor 510 has just used or that are used repeatedly. If the processor 510 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.

[0248] In some embodiments, the processor 510 may include one or more interfaces. These one or more interfaces can be used to connect the processor 510 to the memory 520, the display 530, and the like.

[0249] The memory 520 can be used to store computer executable program code, which includes instructions. The memory 520 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as image playback), etc. The data storage area may store data created during the use of the electronic device 500, etc. The processor 510 executes various functional applications and data processing of the electronic device 500 by running instructions stored in the memory 520 and / or instructions stored in memory located within the processor.

[0250] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include... Figure 8 The diagram shows more or fewer components, or combinations of components, or separate components, or different arrangements of components. The components shown can be implemented in hardware, software, or a combination of both.

[0251] like Figure 9 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. This electronic device 2200 can be used to implement the methods described in the above method embodiments. Specifically, the electronic device 2200 may include a processing unit 2201. Optionally, the electronic device 2200 may include a display unit 2202.

[0252] The processing unit 2201 is used to support the electronic device 2200 in performing operations. Figures 2 to 7 The processing function described in any one of the following.

[0253] The display unit 2202 is used to support the electronic device 2200 in performing its functions. Figures 2 to 7 The display function described in any one of the following statements.

[0254] Optional, Figure 9 The electronic device 2200 shown may also include a communication unit ( Figure 9 (Not shown in the image), this communication unit is used to support electronic device 2200 in performing the steps of communication between electronic device and other electronic devices in the embodiments of this application.

[0255] Optional, Figure 9 The illustrated electronic device 2200 may further include a storage unit 2203 that stores programs or instructions. When the processing unit 2201 executes the program or instructions, it causes... Figure 9 The electronic device 2200 shown can perform the method described in the above-described method embodiments.

[0256] Figure 9The technical effects of the electronic device 2200 shown can be referred to the technical effects of the method shown in the above method embodiments, and will not be repeated here. Figure 9 The processing unit 2201 involved in the illustrated electronic device 2200 can be implemented by a processor or processor-related circuit components, and can be a processor or processing module. The communication unit can be implemented by a transceiver or transceiver-related circuit components, and can be a transceiver or transceiver module. The display unit 2202 can be implemented by display screen-related components.

[0257] This application also provides a chip system, such as... Figure 10 As shown, the chip system includes at least one processor 2301 and at least one interface circuit 2302. The processor 2301 and the interface circuit 2302 are interconnected via lines. For example, the interface circuit 2302 can be used to receive signals from other devices. As another example, the interface circuit 2302 can be used to send signals to other devices (e.g., the processor 2301). Exemplarily, the interface circuit 2302 can read instructions stored in memory and send those instructions to the processor 2301. When the instructions are executed by the processor 2301, the electronic device can perform the various steps performed by the electronic device in the above embodiments. Of course, the chip system may also include other discrete devices, and this application embodiment does not specifically limit this.

[0258] Optionally, the chip system may contain one or more processors. These processors can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in memory.

[0259] Optionally, the chip system may contain one or more memories. The memory may be integrated with the processor or disposed separately from it; this application does not limit this. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or disposed separately on different chips. This application does not specifically limit the type of memory or the arrangement of the memory and processor.

[0260] For example, the chip system can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.

[0261] It should be understood that each step in the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The method steps disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0262] It should be noted that the electronic device provided in this application embodiment belongs to the same concept as the method in the above embodiments. Any method provided in the method embodiments can be run on the electronic device, and the specific implementation process is detailed in the method embodiments, which will not be repeated here. The embodiments, implementation methods and related technical features of this application can be combined and substituted with each other without conflict.

[0263] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in any of the above embodiments.

[0264] In the embodiments of this application, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0265] It should be noted that, for the methods of the embodiments of this application, those skilled in the art will understand that all or part of the processes of the methods of the embodiments of this application can be implemented by a computer program controlling related hardware. This computer program can be stored in a computer-readable storage medium, such as in the memory of an electronic device, and executed by at least one processor within the electronic device. During execution, it can include the processes of the embodiments of the method. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, etc.

[0266] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0267] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. A video processing system, characterized in that, This includes the sending and receiving ends of the communication connection, where: The sending end is used to analyze the video content to obtain the metadata of the video content, and send the metadata to the receiving end; the metadata is used to describe the visual features of the video content in at least one dimension. The receiving end is used to receive the metadata and adjust the display parameters for displaying the video content based on the metadata and the local context information of the receiving end; the local context information of the receiving end is used to indicate at least one of the following: the environmental characteristics of the receiving end, the user characteristics of the viewing user of the receiving end, and the device operation characteristics of the receiving end.

2. The system according to claim 1, characterized in that, The metadata includes one or more of the following: edge features of the video content, local contrast of the video content, RGB channel statistics of low grayscale areas in the video content, lighting data of the video content, depth information of displayed objects in the video content, semantic segmentation data of the video content, source color configuration data of the video content, causal relationships corresponding to visual elements in the video content, motion estimation data of the video content, key visual regions in the video content, creative intent of the video content, saliency score of the video content, sentiment index of the video content, and optimization intent for the video content. The saliency score is used to describe the content weight index of different regions in the video content; the sentiment index is used to describe the emotional tone of the video content.

3. The system according to claim 1, characterized in that, The local context information includes at least one of the following: environmental information, the viewing user's preference information, and the receiving end's device information; the device information includes: the status of the receiving end's display panel and hardware capabilities; the preference information includes one or more of the following: the viewing user's focus, viewing angle, viewing distance, and the viewing user's relevant interaction data or viewing emotions.

4. The system according to claim 1, characterized in that, The display parameters include one or more of the following data: resolution, contrast ratio, brightness, color temperature, color gamut, saturation, or refresh rate.

5. The system according to any one of claims 1-4, characterized in that, The receiving end is specifically used for: Based on the metadata and the local context information of the receiving end, a first target area is determined in the display area of ​​the display panel of the receiving end, and the display parameters of the receiving end for the first target area are adjusted.

6. The system according to any one of claims 1-4, characterized in that, The transmitting end includes a depth and geometry module, used to generate depth information of objects displayed in the video content; The receiving end includes a stereo conversion module, used to render the video content into a stereoscopic image based on the depth information, wherein the perspective of the stereoscopic image is determined based on the position or posture information of the viewing user.

7. The system according to any one of claims 1-4, characterized in that, The receiving end also includes: The color and gamut module is used to determine a tone mapping strategy based on the metadata and the local context information of the receiving end. The tone mapping strategy is used to compress the brightness range of high dynamic range (HDR) video content to the renderable range of the receiving end.

8. The system according to any one of claims 1-4, characterized in that, The receiving end also includes: The language module is used to adjust the rendering effect of the subtitles based on the metadata and the local context information of the receiving end.

9. The system according to any one of claims 1-4, characterized in that, The transmitting end also includes a first noise suppression module, used to generate RGB channel statistics information of low grayscale areas in the video content; The receiving end also includes a second noise suppression module, which is used to perform noise suppression on the low grayscale region based on the RGB channel statistical information of the low grayscale region.

10. The system according to any one of claims 1-4, characterized in that, The transmitting end also includes a lighting analysis module for analyzing the lighting data of the video content; the lighting data includes one or more of the following: shadow direction, lighting direction, lighting area, or flare intensity; The receiving end also includes a light adjustment module, which is used to determine a second target area in the display area of ​​the display panel of the receiving end based on the light data, and adjust the display parameters of the receiving end for the second target area.

11. The system according to any one of claims 1-4, characterized in that, The transmitting end also includes an edge detection module, used to generate edge features and local contrast in the video content; The receiving end also includes a super-resolution module, used to perform upsampling based on the edge features and the local contrast, the upsampling being used to adjust the resolution of the video content.

12. The system according to any one of claims 1-4, characterized in that, The receiving end also includes an orchestration module for instructing the receiving end to adjust the display parameters of the receiving end based on the metadata and the local context information of the receiving end.

13. The system according to claim 12, characterized in that, The orchestration module is specifically used for: Analyze the optimization intent for the video content; Based on the optimization intent and the local context information of the receiving end, determine the modules to be used and the adjustment strategy for the display parameters; The module to be used is instructed to adjust the display parameters of the receiving end according to the adjustment strategy.

14. The system according to claim 2, characterized in that, The transmitting end also includes: The semantic analysis module is used to perform semantic segmentation on the video content and generate the semantic segmentation data; The causal reasoning module is used to generate the optimization intent based on a predefined causal graph, which describes the causal relationships corresponding to visual elements in the video content.

15. The system according to any one of claims 1-4, characterized in that, The display parameters are determined based on the weights of each module; the receiving end also includes a verification module, used for: The rendering effect after adjusting the display parameters is checked to see if it meets the hardware capabilities of the receiving end and the user's perception. If the rendering effect after adjusting the display parameters does not meet the hardware capabilities of the receiving end, or if the rendering effect after adjusting the display parameters does not meet the user's perception, then the display parameters will be adjusted to standby parameters, or the weights of each module will be adjusted.

16. A receiver, communicatively connected to a transmitter in a video processing system, characterized in that, The receiving end is used to receive metadata from the sending end, and adjust the display parameters for displaying video content based on the metadata and the local context information of the receiving end; The metadata is used to describe the visual features of the video content in at least one dimension; the local context information of the receiving end is used to indicate at least one of the following: the environmental characteristics of the receiving end, the user characteristics of the viewing user of the receiving end, and the device operation characteristics of the receiving end.

17. The receiving end according to claim 16, characterized in that, The local context information includes at least one of the following: environmental information, the viewing user's preference information, and the receiving end's device information; the device information includes: the status of the receiving end's display panel and hardware capabilities; the preference information includes one or more of the following: the viewing user's focus, viewing angle, viewing distance, and the viewing user's relevant interaction data or viewing emotions.

18. The receiving end according to claim 16 or 17, characterized in that, The receiving end is specifically used for: Based on the metadata and the local context information of the receiving end, a first target area is determined in the display area of ​​the display panel of the receiving end, and the display parameters of the receiving end for the first target area are adjusted.

19. The receiving end according to claim 16 or 17, characterized in that, The receiving end includes a stereo conversion module, which is used to render the video content into a stereo image based on depth information. The perspective of the stereo image is determined based on the position or posture information of the viewing user.

20. The receiving end according to claim 16 or 17, characterized in that, The receiving end also includes: The color and gamut module is used to determine a tone mapping strategy based on the metadata and the local context information of the receiving end. The tone mapping strategy is used to compress the brightness range of high dynamic range (HDR) video content to the renderable range of the receiving end.

21. The receiving end according to claim 16 or 17, characterized in that, The receiving end also includes: The language module is used to adjust the rendering effect of the subtitles based on the metadata and the local context information of the receiving end.

22. The receiving end according to claim 16 or 17, characterized in that, The receiving end also includes a second noise suppression module, which is used to perform noise suppression on the low grayscale region based on the RGB channel statistical information of the low grayscale region.

23. The receiving end according to claim 16 or 17, characterized in that, The receiving end also includes a light adjustment module, which is used to determine a second target area in the display area of ​​the display panel of the receiving end based on light data, and adjust the display parameters of the receiving end for the second target area.

24. The receiving end according to claim 16 or 17, characterized in that, The receiving end also includes a super-resolution module for performing upsampling based on edge features and local contrast, wherein the upsampling is used to adjust the resolution of the video content.

25. The receiving end according to claim 16 or 17, characterized in that, The receiving end also includes an orchestration module for instructing the receiving end to adjust the display parameters of the receiving end based on the metadata and the local context information of the receiving end.

26. The receiving end according to claim 25, characterized in that, The orchestration module is specifically used for: Analyze the optimization intent for the video content; Based on the optimization intent and the local context information of the receiving end, determine the modules to be used and the adjustment strategy for the display parameters; The module to be used is instructed to adjust the display parameters of the receiving end according to the adjustment strategy.

27. The receiving end according to claim 16 or 17, characterized in that, The display parameters are determined based on the weights of each module; the receiving end also includes a verification module, used for: The rendering effect after adjusting the display parameters is checked to see if it meets the hardware capabilities of the receiving end and the user's perception. If the rendering effect after adjusting the display parameters does not meet the hardware capabilities of the receiving end, or if the rendering effect after adjusting the display parameters does not meet the user's perception, then the display parameters will be adjusted to standby parameters, or the weights of each module will be adjusted.

28. A transmitter, communicatively connected to a receiver in a video processing system, characterized in that, The sending end is used to analyze the video content to obtain the metadata of the video content, and send the metadata to the receiving end; the metadata is used to describe the visual features of the video content in at least one dimension; The metadata includes one or more of the following: edge features of the video content, local contrast of the video content, RGB channel statistics of low grayscale areas in the video content, lighting data of the video content, depth information of displayed objects in the video content, semantic segmentation data of the video content, source color configuration data of the video content, causal relationships corresponding to visual elements in the video content, motion estimation data of the video content, key visual regions in the video content, creative intent of the video content, saliency score of the video content, sentiment index of the video content, and optimization intent for the video content. The saliency score is used to describe the content weight index of different regions in the video content; the sentiment index is used to describe the emotional tone of the video content.

29. The transmitting end according to claim 28, characterized in that, The transmitting end also includes a first noise suppression module, used to generate RGB channel statistics information of low grayscale areas in the video content; The transmitting end also includes a lighting analysis module for analyzing the lighting data of the video content; the lighting data includes one or more of the following: shadow direction, lighting direction, lighting area, or flare intensity; The transmitting end also includes an edge detection module, used to generate edge features and local contrast in the video content; The semantic analysis module is used to perform semantic segmentation on the video content and generate the semantic segmentation data; The causal reasoning module is used to generate the optimization intent based on a predefined causal graph, wherein the causal graph is used to describe the causal relationships corresponding to visual elements in the video content. The transmitting end includes a depth and geometry module, used to generate depth information of objects displayed in the video content.