Video transmission method and apparatus
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-09-12
- Publication Date
- 2026-08-07
AI Technical Summary
由于视频格式多种多样,源设备根据协商的视频制式对不同格式的视频进行统一处理,导致浪费带宽和源设备的功耗,以及宿设备的视频画质差、画面抖动、视频不同步等各种问题
Smart Images

Figure CN119342292B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and more particularly to a video transmission method and device. Background Technology
[0002] Currently, video playback is a fundamental and crucial function in smart devices. In scenarios where source devices (such as personal computers (PCs), mobile phones, set-top boxes, etc.) and destination devices (such as TVs, projectors, monitors, etc.) are interconnected, the source and destination devices negotiate supported video standards (such as resolution, frame rate, pixel format, pixel depth, clock, etc.). The source device then sends the video according to the negotiated video standard. Because video formats are diverse, the source device's uniform processing of different formats based on the negotiated standard leads to wasted bandwidth and power consumption, as well as various problems on the destination device such as poor video quality, image jitter, and video desynchronization. Summary of the Invention
[0003] This application provides a video transmission method and device, thereby reducing bandwidth and power consumption and improving video image quality.
[0004] In a first aspect, a video transmission method is provided, applied to a source device or a chip in a source device. The method includes: obtaining a Device Comprehensive Capability Description (DCCD) of a destination device, wherein the DCCD contains a video capability field; and outputting a video based on flags indicating whether super-resolution (SR) and / or motion estimation and motion compensation (MEMC) capabilities are supported, wherein the video format of the video includes a first resolution and a first frame rate.
[0005] The video transmission method provided in this application adds descriptions of motion estimation and motion compensation capabilities and super-resolution capabilities to the video capability field included in the DCCD. Since the source device learns whether the destination device supports super-resolution and / or motion estimation and motion compensation through the DCCD, the source device does not need to perform video format conversion on the video obtained by the source device according to the video format negotiated between the source device and the destination device before outputting the video. This reduces the bandwidth occupied by the transmitted video and the power consumption of the source device. The destination device performs effective image quality processing on the video to improve the image quality of the video.
[0006] In one possible implementation, outputting video based on flag information indicating whether super-resolution is supported includes: outputting video if the flag information indicating whether the destination device supports super-resolution indicates that the destination device supports super-resolution capability, wherein the first resolution is the original resolution of the video.
[0007] Once the source device learns through DCCD that the destination device supports super-resolution capability, the source device does not need to process the video resolution before outputting the video. That is, it does not need to process the resolution of the video obtained by the source device according to the resolution already negotiated between the source and destination devices. This reduces the bandwidth occupied by the transmitted video. Since there is no need to process the video resolution, the power consumption of the source device is reduced. The destination device performs effective image quality processing on the video at the original resolution, improving the image quality of the video.
[0008] In another possible implementation, the video is output based on the flag information indicating whether super-resolution is supported, including: if the flag information indicating whether the destination device supports super-resolution indicates that the destination device supports super-resolution capability, and the source device supports super-resolution capability, the video is output, and the first resolution is the resolution described by DCCD.
[0009] In another possible implementation, the first resolution is the resolution described by the DCCD, including: the first resolution is the optimal resolution described by the DCCD.
[0010] When both the source and destination devices support super-resolution capabilities, the source device can comprehensively determine the video resolution based on its own capabilities, output the video, and improve the video's image quality.
[0011] In another possible implementation, the video is output based on a flag indicating whether super-resolution is supported, including: if the flag indicating whether the destination device supports super-resolution indicates that the destination device does not support super-resolution capability, the video is output, and the first resolution is the resolution after processing the original resolution of the video.
[0012] Since the host device does not support super-resolution capabilities, the source device can process the video resolution according to its own capabilities to improve the video quality.
[0013] In another possible implementation, the video is output based on flag information indicating whether motion estimation and motion compensation capabilities are supported, including: if the flag information indicating whether the host device supports motion estimation and motion compensation capabilities indicates that the host device supports motion estimation and motion compensation capabilities, the video is output with a first frame rate being the original frame rate of the video.
[0014] Once the source device learns through DCCD that the destination device supports motion estimation and motion compensation, the source device does not need to process the video frame rate before outputting the video. That is, it does not need to process the frame rate of the video obtained by the source device according to the frame rate already negotiated between the source and destination devices, thereby reducing the bandwidth occupied by the transmitted video. Since there is no need to process the video frame rate, the power consumption of the source device is reduced. The destination device performs effective image quality processing on the video at the original frame rate, improving the image quality of the video.
[0015] In another possible implementation, the video is output based on flag information indicating whether motion estimation and motion compensation capabilities are supported, including: if the flag information indicating whether motion estimation and motion compensation capabilities are supported by the destination device indicates that the destination device supports motion estimation and motion compensation capabilities, and the source device supports motion estimation and motion compensation capabilities, the video is output with a first frame rate of the frame rate described by DCCD.
[0016] In another possible implementation, the first frame rate is the frame rate described by the DCCD, including: the first frame rate is the optimal frame rate described by the DCCD.
[0017] When both the source and destination devices support motion estimation and motion compensation, the source device can comprehensively determine the video frame rate based on its own capabilities, output the video, and improve the video's image quality.
[0018] In another possible implementation, the video is output based on flag information indicating whether motion estimation and motion compensation capabilities are supported, including: if the flag information indicating whether motion estimation and motion compensation capabilities are supported by the destination device indicates that the destination device does not support motion estimation and motion compensation capabilities, the video is output with a first frame rate being the frame rate after processing the original frame rate of the video.
[0019] Since the host device does not support motion estimation and motion compensation capabilities, the source device can adjust the video frame rate according to its own capabilities to improve the video quality.
[0020] In a second aspect, a video transmission apparatus is provided, comprising modules for performing operational steps of the method in the first aspect or any possible implementation thereof. For example, the video transmission apparatus includes a receiving module and a transmitting module. The receiving module is configured to acquire a DCCD of a destination device, the DCCD containing a video capability field. The transmitting module is configured to output video based on flag information including whether super-resolution is supported and / or flag information including whether motion estimation and motion compensation capabilities are supported, the video format of which includes a first resolution and a first frame rate.
[0021] In one possible implementation, the sending module, when outputting video based on the flag information indicating whether super-resolution is supported, specifically outputs video when the flag information indicating whether the destination device supports super-resolution indicates that the destination device supports super-resolution capability, wherein the first resolution is the original resolution of the video.
[0022] In another possible implementation, the sending module, when outputting video based on the flag information indicating whether super-resolution is supported, specifically outputs video when the flag information indicating whether the destination device supports super-resolution indicates that the destination device supports super-resolution capability and the source device supports super-resolution capability, with the first resolution being the resolution described by DCCD.
[0023] In another possible implementation, the first resolution is the optimal resolution described by the DCCD.
[0024] In another possible implementation, the sending module, when outputting video based on the flag information indicating whether super-resolution is supported, specifically outputs video when the flag information indicating whether the destination device supports super-resolution indicates that the destination device does not support super-resolution capability, and the first resolution is the resolution after processing the original resolution of the video.
[0025] In another possible implementation, the sending module, when outputting video based on flag information indicating whether motion estimation and motion compensation capabilities are supported, specifically outputs video with the first frame rate being the original frame rate of the video, provided that the flag information indicating whether the destination device supports motion estimation and motion compensation capabilities indicates that the destination device supports motion estimation and motion compensation capabilities.
[0026] In another possible implementation, the sending module, when outputting video based on flag information indicating whether motion estimation and motion compensation capabilities are supported, specifically outputs video if the flag information indicating whether the destination device supports motion estimation and motion compensation capabilities indicates that the destination device supports motion estimation and motion compensation capabilities, and the source device supports motion estimation and motion compensation capabilities, with the first frame rate being the frame rate described by DCCD.
[0027] In another possible implementation, the first frame rate is the optimal frame rate described by the DCCD.
[0028] In another possible implementation, the sending module, when outputting video based on the flag information indicating whether motion estimation and motion compensation capabilities are supported, specifically outputs video when the flag information indicating whether the destination device supports motion estimation and motion compensation capabilities indicates that the destination device does not support motion estimation and motion compensation capabilities, and the first frame rate is the frame rate after processing the original frame rate of the video.
[0029] Thirdly, a video transmission device is provided, comprising: a memory, a transceiver, and a processor; the memory, transceiver, and processor are configured to collaboratively perform the method described in the first aspect and any of its possible implementations.
[0030] Fourthly, a video transmission system is provided, comprising a source device and a destination device. The source device can be used to implement the functions of the source device in any of the first aspect and its possible implementations. The destination device performs effective image quality processing on the resolution and frame rate of the received video, thereby improving the video's image quality. Therefore, this video transmission system can also achieve the beneficial effects of the method in the first aspect, which will not be elaborated upon here.
[0031] Fifthly, a computer-readable storage medium is provided, storing computer instructions that, when executed on a computing device, perform the first aspect and its possible implementations. This computing device may be, for example, the aforementioned source device.
[0032] Sixthly, a computer program product is provided, comprising computer instructions that, when executed on a computing device, perform any one of the methods of the first aspect and its possible implementations. For example, the computing device may be the aforementioned source device.
[0033] A seventh aspect provides a chip system, comprising: a processor for retrieving and running a computer program from memory, causing a computing device equipped with the chip system to execute the first aspect and its possible implementations. This computing device may be a source device or a destination device as described above.
[0034] The beneficial effects achieved by the technical solutions of the second to seventh aspects of this application and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here.
[0035] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0036] Figure 1 A schematic diagram of an audio-visual transmission system provided in this application;
[0037] Figure 2 A schematic diagram of an audio / video encoding / decoding system provided in this application;
[0038] Figure 3 A schematic diagram of a video processing capability negotiation process provided for this application;
[0039] Figure 4 A flowchart illustrating a video transmission method provided in this application;
[0040] Figure 5 A schematic diagram of the structure of a video transmission device provided in this application;
[0041] Figure 6 A schematic diagram of the structure of a video transmission device provided in this application;
[0042] Figure 7 This is a schematic diagram of the structure of a display device provided in this application. Detailed Implementation
[0043] To facilitate understanding, the main terms used in this application will be explained first.
[0044] The Device Comprehensive Capability Description (DCCD) primarily defines a system framework and data structure for declaring the capabilities supported by the device, such as device parameters, performance attributes, and audio / video capabilities. The DCCD uses a variable-length structure; different capabilities supported by the device are encapsulated in different data blocks within the DCCD. The declaration of supported capabilities can be expanded by adding extension data blocks. The DCCD describes the data content structure and is independent of the communication protocols between devices.
[0045] Resolution: Also known as image resolution or image sharpness, it refers to the ability of a measurement or display system to distinguish details. It determines the fineness of details in a bitmap image and is commonly used as a measure of the detail and sharpness displayed in images, videos, or display devices. The higher the resolution, the clearer the image and the richer the details.
[0046] Frame rate: refers to the frequency (rate) at which bitmap images, measured in frames, appear continuously on a display.
[0047] Super-resolution (SR) is the process of creating a high-resolution image from a series of low-resolution images. This technique improves the quality of low-resolution images, making them appear sharper and more detailed. It has applications in many fields, such as digital image processing, video processing, imaging, and satellite remote sensing.
[0048] Motion estimation and motion compensation (MEMC) is a technology widely used in televisions, monitors and other video playback devices. It is an important indicator for improving picture quality. It refers to predicting the trajectory of moving objects and inserting new motion compensation frames into the original image to compensate for the scenes that were not originally in the video source, reducing image ghosting and jitter, thus making the moving picture look clearer and smoother.
[0049] Currently, the source device converts the video format according to the video format negotiated with the destination device. For example, the source device performs super-resolution and / or motion estimation and motion compensation to convert the low resolution of the video to the high resolution required by the destination device, or it processes the video frame rate. However, resolution scaling performed by the source device fails to fully utilize the super-resolution capabilities of the destination device, resulting in poor video quality. Matching the frame rate by the source device through frame interpolation or frame dropping introduces the "motion jitter" problem.
[0050] The receiving device can also restore the resolution of the received video to its original resolution, and then improve the image quality through super-resolution technology and picture quality (PQ) technology. The receiving device can also restore the frame rate of the received video to its original frame rate, and then enable motion estimation and motion compensation techniques to make the motion smoother. This increases the cost of the receiving device.
[0051] The source device processes videos of different formats uniformly according to the negotiated video standard. This requires the source device to perform operations such as resolution scaling, frame dropping, frame interpolation, and pixel-level downsampling. These processes result in the loss of video detail, making it impossible to achieve optimal image quality when restoring the video on the destination device, leading to poor video quality. Resolution scaling also causes various problems such as increased bandwidth and higher device power consumption.
[0052] The video standard mentioned in this application may refer to the video format, such as at least one of resolution, frame rate, pixel format, pixel depth, and clock.
[0053] To address the issues of wasted bandwidth and power consumption, as well as poor video quality, this application provides a video transmission method in which the source device acquires the DCCD of the destination device. The DCCD contains a video capability field, which includes flags indicating whether super-resolution is supported and / or whether motion estimation and motion compensation capabilities are supported. Based on the flags indicating whether super-resolution is supported and / or whether motion estimation and motion compensation capabilities are supported, the video is output. The video format includes a first resolution and a first frame rate. The first resolution can be the original resolution or the processed resolution of the video, and the first frame rate can be the original frame rate or the processed frame rate of the video.
[0054] The video transmission method provided in this application adds descriptions of motion estimation and motion compensation capabilities and super-resolution capabilities to the video capability field included in the DCCD. Since the source device learns whether the destination device supports super-resolution and / or motion estimation and motion compensation through the DCCD, the source device does not need to perform video format conversion on the video obtained by the source device according to the video format negotiated between the source device and the destination device before outputting the video. This reduces the bandwidth occupied by the transmitted video and the power consumption of the source device. The destination device performs effective image quality processing on the video to improve the image quality of the video.
[0055] The technical solutions involved in the embodiments of this application can be applied not only to current audio and video transmission technologies or standards, but also to future audio and video transmission technologies or standards. The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application.
[0056] In this embodiment, "video" is a general term, referring to a sequence of multiple consecutive frames, with each frame corresponding to one image. "Audiovisual" is an information application technology term referring to video, audio, or multimedia content that includes both video and audio.
[0057] Video streaming refers to the transmission of video data. For example, a video stream can be processed over a network as a stable and continuous stream. The video (or video data) captured by a video capture device consists of multiple frames of images. After the video data is encapsulated and packaged, a video stream is obtained. A video stream consists of multiple video frames, each corresponding to one image frame.
[0058] The video transmission method provided in this application supports video formats including but not limited to RGB, YCbCr4:4:4, YCbCr4:2:2, YCbCr4:2:0, ARGB, Y-only, and RAW.
[0059] The video transmission method provided in this application supports video data component bit widths (or color depths) including but not limited to one or more of the following: 8-bit, 10-bit, 12-bit, or 16-bit. For example, for RGB format video data, if the component bit width is 8-bit, then the R component of a pixel occupies 8 bits, the G component occupies 8 bits, and the B component occupies 8 bits. It can also be described as: the video transmission method supports 8-bit per component (bpc), 10-bpc, 12-bpc, and 16-bpc video data.
[0060] The video transmission method provided in this application will be described in detail below with reference to the accompanying drawings.
[0061] Figure 1 This is a schematic diagram of an audio-visual transmission system provided in this application. The video processing process may include, but is not limited to: video acquisition, video encoding, video transmission, video decoding, and playback.
[0062] Figure 1 The audio-visual transmission system includes a set-top box 110, a smart TV 120, multiple audio-visual playback devices, and a server 130. The set-top box 110 connects to the network via a network cable and can receive audio-visual streams from the server 130. The network enables audio-visual transmission and may include one or more network devices, such as a router or switch. In some optional implementations, the set-top box 110 and the server 130 may also communicate wirelessly; this embodiment is not limited to this approach.
[0063] Set-top box 110 is an audio-visual processing device used to receive, process, and push video or audio streams. In some possible scenarios, set-top box 110 may also be called an internet TV set-top box, a network HD media player, or something similar. For example, set-top box 110 can refer to a TV box provided by a network operator or a TV box purchased by the user. The hardware implementation of set-top box 110 can be found below. Figure 6 The description of that will not be repeated here.
[0064] The smart TV 120 is a display device with audio-visual processing capabilities, enabling functions such as receiving, processing, pushing, and playing video or audio-visual streams. In some possible cases, the smart TV 120 may refer to audio-visual devices such as conference tablets, smart TVs, or projectors; this application embodiment does not limit this. The hardware implementation of the smart TV 120 can be referred to below. Figure 6 or Figure 7 The description of that will not be repeated here.
[0065] Multiple audio-visual playback devices include audio-visual playback devices 121 to 124. For example, these audio-visual playback devices may include, but are not limited to, multimedia control platforms or other devices supporting audio-visual playback functions, such as virtual reality (VR) terminal devices or augmented reality (AR) terminal devices, etc. The hardware implementation of the audio-visual playback devices can be referred to below. Figure 7 The description of that will not be repeated here.
[0066] In this embodiment, the set-top box 110 and the smart TV 120 are connected via a network, for example, through an audio / video interface network. The set-top box 110 and various audio / video playback devices can also be connected via an audio / video interface network, as can the smart TV 120 and various audio / video playback devices. Exemplarily, this audio / video interface network can be a wired network supporting video and audio / video transmission. Optionally, this audio / video interface network can be called a unified multimedia interconnection network. In some implementations, the unified multimedia interconnection network can also have other names.
[0067] The Unified Multimedia Interconnection Network (UMIN) supports both uncompressed and compressed video transmission, as well as advanced features such as Quick Video Transport (QVT), Auto Low Latency Mode (ALLM), and Dynamic Frame Rate Refresh (DFR). It also supports LPCM format audio and video as defined by IEC 60958, and various HDR protocols, such as those specified in T / UWA 005.1-2022, including HDR Vivid. Furthermore, the UMIN supports encryption control and protection for audio and video data transmission.
[0068] Server 130 can be an application server or an authentication and authorization server. Server 130 can provide video services, game services, messaging services, music services, authentication and authorization services, etc. In one example, the functions of multiple services can be integrated on server 130; for example, game services and music services can be deployed on server 130. In another example, server 130 can integrate the functions of some services; for example, server 130 can deploy parts of game services and parts of video services. Server 130 can also utilize virtualization technology to provide multiple virtual machines, which provide various services. This application embodiment does not limit the deployment form of the server. Network device 131 is connected to server 130 wirelessly or via a wired connection. Figure 1The diagram is just one example; the network may also include other devices. Figure 1 It is not shown in the middle.
[0069] Figure 1 This is just an illustration; the audio-visual transmission system may also include other devices. Figure 1 Not shown in the diagram. This application does not limit the number or type of devices included in the system.
[0070] exist Figure 1 Based on the audio and video transmission system shown, Figure 2 This is a schematic diagram of an audio-visual encoding and decoding system provided in this application. The audio-visual encoding and decoding system includes a source device 210 and a destination device 220. The source device 210 establishes a communication connection with the destination device 220 through a unified multimedia interconnection network.
[0071] The aforementioned source device 210 can perform audio and video encoding functions, such as... Figure 1 and Figure 2 As shown, the source device 210 can be the one described above. Figure 1 The set-top box 110 or smart TV 120 shown, and the source device 210 can also be an audio-visual control center with audio-visual encoding capabilities, for example, the audio-visual control center includes one or more servers.
[0072] The source device 210 may include a data source 211, a preprocessing module 212, an audio / video transmission adapter 213, and a communication interface 214.
[0073] Data source 211 may include or may be any type of electronic device for acquiring audio and video, and / or any type of source audio and video generation device, such as a computer graphics processor for generating computer animation scenes or any type of device for acquiring and / or providing source audio and video, or computer-generated source audio and video. Data source 211 may be any type of memory or storage device for storing the aforementioned source audio and video. The aforementioned source audio and video may include multiple audio and video streams or images acquired by multiple audio and video acquisition devices (such as cameras), such as Ultra High Definition (UHD) video, High Definition (HD) video, 4K video, 8K video, etc.
[0074] The preprocessing module 212 is used to receive source audio and video and preprocess the source audio and video to obtain audio and video or multi-frame images. For example, the preprocessing performed by the preprocessing module 212 may include color format conversion (e.g., from RGB to YCbCr), octree structuring, audio and video splicing, audio track merging and deletion, or channel number adjustment, etc.
[0075] The audio / video transmitter adapter 213 is used to receive audio / video or images and encode them to obtain encoded data. In some optional cases, the encoded bitstream (encoded data) may also be referred to as a bitstream. If the encoded data is obtained by encoding audio / video data, then the bitstream refers to the audio / video stream; if the encoded data is obtained by encoding video data, then the bitstream refers to the video stream.
[0076] The communication interface 214 in the source device 210 can be used to: receive encoded data (such as video streams or audio / video streams) and send encoded data (or a version of the encoded data after any other processing) to another device such as the destination device 220 or any other device through a unified multimedia interconnection network for storage, display, playback or image reconstruction, etc.
[0077] Optionally, the source device 210 includes a bitstream buffer for storing bitstreams corresponding to one or more coding units.
[0078] The aforementioned receiver device 220 can perform audio and video decoding functions, such as... Figure 1 As shown, when the source device 210 is a set-top box 110, the destination device 220 can be... Figure 1 When the source device 210 is a smart TV 120 or any of the audio-visual playback devices shown, the destination device 220 can be any of the audio-visual playback devices.
[0079] The receiving device 220 may include an audio / video playback unit 221, a post-processing module 222, an audio / video receiver adapter 223, and a communication interface 224.
[0080] The communication interface 224 in the receiving device 220 is used to receive encoded data (or a version of the encoded data after any other processing) from the source device 210 or from any other source device such as a storage device.
[0081] Communication interfaces 214 and 224 can be used for a direct communication link between source device 210 and destination device 220, such as a direct wired connection. Figure 2 The Unified Multimedia Interconnection Network. For more information on the Unified Multimedia Interconnection Network, please refer to... Figure 1 The description will not be repeated here.
[0082] Communication interface 224 corresponds to communication interface 214. For example, it can be used to transmit data and process the data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain audio and video data.
[0083] Both communication interface 224 and communication interface 214 can be configured as follows: Figure 2The one-way or two-way communication interface indicated by the arrow pointing from the source device 210 to the destination device 220 of the corresponding unified multimedia interconnection network can be used to send and receive messages, establish connections, and transmit other information, such as information related to the communication link, or information related to audio and video data (e.g., descriptive information DIP for audio and video).
[0084] The audio / video receiver adapter 223 is used to receive encoded data and decode the encoded data to obtain decoded data (video or audio / video, etc.).
[0085] The post-processing module 222 is used to post-process the decoded data to obtain post-processed data (such as an image to be displayed or audio / video to be played). The post-processing performed by the post-processing module 222 may include, for example, color format conversion (e.g., from YCbCr to RGB), octree reconstruction, audio / video splitting and merging, or any other processing to generate data for output by the audio / video playback unit 221.
[0086] The audio-visual playback unit 221 is used to receive post-processed data for display or playback to users or viewers. The audio-visual playback unit 221 can be or includes any type of display for representing the reconstructed image, such as an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen. The audio-visual playback unit 221 may also include one or more audio-visual playback modules, each of which may refer to a speaker, a smart speaker, or an amplifier, etc.
[0087] As an optional implementation, the source device 210 and the destination device 220 can transmit encoded data using a data forwarding device. For example, the data forwarding device could be a router or a switch. It is worth noting that this data forwarding device needs to support a unified multimedia interconnection network.
[0088] In the embodiments of this application, the source device may also be referred to as an audio / video transmitting device, audio / video transmitting end, etc., and the destination device may also be referred to as an audio / video receiving device, audio / video playback device, etc. In this embodiment, the source device and the destination device are connected through a unified multimedia interconnection network.
[0089] In the first possible application scenario, the source device can be Figure 1 The set-top box 110 in the middle, the receiving device can be Figure 1 The smart TV 120 in the set-top box. If the set-top box negotiates with the smart TV, that is, the set-top box obtains the smart TV's DCCD, the video capability fields contained in the DCCD are expanded to include descriptions of motion estimation and motion compensation capabilities and super-resolution capabilities. The set-top box pushes video to the smart TV based on the smart TV's super-resolution capabilities and / or motion estimation and motion compensation capabilities.
[0090] In the second possible application scenario, the source device can be Figure 1 The set-top box 110 in the middle, the receiving device can be Figure 1 Any one of the audio / video playback devices, such as any one of audio / video playback devices 121 to 124. For example, the set-top box negotiates with the audio / video playback device, that is, the set-top box obtains the DCCD of the audio / video playback device, and the set-top box pushes audio / video data to the audio / video playback device.
[0091] In the third possible application scenario, the source device can be Figure 1 The smart TV 120 in the middle, the host device can be Figure 1 Any one of the audio / video playback devices, such as any one of audio / video playback devices 121 to 124. For example, the smart TV negotiates with the audio / video playback device, that is, the smart TV obtains the DCCD of the audio / video playback device, and the smart TV pushes audio / video data to the audio / video playback device.
[0092] The three possible application scenarios described above are merely examples provided in this embodiment and should not be construed as limiting this application. In other possible examples, the source device may be... Figure 1 The host device is any one of the audio-visual playback devices (such as audio-visual playback device 121), and the host device is another audio-visual playback device (such as audio-visual playback device 122) that is different from the aforementioned audio-visual playback devices.
[0093] The process of negotiating DCCD between the source and destination devices is explained below.
[0094] Figure 3 This is a flowchart illustrating a video processing capability negotiation method provided in this application. Here, the video transmission method of this application is used as an example. Figure 2 The following explanation uses the source device 210 or the chip in the source device 210, and the destination device 220 or the chip in the destination device 220 as examples. (See reference...) Figure 3 The video transmission method provided in this application includes steps 310 to 320.
[0095] Step 310: The source device acquires the DCCD of the destination device.
[0096] Step 320: The destination device acquires the DCCD of the source device.
[0097] The following explanation uses the example of a source device acquiring the DCCD of a destination device. For details on how a destination device acquires the DCCD of a source device, please refer to the description of how a source device acquires the DCCD of a destination device.
[0098] The destination device provides information indicating its capabilities. The source device queries the capabilities of the destination device, which includes the source device obtaining the capability description information of the destination device. The process of the source device obtaining the capability information of the destination device includes: the source device sending a query message to the destination device, which is used to query the capabilities of the destination device; then, the destination device sending (feedback) the capability description information of the destination device to the source device; and the source device receiving the capability description information sent by the destination device, which can indicate the capabilities of the destination device.
[0099] The capability description information of the destination device may include information indicating the video processing capabilities of the destination device. For example, video processing capabilities may include whether the destination device supports super-resolution capabilities and / or whether it supports motion estimation and motion compensation capabilities. In some possible cases, this capability description information may also be referred to as a capability descriptor, which may include one or more bits of information.
[0100] In some cases, the capability description information of the destination device can also be called a device comprehensive capability description, and the area in the destination device that stores the DCCD is called the DCCD area or simply DCCD, which is not limited in this application.
[0101] In this embodiment of the application, the DCCD makes a capability declaration regarding whether it supports super-resolution capability and / or whether it supports motion estimation and motion compensation capability. For example, the receiving device supports any one of the following: super-resolution capability, motion estimation and motion compensation capability, or does not support super-resolution capability or motion estimation and motion compensation capability.
[0102] The source device queries the destination device's super-resolution capabilities, motion estimation, and motion compensation capabilities via DCCD. Subsequently, the destination device feeds back its own super-resolution capabilities, motion estimation, and motion compensation capabilities to the source device.
[0103] In some embodiments, the DCCD includes a video capability field, which primarily describes the video-related capabilities supported by the display. The video capability field includes flags indicating whether super-resolution is supported and / or whether motion estimation and motion compensation capabilities are supported.
[0104] For example, the data structure of the video capability field is shown in Table 1.
[0105] Table 1
[0106]
[0107] As shown in Table 1, the offset 06h in the video capability field includes flags indicating whether MEMC and SR are supported. In some embodiments, different bit values can indicate whether super-resolution and / or motion estimation and motion compensation are supported.
[0108] For example, a bit value of 1 indicates support for super-resolution or motion estimation and motion compensation, while a bit value of 0 indicates no support for super-resolution or motion estimation and motion compensation. Similarly, a bit value of 0 indicates support for super-resolution or motion estimation and motion compensation, while a bit value of 1 indicates no support for super-resolution or motion estimation and motion compensation.
[0109] Optionally, the DCCD may also store other content describing the audio and video processing capabilities of the receiving device, such as audio processing capabilities (whether the receiving device supports processing pure audio data packets), HDR display capabilities (whether the receiving device supports displaying specific types of HDR video, such as HDR), or others.
[0110] In one alternative example, the receiving device includes a memory storing the aforementioned capability description information. This memory may include, but is not limited to: random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art.
[0111] In another alternative example, the destination device includes a dedicated register whose state indicates the capability description information described above. If the register indicates a value of 1, the capability description indicates that the destination device has super-resolution capability; if the register indicates a value of 0, the capability description indicates that the destination device does not have super-resolution capability. If the register indicates a value of 1, the capability description indicates that the destination device has motion estimation and motion compensation capability; if the register indicates a value of 0, the capability description indicates that the destination device does not have motion estimation and motion compensation capability.
[0112] The process of transmitting video between the source and destination devices is explained below.
[0113] Upon receiving the DCCD from the destination device, the source device parses the DCCD to determine whether the destination device supports it and / or whether motion estimation and motion compensation capabilities are supported. If the source device outputs video, based on the super-resolution support and / or motion estimation and motion compensation support information, it determines whether to process the video resolution and frame rate, outputs the video, and transmits it to the destination device.
[0114] Figure 4 This is a flowchart illustrating a video transmission method provided in this application. Here, the video transmission method of this application is described as follows: Figure 2 The following explanation uses the source device 210 or the chip in the source device 210, and the destination device 220 or the chip in the destination device 220 as examples. (See reference...) Figure 4 The video transmission method provided in this application includes steps 410 to 420.
[0115] Step 410: The source device outputs video based on the flags indicating whether super-resolution is supported and / or whether motion estimation and motion compensation capabilities are supported. The video format includes a first resolution and a first frame rate. Correspondingly, the destination device receives the video from the source device.
[0116] The video sent from the source device to the destination device can be a frame of video data to be sent by the source device. The video corresponds to the valid data (or valid pixels, or valid pixel data) of a frame image in the video.
[0117] This application does not limit the source from which the source device acquires the video. The source device may acquire the video from a storage device, or receive video sent by other devices.
[0118] In some embodiments, after acquiring the video, the source device can determine a video output strategy based on flags indicating whether the destination device supports super-resolution and / or motion estimation and motion compensation capabilities. The output strategy indicates the video format of the video. The source device outputs the video and transmits it to the destination device. The video format of the video can be the format indicated by the output strategy. For example, the output strategy indicates whether to process the video's resolution and / or frame rate. If the output strategy indicates processing the video's resolution and / or frame rate, it can also indicate the processed resolution and / or processed frame rate.
[0119] The following explains how the source device outputs video based on the flag information indicating whether the destination device supports super-resolution.
[0120] In the first possible implementation, if the host device supports super-resolution, the output video is provided if the host device's super-resolution flag indicates that the host device supports super-resolution capability. The video format of the video includes a first resolution that is the original resolution of the video.
[0121] The source device learns that the destination device has super-resolution capabilities. After acquiring the video, the source device does not need to process the video resolution and directly outputs the video at its original resolution. For example, if the original resolution of the video acquired by the source device is 4K, the first resolution of the video output by the source device will also be 4K.
[0122] In some embodiments, the DCCD acquired by the source device also includes a CTA Timing field, which describes the video timing information contained in the DCCD that supports the CTA protocol. The data structure of the CTA Timing field is shown in Table 2.
[0123] Table 2
[0124]
[0125] The CTA Timing Code refers to the VIC number in Table 1, Video Format Timings, of the CTA-861-G protocol. If an unknown VIC is detected during parsing on the other end, the VIC Code can be ignored, and subsequent parsing can continue.
[0126] For example, the VIC codes in the CTA Timing field are arranged in a preferred order, that is, the first CTA Timing is the most recommended timing in the CTA Timing field, the second is the second most recommended, and so on.
[0127] The original resolution can also be the resolution that the receiving device can recognize. For example, the original resolution can be the resolution indicated by any of the video standards shown in Table 2.
[0128] In the second possible implementation, if the destination device supports super-resolution and the source device supports super-resolution, the output video is provided with a video format including a first resolution described by DCCD.
[0129] The source device supporting super-resolution capability means that the source device has the ability to process resolution data. When the source device is aware that the destination device also has super-resolution capability, after acquiring the video, it makes a comprehensive decision based on its own capabilities and outputs the video according to the resolution described by DCCD. In other words, the video resolution is processed according to the resolution described by DCCD, and the output video resolution is the processed resolution. For example, if the original resolution of the video acquired by the source device is 4K, the first resolution of the video output by the source device will be 8K.
[0130] In some embodiments, the video format includes a first resolution that is the optimal resolution described by DCCD. For example, Table 2 indicates that the first CTA Timing is the most recommended Timing in the CTA Timing field, converting the original resolution of the video to the resolution indicated by the most recommended Timing.
[0131] Optionally, the video format includes a first resolution of the suboptimal resolution described by DCCD.
[0132] In the third possible implementation, if the flag information indicating whether the destination device supports super-resolution indicates that the destination device does not support super-resolution capability, the output video is provided, and the video format of the video includes the first resolution as the resolution after processing the original resolution of the video.
[0133] If the source device learns that the destination device lacks super-resolution capabilities, it will process the video resolution after acquiring it, and output the video at the processed resolution. For example, the source device can process the original video resolution according to a resolution negotiated with the destination device; that is, the source device will output the video at the negotiated resolution. The negotiated resolution can be any of the resolutions shown in Table 2. Alternatively, the source device can process the original video resolution according to the resolution described by DCCD; the source device will output the video at the suboptimal resolution described by DCCD. For example, if the original resolution of the video acquired by the source device is 2K, the first resolution of the video output by the source device is 4K.
[0134] In the above scheme, the original resolution of the video can be any of the resolutions shown in Table 2. Optionally, the resolution of the video acquired by the source device may not be any of the resolutions shown in Table 2. In this case, the source device can process the original resolution of the video according to the resolution negotiated with the destination device, that is, convert the resolution of the video output by the source device to the resolution negotiated with the destination device. Alternatively, the source device can convert the original resolution of the video to any of the resolutions shown in Table 2.
[0135] The following explains how the source device outputs video based on the flag information indicating whether the motion estimation and motion compensation capabilities of the destination device support it.
[0136] In the first possible implementation, if the host device supports motion estimation and motion compensation capabilities, indicated by flag information indicating that the host device supports motion estimation and motion compensation capabilities, the output video is provided, and the video format of the video includes the first frame rate as the original frame rate of the video.
[0137] The source device learns that the destination device has motion estimation and motion compensation capabilities. After acquiring the video, the source device does not need to process the video's frame rate and directly outputs the video at the original frame rate. For example, if the source device acquires a video with an original frame rate of 60 frames per second, the first frame rate of the video output by the source device will be 4K 60 frames per second.
[0138] The raw frame rate can also be a frame rate that the receiving device can recognize. For example, the raw frame rate can be the frame rate indicated by any of the video standards shown in Table 2.
[0139] In the second possible implementation, if the destination device supports motion estimation and motion compensation capabilities, and the source device supports motion estimation and motion compensation capabilities, then the output video is provided. The video format of the video includes the first frame rate as described by DCCD.
[0140] The source device supports motion estimation and motion compensation capabilities, indicating that it has frame rate processing capabilities. Given that the source device is aware that the destination device also possesses these capabilities, after acquiring the video, the source device makes a comprehensive decision based on its own capabilities and outputs the video at the frame rate described by the DCCD. In other words, the video's frame rate is processed according to the frame rate described by the DCCD, and the output video's frame rate is the processed frame rate. For example, if the original frame rate of the video acquired by the source device is 30 frames per second, the first frame rate of the video output by the source device will be 60 frames per second.
[0141] In some embodiments, the video format includes a first frame rate that is the optimal frame rate described by DCCD. For example, Table 2 indicates that the first CTA Timing is the most recommended Timing in the CTA Timing field, converting the original frame rate of the video to the frame rate indicated by the most recommended Timing.
[0142] Optionally, the video format includes a first frame rate and a second-best frame rate included in DCCD.
[0143] In a third possible implementation, if the host device does not support motion estimation and motion compensation, the output video is provided if the host device does not support motion estimation and motion compensation. The video format of the output video includes a first frame rate which is the frame rate after processing the original frame rate of the video.
[0144] Knowing that the destination device lacks motion estimation and motion compensation capabilities, the source device, after acquiring the video, processes the video's frame rate, outputting a video with the processed frame rate. For example, the source device can process the original video frame rate according to a frame rate negotiated with the destination device; that is, the source device outputs a video with the negotiated frame rate. The negotiated frame rate can be any of the frame rates shown in Table 2. Alternatively, the source device can process the original video frame rate according to the frame rate described by DCCD; the source device outputs a video with the suboptimal frame rate described by DCCD. For example, if the original frame rate of the video acquired by the source device is 30 frames per second, the first frame rate of the video output by the source device is 60 frames per second.
[0145] In the above scheme, the original frame rate of the video can be any of the frame rates shown in Table 2. Optionally, the frame rate of the video obtained by the source device may not be any of the frame rates shown in Table 2. In this case, the source device can process the original frame rate of the video according to the frame rate negotiated with the destination device, that is, convert the frame rate of the video output by the source device into the frame rate negotiated with the destination device. Alternatively, the source device can convert the original frame rate of the video into any of the frame rates shown in Table 2.
[0146] The above-mentioned possible implementations of super-resolution capability, motion estimation, and motion compensation capability can be combined arbitrarily, and this application does not limit them.
[0147] For example, if the host device supports super-resolution capabilities and motion estimation and motion compensation capabilities, the output video has a video format including a first resolution that is the original resolution of the video and a first frame rate that is the original frame rate of the video.
[0148] For example, if the host device does not support super-resolution capability and supports motion estimation and motion compensation capability, the output video has the first resolution as the resolution converted from the original resolution of the video, and the first frame rate as the original frame rate of the video.
[0149] For example, if the host device supports super-resolution capabilities but does not support motion estimation and motion compensation capabilities, the output video has the first resolution as the original resolution of the video and the first frame rate as the frame rate converted from the original frame rate of the video.
[0150] For example, if the host device does not support super-resolution capabilities and does not support motion estimation and motion compensation capabilities, the output video has the following first resolution: the resolution after conversion of the original resolution of the video, and the first frame rate: the frame rate after conversion of the original frame rate of the video.
[0151] The source device transmits video to the destination device at its original resolution and frame rate, effectively reducing transmission bandwidth and saving power. The destination device can then process the image quality more effectively based on the original resolution and frame rate, achieving optimal results.
[0152] Step 420: The destination device processes the video from the source device.
[0153] When the destination device supports super-resolution capabilities, after receiving video from the source device, the destination device performs super-resolution processing on the video's initial resolution. For example, it converts the video's initial resolution to a second resolution. The initial resolution could be 4K, and the second resolution could be 8K.
[0154] The video standard, including its first resolution, can be the original resolution of the video. Alternatively, the video standard, including its first resolution, can also be the resolution obtained after processing the original resolution.
[0155] In some embodiments, the source device compares the resolution of the received video with the target resolution. If the resolution of the received video is lower than the target resolution, super-resolution processing is performed on the video. The target resolution may be the resolution of the display.
[0156] For example, the super-resolution processing mainly includes the following steps.
[0157] 1. Feature Extraction: First, the low-resolution image is processed using a convolutional neural network to extract features, resulting in a feature map. The purpose of this step is to extract useful feature information from the low-resolution image, preparing for the subsequent generation of high-resolution images.
[0158] 2. Generating High-Resolution Images: Next, the feature maps are input into the generator, which generates high-resolution images layer by layer through a series of convolutional and deconvolutional layers. This step is the core of super-resolution processing; using the learned feature information, the generator attempts to reconstruct a high-resolution image.
[0159] 3. Calculate perspective loss: Compare the generated high-resolution image with the real high-resolution image to calculate the perspective loss. Perspective loss is a loss function used to measure the difference between the image generated by the generator and the real image, taking into account factors such as image brightness, contrast, and structure.
[0160] 4. Parameter Optimization: Finally, by optimizing the parameters of the generator and discriminator, the generated high-resolution image becomes increasingly closer to the real high-resolution image. This step, by adjusting the parameters of the generator and discriminator, continuously improves the quality of the generated image, ultimately achieving an effect similar to the real high-resolution image.
[0161] The entire super-resolution process is an iterative optimization process. By continuously adjusting and optimizing parameters, the quality of the generated image is continuously improved, ultimately achieving the goal of super-resolution. This process involves not only image processing techniques but also deep learning and machine learning methods, making it possible to generate high-resolution images from low-resolution images.
[0162] For example, the image quality is improved by the scaler module of the host device (such as a TV). High-end TVs can perform super-resolution processing through AI SR, resulting in better picture quality.
[0163] When the receiving device supports motion estimation and motion compensation capabilities, after receiving video from the source device, the receiving device performs motion estimation and motion compensation processing on the first frame rate of the video. For example, the receiving device enables motion estimation and motion compensation capabilities and performs frame interpolation processing on the video. For example, the first frame rate of the video is converted to a second frame rate. The first frame rate can be 30 frames per second, and the second frame rate can be 60 frames per second.
[0164] In this context, the video format includes the original frame rate of the video. Alternatively, the video format may also include the frame rate after processing the original frame rate.
[0165] In some embodiments, the source device compares the frame rate of the received video with a target frame rate. If the frame rate of the received video is lower than the target frame rate, motion estimation and motion compensation processing are performed on the video frame rate. The target frame rate may be the frame rate of the display.
[0166] The primary purpose of motion estimation and compensation (MEP) technology is to improve the smoothness and clarity of video playback, especially in fast-moving scenes. MEP predicts the trajectory of an object by analyzing changes between two consecutive frames, and then generates one or more intermediate frames to fill the gaps between the original video frames—essentially, the device uses motion estimation and compensation capabilities for frame interpolation. This reduces motion blur, improves dynamic clarity, and makes high-speed motion appear smoother and more natural.
[0167] The motion estimation process primarily involves finding the best match between a block in the current frame and a block in the reference frame. This is achieved by calculating the similarity between the current block and blocks at various locations in the reference frame. Similarity metrics include MAD, MSE, and NCCF, while the H.264 standard uses the Sum of Absolute Transformed Differences (SATD). After finding the best match, motion estimation outputs a motion vector (MV), which represents the position coordinates of the reference block relative to the current block. This motion vector is then used in the motion compensation process.
[0168] Motion compensation involves using motion vectors obtained from motion estimation to extract corresponding blocks from a reference frame and comparing them with the corresponding blocks in the current frame to obtain residual blocks. If the matching degree between two blocks is high, the residual block is essentially zero, meaning that data compression efficiency is high. Motion compensation includes two types: global motion compensation and block motion compensation. It is a method for describing the differences between adjacent frames and is used to reduce spatial redundancy in video sequences. This method is used in video compression / video codecs to reduce spatial redundancy in video sequences and can also be used for deinterleaving operations.
[0169] The receiving device outputs the original frame rate or performs motion compensation by the motion estimation and motion compensation module of the receiving device (such as a TV), resulting in smoother motion.
[0170] The video transmission method provided in this application allows the source device to provide the best video output strategy by comprehensively considering the capabilities of the source device itself, based on the flag information such as whether MEMC and SR are supported in the video capability field 06h of the DCCD. The strategy indicates the video standard of the output video, which includes resolution and frame rate.
[0171] When the source device negotiates with the destination device, it obtains the destination device's DCCD and analyzes information such as whether it supports MEMC and SR video processing capabilities.
[0172] If the destination device supports MEMC and / or SR, the source device outputs the original resolution and / or original frame rate when outputting video, and the destination device enables MEMC for frame interpolation or / and enables SR for resolution scaling.
[0173] If the destination device supports MEMC and / or SR, and the source device supports MEMC and / or SR, the source device can make a comprehensive decision based on its own capabilities when outputting video, and output video according to the resolution and / or frame rate described by DCCD. For example, output video according to the optimal resolution and / or optimal frame rate described by DCCD.
[0174] If the source device does not support MEMC or / and SR, the source device outputs the processed resolution and / or the processed frame rate when outputting video.
[0175] Thus, the video capability field included in the DCCD adds descriptions of motion estimation and motion compensation capabilities, as well as super-resolution capabilities. Since the source device learns the video processing capabilities of the destination device through the DCCD, including super-resolution, motion estimation, and motion compensation capabilities, the source device outputs video based on the video processing capabilities of the destination device. This eliminates the need to convert the video obtained by the source device to a video format that has been negotiated between the source and destination devices. Consequently, the bandwidth occupied by the transmitted video is reduced, as is the power consumption of the source device. The destination device then performs effective image quality processing on the video, improving its image quality.
[0176] It is understood that, in order to achieve the functions in the above embodiments, the source device and the destination device include hardware structures and / or software modules corresponding to perform each function. Those skilled in the art should readily recognize that, based on the units and method steps of the various examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0177] The above text combines Figures 1 to 4 The video transmission method provided according to this embodiment is described in detail below, and will be combined with Figure 5 This describes the video transmission apparatus provided according to this embodiment.
[0178] Figure 5 This is a schematic diagram of a video transmission device provided in an embodiment of this application. This video transmission device can be used to implement the function of any one of the devices in the above method embodiments, and therefore can also achieve the beneficial effects of the above method embodiments. In this embodiment, the video transmission device can be as follows: Figure 1 The set-top box 110, smart TV 120, or any display device shown can also be Figure 2 The source device 210 or destination device 220 shown can also be the source device or destination device provided in subsequent embodiments. It should be understood that the video transmission device can also be a module (such as a chip) applied to any of the aforementioned devices.
[0179] like Figure 5As shown, the video transmission device includes a transceiver module 510 and a processing module 520. The transceiver module 510 and the processing module 520 can work together to implement the various steps in the above method embodiments. A more detailed description of the transceiver module 510 and the processing module 520 can be obtained directly from the relevant description of the device in the method embodiments shown in the foregoing figures, and will not be repeated here.
[0180] For example, transceiver module 510 is used to execute steps 310, 320, 410, and 420. Transceiver module 510 acquires the DCCD of the destination device, which contains a video capability field. Based on the flag information regarding whether super-resolution is supported and / or whether motion estimation and motion compensation capabilities are supported, contained in the video capability field, it outputs video. The video format of the video includes a first resolution and a first frame rate. Processing module 520 parses the DCCD to acquire the flag information regarding whether super-resolution is supported and / or whether motion estimation and motion compensation capabilities are supported, contained in the video capability field.
[0181] Optionally, the video transmission device may also include a storage module 530 for storing DCCDs and videos, etc.
[0182] When a video transmission device implements any of the video transmission methods shown in the foregoing figures through software, the video transmission device and its various units can also be software modules. The video transmission method described above is implemented by a processor calling this software module. This processor can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0183] Understandable. Figure 5 The video transmission device shown is merely an example provided in this embodiment. Depending on the video transmission process, the video transmission device may include more or fewer units, and this application does not limit it in this regard.
[0184] When the video transmission device is implemented in hardware, the hardware can be implemented using a processor or a chip system. The chip system includes one or more chips, each chip including interface circuitry and control circuitry. The interface circuitry is used to receive data from other devices outside the chip and transmit it to the control circuitry, or to send data from the control circuitry to other devices outside the chip. The control circuitry and interface circuitry are used through logic circuitry or executable code instructions to implement the method of any of the possible implementations in the above embodiments. The beneficial effects can be found in the description of any aspect of the above embodiments, and will not be repeated here.
[0185] It is understood that the processor in the embodiments of this application can be a CPU, or other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.
[0186] Figure 5 The video transmission device shown can also be implemented using video transmission equipment. Figure 6 The present application provides a schematic diagram of the structure of a video transmission device, which includes a memory 601 and at least one processor 602. The processor 602 can implement the video transmission method provided in the above embodiments, and the memory 601 is used to store the software instructions corresponding to the above video transmission method.
[0187] As an optional implementation, in hardware implementation, the video transmission device can refer to a chip or chip system that encapsulates one or more processors 602. For example, when the video transmission device is used to implement the method steps in the above embodiments, the processor 602 included in the video transmission device executes the steps of the source device and its possible sub-steps in the above method. In an optional case, the video transmission device may also include a communication interface 603, which can be used to send and receive data. For example, the communication interface 603 is used to receive DCCD, audio / video data, or send audio / video streams, etc.; the communication interface 603 can be implemented through interface circuitry included in the video transmission device. Therefore, in some examples, the communication interface 603 can also be referred to as the transceiver of the video transmission device. In this embodiment, the communication interface 603 supports the use of a unified multimedia interconnection network.
[0188] In the embodiments of this application, the communication interface 603, the processor 602, and the memory 601 can be connected via a bus 604. The bus 604 can be an address bus, a data bus, a control bus, etc. The bus 604 can be a Peripheral Component Interconnect Express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cachecoherent interconnect for accelerators (CCIX), or other types of buses, etc.
[0189] It is worth noting that video transmission devices can also perform Figure 5 The functions of the video transmission device shown are not described in detail here.
[0190] The video transmission device provided in this embodiment can be the set-top box 110, smart TV 120, source device 210, etc., or other devices with video processing functions. This application does not limit this. For example, when the aforementioned display device also has video processing functions, the video transmission device can refer to any of the aforementioned display devices.
[0191] in addition, Figure 5 The video transmission device shown can also be implemented using a display device. When the video transmission device is implemented using a display device, this embodiment provides a possible example, such as... Figure 7 As shown, Figure 7 This is a schematic diagram of the display device provided in this application. The display device includes: a processor 710, an external memory interface 720, an internal memory 721, a universal serial bus (USB) interface 730, a unified multimedia interconnection interface 731, an antenna 1, an antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, a sensor module 780, buttons 790, an indicator 792, a camera 793, a display screen 794, and a subscriber identification module (SIM) card interface 1-N 795, etc.
[0192] The aforementioned sensor module 780 may include sensors such as pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors.
[0193] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the display device. In other embodiments, the display device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0194] The processor 710 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0195] The controller can serve as the nerve center and command center of the display device. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0196] The processor 710 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 710 is a cache memory. This memory can store instructions or data that the processor 710 has just used or that are used repeatedly. If the processor 710 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 710, and thus improves the efficiency of the system.
[0197] In some embodiments, the processor 710 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, a unified multimedia interconnect interface, etc.
[0198] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the display device. In other embodiments, the display device may also employ different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0199] The wireless communication function of the display device can be implemented through antenna 1, antenna 2, mobile communication module 750, wireless communication module 760, modem processor, and baseband processor. In some embodiments, antenna 1 and mobile communication module 750 of the display device are coupled, and antenna 2 and wireless communication module 760 are coupled, enabling the display device to communicate with networks and other devices through wireless communication technology.
[0200] Wired communication functionality of the display device can be achieved through the USB interface 730 or the Unified Multimedia Interconnect Interface 731. For example, the display device can receive or send DCCDs, video streams, etc., through a bus connected via the Unified Multimedia Interconnect Interface 731.
[0201] The display device implements display functions through a GPU, a display screen 794, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 794 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. The processor 710 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0202] The display screen 794 is used to display images, videos, etc. The display screen 794 includes a display panel.
[0203] The display device can implement shooting functions through an ISP, a camera 793, a video codec, a GPU, a display screen 794, and an application processor. The ISP is used to process the data fed back by the camera 793. The camera 793 is used to capture still images or videos. In some embodiments, the display device may include one or N cameras 793, where N is a positive integer greater than 1.
[0204] In this embodiment, the above-mentioned display screen 794, video codec, GPU, display screen 794 and application processor can also be collectively referred to as the display unit of the receiving device, which is used to process and display the received video stream.
[0205] The external storage interface 720 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the display device. The external storage card communicates with the processor 710 through the external storage interface 720 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0206] Internal memory 721 can be used to store computer executable program code, which includes instructions. Processor 710 executes various functional applications and data processing of the display device by running the instructions stored in internal memory 721. For example, in this embodiment, processor 710 can execute instructions stored in internal memory 721, which may include a program storage area and a data storage area.
[0207] The program storage area can store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area can store data created during the use of the display device (such as audio and video data, phonebook, etc.). Furthermore, the internal memory 721 can include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0208] The display device can implement audio functions through an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, and an application processor. Examples include music playback and recording.
[0209] Buttons 790 include a power button, volume buttons, etc. Buttons 790 can be mechanical buttons or touch-sensitive buttons. Indicator 792 can be an indicator light, used to indicate charging status, battery level changes, or to indicate messages, missed calls, notifications, etc.
[0210] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0211] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0212] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0213] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0214] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0215] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0216] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video transmission method, characterized in that, The method, applied to a source device or a chip in the source device, includes: Obtain the Device Comprehensive Capability Description (DCCD) of the destination device. The DCCD includes a video capability field, which contains flags indicating whether motion estimation and motion compensation capabilities are supported. If the flag information indicating whether the destination device supports motion estimation and motion compensation capabilities indicates that the destination device supports motion estimation and motion compensation capabilities, a video is output, wherein the video format of the video includes a first frame rate, and the first frame rate is the original frame rate of the video.
2. The method according to claim 1, characterized in that, The video capability field also includes flags indicating whether super-resolution is supported; If the flag information indicating whether the host device supports super-resolution indicates that the host device supports super-resolution capability, the video format of the video also includes a first resolution, which is the original resolution of the video.
3. The method according to claim 2, characterized in that, If the flag information indicating whether the destination device supports super-resolution indicates that the destination device supports super-resolution capability, and the source device supports super-resolution capability, then the first resolution is the resolution described by the DCCD.
4. The method according to claim 3, characterized in that, The first resolution is the resolution described by the DCCD, including: The first resolution is the optimal resolution described by the DCCD.
5. The method according to any one of claims 2-4, characterized in that, If the flag information indicating whether the host device supports super-resolution indicates that the host device does not support super-resolution capability, the first resolution is the resolution after processing the original resolution of the video.
6. The method according to claim 1, characterized in that, If the destination device supports motion estimation and motion compensation capabilities, and the source device supports motion estimation and motion compensation capabilities, then the first frame rate is the frame rate described by the DCCD.
7. The method according to claim 6, characterized in that, The first frame rate is the frame rate described by the DCCD, including: The first frame rate is the optimal frame rate described by the DCCD.
8. The method according to claim 6 or 7, characterized in that, If the flag information indicating whether the host device supports motion estimation and motion compensation capabilities indicates that the host device does not support motion estimation and motion compensation capabilities, the first frame rate is the frame rate after processing the original frame rate of the video.
9. A video transmission device, characterized in that, include: A memory, a transceiver, and a processor; the memory, the transceiver, and the processor are configured to collaboratively perform the method of any one of claims 1 to 8.
10. A chip, characterized in that, The device includes one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of an electronic device and send the signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, it causes the processor to perform the operational steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The device stores computer instructions that, when executed on a computing device, perform the operational steps of the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Data structure and analysis method for describing comprehensive capability of equipment
CN115955515A
Equipment capability negotiation method and related equipment
CN118450012A