A video program playing system and method adapted to a digital retina architecture
By adopting an end-side node-edge node-cloud node architecture, combined with precoding and feature extraction technologies, the problem of insufficient IPTV video quality and functional module scalability is solved. It realizes high-quality video playback and flexible auxiliary services under low bandwidth, making it an innovative architecture suitable for the IPTV field.
Patent Information
- Application Number
- CN202510923085.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Existing IPTV technology has shortcomings in video quality and functional module scalability. It consumes a lot of bandwidth when the encoding accuracy is high, and it is difficult to flexibly adjust auxiliary services, thus failing to meet users' real-time interactive needs.
The architecture adopts a terminal node-edge node-cloud node approach, combining precoding technology and feature extraction. The cloud node determines the target video program and the associated program screen feature set, and transmits the encoded video stream and feature stream to the terminal node. The edge node performs feature analysis and provides auxiliary services, while the terminal node performs synchronization and hybrid rendering.
It improves video quality under low bandwidth conditions, flexibly adjusts and expands ancillary services, is compatible with the existing IPTV architecture, supports digital retina technology, and provides a high-quality video playback experience and flexible function adjustment mechanism.
Smart Images

Figure CN120475205B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of IPTV, in particular to a video program playing system and method adapted to digital retina architecture. BACKGROUND
[0002] IPTV (Internet Protocol Television) is a technology based on IP network transmission of video content, which realizes on-demand, live and interactive services of video programs by digitizing television signals and packaging them into IP data packets.
[0003] Traditional IPTV technology mainly adopts coding standards such as H.264 / AVC and H.265 / HEVC, reduces bandwidth demand through compression, and directly pushes video content to the terminal after pre-coding in the cloud. The video quality depends on the coding accuracy, the higher the coding accuracy, the higher the video quality, but the bandwidth occupation is also higher. In addition, the functional modules (such as advertisements) are fixed in the coding process, which is difficult to dynamically expand or adjust. If new functions (such as intelligent analysis of characters, objects, etc. in the program) are added, the existing scheme fixed to the coding process is obviously not applicable, and it is not flexible (it cannot be flexibly turned on and off for such auxiliary services), and it is difficult to meet the real-time interaction needs of users. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a video program playing system and method adapted to digital retina architecture, which builds an architecture of end-side node-edge node-cloud node, uses the combination of pre-coding technology and feature extraction, and improves the video quality of IPTV video programs in the form of relatively low bandwidth, and can flexibly adjust and expand auxiliary services, has good compatibility with the existing IPTV architecture, and can provide architectural support for introducing "digital retina" technology.
[0005] In order to achieve the above purpose, the embodiments of the present application are implemented by the following ways:
[0006] In a first aspect, the embodiments of the present application provide a video program playing system adapted to a digital retina architecture, comprising a cloud node, an edge node and an end-side node, the end-side node is configured to initiate a video program playing request; the cloud node is configured to determine a target video program and an associated program picture feature set based on the video program playing request, and transmit a video stream encoded based on the target video program and the associated program picture feature set to the end-side node, and transmit a feature stream encoded based on the program picture feature set to the edge node, wherein the program picture feature is obtained by performing feature extraction on a corresponding video program image in the target video program; the edge node is configured to receive the feature stream transmitted by the cloud node and perform feature analysis, determine display information according to the feature analysis result and a service type provided by the edge node, and transmit the display information to the end-side node, wherein the service type comprises at least one of intelligent role analysis, intelligent article analysis and intelligent advertisement placement; the end-side node is configured to receive the video stream transmitted by the cloud node and decode the video stream to obtain a video program image, receive the display information transmitted by the edge node and decode the display information, perform synchronous and mixed rendering based on the decoded video program image and the display information, obtain a program picture and play the program picture on the end-side node or a playing device connected to the end-side node.
[0007] With reference to the first aspect, in a first possible implementation manner of the first aspect, the cloud node is specifically configured to: determine a type of the end-side node based on the video program playing request, wherein the type of the end-side node comprises a set-top box and a smart terminal; determine the target video program and the associated program picture feature set based on the video program playing request; determine the video stream encoded based on the target video program and the associated program picture feature set by using a target encoding strategy based on the type of the end-side node; determine the feature stream encoded based on the target video program and the associated program picture feature set; transmit the video stream to the end-side node, and transmit the feature stream to the edge node.
[0008] In a second possible implementation manner of the first aspect, in the first possible implementation manner of the first aspect, the target coding strategy is: determining a correspondence between a video program image in the target video program and a program picture feature in the program picture feature set; for each frame of the video program image: determining a region of interest and a transition region based on the program picture feature corresponding to the frame of the video program image, performing low-precision coding on the video program image by using a region block of a first size to form a basic quality block, performing medium-precision coding on the transition region by using a region block of a second size to form an edge transition block, and performing high-precision coding on the region of interest by using a region block of a third size to form a ROI enhancement block, wherein the transition region is located at an edge of the region of interest, the first size is greater than the second size, and the second size is greater than the third size, wherein ROI represents the region of interest; if the type of the end-side node is a set-top box, uniformly packaging the basic quality block, the ROI enhancement block, and the edge transition block to form a single video stream in a frame order of the target video program; if the type of the end-side node is a smart terminal, uniformly packaging the edge transition block and the ROI enhancement block to form a comprehensive enhancement block, and hierarchically packaging the basic quality block and the comprehensive enhancement block to form a basic video stream and an enhancement video stream in the frame order of the target video program.
[0009] In a third possible implementation manner of the first aspect, in the second possible implementation manner of the first aspect, if the type of the end-side node is a set-top box, the end-side node is specifically configured to: receive the single video stream transmitted by the cloud-side node; for the packaging data of each frame of the video program image in the single video stream: unpacking the packaging data of the frame of the video program image to extract coding data and meta information of each block, wherein the meta information includes a block type identifier, a block position coordinate, and a block size; decoding the coding data of the basic quality block to generate a full-image low-precision video frame, decoding the coding data of the ROI enhancement block and covering the ROI enhancement block to the region of interest of the full-image low-precision video frame based on the block position coordinate and the block size of the ROI enhancement block, decoding the coding data of the edge transition block and covering the edge transition block to the transition region of the full-image low-precision video frame based on the block position coordinate and the block size of the edge transition block, and performing filtering and color space conversion to obtain the video program image.
[0010] In a fourth possible implementation of the first aspect, in the second possible implementation of the first aspect, if the type of the end-side node is a smart terminal, the end-side node is specifically configured to: receive the base video stream and the enhancement video stream transmitted by the cloud-side node; for the encapsulated data of each frame of video program image in the base video stream: decapsulate the encapsulated data of the frame of video program image to generate a full-image low-precision video frame after decoding; for the encapsulated data of a corresponding frame of video program image in the enhancement video stream: decapsulate the encapsulated data of the frame of video program image to extract the encoded data of the comprehensive enhancement block; perform secondary decapsulation on the comprehensive enhancement block to extract the encoded data and meta information of the ROI enhancement block and the edge transition block, wherein the meta information includes block type identification, block position coordinates and block size; decode the encoded data of the ROI enhancement block and overlay the ROI enhancement block to a region of interest of the full-image low-precision video frame of the corresponding frame of video program image based on the block position coordinates and the block size of the ROI enhancement block, and decode the encoded data of the edge transition block and overlay the edge transition block to a transition region of the full-image low-precision video frame of the corresponding frame of video program image based on the block position coordinates and the block size of the edge transition block, to obtain a video program image after filtering and color space conversion.
[0011] In a fifth possible implementation of the first aspect, in the second possible implementation of the first aspect, if the type of the end-side node is a set-top box, the cloud-side node is further configured to: determine version information of the set-top box from the video program playing request; determine a target encapsulation mode matched with the set-top box based on the version information of the set-top box, wherein the target encapsulation mode includes a first encapsulation mode and a second encapsulation mode, the first encapsulation mode indicates encapsulation of a base quality block, and the second encapsulation mode indicates unified encapsulation of the base quality block, the ROI enhancement block and the edge transition block; and determine a single video stream encapsulated based on the target encapsulation mode as a video stream to be delivered to the set-top box.
[0012] In a sixth possible implementation form of the first aspect, in combination with the second possible implementation form of the first aspect, if the type of the end-side node is a smart terminal, the cloud-side node is further configured to: determine a terminal network state carried in the video program playing request; determine a target encapsulation mode and a ROI proportion mode matched by the smart terminal based on the terminal network state, wherein the target encapsulation mode comprises a third encapsulation mode and a fourth encapsulation mode, the third encapsulation mode indicates that only the basic quality blocks are encapsulated, and the fourth encapsulation mode indicates that the edge transition blocks and the ROI enhancement blocks are uniformly encapsulated to form comprehensive enhancement blocks, and the basic quality blocks and the comprehensive enhancement blocks are encapsulated in layers, the ROI proportion mode comprises a ROI quantitative mode and a ROI full-amount mode, the ROI quantitative mode indicates that no more than a set number of ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of video program image are reserved, and the ROI full-amount mode indicates that all ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of video program image are reserved; encapsulate based on the target encapsulation mode and the ROI proportion mode to form a basic video stream as a video stream to be delivered to the smart terminal, or to form the basic video stream and an enhancement video stream as the video stream to be delivered to the smart terminal.
[0013] In a seventh possible implementation form of the first aspect, in combination with the first aspect, the edge node is specifically configured to: receive a feature stream sent by the cloud-side node, wherein the feature stream comprises program picture features corresponding to each frame of video program image, each program picture feature has a corresponding scene label, and each program picture feature comprises a plurality of types of sub-feature, the types of sub-feature comprising a character feature and an article feature; if the type of the auxiliary service provided by the edge node is intelligent character analysis, determine a target character based on the character feature, and determine target character information as display information based on the target character; if the type of the auxiliary service provided by the edge node is intelligent article analysis, determine a target article based on the article feature, and determine target article information as display information based on the target article; if the type of the auxiliary service provided by the edge node is intelligent advertisement delivery, determine a target article based on the article feature, determine corresponding target advertisement information based on the target article, and take the target advertisement information as the display information.
[0014] In a possible implementation form of the seventh implementation form of the first aspect, after obtaining the video program image, the end-side node is further configured to decode the received display information, extract a timestamp and a display information type identifier therefrom; determine a target display interface corresponding to the display information type identifier, and fill the display information into the target display interface, wherein the target display interface comprises position coordinates and interface size of the interface; synchronize the target display interface with the video program image of the corresponding frame based on the timestamp of the display information; superimpose the target display interface in the form of a transparent layer to a specified area of the video program image according to the position coordinates and the interface size, and perform pixel fusion through an Alpha blending algorithm; and render and output the fused program picture.
[0015] In a second aspect, the embodiments of the present application provide a video program playing method for adapting to a digital retina architecture, which is applied to the video program playing system for adapting to a digital retina architecture in the first aspect or any of the possible implementation forms of the first aspect. The method comprises: initiating a video program playing request by the end-side node; determining a target video program and an associated program picture feature set based on the video program playing request by the cloud-side node, and transmitting a video stream encoded based on the target video program and the associated program picture feature set to the end-side node, and transmitting a feature stream encoded based on the program picture feature set to the edge node, wherein the program picture feature is obtained by performing feature extraction on a corresponding video program image in the target video program; receiving the feature stream transmitted by the cloud-side node and performing feature analysis by the edge node, determining display information according to the feature analysis result and a provided auxiliary service type, and transmitting the display information to the end-side node, wherein the auxiliary service type comprises at least one of intelligent character analysis, intelligent article analysis, and intelligent advertisement placement; receiving the video stream transmitted by the cloud-side node and decoding to obtain a video program image by the end-side node, receiving the display information transmitted by the edge node and decoding, synchronizing and hybrid rendering based on the decoded video program image and the display information, obtaining a program picture, and playing the program picture on the end-side node or a playing device connected to the end-side node.
[0016] Advantages:
[0017] The video program playing system provided by the scheme is adapted to the digital retina architecture, and the cloud node-edge node (several)-terminal node architecture is built, and the video program playing request is initiated through the terminal node. Since the IPTV video program (non-live scene) is in the cloud, the feature extraction can be performed in advance, the corresponding features can be extracted for each frame image of each video program, the program picture feature set associated with the video program is formed and stored, and whether to perform pre-encoding is selected according to the need (pre-encoding can save resource occupation during distribution, and without pre-encoding, the feature encoding strategy can be flexibly adjusted), and the video program can be pre-encoded to improve the distribution efficiency. Therefore, the cloud node can determine the target video program and the associated program picture feature set (the program picture feature is obtained by performing feature extraction on the corresponding video program image in the target video program) based on the video program playing request, and transmit the video stream encoded based on the target video program and the associated program picture feature set to the terminal node, and transmit the feature stream encoded based on the program picture feature set to the edge node. The edge node can receive the feature stream sent by the cloud node and perform feature analysis, determine the display information according to the feature analysis result and the auxiliary service type (such as intelligent character analysis, intelligent article analysis, intelligent advertisement placement, etc.) provided by the edge node, and send the display information to the terminal node. The terminal node can receive the video stream transmitted by the cloud node and decode to obtain the video program image, receive the display information sent by the edge node and decode, and perform synchronous and mixed rendering based on the decoded video program image and the display information to obtain the program picture and play the program picture on the terminal node or the playing device connected to the terminal node. The architecture can utilize the combination of pre-encoding technology and feature extraction, flexibly adjust and expand auxiliary services (and can also well compatible with auxiliary services provided by external suppliers, the external suppliers can join the architecture in the form of the edge node, flexibly expand auxiliary services), and the architecture design is good in compatibility with the existing IPTV architecture, and the existing architecture can be slightly modified to complete (for old set-top boxes, the upgrade can be realized to support the playing of higher quality video, and if not upgraded, the original architecture can be used for video program playing, that is, high bandwidth occupation is used to watch high-definition video programs), and the architecture can also provide architecture support for introducing the “digital retina” technology (the digital retina technology is an idea proposed in recent years, the “video stream”+“feature stream” transmission scheme is realized by deploying a light-weight model in the terminal node during video transmission to compress the video and extract the features, and the bandwidth occupation is reduced).Therefore, the architecture of the system is an innovative architecture suitable for the IPTV field, which can not only be well compatible with the original IPTV architecture, but also provide a flexible function adjustment mechanism (flexibly expand or adjust auxiliary services through edge nodes), and can also serve as the basis for introducing the "digital retina" technology.
[0018] In order to improve the video quality of the IPTV video program in the form of relatively low bandwidth, different pre-encoding strategies are designed for different end-side node types: for each frame of the video program image: based on the program picture features corresponding to the frame of the video program image, the region of interest and the transition region are determined, the video program image is encoded with a first size of region block for low-precision encoding to form a basic quality block, the transition region is encoded with a second size of region block for medium-precision encoding to form an edge transition block, and the region of interest is encoded with a third size of region block for high-precision encoding to form a ROI enhancement block, wherein the transition region is located at the edge of the region of interest, the first size is greater than the second size, and the second size is greater than the third size; if the type of the end-side node is a set-top box, the basic quality block, the ROI enhancement block and the edge transition block are uniformly packaged to form a single video stream in the order of frames of the target video program; if the type of the end-side node is a smart terminal, the edge transition block and the ROI enhancement block are uniformly packaged to form a comprehensive enhancement block, and the basic quality block and the comprehensive enhancement block are hierarchically packaged to form a basic video stream and an enhancement video stream. Such a pre-encoding scheme can maintain high quality (high definition) video playback as much as possible under the condition of effectively reducing bandwidth, because there is usually a focal point area when users watch, and some important parts (such as people, objects, etc.) are enhanced with ROI, and the background is reduced in definition, which basically does not affect the user's perception. In order to eliminate the perception difference at the junction of the ROI and the background region (which may form a "jagged edge"), an edge transition block is also added for transition, effectively avoiding the perception difference caused by the "jagged edge" to give users a better viewing experience. Different encoding strategies are designed according to the characteristics of the decoding technology of different end-side nodes (set-top box or smart terminal). The network of the set-top box is stable and suitable for decoding of a single video stream, so a uniform packaging encoding scheme is adopted. The network of the smart terminal is usually unstable (may be good or bad at times), and is suitable for decoding of multiple video streams with hierarchical encoding, so the code stream distribution of the multiple video streams can be flexibly adjusted according to different network states, and part of the ROI enhancement block + edge transition block can be adaptively reserved to adapt to the network situation of the smart terminal.
[0019] The edge node can determine corresponding display information (such as target role information, target article information, target advertisement information, etc.) according to the type of the auxiliary service provided by the edge node after receiving the feature stream sent by the cloud node (the feature stream contains program picture features corresponding to each frame of video program image, each program picture feature has a corresponding scene label, and each program picture feature contains a plurality of types of subdivision features, the types of subdivision features including character features and article features), and sends the display information to the terminal side node. The terminal side node can decode the received display information, extract the timestamp and display information type identifier thereof, determine the corresponding target display interface based on the display information type identifier, and fill the display information into the target display interface, wherein the target display interface contains the position coordinates and interface size of the interface. The target display interface is synchronized with the video program image of the corresponding frame based on the timestamp of the display information, and the target display interface is superimposed on the specified area of the video program image in the form of a transparent layer based on the position coordinates and interface size, and pixel fusion is performed through an Alpha blending algorithm. The fused program picture is rendered and output. Accordingly, the edge node can provide the auxiliary service it needs to provide, and can flexibly interact (such as starting the service, stopping the service) according to the needs of the user, without the need to solidify the service in the pre-encoding stage (if the service is solidified in the pre-encoding stage, it is difficult to adjust and difficult to flexibly respond to the real-time needs of the user).
[0020] In order to make the above objectives, features and advantages of the present application more apparent, the following will describe a preferred embodiment in detail, and the accompanying drawings will be described as follows. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0022] Figure 1 The schematic diagram of the architecture of the video program playing system adapted to the digital retina architecture provided by the embodiments of the present application.
[0023] Figure 2 The interaction diagram of the nodes of the video program playing system adapted to the digital retina architecture.
[0024] Figure 3 The effect diagram of providing auxiliary service in the video program playing interface. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0026] Please refer to Figure 1 , Figure 1 The architecture schematic diagram of the video program playing system adapting to the digital retina architecture provided in the embodiments of the present application is shown in the following figure.
[0027] In the embodiments, the architecture design of the video program playing system adapting to the digital retina architecture is to include a cloud node, an edge node and an end node: the cloud node is regarded as a center node (including cloud storage, processing, distribution and the like); the edge node is a node providing auxiliary services (for example, intelligent person analysis, intelligent article analysis, intelligent advertisement delivery and the like), which can be a server deployed by the unit, or a server deployed by other external suppliers, and can be a same edge node providing multiple different functions, or multiple edge nodes providing one auxiliary service or multiple different auxiliary services; the end node is a device on the user side, and in the embodiments, two types of end nodes are considered: a set top box and a smart terminal (such as a smart phone, a tablet computer and the like), and the number of end nodes is large. Each end node is in communication connection with the cloud node, each edge node is in communication connection with the cloud node, and the connection between the edge node and the end node can be in communication connection according to needs (for example, some end nodes do not need to be in communication connection for some auxiliary services, and need to be in communication connection with the edge node providing the auxiliary service if the auxiliary service is opened).
[0028] In the existing IPTV architecture, video program contents are stored in the cloud node (stored by a special database server), and before the architecture is built, if some video program contents need to be analyzed, feature extraction needs to be performed. Therefore, under the architecture of the present system, feature extraction needs to be performed on the current IPTV video program (which can be offline extraction), and the feature extraction technology can adopt an existing feature extraction scheme or a specially developed feature extraction scheme, which is not limited here. Each frame of each video program (for example, a single movie or a single episode of a TV series) will extract corresponding features to form a program picture feature set associated with the video program, and through analysis, the scene to which each program picture feature belongs (for example, the scene is divided according to lens switching) and the subdivided features (person features and article features, such as face, body and the like) contained in each program picture feature will be determined. For those that do not contain person features and / or article features (for example, only background features), they do not need to be identified separately.
[0029] It should be noted that the IPTV video program is provided by a content source, most of which is a static content source (i.e. a video program uploaded after being shot, edited and reviewed), and there is also a live content source (provided by a live camera), and the "digital retina" architecture is proposed, which provides a video stream and a feature stream by an end-side node (such as a live camera providing a content source in a live scene) providing content. A lightweight feature extraction model is deployed at the end-side node to reduce data transmission and reduce the burden of the cloud node, because the cloud node does not need to perform feature extraction. Therefore, the architecture of the video program playing system provided by the embodiment adapting to the digital retina architecture can provide architectural support for accessing the digital retina technology.
[0030] In addition, the cloud node also needs to precode the video program and the associated program picture feature set, and the precoding strategy is designed as follows:
[0031] First, the corresponding relationship between the video program image in the target video program and the program picture feature in the program picture feature set is determined, and then for each frame of video program image: based on the program picture feature corresponding to this frame of video program image, the region of interest and the transition region (the transition region is located at the edge of the region of interest, and is mainly used to smooth the difference in definition between the region of interest and the background region, to avoid "jagged edges", the program picture feature can guide the determination of the region of interest, and the transition region depends on the determination of the region of interest, for example, the boundary of the region of interest is used as a limit to expand a certain size to obtain an edge region) are determined. Specifically, the video program image is low-precision coded by region block of a first size (such as 32*32, for the boundary part, an asymmetric size can be used, and the existing technology also does this for the coding of this boundary part, which will not be specifically described here) to form a basic quality block; the transition region is medium-precision coded by region block of a second size (such as 16*16) to form an edge transition block; and the region of interest is high-precision coded by region block of a third size (such as 8*8) to form a ROI (Region of Interest, region of interest) enhancement block. In the embodiment, the first size is greater than the second size, and the second size is greater than the third size, and this size relationship needs to be met.
[0032] For the case that the type of the end-side node is a set-top box, single video stream encoding is needed, therefore, the cloud node needs to uniformly encapsulate the basic quality block, the ROI enhancement block and the edge transition block, and form a single video stream according to the frame order of the target video program. This scheme can support existing coding protocols such as H.265 / HEVC, AVS3 and the like, allow using variable block sizes in the same video stream, and mark the block type (basic / edge / ROI) through SEI (Supplemental Enhancement Information) messages, so that the decoder can process the blocks in different areas according to the SEI information; the marking method of AVS3 is different from that of H.265, and AVS3_user_data can be used to customize the SEI syntax. In order to adapt to different versions of set-top boxes, different encapsulation modes can also be determined: a first encapsulation mode and a second encapsulation mode, the first encapsulation mode represents encapsulating the basic quality block (for example, an old version of set-top box only supports H.264, so a single video stream based on the pre-encoding of the basic quality block can be loaded), and the second encapsulation mode represents uniformly encapsulating the basic quality block, the ROI enhancement block and the edge transition block.
[0033] For the case that the type of the end-side node is a smart terminal, layered encoding can be performed, in order to ensure the quality of the region of interest, the edge transition block and the ROI enhancement block are uniformly encapsulated to form a comprehensive enhancement block, and then the basic quality block and the comprehensive enhancement block are layered encapsulated to form a basic video stream and an enhancement video stream according to the frame order of the target video program. This scheme can also support existing coding protocols such as H.265 / HEVC, AVS3 and the like, and the SHVC (scalable encoding) hierarchy of H.265 can realize layered encapsulation.
[0034] In order to adapt to the changing network conditions of intelligent terminals, such as good network state, moderate network state, poor network state (this standard can be defined according to the actual network test situation to define the network condition level of the specific parameter, which is only the current network state classification, and other classification schemes can also be used, which is not limited here) and different network states, multiple precoding can be performed: such as precoding the basic quality block to form a basic video stream (suitable for poor network state); precoding the basic quality block to form a basic video stream, and uniformly packaging the edge transition block and the ROI enhancement block to form a comprehensive enhancement block, and then precoding all the comprehensive enhancement blocks to form an enhanced video stream (suitable for good network state); and precoding the basic quality block to form a basic video stream, and uniformly packaging the edge transition block and the ROI enhancement block to form a comprehensive enhancement block, and then precoding part of the comprehensive enhancement blocks (such as retaining no more than 5 or no more than 10 comprehensive enhancement blocks for each frame of video program image) to form an enhanced video stream (suitable for moderate network state).
[0035] Specifically, it can be a division and packaging mode and an ROI ratio mode, wherein the packaging mode can include a third packaging mode and a fourth packaging mode, the third packaging mode indicating that only the basic quality block is packaged, and the fourth packaging mode indicating that the edge transition block and the ROI enhancement block are uniformly packaged to form a comprehensive enhancement block, and then the basic quality block and the comprehensive enhancement block are hierarchically packaged, and the ROI ratio mode includes an ROI quantitative mode and an ROI full-quantity mode, the ROI quantitative mode indicating that no more than a set number of ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of video program image are retained, and the ROI full-quantity mode indicating that all ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of video program image are retained.
[0036] The packaging mode (first packaging mode, second packaging mode, third packaging mode, fourth packaging mode) and the ROI ratio mode (ROI quantitative mode and ROI full-quantity mode) mentioned above are precoded in each case to form a precoded video stream, so that real-time encoding is not required when it is delivered, resource occupation is saved, and efficiency is improved.
[0037] The precoding of the feature stream is simple, and the program picture features are encoded according to the frame order corresponding to the video program to form a feature stream, which is sent to the corresponding edge node when needed to provide auxiliary services by the edge node.
[0038] The pre-encoding scheme can effectively reduce the bandwidth while maintaining high quality (high definition) video playback as much as possible, because when users watch, there is usually a focal point area, and some important parts (such as characters, objects, etc.) are enhanced by ROI, and the background is reduced. The requirement for reducing the definition of the background basically does not affect the user's visual perception, and in order to eliminate the visual difference at the junction of the ROI and the background area (which may form a "jagged edge"), an edge transition block is also added for transition, effectively avoiding the visual difference caused by the "jagged edge" to give users a better viewing experience. Different encoding strategies are designed according to the characteristics of the decoding technology of different end node types (set-top boxes or smart terminals). The set-top box network is stable and suitable for single video stream decoding, so a unified encapsulated encoding scheme is adopted. The smart terminal network is usually unstable (may be good or bad at times), and is suitable for layered encoding of multiple video streams, so the code stream allocation of multiple video streams can be flexibly adjusted according to different network states, and part of the ROI enhancement block + edge transition block can be adaptively reserved to adapt to the network situation of the smart terminal.
[0039] Moreover, the scheme can well consider different end node types and support existing protocols, and for old set-top boxes that do not support H.265, the unit can also use various commercial promotion schemes to gradually update old set-top boxes to gradually improve the architecture deployment of the video program playback system adapted to the digital retina architecture.
[0040] Therefore, the architecture design of the video program playing system adapting to the digital retina architecture can utilize the pre-encoding technology and the feature extraction technology, can flexibly adjust and expand the auxiliary services (and can also be well compatible with the auxiliary services provided by external suppliers, the external suppliers can join the architecture in the form of edge nodes, and flexibly expand the auxiliary services), and the architecture design is well compatible with the existing IPTV architecture, and can be completed by making slight modifications to the existing architecture (for old set-top boxes, the upgrade can be realized to support the playing of higher quality videos, and if not upgraded, the original architecture can be used for playing video programs, that is, high bandwidth occupation is used to watch high-definition video programs), and the architecture can also provide architecture support for introducing the "digital retina" technology (the digital retina technology is an idea proposed in recent years, which compresses the video and extracts the features for the transmission scheme of "video stream" + "feature stream" during the transmission of the video, and reduces the bandwidth occupation), for example, in a live scene, the "video stream" and "feature stream" transmitted from the end-side node to the cloud node can be synchronously obtained, without additional feature extraction by the cloud. Therefore, the architecture of the system is an innovative architecture suitable for the IPTV field, which can not only be well compatible with the original IPTV architecture, but also provide a flexible function adjustment mechanism (flexibly expand or adjust the auxiliary services through the edge nodes), and can also serve as the basis for introducing the "digital retina" technology.
[0041] In order to facilitate the understanding of the video program playing system adapting to the digital retina architecture, the following will be introduced in combination with the accompanying drawings. Figure 2 The interaction process of each node in the video program playing system adapting to the digital retina architecture will be introduced. Please refer to Figure 2 , Figure 2 The interaction diagram of each node in the video program playing system adapting to the digital retina architecture.
[0042] Firstly, the end-side node can initiate a video program playing request. In the video program playing request, the video program number to be played, the type of the end-side node (for example, distinguished by a type identifier), and the version information of the set-top box when the type of the end-side node is the set-top box, and the current network status when the type of the end-side node is the intelligent terminal (and the network status will also be obtained in real time during the subsequent video stream transmission, so as to dynamically adjust the pre-encoded video stream).
[0043] After receiving the video program playing request, the cloud node can analyze the request to determine the target video program and the associated program picture feature set, and transmit the video stream encoded based on the target video program and the associated program picture feature set to the end-side node, and transmit the feature stream encoded based on the program picture feature set to the edge node.
[0044] Exemplarily, the cloud-side node can determine the type of the end-side node, the video program number, and the network status of the intelligent terminal or the version information of the set-top box based on the video program playing request, and then determine the target video program (and the associated program picture feature set) based on the video program number in the video program playing request, and determine the video stream encoded based on the target video program and the associated program picture feature set by using the target encoding strategy based on the type of the end-side node.
[0045] Specifically, when the type of the end-side node is the set-top box, the cloud-side node can determine the version information of the set-top box from the video program playing request, and then determine the target packaging mode matched by the set-top box based on the version information of the set-top box (for example, if the set-top box version supports H.265, the second packaging mode is determined as the target packaging mode, which means that the basic quality block, the ROI enhancement block, and the edge transition block are uniformly packaged; if the set-top box version does not support H.265, the first packaging mode is determined as the target packaging mode, which means that the basic quality block is packaged), and then determine the single video stream packaged based on the target packaging mode as the video stream delivered to the set-top box.
[0046] Specifically, when the type of the end-side node is the intelligent terminal, the cloud-side node can determine the terminal network status (such as good network status, moderate network status, poor network status, etc.) carried in the video program playing request, and then determine the target packaging mode and the ROI ratio mode matched by the intelligent terminal based on the terminal network status. For example, when the network status is good, the fourth packaging mode is determined as the target packaging mode, and the ROI full-amount mode is determined as the ROI ratio mode, and then the basic video stream and the enhancement video stream formed by packaging based on the target packaging mode (the fourth packaging mode) and the ROI full-amount mode can be determined as the video stream delivered to the intelligent terminal. For another example, when the network status is moderate, the cloud-side node can determine the fourth packaging mode as the target packaging mode, and determine the ROI quantitative mode as the ROI ratio mode, and then the basic video stream and the enhancement video stream formed by packaging based on the target packaging mode (the fourth packaging mode) and the ROI quantitative mode can be determined as the video stream delivered to the intelligent terminal. For another example, when the network status is poor, the cloud-side node can determine the third packaging mode as the target packaging mode, and then the basic video stream formed by packaging based on the target packaging mode (the third packaging mode) can be determined as the video stream delivered to the intelligent terminal.
[0047] At the same time, the cloud-side node also needs to determine the program picture feature set associated with the target video program, so as to determine the feature stream encoded based on the program picture feature set and transmitted to the edge node.
[0048] The edge node can receive the feature stream sent by the cloud node and perform feature analysis, determine the display information according to the feature analysis result and the auxiliary service type provided by the edge node, and send the display information to the terminal side node. The auxiliary service type includes at least one of intelligent role analysis, intelligent article analysis, and intelligent advertisement placement.
[0049] For example, the edge node can receive the feature stream sent by the cloud node, the feature stream contains program picture features corresponding to each frame of video program image, each program picture feature has a corresponding scene label, and each program picture feature contains a plurality of types of sub-features, including character features and article features.
[0050] If the auxiliary service type provided by the edge node is intelligent role analysis, the edge node can determine a target role based on the character features, and determine target role information as the display information based on the target role.
[0051] If the auxiliary service type provided by the edge node is intelligent article analysis, the edge node can determine a target article based on the article features, and determine target article information as the display information based on the target article.
[0052] If the auxiliary service type provided by the edge node is intelligent advertisement placement, the edge node can determine a target article based on the article features, and determine corresponding target advertisement information (which may not be unique advertisement information, and may involve priority recommendation) based on the target article, and send the target advertisement information as the display information.
[0053] Since such auxiliary services have high real-time requirements, the feature analysis model (or feature matching model) deployed on the edge node is preferably a lightweight model with short processing time. For the program picture features of continuous multiple frames of video program images under the same scene (same scene label), each sub-feature can use the display information determined based on the sub-feature under the same scene label (if the scene label changes, the display information needs to be determined again, and the interval scene also needs to be determined again, and the previously determined display information cannot be used). The display information can also determine the number of continuous frames to reduce the amount of data to be transmitted, but in this case, the strict correspondence between the number of continuous frames of the display information and the number of frames of the video program image needs to be designed to avoid errors.
[0054] After the edge node determines the display information, the edge node can send the determined display information (with a timestamp and a display information type identifier) to the terminal side node.
[0055] The end-side node can receive the video stream transmitted by the cloud-side node and decode the video program image, and receive the display information sent by the edge node and decode the display information, and perform synchronous and mixed rendering based on the decoded video program image and the display information to obtain a program picture and play the program picture on the end-side node (such as a smart terminal) or a playing device (such as a television connected with a set-top box) connected with the end-side node.
[0056] When the end-side node is a set-top box, the end-side node can receive a single video stream transmitted by the cloud-side node, and then, for the encapsulated data of each frame of the video program image in the single video stream: decapsulate the encapsulated data of the frame of the video program image, extract the encoded data and the meta information (including the block type identifier, the block position coordinates, and the block size) of each block; then decode the encoded data of the basic quality block to generate a full-image low-precision video frame, decode the encoded data of the ROI enhancement block, and cover the ROI enhancement block to the region of interest of the full-image low-precision video frame based on the block position coordinates and the block size of the ROI enhancement block, and decode the encoded data of the edge transition block, and cover the edge transition block to the transition region of the full-image low-precision video frame based on the block position coordinates and the block size of the edge transition block, and then perform filtering and color space conversion to obtain the video program image.
[0057] It should be noted that the manner in which the ROI enhancement block is covered to the region of interest of the full-image low-precision video frame and the manner in which the edge transition block is covered to the transition region of the full-image low-precision video frame can be incremental superimposition or replacement coverage. If the incremental superimposition scheme is adopted, the encoding of the region of interest and the encoding of the edge region need to be designed as an incremental encoding scheme during pre-encoding.
[0058] In addition, for the case of a set-top box compatible with an old version (if the old version set-top box does not support the H.265 protocol), the set-top box can directly decode a single video stream (only containing the basic quality block, and the full-image low-precision video frame can be obtained by splicing), and does not need to perform the coverage of the ROI enhancement block and the edge transition block to the full-image low-precision video frame. The full-image low-precision video frame obtained is subjected to color space conversion to obtain the video program image.
[0059] If the type of the end-side node is a smart terminal, the end-side node can receive the base video stream and the enhanced video stream transmitted by the cloud-side node. For the encapsulated data of each frame of the video program image in the base video stream, the end-side node can decapsulate the encapsulated data of the frame of the video program image to generate a full-image low-precision video frame after decoding; for the encapsulated data of the corresponding frame of the video program image in the enhanced video stream, the end-side node can decapsulate the encapsulated data of the frame of the video program image to extract the encoded data of the comprehensive enhancement block; the comprehensive enhancement block is decapsulated again to extract the encoded data and meta information of the ROI enhancement block and the edge transition block, wherein the meta information includes block type identification, block position coordinates and block size. The encoded data of the ROI enhancement block is decoded again, and is overlaid to the region of interest of the full-image low-precision video frame of the corresponding frame of the video program image based on the block position coordinates and the block size of the ROI enhancement block, and the encoded data of the edge transition block is decoded and overlaid to the transition region of the full-image low-precision video frame of the corresponding frame of the video program image based on the block position coordinates and the block size of the edge transition block, to obtain the video program image after filtering and color space conversion.
[0060] Similarly, the way in which the ROI enhancement block is overlaid to the region of interest of the full-image low-precision video frame and the edge transition block is overlaid to the transition region of the full-image low-precision video frame can be incremental superimposition or replacement overlay. If the incremental superimposition scheme is adopted, the encoding of the region of interest and the encoding of the edge region need to be designed as an incremental encoding scheme during pre-encoding. For a smart terminal, an incremental encoding scheme is usually adopted. Since such an incremental encoding scheme is mature in existing technologies, the process thereof will not be described herein.
[0061] In addition, for the case where the network state is good or the network state is moderate, the processing process is similar (only the number of comprehensive enhancement blocks is different), and for the case where the network state is poor, the video stream decoded by the smart terminal is only the base video stream, and there is no need to overlay the ROI enhancement block and the edge transition block to the full-image low-precision video frame, and the full-image low-precision video frame obtained after the overlaying is subjected to color space conversion to obtain the video program image.
[0062] After obtaining the video program image (the decoding of the display information is actually performed synchronously with the decoding of the video stream after obtaining the video program image), the end-side node can further decode the received display information to extract the timestamp and the display information type identification. Then, the end-side node can determine the corresponding target display interface based on the display information type identification, and fill the display information (such as text information, advertisement images, etc.) into the target display interface, wherein the target display interface includes the position coordinates and the interface size of the interface.
[0063] Afterwards, the end-side node can synchronize the target display interface with the video program image of the corresponding frame based on the timestamp of the display information; superimpose the target display interface in the form of a transparent layer to the specified area of the video program image according to the position coordinates and interface size, and perform pixel fusion through an Alpha blending algorithm, so as to render and output the fused program picture, and realize the playing of the video program. Figure 3 As shown in FIG. 13, it is a schematic diagram of the effect of providing auxiliary services. This is only a schematic diagram, and as the auxiliary services develop and the edge node deployment is improved, and the service providers participating in the architecture to provide auxiliary services increase and the technology develops, it is expected to provide very effective and fine auxiliary services, in addition to role information and item information, even to identify prop information and character moves in the play, to introduce more comprehensively and to make the program more interactive. The schematic diagram is not considered as a limitation to the present application.
[0064] For the user side, a certain auxiliary service can be started or stopped according to the needs, and through the operation instruction of starting or stopping, feedback to the cloud node (which can be fed back to the cloud node through the edge node, or fed back to the cloud node and then transmitted to the edge node), to realize the stop of the feature flow and the closing of the display information. Of course, for the existing advertising placement strategy, which is not started and stopped by the user, in order to comply with the relevant regulations of advertising placement, it is necessary to design the placement strategy and the placement time length, and how to deal with it is not described here. Of course, this situation should not be considered as a limitation to the present application.
[0065] In summary, the edge node can provide the auxiliary services it needs to provide, and can flexibly interact according to the needs of the user (for example, starting the service, stopping the service), without being fixed in the pre-encoding stage (if the service is fixed in the pre-encoding stage, it is difficult to adjust and difficult to flexibly respond to the real-time needs of the user).
[0066] Based on the same inventive concept, the embodiments of the present application also provide a video program playing method adapted to the digital retina architecture, which is applied to the video program playing system adapted to the digital retina architecture, and the method comprises the following steps:
[0067] Step S10: initiating a video program playing request through the end-side node.
[0068] Step S20: determining a target video program and an associated program picture feature set based on the video program playing request through the cloud node, and transmitting a video stream encoded based on the target video program and the associated program picture feature set to the end-side node, and transmitting a feature stream encoded based on the program picture feature set to the edge node, wherein the program picture feature is obtained by feature extraction on the corresponding video program image in the target video program.
[0069] Step S30: receiving the feature stream sent by the cloud node through the edge node and performing feature analysis, determining the display information according to the feature analysis result and the auxiliary service type provided by itself, and sending to the end-side node, wherein the auxiliary service type includes at least one of intelligent role analysis, intelligent article analysis and intelligent advertisement placement.
[0070] Step S40: receiving the video stream transmitted by the cloud node through the end-side node and decoding to obtain the video program image, and receiving the display information sent by the edge node and decoding, synchronizing and mixing rendering based on the decoded video program image and the display information, obtaining the program picture and playing on the end-side node or the playing device connected to the end-side node.
[0071] The specific processing process of the video program playing method adapted to the digital retina architecture can be referred to in the foregoing, and will not be described here.
[0072] In summary, the embodiment of the present application provides a video program playing system and method adapting to digital retina architecture, the video program playing system adapting to digital retina architecture provided by the present solution builds the architecture of cloud node-edge node (several)-end side node, and initiates the video program playing request through the end side node. Since the IPTV video program (non-live scene) is in the cloud, the pre-extraction of features can be performed, the corresponding features of each frame image of each video program can be extracted, the program picture feature set associated with the video program is formed and stored, and whether to perform pre-encoding is selected according to the need (pre-encoding can save resource occupation during distribution, and without pre-encoding, the feature encoding strategy can be flexibly adjusted), and the video program can be pre-encoded to improve the distribution efficiency. Therefore, the cloud node can determine the target video program and the associated program picture feature set (the program picture feature is obtained by performing feature extraction on the corresponding video program image in the target video program) based on the video program playing request, and transmit the video stream encoded based on the target video program and the associated program picture feature set to the end side node, and transmit the feature stream encoded based on the program picture feature set to the edge node. The edge node can receive the feature stream sent by the cloud node and perform feature analysis, determine the display information according to the feature analysis result and the auxiliary service type (such as intelligent character analysis, intelligent article analysis, intelligent advertisement placement, etc.) provided by the edge node, and send the display information to the end side node. The end side node can receive the video stream transmitted by the cloud node and decode to obtain the video program image, receive the display information sent by the edge node and decode, and perform synchronous and mixed rendering based on the decoded video program image and the display information to obtain the program picture and play the program picture on the end side node or the playing device connected to the end side node. This architecture can utilize the combination of pre-encoding technology and feature extraction, flexibly adjust and expand auxiliary services (it can also well compatible with the auxiliary services provided by external suppliers, the external suppliers can join the present architecture in the form of edge node to flexibly expand auxiliary services), and the present architecture design has good compatibility with the existing IPTV architecture, and can be completed by slight modification of the existing architecture (for old set-top boxes, the upgrade can be realized to support the playing of higher quality video, if not upgraded, the original architecture can be used for video program playing, that is, high bandwidth occupation is used to watch high definition video programs), and the present architecture can also provide architecture support for introducing the "digital retina" technology (the digital retina technology is an idea proposed in recent years, the "video stream" + "feature stream" transmission scheme of compressing video and extracting features during video transmission can be realized by deploying a light weight model in the end side node to reduce bandwidth occupation).Therefore, the architecture of the system is an innovative architecture suitable for the IPTV field, which can not only be compatible with the original IPTV architecture, but also provide a flexible function adjustment mechanism (flexibly expand or adjust auxiliary services through edge nodes), and can also serve as the basis for introducing the "digital retina" technology.
[0073] In order to improve the video quality of the IPTV video program in the form of relatively low bandwidth, different pre-encoding strategies are designed for different end-side node types: for each frame of the video program image: based on the program picture features corresponding to the frame of the video program image, the region of interest and the transition region are determined, the video program image is encoded with a first size of region block for low-precision encoding to form a basic quality block, the transition region is encoded with a second size of region block for medium-precision encoding to form an edge transition block, and the region of interest is encoded with a third size of region block for high-precision encoding to form a ROI enhancement block, wherein the transition region is located at the edge of the region of interest, the first size is greater than the second size, and the second size is greater than the third size; if the type of the end-side node is a set-top box, the basic quality block, the ROI enhancement block and the edge transition block are uniformly packaged to form a single video stream according to the frame order of the target video program; if the type of the end-side node is a smart terminal, the edge transition block and the ROI enhancement block are uniformly packaged to form a comprehensive enhancement block, and the basic quality block and the comprehensive enhancement block are hierarchically packaged to form a basic video stream and an enhancement video stream. Such a pre-encoding scheme can maintain high quality (high definition) video playback as much as possible while effectively reducing bandwidth, because when users watch, there is usually a focal point area, and for some important parts (such as people, objects, etc.), ROI enhancement is used, and for the background, the requirement is reduced, and the background with reduced definition basically does not affect the user's perception. And in order to eliminate the perception difference at the junction of the ROI and the background region (which may form a "jagged edge"), an edge transition block is added for transition, effectively avoiding the perception difference caused by the "jagged edge" to give users a better viewing experience. And according to the characteristics of the decoding technology of different end-side nodes (set-top box or smart terminal), different encoding strategies are designed, the network of the set-top box is stable, and it is suitable for decoding of a single video stream, so the uniform packaging encoding scheme is adopted, while the network of the smart terminal is usually unstable (may be good or bad at times), and is suitable for decoding of multiple video streams of hierarchical encoding, so the code stream allocation of the multiple video streams can be flexibly adjusted according to different network states, and part of the ROI enhancement block + edge transition block can be adaptively reserved to adapt to the network situation of the smart terminal.
[0074] After receiving the feature stream sent by the cloud node (the feature stream contains program picture features corresponding to each frame of video program image, each program picture feature has a corresponding scene label, and each program picture feature contains several types of subdivision features, and the types of subdivision features include character features and article features), the edge node can determine the corresponding display information (such as target role information, target article information, target advertisement information, etc.) according to the type of auxiliary service provided by the edge node, and send it to the end-side node. The end-side node can decode the received display information, extract its timestamp and display information type identifier; determine the corresponding target display interface based on the display information type identifier, and fill the display information into the target display interface, wherein the target display interface includes the position coordinates and interface size of this interface; synchronize the target display interface with the corresponding frame of video program image based on the timestamp of the display information; superimpose the target display interface in the form of a transparent layer to the specified area of the video program image according to the position coordinates and interface size, and perform pixel fusion through the Alpha blending algorithm; and render and output the fused program picture. Accordingly, the edge node can provide the auxiliary service it needs to provide, and can interact flexibly according to the needs of the user (for example, starting the service, closing the service), without the need to solidify the service in the pre-encoding stage (if the service is solidified in the pre-encoding stage, it is difficult to adjust and respond to the real-time needs of the user flexibly).
[0075] The above only describes the embodiments of the present application and is not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A video program playback system adapted to a digital retina architecture, characterized in that, Including cloud nodes, edge nodes, and endpoint nodes. The endpoint node is used to initiate a video program playback request; The cloud node is used to determine the target video program and the associated program screen feature set based on the video program playback request, and transmit the video stream encoded based on the target video program and the associated program screen feature set to the end node, and transmit the feature stream encoded based on the program screen feature set to the edge node, wherein the program screen feature is obtained by feature extraction of the corresponding video program image in the target video program. The edge node is used to receive the feature stream sent by the cloud node and perform feature analysis. Based on the feature analysis results and the auxiliary service type it provides, it determines the display information and sends it to the end-side node. The auxiliary service type includes at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising. The edge node is used to receive the video stream transmitted by the cloud node and decode it to obtain a video program image, and to receive and decode the display information sent by the edge node, and to perform synchronous and mixed rendering based on the decoded video program image and display information to obtain a program picture and play it on the edge node or a playback device connected to the edge node.
2. The video program playback system adapted to digital retina architecture according to claim 1, characterized in that, The cloud node is specifically used for: Based on the video program playback request, the type of the end-side node is determined, wherein the type of the end-side node includes set-top boxes and smart terminals; Based on the video program playback request, the target video program and the associated program screen feature set are determined; Based on the type of the end-side node, a video stream is determined that is encoded using a target coding strategy based on the target video program and the associated program screen feature set; The feature stream is determined and encoded based on the program image feature set associated with the target video program; The video stream is transmitted to the end-side node, and the feature stream is transmitted to the edge node.
3. The video program playback system adapted to digital retina architecture according to claim 2, characterized in that, The target encoding strategy is: Determine the correspondence between video program images in the target video program and program screen features in the program screen feature set; For each frame of video program image: Based on the program picture features corresponding to this frame of video program image, the region of interest and transition region are determined. The video program image is encoded with low precision using region blocks of the first size to form a basic quality block. The transition region is encoded with medium precision using region blocks of the second size to form an edge transition block. The region of interest is encoded with high precision using region blocks of the third size to form an ROI enhancement block. The transition region is located at the edge of the region of interest. The first size is larger than the second size, and the second size is larger than the third size. The ROI represents the region of interest. If the type of the end-side node is a set-top box, the basic quality block, ROI enhancement block and edge transition block are uniformly encapsulated to form a single video stream according to the frame order of the target video program; If the type of the end-side node is a smart terminal, the edge transition block and ROI enhancement block are uniformly encapsulated to form a comprehensive enhancement block, and the basic quality block and comprehensive enhancement block are layered and encapsulated to form a basic video stream and an enhanced video stream according to the frame order of the target video program.
4. The video program playback system adapted to digital retina architecture according to claim 3, characterized in that, If the type of the terminal node is a set-top box, the terminal node is specifically used for: Receive a single video stream transmitted from the cloud node; For the encapsulated data of each frame of video program image in a single video stream: decapsulate the encapsulated data of this frame of video program image, and extract the encoded data and metadata of each block. The metadata includes block type identifier, block position coordinates, and block size. The encoded data of the basic quality block is decoded to generate a full-image low-precision video frame. The encoded data of the ROI enhancement block is decoded and covered to the region of interest of the full-image low-precision video frame based on the block position coordinates and block size of the ROI enhancement block. The encoded data of the edge transition block is decoded and covered to the transition region of the full-image low-precision video frame based on the block position coordinates and block size of the edge transition block. Filtering and color space conversion are then performed to obtain the video program image.
5. The video program playback system adapted to digital retina architecture according to claim 3, characterized in that, If the type of the endpoint node is a smart terminal, the endpoint node is specifically used for: Receive the base video stream and enhanced video stream transmitted by the cloud node; For the encapsulated data of each frame of video program image in the basic video stream: decapsulate the encapsulated data of this frame of video program image, and generate a full-image low-precision video frame after decoding. For the encapsulated data of the corresponding frame video program image in the enhanced video stream: decapsulate the encapsulated data of this frame video program image and extract the encoded data of the integrated enhancement block; The integrated enhancement block is decapsulated twice to extract the encoded data and metadata of the ROI enhancement block and edge transition block. The metadata includes block type identifier, block position coordinates, and block size. The encoded data of the ROI enhancement block is decoded, and the ROI enhancement block is overlaid onto the region of interest of the full-image low-precision video frame of the corresponding frame video program image based on the block position coordinates and block size of the ROI enhancement block. The encoded data of the edge transition block is decoded, and the edge transition block is overlaid onto the transition region of the full-image low-precision video frame of the corresponding frame video program image based on the block position coordinates and block size of the edge transition block. Filtering and color space conversion are then performed to obtain the video program image.
6. The video program playback system adapted to digital retina architecture according to claim 3, characterized in that, If the type of the end-side node is a set-top box, the cloud node is also used for: The version information of the set-top box is determined from the video program playback request; Based on the version information of the set-top box, the target packaging mode matching the set-top box is determined. The target packaging mode includes a first packaging mode and a second packaging mode. The first packaging mode represents packaging the basic mass block, and the second packaging mode represents uniformly packaging the basic mass block, ROI enhancement block, and edge transition block. A single video stream encapsulated based on the target encapsulation mode is determined as the video stream to be sent to the set-top box.
7. The video program playback system adapted to digital retina architecture according to claim 3, characterized in that, If the type of the end-side node is a smart terminal, the cloud node is also used for: The terminal network status carried in the video program playback request is determined; Based on the terminal network status, the target encapsulation mode and ROI ratio mode matching the smart terminal are determined. The target encapsulation mode includes a third encapsulation mode and a fourth encapsulation mode. The third encapsulation mode indicates that only the basic quality block is encapsulated. The fourth encapsulation mode indicates that the edge transition block and ROI enhancement block are uniformly encapsulated to form a comprehensive enhancement block. Then, the basic quality block and the comprehensive enhancement block are encapsulated in layers. The ROI ratio mode includes a quantitative ROI mode and a full ROI mode. The quantitative ROI mode indicates that no more than a set number of ROI enhancement blocks and corresponding edge transition blocks are retained for each frame of video program image. The full ROI mode indicates that all ROI enhancement blocks and corresponding edge transition blocks are retained for each frame of video program image. Encapsulation is performed based on the target encapsulation mode and ROI ratio mode to form a basic video stream as the video stream sent to the smart terminal, or a basic video stream and an enhanced video stream are formed as the video stream sent to the smart terminal.
8. The video program playback system adapted to digital retina architecture according to claim 1, characterized in that, The edge node is specifically used for: The feature stream sent by the cloud node is received. The feature stream contains program screen features corresponding to each frame of video program image. Each program screen feature has a corresponding scene tag and each program screen feature contains several types of sub-features. The types of sub-features include human features and object features. If the auxiliary service provided by the edge node is intelligent role analysis, the target role is determined based on the character's characteristics, and the target role information is determined as the display information based on the target role. If the auxiliary service provided by the edge node is intelligent item analysis, the target item is determined based on the item characteristics, and the target item information is determined as the display information based on the target item. If the auxiliary service provided by the edge node is intelligent advertising, the target item is determined based on the item's characteristics, the corresponding target advertising information is determined based on the target item, and the target advertising information is used as the display information.
9. The video program playback system adapted to digital retina architecture according to claim 8, characterized in that, After obtaining the video program image, the end-side node is also used for: The received display information is decoded, and its timestamp and display information type identifier are extracted. The corresponding target display interface is determined based on the display information type identifier, and the display information is filled into the target display interface. The target display interface includes the position coordinates and interface size of this interface. Synchronize the target display interface with the corresponding frame of video program image based on the timestamp of the displayed information; Based on the location coordinates and interface size, the target display interface is overlaid as a transparent layer onto a specified area of the video program image, and pixel fusion is performed using an Alpha blending algorithm. The merged program footage will be rendered and output.
10. A video program playback method adapted to a digital retina architecture applied to a video program playback system adapted to a digital retina architecture as described in any one of claims 1-9, characterized in that, The method includes: Initiate a video program playback request through the aforementioned end-side node; Based on the video program playback request, the cloud node determines the target video program and the associated program screen feature set, and transmits the video stream encoded based on the target video program and the associated program screen feature set to the end node, and transmits the feature stream encoded based on the program screen feature set to the edge node, wherein the program screen feature is obtained by feature extraction of the corresponding video program image in the target video program. The edge node receives the feature stream sent by the cloud node and performs feature analysis. Based on the feature analysis results and the auxiliary service types it provides, it determines the display information and sends it to the edge node. The auxiliary service types include at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising. The edge node receives the video stream transmitted by the cloud node and decodes it to obtain a video program image. It also receives and decodes the display information sent by the edge node. Based on the decoded video program image and display information, it performs synchronous and mixed rendering to obtain a program picture and plays it on the edge node or a playback device connected to the edge node.
Citation Information
Patent Citations
Synchronous transmission control method for digital retina video stream and characteristic stream
CN110719438A
Telescopic visual computing system
CN112804188A