Video program playing system and method adaptive to digital retina architecture
Through the end-side node-edge node-cloud node architecture, combined with precoding and feature extraction, the problem of insufficient video quality and functional scalability in IPTV technology is solved, and high-quality video playback and flexible auxiliary services are realized under low bandwidth, adapting to different network conditions, and supporting "digital retina" technology.
Patent Information
- Application Number
- CN202510923085.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The existing IPTV technology has shortcomings in video quality and functional scalability. High encoding accuracy leads to high bandwidth usage, and it is difficult to dynamically adjust the solidification of functional modules, which cannot meet the user's real-time interaction needs.
The architecture of end-side node-edge node-cloud node is adopted, combined with precoding technology and feature extraction, video and feature streams are determined through cloud nodes. Edge nodes provide auxiliary services, and end-side nodes perform synchronous and hybrid rendering, supporting flexible adjustment and amplification of auxiliary services.
Improve video quality at lower bandwidth, provide flexible auxiliary services, is compatible with existing IPTV architecture, supports "digital retina" technology, adapts to different network conditions, and improves user experience.
Smart Images

Figure CN120475205A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of IPTV technology, and in particular to a video program playback system and method adapted to a digital retina architecture. Background Art
[0002] Internet Protocol Television (IPTV) is a technology for transmitting video content over IP networks. It digitizes television signals and encapsulates them into IP data packets to enable on-demand, live broadcast, and interactive services for video programs.
[0003] Traditional IPTV technology primarily uses encoding standards such as H.264 / AVC and H.265 / HEVC. This reduces bandwidth requirements through compression, pre-encoding video content in the cloud and pushing it directly to the device. Video quality depends on encoding accuracy: higher encoding accuracy results in higher video quality, but also increases bandwidth usage. Furthermore, functional modules (such as advertising) are embedded within the encoding process, making them difficult to dynamically expand or adjust. Adding new features (such as intelligent analysis of characters and objects in a program) is inherently difficult to implement due to the inflexibility of existing solutions, which prevent the flexible activation and deactivation of these ancillary services and hinder the real-time interaction needs of users. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a video program playback system and method adapted to the digital retina architecture, so as to improve the video quality of IPTV video programs in a relatively low bandwidth form by building an architecture of terminal node-edge node-cloud node and utilizing a combination of precoding technology and feature extraction, and to flexibly adjust and expand auxiliary services, have good compatibility with the existing IPTV architecture, and provide architectural support for the introduction of "digital retina" technology.
[0005] In order to achieve the above objectives, the embodiments of the present application are implemented in the following manner:
[0006] In the first aspect, an embodiment of the present application provides a video program playback system adapted to a digital retina architecture, including a cloud node, an edge node, and an end-side node, wherein the end-side node is used to initiate a video program playback request; the cloud node is used to determine a target video program and an associated program screen feature set based on the video program playback request, and transmit a video stream encoded based on the target video program and the associated program screen feature set to the end-side node, and transmit a feature stream encoded based on the program screen feature set to the edge node, wherein the program screen feature is obtained by extracting features from the corresponding video program image in the target video program; the edge node is used to determine a target video program and an associated program screen feature set based on the video program playback request ... The edge node is used to receive the feature stream sent by the cloud node and perform feature analysis, determine the display information based on the feature analysis results and the auxiliary service type provided by itself, and send it to the end-side node, wherein the auxiliary service type includes at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising delivery; the end-side node is used to receive the video stream transmitted by the cloud node and decode it to obtain a video program image, and receive the display information sent by the edge node and decode it, perform synchronization and mixed rendering based on the decoded video program image and display information, obtain the program picture and play it on the end-side node or the playback device connected to the end-side node.
[0007] In combination with the first aspect, in a first possible implementation method of the first aspect, the cloud node is specifically used to: determine the type of the end-side node based on the video program playback request, wherein the types of the end-side node include set-top boxes and smart terminals; determine the target video program and the associated program picture feature set based on the video program playback request; determine the video stream encoded based on the target video program and the associated program picture feature set using the target encoding strategy based on the type of the end-side node; determine the feature stream encoded based on the program picture feature set associated with the target video program; transmit the video stream to the end-side node, and transmit the feature stream to the edge node.
[0008] In combination with the first possible implementation method of the first aspect, in the second possible implementation method of the first aspect, the target encoding strategy is: determining the correspondence between the video program image in the target video program and the program screen features in the program screen feature set; for each frame of the video program image: based on the program screen features corresponding to this frame of the video program image, determining the area of interest and the transition area, using a first size of regional block to perform low-precision encoding on the video program image to form a basic quality block, using a second size of regional block to perform medium-precision encoding on the transition area to form an edge transition block, and using a third size of regional block to perform high-precision encoding on the area of interest. Encoding to form an ROI enhancement block, wherein the transition area is located at the edge of the region of interest, the first size is larger than the second size, and the second size is larger than the third size, wherein ROI represents the region of interest; if the type of the end-side node is a set-top box, the basic quality block, the ROI enhancement block, and the edge transition block are uniformly encapsulated to form a single video stream according to the frame sequence of the target video program; if the type of the end-side node is a smart terminal, the edge transition block and the ROI enhancement block are uniformly encapsulated to form a comprehensive enhancement block, the basic quality block and the comprehensive enhancement block are layered and encapsulated to form a basic video stream and an enhanced video stream according to the frame sequence of the target video program.
[0009] In combination with the second possible implementation manner of the first aspect, in a third possible implementation manner of the first aspect, if the type of the end-side node is a set-top box, the end-side node is specifically used to: receive a single video stream transmitted by the cloud node; for the encapsulated data of each frame of video program image in the single video stream: decapsulate the encapsulated data of this frame of video program image, and extract the encoded data and metadata of each block, wherein the metadata includes a block type identifier, block position coordinates, and block size; decode the encoded data of the basic quality block to generate a full-image low-precision video frame, decode the encoded data of the ROI enhancement block, and cover the area of interest of the full-image low-precision video frame based on the block position coordinates and block size of the ROI enhancement block, and decode the encoded data of the edge transition block, and cover the transition area of the full-image low-precision video frame based on the block position coordinates and block size of the edge transition block, and perform filtering and color space conversion to obtain a video program image.
[0010] In combination with the second possible implementation method of the first aspect, in the fourth possible implementation method of the first aspect, if the type of the end-side node is a smart terminal, the end-side node is specifically used to: receive the basic video stream and the enhanced video stream transmitted by the cloud node; for the encapsulation data of each frame of video program image in the basic video stream: decapsulate the encapsulation data of this frame of video program image, and generate a full-image low-precision video frame after decoding; for the encapsulation data of the corresponding frame of video program image in the enhanced video stream: decapsulate the encapsulation data of this frame of video program image, and extract the encoding data of the comprehensive enhancement block; and perform two-step decoding on the comprehensive enhancement block. The method further comprises: decapsulating the ROI enhancement block and extracting the encoded data and meta information of the ROI enhancement block and the edge transition block, wherein the meta information includes a block type identifier, a block position coordinate, and a block size; decoding the encoded data of the ROI enhancement block, and overlaying the ROI enhancement block onto a region of interest of a full-image low-precision video frame of a corresponding frame video program image based on the block position coordinates and the block size of the ROI enhancement block; and decoding the encoded data of the edge transition block, and overlaying the ROI enhancement block onto a transition region of a full-image low-precision video frame of a corresponding frame video program image based on the block position coordinates and the block size of the edge transition block, performing filtering and color space conversion to obtain a video program image.
[0011] In combination with the second possible implementation method of the first aspect, in the fifth possible implementation method of the first aspect, if the type of the end-side node is a set-top box, the cloud node is also used to: determine the version information of the set-top box from the video program playback request; based on the version information of the set-top box, determine the target encapsulation mode matched by the set-top box, wherein the target encapsulation mode includes a first encapsulation mode and a second encapsulation mode, the first encapsulation mode indicates the encapsulation of the basic mass block, and the second encapsulation mode indicates the unified encapsulation of the basic mass block, the ROI enhancement block and the edge transition block; determine the single video stream encapsulated based on the target encapsulation mode as the video stream sent down to the set-top box.
[0012] In combination with the second possible implementation manner of the first aspect, in the sixth possible implementation manner of the first aspect, if the type of the end-side node is a smart terminal, the cloud node is further used to: determine the terminal network status carried therein from the video program playback request; determine the target encapsulation mode and ROI ratio mode matching the smart terminal based on the terminal network status, wherein the target encapsulation mode includes a third encapsulation mode and a fourth encapsulation mode, the third encapsulation mode indicates that only the basic quality block is encapsulated, and the fourth encapsulation mode indicates that the edge transition block and the ROI enhancement block are uniformly encapsulated to form a comprehensive enhancement block, and then the basic quality block and the comprehensive enhancement block are layered encapsulated, and the ROI ratio mode includes an ROI quantitative mode and an ROI full mode. The ROI quantitative mode indicates that no more than a set number of ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of the video program image are retained, and the ROI full mode indicates that all ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of the video program image are retained; encapsulation is performed based on the target encapsulation mode and the ROI ratio mode to form a basic video stream as the video stream sent down to the smart terminal, or a basic video stream and an enhanced video stream are formed as the video stream sent down to the smart terminal.
[0013] In combination with the first aspect, in the seventh possible implementation method of the first aspect, the edge node is specifically used to: receive a feature stream sent by the cloud node, wherein the feature stream includes program screen features corresponding to each frame of the video program image, each program screen feature has a corresponding scene label, and each program screen feature includes several types of subdivision features, and the types of subdivision features include character features and object features; if the auxiliary service type provided by the edge node is intelligent character analysis, the target character is determined based on the character features, and the target character information is determined as display information based on the target character; if the auxiliary service type provided by the edge node is intelligent item analysis, the target item is determined based on the item features, and the target item information is determined as display information based on the target item; if the auxiliary service type provided by the edge node is intelligent advertising delivery, the target item is determined based on the item features, the corresponding target advertising information is determined based on the target item, and the target advertising information is used as display information.
[0014] In combination with the seventh possible implementation manner of the first aspect, in the eighth possible implementation manner of the first aspect, after obtaining the video program image, the end-side node is further used to: decode the received display information and extract its timestamp and display information type identifier; determine the corresponding target display interface based on the display information type identifier, and fill the display information into the target display interface, wherein the target display interface includes the position coordinates and interface size of this interface; synchronize the target display interface with the video program image of the corresponding frame based on the timestamp of the display information; superimpose the target display interface on the specified area of the video program image in the form of a transparent layer according to the position coordinates and interface size, and perform pixel fusion through an Alpha blending algorithm; and render and output the fused program screen.
[0015] In the second aspect, an embodiment of the present application provides a method for playing video programs adapted to a digital retina architecture of a video program playing system adapted to a digital retina architecture applied to the first aspect or any one of the possible implementation methods of the first aspect, the method comprising: initiating a video program playing request through the end-side node; determining a target video program and an associated program screen feature set based on the video program playing request through the cloud node, and transmitting a video stream encoded based on the target video program and the associated program screen feature set to the end-side node, and transmitting a feature stream encoded based on the program screen feature set to the edge node, wherein the program screen feature is a feature of the target video program corresponding to the target video program. The video program image is obtained by extracting features; the feature stream sent by the cloud node is received by the edge node and feature analysis is performed, and display information is determined according to the feature analysis result and the auxiliary service type provided by itself, and is sent to the end-side node, wherein the auxiliary service type includes at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising delivery; the video stream transmitted by the cloud node is received by the end-side node and decoded to obtain the video program image, and the display information sent by the edge node is received and decoded, and synchronization and mixed rendering are performed based on the decoded video program image and display information to obtain the program picture and play it on the end-side node or the playback device connected to the end-side node.
[0016] Beneficial effects:
[0017] This solution provides a video program playback system adapted to the digital retina architecture. It builds an architecture consisting of cloud nodes, edge nodes (several), and end-side nodes. Video program playback requests are initiated through the end-side nodes. Because IPTV video programs (non-live broadcast scenarios) are hosted in the cloud, pre-processing feature extraction can be performed. For each frame of each video program, corresponding features can be extracted to form a program image feature set associated with the video program and stored. Pre-encoding can be selected as needed (pre-encoding saves resources during distribution, while not pre-encoding allows for flexible adjustment of feature encoding strategies). Furthermore, pre-encoding can be performed to improve distribution efficiency. Therefore, based on the video program playback request, the cloud node can determine the target video program and its associated program image feature set (wherein the program image features are extracted from the corresponding video program images in the target video program). The cloud node then transmits a video stream encoded based on the target video program and its associated program image feature set to the end-side node, and transmits a feature stream encoded based on the program image feature set to the edge node. Edge nodes receive and analyze feature streams from cloud nodes. Based on the analysis results and the types of auxiliary services they provide (such as intelligent character analysis, intelligent item analysis, and intelligent advertising), they determine display information and send it to the client-side node. Client-side nodes receive and decode video streams from cloud nodes to obtain video program images. They also receive and decode display information from edge nodes. They synchronize and blend the decoded video program images with the display information to render the resulting program images, which are then played on the client-side node or a connected playback device. This architecture leverages a combination of precoding technology and feature extraction to flexibly adjust and expand auxiliary services (it is also highly compatible with auxiliary services provided by external suppliers, who can join this architecture as edge nodes to flexibly expand auxiliary services). Furthermore, this architectural design is highly compatible with existing IPTV architectures and can be implemented with minor modifications to the existing architecture (older set-top boxes can be upgraded to support higher-quality video playback. If not, the existing architecture can still be used for video playback, consuming high bandwidth to watch high-definition video programs). Furthermore, this architecture provides architectural support for the introduction of "digital retina" technology (a concept proposed in recent years that deploys lightweight models on client-side nodes to compress video and extract features during video transmission, creating a "video stream" + "feature stream" transmission solution, reducing bandwidth usage). For example, in live broadcast scenarios, the "video stream" and "feature stream" transmitted from the client-side node to the cloud node can be simultaneously obtained, without the need for additional feature extraction on the cloud.Therefore, the architecture of this system is an innovative architecture suitable for the IPTV field. It is not only well compatible with the original IPTV architecture, but also provides a flexible function adjustment mechanism (flexible expansion or adjustment of auxiliary services through edge nodes), and can also serve as the basis for introducing "digital retina" technology.
[0018] To improve the video quality of IPTV video programs at relatively low bandwidth, this system designs different pre-encoding strategies for different end-side node types. For each video program image, based on the corresponding program image features of this frame, the region of interest and transition region are determined. The video program image is low-precision encoded using a first-size region block to form a basic quality block. The transition region is medium-precision encoded using a second-size region block to form an edge transition block. The region of interest is high-precision encoded using a third-size region block to form a ROI enhancement block. The transition region is located at the edge of the region of interest, and the first size is larger than the second size, and the second size is larger than the third size. If the end-side node is a set-top box, the basic quality block, ROI enhancement block, and edge transition block are uniformly packaged to form a single video stream according to the frame sequence of the target video program. If the end-side node is a smart terminal, the edge transition block and ROI enhancement block are uniformly packaged to form a comprehensive enhancement block. The basic quality block and comprehensive enhancement block are layered and packaged to form a basic video stream and an enhanced video stream according to the frame sequence of the target video program. This pre-encoding scheme maintains high-quality (high-definition) video playback while effectively reducing bandwidth. Because users typically focus on specific areas of focus, ROI enhancement is used for key features (such as people and objects), while lowering the definition for the background. Reduced background clarity barely impacts the user experience. Furthermore, to eliminate visual differences at the interface between the ROI and background (which can create jagged edges), edge transition blocks are added to effectively prevent these artifacts, providing a better viewing experience. Different encoding strategies are designed for different end-point node types (set-top boxes or smart terminals) based on their decoding technology. Set-top boxes, with their stable networks, are suited for decoding single video streams, so a unified encapsulation encoding scheme is used. Smart terminals, however, often experience unstable networks (possibly with intermittent performance), and are therefore suited for decoding multiple video streams using layered encoding. This allows for flexible bitrate allocation across multiple streams based on varying network conditions, and adaptively retains some ROI enhancement blocks and edge transition blocks to accommodate the network conditions of smart terminals.
[0019] After receiving the feature stream sent by the cloud node (the feature stream contains program screen features corresponding to each frame of the video program image, each program screen feature has a corresponding scene label, and each program screen feature includes several types of detailed features, including character features and object features), the edge node can determine the corresponding display information (such as target character information, target object information, target advertising information, etc.) based on the auxiliary service type provided by the edge node and send it to the end-side node. The end-side node can decode the received display information, extract its timestamp and display information type identifier; based on the display information type identifier, determine the corresponding target display interface and fill the display information into the target display interface, where the target display interface includes the location coordinates and interface size of this interface; synchronize the target display interface with the video program image of the corresponding frame based on the display information timestamp; overlay the target display interface as a transparent layer on the specified area of the video program image based on the location coordinates and interface size, and perform pixel fusion using the Alpha blending algorithm; and render the fused program image for output. Based on this, edge nodes can provide the auxiliary services they need to provide and can flexibly interact according to user needs (for example, turning on or off the service) without being fixed in the precoding stage (if the service is fixed in the precoding stage, it is difficult to adjust and it is difficult to flexibly respond to users' real-time needs).
[0020] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a schematic diagram of the architecture of a video program playback system adapted to the digital retina architecture provided in an embodiment of the present application.
[0023] Figure 2 This is an interactive diagram of the various nodes in the video program playback system adapted to the digital retina architecture.
[0024] Figure 3 This is a rendering of the auxiliary services provided on the video program playback interface. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0026] See also Figure 1 , Figure 1 This is a schematic diagram of the architecture of a video program playback system adapted to the digital retina architecture provided in an embodiment of the present application.
[0027] In this embodiment, the architecture of a video program playback system adapted to the digital retina architecture is designed to include cloud nodes, edge nodes, and end-side nodes. Cloud nodes are considered central nodes (including cloud storage, processing, and distribution). Edge nodes, as nodes providing auxiliary services (such as intelligent person analysis, intelligent item analysis, and intelligent advertising placement), can be servers deployed by the organization or by external vendors. A single edge node can provide multiple functions, or multiple edge nodes can provide a single auxiliary service or multiple different auxiliary services. End-side nodes are user-side devices. In this embodiment, two types of end-side nodes are considered: set-top boxes and smart terminals (such as smartphones and tablets). There are numerous end-side nodes. Each end-side node is connected to a cloud node, and each edge node is connected to a cloud node. Connections between edge nodes and end-side nodes can be as needed (for example, if some end-side nodes do not have certain auxiliary services enabled, they do not require a communication connection. However, if the auxiliary service is enabled, a communication connection with the edge node providing the auxiliary service is required).
[0028] In existing IPTV architectures, video program content is stored in cloud nodes (with dedicated database servers). Prior to implementing this architecture, feature extraction is required to analyze certain video program content. Therefore, within this system, feature extraction (which can be performed offline) is required for existing IPTV video programs. The feature extraction technology can utilize existing or custom-developed solutions, but this is not limited here. For each video program (e.g., a single movie or TV series episode), corresponding features are extracted from each frame to form a set of program features associated with that video program. This analysis identifies the scene to which each program feature belongs (e.g., by shot cuts), as well as the sub-features contained within each program feature (e.g., human and object features, such as faces and bodies). Images that do not contain human and / or object features (e.g., only background features) do not need to be individually identified.
[0029] It should be noted that IPTV video programs are provided by content sources. Most content sources are static content sources (i.e., video programs that have been filmed, edited, and uploaded after review). There are also live content sources (provided by live cameras). The "Digital Retina" architecture proposes that the end-side node that provides content (such as the live camera that provides the content source in the live scene) provides video streams and feature streams (a lightweight feature extraction model is deployed on the end-side node to reduce the amount of data transmission and reduce the burden on the cloud node, because the cloud node does not need to perform feature extraction). Therefore, the architecture of the video program playback system adapted to the digital retina architecture provided in this embodiment can provide architectural support for accessing the "Digital Retina" technology.
[0030] Furthermore, the cloud node also needs to pre-encode the video program and its associated program image feature set. The pre-encoding strategy is designed as follows:
[0031] First, the correspondence between the video program image and the program screen features in the program screen feature set in the target video program is determined. Then, for each frame of the video program image, a region of interest (ROI) and a transition region are determined based on the program screen features corresponding to the frame of the video program image. (The transition region is located at the edge of the ROI and is primarily used to smooth the difference in clarity between the ROI and the background area to avoid jagged edges. The program screen features can guide the determination of the ROI, while the transition region depends on the ROI. For example, the edge region is obtained by expanding the ROI boundary outward by a certain size.) Specifically, the video program image is low-precision encoded using region blocks of a first size (e.g., 32*32; asymmetric sizes can be used for the boundary portion, as is done in the prior art for encoding such boundary portions, which will not be described in detail here). A second size (e.g., 16*16) is used to medium-precision encode the transition region to form an edge transition block. A third size (e.g., 8*8) is used to high-precision encode the ROI to form an ROI (Region of Interest) enhancement block. In this embodiment, the first size is larger than the second size, and the second size is larger than the third size. This size relationship needs to be met.
[0032] For set-top boxes (STBs), encoding a single video stream is required. Therefore, the cloud node must uniformly encapsulate the base quality block, ROI enhancement block, and edge transition block to form a single video stream in the target video program's frame order. This solution supports existing codec protocols such as H.265 / HEVC and AVS3, allowing for variable block sizes within the same video stream. The block type (base / edge / ROI) can be marked using SEI (Supplemental Enhancement Information) messages, allowing the decoder to process blocks in different regions based on the SEI information. AVS3's marking method differs from H.265, allowing the use of AVS3_user_data and customizable SEI syntax. In order to adapt to different set-top box versions, different encapsulation modes can also be determined: the first encapsulation mode and the second encapsulation mode. The first encapsulation mode indicates the encapsulation of the basic quality block (for example, the old version of the set-top box only supports H.264, then only a single video stream pre-encoded based on the basic quality block can be loaded), and the second encapsulation mode indicates the unified encapsulation of the basic quality block, ROI enhancement block and edge transition block.
[0033] For smart terminals, layered encoding is possible. To ensure the quality of the region of interest (ROI), this embodiment uniformly encapsulates the edge transition block and the ROI enhancement block to form a comprehensive enhancement block. The base quality block and the comprehensive enhancement block are then layered, forming a base video stream and an enhanced video stream according to the frame sequence of the target video program. This solution also supports existing codec protocols such as H.265 / HEVC and AVS3, and the SHVC (Scalable Coding) layer structure of H.265, enabling layered encapsulation.
[0034] In order to adapt to the changing network conditions of smart terminals, such as good network status, moderate network status, and poor network status (this standard can define the network status level of specific parameters based on actual network test conditions. This is only a tentative network status classification at present, and other classification schemes can be proposed and are not limited here), multiple pre-coding methods can be used: for example, pre-coding basic quality blocks to form a basic video stream (suitable for poor network conditions); pre-coding basic quality blocks to form a basic video stream, and uniformly encapsulating edge transition blocks and ROI enhancement blocks to form a comprehensive enhancement block, and then pre-coding all comprehensive enhancement blocks to form an enhanced video stream (suitable for good network conditions); and pre-coding basic quality blocks to form a basic video stream, and uniformly encapsulating edge transition blocks and ROI enhancement blocks to form a comprehensive enhancement block, and then pre-coding some comprehensive enhancement blocks (for example, retaining no more than 5 or no more than 10 comprehensive enhancement blocks per frame of video program image) to form an enhanced video stream (suitable for moderate network conditions).
[0035] Specifically, it can be divided into a packaging mode and an ROI proportion mode, wherein the packaging mode can include a third packaging mode and a fourth packaging mode. The third packaging mode indicates that only the basic mass block is packaged, and the fourth packaging mode indicates that the edge transition block and the ROI enhancement block are uniformly packaged to form a comprehensive enhancement block, and then the basic mass block and the comprehensive enhancement block are layered packaged. The ROI proportion mode includes an ROI quantitative mode and an ROI full mode. The ROI quantitative mode indicates that no more than a set number of ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of video program image are retained. The ROI full mode indicates that all ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of video program image are retained.
[0036] The encapsulation modes (first encapsulation mode, second encapsulation mode, third encapsulation mode, and fourth encapsulation mode) and ROI ratio modes (ROI quantitative mode and ROI full mode) mentioned above are all pre-encoded to form a pre-encoded video stream. This eliminates the need for real-time encoding when delivering the video, saving resources and improving efficiency.
[0037] The pre-coding of feature streams is simple. The program picture features are encoded according to the frame sequence corresponding to the video program to form a feature stream, which is sent to the corresponding edge node when needed so that the edge node can provide auxiliary services.
[0038] This pre-encoding scheme maintains high-quality (high-definition) video playback while effectively reducing bandwidth. Because users typically focus on specific areas of focus, ROI enhancement is used for key features (such as people and objects), while lowering the definition for the background. Reduced background clarity barely impacts the user experience. Furthermore, to eliminate visual differences at the interface between the ROI and background (which can create jagged edges), edge transition blocks are added to effectively prevent these artifacts, providing a better viewing experience. Different encoding strategies are designed for different end-point node types (set-top boxes or smart terminals) based on their decoding technology. Set-top boxes, with their stable networks, are suited for decoding single video streams, so a unified encapsulation encoding scheme is used. Smart terminals, however, often experience unstable networks (possibly with intermittent performance), and are therefore suited for decoding multiple video streams using layered encoding. This allows for flexible bitrate allocation across multiple streams based on varying network conditions, and adaptively retains some ROI enhancement blocks and edge transition blocks to accommodate the network conditions of smart terminals.
[0039] Moreover, this solution can take into account different end-side node types and support existing protocols. For set-top boxes with older versions that do not support H.265, units can also adopt a variety of commercial promotion plans to gradually update the old set-top boxes to gradually improve the architectural deployment of the video program playback system adapted to the digital retina architecture.
[0040] Therefore, the architectural design of a video program playback system adapted to the digital retina architecture can leverage precoding and feature extraction technologies to flexibly adjust and expand auxiliary services (it can also be well compatible with auxiliary services provided by external suppliers, who can join this architecture as edge nodes to flexibly expand auxiliary services). This architectural design is also highly compatible with the existing IPTV architecture and can be completed with minor modifications to the existing architecture (for older set-top boxes, they can be upgraded to support higher-quality video playback. If not, the existing architecture can still be used for video program playback, i.e., high-bandwidth viewing of high-definition video programs). Furthermore, this architecture can also provide architectural support for the introduction of "digital retina" technology (digital retina technology is an idea proposed in recent years. By deploying lightweight models on edge nodes, video can be compressed and features extracted during video transmission to achieve a "video stream" + "feature stream" transmission solution, reducing bandwidth usage). For example, in live broadcast scenarios, the "video stream" and "feature stream" transmitted from the edge node to the cloud node can be simultaneously obtained, without the need for additional feature extraction on the cloud. Therefore, the architecture of this system is an innovative architecture suitable for the IPTV field. It is not only well compatible with the original IPTV architecture, but also provides a flexible function adjustment mechanism (flexible expansion or adjustment of auxiliary services through edge nodes), and can also serve as the basis for introducing "digital retina" technology.
[0041] In order to facilitate the understanding of the video program playback system adapted to the digital retina architecture, Figure 2 , which introduces the interaction process of each node in the video program playback system adapted to the digital retina architecture. Figure 2 , Figure 2 This is an interactive diagram of the various nodes in the video program playback system adapted to the digital retina architecture.
[0042] First, the client node can initiate a video playback request. This request includes the video program ID and the client node type (e.g., distinguished by a type identifier). For set-top boxes, this request also includes the set-top box version information. For smart terminals, this request also includes the current network status. (During subsequent video streaming, the network status is also obtained in real time to dynamically adjust the pre-encoded video stream being transmitted.)
[0043] After receiving a video program playback request, the cloud node can parse the request to determine the target video program and the associated program picture feature set, and transmit the video stream encoded based on the target video program and the associated program picture feature set to the end-side node, and transmit the feature stream encoded based on the program picture feature set to the edge node.
[0044] Exemplarily, the cloud node can determine the type of the end-side node, the video program number, and the network status of the smart terminal or the version information of the set-top box based on the video program playback request, and then determine the target video program (and the associated program picture feature set) based on the video program number in the video program playback request, and then, based on the type of the end-side node, determine the video stream to be encoded based on the target video program and the associated program picture feature set using the target encoding strategy.
[0045] Specifically, when the type of the end-side node is a set-top box, the cloud node can determine the version information of the set-top box from the video program playback request; then, based on the version information of the set-top box, determine the target encapsulation mode that matches the set-top box (for example, if the set-top box version supports H.265, then determine the second encapsulation mode as the target encapsulation mode, and the second encapsulation mode indicates that the basic mass block, ROI enhancement block, and edge transition block are uniformly encapsulated; if the set-top box version does not support H.265, then determine the first encapsulation mode as the target encapsulation mode, and the first encapsulation mode indicates that the basic mass block is encapsulated), and then determine the single video stream encapsulated based on the target encapsulation mode as the video stream sent down to the set-top box.
[0046] Specifically, when the end-side node is a smart terminal, the cloud node can determine the terminal network status (e.g., good network status, moderate network status, poor network status, etc.) contained in the video program playback request. Based on the terminal network status, the cloud node can then determine the target encapsulation mode and ROI ratio mode that match the smart terminal. For example, when the network status is good, the fourth encapsulation mode is determined as the target encapsulation mode, and the ROI ratio mode is determined as the ROI full mode. The cloud node can then determine the base video stream and enhanced video stream formed by encapsulating the target encapsulation mode (the fourth encapsulation mode) and the ROI full mode as the video streams to be delivered to the smart terminal. For another example, when the network status is moderate, the cloud node can determine the fourth encapsulation mode as the target encapsulation mode, and the ROI ratio mode is determined as the ROI quantitative mode. The cloud node can then determine the base video stream and enhanced video stream formed by encapsulating the target encapsulation mode (the fourth encapsulation mode) and the ROI quantitative mode as the video streams to be delivered to the smart terminal. For example, when the network status is poor, the cloud node can determine the third encapsulation mode as the target encapsulation mode, and then use the basic video stream formed by encapsulation based on the target encapsulation mode (the third encapsulation mode) as the video stream sent to the smart terminal.
[0047] At the same time, the cloud node also needs to determine the program picture feature set associated with the target video program, thereby determining the feature stream encoded with the program picture feature set and transmitting it to the edge node.
[0048] The edge node can receive the feature stream sent by the cloud node and perform feature analysis. Based on the feature analysis results and the auxiliary service type provided by itself, it determines the display information and sends it to the terminal node. The auxiliary service type includes at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising delivery.
[0049] Exemplarily, the edge node can receive a feature stream sent by the cloud node, where the feature stream includes program screen features corresponding to each frame of video program image, each program screen feature has a corresponding scene label, and each program screen feature includes several types of subdivided features, including character features and object features.
[0050] If the auxiliary service type provided by the edge node is intelligent role analysis, then the edge node can determine the target role based on the character characteristics, and determine the target role information as display information based on the target role.
[0051] If the auxiliary service type provided by the edge node is intelligent item analysis, then the edge node can determine the target item based on the item characteristics, and determine the target item information as display information based on the target item.
[0052] If the auxiliary service type provided by the edge node is intelligent advertising delivery, then the edge node can determine the target item based on the item characteristics, and then determine the corresponding target advertising information based on the target item (which may not be unique advertising information and may involve priority recommendation), and use the target advertising information as display information.
[0053] Due to the high real-time requirements of such auxiliary services, lightweight feature analysis models (or feature matching models) deployed at edge nodes are preferred, as they require minimal processing time. For the program features of multiple consecutive frames of video program images within the same scene (with the same scene label), the display information determined based on the segmented features under that scene label can be used for each of these segmented features (if the scene label changes, the display information needs to be re-determined, and for interval scenes, the display information needs to be re-determined, and the previously determined display information cannot be used). The number of consecutive frames can also be determined in the display information to reduce the amount of data required to be transmitted. However, in this case, a strict correspondence between the number of consecutive frames of the display information and the number of frames of the video program image must be designed to avoid incorrect correspondence issues.
[0054] After the edge node determines the display information, it can send the determined display information (with a timestamp and a display information type identifier) to the end-side node.
[0055] The end-side node can receive the video stream transmitted by the cloud node and decode it to obtain the video program image, and receive the display information sent by the edge node and decode it. It synchronizes and mixes the decoded video program image and display information to render the program image and play it on the end-side node (such as a smart terminal) or the playback device connected to the end-side node (such as a TV connected to a set-top box).
[0056] When the end-side node is a set-top box, the end-side node can receive a single video stream transmitted by the cloud node; then, for the encapsulated data of each frame of the video program image in the single video stream: the encapsulated data of this frame of the video program image is decapsulated, and the encoding data and metadata of each block (including block type identifier, block position coordinates, and block size) are extracted; then, the encoding data of the basic quality block is decoded to generate a full-image low-precision video frame, and then the encoding data of the ROI enhancement block is decoded, and based on the block position coordinates and block size of the ROI enhancement block, it is covered to the area of interest of the full-image low-precision video frame, and the encoding data of the edge transition block is decoded, and based on the block position coordinates and block size of the edge transition block, it is covered to the transition area of the full-image low-precision video frame, and then filtering and color space conversion are performed to obtain the video program image.
[0057] It should be noted that the way in which the ROI enhancement block covers the area of interest of the full-image low-precision video frame and the edge transition block covers the transition area of the full-image low-precision video frame can be incremental overlay or replacement overlay. If an incremental overlay scheme is adopted, during pre-encoding, the encoding of the area of interest and the encoding of the edge area need to be designed as an incremental encoding scheme.
[0058] In addition, for compatibility with old versions of set-top boxes (if the old version of the set-top box does not support the H.265 protocol), the set-top box can directly decode a single video stream (containing only basic quality blocks, which can be spliced to obtain a full-image low-precision video frame). There is no need to overwrite the full-image low-precision video frame with ROI enhancement blocks and edge transition blocks. The obtained full-image low-precision video frame can be converted into a color space to obtain the video program image.
[0059] If the end-side node is a smart terminal, the end-side node can receive the basic video stream and enhanced video stream transmitted by the cloud node. For the encapsulated data of each frame of video program image in the basic video stream: the end-side node can decapsulate the encapsulated data of this frame of video program image, and generate a full-image low-precision video frame after decoding; for the encapsulated data of the corresponding frame of video program image in the enhanced video stream: decapsulate the encapsulated data of this frame of video program image, extract the coded data of the comprehensive enhancement block; decapsulate the comprehensive enhancement block twice, extract the coded data and meta information of the ROI enhancement block and the edge transition block, wherein the meta information includes the block type identifier, block position coordinates, and block size. Then decode the coded data of the ROI enhancement block, and based on the block position coordinates and block size of the ROI enhancement block, cover the area of interest of the full-image low-precision video frame of the corresponding frame of video program image, and decode the coded data of the edge transition block, and based on the block position coordinates and block size of the edge transition block, cover the transition area of the full-image low-precision video frame of the corresponding frame of video program image, perform filtering and color space conversion to obtain the video program image.
[0060] Similarly, the ROI enhancement block covers the region of interest of the full low-precision video frame, and the edge transition block covers the transition region of the full low-precision video frame, either through incremental overlay or replacement. If an incremental overlay scheme is adopted, the coding of the region of interest and the edge region needs to be designed as an incremental coding scheme during pre-encoding. Incremental coding schemes are commonly used for smart terminals. Since this type of incremental coding technology is already mature and applicable, the process will not be detailed here.
[0061] In addition, for the cases of good network status and moderate network status, the processing process is similar (only the number of comprehensive enhancement blocks is different). For the case of poor network status, the video stream decoded by the smart terminal only has the basic video stream, and there is no need to cover the full-image low-precision video frame with ROI enhancement blocks and edge transition blocks. The obtained full-image low-precision video frame can be converted into a color space to obtain the video program image.
[0062] After receiving the video program image (the presentation information decoding mentioned here refers to the process after receiving the video program image; in practice, presentation information decoding and video stream decoding can be performed simultaneously), the client-side node can also decode the received presentation information, extracting its timestamp and presentation information type identifier. The client-side node can then determine the corresponding target presentation interface based on the presentation information type identifier and populate the target presentation interface with the presentation information (e.g., text information, advertising images, etc.). The target presentation interface includes its location coordinates and dimensions.
[0063] Afterwards, the client-side node can synchronize the target display interface with the video program image of the corresponding frame based on the timestamp of the display information; superimpose the target display interface as a transparent layer on the designated area of the video program image according to the position coordinates and interface size, and perform pixel fusion through the Alpha blending algorithm, thereby rendering and outputting the fused program image to achieve video program playback. Figure 3 The following is a schematic diagram of the effect of providing auxiliary services. This is for illustrative purposes only. With the development of auxiliary services and the improvement of edge node deployment, as well as the increase in service providers participating in this architecture to provide auxiliary services and technological development, it is expected that extremely effective and sophisticated auxiliary services will be provided. In addition to character information and item information, it can even identify prop information and character moves in the play, providing more comprehensive introductions and stronger program interactivity. The schematic diagram here is not considered to limit this application.
[0064] On the user side, an auxiliary service can be turned on or off as needed, and the on or off operation instructions can be fed back to the cloud node (which can be fed back to the cloud node through the edge node, or fed back to the cloud node and then transmitted to the edge node) to stop the transmission of the feature flow and turn off the display information. Of course, for the current advertising delivery strategy, it is not the user who actively operates to turn it on and off. In order to comply with the relevant regulations on advertising delivery, it is necessary to design the delivery strategy and delivery duration in a targeted manner. The specific handling of this will not be elaborated here. Of course, this situation should not be regarded as a limitation of this application.
[0065] In summary, edge nodes can provide the auxiliary services they need to provide and can flexibly interact according to user needs (for example, turning the service on and off) without being fixed in the precoding stage (if the service is fixed in the precoding stage, it is difficult to adjust and it is difficult to flexibly respond to users' real-time needs).
[0066] Based on the same inventive concept, the embodiment of the present application further provides a video program playback method adapted to the digital retina architecture, which is applied to the video program playback system adapted to the digital retina architecture of the present embodiment. The method includes:
[0067] Step S10: Initiate a video program playback request through the terminal-side node.
[0068] Step S20: Based on the video program playback request, the cloud node determines the target video program and the associated program screen feature set, and transmits the video stream encoded based on the target video program and the associated program screen feature set to the end-side node, and transmits the feature stream encoded based on the program screen feature set to the edge node, wherein the program screen feature is obtained by feature extraction of the corresponding video program image in the target video program.
[0069] Step S30: The edge node receives the feature stream sent by the cloud node and performs feature analysis. Based on the feature analysis results and the auxiliary service type provided by itself, the display information is determined and sent to the terminal node. The auxiliary service type includes at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising delivery.
[0070] Step S40: Receive the video stream transmitted by the cloud node through the end-side node and decode it to obtain a video program image, and receive and decode the display information sent by the edge node, synchronize and mix the decoded video program image and display information, obtain the program picture, and play it on the end-side node or the playback device connected to the end-side node.
[0071] The specific processing process of the video program playback method adapted to the digital retina architecture can be found in the previous article and will not be repeated here.
[0072] In summary, embodiments of the present application provide a video program playback system and method adapted for a digital retina architecture. The video program playback system adapted for a digital retina architecture provided by this solution establishes a cloud node-edge node(s)-end-side node architecture, with video program playback requests initiated by the end-side nodes. Because IPTV video programs (non-live broadcast scenarios) are hosted in the cloud, pre-processing feature extraction can be performed. Corresponding features can be extracted from each frame of each video program, forming a program image feature set associated with the video program and storing it. Pre-encoding can be selected as needed (pre-encoding can save resources during distribution, while not pre-encoding allows for flexible adjustment of feature encoding strategies). Furthermore, pre-encoding can be performed for video programs to improve distribution efficiency. Therefore, based on the video program playback request, the cloud node can determine the target video program and the associated program image feature set (wherein the program image features are obtained by extracting features from the corresponding video program images in the target video program), transmit a video stream encoded based on the target video program and the associated program image feature set to the end-side node, and transmit a feature stream encoded based on the program image feature set to the edge node. Edge nodes receive and analyze feature streams from cloud nodes. Based on the analysis results and the types of auxiliary services they provide (such as intelligent character analysis, intelligent item analysis, and intelligent advertising), they determine display information and send it to the client-side node. Client-side nodes receive and decode video streams from cloud nodes to obtain video program images. They also receive and decode display information from edge nodes. They synchronize and blend the decoded video program images with the display information to render the resulting program images, which are then played on the client-side node or a connected playback device. This architecture leverages a combination of precoding technology and feature extraction to flexibly adjust and expand auxiliary services (it is also highly compatible with auxiliary services provided by external suppliers, who can join this architecture as edge nodes to flexibly expand auxiliary services). Furthermore, this architectural design is highly compatible with existing IPTV architectures and can be implemented with minor modifications to the existing architecture (older set-top boxes can be upgraded to support higher-quality video playback. If not, the existing architecture can still be used for video playback, consuming high bandwidth to watch high-definition video programs). Furthermore, this architecture provides architectural support for the introduction of "digital retina" technology (a concept proposed in recent years that deploys lightweight models on client-side nodes to compress video and extract features during video transmission, creating a "video stream" + "feature stream" transmission solution, reducing bandwidth usage). For example, in live broadcast scenarios, the "video stream" and "feature stream" transmitted from the client-side node to the cloud node can be simultaneously obtained, without the need for additional feature extraction on the cloud.Therefore, the architecture of this system is an innovative architecture suitable for the IPTV field. It is not only well compatible with the original IPTV architecture, but also provides a flexible function adjustment mechanism (flexible expansion or adjustment of auxiliary services through edge nodes), and can also serve as the basis for introducing "digital retina" technology.
[0073] To improve the video quality of IPTV video programs at relatively low bandwidth, this system designs different pre-encoding strategies for different end-side node types. For each video program image, based on the corresponding program image features of this frame, the region of interest and transition region are determined. The video program image is low-precision encoded using a first-size region block to form a basic quality block. The transition region is medium-precision encoded using a second-size region block to form an edge transition block. The region of interest is high-precision encoded using a third-size region block to form a ROI enhancement block. The transition region is located at the edge of the region of interest, and the first size is larger than the second size, and the second size is larger than the third size. If the end-side node is a set-top box, the basic quality block, ROI enhancement block, and edge transition block are uniformly packaged to form a single video stream according to the frame sequence of the target video program. If the end-side node is a smart terminal, the edge transition block and ROI enhancement block are uniformly packaged to form a comprehensive enhancement block. The basic quality block and comprehensive enhancement block are layered and packaged to form a basic video stream and an enhanced video stream according to the frame sequence of the target video program. This pre-encoding scheme maintains high-quality (high-definition) video playback while effectively reducing bandwidth. Because users typically focus on specific areas of focus, ROI enhancement is used for key features (such as people and objects), while lowering the definition for the background. Reduced background clarity barely impacts the user experience. Furthermore, to eliminate visual differences at the interface between the ROI and background (which can create jagged edges), edge transition blocks are added to effectively prevent these artifacts, providing a better viewing experience. Different encoding strategies are designed for different end-point node types (set-top boxes or smart terminals) based on their decoding technology. Set-top boxes, with their stable networks, are suited for decoding single video streams, so a unified encapsulation encoding scheme is used. Smart terminals, however, often experience unstable networks (possibly with intermittent performance), and are therefore suited for decoding multiple video streams using layered encoding. This allows for flexible bitrate allocation across multiple streams based on varying network conditions, and adaptively retains some ROI enhancement blocks and edge transition blocks to accommodate the network conditions of smart terminals.
[0074] After receiving the feature stream sent by the cloud node (the feature stream contains program screen features corresponding to each frame of the video program image, each program screen feature has a corresponding scene label, and each program screen feature includes several types of detailed features, including character features and object features), the edge node can determine the corresponding display information (such as target character information, target object information, target advertising information, etc.) based on the auxiliary service type provided by the edge node and send it to the end-side node. The end-side node can decode the received display information, extract its timestamp and display information type identifier; based on the display information type identifier, determine the corresponding target display interface and fill the display information into the target display interface, where the target display interface includes the location coordinates and interface size of this interface; synchronize the target display interface with the video program image of the corresponding frame based on the display information timestamp; overlay the target display interface as a transparent layer on the specified area of the video program image based on the location coordinates and interface size, and perform pixel fusion using the Alpha blending algorithm; and render the fused program image for output. Based on this, edge nodes can provide the auxiliary services they need to provide and can flexibly interact according to user needs (for example, turning on or off the service) without being fixed in the precoding stage (if the service is fixed in the precoding stage, it is difficult to adjust and it is difficult to flexibly respond to users' real-time needs).
[0075] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A video program playback system adapted to a digital retina architecture, characterized in that: Including cloud nodes, edge nodes and end nodes, The terminal-side node is used to initiate a video program playback request; The cloud node is configured to determine a target video program and an associated program image feature set based on a video program playback request, and transmit a video stream encoded based on the target video program and the associated program image feature set to the end-side node, and transmit a feature stream encoded based on the program image feature set to the edge node, wherein the program image features are obtained by extracting features from corresponding video program images in the target video program; The edge node is configured to receive the feature stream sent by the cloud node and perform feature analysis, determine display information based on the feature analysis results and the type of auxiliary service provided by the edge node, and send the display information to the end-side node, wherein the auxiliary service type includes at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising delivery; The end-side node is configured to receive the video stream transmitted by the cloud node and decode it to obtain a video program image, and to receive and decode the presentation information sent by the edge node, perform synchronization and mixed rendering based on the decoded video program image and presentation information, obtain a program screen, and play it on the end-side node or a playback device connected to the end-side node.
2. The video program playback system adapted to the digital retina architecture according to claim 1, characterized in that: The cloud node is specifically used for: Determining, based on the video program playback request, a type of the end-side node, wherein the type of the end-side node includes a set-top box and a smart terminal; Based on the video program playback request, determining a target video program and an associated program picture feature set; Based on the type of the end-side node, determining a video stream to be encoded based on a target video program and an associated program picture feature set using a target encoding strategy; Determining a feature stream to be encoded based on a program picture feature set associated with a target video program; The video stream is transmitted to the end-side node, and the feature stream is transmitted to the edge node.
3. The video program playback system adapted to the digital retina architecture according to claim 2, characterized in that: The target encoding strategy is: Determining a correspondence between a video program image in a target video program and program picture features in a program picture feature set; For each frame of video program image: based on the program screen features corresponding to this frame of video program image, determine the region of interest and the transition region, use a first-size regional block to perform low-precision encoding on the video program image to form a basic quality block, use a second-size regional block to perform medium-precision encoding on the transition region to form an edge transition block, and use a third-size regional block to perform high-precision encoding on the region of interest to form an ROI enhancement block, wherein the transition region is located at the edge of the region of interest, the first size is larger than the second size, and the second size is larger than the third size, wherein ROI represents the region of interest; If the type of the end-side node is a set-top box, the basic quality block, the ROI enhancement block, and the edge transition block are uniformly encapsulated to form a single video stream according to the frame sequence of the target video program; If the type of the end-side node is a smart terminal, the edge transition block and the ROI enhancement block are uniformly encapsulated to form a comprehensive enhancement block, the basic quality block and the comprehensive enhancement block are layered and encapsulated to form a basic video stream and an enhanced video stream according to the frame order of the target video program.
4. The video program playback system adapted to the digital retina architecture according to claim 3, characterized in that: If the type of the end-side node is a set-top box, the end-side node is specifically configured to: receiving a single video stream transmitted by the cloud node; For each frame of video program image encapsulation data in a single video stream: decapsulate the encapsulation data of this frame of video program image and extract the coded data and meta information of each block, wherein the meta information includes block type identification, block position coordinates, and block size; The encoded data of the basic quality block is decoded to generate a full-image low-precision video frame, the encoded data of the ROI enhancement block is decoded, and based on the block position coordinates and block size of the ROI enhancement block, the region of interest of the full-image low-precision video frame is covered, and the encoded data of the edge transition block is decoded, and based on the block position coordinates and block size of the edge transition block, the transition region of the full-image low-precision video frame is covered, and filtering and color space conversion are performed to obtain a video program image.
5. The video program playback system adapted to the digital retina architecture according to claim 3, characterized in that: If the type of the end-side node is a smart terminal, the end-side node is specifically used to: Receiving the basic video stream and the enhanced video stream transmitted by the cloud node; For the encapsulated data of each frame of video program image in the basic video stream: decapsulate the encapsulated data of this frame of video program image, and generate a full-image low-precision video frame after decoding; For the encapsulated data of the corresponding frame of video program image in the enhanced video stream: decapsulating the encapsulated data of the frame of video program image and extracting the coded data of the integrated enhancement block; Perform secondary decapsulation on the comprehensive enhancement block to extract the coded data and meta information of the ROI enhancement block and the edge transition block, wherein the meta information includes the block type identifier, block position coordinates, and block size; The encoded data of the ROI enhancement block is decoded, and based on the block position coordinates and block size of the ROI enhancement block, the ROI enhancement block is covered to the region of interest of the full-image low-precision video frame of the corresponding frame video program image. Furthermore, the encoded data of the edge transition block is decoded, and based on the block position coordinates and block size of the edge transition block, the ROI enhancement block is covered to the transition region of the full-image low-precision video frame of the corresponding frame video program image. Filtering and color space conversion are performed to obtain a video program image.
6. The video program playback system adapted to the digital retina architecture according to claim 3, characterized in that: If the type of the client-side node is a set-top box, the cloud-side node is further configured to: Determining version information of the set-top box from the video program playback request; Determining, based on version information of the set-top box, a target encapsulation mode matched by the set-top box, wherein the target encapsulation mode includes a first encapsulation mode and a second encapsulation mode, the first encapsulation mode indicates encapsulating a basic mass block, and the second encapsulation mode indicates uniformly encapsulating the basic mass block, the ROI enhancement block, and the edge transition block; A single video stream encapsulated based on the target encapsulation mode is determined as the video stream sent to the set-top box.
7. The video program playback system adapted to the digital retina architecture according to claim 3, characterized in that: If the type of the terminal-side node is a smart terminal, the cloud-side node is further used to: Determining the terminal network status carried in the video program playback request from the video program playback request; Determining a target encapsulation mode and an ROI ratio mode matching the smart terminal based on the terminal network state, wherein the target encapsulation mode includes a third encapsulation mode and a fourth encapsulation mode, the third encapsulation mode indicates that only the basic mass block is encapsulated, and the fourth encapsulation mode indicates that the edge transition block and the ROI enhancement block are uniformly encapsulated to form a comprehensive enhancement block, and then the basic mass block and the comprehensive enhancement block are layered encapsulated, and the ROI ratio mode includes an ROI quantitative mode and an ROI full mode, the ROI quantitative mode indicates that no more than a set number of ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of the video program image are retained, and the ROI full mode indicates that all ROI enhancement blocks and corresponding edge transition blocks corresponding to each frame of the video program image are retained; Encapsulation is performed based on the target encapsulation mode and the ROI ratio mode to form a basic video stream as the video stream sent to the smart terminal, or a basic video stream and an enhanced video stream are formed as the video stream sent to the smart terminal.
8. The video program playback system adapted to the digital retina architecture according to claim 1, characterized in that: The edge node is specifically used to: Receive a feature stream sent by the cloud node, wherein the feature stream includes program screen features corresponding to each frame of the video program image, each program screen feature has a corresponding scene label, and each program screen feature includes several types of subdivided features, and the types of subdivided features include character features and object features; If the auxiliary service type provided by the edge node is intelligent role analysis, a target role is determined based on the character characteristics, and target role information is determined based on the target role as display information; If the auxiliary service type provided by the edge node is intelligent item analysis, a target item is determined based on item characteristics, and target item information is determined based on the target item as display information; If the auxiliary service type provided by the edge node is intelligent advertising, a target item is determined based on item features, corresponding target advertising information is determined based on the target item, and the target advertising information is used as display information.
9. The video program playback system adapted to the digital retina architecture according to claim 8, characterized in that: After obtaining the video program image, the terminal-side node is further configured to: Decode the received display information and extract its timestamp and display information type identifier; Determine the corresponding target display interface based on the display information type identifier, and fill the display information into the target display interface, wherein the target display interface includes the position coordinates and interface size of the interface; Synchronize the target display interface with the video program image of the corresponding frame based on the timestamp of the display information; The target display interface is superimposed on the designated area of the video program image in the form of a transparent layer according to the position coordinates and interface size, and pixel fusion is performed using the Alpha blending algorithm; Render and output the fused program images.
10. A method for playing video programs adapted to a digital retinal architecture, applied to a video program playing system adapted to a digital retinal architecture according to any one of claims 1 to 9, characterized in that: The method comprises: Initiating a video program playback request through the terminal-side node; The cloud node determines a target video program and an associated program picture feature set based on a video program playback request, and transmits a video stream encoded based on the target video program and the associated program picture feature set to the end-side node, and transmits a feature stream encoded based on the program picture feature set to the edge node, wherein the program picture feature is obtained by extracting features of corresponding video program images in the target video program; The edge node receives the feature stream sent by the cloud node and performs feature analysis, determines display information based on the feature analysis result and the type of auxiliary service provided by the edge node, and sends the display information to the end-side node, wherein the auxiliary service type includes at least one of intelligent role analysis, intelligent item analysis, and intelligent advertising delivery; The end-side node receives the video stream transmitted by the cloud node and decodes it to obtain a video program image, and receives the display information sent by the edge node and decodes it. Synchronization and mixed rendering are performed based on the decoded video program image and display information to obtain a program picture and play it on the end-side node or a playback device connected to the end-side node.
Citation Information
Patent Citations
Synchronous transmission control method for digital retina video stream and characteristic stream
CN110719438A
Digital retina software definition camera method and system
CN111541864A
Telescopic visual computing system
CN112804188A
Video coding and decoding method and device based on intelligent digital retina
CN114630129A
Novel camera system and intelligent perception-based video coding method thereof
CN118354083A