Video code stream processing method, and device, medium and program product

By encapsulating and mapping the primary bitstream and long-term reference frame bitstream of the video bitstream in the CDN system, the problems of low caching efficiency and high bandwidth cost in the CDN system are solved, achieving efficient video content transmission and quality improvement, and reducing operating costs.

WO2025261101A1PCT designated stage Publication Date: 2025-12-26ZTE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/097209
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-21
Filing Date
2025-05-26
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing CDN systems suffer from low video content caching efficiency, long loading times, and slow distribution speeds. They cannot balance video stream transmission efficiency and video quality, and the large amount of data transmitted results in high bandwidth costs, impacting the viewing experience for end users and increasing operating costs.

Method used

By receiving the raw video bitstream, determining the primary stream and long-term reference frame bitstream, performing encapsulation processing, and generating a media index file, the system supports mapping between the primary stream and the long-term reference frame bitstream, provides long-term caching, responds to index information, indexes media data transmission, and optimizes content storage and transmission methods.

Benefits of technology

Improve video stream caching and transmission efficiency, enhance video quality, reduce operating costs, improve the viewing experience for end users, and reduce bandwidth consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025097209_26122025_PF_FP_ABST
    Figure CN2025097209_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a video code stream processing method, and a device, a medium and a program product. The method comprises: receiving an original video code stream (S100); on the basis of the original video code stream, determining a main bit stream and a long-term reference frame bit stream corresponding to the main bit stream (S200); encapsulating the main bit stream and the long-term reference frame bit stream to obtain media data, generating a media index file, and storing the media data and the media index file, wherein the media index file comprises index information (S300); and in response to receiving a media data request carrying the index information, performing indexing on the basis of the index information to obtain the media data, and sending the media data (S400).
Need to check novelty before this filing date? Find Prior Art

Description

A video stream processing method, device, medium, and program product.

[0001] Cross-reference to related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 202410817084.1, filed on June 21, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of video processing technology, and in particular to a video stream processing method, device, medium, and program product. Background Technology

[0004] A Content Delivery Network (CDN) is a distributed content delivery network built on a data network. The role of a CDN is to use streaming media server cluster technology to overcome the shortcomings of insufficient output bandwidth and concurrency of a single system, which can greatly increase the number of concurrent streams supported by the system and reduce or avoid the adverse effects of single point of failure.

[0005] With the advent of the information age, CDN is being used more widely in audio and video services. However, existing CDN systems have the following drawbacks: on the one hand, video content has low caching efficiency, long loading time, and slow distribution speed, making it impossible to balance the transmission efficiency and video quality of the video stream, which affects the viewing experience of end users; on the other hand, CDN systems transmit large amounts of data and have high bandwidth costs, which increases operating costs. Summary of the Invention

[0006] This application provides a video stream processing method, an electronic device, a computer-readable storage medium, and a computer program product.

[0007] In a first aspect, embodiments of this application provide a video stream processing method, the method comprising: receiving an original video stream; determining a primary bitstream and a long-term reference frame bitstream corresponding to the primary bitstream based on the original video stream; encapsulating the primary bitstream and the long-term reference frame bitstream to obtain media data, generating a media index file, saving the media data and the media index file, the media index file including index information; and, in response to receiving a media data request carrying the index information, indexing the media data according to the index information to obtain the media data, and sending the media data.

[0008] Secondly, embodiments of this application provide a video stream processing method, the method comprising: acquiring a media index file, wherein the media index file includes index information; in response to receiving the index information sent by a terminal, indexing media data according to the index information, wherein the media data is obtained by encapsulating the main bitstream and long-term reference frame bitstream in the original video stream; and sending the media data to the terminal.

[0009] Thirdly, embodiments of this application provide an electronic device, including: one or more processors; and a memory storing one or more programs thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the video stream processing method as described in the first or second aspect.

[0010] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video stream processing method as described in the first or second aspect.

[0011] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the video stream processing method as described in the first or second aspect. Attached Figure Description

[0012] The accompanying drawings are used to provide an understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0013] Figure 1 is a schematic diagram of a network architecture for a video stream processing method provided in an embodiment of this application;

[0014] Figure 2 is a flowchart illustrating a video stream processing method provided in an embodiment of this application;

[0015] Figure 3 is a schematic block diagram of a video stream processing method provided in an embodiment of this application;

[0016] Figure 4 is a schematic diagram of the process of encapsulating the main bit stream and the long-term reference frame bit stream to obtain media data, as shown in Figure 2.

[0017] Figure 5 is a schematic diagram of the process of sending media data in Figure 2;

[0018] Figure 6 is a flowchart illustrating a video stream processing method according to another embodiment of this application;

[0019] Figure 7 is a flowchart of step S200 in Figure 2;

[0020] Figure 8 is a schematic diagram of another video stream processing method provided in this application;

[0021] Figure 9 is a flowchart of step S600 in Figure 8;

[0022] Figure 10 is a flowchart of step S700 in Figure 8;

[0023] Figure 11 is a flowchart illustrating another video stream processing method provided in another embodiment of this application;

[0024] Figure 12 is a schematic diagram of the data flow of the video stream processing method provided in the embodiment of this application;

[0025] Figure 13 is a structural diagram of the video stream processing system provided in the embodiment of this application, divided by functional modules;

[0026] Figure 14 is a structural diagram of another video stream processing system provided in this application embodiment, divided by functional modules;

[0027] Figure 15 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0028] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions provided in this application will be described in detail below with reference to the accompanying drawings.

[0029] Exemplary embodiments will be described more fully below with reference to the accompanying drawings; however, the described exemplary embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will enable those skilled in the art to fully understand the scope of this application.

[0030] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0031] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of a feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded.

[0032] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0033] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined in the embodiments of this application.

[0034] To facilitate a better understanding of the solutions in the embodiments of this application, the relevant technologies will be introduced first below.

[0035] A Content Delivery Network (CDN) is a distributed content delivery network built on a data network. The role of a CDN is to use streaming media server cluster technology to overcome the shortcomings of insufficient output bandwidth and concurrency of a single system, which can greatly increase the number of concurrent streams supported by the system and reduce or avoid the adverse effects of single point of failure.

[0036] Long span predictive coding is an advanced feature of video coding that allows the encoder and decoder to reference frames in the sequence that are farthest from the current frame. This technique can significantly improve video compression efficiency and quality, especially when the video content does not change much or has periodically repeating content. The latest video coding standards, such as VVC (Versatile Video Coding) and AVS3 (Audio Video Coding Standard 3), have adopted this technique.

[0037] VVC is a newer video coding standard designed to provide higher compression efficiency than HEVC (High Efficiency Video Coding). VVC introduces Long-Term Reference Pictures (LDPs) to increase coding efficiency and video quality. A LDP is a reference frame with a relatively long duration within a video sequence. In VVC, LDPs can span multiple groups of consecutive frames (GOPs), providing a larger reference time span and thus offering better coding results.

[0038] AVS3 is the world's first operational audio and video source coding standard for 8K and 5G industrial applications. It is also the first officially released standard of its kind internationally, and my country possesses complete independent intellectual property rights. Compared to the international video coding standard HEVC, AVS3's coding performance is nearly 30% higher, and at the same bitrate, AVS3 video quality is significantly higher than H.265 / HEVC, thus receiving widespread industry-wide promotion. AVS3 video is designed for various application scenarios, particularly in the surveillance and security field. AVS3 video coding supports encoding and decoding of information with large spans, using knowledge images transmitted additionally at the system layer as reference frames. For example, in surveillance video, background frames are used as knowledge images for reference. Combined with efficient management of knowledge images, this provides more accurate references and improves compression efficiency. For AVS3 video streams using large-span predictive coding technology, the main bitstream and knowledge image bitstream can be merged into the same bitstream for transmission, or they can be transmitted as two independent bitstreams. Similarly, the AVS3 video main bitstream and knowledge image bitstream can also be stored as two independent bitstreams (files).

[0039] With the advent of the information age, CDN applications have become more widespread in audio and video services. However, existing CDN systems have the following drawbacks: Firstly, video content caching efficiency is low, loading time is long, and distribution speed is slow, failing to balance video stream transmission efficiency and video quality, thus affecting the end-user's viewing experience. Secondly, CDN systems transmit large amounts of data, resulting in high bandwidth costs and increased operating costs. However, for systems combining CDN technology and long-span predictive coding, the traditional caching and update mechanisms on the server side of CDN systems are not suitable for storing long-span reference frames because video stream decoding and playback rely on them for extended periods. Currently, there is no effective solution for this scenario.

[0040] Based on this, embodiments of this application provide a video stream processing method, device, medium, and program product. First, the original video stream is received. Then, the main bitstream and the corresponding long-term reference frame bitstream are determined based on the original video stream. The main bitstream and the long-term reference frame bitstream are then encapsulated to obtain and store media data. This method can combine CDN technology and large-span predictive coding technology to provide long-term caching for the long-term reference frame bitstream. By generating a media index file including index information, and finally responding to a received media data request carrying index information, the media data is obtained based on the index information and sent. This method can support the mapping between the main bitstream and the long-term reference frame bitstream during service, enabling the terminal to efficiently obtain long-term reference frames, avoiding latency and stuttering. It can improve the caching and transmission efficiency of the video stream, while providing higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0041] The following sections introduce several application scenarios and advantages of the video stream processing method, device, medium, and program products according to embodiments of this application:

[0042] 1. Video surveillance system: In video surveillance systems, scenes often last for a long time and do not change much (e.g., monitoring a corridor where little movement occurs). In such cases, using long-term reference frames can greatly improve coding efficiency.

[0043] 2. Live Sports Events: In live sports events, some images (such as the stands, panoramic views of the venue, etc.) may remain unchanged for a long time. Using knowledge images or long-term reference frames can effectively reduce the amount of data in these static areas, while reducing encoding and decoding processing time, thereby reducing latency, which is especially important for live broadcasts.

[0044] 3. Cloud gaming and virtual reality: In cloud gaming and VR scenarios, the background may be relatively static while the foreground (such as characters and interactive objects) changes significantly. By using long-term reference frames for the background, the coding focus can be placed on the dynamic content of the foreground.

[0045] 4. Video conferencing: In video conferencing, the background usually remains unchanged or changes only slightly, while the participants' faces and gestures change more frequently. Using long-term reference frame technology can improve the encoding efficiency of the background.

[0046] The combination of CDN and wide-span predictive coding technology primarily improves the transmission efficiency and viewing experience of video streams by optimizing content storage and transmission methods. The following are several ways to illustrate how these two technologies can be combined and demonstrate their synergistic effects:

[0047] 1. Optimize encoding and transmission

[0048] Large-span predictive coding technology significantly improves data compression rates by allowing video encoders to use frames that are temporally distant as references, thereby reducing video file size. CDNs can leverage this efficient coding method to store and distribute video content more efficiently, thus reducing the amount of data transmitted through the CDN, helping to save bandwidth costs, speeding up video content loading time, and improving the viewing experience for end users.

[0049] 2. Improve caching efficiency

[0050] Video files generated by large-span predictive coding technology are smaller and save more space, allowing CDN servers to cache more content under the same hardware conditions. This increases the CDN cache hit rate, making it more likely that users will obtain data from the nearest CDN node, reducing the load on the origin server, and improving the distribution speed of video content. It can effectively handle a large number of concurrent requests, especially during peak hours.

[0051] 3. Improve live streaming and real-time services

[0052] In live streaming or real-time video transmission, wide-span predictive coding technology can provide high video quality while maintaining low latency, thereby reducing live streaming latency and improving the real-time interactive experience (such as cloud gaming, online education, and real-time conferencing). Combined with the global node distribution of CDN, it can achieve rapid content distribution by distributing traffic across multiple CDN nodes, reducing the pressure on individual nodes and improving service stability.

[0053] 4. Enhance Video on Demand (VOD) services

[0054] For video-on-demand (VOD) services, wide-span predictive encoding can be stored long-term on CDN nodes for repeated playback. This encoding technology ensures high compression rates and quality even for videos stored for extended periods, thereby improving the speed and quality of VOD content access, especially in areas with high user demand. It also reduces data transmission volume, lowers operating costs, and enhances user satisfaction.

[0055] The combination of wide span predictive coding (BSC) and CDN is not merely a technological synergy, but also a strategic complementarity. Together, they improve the efficiency and quality of video transmission. This combination leverages the strengths of each technology to provide end users with a faster and higher-quality video viewing experience, while also reducing the network and storage burden on content providers.

[0056] Please refer to Figure 1, which is a schematic diagram of a network architecture for a video stream processing method provided in this application embodiment. As shown in Figure 1, the network architecture includes a server, a CDN system, and terminals. The server includes a business system and a CP / SP origin server. A CP (Content Provider) is a content provider that provides media content, such as music, movies, and TV programs; an SP (Service Provider) is a service provider that provides resources and services, such as network and communication services. The CDN system includes a management module, a content center node, edge service nodes, and a global content routing module. The management module is mainly responsible for the business operation management and maintenance management of the CDN internal network. The global content routing module, as the CDN entry point, is mainly responsible for the unified scheduling of terminal requests according to the scheduling policy. In Figure 1, RR (Route Reflector) is the route reflector, and GSLB (Global Server Load Balance) is the global load balancer. The content center node is mainly responsible for interfacing with the business system to realize CDN content access, management, storage, and proactive distribution to various edge service nodes of the CDN network. Edge service nodes, as the main entities of the CDN service, are primarily responsible for receiving terminal requests, verifying them, and providing locally cached content services to the terminals. If the content is not found, they retrieve and cache the content from the superior content center node, or redirect the request to the superior node and then provide the service. The content center node includes a first content routing module, a first content access module 1310, a first content processing module, a first content storage module 1340, and a first content distribution module 1350; the edge service nodes include a second content routing module, a second content access module, a second content processing module, a second content storage module, and a second content distribution module. Terminals include, but are not limited to, set-top boxes (STBs), PCs, internet TVs, and mobile terminals. It should be noted that the basic principles of the above network architecture are existing technologies known to those skilled in the art and will not be explained in detail here.

[0057] Please refer to Figure 2, which is a schematic flowchart of a video stream processing method provided in an embodiment of this application. The execution subject of this method can be the content center node shown in Figure 1. As shown in Figure 2, the video stream processing method provided in this embodiment includes, but is not limited to, steps S100 to S400:

[0058] Step S100: Receive the original video stream.

[0059] Referring to Figure 1, the original video stream is transmitted from the CP / SP origin server. This application embodiment supports the injection of original video streams in various application scenarios, including but not limited to the injection of live streams in live streaming services and the injection of on-demand source files. In one embodiment, after receiving a live channel creation message or an on-demand content injection message, the content center node obtains the channel's live stream or on-demand source file from the CP / SP origin server.

[0060] It is understood that live channel streams include, but are not limited to, the following stream formats: Network IP Stream (TS over UDP) / Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), Web Real-Time Communications (WebRTC), and HTTP Live Streaming (HLS); on-demand source files include, but are not limited to, the following file formats: Moving Picture Experts Group 4 (MP4), Flash Video (FLV), Matroska multimedia container file (MKV file format), and MPEG-TS transport stream (MPEG-TS).

[0061] Step S200: Determine the main bitstream and the long-term reference frame bitstream corresponding to the main bitstream based on the original video bitstream.

[0062] After receiving the raw video stream from the CP / SP source station, the receiver parses the raw video stream using a parser. If the encoding format of the raw video stream is not a large-span predictive coding format, the encoder performs large-span predictive coding on the raw video stream to obtain the main bit stream and its corresponding long-term reference frame bit stream. The main bit stream and its corresponding long-term reference frame bit stream can be merged into the same bit stream for transmission, or they can be transmitted independently as two bit streams. Similarly, the main bit stream and the long-term reference frame bit stream can also be stored as two independent bit streams (files).

[0063] In this embodiment of the application, converting the video stream into a video stream with a large span predictive coding format can provide long-term caching for the long-term reference frame bitstream and support the mapping between the main bitstream and the long-term reference frame bitstream during the service process. This enables the terminal to efficiently obtain long-term reference frames, avoids latency and stuttering, improves the caching efficiency and transmission efficiency of the video stream, provides higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0064] Step S300: Encapsulate the main bit stream and the long-term reference frame bit stream to obtain media data, generate a media index file, and save the media data and the media index file. The media index file includes index information.

[0065] In one embodiment of this application, the main bitstream and the long-term reference frame bitstream can be encapsulated and stored separately, enabling the terminal to independently and efficiently acquire the long-term reference frame, avoiding latency and stuttering, improving the caching and transmission efficiency of the video stream, and providing higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0066] In the above embodiments, the primary bitstream typically changes significantly; therefore, all primary bitstreams are usually encapsulated within a single primary bitstream streaming segment. In relatively static scenarios such as monitoring and video-on-demand, where the long-term reference frame remains unchanged for an extended period, the long-term reference frame can be considered a still image. Since an image occupies significantly less memory than a media segment, the long-term reference frame bitstream can be encapsulated in image format, reducing the amount of data transmitted via CDN, saving bandwidth costs, accelerating video content loading time, and improving the access speed and quality of video-on-demand content. In video-on-demand scenarios, where the long-term reference frame size is small, all long-term reference frames in the bitstream can be encapsulated within the same media segment for one-time distribution. In live streaming scenarios, due to the larger size of the long-term reference frame, the bitstream can be encapsulated into several media segments for distribution, thereby providing high video quality while maintaining low latency and improving the real-time interactive experience.

[0067] In another embodiment of this application, the main bitstream and the long-term reference frame bitstream can also be merged and encapsulated. When both the main bitstream and the long-term reference frame bitstream change significantly, merging and encapsulating them can reduce data transmission volume, optimize encoding efficiency, and save bandwidth costs. Simultaneously, distinguishing between the main bitstream sub-segments and the long-term reference frame sub-segments provides a long-term cache for the long-term reference frame bitstream, and supports mapping between the main bitstream and the long-term reference frame bitstream during service. This enables the terminal to efficiently acquire long-term reference frames, avoiding latency and stuttering, improving the caching and transmission efficiency of the video stream, providing higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0068] In one possible implementation of this application, in the above embodiments, the primary stream sub-fragment and the long-term reference frame sub-fragment are located on different tracks of the media segment; or, the primary stream sub-fragment and the long-term reference frame sub-fragment have different identifiers in the media segment. It should be noted that the primary stream sub-fragment and the long-term reference frame sub-fragment in the same media segment can also be distinguished in other ways, and this should not be regarded as a limitation of this application.

[0069] Media data is generated by the first content processing module and stored in the first content storage module 1340. Additionally, predictable trending content is pushed directly down from the content center node to the edge service nodes and stored in the second content storage module.

[0070] In live streaming scenarios, the caching time for media data can be preset, retaining only the media data within a specified time period of the current live stream, thereby reducing the amount of data stored.

[0071] Step S400: In response to receiving a media data request carrying index information, retrieve the media data based on the index information and send the media data.

[0072] In one embodiment of this application, the index information includes a primary streaming media data index and a long-term reference frame media data index. The primary streaming media data index is used to index and obtain the primary streaming media data, and the long-term reference frame media data index is used to index and obtain the long-term reference frame media data. The primary streaming media data index and the long-term reference frame media data index can respectively index and obtain the primary streaming media data and the long-term reference frame streaming media data, enabling the terminal to independently and efficiently acquire the long-term reference frame, avoiding latency and stuttering. This improves the caching and transmission efficiency of the video stream, while providing higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0073] It is understood that the primary streaming media data index and the long-term reference frame media data index can be in the same media index file or in two different media index files, and this should not be regarded as a limitation of this application.

[0074] Please refer to Figure 4, which is a flowchart of step S300 in Figure 2. In this embodiment, the media data includes primary streaming media data and long-term reference frame media data. The encapsulation process of the primary stream and the long-term reference frame bit stream in step S300 to obtain media data includes, but is not limited to, the following steps S310 to S320: Step S310: Perform a first encapsulation process on the primary stream to obtain primary streaming media data; Step S320: Perform a second encapsulation process on the long-term reference frame bit stream to obtain long-term reference frame media data.

[0075] In one embodiment of this application, the main bitstream and the long-term reference frame bitstream are encapsulated and stored separately, which enables the terminal to independently and efficiently obtain the long-term reference frame, avoid latency and stuttering, improve the caching efficiency and transmission efficiency of the video stream, and provide higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0076] The main stream usually varies greatly, so all the main streams are usually encapsulated in a single main streaming segment.

[0077] In relatively static scenarios such as surveillance and video-on-demand, when the long-term reference frame remains unchanged for an extended period, it can be considered as a still image. Since an image occupies much less memory than a media segment, the long-term reference frame bitstream can be encapsulated in an image format. This reduces the amount of data transmitted through the CDN, helps save bandwidth costs, speeds up video content loading time, and improves the access speed and quality of video-on-demand content.

[0078] In video-on-demand scenarios, when the capacity of long-term reference frames is not large, all long-term reference frames in the long-term reference frame bitstream can be encapsulated in the same media segment for one-time distribution.

[0079] In live streaming scenarios, since long-term reference frames (LTFFs) have a large capacity, the LFF bitstream can be encapsulated into several media segments for distribution. This can provide high video quality while maintaining low latency and improve the real-time interactive experience.

[0080] In this embodiment of the application, the index information includes a primary streaming media data index and a long-term reference frame media data index, wherein the primary streaming media data index is used to index the primary streaming media data, and the long-term reference frame media data index is used to index the long-term reference frame media data.

[0081] The primary streaming media data index and the long-term reference frame media data index can respectively index the primary streaming media data and the long-term reference frame streaming media data, enabling the terminal to independently and efficiently obtain the long-term reference frame, avoiding latency and stuttering. This can improve the caching efficiency and transmission efficiency of the video stream, while providing higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0082] It is understood that the primary streaming media data index and the long-term reference frame media data index can be in the same media index file or in two different media index files, and this should not be regarded as a limitation of this application.

[0083] In this embodiment of the application, step S400, in response to receiving a media data request carrying index information, retrieves media data according to the index information and sends the media data, including but not limited to at least one of steps S410 and S420: Step S410, in response to receiving a media data request carrying a primary streaming media data index, retrieves the primary streaming media data corresponding to the primary streaming media data index and sends the primary streaming media data; Step S420, in response to receiving a media data request carrying a long-term reference frame media data index, retrieves the long-term reference frame media data corresponding to the long-term reference frame media data index and sends the long-term reference frame media data.

[0084] In one embodiment of this application, after obtaining the media index file, the edge service node parses it to obtain the primary streaming media data index and the long-term reference frame media data index. Then, it continues to request the required primary streaming media data and long-term reference frame media data from the content center node. If the content center node finds that the first content storage module 1340 contains the primary streaming media data and long-term reference frame media data requested by the edge service node, it sends the primary streaming media data and long-term reference frame media data to the edge service node.

[0085] In one possible implementation of this application, step S310 involves performing a first encapsulation process on the primary stream to obtain primary streaming media data, including but not limited to step S311: Step S311 encapsulates the primary stream in a media segment format to obtain primary streaming media data, wherein the primary streaming media data includes at least one primary streaming media segment.

[0086] Understandably, the main stream usually varies greatly, so all the main streams are usually encapsulated in a single main streaming segment.

[0087] In one possible implementation of this application, step S320 involves performing a second encapsulation process on the long-term reference frame bitstream to obtain long-term reference frame media data, including but not limited to at least one of steps S321, S322, and S323:

[0088] Step S321: Encapsulate the long-term reference frame bitstream in image format to obtain long-term reference frame media data, which includes at least one long-term reference frame in image format.

[0089] Understandably, in relatively static scenarios such as surveillance and video-on-demand, where the long-term reference frame remains unchanged for an extended period, the long-term reference frame can be considered a still image. Since an image occupies much less memory than a media segment, encapsulating the long-term reference frame bitstream in image format can reduce the amount of data transmitted through CDN, helping to save bandwidth costs, speed up video content loading time, and improve the access speed and quality of video-on-demand content.

[0090] In one embodiment, a subset of long-term reference frames from multiple image formats can be selected for distribution, instead of sending all images at once. For example, only the long-term reference frames used in the first 3 minutes of the on-demand content can be distributed.

[0091] Step S322: Encapsulate the long-term reference frame bitstream in media segment format to obtain long-term reference frame media data. The long-term reference frame media data includes a long-term reference frame media segment, which contains all the long-term reference frames in the long-term reference frame bitstream.

[0092] Understandably, in video-on-demand scenarios, when the capacity of long-term reference frames is not large, all long-term reference frames are usually distributed at once. Therefore, all long-term reference frames in the long-term reference frame bitstream can be encapsulated in the same media segment for distribution.

[0093] Step S323: Encapsulate the long-term reference frame bitstream in media segment format to obtain long-term reference frame media data. The long-term reference frame media data includes multiple long-term reference frame media segments, and each long-term reference frame media segment contains a portion of the long-term reference frame in the long-term reference frame bitstream.

[0094] Understandably, in live streaming scenarios, due to the large capacity of long-term reference frames (LTFFs), the LFF bitstream can be encapsulated into several media segments for distribution. This can provide high video quality while maintaining low latency and improving the real-time interactive experience.

[0095] It should be noted that the encapsulation methods in steps S311, S321, S322, and S323 can be flexibly expanded, and this should not be regarded as a limitation of this application.

[0096] Figure 5 shows a flowchart of step S400 in Figure 2. In one possible implementation of this application, the long-term reference frame media data includes multiple long-term reference frames in image formats. The media data to be transmitted in step S400 includes, but is not limited to, steps S401 to S402: Step S401: Select at least one long-term reference frame as the target long-term reference frame from the multiple long-term reference frames, based on the current video scene as the first scene; Step S402: Transmit the target long-term reference frame; wherein, the first scene includes a video-on-demand scene or a monitoring scene.

[0097] Understandably, in relatively static scenarios such as surveillance and video-on-demand, a subset of long-term reference frames from multiple image formats can be selected for distribution, rather than sending all images at once. For example, only the long-term reference frames used in the first 3 minutes of video-on-demand content can be distributed, thereby reducing the amount of data transmitted through the CDN, helping to save bandwidth costs, speeding up video content loading time, and improving the access speed and quality of video-on-demand content.

[0098] It should be noted that the first scenario can also be other relatively static scenarios and should not be regarded as a limitation of this application.

[0099] In one possible implementation of this application, the long-term reference frame media data includes multiple long-term reference frame media segments, and the transmission media data in step S400 includes, but is not limited to, step S403:

[0100] Step S403: Based on the current video scene being the second scene, send long-term reference frame media segments one by one; wherein, the second scene includes the live broadcast scene.

[0101] Understandably, in live streaming scenarios, due to the large capacity of long-term reference frames, they can be split into several media segments for transmission, thereby providing high video quality and improving the real-time interactive experience while maintaining low latency.

[0102] It should be noted that the second scenario can also be other scenarios with a large long-term reference frame capacity, and should not be regarded as a limitation on this application.

[0103] In one possible implementation of this application, the long-term reference frame media data includes a long-term reference frame media segment, which contains all the long-term reference frames in the long-term reference frame bitstream. The transmission of media data in step S400 includes, but is not limited to: transmitting the long-term reference frame media segment based on the current video scene being a first scene; wherein the first scene includes a video-on-demand scene.

[0104] Understandably, in video-on-demand scenarios, when the capacity of long-term reference frames is not large, all long-term reference frames are usually distributed at once. Therefore, all long-term reference frames in the long-term reference frame bitstream can be encapsulated in the same media segment for distribution.

[0105] It should be noted that the first scenario can also be other scenarios where the long-term reference frame capacity is not large, and should not be regarded as a limitation on this application.

[0106] Please refer to Figure 6. The method of this application embodiment also includes, but is not limited to, steps S510 to S520: Step S510, in response to the storage time of the primary streaming media data exceeding the first preset aging time, delete the primary streaming media data; Step S520, in response to the storage time of the long-term reference frame media data exceeding the second preset aging time, delete the long-term reference frame media data; wherein, the first preset aging time is less than the second preset aging time.

[0107] In live streaming scenarios, as the live stream continues to play, the primary streaming media data whose cache time exceeds the preset time and the long-term reference frame media data it depends on can be deleted.

[0108] In one embodiment of this application, when long-term reference frame media data and primary streaming media data are stored separately, the primary streaming media data can be set to a shorter cache aging time, while the long-term reference frame media data, which requires long-term reliance, can be set to a longer aging time. This tiered aging mechanism can flexibly address various application scenarios and ensure the reasonable release of storage space.

[0109] In another embodiment, when long-term reference frame media data and primary streaming media data are stored together, a shorter cache aging time will be set to ensure that the cached content is up-to-date.

[0110] It should be noted that the deletion of tasks is issued by the management module.

[0111] In one possible implementation of this application, the encapsulation process of the main bit stream and the long-term reference frame bit stream in step S300 to obtain media data includes, but is not limited to, step S301: Step S301, performing a merging and encapsulation process on the main bit stream and the long-term reference frame bit stream to obtain at least one media segment, wherein each media segment includes a main bit stream sub-segment and a long-term reference frame sub-segment.

[0112] When both the main stream and the long-term reference frame (LTFrame) bitstream undergo significant changes, merging and encapsulating them can reduce data transmission volume, optimize encoding efficiency, and save bandwidth costs. Simultaneously, distinguishing between main stream sub-fragments and LFrame sub-fragments provides a long-term buffer for the LFrame bitstream, and supporting mapping between the main stream and LFrame bitstream during service allows the terminal to efficiently acquire LFrames, avoiding latency and stuttering. This improves video stream buffering and transmission efficiency, provides higher video quality, enhances the user's viewing experience, and reduces operating costs.

[0113] In one possible implementation of this application, in step S301, the main bitstream sub-fragment and the long-term reference frame sub-fragment are located on different tracks of the media segment; or, the main bitstream sub-fragment and the long-term reference frame sub-fragment have different identifiers in the media segment.

[0114] It should be noted that the merging and encapsulation of the main bitstream sub-fragments and long-term reference frame sub-fragments in the same media segment can also be distinguished in other ways, and should not be regarded as a limitation of this application.

[0115] Please refer to Figure 7, which is a flowchart of step S200 in Figure 2. In this embodiment, step S200 determines the primary bitstream and the corresponding long-term reference frame bitstream based on the original video bitstream, including but not limited to the following steps S210 to S220: Step S210: Determine the encoding format of the original video bitstream; Step S220: If the encoding format of the original video bitstream is not a large span prediction encoding format, perform large span prediction encoding on the original video bitstream to obtain the primary bitstream and the corresponding long-term reference frame bitstream.

[0116] After receiving the raw video stream from the CP / SP source station, the receiver parses the raw video stream using a parser. If the encoding format of the raw video stream is not a large-span predictive coding format, the encoder performs large-span predictive coding on the raw video stream to obtain the main bit stream and its corresponding long-term reference frame bit stream. The main bit stream and its corresponding long-term reference frame bit stream can be merged into the same bit stream for transmission, or they can be transmitted independently as two bit streams. Similarly, the main bit stream and the long-term reference frame bit stream can also be stored as two independent bit streams (files).

[0117] In this embodiment of the application, converting the video stream into a video stream with a large span predictive coding format can provide long-term caching for the long-term reference frame bitstream and support the mapping between the main bitstream and the long-term reference frame bitstream during the service process. This enables the terminal to efficiently obtain long-term reference frames, avoids latency and stuttering, improves the caching efficiency and transmission efficiency of the video stream, provides higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0118] Referring to Figure 8, which is a schematic flowchart of another video stream processing method provided in this application, the execution entity of this method can be the edge service node shown in Figure 1. As shown in Figure 8, the other video stream processing method provided in this application includes, but is not limited to, steps S500 to S700:

[0119] Step S500: Obtain the media index file, wherein the media index file includes index information.

[0120] Step S600: In response to receiving the index information sent by the terminal, the media data is obtained by indexing according to the index information. The media data is obtained by encapsulating the main bit stream and long-term reference frame bit stream in the original video bit stream.

[0121] Step S700: Send media data to the terminal.

[0122] In one embodiment of this application, the terminal requests a service from the global content routing module. The global content routing module schedules the terminal's service request to a specific edge service node to provide the corresponding service. The terminal sends a media index file request to the edge service node to initiate playback. If the second content distribution module of the edge service node finds that the second content storage module exists, it directly hits the target and sends the media index file to the terminal; if it does not exist, it requests the media index file from the upper-level content center node, caches it in the second content storage module, and simultaneously sends it to the terminal. After obtaining the media index file, the terminal parses it to obtain the index information, and then continues to request the required media data from the edge service node. If the edge service node finds that the media data requested by the terminal exists in the second content storage module, it sends the media data to the terminal. If the edge service node does not find the media data requested by the terminal in the second content storage module, it requests the media data from the upper-level content center node. The content center node also sends the request to the edge service node in the same manner as described above. After receiving the request, the edge service node caches it in the second content storage module and simultaneously sends it to the terminal.

[0123] In this embodiment of the application, converting the video stream into a video stream with a large span predictive coding format can provide long-term caching for the long-term reference frame bitstream and support the mapping between the main bitstream and the long-term reference frame bitstream during the service process. This enables the terminal to efficiently obtain long-term reference frames, avoids latency and stuttering, improves the caching efficiency and transmission efficiency of the video stream, provides higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0124] Please refer to Figure 9, which is a flowchart of step S600 in Figure 8. In this embodiment, the media data includes primary streaming media data and long-term reference frame media data, and the index information includes primary streaming media data index and long-term reference frame media data index. Step S600, in response to receiving the index information sent by the terminal, obtains the media data according to the index information, including but not limited to steps S610 and S620: Step S610, in response to receiving the primary streaming media data index sent by the terminal, obtains the primary streaming media data according to the primary streaming media data index; Step S620, in response to receiving the long-term reference frame media data index sent by the terminal, obtains the long-term reference frame media data according to the long-term reference frame media data index.

[0125] The primary streaming media data index and the long-term reference frame media data index can respectively index the primary streaming media data and the long-term reference frame streaming media data, enabling the terminal to efficiently obtain the long-term reference frame, avoiding latency and stuttering, improving the caching efficiency and transmission efficiency of the video stream, while providing higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0126] In one embodiment of this application, the terminal requests a service from the global content routing module. The global content routing module schedules the terminal's service request to a certain edge service node to provide the corresponding service. The terminal sends a media index file request to the edge service node to initiate playback. If the second content distribution module of the edge service node finds that the second content storage module exists, it directly hits the target and sends it to the terminal; if it does not exist, it requests the upper-level content center node, obtains the media index file, caches it in the second content storage module, and simultaneously sends it to the terminal. After obtaining the media index file, the terminal parses it to obtain the primary streaming media data index and the long-term reference frame media data index. Then, it continues to request the required primary streaming media data and long-term reference frame media data from the edge service node. If the edge service node finds that the primary streaming media data and long-term reference frame media data requested by the terminal exist in the second content storage module, it sends the primary streaming media data and long-term reference frame media data to the terminal. If the edge service node does not find the primary streaming media data and long-term reference frame media data requested by the terminal in the second content storage module, it requests the data from the upper-level content center node. The content center node also sends the data to the edge service node in the above manner. After receiving the data, the edge service node caches it in the second content storage module and simultaneously sends it to the terminal.

[0127] In the embodiments of this application, the primary bitstream streaming data includes at least one primary bitstream streaming segment, and the long-term reference frame media data includes one of the following: at least one long-term reference frame in an image format; a long-term reference frame media segment containing all long-term reference frames in the long-term reference frame bitstream; or multiple long-term reference frame media segments, each containing a portion of the long-term reference frames in the long-term reference frame bitstream.

[0128] In one embodiment of this application, in relatively static scenarios such as monitoring scenarios and on-demand scenarios, if the long-term reference frame remains unchanged for a long time, the long-term reference frame during this period can be regarded as a still image. Since the memory occupied by an image is much smaller than that occupied by a media segment, encapsulating the long-term reference frame bitstream in image format can reduce the amount of data transmitted through CDN, help save bandwidth costs, speed up the loading time of video content, and improve the access speed and quality of on-demand content.

[0129] In another embodiment, in the on-demand scenario, when the long-term reference frame capacity is not large, all long-term reference frames are usually distributed at once. Therefore, all long-term reference frames in the long-term reference frame bitstream can be encapsulated in the same media segment for distribution.

[0130] In another embodiment, in a live streaming scenario, since the long-term reference frame has a large capacity, the long-term reference frame bitstream can be encapsulated into several media segments for distribution, thereby providing high video quality while maintaining low latency and improving the real-time interactive experience.

[0131] In one possible implementation of this application, the long-term reference frame media data includes multiple long-term reference frames in image formats, as shown in Figure 10. Step S700 involves sending the media data to the terminal, including but not limited to steps S710 to S720: Step S710: Based on the current video scene as the first scene, select at least one long-term reference frame from the multiple long-term reference frames as the target long-term reference frame; Step S720: Send the target long-term reference frame to the terminal; wherein, the first scene includes a video-on-demand scene or a monitoring scene.

[0132] Understandably, in relatively static scenarios such as surveillance and video-on-demand, a subset of long-term reference frames from multiple image formats can be selected for distribution, rather than sending all images at once. For example, only the long-term reference frames used in the first 3 minutes of video-on-demand content can be distributed, thereby reducing the amount of data transmitted through the CDN, helping to save bandwidth costs, speeding up video content loading time, and improving the access speed and quality of video-on-demand content.

[0133] In one possible implementation of this application, the long-term reference frame media data includes multiple long-term reference frame media segments. Step S700 involves sending the media data to the terminal, including but not limited to step S730: Step S730, based on the current video scene as a second scene, sending the long-term reference frame media segments to the terminal one by one; wherein, the second scene includes a live broadcast scene.

[0134] Understandably, in live streaming scenarios, due to the large capacity of long-term reference frames, they can be split into several media segments for transmission, thereby providing high video quality and improving the real-time interactive experience while maintaining low latency.

[0135] Please refer to Figure 11. The method of this application embodiment also includes, but is not limited to, steps S810 to S820: Step S810, in response to the primary streaming media data not being accessed for more than a third preset aging time, delete the primary streaming media data; Step S820, in response to the long-term reference frame media data not being accessed for more than a fourth preset aging time, delete the long-term reference frame media data; wherein, the third preset aging time is less than the fourth preset aging time.

[0136] In video-on-demand scenarios, primary streaming media data can be set to automatically age out if there are no user accesses for 7 days. Long-term reference frame streams, with their smaller storage space, can be set to automatically age out if there are no user accesses for 14 days. This tiered aging-out mechanism flexibly addresses various application scenarios, ensuring the reasonable release of storage space.

[0137] It should be noted that the third and fourth preset aging times can be varied according to actual circumstances and should not be regarded as limitations on this application.

[0138] The following is a detailed description of a video stream processing method provided in this application, using an embodiment as an example.

[0139] In one embodiment of this application, referring to Figure 3, the receiver receives the RTMP stream from the CP / SP source station and sends it to the parser. After parsing, the parser finds that it is an H.264 stream and determines that it needs to be re-transcoded. After being transcoded by the encoder, it is converted into an AVS3 video stream using wide span predictive coding. The packer then encapsulates the AVS3 video stream into at least one MP4 format master streaming media segment, at least one knowledge image of High Efficiency Image File Format (HEIF), and a media index file describing them, and saves them in the content storage module.

[0140] Understandably, if the parser finds that the video stream is already an AVS3 video stream using large span predictive coding after parsing, then no further encoding is needed.

[0141] It should be noted that both the content center node and the edge service nodes can process the original video stream in the manner described above to obtain an AVS3 video stream with knowledge images and a primary bitstream. The format after re-encoding using wide span predictive coding technology includes, but is not limited to, VVC video streams and AVS3 video streams, and can also be video streams in other wide span predictive coding formats; this should not be considered a limitation of this application. Furthermore, the media index file can also be generated in real time according to terminal requests, and this should not be considered a limitation of this application.

[0142] Referring to Figure 12, in the above embodiment, the distribution process for the current video scenario, whether it is a live streaming scenario or a video-on-demand scenario, is as follows: The terminal requests a live streaming or video-on-demand service playback request from the global content routing module. The global content routing module will schedule the terminal's service request to a certain edge service node to provide the corresponding service. The terminal sends a media index file request to the edge service node to start playback. If the second content distribution module of the edge service node finds that the second content storage module exists, it will directly hit the target and send it to the terminal; if it does not exist, it will request from the upper-level content center node, obtain the media index file, cache it in the second content storage module, and send it to the terminal. After obtaining the media index file, the terminal parses it to obtain the primary streaming media data index and the knowledge image media data index. Then, it continues to request the required primary streaming media segment and multiple HEIF format knowledge images from the edge service node. If the edge service node finds that the second content storage module contains the primary streaming media segment and multiple HEIF format knowledge images requested by the terminal, it will send the multiple HEIF format knowledge images to the terminal one by one, and at the same time send the primary streaming media segment to the terminal. If the edge service node does not find the requested master streaming media segment and multiple HEIF format knowledge images in the second content storage module, it sends a request to the superior content center node. The content center node also sends the request to the edge service node in the same manner. After receiving the request, the edge service node caches it in the second content storage module and simultaneously sends it to the terminal. Once the terminal obtains the master streaming media segment and multiple HEIF format knowledge images from the edge service node, it can render and play them.

[0143] The embodiments of this application can combine CDN technology and large span predictive coding technology to provide long-term caching for long-term reference frame bitstreams, and support the mapping between main bitstreams and long-term reference frame bitstreams during the service process. This enables terminals to efficiently obtain long-term reference frames, avoid latency and stuttering, improve the caching efficiency and transmission efficiency of video streams, and provide higher video quality, thereby enhancing the user's viewing experience and reducing operating costs.

[0144] This application also provides a video transmission system, comprising: a server side, including a business system and a CP / SP origin server connected to the business system, the server side being used to provide the original video stream; and a CDN system, including a management module, a content center node, edge service nodes, and a global content routing module. The management module is used for business operation management and maintenance management of the CDN internal network. The global content routing module, as the CDN entry point, is used to uniformly schedule terminal requests according to scheduling policies. The content center node is used to interface with the business system to realize CDN content access, management, storage, and proactive distribution to various edge service nodes in the CDN network. Edge service nodes, as the main entities of the CDN service, are used to receive terminal requests, verify them, and provide locally cached content services to the terminals. Terminals are used to request unified scheduling from the global content routing module and request content services from the corresponding edge service nodes. It is understood that terminals include, but are not limited to, STBs (Set Top Boxes), PCs, internet TVs, and mobile terminals.

[0145] Referring to Figure 13, based on the functions of the devices in the aforementioned video transmission system, the content center node can be divided into a first content access module 1310, a large-span predictive coding module 1320, an encapsulation processing module 1330, a first content storage module 1340, and a first content distribution module 1350. The first content access module 1310 receives the original video bitstream; the large-span predictive coding module 1320 determines the main bitstream and the corresponding long-term reference frame bitstream based on the original video bitstream; the encapsulation processing module 1330 encapsulates the main bitstream and the long-term reference frame bitstream to obtain media data and generates a media index file; the first content storage module 1340 stores the media data and the media index file, the media index file including index information; and the first content distribution module 1350 responds to a media data request carrying index information, retrieves the media data based on the index information, and sends the media data.

[0146] It should be noted that both the large-span prediction encoding module 1320 and the encapsulation processing module 1330 are sub-modules of the first content processing module of the content center node.

[0147] Referring to Figure 14, based on the functions of the devices in the video transmission system described above, the edge service node can be divided into a media index file acquisition module 1410, an index module 1420, and a media data distribution module 1430.

[0148] The media index file acquisition module 1410 is used to acquire a media index file, wherein the media index file includes index information; the index module 1420 is used to obtain media data by indexing according to the index information in response to receiving the index information sent by the terminal, wherein the media data is obtained by encapsulating the main bit stream and long-term reference frame bit stream in the original video bit stream; the media data distribution module 1430 is used to send the media data to the terminal.

[0149] It should be noted that the media index file acquisition module 1410, the index module 1420, and the media data distribution module 1430 are all sub-modules of the second content distribution module of the edge service node.

[0150] This application also provides an electronic device, as shown in FIG15. The electronic device 1500 includes: one or more processors 1510; and a memory 1520 storing one or more programs. When the one or more programs are executed by the one or more processors 1510, the one or more processors 1510 implement the video stream processing method provided in any embodiment of this application.

[0151] Memory 1520, as a non-transitory network system, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory 1520 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 1520 may include remotely located memories 1520 relative to processor 1510, which can be connected to processor 1510 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0152] The memory 1520 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1520 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1520 and is called and executed by the processor 1510.

[0153] The processor 1510 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0154] In some embodiments, the electronic device further includes: an input / output interface for inputting and outputting information; a communication interface for communication and interaction between the device and other devices, which can be implemented via wired means (e.g., USB, Ethernet cable, etc.) or wireless means (e.g., mobile network, WIFI, Bluetooth, etc.); and a bus for transmitting information between various components of the device (e.g., processor 1510, memory 1520, input / output interface, and communication interface); wherein the processor 1510, memory 1520, input / output interface, and communication interface can be interconnected within the device via the bus.

[0155] One embodiment of this application also provides a computer-readable storage medium storing computer-executable instructions for executing the video stream processing method provided in any embodiment of this application.

[0156] An embodiment of this application also provides a computer program product, including a computer program or computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium and executes the computer program or computer instructions, causing the computer device to perform the video stream processing method provided in any embodiment of this application.

[0157] The system architecture and application scenarios described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will understand that as system architectures evolve and new application scenarios emerge, the technical solutions provided in this application are also applicable to similar technical problems.

[0158] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0159] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0160] The above description, with reference to the accompanying drawings, illustrates some embodiments of this application, but does not limit the scope of this application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of this application shall be within the scope of this application.

Claims

1. A video stream processing method, comprising: Receive the raw video stream; The primary bitstream and the corresponding long-term reference frame bitstream are determined based on the original video bitstream; The main bitstream and the long-term reference frame bitstream are encapsulated to obtain media data, and a media index file is generated. The media data and the media index file are saved, and the media index file includes index information. In response to receiving a media data request carrying the index information, the media data is obtained by indexing according to the index information and then sent.

2. The method according to claim 1, wherein, The media data includes primary stream data and long-term reference frame (LTFF) media data; the encapsulation process of the primary stream and the LFF bitstream to obtain the media data includes: The primary stream is subjected to a first encapsulation process to obtain the primary streaming media data; The long-term reference frame bitstream is subjected to a second encapsulation process to obtain the long-term reference frame media data.

3. The method according to claim 2, wherein, The index information includes a primary streaming media data index and a long-term reference frame media data index, wherein the primary streaming media data index is used to index the primary streaming media data, and the long-term reference frame media data index is used to index the long-term reference frame media data.

4. The method according to claim 3, wherein, The step of responding to receiving a media data request carrying the index information, indexing the media data according to the index information, and sending the media data includes at least one of the following: In response to receiving a media data request carrying the primary streaming media data index, the primary streaming media data is sent according to the primary streaming media data index; In response to receiving a media data request carrying the long-term reference frame media data index, the long-term reference frame media data corresponding to the long-term reference frame media data index is sent.

5. The method according to claim 2, wherein, The first encapsulation process of the primary stream to obtain the primary streaming media data includes: The master stream is encapsulated in a media segment format to obtain the master streaming media data, which includes at least one master streaming media segment.

6. The method according to claim 2, wherein, The second encapsulation process performed on the long-term reference frame bitstream to obtain the long-term reference frame media data includes at least one of the following: The long-term reference frame bitstream is encapsulated in an image format to obtain the long-term reference frame media data, which includes at least one long-term reference frame in an image format. The long reference frame bitstream is encapsulated in a media segment format to obtain long reference frame media data. The long reference frame media data includes a long reference frame media segment, which contains all the long reference frames in the long reference frame bitstream. The long-term reference frame bitstream is encapsulated in a media segment format to obtain long-term reference frame media data. The long-term reference frame media data includes multiple long-term reference frame media segments, and each long-term reference frame media segment contains a portion of the long-term reference frame in the long-term reference frame bitstream.

7. The method according to claim 6, wherein, The long-term reference frame media data includes multiple long-term reference frames in image formats, and transmitting the media data includes: Based on the current video scene as the first scene, at least one of the long-term reference frames is selected from multiple long-term reference frames as the target long-term reference frame. Send the target long-term reference frame; The first scenario includes either a video-on-demand scenario or a monitoring scenario.

8. The method according to claim 6, wherein, The long-term reference frame media data includes multiple long-term reference frame media segments, and transmitting the media data includes: Based on the current video scene being the second scene, the long-term reference frame media segments are sent one by one; The second scenario includes a live streaming scenario.

9. The method according to claim 2, further comprising: In response to the primary streaming media data being stored for a duration exceeding a first preset aging time, the primary streaming media data is deleted. In response to the fact that the storage time of the long-term reference frame media data exceeds the second preset aging time, the long-term reference frame media data is deleted. Wherein, the first preset aging time is less than the second preset aging time.

10. The method according to claim 1, wherein, The encapsulation process of the primary bitstream and the long-term reference frame bitstream to obtain media data includes: The primary bitstream and the long-term reference frame bitstream are merged and encapsulated to obtain at least one media segment, wherein each media segment includes a primary bitstream sub-segment and a long-term reference frame sub-segment.

11. The method according to claim 10, wherein, The primary stream sub-fragment and the long-term reference frame sub-fragment are located on different tracks of the media segment; Alternatively, the primary stream sub-fragment and the long-term reference frame sub-fragment may have different identifiers in the media segment.

12. The method according to claim 1, wherein, The step of determining the primary bitstream and the corresponding long-term reference frame bitstream based on the original video bitstream includes: Determine the encoding format of the original video stream; If the encoding format of the original video bitstream is not a large span prediction encoding format, large span prediction encoding is performed on the original video bitstream to obtain the main bitstream and the long-term reference frame bitstream corresponding to the main bitstream.

13. A video stream processing method, comprising: Obtain a media index file, wherein the media index file includes index information; In response to receiving the index information sent by the terminal, media data is obtained by indexing according to the index information, wherein the media data is obtained by encapsulating the main bit stream and the long-term reference frame bit stream in the original video bit stream; The media data is sent to the terminal.

14. The method according to claim 13, wherein, The media data includes primary streaming media data and long-term reference frame media data, and the index information includes primary streaming media data index and long-term reference frame media data index. The step of responding to receiving the index information sent by the terminal and obtaining media data based on the index information includes: In response to receiving the primary streaming media data index sent by the terminal, the primary streaming media data is obtained according to the primary streaming media data index; In response to receiving the long-term reference frame media data index sent by the terminal, the long-term reference frame media data is obtained according to the long-term reference frame media data index.

15. The method according to claim 14, wherein, The primary streaming media data includes at least one primary streaming media segment, and the long-term reference frame media data includes one of the following: At least one long-term reference frame in an image format; A long-term reference frame media segment, the long-term reference frame media segment containing all long-term reference frames in the long-term reference frame bitstream; Multiple long-term reference frame media segments, each of the long-term reference frame media segments containing a portion of the long-term reference frame in the long-term reference frame bitstream.

16. The method according to claim 15, wherein, The long-term reference frame media data includes multiple long-term reference frames in image formats, and sending the media data to the terminal includes: Based on the current video scene as the first scene, at least one of the long-term reference frames is selected from multiple long-term reference frames as the target long-term reference frame. Send the target long-term reference frame to the terminal; The first scenario includes either a video-on-demand scenario or a monitoring scenario.

17. The method according to claim 15, wherein, The long-term reference frame media data includes multiple long-term reference frame media segments, and sending the media data to the terminal includes: Based on the current video scene being the second scene, the long-term reference frame media segments are sent to the terminal one by one; The second scenario includes a live streaming scenario.

18. The method of claim 14, further comprising: In response to the primary streaming media data not being accessed for more than a third preset aging time, the primary streaming media data is deleted; In response to the fact that the long-term reference frame media data has not been accessed for more than a fourth preset aging time, the long-term reference frame media data is deleted. The third preset aging time is less than the fourth preset aging time.

19. An electronic device comprising: One or more processors; A memory having stored thereon one or more programs that, when executed by one or more processors, cause the one or more processors to implement: the video stream processing method as described in any one of claims 1 to 12, or the video stream processing method as described in any one of claims 13 to 18.

20. A computer-readable storage medium having a computer program stored thereon, the program, when executed by a processor, implementing: the video stream processing method as described in any one of claims 1 to 12, or the video stream processing method as described in any one of claims 13 to 18.

21. A computer program product comprising a computer program that, when executed by a processor, implements: the video stream processing method as described in any one of claims 1 to 12, or the video stream processing method as described in any one of claims 13 to 18.

Citation Information

Patent Citations

  • Transpression method of video code stream and system of same

    CN101969559A

  • Video coding

    CN113905241A

  • Video bit stream packaging, decoding and accessing method and device

    CN115842922A

  • Method and system for supporting media data of multi-coding formats

    CN1949876A

  • Stream handling using an intermediate format

    US20170324796A1