A video playing method, device and storage medium

The method enhances web-based video playback by parallel processing of media slices to support advanced codecs and allow independent audio and video editing, addressing bandwidth and compatibility issues.

CN115086282BActive Publication Date: 2025-07-15XUNLEI NETWORKING TECHNOLOGIES LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110281970.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-16
Publication Date
2025-07-15
Estimated Expiration
2041-03-16

AI Technical Summary

Technical Problem

The prior art cannot process audio and video data separately when playing audio and video on the web side, resulting in insufficient editing capabilities and excessive bandwidth resources.

Method used

By obtaining media slice data, decapsulation and decoding processing is used to use parallel threads to separate video data and audio data, and play it simultaneously, supporting the H265 encoding format.

Benefits of technology

It realizes the separation of audio and video data, enhances editing capabilities, and is especially suitable for the separate switching and editing of multilingual videos, saving bandwidth resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115086282B_ABST
    Figure CN115086282B_ABST
Patent Text Reader

Abstract

The present application discloses a video playback method, device, and storage medium. The video playback method includes obtaining current media slice data, performing de-encapsulation and decoding processing on the current media slice data to obtain video data and audio data, and synchronously playing the audio data and the video data. By the above method, the present application can process the audio data and the video data separately, enhancing the audio-visual editing ability; and slicing the media data, resulting in higher playback fluency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio - video data processing, and particularly to a video playing method, device, and storage medium. Background Art

[0002] With the development of network technology, the demand for playing audio - video on the web - end is increasing. Currently, the solutions for playing audio - video on the web - end mainly include using Flash plugins, the video tag of HTML5, and server - side decoding, etc.

[0003] Among them, using Flash plugins does not support H265, and mainstream browsers no longer update it and gradually cancel their support; the video encoding types that can be decoded using the video tag of HTML5 are limited, for example, it cannot support H265; while after server - side decoding, the data volume is too large, occupying too much bandwidth resource; and currently, the audio and video in the audio - video played on the web - end are an integral whole, and it is impossible to edit and switch the audio or video separately. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a video playing method, device, and storage medium, which can process audio data and video data separately and enhance the editing ability of audio - video.

[0005] To solve the above - mentioned technical problem, a technical solution adopted by this application is: to provide a video playing method, including obtaining current media slice data, performing demultiplexing and decoding processing on the current media slice data to obtain video data and audio data, and synchronously playing the audio data and video data.

[0006] Among them, obtain an index file of the slice position of the video file, and sequentially obtain media slice data according to the index file.

[0007] Among them, after obtaining the index file of the slice position of the video file, it further includes creating a first thread and a second thread, using the first thread to obtain media slice data, and using the second thread to perform demultiplexing and decoding processing on the obtained current media slice data.

[0008] Among them, the first thread and the second thread work in parallel to concurrently use the second thread to perform demultiplexing and decoding processing on the obtained current media slice data, and use the first thread to obtain the next media slice data according to the index file.

[0009] Among them, performing demultiplexing and decoding processing on the current media slice data includes creating a first object and a second object, demultiplexing the current media slice data into video encoded data and audio encoded data through the first object, and decoding the video encoded data and audio encoded data into video data and audio data through the second object.

[0010] Among them, the video encoding data is an H265 encoded video file. The process of demultiplexing and decoding the current media slice data includes demultiplexing the media slice data into video encoding data and audio encoding data through demuxe.js of the first object, and decoding the video encoding data and audio encoding data into video data and audio data through ffmpeg.wasm of the second object.

[0011] Among them, the video file is a TS video file. Decoding the video encoding data and audio encoding data into video data and audio data through the second object includes decoding and converting the video encoding data of the media slice data into yuv data using Web Assembly, and drawing the yuv data into a picture using yuv-canvas.

[0012] Among them, decoding the video encoding data and audio encoding data into video data and audio data through the second object includes decoding the audio encoding data of the media slice data into audio data using the Web Audio API.

[0013] Among them, the first object and the second object work in parallel to concurrently decode the video encoding data and audio encoding data obtained by demultiplexing the current media slice into video data and audio data using the second object, and demultiplex the next media slice data using the first object.

[0014] Among them, synchronously playing the audio data and the video data includes controlling the video data to synchronize the timestamp of the audio data for playing the picture.

[0015] The beneficial effect of this application is: Different from the prior art situation, this application obtains the current media slice data, performs demultiplexing and decoding processing on the current media slice data to obtain video data and audio data, and synchronously plays the audio data and the video data. Through the above method, the audio and video data can be divided into audio data and video data, especially for multi-language videos, the audio or video can be separately switched and edited, enhancing the editing ability of the audio and video. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic flowchart of an embodiment of the video playing method of this application;

[0017] Figure 2 is a schematic flowchart of another embodiment of the video playing method of this application;

[0018] Figure 3 is a schematic structural diagram of a video playing device in an embodiment of this application;

[0019] Figure 4It is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present application;

[0020] Figure 5 It is a schematic structural diagram of a video playback device in an embodiment of the present application. Detailed implementation manners

[0021] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples.

[0022] To facilitate the understanding of the technical solutions provided in the embodiments of the present application, the following will first briefly explain the technical terms involved in the embodiments of the present application.

[0023] Coding format: The coding format usually includes two types, namely video coding format and audio coding format. Among them, the video coding format is also known as the video coding specification. Since the original video data is very large and inconvenient for transmission and storage, the original video is compressed by means of compression coding. The video coding format defines the specifications of video data during storage and transmission. Common video compression formats include H.264 and H.265.

[0024] The audio coding format is also known as the audio coding specification, which compresses the original audio data and defines the specifications of audio data during storage and transmission. Common audio compression formats include AAC and MP3.

[0025] The encapsulation format (also called a container) encapsulates the original video data and audio data into a file after compression coding, such as TS, AVI, RMVB, MP4, etc.

[0026] It should be noted that the method provided in the embodiments of the present application can be applied to any scenario on the web side and is not limited to the web side player. For the sake of easy understanding, the following will take the application to the web side player as an example for explanation.

[0027] Refer to Figure 1 , Figure 1 It is a schematic flowchart of a video playback method according to an embodiment of the present application. The video playback method includes the following steps:

[0028] Step S101: Obtain the current media slice data;

[0029] The web end sends a network data request to the server end to obtain the current media slice data. Media slice data is to divide the media data into several slice data. When the user watches a segment, it loads a segment, and the video is seamlessly played. When it is paused or jumped out, it will not continue to load, which saves bandwidth to a greater extent. Specifically, media slice data is the audio and video data that has been encoded and encapsulated. There are many types of encapsulation formats, such as MP4, MKV, RMVB, TS, FLV, AVI, etc. Its function is to put the compressed and encoded video data and audio data together in a certain format.

[0030] Step S102: decapsulate and decode the current media slice data to obtain video data and audio data;

[0031] The acquired media slice data is decapsulated and decoded to obtain video data and audio data. Decapsulation is the reverse process of encapsulation, which involves disassembling the protocol packet, processing the information in the packet header, and extracting the business information data in the payload. Encapsulation and decapsulation are a pair of reverse processes. Decapsulation can separate the input encapsulation format data into audio compression encoded data and video compression encoded data; decoding is decoding the video / audio compression encoded data into uncompressed video / audio original data. Decoding can include soft decoding and hard decoding. Soft decoding means that the CPU is used to decode the video through software, and then the GPU is called to render and merge the video and display it on the screen. Hard decoding means that the video decoding task is completed independently through a dedicated sub-card device without the help of the CPU.

[0032] Step S103: synchronously playing the audio data and the video data.

[0033] The video and audio data are obtained by decapsulating and decoding the media slice data, and the audio and video data are played synchronously to complete the video playback on the Web side.

[0034] The present application obtains media slice data, decapsulates and decodes the current media slice data, obtains video data and audio data, and synchronously plays the audio data and video data. In the above manner, the audio and video data can be divided into audio data and video data. In particular, for multi-language videos, the audio or video can be switched and edited separately, thereby enhancing the audio and video editing capabilities. The media data is sliced, and a section is loaded when the user watches a section, so that the video has a good seamless playback experience. When the video is paused or jumped out, the loading is no longer continued, which saves bandwidth to a greater extent.

[0035] In one embodiment, the demuxing of media slice data is implemented through the demuxe.js library, and the media slice data is demuxed into video encoded data and audio encoded data; the decoding is implemented through the wasm software decoding generated by ffmpeg packaging, and the video encoded data and audio encoded data are decoded into video data and audio data, with a relatively high CPU usage rate.

[0036] Specifically, the decoding process of the media slice data includes decoding and converting the video encoded data of the media slice data into yuv data by using Web Assembly, and using yuv-canvas to draw the yuv data into a picture. WebAssembly is a bytecode standard that runs in the browser relying on a virtual machine in the form of bytecode. It can rely on compilers such as Emscripten to compile strongly typed languages such as C++ / Golang / Rust / Kotlin into WebAssembly bytecode (.wasm file). Therefore, WebAssembly is not Assembly (assembly). In browsers that do not support containers and codecs, an efficient video decoding module (C / C++ code) is compiled into WebAssembly to decode the RTP media stream into yuv data in real time. Among them, on the Web side, the yuv video data is converted into rgb data by using WebGL, and then the video picture is drawn on the canvas; during the decoding process, the audio and video timestamps for decoding this frame of the picture are generated, which is convenient for the Web side to accurately achieve audio and video synchronization during playback. Specifically, the decoding process of the media slice data includes decoding the audio encoded data of the media slice data into audio data and playing it by using the Web Audio API.

[0037] Specifically, synchronously playing the audio data and the video data includes: controlling the video data to synchronize the timestamp of the audio data for playing the picture. After the audio data is sent into the player constructed by the Web Audio API, the timestamp of the currently playing audio can be obtained through the Web Audio API, and the video frames are synchronized based on this timestamp. If the time of the current video frame has fallen behind, it is immediately rendered; if it is earlier, it needs to be postponed.

[0038] Next, taking the playback of an MPEG-TS (TS) video file by a Web player as an example, the present application will be further described. Please refer to Figure 2 , Figure 2 is a schematic flowchart of another embodiment of the video playback method of the present application. The video playback method includes the following steps:

[0039] S201: The Web player creates a worker thread;

[0040] A thread is the smallest unit that an operating system can perform operation scheduling on. It is contained within a process and is the actual operating unit within the process. A thread refers to a single sequential control flow within a process. Multiple threads can run concurrently within a process, and each thread executes different tasks in parallel. The role of a Worker thread is to create a multi-threaded environment. It allows the main thread to create Worker threads and assign some tasks to the latter for execution. While the main thread is running, the Worker thread runs in the background without interfering with each other. When the Worker thread completes the calculation task, it returns the result to the main thread. The advantage of this is that some computationally intensive or high-latency tasks are borne by the Worker thread, and the main thread (usually responsible for UI interaction) will be very smooth without being blocked or slowed down. In this embodiment, the web player is initialized and two worker threads are created, namely httpWorker and demuxWorker. Among them, httpWorker is responsible for network data requests, that is, for the Web player to obtain media slice data from the server; demuxWorker is responsible for unpacking and decoding the media slice data obtained by the Web player. Specifically, demuxWorker is initialized, and two objects, demuxer and decode, are created to unpack and decode the obtained media slice data. Among them, demuxer realizes the unpacking of media slice data through the demuxe.js library to obtain video encoded data and audio encoded data; decode decodes the video encoded data and audio encoded data. Among them, the common decoding methods are soft decoding (ffmpeg) and hard decoding (MediaCodec, MediaPlayer). Soft decoding means that the CPU is used to decode the video through software, and hard decoding means that part of the video data that was originally all processed by the CPU is handed over to the GPU for processing. Hard decoding is very efficient. This can not only reduce the burden on the CPU, but also has the characteristics of low power consumption and less heat generation. However, due to the relatively late start of hard decoding, the software and drivers have very low support for it, and compatibility problems often occur. In addition, the filters, subtitles, and picture quality of hard decoding are not ideal enough. Soft decoding requires a large amount of video information to be calculated, so it has very high requirements for the CPU processing performance. The huge amount of calculation will cause problems such as low conversion efficiency and high heat generation. However, soft decoding does not require much hardware support and has very high compatibility. Moreover, soft decoding has rich filter, subtitle, and picture processing optimization effects, and can achieve more excellent picture effects. In this embodiment, decode is implemented by the wasm soft decoding generated by ffmpeg packaging.

[0041] Step S203: Request TS video slices based on the range.

[0042] The server generates an m3u8 file based on the byte slice positions of the file, but does not generate the specific slice data. Specifically, the m3u8 file refers to an m3u file in UTF-8 encoding format. The m3u file is a plain text file that records an index. When opened, the playback software does not play it directly, but finds the network address of the corresponding audio and video files according to its index for online playback. Multi-bitrate adaptation can be performed. According to the network bandwidth, the client will automatically select a file with a suitable bitrate for playback to ensure the smoothness of the video stream. Specifically, through event delegation and postMessage, the httpworker requests and obtains the corresponding slice data of the TS file according to the range parameter recorded in the m3u8 file.

[0043] Step 205: Demux and decode the obtained TS file slices.

[0044] After the Web player obtains the TS file slices, they are sent into the demuxWorker thread, and the demuxer demuxes the audio and video. Among them, the audio data is sent into the player constructed by the Web Audio API. The timestamp of the currently playing audio can be obtained through the Web Audio API, and the video frames are synchronized based on this timestamp. For the video, it relies on webAssembly to decode and convert it into yuv data, and uses yuv-canvas to draw the yuv data into a picture to obtain a set of yuv videos. Among them, on the Web side, the yuv video data is converted into rgb data using WebGL and the video picture is drawn on the canvas. During the decoding process, the audio and video timestamps for decoding this frame of the picture are generated, which facilitates the Web side to accurately achieve audio and video synchronization during playback.

[0045] Step 207: Synchronously play the audio data and video data

[0046] After the audio data is sent into the player constructed by the Web Audio API, the timestamp of the currently playing audio can be obtained through the Web Audio API. Based on this timestamp, the video frames are synchronized. If the time of the current video frame has fallen behind, it is immediately rendered. If it is earlier, it needs to be postponed.

[0047] Furthermore, the request to load the next TS slice file can be considered based on the change in the number within the yuv video set. Specifically, when the number of videos within the yuv video set is lower than a certain threshold, the next TS slice file is obtained to ensure the continuity of the audio and video playback by the Web player.

[0048] In the above embodiments, an m3u8 file is generated on the server based on the byte slice positions of the file, but no specific slice data is generated. The m3u8 file is a plain text file that records an index and does not retain a large number of slice files on the server, saving space. After obtaining the TS file slice data, the Web player unpacks and decodes it, and then synchronously plays the video data and audio data, enabling direct playback of TS video files and H265 format videos on the Web. When playing a TS video, it can directly start playing from the corresponding slice at the time of jumping, reducing the waiting time and improving the response efficiency. In the above solution, by constructing a player through the Web Audio API, API methods for controlling the player can be provided externally, facilitating the customization of UI operation controls. After dividing the TS file data into audio data and video data, the audio data or video data can be processed separately. For multilingual videos, the audio can be switched separately without the need to switch the entire audio and video. It is also possible to control the picture quality / rate, etc. of the audio and video, enhancing the editing ability of the audio and video.

[0049] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a video playback device in an embodiment of the present application. Specifically, the video playback device includes an acquisition module 31 for acquiring media slice data generated by the server, a parsing module 32 for unpacking and decoding the media slice data, and a playback module 33 for playing video data and audio data.

[0050] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present application. The computer-readable storage medium 41 of the embodiment of the present application stores instructions / program data 42, and when the instructions / program data 42 are executed, the methods provided by any embodiment of the video playback method of the present application and any non-conflicting combination are implemented. Among them, the instructions / program data 42 can form a program file and be stored in the above storage medium 41 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods of various embodiments of the present application. The foregoing storage medium 41 includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.

[0051] Please refer to Figure 5 , Figure 5It is a schematic structural diagram of a video playback device in an embodiment of the present application. In this embodiment, the video playback device 50 includes a processor 51.

[0052] The processor 51 can also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 51 can also be any conventional processor, etc.

[0053] The video playback device 50 may further include a memory (not shown in the figure) for storing instructions and data required for the operation of the processor 51.

[0054] The processor 51 is used to execute instructions to implement the method provided by any embodiment and any non-conflicting combination of the above-mentioned field strength monitoring methods of the present application.

[0055] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the shown or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0056] In addition, the functional units in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0057] The above is only the embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the description and drawings of the present invention, or directly or indirectly applied to other related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. A video playing method, characterized in that, Including: Obtaining an index file of the slice positions of a video file, and sequentially obtaining media slice data according to the index file; Obtaining current media slice data, and performing demultiplexing and decoding processing on the current media slice data to obtain video data and audio data; Synchronously playing the audio data and the video data; Wherein, the obtaining of the current media slice data and the performing of demultiplexing and decoding processing on the current media slice data includes: Creating a first thread and a second thread, using the first thread to obtain media slice data, and using the second thread to perform demultiplexing and decoding processing on the obtained current media slice data; the first thread is httpWorker, and the second thread is demuxWorker; The first thread and the second thread work in parallel to concurrently use the second thread to perform demultiplexing and decoding processing on the obtained current media slice data, and use the first thread to obtain the next media slice data according to the index file; The performing of demultiplexing and decoding processing on the current media slice data includes: Creating a first object and a second object; Demultiplexing the current media slice data into video encoded data and audio encoded data through the first object; Decoding the video encoded data and the audio encoded data into video data and audio data through the second object; Wherein, the first object is demuxer, the second object is decode, and the decode performs software decoding on the video encoded data and the audio encoded data through ffmpeg.wasm to obtain video data and audio data.

2. The video playback method according to claim 1, wherein The video encoded data is an H265 encoded video file, and the performing of demultiplexing and decoding processing on the current media slice data includes: Demultiplexing the media slice data into video encoded data and audio encoded data through demuxe.js of the first object.

3. The video playing method according to claim 1, wherein The video file is a TS video file, and the decoding of the video encoded data and the audio encoded data into video data and audio data through the second object includes: Using WebAssembly to decode and convert the video encoded data of the media slice data into yuv data, and using yuv-canvas to draw the yuv data into a picture.

4. The video playback method according to claim 1, wherein The first object and the second object work in parallel to concurrently use the second object to decode the video encoded data and the audio encoded data of the current media slice demultiplexed into video data and audio data, and use the first object to demultiplex the next media slice data.

5. The video playing method according to claim 1, wherein The synchronously playing the audio data and the video data includes: Controlling the video data to synchronize the time stamp of the audio data for playing the picture.

6. A video player, characterized in that, Including a processor, the processor is used to execute instructions to implement the video playing method according to any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store instructions / program data, and the instructions / program data can be executed to implement the video playing method according to any one of claims 1-5.

Citation Information

Patent Citations

  • A method and apparatus for decoding streaming media

    CN109088887A

  • Video processing method and device, electronic equipment and computer readable storage medium

    CN110996160A

  • Browser video playing method and device thereof and computer storage medium

    CN111641838A