Audio and video synchronization method during playing mpegts at browser side
The synchronization system adjusts video frame playback speed or drops frames to maintain audio-visual synchronization in browser-based MPEG2-TS playback, addressing desynchronization issues caused by CPU and GPU fluctuations.
Patent Information
- Application Number
- CN202510463629.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
Existing browser-based MPEG2-TS video playback technologies face challenges in maintaining audio-visual synchronization due to CPU and GPU usage, leading to unpredictable frame rate fluctuations and audio-visual desynchronization.
A method involving a synchronization system that decodes audio and video frames, records their presentation timestamps, and adjusts video frame playback based on a threshold to maintain synchronization, including separate decoding on the browser and audio processor sides.
Effectively synchronizes audio and video by adjusting video frame playback speed or dropping frames when necessary, ensuring accurate synchronization despite CPU and GPU fluctuations.
Smart Images

Figure CN120302094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio and video playback, and specifically to an audio and video synchronization method for playing mpegts on the browser side. Background Art
[0002] MPEG2-TS (Transport Stream) is a communication protocol for audio, images, and data. The mpegts video format is widely used in the field of radio and television. Currently, major browsers do not directly support playing this type of video data. Therefore, to play mpegts videos on the browser side, it is necessary to implement the decoding of mpegts audio and video and render the decoded audio and video data. When we implement a player using software decoding, the decoded data is pcm audio data and yuv420p video data. The pcm audio data can be implemented using AudioContext, and the yuv420p video data is rendered using webgl.
[0003] However, in the prior art, most local players synchronize audio with video. On the browser side, we adopt the same strategy. First, AudioContext is directly controlled by the CPU, and the timeline is relatively stable and not easily affected by decoding, etc. Second, there are not many APIs provided by AudioContext that we can control, and it is not easy to achieve audio acceleration, deceleration, and time shift. Ideally, audio and video can be synchronized when played according to two timelines respectively. Therefore, some implementations do not perform actual audio and video synchronization. Instead, the audio is played linearly, and the video is played according to the frame rate. Under the condition that the computer or browser is not affected by other factors, this situation is fine. However, in reality, the browser is affected by decoding, CPU, GPU occupancy, etc., resulting in too long video rendering time and unexpected increase in frame rate, and finally audio and video deviation. Summary of the Invention
[0004] The purpose of the present invention is to provide an audio and video synchronization method for playing mpegts on the browser side, so as to solve the problem that the browser is affected by decoding, CPU, GPU occupancy, etc., resulting in too long video rendering time, unexpected increase in frame rate, and finally audio and video deviation proposed in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: An audio and video synchronization method for playing mpegts on the browser side, including the following steps:
[0006] Step 1: Decode using the synchronization system module, retain the PTS of the decoded audio frames and video frames, and save them together with the frame data;
[0007] Step 2: Record the PTS when the audio starts playing, denoted as start_pts, with the unit of ms;
[0008] Step 3: After starting to play, obtain the currentTime value of the current AudioContext and convert it to milliseconds, denoted as current_time;
[0009] Step 4: Obtain the PTS of the next video frame and convert it to milliseconds, denoted as next_pts;
[0010] Step 5: Set a threshold in the browser's player, denoted as threshold, with the unit of milliseconds, and it is allowed to be negative;
[0011] Step 6: Among them, next_pts - start_pts - current_time is the time in milliseconds when the next video frame will be played after the current time, denoted as n_ms, with the unit of milliseconds;
[0012] Step 7: If n_ms is negative, judge according to the set threshold. If it exceeds the set threshold, discard this frame and fetch the next frame for preparation to play. Otherwise, set the playing time according to the value of n_ms;
[0013] Step 8: After playing one frame, repeat S3 to S8 until the audio or video finishes playing.
[0014] Preferably, the working process of the synchronization system module is specifically divided into:
[0015] Step 1: Demultiplex the audio and video;
[0016] Step 2: Convert the format of the audio stream and download it;
[0017] Step 3: The audio processor combines the collected human voice data and the audio stream downloaded by the browser;
[0018] Step 4: Perform synchronous correction processing on the video playback data and the new audio playback data during the playback process;
[0019] Step 5: Synthesize and save the new audio and video streams.
[0020] Preferably, the synchronization system module includes a demultiplexing module, a video decoding and filtering module, a first format conversion module, a first transmission module, an audio and video stream merging module, a user interaction module, and a synchronization module;
[0021] The demultiplexing module is used to obtain and demultiplex the media source file to obtain a video stream and an audio stream;
[0022] The video decoding and filtering module is used to perform video decoding and video filtering on the video stream to obtain video playback data and play the video playback data;
[0023] The first format conversion module is used to perform format conversion and encoding compression on the audio stream;
[0024] The first transmission module is used to transmit the audio stream after format conversion and encoding compression to the audio processor. The browser converts the obtained audio stream into a format suitable for streaming transmission and then transmits it to the audio processor. At this time, the transmitted data is the encoded and compressed audio stream data, reducing the network bandwidth occupancy. If only the audio data at the current playback moment is transmitted in real time, when the network is congested and unstable, the audio data cannot reach the audio processor in real time. The method adopted by the first transmission module is to serve the audio stream transmission process to the greatest extent, transmit as much audio stream data as possible to the audio processor in advance and save it as a file, reducing the probability of audio-video desynchronization;
[0025] The audio-video stream merging module is used to receive the new audio playback data. The audio processor simultaneously performs audio encoding on the new audio playback data, and then converts it into an audio stream suitable for streaming transmission through the first format conversion module and transmits it to the browser in real time. At this time, the transmitted data is the encoded audio data, greatly reducing the bandwidth occupancy. After the video playback data and the new audio playback data are played, the two are combined for audio-video synthesis to generate a new media file;
[0026] The user interaction module is used to implement the interaction between the synchronization system module and the user;
[0027] The synchronization module is used to perform initial synchronization processing on the video playback data and the audio playback data at the start, and perform synchronization correction processing during the playback of the video playback data and the audio playback data.
[0028] Preferably, the user interaction module includes the following:
[0029] The audio decoding and filtering module is used to call the decoder to decode and filter the audio stream to obtain audio playback data and play the audio playback data;
[0030] The mixing and sound effect adjustment module is used to perform synthesis processing on the audio playback data and the human voice data to obtain new audio playback data;
[0031] The second format conversion module is used to perform audio encoding and format conversion on the new audio playback data;
[0032] The second transmission module is used to transmit the new audio playback data after audio encoding and format conversion to the browser;
[0033] A voice collection device for collecting human voice data.
[0034] Preferably, the synchronization module includes the following:
[0035] An initial synchronization unit module, which is used to calculate the average value Δt of the time difference Δt=t2 - t1 according to the frame number in the video start frame data packet, the time t1 corresponding to the video start frame, the frame number in the audio start frame data packet, and the time t2 corresponding to the audio start frame by using the averaging method, and perform frame rate reduction processing on the video start frame according to the average value Δt.
[0036] A synchronization correction unit module, which is used to estimate the time that the current video frame leads the current audio frame according to the current playback time t of the current video frame v and the playback time t of the current audio frame corresponding to the current video frame a obtain the average time Δt' that the current video frame leads the current audio frame by using the averaging method va , and according to the average time Δt' va reduce the frame rate of video playback or perform frame skipping processing.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. In the present invention, when the video is not synchronized with the audio, the playback of the next video frame is accelerated or decelerated, and a threshold is set. When the playback time pts of the video exceeds the set threshold, the outdated video frames are discarded, and then the rendering time of the next video frame is calculated. In this way, it is possible to effectively adjust the playback time of the video frames when the audio and video are not synchronized, so as to achieve audio and video synchronization.
[0039] 2. In the present invention, the decoding of the audio and video is separately realized on the browser side and the audio processor, and they are not executed synchronously. The browser side is responsible for playing the video picture, and the audio stream is sent to the audio processor for playback. The synchronization module is used to ensure the accurate synchronization of the audio and video.
[0040] 3. In the present invention, the frame rate of the video frame is reduced instead of processing on the audio processor in order to obtain a better user experience. At the moment of initial synchronization correction, the slight delay of the picture has a very slight impact on the user. When Δt' is small enough, the user may not even notice it. After ensuring the initial synchronization of the audio and video, the audio and video are synchronized and corrected during the decoding and playback process. Description of the Drawings
[0041] Figure 1 It is a flow chart of an audio and video synchronization method for playing mpegts on the browser side according to the present invention.
[0042] In the figure: 1. Synchronization system module; 11. Demultiplexing module; 12. Video decoding and filtering module; 13. First format conversion module; 14. First transmission module; 15. Audio and video stream merging module; 16. User interaction module; 161. Audio decoding and filtering module; 162. Mixing and sound effect adjustment module; 163. Second format conversion module; 164. Second transmission module; 165. Sound acquisition device; 17. Synchronization module; 171. Initial synchronization unit module; 172. Synchronization correction unit module. Detailed implementation mode
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment 1
[0045] Refer to Figure 1 As shown: An audio and video synchronization method when playing mpegts on the browser side includes the following steps:
[0046] (1) Decode using the synchronization system module 1, retain the pts of the decoded audio frames and video frames, and save them together with the frame data;
[0047] (2) Record the pts when the audio starts playing, denoted as start_pts, with the unit of ms;
[0048] (3) After starting to play, obtain the currentTime value of the current AudioContext and convert it to milliseconds, denoted as current_time;
[0049] (4) Obtain the pts of the next video frame and convert it to milliseconds, denoted as next_pts;
[0050] (5) Set a threshold in the browser player, denoted as threshold, with the unit of ms, and it is allowed to be negative;
[0051] (6) Among them, next_pts - start_pts - current_time is the time when the next video frame will be played after the current time in milliseconds, denoted as n_ms, with the unit of ms;
[0052] (7) If n_ms is negative, it is judged according to the set threshold. If it exceeds the set threshold, this frame is discarded and the next frame is taken for playback preparation. Otherwise, the playback time is set according to the value of n_ms;
[0053] (8) After playing one frame, repeat S3 to S8 until the audio or video is finished playing.
[0054] In this embodiment, the pts when the audio starts playing is recorded and converted into ms. Then the decoded pcm buffer data is added to the AudioContext. When starting to play, the currentTime of the audio context is taken out, and the pts converted into milliseconds and the currentTime are added together to obtain the time corresponding to the currently playing video. Then we take out the pts of the next video frame and convert it into milliseconds, and then subtract the milliseconds pts obtained from the previous addition from the milliseconds pts of this frame, which is how many seconds there are until the next video frame starts to be rendered from the current time;
[0055] This method accelerates or decelerates the playback of the next video frame when the video is not synchronized with the audio, and sets a threshold. When the playback time pts of the video exceeds the set threshold, the outdated video frames are discarded, and then the rendering time of the next video frame is calculated. In this way, it can effectively adjust the playback time of the video frame when the audio and video are not synchronized, so as to achieve audio and video synchronization.
[0056] Embodiment 2
[0057] The working process of the synchronization system module is specifically divided into:
[0058] (1) Demultiplex the audio and video;
[0059] (2) Convert the format of the audio stream and download it;
[0060] (3) The audio processor combines the collected human voice data and the audio stream downloaded by the browser;
[0061] (4) Synchronization correction processing for the video playback data and the new audio playback data during the playback process;
[0062] (5) Synthesize and save the new audio and video stream.
[0063] In this embodiment, the decoding of the audio and video is separately implemented on the browser side and the audio processor, and they are not executed synchronously. The browser side is responsible for playing the video picture, and the audio stream is sent to the audio processor for playback, and the synchronization module is used to ensure the accurate synchronization of the audio and video;
[0064] When the user interacts with the browser side, the audio processor records the user's voice and sends it back to the browser side, and finally the audio and video are merged on the browser side.
[0065] Embodiment III
[0066] The synchronization system module includes a demultiplexing module 11, a video decoding and filtering module 12, a first format conversion module 13, a first transmission module 14, an audio and video stream merging module 15, a user interaction module 16, and a synchronization module 17;
[0067] The demultiplexing module 11 is used to obtain and demultiplex the media source file to obtain a video stream and an audio stream;
[0068] The video decoding and filtering module 12 is used to perform video decoding and video filtering on the video stream to obtain video playback data and play the video playback data;
[0069] The first format conversion module 13 is used to perform format conversion and encoding compression on the audio stream;
[0070] The first transmission module 14 is used to transmit the format-converted and encoded-compressed audio stream to the audio processor. The browser converts the obtained audio stream into a format suitable for stream transmission and then transmits it to the audio processor. At this time, the transmitted data is the encoded-compressed audio stream data, which reduces the network bandwidth occupancy. If only the audio data at the current playback moment is transmitted in real time, when the network is congested and unstable, the audio data cannot reach the audio processor in real time. The method adopted by the first transmission module 14 is to serve the audio stream transmission process to the greatest extent, transmit as much audio stream data as possible to the audio processor in advance and save it as a file, reducing the probability of audio and video desynchronization;
[0071] The audio and video stream merging module 15 is used to receive the new audio playback data. The audio processor simultaneously performs audio encoding on the new audio playback data, and then converts it into an audio stream suitable for stream transmission through the first format conversion module 13 and transmits it to the browser in real time. At this time, the transmitted data is the encoded audio data, which greatly reduces the bandwidth occupancy. After the video playback data and the new audio playback data are played, the two are combined for audio and video synthesis to generate a new media file;
[0072] The user interaction module 16 is used to realize the interaction between the synchronization system module and the user, and it includes the following:
[0073] An audio decoding and filtering module 161, which is used to call a decoder to decode and filter the audio stream to obtain audio playback data and play the audio playback data;
[0074] A mixing and sound effect adjustment module 162, which is used to perform synthesis processing on the audio playback data and the human voice data to obtain new audio playback data;
[0075] The second format conversion module 163 is used to perform audio encoding and format conversion on the new audio playback data;
[0076] The second transmission module 164 is used to transmit the new audio playback data after audio encoding and format conversion to the browser;
[0077] The voice acquisition device 165 is used to acquire voice data;
[0078] The synchronization module 17 is used to perform initial synchronization processing on the video playback data and the audio playback data at the start, and perform synchronization correction processing during the playback of the video playback data and the audio playback data. It includes the following:
[0079] The initial synchronization unit module 171 is used to calculate the average value Δt of the time difference Δt = t2 - t1 according to the frame number in the video start frame data packet and the time t1 corresponding to the video start frame, the frame number in the audio start frame data packet and the time t2 corresponding to the audio start frame by using the averaging method, and perform frame rate reduction processing on the video start frame according to the average value Δt.
[0080] The synchronization correction unit module 172 is used to estimate the time by which the current video frame is ahead of the current audio frame according to the current playback time t of the current video frame v and the playback time t of the current audio frame corresponding to the current video frame a obtain the average time Δt′ by which the current video frame is ahead of the current audio frame by using the averaging method va and reduce the frame rate of the video playback or perform frame skipping processing according to the average time Δt′ va .
[0081] In this embodiment, when the video starts playing, a frame data packet including the frame number and the corresponding time t1 (ms) of the frame is sent while the first frame of the video is being played. When this frame data packet arrives at the audio processor, the audio processor immediately decodes and plays the audio stream, and at the same time sends back a frame data packet to the browser side in the same format, with the frame number being the same as the frame number sent from the browser side. When the browser side receives the frame data packet sent back by the audio processor, it obtains the corresponding time t1 of the frame in the data packet, and the difference from the current system time t2 is Δt = t2 - t1. To more accurately and reasonably calculate this time difference and exclude the interference of some accidental situations, the method of sampling average is used to repeat this process. This repetition operation is performed on the first 10 frames of the video, and the statistical average value of the 10 groups obtained is calculated to obtain the average value Δt' of the time difference. If Δt' exceeds a certain preset threshold, for example, 50 ms, the frame rate of the video frame playback is slowed down. Suppose the current video frame is p1, and this frame is maintained within the time of Δt' / 2 to achieve the purpose of audio-video start synchronization. The reason for slowing down the frame rate of the video frame instead of processing it in the audio processor is to obtain a better user experience. At the moment of start synchronization correction, the impact of a certain degree of delay in the picture on the user is very subtle, and when Δt' is small enough, the user may not even notice it. After ensuring the start synchronization of the audio and video, the audio and video are synchronized and corrected during the decoding and playback process.
[0082] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An audio and video synchronization method when playing mpegts on the browser side, characterized in that, It includes the following steps: S1. Decode using the soft decoding module (1). The decoding module outputs yuv420p video frames and display time pts, PCM audio frames and playback time pts; S2. Combine multiple audio frames with consecutive playback times into one audio node in sequence and place it into the AudioContext for playback. S3. During playback, the synchronization time is based on the pts of the first frame of the audio node and the currentTime of the AudioContext at the end of the playback of the previous audio node. When the audio is played forward linearly, the pts corresponding to the current audio is calculated based on the above two times; S4. Take out the pts of the video frames that are close and not yet played; S5. Set a threshold in the browser player; S6. Compare the PTS of the video frames taken out in S4 with the corresponding audio PTS in S3; S7. If the compared pts in S6 are within the set threshold, play this video frame, otherwise discard the rendering of this frame; S8. When the audio node in S3 finishes playing, generate a new node according to S3 for playback.
2. The audio - video synchronization method during mpegts playback on the browser side according to claim 1, wherein The working process of the soft decoding system module is specifically divided into: A1. Segmentedly load the mpegts video file; A2. Demultiplex the downloaded audio and video; A3. After demultiplexing, find out the encoding methods used for the audio and video and select the corresponding decoders; A4. The soft decoder decodes the corresponding video and audio files; A5. Output audio and video frames and the corresponding playback time pts.
3. A method for audio - video synchronization when playing mpegts on the browser side according to claim 2, characterized in that, The soft decoding system module includes a demultiplexing module (11), a video decoding and video filtering and audio decoding and filtering module; (Only keep 3 modules, other modules are not needed) The demultiplexing module (11) is used to obtain and demultiplex the media source file to obtain a video stream and an audio stream; The video decoding and filtering module (12) is used to perform video decoding and video filtering on the video stream to obtain video playback data and play the video playback data; The first format conversion module (13) is used to perform format conversion and encoding compression on the audio stream; The first transmission module (14) is used to transmit the audio stream after format conversion and encoding compression to the audio processor. The browser performs format conversion on the obtained audio stream, adjusts it into a format suitable for streaming transmission, and then transmits it to the audio processor; The audio and video stream merging module (15) is used to receive the new audio playback data. The audio processor simultaneously performs audio encoding on the new audio playback data, and then converts it into an audio stream suitable for streaming transmission through the first format conversion module (13) and transmits it to the browser in real time; The user interaction module (16) is used to realize the interaction between the synchronization system module and the user; The synchronization module (17) is used to perform initial synchronization processing on the video playback data and the audio playback data at the start, and perform synchronization correction processing during the playback of the video playback data and the audio playback data.
4. The audio - video synchronization method when playing mpegts on the browser side according to claim 3, characterized in that, The user interaction module (16) includes the following: The audio decoding and filtering module (161) is used to call the decoder to decode and filter the audio stream to obtain audio playback data and play the audio playback data; The mixing and sound effect adjustment module (162) is used to synthesize the audio playback data and the human voice data to obtain new audio playback data; The second format conversion module (163) is used to perform audio encoding and format conversion on the new audio playback data; The second transmission module (164) is used to transmit the new audio playback data after audio encoding and format conversion to the browser; The voice acquisition device (165) is used to acquire the human voice data.
5. The audio - video synchronization method when playing mpegts on the browser side according to claim 3, characterized in that, The synchronization module (17) includes the following: The initial synchronization unit module (171) is used to calculate the average value Δt of the time difference Δt = t2 - t1 by using the average method according to the frame number in the video start frame data packet and the time t1 corresponding to the video start frame, the frame number in the audio start frame data packet and the time t2 corresponding to the audio start frame, and perform frame rate reduction processing on the video start frame according to the average value Δt; The synchronization correction unit module (172) is used to, according to the current playback time t of the current video frame v , the playback time t of the current audio frame corresponding to the current video frame a estimate the time by which the current video frame leads the current audio frame, and use the averaging method to obtain the average time Δt′ by which the current video frame leads the current audio frame va , and according to the average time Δt′ va slow down the frame rate of video playback or perform frame skipping processing.