Media file processing method, device, electronic device and storage medium

By acquiring and decoding media files, determining the timestamps of audio and video frames in a frame set, and calculating the time difference between frames, the problem of low efficiency in audio and video synchronization detection is solved, efficient and accurate audio and video synchronization detection is achieved, and the user experience is improved.

CN114339212BActive Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111669944.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-10-03
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The existing technology has low efficiency in audio-visual synchronization detection, and the subjective evaluation method is time-consuming and inaccurate, which makes it difficult to meet the needs of audio and video evaluation.

Method used

By acquiring and decoding media files, determining the timestamps of audio and video frames in a frame set, calculating the time difference between frames, excluding silent frames, using audio feature points to determine video frames, and calculating the time difference of audio and video synchronization deviation after full-link processing, detection accuracy and efficiency are improved.

Benefits of technology

It achieves efficient and accurate audio and video synchronization detection, improves user experience, and is suitable for audio and video synchronization evaluation of live broadcasts and video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114339212B_ABST
    Figure CN114339212B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method device, electronic device and storage medium for processing media files, which relates to the field of computers, and in particular to the field of voice technology. A specific implementation scheme is as follows: obtaining a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file; determining a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame in the first media file with the same timestamp as the first audio frame, the second frame set includes a second video frame in the second media file with the same timestamp as the second audio frame, and the audio content of the first audio frame is the same as the audio content of the second audio frame; based on the first frame set and the second frame set, determining the time difference between the timestamp of the first video frame and the timestamp of the second video frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method and apparatus for processing media files, an electronic device, and a storage medium in the field of voice technology. Background Art

[0002] At present, the demand for analyzing user experience data such as audio and video evaluation is gradually increasing. Among them, audio and video synchronization is one of the more intuitive user experience indicators.

[0003] In the related art, the method of evaluating the matching degree of sound and picture is usually through subjective viewing, which has the technical problem of low efficiency in audio and picture synchronization detection. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, device, and storage medium for processing media files.

[0005] According to one aspect of the present disclosure, a method for processing media files is provided. The method may include: obtaining a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file; determining a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame in the first media file having the same timestamp as the first audio frame, the second frame set includes a second video frame in the second media file having the same timestamp as the second audio frame, and the audio content of the first audio frame is the same as the audio content of the second audio frame; and determining a time difference between the timestamp of the first video frame and the timestamp of the second video frame based on the first frame set and the second frame set.

[0006] According to another aspect of the present disclosure, a media file processing device is also provided. The device may include: an acquisition unit for acquiring a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file; a first determination unit for determining a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame having the same timestamp as the first audio frame in the first media file, the second frame set includes a second video frame having the same timestamp as the second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame; a second determination unit for determining the time difference between the timestamp of the first video frame and the timestamp of the second video frame based on the first frame set and the second frame set.

[0007] According to another aspect of the present disclosure, an electronic device is provided. The electronic device may include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the media file processing method of an embodiment of the present disclosure.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided, wherein the computer instructions are used to enable a computer to execute the media file processing method of an embodiment of the present disclosure.

[0009] According to another aspect of the present disclosure, a computer program product is provided, which may include a computer program. When the computer program is executed by a processor, the media file processing method of the embodiment of the present disclosure is implemented.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 is a flow chart of a method for processing a media file according to an embodiment of the present disclosure;

[0013] Figure 2 This is a flowchart of a full-link processing of audio and video synchronization detection according to an embodiment of the present disclosure;

[0014] Figure 3 is a flow chart of audio-video synchronization detection according to an embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of an overall system for detecting audio and video synchronization according to an embodiment of the present disclosure;

[0016] Figure 5 This is a diagram of the duration of audio and video asynchrony and user perception criteria defined according to international standards;

[0017] Figure 6 is a schematic diagram of a media file processing device according to an embodiment of the present disclosure;

[0018] Figure 7 The present invention is a block diagram of an electronic device that implements a method for processing a media file according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0020] Figure 1 is a flow chart of a method for processing a media file according to an embodiment of the present disclosure. Figure 1 As shown, the method may include the following steps:

[0021] Step S102: Acquire a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file.

[0022] In the technical solution provided in step S102 above of the present disclosure, the second media file is subjected to full-link processing to obtain a first media file, wherein the first media file may be audio and video data collected at the playback end, or may be audio and video data collected from video shooting and live broadcast recording, and may also be referred to as playback end media file B, and may include video data and audio data; the second media file may be preset audio and video data, or may be customized video content, and may also be referred to as reference media file A, and may include video data and audio data.

[0023] Optionally, the second media file is transmitted to the playback end after full-link processing to generate the first media file. That is, after the reference media file A is processed through the full-link, the playback end media file B is collected at the playback end. The full-link processing can be the entire processing process of replacing, beautifying, and other forward processing and backward processing of the data in the second media file to obtain the first media file.

[0024] Step S104: Determine a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame having the same timestamp as the first audio frame in the first media file, the second frame set includes a second video frame having the same timestamp as the second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame.

[0025] In the technical solution provided in step S104 of the present disclosure, the second media file and the first media file obtained after full-link processing are analyzed and processed separately to obtain a first frame set of the first media file and a second frame set of the second media file. The first frame set may be a set of frames containing sound in the first media file, and the second frame set may be a set of frames containing sound in the second media file.

[0026] Optionally, the first frame set includes a first video frame in the first media file having the same timestamp as the first audio frame; the second frame set includes a second video frame in the second media file having the same timestamp as the second audio frame, wherein the audio content of the first audio frame is the same as the audio content of the second audio frame.

[0027] Optionally, the first media file is analyzed and processed to determine a set of video frames containing audio frames in the first media file, that is, to determine a first video frame in the first media file having the same timestamp as the first audio frame, thereby obtaining a first frame set. The second media file is analyzed and processed to determine a set of video frames containing audio frames in the second media file, that is, to determine a second video frame in the second media file having the same timestamp as the second audio frame, thereby obtaining a second frame set. Since the first media file is obtained by processing the second media file, the audio content of the first audio frame in the first media file and the second audio frame in the second media file are the same.

[0028] Step S106 : determining a time difference between a timestamp of the first video frame and a timestamp of the second video frame based on the first frame set and the second frame set.

[0029] In the technical solution provided in step S106 of the present disclosure, a timestamp corresponding to each video frame in the first frame set and a timestamp corresponding to each video frame in the second frame set are determined, and the timestamp corresponding to each video frame in the first frame set and the timestamp corresponding to each video frame in the second frame set are analyzed and compared to obtain a time difference between the timestamps of each video frame, that is, a time difference between the frames containing the sound. The time difference can be used to represent a synchronization deviation time difference between the first media file and the second media file, which can also be called an audio-video time difference.

[0030] Optionally, by analyzing and comparing the time difference between the timestamps of each video frame, the synchronization deviation time difference between the first media file and the second media file can be determined, thereby improving the efficiency of audio and video synchronization detection.

[0031] Through the above steps S102 to S106, the first media file and the second media file are obtained, wherein the first media file is obtained by performing full-link processing on the second media file; a first frame set of the first media file and a second frame set of the second media file are determined, wherein the first frame set includes a first video frame with the same timestamp as the first audio frame in the first media file, and the second frame set includes a second video frame with the same timestamp as the second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame; based on the first frame set and the second frame set, the time difference between the timestamp of the first video frame and the timestamp of the second video frame is determined. In other words, the present disclosure obtains the time difference of audio and video synchronization deviation between the first media file and the second media file by comparative analysis of the second media file and the first media file, thereby solving the technical problem of low efficiency of audio and video synchronization detection and achieving the technical effect of improving the efficiency of audio and video synchronization detection.

[0032] The above method of this embodiment is further described in detail below.

[0033] As an optional embodiment, the method also includes: step S106, determining the time difference between the timestamp of the first video frame and the timestamp of the second video frame based on the first frame set and the second frame set, including: determining the first identifier of the first video frame in the first frame set; determining the second identifier of the second video frame in the second frame set; determining the time difference based on the first identifier and the second identifier.

[0034] In this embodiment, a first identifier of a first video frame in a first frame set is determined; a second identifier of a first video frame in a second frame set is determined; and a time difference is determined based on the first identifier and the second identifier, wherein the first identifier is used to identify a video frame corresponding to a sound frame that first appears in the second media file. For example, taking a video with a frame rate of 10 fps as an example, assuming that the second identifier of the video frame of the second media file is 10, it means that the video frame corresponding to the audio frame that first appears in the second media file is the 10th frame in a cycle; the second identifier is used to identify a video frame corresponding to an audio frame that first appears in the first media file. For example, taking a media file with a frame rate of 10 fps as an example, assuming that the first identifier of the video frame of the first media file is 2, it means that the video frame corresponding to the audio frame that first appears in the first media file is the 2nd frame in a cycle.

[0035] Optionally, by processing the first media file and the second media file, a first identifier of the first video frame in the first frame set in the second media file (the video frame rate of the second media file) and a second identifier of the first video frame in the second frame set in the first media file (the video frame rate of the first media file) are obtained. Based on the obtained video frame rates, the total duration is calculated according to the frame interval to determine the synchronization offset time difference between the first media file and the second media file. For example, taking a 10fps media file as an example, assuming that the original frame sequence of the video frames containing audio in the reference media file is [10,10,10,10,10,10,10,10,10,10,…], and the sequence of video frames containing audio in the playback-end media file is [2,3,2,2,3,2,2,2,2,2,2,2,2,2,2,2,…], then the synchronization offset time difference between the first media file and the second media file is ≈[(2+10)-10]*(1000ms / 10).

[0036] It should be noted that the time difference calculated based on the first identifier and the second identifier is within the protection scope of the present disclosure, and no examples are given here one by one.

[0037] It should be noted that the frequency of the media files collected at the playback end may be 15fps, 24fps or other frequencies. The frame rate will change to a certain extent after transcoding, network compatibility packet loss and other processing during transmission. Therefore, there will be a certain error in the comparison between the inter-frame duration and the media file before and after the reference media file and the media file on the playback end. For each scene, the error needs to be analyzed and processed separately based on the actual playback frame interval and frame loss and frame tracking situation to reduce the error. In addition, high-frequency video is used as much as possible in the selection of reference video to narrow the error range to below 25ms.

[0038] As an optional implementation, the method further includes: determining the first identifier of the first video frame in the first frame set includes: determining a value of the first identifier of the first video frame in the first frame set based on a first frame rate of the first media file.

[0039] In this embodiment, the value of the first identifier of the first video frame in the first frame set is determined based on the first frame rate of the first media file, wherein the first frame rate is used to represent the frequency of the first media file and can be the frequency collected at the collection end.

[0040] As an optional implementation, the method further includes: determining the second identifier of the first video frame in the second frame set based on the second frame rate of the second media file, including: determining the value of the second identifier of the first video frame in the second frame set based on the second frame rate of the second media file.

[0041] In this embodiment, the value of the second identifier of the first video frame in the second frame set is determined based on the second frame rate of the second media file, wherein the second frame rate is used to represent the frequency of the second media file, which can be the frequency of live broadcast or video scene playback.

[0042] As an optional embodiment, the method also includes: in response to the value of the first identifier being greater than the value of the second identifier, determining that the timestamp of the first video frame is ahead of the timestamp of the second video frame; in response to the value of the first identifier being less than the value of the second identifier, determining that the timestamp of the first video frame is behind the timestamp of the second video frame.

[0043] In this embodiment, if the value of the first identifier is greater than the value of the second identifier, it can be determined that the timestamp of the first video frame is ahead of the timestamp of the second video frame, that is, the audio and video data collected by the playback end is ahead of the preset audio data; if the value of the first identifier is less than the value of the second identifier, it can be determined that the timestamp of the first video frame lags behind the timestamp of the second video frame, that is, the audio and video data collected by the playback end lags behind the preset audio data.

[0044] Optionally, within the same cycle, if the value of the first identifier of the audio and video data collected by the playback end is greater than the value of the second identifier of the preset audio data, it may indicate that the sound of the playback end media file obtained after full-link processing is ahead of the preset picture by a certain length of time; conversely, if the value of the first identifier of the audio and video data collected by the playback end is less than the value of the second identifier of the preset audio data, it may indicate that the sound of the playback end media file obtained after full-link processing lags behind the preset picture by a certain length of time.

[0045] As an optional implementation, the method further includes: determining a first frame set of the first media file, including: decoding the first media file to obtain first audio data and first video data; and determining the first frame set based on the first audio data and the first video data.

[0046] In this embodiment, the first media file is decoded to obtain first audio data and first video data. The first audio data is analyzed, and then the first video data is judged. If the video frame timestamp set of the first video data is not within the silent frame timestamp set of the first audio data, a first frame set is obtained. The first media file may also be a data encoding file.

[0047] Optionally, the decoder address can be obtained according to the editor, and the decoder address can be used to decode the first media file to obtain decoding information such as the context, frame rate, and time base of the first media file. Based on the decoding information such as the context, frame rate, and time base of the first media file, the audio data and video data of the first media file can be determined.

[0048] As an optional implementation, the method also includes: step S104, determining a first frame set based on the first audio data and the first video data, including: determining a first target video frame in the first video data based on the target audio parameters of the first media file, wherein the first target video frame does not match an audio frame in the first audio data; determining the first frame set based on the timestamp of the video frame in the first video data and the timestamp of the first target video frame.

[0049] In this embodiment, the target audio parameters of the first media file can be used to judge the first video data in the first media file, determine that there is no matching audio frame in the first audio data, and obtain the first target video frame, wherein the target audio parameters can be the wavelength and amplitude of the preset video sound, used to determine the timestamp range of the first target video frame; the first target video frame can be a silent frame, used to indicate that there is no matching audio frame in the first audio data; the timestamp of the video frame in the first video data can be a video frame timestamp set, used to indicate the video frame set corresponding to the silent frame.

[0050] Optionally, based on the target audio parameters, it is determined in the first video data that there is no match for an audio frame in the first audio data to obtain a first target video frame, and a first frame set is obtained based on the timestamp of the video frame in the first video data and the timestamp of the first target video frame.

[0051] For example, preset audio feature points may be determined based on feature points of wavelength and amplitude of preset video sound, and then a range of silent frame timestamps may be acquired based on the preset audio feature points to obtain the first frame set.

[0052] As an optional embodiment, the method also includes: step S104, determining the first frame set based on the timestamp of the video frame in the first video data and the timestamp of the first target video frame, including: determining the video frame in the first video data whose timestamp exceeds the timestamp of the first target video frame as the first video frame in the first frame set.

[0053] In this embodiment, the video frame in the first video data whose timestamp exceeds the timestamp of the first target video frame is determined as the first video frame in the first frame set. That is, if the video frame timestamp set is not in the silent frame timestamp set, the set of video frames where the audio is located can be inferred to obtain the first frame set.

[0054] As an optional implementation, the method further includes: determining a second frame set of the second media file, including: decoding the second media file to obtain second audio data and second video data; and determining the second frame set based on the second audio data and the second video data.

[0055] In this embodiment, the second media file is decoded to obtain second audio data and second video data. The second audio data is analyzed, and then the second video data is judged. If the video frame timestamp set of the second video data is not within the silent frame timestamp set in the second audio data, a second frame set is obtained. The second media file may also be a data encoding file.

[0056] Optionally, the decoder address can be obtained according to the editor, and the decoder address can be used to decode the second media file to obtain decoding information such as the context, frame rate, and time base of the second media file. The audio data and video data of the second media file can be determined based on the decoding information such as the context, frame rate, and time base of the second media file.

[0057] As an optional embodiment, the method also includes: determining a second frame set based on the second audio data and the second video data, including: determining a second target video frame in the second video data based on the target audio parameters of the second media file, wherein the second target video frame does not match the audio frame in the second audio data; determining the second frame set based on the timestamp of the video frame in the second video data and the timestamp of the second target video frame.

[0058] In this embodiment, the target audio parameters of the second media file can be used to judge the second video data in the second media file, determine that there is no matching audio frame in the second audio data, and obtain a second target video frame, wherein the target audio parameters can be the wavelength and amplitude of the preset video sound, used to determine the timestamp range of the second target video frame; the second target video frame can be a silent frame, used to indicate that there is no matching audio frame in the second audio data; the timestamp of the video frame in the second video data can be a video frame timestamp set, used to indicate the video frame set corresponding to the silent frame.

[0059] Optionally, based on the target audio parameters, it is determined that there is no matching audio frame in the second audio data in the first video data to obtain a second target video frame, and a second frame set is obtained based on the timestamp of the video frame in the second video data and the timestamp of the second target video frame.

[0060] For example, preset audio feature points may be determined based on feature points of wavelength and amplitude of preset video sound, and then the range of silent frame timestamps may be acquired based on the preset audio feature points to obtain the second frame set.

[0061] As an optional embodiment, the method also includes: determining a second frame set based on the timestamp of the video frame in the second video data and the timestamp of the second target video frame, including: determining the video frame in the second video data whose timestamp exceeds the timestamp of the second target video frame as the second video frame in the second frame set.

[0062] In this embodiment, the video frame in the second video data whose timestamp exceeds the timestamp of the second target video frame is determined as the second video frame in the second frame set. That is, if the video frame timestamp set is not in the silent frame timestamp set, the set of video frames where the audio is located can be inferred to obtain the second frame set.

[0063] In this embodiment, by comparing and analyzing the second media file with the first media file, the time difference between the audio and video synchronization deviations of the first media file and the second media file is obtained, thereby solving the technical problem of low audio and video synchronization detection efficiency and achieving the technical effect of improving audio and video synchronization detection efficiency.

[0064] The above technical solutions of the embodiments of the present disclosure are further introduced below with reference to preferred embodiments.

[0065] As the popularity of video and live streaming software continues to rise, interactive scenarios such as beauty and motion effects are becoming more and more abundant. The demand for user experience data analysis such as effectiveness evaluation of various scenarios and tools, audio and video evaluation, is also gradually increasing. Among them, picture synchronization in live streaming and video viewing is one of the more intuitive user experience indicators.

[0066] In relevant situations, the main situations of audio and video asynchrony may include: audio and video are collected simultaneously at the acquisition end, but after forward processing such as beauty, filters, transcoding, encoding, etc., certain reasons cause the audio and video timestamps to be asynchronized, or the timestamps do not increase gradually, resulting in freezes or frame loss on the playback end, resulting in audio and video asynchrony; network transmission problems, due to packet loss, delay, etc., resulting in audio and video data packets not being decoded synchronously on the playback end; decoding logic problems on the playback end or the performance of the playback end cannot support high-definition stream decoding, as well as frame loss processing problems, etc., resulting in audio and video asynchrony.

[0067] In related technologies, the playback effects of video shooting and live recording are usually evaluated by subjectively evaluating the matching degree of sound and picture. However, this method has the technical problem of low efficiency in audio and picture synchronization detection.

[0068] To address the problem of low efficiency in audio-visual synchronization detection, a method is proposed to respectively obtain the audio frame acquisition time point and image frame of the received audio and video signals, encode the image frame and audio frame to generate a detection video, and compare the time difference between the time point of the image frame and the audio acquisition point to determine whether the audio and video are synchronized. However, this method takes the time points of audio and video separately, and the selection of reference time points may lead to certain errors.

[0069] A method is also proposed, which inserts marker codes into sound and video, calculates the transmission delay of audio and video respectively according to the arrival time of the audio and video marker codes, and calculates the time difference between audio and video based on the delay. However, if this method is used on mobile software, it requires code level analysis, which may deviate from the actual user experience.

[0070] A liveness recognition method has also been proposed, that is, detecting the time difference between sound and motion picture through lip reading. However, this method has the problem of requiring a high accuracy rate for liveness recognition.

[0071] A lip shape recognition method was also proposed, that is, the length and size of the human mouth are detected based on the picture to obtain the predicted time point, the sound is denoised by machine learning to retain the human voice, and the two time points are compared. However, this method has the problem of high accuracy requirements for lip shape recognition.

[0072] A method is also proposed to export the audio stream through an external hardware device or software processing method, scan the sound spectrum, and compare the sound spectrum within the estimated range with the video frame timestamp alignment time difference. However, this method requires intermediate conversion of the sound during the process of exporting the audio stream through an external device or software. The conversion process can easily lead to timestamp deviation, thereby affecting the final test results. Although it solves the frame loss problem, it is not conducive to automated processing.

[0073] Therefore, in order to solve the above problems, the present disclosure proposes a method for detecting audio and video synchronization, which is mainly used to detect the audio and video synchronization time difference of the video content watched by users after live broadcast and video have been processed through the whole link. By excluding silent video frames, the video screen where the sound is located is inferred, and compared with the reference video, the time difference of the sound and picture synchronization deviation after the whole link processing such as acquisition, encoding, transmission, decoding, enhancement and playback is calculated, thereby improving the efficiency of audio and video synchronization detection and user experience.

[0074] Figure 2 This is a flowchart of a full-link processing of audio and video synchronization detection according to an embodiment of the present disclosure. Figure 2 As shown, the full-link processing of audio and video synchronization detection can include.

[0075] Step S201: perform audio and video acquisition and data processing in the data stream to complete data acquisition redirection.

[0076] The camera data replacement assists in audio and video synchronization detection. The mobile software audio and video synchronization detection requires pre-processing through camera imitation (mock) processing technology. The live broadcast and video scene change upload content, and audio and video are collected from the data stream uploaded by the camera. The collected data is redirected to complete the data collection redirection. Among them, data redirection specifies the replacement of audio and video collection content.

[0077] Step S202: intercept the audio and video acquisition data and direct it to the target data source to replace the data.

[0078] Intercept audio and video acquisition data and point to the preset video with the inserted tag to complete the data replacement.

[0079] Step S203: retain the forward processing logic of the business layer camera and perform forward processing such as beautification.

[0080] The forward processing logic of the business layer camera is retained to perform forward processing such as beautification on the video data.

[0081] Step S204: transmit to the playback end for backward processing.

[0082] The data after a series of processing is transmitted to the playback end, and the data transmitted to the playback end is processed backward.

[0083] Step S205: collect data at the playback end, and then perform audio and video synchronization detection.

[0084] Collect the data played on the playback end, and draw the final conclusion by performing audio and video synchronization detection on the preset video and the playback content collected by the playback end.

[0085] It should be noted that on other direct streaming tools, the data replacement step can be omitted, and customized video content can be pushed directly, and then subsequent data comparison and analysis on the playback end can be performed.

[0086] The following is a further introduction to audio and video synchronization detection. Figure 3 This is a flow chart of an audio-video synchronization detection according to an embodiment of the present disclosure. Figure 3 As shown, audio-visual synchronization detection may include.

[0087] Step S301: noise filtering is performed on the reference video and the playback video.

[0088] Step S302: Compare the audio and video of the reference video and the playback video.

[0089] Analyze the audio spectrum, obtain the corresponding video frames, perform inter-frame comparison and frame alignment, and calculate the time difference between audio and video.

[0090] Taking a 10fps reference video as an example, the reference video has 10 frames per second, and customized sound output is at the end of the cycle. After reference video A undergoes full-link processing, audio and video data B is obtained on the playback end. Audio and video analysis is performed on reference video A and video B on the playback end to determine the frames containing the actual sound playback and the frames containing the sound playback of the reference video. These frames are then compared and calculated to determine the time difference between the sound and picture synchronization after full-link processing.

[0091] The following is a further introduction to the overall audio and video synchronization detection system. Figure 4 is a schematic diagram of an overall system for detecting audio and video synchronization according to an embodiment of the present disclosure. Figure 4 As shown, a video de-frame set 1 is performed on the reference video A and the video B played by the terminal respectively, and a silent frame set 2 is detected. The frame where the sound is located is obtained by elimination method, and the time difference between the frame where the sound is located between the reference video A and the video B played by the terminal is calculated by the frame label in the obtained frame, thereby obtaining the time difference of the audio and video synchronization deviation of the video B played by the terminal.

[0092] The above technical solution will be further introduced as a whole below.

[0093] Step S401: decode the video.

[0094] Get the decoder address (IcodecID) according to the editor context, use different decoder addresses to decode the original video A and the terminal video B respectively, and obtain the context, frame rate, time base and other information of the original video A and the context, frame rate, time base and other information of the terminal video B.

[0095] Step S402: Perform feature analysis on the video and audio.

[0096] The audio data and video data of the decoding information such as context, frame rate, and time base obtained after decoding the original video A are analyzed, and the audio data and video data of the decoding information such as context, frame rate, and time base obtained after decoding the terminal video B are analyzed.

[0097] Step S403: Obtain reference video and terminal video sound frame images.

[0098] The preset audio feature points are determined according to the wavelength and amplitude of the reference video sound, and the range of the silent frame timestamp of the reference video is obtained according to the preset audio feature points; the preset audio feature points are determined according to the wavelength and amplitude of the terminal video sound, and the range of the silent frame timestamp of the terminal video is obtained according to the preset audio feature points; the audio feature point recognition method can be to analyze the sound amplitude, the preset video sound amplitude is controllable, and the noise interference is eliminated by the threshold, and then the video frame interval is judged. If the video frame timestamp set is not in the silent frame timestamp set, the frame picture set where the sound is located is inferred, and this set is the picture with sound, thereby obtaining the sound frame picture set of the reference video and the terminal video.

[0099] Step S404: Calculate the audio and video synchronization difference.

[0100] After obtaining a set of video frames with sounds with preset feature points, the audio and video synchronization difference between the reference video A and the terminal video B is calculated. The distance difference between the two video frames is converted into time, that is, the total duration is calculated according to the frame interval, so as to obtain the audio and video synchronization time.

[0101] During the same period, if the audio frame label of terminal video B is greater than the audio frame label of reference video A, it means that the sound after full-link processing is ahead of the picture by a certain length. Conversely, if the audio frame label of video B is less than the audio frame label of video A, the sound is lagging behind.

[0102] Taking a reference video with a frame rate of 10fps as an example, assuming that the original frame sequence of the video frames with sound in the reference media file is [10, ... 2,2,2,2,2,2,2,2,2,1,2,2,1,2,2,3…], then the audio and video time difference ≈[(2+10)-10]*(1000ms / 10), where the frame sequence represents the video frames where sound appears. Referring to the frame sequence where the video sound is located, the first 10 in the original sequence indicates that the video frame corresponding to the first sound frame is the 10th frame of the first cycle, and the first 2 in the frame sequence where the sound is located on the playback end indicates that the video frame corresponding to the first sound frame is the 2nd frame of the second cycle. It can be inferred that the sound frame of video B lags behind the sound frame of video A.

[0103] One thing to note here is that the frequency of live broadcast or video scenes captured at the acquisition end may be 15fps, 24fps, or other frequencies. The frame rate will change after transcoding, network compatibility with packet loss, and other processing during transmission. Therefore, there will be a certain error in the inter-frame duration when comparing video A and terminal video B. The errors in each scene need to be handled separately. For example, frame loss during transmission will cause the frame interval to lengthen, or frame chasing will cause the frame interval to shorten. It is necessary to analyze the actual playback frame interval and the frame loss and frame chasing situations.

[0104] In addition, high-frequency videos are used as much as possible when selecting reference videos to reduce the error range to below 25ms.

[0105] Step S405: result analysis.

[0106] After the time difference is obtained, the audio and video time difference can be analyzed to determine whether it is within an acceptable range based on the definition of the international standard (RFC-1359). The industry has three standards: the International Telecommunication Union standard (ITU-R BT.1359 (1998)), the digital television standard (ATSC IS / 191 (2003)), and the digital equipment interface protocol for transmitting and receiving digital audio signals (EBU R37 (2007)). Among them, the most influential one is the audio and video synchronization standard "ITU-R BT.1359-1" revised by the International Telecommunication Union in 1998 for television muting, which is currently also widely used in the quality assessment of Internet live broadcasts. Figure 5 This is a diagram of the duration of audio and video asynchrony and user sensory standards defined according to international standards, such as Figure 5 As shown, the audio lags behind the video as a negative value, and vice versa as a positive value. When the audio and video time deviation is in the range of -100ms to 25ms, it cannot be perceived. 125ms to 45ms can be recognized. Less than 185ms and greater than 90ms are unacceptable. This allows analysis of whether the video audio and video asynchrony is within the acceptable range.

[0107] The present disclosure performs video deframing on reference video A and video B played on the terminal respectively, obtains the range of the silent frame timestamp according to the preset audio feature points, obtains the silent frame picture set 2 of the reference video A and video B played on the terminal, analyzes the sound amplitude, obtains the frame picture where the sound is located by elimination method, thereby obtaining the frame picture set 1 where the sound is located of the reference video A and video B played on the terminal, calculates the time difference between the frame picture where the sound is located of the reference video A and video B played on the terminal through the frame label in the obtained picture, thereby obtaining the audio and video synchronization deviation time difference of the terminal playing video B, thereby solving the technical problem of low efficiency of audio and video synchronization detection and achieving the technical effect of improving the efficiency of audio and video synchronization detection.

[0108] The present disclosure also provides a method for executing Figure 1 The media file processing method and the media file processing device of the illustrated embodiment are provided.

[0109] Figure 6 FIG. 1 is a schematic diagram of a device for processing media files according to an embodiment of the present disclosure. Figure 6 As shown, the media file processing device 60 may include: an acquisition unit 61 , a first determination unit 62 and a second determination unit 63 .

[0110] The acquiring unit 61 is configured to acquire a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file.

[0111] The first determination unit 62 is configured to determine a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame having the same timestamp as the first audio frame in the first media file, the second frame set includes a second video frame having the same timestamp as the second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame.

[0112] The second determining unit 63 is configured to determine a time difference between a timestamp of the first video frame and a timestamp of the second video frame based on the first frame set and the second frame set.

[0113] Optionally, the second determining unit 63 includes: a first determining module, configured to determine a first identifier of a first video frame in a first frame set; determine a second identifier of a second video frame in a second frame set; and determine a time difference based on the first identifier and the second identifier.

[0114] Optionally, the first determining module includes: a first determining submodule, configured to determine a value of a first identifier of the first video frame in the first frame set based on a first frame rate of the first media file.

[0115] Optionally, the first determination module includes: a second determination submodule, which is used to determine the second identifier of the first video frame in the second frame set based on the second frame rate of the second media file, including: determining the value of the second identifier of the first video frame in the second frame set based on the second frame rate of the second media file.

[0116] Optionally, the first determination module includes: a third determination submodule, used to determine that the timestamp of the first video frame is ahead of the timestamp of the second video frame in response to the value of the first identifier being greater than the value of the second identifier; and to determine that the timestamp of the first video frame lags behind the timestamp of the second video frame in response to the value of the first identifier being less than the value of the second identifier.

[0117] Optionally, the second determining unit 63 includes: a second determining module, configured to decode the first media file to obtain first audio data and first video data; and determine the first frame set based on the first audio data and the first video data.

[0118] Optionally, the second determination module includes: a fourth determination submodule, used to determine a first target video frame in the first video data based on the target audio parameters of the first media file, wherein the first target video frame does not match an audio frame in the first audio data; and determine a first frame set based on the timestamp of the video frame in the first video data and the timestamp of the first target video frame.

[0119] Optionally, the second determining module includes: a fifth determining submodule, configured to determine a video frame in the first video data whose timestamp exceeds the timestamp of the first target video frame as the first video frame in the first frame set.

[0120] Optionally, the first determining unit includes: a third determining module, configured to decode the second media file to obtain second audio data and second video data; and determine the second frame set based on the second audio data and the second video data.

[0121] Optionally, the third determination module includes: a sixth determination submodule, used to determine a second target video frame in the second video data based on the target audio parameters of the second media file, wherein the second target video frame does not match the audio frame in the second audio data; and determine the second frame set based on the timestamp of the video frame in the second video data and the timestamp of the second target video frame.

[0122] Optionally, the third determination module includes: a seventh determination submodule, configured to determine a video frame in the second video data whose timestamp exceeds the timestamp of the second target video frame as the second video frame in the second frame set.

[0123] In the media file processing device of the application of this embodiment, an acquisition unit is used to acquire a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file; a first determination unit is used to determine a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame having the same timestamp as the first audio frame in the first media file, the second frame set includes a second video frame having the same timestamp as the second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame; and a second determination unit is used to determine a time difference between the timestamp of the first video frame and the timestamp of the second video frame based on the first frame set and the second frame set, thereby obtaining a time difference in audio and video synchronization deviation between the first media file and the second media file, thereby solving the technical problem of low audio and video synchronization detection efficiency and achieving the technical effect of improving audio and video synchronization detection efficiency.

[0124] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0125] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0126] An embodiment of the present disclosure provides an electronic device, which may include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the media file processing method of the embodiment of the present disclosure.

[0127] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0128] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the media file processing method of the embodiment of the present disclosure.

[0129] Optionally, in this embodiment, the non-volatile storage medium may be configured to store a computer program for executing the following steps:

[0130] S1, obtaining a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file;

[0131] S2, determining a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame having the same timestamp as the first audio frame in the first media file, the second frame set includes a second video frame having the same timestamp as the second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame;

[0132] S3 : Determine a time difference between a timestamp of the first video frame and a timestamp of the second video frame based on the first frame set and the second frame set.

[0133] Alternatively, in this embodiment, the non-transitory computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any suitable combination of the above. More specific examples of readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0134] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0135] S1, obtaining a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file;

[0136] S2, determining a first frame set of the first media file and a second frame set of the second media file, wherein the first frame set includes a first video frame having the same timestamp as the first audio frame in the first media file, the second frame set includes a second video frame having the same timestamp as the second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame;

[0137] S3 : Determine a time difference between a timestamp of the first video frame and a timestamp of the second video frame based on the first frame set and the second frame set.

[0138] Figure 7 1 is a block diagram of an electronic device for a method of processing media files according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0139] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0140] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0141] The computing unit 701 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the method for processing media files. For example, in some embodiments, the method for processing media files can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method for processing media files described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the data processing method by any other suitable means (e.g., by means of firmware).

[0142] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0146] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0147] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0148] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0149] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for processing a media file, comprising: Obtaining a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file; Analyzing and processing the first media file to obtain a first frame set, and analyzing and processing the second media file to obtain a second frame set, wherein the first frame set is a set of frames containing sound in the first media file, the first frame set includes a first video frame having the same timestamp as a first audio frame in the first media file, the second frame set is a set of frames containing sound in the second media file, the second frame set includes a second video frame having the same timestamp as a second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame; determining, based on the first frame set and the second frame set, a time difference between a timestamp of the first video frame and a timestamp of the second video frame; Wherein, determining the time difference between the timestamp of the first video frame and the timestamp of the second video frame based on the first frame set and the second frame set includes: determining the value of a first identifier of the first video frame in the first frame set based on the first frame rate of the first media file, wherein the first identifier is used to identify the first video frame corresponding to the first appearance of the first audio frame in the first media file; determining the value of a second identifier of the second video frame in the second frame set based on the second frame rate of the second media file, wherein the second identifier is used to identify the second video frame corresponding to the first appearance of the second audio frame in the second media file; and determining the time difference based on the value of the first identifier and the value of the second identifier.

2. The method according to claim 1, wherein Also includes: In response to the value of the first identifier being greater than the value of the second identifier, determining that the timestamp of the first video frame is ahead of the timestamp of the second video frame; In response to the value of the first identifier being smaller than the value of the second identifier, it is determined that the timestamp of the first video frame lags behind the timestamp of the second video frame.

3. The method according to claim 1, wherein Analyzing and processing the first media file to obtain the first frame set includes: Decoding the first media file to obtain first audio data and first video data; The first set of frames is determined based on the first audio data and the first video data.

4. The method according to claim 3, wherein: Determining the first frame set based on the first audio data and the first video data includes: determining a first target video frame in the first video data based on a target audio parameter of the first media file, wherein the first target video frame does not match an audio frame in the first audio data; The first frame set is determined based on time stamps of video frames in the first video data and time stamps of the first target video frames.

5. The method according to claim 4, wherein Determining the first frame set based on timestamps of video frames in the first video data and timestamps of the first target video frame includes: A video frame in the first video data whose timestamp exceeds the timestamp of the first target video frame is determined as the first video frame in the first frame set.

6. The method according to claim 1, wherein Analyzing and processing the second media file to obtain the second frame set includes: Decoding the second media file to obtain second audio data and second video data; The second set of frames is determined based on the second audio data and the second video data.

7. The method according to claim 6, wherein: Determining the second frame set based on the second audio data and the second video data includes: determining a second target video frame in the second video data based on a target audio parameter of the second media file, wherein the second target video frame does not match an audio frame in the second audio data; The second frame set is determined based on the time stamps of the video frames in the second video data and the time stamps of the second target video frames.

8. The method according to claim 7, wherein: Determining the second frame set based on the timestamp of the video frame in the second video data and the timestamp of the second target video frame includes: A video frame in the second video data whose time stamp exceeds the time stamp of the second target video frame is determined as the second video frame in the second frame set.

9. A media file processing device, comprising: an acquiring unit, configured to acquire a first media file and a second media file, wherein the first media file is obtained by performing full-link processing on the second media file; a first determining unit configured to analyze and process the first media file to obtain a first frame set, and to analyze and process the second media file to obtain a second frame set, wherein the first frame set is a set of frames containing sound in the first media file, the first frame set includes a first video frame having the same timestamp as a first audio frame in the first media file, the second frame set is a set of frames containing sound in the second media file, the second frame set includes a second video frame having the same timestamp as a second audio frame in the second media file, and the audio content of the first audio frame is the same as the audio content of the second audio frame; a second determining unit, configured to determine a time difference between a timestamp of the first video frame and a timestamp of the second video frame based on the first frame set and the second frame set; The second determination unit is configured to determine, based on the first frame set and the second frame set, the time difference between the timestamp of the first video frame and the timestamp of the second video frame by performing the following steps, including: determining, based on a first frame rate of the first media file, a value of a first identifier of the first video frame in the first frame set, wherein the first identifier is used to identify the first video frame corresponding to the first appearance of the first audio frame in the first media file; determining, based on a second frame rate of the second media file, a value of a second identifier of the second video frame in the second frame set, wherein the second identifier is used to identify the second video frame corresponding to the first appearance of the second audio frame in the second media file; and determining the time difference based on the values ​​of the first identifier and the second identifier.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Live broadcast data processing method and device, computer equipment and storage medium

    CN111918093A

  • Play delay difference measurement method, device, equipment, system and storage medium

    CN112040225A

  • Audio and video timestamp processing method and device, electronic equipment and storage medium

    CN112423075A