An audio-video bitstream time calibration method and an electronic device

The audio and video stream time calibration method synchronizes separate audio and video streams by adjusting time stamps based on synchronized cues, addressing the misalignment issue and improving playback synchronization.

CN115426501BActive Publication Date: 2025-07-15ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210943588.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-07-15
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

During the audio and video encoding process, the independent audio and video encoding system is not synchronized by the clock, resulting in millisecond-level accurate synchronization during decoding, which affects the matching of the playback screen and the sound.

Method used

By inserting a timestamp calibration module in the audio and video encoding system, synchronous calibration is performed using the acousto-optical signals emitted by the calibration device, the timestamp calibration value is determined, and the timestamps of the audio and video frames are adjusted to achieve synchronization.

Benefits of technology

It realizes accurate synchronization of audio and video decoding, ensures the matching of the playback screen and the sound content, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115426501B_ABST
    Figure CN115426501B_ABST
Patent Text Reader

Abstract

The present application discloses an audio-video stream time calibration method and an electronic device, which are used to ensure the precise synchronization of audio decoding and video decoding, thereby ensuring the matching of the playback picture and the sound content and improving the user experience. The audio-video stream time calibration method provided by the present application includes: in the first working mode, obtaining the currently input audio stream frame generated by the first system clock and the currently input video stream frame generated by the second system clock; wherein, the first system clock and the second system clock are in different systems, and the first system clock and the second system clock are different; extracting the timestamp of the audio stream frame or the video stream frame; and calibrating the timestamp based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, so that the timestamps of the audio stream frame and the video stream frame are synchronized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio - video technology, and in particular, to an audio - video bitstream time calibration method and an electronic device. Background Art

[0002] In order to save network transmission bandwidth, when transmitting multimedia data, it is necessary to compress the original digital audio and video data. This method of compressing digital audio and video is called audio encoding and video encoding. Common audio encoding formats include G711A, G711U, AAC, etc. Common video encoding formats include H264, H265, MPEG - 4, etc. Different encoding formats have different compression effects.

[0003] In the field of multimedia applications, video and audio information generally appears in pairs. For example, when playing movies, videos, etc., the audio is the accompanying sound of the video and matches the video content. Audio - encoded data and video - encoded data are generally alternately mixed and packed for transmission, such as the TS stream (Transport Stream). During decoding, the audio bitstream and the video bitstream are processed in independent processes. The audio data to be decoded given to the audio decoder and the video data to be decoded given to the video decoder at the same moment may not be the content at the same moment during encoding. Therefore, a mechanism is needed to ensure the precise synchronization of audio and video decoding, and thus ensure the matching of the playing picture and the sound content. However, if two systems are synchronized only by time, most of them cannot meet the audio - video synchronization requirements because time synchronization is generally at the second level; while audio - video synchronization requires synchronization to the millisecond level; and if synchronization is performed after encoding, the processing time error caused by encoding will be introduced. Summary of the Invention

[0004] Embodiments of this application provide an audio - video bitstream time calibration method and an electronic device to ensure the precise synchronization of audio decoding and video decoding, thus ensuring the matching of the playing picture and the sound content and improving the user experience.

[0005] An audio - video bitstream time calibration method provided by an embodiment of this application includes:

[0006] In the first working mode, obtain the currently input audio bitstream frame generated by using the first system clock and the currently input video bitstream frame generated by using the second system clock; wherein, the first system clock and the second system clock are in different systems, and the first system clock and the second system clock are different;

[0007] Extract the timestamps of the audio stream frames and / or video stream frames; and calibrate the timestamps based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, so that the timestamps of the audio stream frames and the video stream frames are synchronized.

[0008] In the embodiments of the present application, in the first working mode, the audio stream frames generated by using the first system clock in the currently input, and the video stream frames generated by using the second system clock in the currently input are obtained, and the timestamps of the audio stream frames and / or video stream frames are extracted; based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, the timestamps are calibrated, so that the timestamps of the audio stream frames and the video stream frames are synchronized, thereby ensuring the precise synchronization of audio decoding and video decoding, and further ensuring the matching of the playing picture and the sound content, and improving the user experience.

[0009] In some embodiments, the timestamp calibration value is determined in the following manner:

[0010] In the second working mode, the audio acquisition system and the video acquisition system are simultaneously triggered to collect information on a preset calibration device, and an audio stream generated by using the first system clock and a collected video stream generated by using the second system clock are respectively obtained; wherein, a lamp with a preset shape is arranged on the calibration device, and the calibration device periodically emits a sound with a preset duration, and lights the lamp with the preset shape while emitting the sound;

[0011] According to the preset standard audio, perform sound recognition on the audio stream collected from the calibration device, and determine the audio frame sequence corresponding to the standard audio through the sound recognition; and according to the preset standard image, perform image recognition on the video stream collected from the calibration device, and determine the video frame sequence corresponding to the standard image through the image recognition;

[0012] Determine the timestamp calibration value according to the audio frame sequence and the video frame sequence.

[0013] In some embodiments, determining the timestamp calibration value according to the audio frame sequence and the video frame sequence specifically includes:

[0014] Respectively determine the timestamp of the last frame of the audio frame sequence and the timestamp of the last frame of the video frame sequence;

[0015] Take the difference between the timestamp of the last frame of the video frame sequence and the timestamp of the last frame of the audio frame sequence as the timestamp calibration value.

[0016] In some embodiments, the video frame sequence is a video frame sequence of N consecutive frames including the standard image, where t*f - c < N < t*f + c, c is a preset constant, f is the frame rate of the video encoding of the second system, and t is the preset duration.

[0017] In some embodiments, according to a preset standard audio, voice recognition is performed on the audio bitstream collected by the first system from the calibration device, and the audio frame sequence corresponding to the standard audio is determined through the voice recognition. Specifically, it includes:

[0018] The audio bitstream collected from the calibration device is decoded and then placed in a preset first-in-first-out (FIFO) buffer queue, and the size of the FIFO buffer queue is equal to the data size of the standard audio.

[0019] The data in the FIFO buffer queue is matched with the standard audio. When the matching is successful, the audio frame sequence in the FIFO buffer queue is determined as the audio frame sequence corresponding to the standard audio.

[0020] In some embodiments, a user instruction is received through a user interface output to the user to implement the switching between the first working mode and the second working mode.

[0021] In some embodiments, the timestamps of the audio bitstream frames and / or video bitstream frames are extracted; and, based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, the timestamps are calibrated. Specifically, it includes:

[0022] Extract the timestamp of the audio bitstream frame;

[0023] Use the sum of the timestamp and the timestamp calibration value as the new timestamp and update it to the audio bitstream frame.

[0024] Another embodiment of the present application provides an electronic device, which includes a memory and a processor. Among them, the memory is used to store program instructions, and the processor is used to call the program instructions stored in the memory and execute any of the above methods according to the obtained program.

[0025] In addition, according to an embodiment, for example, a computer program product for a computer is provided, which includes a software code part. When the product runs on a computer, these software code parts are used to execute the steps of the method defined above. The computer program product may include a computer-readable medium on which the software code part is stored. In addition, the computer program product can be directly loaded into the internal memory of the computer and / or sent via a network through at least one of an upload process, a download process, and a push process.

[0026] Another embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing the computer to execute any of the above methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is a schematic diagram of the audio and video coding transmission processing flow provided by the embodiment of the present application;

[0029] Figure 2 It is a schematic diagram of the working process of the recording and playing host in the normal mode provided by the embodiment of the present application;

[0030] Figure 3 It is a schematic diagram of the working process of the recording and playing host in the calibration mode provided by the embodiment of the present application;

[0031] Figure 4 It is a schematic diagram of the video matching process provided by the embodiment of the present application;

[0032] Figure 5 It is a schematic diagram of the audio matching process provided by the embodiment of the present application;

[0033] Figure 6 It is a schematic diagram of the process of an audio and video bitstream time calibration method provided by the embodiment of the present application;

[0034] Figure 7 It is a schematic diagram of the structure of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all of them. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0036] The embodiments of the present application provide an audio and video bitstream time calibration method and an electronic device to ensure the precise synchronization of audio decoding and video decoding, thereby ensuring the matching of the playback picture and the sound content and improving the user experience.

[0037] Among them, the method and the device are based on the same inventive concept. Since the principles of the method and the device for solving problems are similar, the implementation of the device and the method can be referred to each other, and the repeated parts will not be elaborated here.

[0038] In the description of the embodiments of this application, the terms "first", "second", etc. (if any) in the specification, claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0039] The following examples and embodiments are only to be understood as illustrative examples. Although this specification may mention "one", "a" or "some" examples or embodiments in several places, this does not mean that each such mention is related to the same example or embodiment, nor does it mean that the feature only applies to a single example or embodiment. The individual features of different embodiments can also be combined to provide other embodiments. In addition, terms such as "comprising" and "including" should be understood not to limit the described embodiments to only the features already mentioned; such examples and embodiments may also include features, structures, units, modules, etc. not specifically mentioned.

[0040] The following will describe each embodiment of this application in detail with reference to the drawings in the specification. It should be noted that the display order of the embodiments of this application only represents the sequence of the embodiments, and does not represent the superiority or inferiority of the technical solutions provided by the embodiments.

[0041] To play the effect of matching the picture and the sound content, the general processing method is to perform timestamp synchronization processing during audio and video encoding. That is, the video frames are encoded by video encoding to obtain a video bitstream, and the audio frames are encoded by audio encoding to obtain an audio bitstream. System clock information is added to the video bitstream and the audio bitstream. Specifically:

[0042] Video frames: Derived from the video acquisition module of the digital video recorder, the pictures of the external camera are collected into the memory, and the format is, for example, YUV420;

[0043] Audio frames: Derived from the audio data collected by the audio codec, and the format is, for example, PCM;

[0044] Video encoding: Encoding video frames according to a preset encoding format (such as H265), and outputting the encoded video bitstream.

[0045] Audio encoding: Encoding audio frames according to a preset encoding format (such as AAC), and outputting the encoded audio bitstream.

[0046] System clock: A high-precision physical clock in the system, such as RTC (Real-Time Clock), which represents the passage of time in the current encoding system.

[0047] In order to enable the audio and video bitstreams to have unified time reference information when played by a player, when the audio encoding and video encoding modules generate each frame of bitstream, they will obtain the current system clock, generate a timestamp according to a certain correspondence relationship for the system clock, and the timestamp is encapsulated in the auxiliary information of the bitstream. Each frame of video and audio bitstream carries a timestamp, and the timestamp represents the system time of video and audio encoding. When the video bitstream and audio bitstream are played by a player (i.e., video and audio decoding), the video decoder and audio decoder refer to the timestamp to achieve synchronization of video content and audio content.

[0048] The existing method for obtaining audio and video encoding timestamps based on the encoding system clock is only applicable to the scenario where the audio encoder and video encoder are in the same system. When the audio encoder and video encoder are two independent systems and the system times of these two systems are not synchronized, that is, the clocks of these two independent systems are not precisely aligned, it will cause a certain error in the timestamps of the audio and the timestamps of the video bitstream, so the decoded video content and audio content will be out of sync.

[0049] For example, see Figure 1 , the IPC is a network camera, and the recording and playback host is a device installed in the classroom for recording courses. In a business scenario, the video scenario monitored by the IPC and the audio scenario collected by the microphone (MIC) of the recording and playback host are the same classroom. The recording and playback host needs to package the video bitstream of the IPC and the audio bitstream collected by the local MIC of the recording and playback host, and synthesize a single audio-visual stream for storage or network transmission after that.

[0050] Among them, the timestamp of the audio bitstream is marked by the recording and playback host. That is to say, the audio acquisition system uses the first system clock to generate audio bitstream frames, and each frame of the audio bitstream carries the timestamp generated by the first system clock.

[0051] The video bitstream is encoded by the IPC, and the timestamp is also marked by the IPC. That is to say, the video acquisition system uses the second system clock to generate video bitstream frames, and each frame of the video bitstream carries the timestamp generated by the second system clock.

[0052] The system times of the audio acquisition system and the video acquisition system are not accurately aligned. And since the IPC bitstream is transmitted over the network with a relatively large delay, it will cause the actual video content and audio content in the audio-visual bitstream frames with the same or similar timestamps to not match, and there is a problem that the video image will be slightly slower than the audio content by one beat. Since the audio-visual synchronization on the player side heavily relies on timestamps, the existing solutions cannot solve the above problems.

[0053] The above scenario is only an example. In fact, as long as the audio and video encodings are not in the same system and there is a need for audio-visual synchronization during playback, this problem exists.

[0054] In the technical solution provided by the embodiments of the present application, to solve the problem of the mismatch of the audio-visual encoding timestamps in such a scenario, the audio-visual encoding timestamps are corrected so that the actual video frames and audio frames match under the same or similar encoding timestamps. The specific example is as follows:

[0055] As Figure 1 shown, in the embodiments of the present application, a calibration device is placed at a position that can be monitored by both the IPC and the recording and playback host MIC, and timestamp calibration processing is performed before the bitstream is packaged.

[0056] Regarding the calibration device:

[0057] For example, there is a light box with a specific shape on the surface of the calibration device. After the light box is lit, a pattern of a fixed color can be observed on the surface of the calibration device. For example, a triangular light box will emit green light when lit, and the fixed-color pattern will disappear when the light box is extinguished. That is to say, when the calibration device lights up the light box, if an image is collected through an image acquisition device, there will be an image of a green triangle in the collected image.

[0058] The calibration device also has a buzzer that can emit a beep with a fixed frequency. That is to say, the calibration device emits a beep with a fixed frequency while the light box is lit. If the sound is collected through an audio acquisition device at this time, there will be a beep with this fixed frequency in the collected audio.

[0059] In summary, after the calibration device is started, it can light up the light box and turn on the buzzer at a fixed time period, and after a fixed duration, turn off the light box and turn off the buzzer. And so on in a cycle.

[0060] As Figure 1 shown, in the embodiments of the present application, in the recording and playback host, a timestamp calibration module is inserted before the video bitstream and the audio bitstream are packaged and transmitted or stored.

[0061] The timestamp calibration module has two working modes: normal mode (which can also be called the first working mode) and calibration mode (which can also be called the second working mode).

[0062] Normal mode: As Figure 2 shown, for each input audio frame, a timestamp Ta is extracted, and then the newly calculated timestamp Tb is written back to the audio frame, where Tb = Ta + DT, and DT represents the timestamp calibration value determined based on a preset standard image (i.e., a specific image contained in the image collected when the light box of the above calibration device is lit (such as a green triangle image)) and a standard audio (i.e., the beeping sound emitted by the above calibration device simultaneously when the light box is lit) in the calibration mode. DT is initially 0, and its value is refreshed after the calibration mode runs to completion. In the normal mode, for the input video frames, no processing is performed and they are directly passed through to the subsequent module, that is, the timestamps of the video frames do not need to be calibrated. The timestamp calibration module defaults to the normal mode.

[0063] Of course, in this embodiment, the timestamp calibration of audio frames is taken as an example for illustration, and the timestamp calibration of video frames is not performed. In practical applications, similarly, the timestamp calibration of video frames can be performed, while the timestamp calibration of audio frames is not performed. Or, the timestamp calibration of both video frames and audio frames can be performed (for example, one is adjusted faster and the other is adjusted slower). How to specifically adjust the timestamps of video and audio frames using the timestamp calibration value is not limited in the embodiments of this application.

[0064] Calibration mode: As Figure 3 shown, the calibration mode is triggered by the user from the operation interface of the recording and broadcasting host. It is used to calculate DT. After obtaining DT, the calibration mode automatically exits and switches to the normal mode.

[0065] That is to say, in the embodiments of this application, the recording and broadcasting host can receive user instructions through the user interface output to the user to achieve the switching between the calibration mode and the normal mode. The recording and broadcasting host defaults to the normal mode. When it switches to the calibration mode according to the user instructions and obtains the latest timestamp calibration value, it can automatically switch back to the normal mode, or it can also switch back to the normal mode again according to the user instructions.

[0066] Before enabling the calibration mode, the user correctly places the calibration device (i.e., an acoustic-optical signal that can be collected by the image and sound acquisition devices) and turns on the calibration device. The specific position of the calibration device is not limited in the embodiments of the present application and can be determined according to actual needs. As long as in the calibration mode, the user simultaneously triggers the audio acquisition system (such as a recording and broadcasting host) and the video acquisition system (such as an IPC) to collect information about the calibration device, respectively obtaining an audio bitstream generated using the first system clock and a captured video bitstream generated using the second system clock. Among them, the audio acquisition system corresponds to the first system clock, and the video acquisition system corresponds to the second system clock.

[0067] The first step of the calibration mode is to separately identify and match the characteristic acoustic-optical signals of the calibration device for the input audio and video bitstreams. The calibration device periodically emits synchronous and fixed-length acoustic-optical characteristic signals. In the corresponding video bitstream, there will be a corresponding video frame sequence, and in the corresponding audio bitstream, there will be a corresponding audio frame sequence. After successful matching, the corresponding video frame sequence and audio frame sequence will be obtained, and finally T0 and T1 will be obtained. That is to say, in the embodiments of the present application, according to a preset standard audio, voice recognition is performed on the audio bitstream collected through the calibration device, and the audio frame sequence corresponding to the standard audio is determined through the voice recognition; and, according to a preset standard image, image recognition is performed on the video bitstream collected through the calibration device, and the video frame sequence corresponding to the standard image is determined through the image recognition; according to the audio frame sequence and the video frame sequence, the timestamp calibration value is determined.

[0068] In the calibration mode, the matching processes of audio and video are respectively as Figure 4 , Figure 5 shown, where Figure 4 , Figure 5 the video matching and audio matching processes are a parallel process and can be carried out simultaneously. A specific example is described as follows:

[0069] As Figure 4 shown, for the video matching process, the description is as follows:

[0070] Step 1: Identify the characteristic pattern of the calibration box (the pattern after the light box is lit) in each frame of the image.

[0071] Among them, the characteristic pattern, that is, the preset standard image. Therefore, it is easy to detect whether the preset standard image exists in the image. For example, it is determined whether a green triangle exists in each frame of the image.

[0072] Step 2: Determine whether a continuous video frame sequence containing the preset standard image for a preset duration is recognized, that is, whether there is a continuous N-frame video frame sequence containing the preset standard image, and N is approximately equal to t*f.

[0073] Among them, the preset duration is at least equivalent to the duration that the calibration device continuously lights the box each time.

[0074] The specific judgment method is as follows:

[0075] The feature pattern is recognized in N consecutive frames, and N is approximately equal to t*f. Where f is the frame rate of video coding, and t is the actual duration of a single light box lighting. The condition for the judgment that N is approximately equal to t*f to be true is t*f - c < N < t*f + c, where c is a preset empirical value, for example, equal to 2. That is to say, the N approximately equal to t*f satisfies the following condition: t*f - c < N < t*f + c.

[0076] Step 3: Obtain the corresponding video bitstream frame sequence, and the video content of this frame sequence is the continuous picture set corresponding to the actual single light box lighting.

[0077] Step 4: T0 takes the coding timestamp of the last frame of the video frame sequence. Of course, according to needs, T0 can also take the timestamps of other frames of the video frame sequence. The last frame is just easier to implement.

[0078] As Figure 5 shown, the audio matching process is described as follows:

[0079] Step 1: Identify whether the audio data matches the beep emitted by the calibration device. In the recording and broadcasting host, the audio data of a single beep of the calibration device is pre-recorded in advance, which can be called the preset standard audio. The input audio bitstream is decoded and placed in a first-in-first-out (FIFO) buffer queue with a total size equal to the size of the preset standard audio data (that is, the audio emitted by the calibration device), and the audio frame sequence in the FIFO buffer queue is matched with the preset standard audio data. If the match is successful, it indicates that the audio input sequence and the preset standard audio data are the same content.

[0080] Step 2: In the same way as the T0 value-taking method, T1 takes the coding timestamp of the last frame of the audio frame sequence (that is, the audio sequence in the FIFO buffer queue where the match is successful). Of course, as long as it is the same as the T0 value-taking method, that is to say, T1 can also take the timestamps of other frames.

[0081] In the embodiment of the present application, according to the working mechanism of the actual calibration device, the sound and light are synchronized, that is, the beep and the feature pattern appear and end at the same time. In theory, T1 and T0 should be very close. However, since the audio coding and video coding are not in the same time system, and there are delay differences in the processing process, therefore DT = T0 - T1, and DT is the actual audio-visual frame error time before the bitstream packaging module. Using DT to correct the audio coding timestamp can eliminate this actual time error.

[0082] The technical solution provided by the embodiments of the present application is applicable to scenarios where audio and video coding are not in the same system and audio-video synchronization is required during playback. The prior art cannot solve the audio-video synchronization problem in this scenario, and the technical solution provided by the embodiments of the present application can perfectly solve this problem.

[0083] In summary, referring to Figure 6 , a method for calibrating the time of an audio-video code stream provided by the embodiments of the present application can be applied to the above-mentioned recording and playback host side or the IPC side of the video acquisition system, and is not specifically limited. The method includes:

[0084] S101. In the first working mode, obtain the currently input audio code stream frame generated using the first system clock and the currently input video code stream frame generated using the second system clock; wherein, the first system clock and the second system clock are in different systems, and the first system clock and the second system clock are different;

[0085] The first working mode is, for example, the above-mentioned normal mode.

[0086] The first system clock, that is, the system clock of the audio acquisition system, is, for example, the system clock of the above-mentioned recording and playback host.

[0087] The second system clock, that is, the system clock of the video acquisition system, is, for example, the system clock of the above-mentioned IPC.

[0088] S102. Extract the timestamps of the audio code stream frame and / or the video code stream frame; and, based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, calibrate the timestamps so that the timestamps of the audio code stream frame and the video code stream frame are synchronized.

[0089] The second working mode is, for example, the above-mentioned calibration mode.

[0090] The timestamp calibration value is, for example, the above-mentioned DT.

[0091] In some embodiments, the timestamp calibration value is determined in the following manner:

[0092] In the second working mode, simultaneously trigger the audio acquisition system and the video acquisition system to collect information from a preset calibration device, and respectively obtain the audio code stream generated using the first system clock and the collected video code stream generated using the second system clock; wherein, a lamp with a preset shape is provided on the calibration device, and the calibration device periodically emits a sound for a preset duration and lights the lamp with the preset shape while emitting the sound;

[0093] Perform voice recognition on the audio bitstream collected by the calibration device according to a preset standard audio, and determine the audio frame sequence corresponding to the standard audio through the voice recognition; and perform image recognition on the video bitstream collected by the calibration device according to a preset standard image, and determine the video frame sequence corresponding to the standard image through the image recognition;

[0094] Determine the timestamp calibration value according to the audio frame sequence and the video frame sequence.

[0095] In some embodiments, determining the timestamp calibration value according to the audio frame sequence and the video frame sequence specifically includes:

[0096] Respectively determine the timestamp (such as T1 above) of the last frame of the audio frame sequence and the timestamp (such as T0 above) of the last frame of the video frame sequence;

[0097] Use the difference between the timestamp of the last frame of the video frame sequence and the timestamp of the last frame of the audio frame sequence as the timestamp calibration value.

[0098] In some embodiments, the video frame sequence is a video frame sequence of consecutive N frames containing the standard image, where t*f - c < N < t*f + c, c is a preset constant, f is the frame rate of the video encoding of the second system, and t is the preset duration.

[0099] In some embodiments, performing voice recognition on the audio bitstream collected by the calibration device through the first system according to a preset standard audio, and determining the audio frame sequence corresponding to the standard audio through the voice recognition specifically includes:

[0100] Put the audio bitstream collected by the calibration device into a preset first-in-first-out (FIFO) buffer queue after decoding, and the size of the FIFO buffer queue is equal to the data size of the standard audio;

[0101] Match the data in the FIFO buffer queue with the standard audio, and when the match is successful, determine the audio frame sequence in the FIFO buffer queue as the audio frame sequence corresponding to the standard audio.

[0102] In some embodiments, receive a user instruction through a user interface output to the user to implement the switching between the first working mode and the second working mode.

[0103] In some embodiments, the timestamps of the audio stream frames and / or video stream frames are extracted; and, based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, the timestamps are calibrated, which specifically includes:

[0104] Extract the timestamp of the audio stream frame;

[0105] Use the sum of the timestamp and the timestamp calibration value as the new timestamp and update it into the audio stream frame.

[0106] Next, the device or apparatus provided in the embodiments of the present application will be introduced. For the explanations or examples of the same or corresponding technical features as those in the above method, they will not be repeated hereinafter.

[0107] See Figure 7 , the embodiments of the present application provide an electronic device, which can also be referred to as a computing device. The device can specifically be a desktop computer, a portable computer, a smart phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, PDA), etc. The computing device can include a central processing unit (Center Processing Unit, CPU), a memory, and further can also include input / output devices, etc. The input devices can include a keyboard, a mouse, a touch screen, etc., and the output devices can include display devices, such as a liquid crystal display (Liquid Crystal Display, LCD), a cathode ray tube (Cathode Ray Tube, CRT), etc.

[0108] The memory can include a read-only memory (ROM) and a random access memory (RAM), and provide program instructions and data stored in the memory to the processor. In the embodiments of the present application, the memory can be used to store the program of any of the methods provided in the embodiments of the present application.

[0109] The processor is used to execute any of the methods provided in the embodiments of the present application by calling the program instructions stored in the memory.

[0110] The electronic device, for example, can be the above-mentioned live recording host, which has functions such as audio acquisition and encoding, timestamp calibration, and stream packaging, transmission, storage, etc.

[0111] The electronic device can also specifically be the timestamp calibration module included in the above-mentioned live recording host.

[0112] Alternatively, the electronic device can also be a video acquisition system, which has a video acquisition and encoding function, such as an IPC. The IPC can also perform timestamp calibration and has the above-mentioned timestamp calibration module.

[0113] Specifically, for example, referring to Figure 7 , the electronic device provided by the embodiment of the present application includes:

[0114] A memory 620 for storing program instructions;

[0115] A processor 600 for calling the program instructions stored in the memory 620 and executing according to the obtained program:

[0116] In the first working mode, acquire the currently input audio stream frame generated by the first system clock and the currently input video stream frame generated by the second system clock; wherein, the first system clock and the second system clock are in different systems, and the first system clock and the second system clock are different;

[0117] Extract the timestamps of the audio stream frame and / or the video stream frame; and, based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, calibrate the timestamps so that the timestamps of the audio stream frame and the video stream frame are synchronized.

[0118] In some embodiments, the timestamp calibration value is determined in the following manner:

[0119] In the second working mode, simultaneously trigger the audio acquisition system and the video acquisition system to collect information about a preset calibration device, and respectively obtain the audio stream generated by the first system clock and the acquired video stream generated by the second system clock; wherein, a lamp with a preset shape is provided on the calibration device, and the calibration device periodically emits a sound for a preset duration, and lights the lamp with the preset shape while emitting the sound;

[0120] According to the preset standard audio, perform voice recognition on the audio stream collected from the calibration device, and determine the audio frame sequence corresponding to the standard audio through the voice recognition; and, according to the preset standard image, perform image recognition on the video stream collected from the calibration device, and determine the video frame sequence corresponding to the standard image through the image recognition;

[0121] Determine the timestamp calibration value according to the audio frame sequence and the video frame sequence.

[0122] In some embodiments, determining the timestamp calibration value according to the audio frame sequence and the video frame sequence specifically includes:

[0123] Respectively determine the timestamp of the last frame of the audio frame sequence and the timestamp of the last frame of the video frame sequence;

[0124] Use the difference between the timestamp of the last frame of the video frame sequence and the timestamp of the last frame of the audio frame sequence as the timestamp calibration value.

[0125] In some embodiments, the video frame sequence is a video frame sequence of N consecutive frames containing the standard image, where t*f - c < N < t*f + c, c is a preset constant, f is the frame rate of the second system video encoding, and t is the preset duration.

[0126] In some embodiments, perform voice recognition on the audio bitstream collected by the calibration device through the first system according to a preset standard audio, and determine the audio frame sequence corresponding to the standard audio through the voice recognition, specifically including:

[0127] Put the audio bitstream collected from the calibration device into a preset first-in-first-out (FIFO) buffer queue after decoding, and the size of the FIFO buffer queue is equal to the data size of the standard audio;

[0128] Match the data in the FIFO buffer queue with the standard audio, and when the match is successful, determine the audio frame sequence in the FIFO buffer queue as the audio frame sequence corresponding to the standard audio.

[0129] In some embodiments, receive a user instruction through a user interface output to the user to implement the switching between the first working mode and the second working mode.

[0130] In some embodiments, extract the timestamps of the audio bitstream frames and / or video bitstream frames; and calibrate the timestamps based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode, specifically including:

[0131] Extract the timestamp of the audio bitstream frame;

[0132] Use the sum of the timestamp and the timestamp calibration value as the new timestamp, and update it to the audio bitstream frame.

[0133] The transceiver 610 is used to receive and send data under the control of the processor 600.

[0134] Among them, in Figure 7Among them, the bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits of one or more processors represented by the processor 600 and the memory represented by the memory 620. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and thus will not be further described herein. The bus interface provides an interface. The transceiver 610 may be a plurality of components, that is, including a transmitter and a receiver, providing a unit for communicating with various other devices on the transmission medium. For different user devices, the user interface 630 may also be an interface capable of externally or internally connecting required devices, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, etc.

[0135] The processor 600 is responsible for managing the bus architecture and general processing, and the memory 620 may store data used by the processor 600 when performing operations.

[0136] Optionally, the processor 600 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a CPLD (Complex Programmable Logic Device).

[0137] It should be noted that the division of units in the embodiments of the present application is illustrative, merely a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, the functional units may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated units may be implemented in the form of hardware or in the form of software functional units.

[0138] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0139] Embodiments of this application also provide a computer program product or a computer program. This computer program product or computer program includes computer instructions, and these computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads these computer instructions from the computer-readable storage medium, and the processor executes these computer instructions to cause the computer device to execute any of the methods described in the above embodiments. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0140] Embodiments of this application provide a computer-readable storage medium for storing the computer program instructions used for the device provided in the above embodiments of this application. It contains a program for executing any of the methods provided in the above embodiments of this application. The computer-readable storage medium can be a non-transitory computer-readable medium.

[0141] The computer-readable storage medium can be any available medium or data storage device accessible by a computer, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROM, EPROM, EEPROM, non-volatile memories (NANDFLASH), solid-state drives (SSD)), etc.

[0142] It should be understood that:

[0143] The access technology through which entities in a communication network transmit traffic to and from each other can be any suitable current or future technology, such as WLAN (Wireless Local Area Network), WiMAX (Worldwide Interoperability for Microwave Access), LTE, LTE-A, 5G, Bluetooth, infrared, etc.; additionally, the embodiments can also apply wired technologies, for example, IP-based access technologies, such as wired networks or fixed lines.

[0144] Embodiments suitable for being implemented as software code or a part thereof and running using a processor or processing function are independent of the software code and can be specified using any known or future-developed programming language, such as high-level programming languages, such as objective-C, C, C++, C#, Java, Python, Javascript, other scripting languages, etc., or low-level programming languages, such as machine language or assembler.

[0145] The implementation of the embodiments is independent of hardware and can be implemented using any known or future-developed hardware technology or any combination thereof, such as microprocessors or CPUs (Central Processing Units), MOS (Metal Oxide Semiconductor), CMOS (Complementary MOS), BiMOS (Bipolar MOS), BiCMOS (Bipolar CMOS), ECL (Emitter Coupled Logic), and / or TTL (Transistor-Transistor Logic).

[0146] The embodiments can be implemented as a separate device, apparatus, unit, component, or function, or in a distributed manner. For example, one or more processors or processing functions can be used or shared in the processing, or one or more processing segments or processing parts can be used and shared in the processing, where one physical processor or more than one physical processor can be used to implement one or more processing parts dedicated to a specific processing as described.

[0147] The apparatus can be implemented by a semiconductor chip, a chipset, or a (hardware) module including such a chip or chipset.

[0148] The embodiments can also be implemented as any combination of hardware and software, such as ASIC (Application Specific IC (Integrated Circuit)) components, FPGA (Field Programmable Gate Array) or CPLD (Complex Programmable Logic Device) components or DSP (Digital Signal Processor) components.

[0149] The embodiments can also be implemented as a computer program product, including a computer-usable medium having computer-readable program code embodied therein, the computer-readable program code being adapted to perform the processes as described in the embodiments, wherein the computer-usable medium can be a non-transitory medium.

[0150] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0151] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0152] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, the instruction means realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are performed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocksFigure 1 Steps of functions specified in one or more boxes.

[0154] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.

Claims

1. An audio-video stream time calibration method, characterized in that, The method includes: In a first working mode, obtaining an audio stream frame generated by using a first system clock and currently input, and a video stream frame generated by using a second system clock and currently input; wherein, the first system clock and the second system clock are in different systems, and the first system clock and the second system clock are different; Extracting timestamps of the audio stream frame and / or the video stream frame; and calibrating the timestamps based on a timestamp calibration value determined by using a preset standard image and standard audio in a second working mode, so that the timestamps of the audio stream frame and the video stream frame are synchronized; Wherein, the timestamp calibration value is determined by the following method: In the second working mode, simultaneously triggering an audio acquisition system and a video acquisition system to collect information from a preset calibration device, respectively obtaining an audio stream generated by using the first system clock and a collected video stream generated by using the second system clock; wherein, a lamp with a preset shape is arranged on the calibration device, and the calibration device periodically emits a sound with a preset duration, and lights the lamp with the preset shape while emitting the sound; According to the preset standard audio, performing voice recognition on the audio stream collected from the calibration device, and determining an audio frame sequence corresponding to the standard audio through the voice recognition; and according to the preset standard image, performing image recognition on the video stream collected from the calibration device, and determining a video frame sequence corresponding to the standard image through the image recognition; Determining the timestamp calibration value according to the audio frame sequence and the video frame sequence.

2. The method according to claim 1, wherein Determining the timestamp calibration value according to the audio frame sequence and the video frame sequence specifically includes: Respectively determining the timestamp of the last frame of the audio frame sequence and the timestamp of the last frame of the video frame sequence; Taking the difference between the timestamp of the last frame of the video frame sequence and the timestamp of the last frame of the audio frame sequence as the timestamp calibration value.

3. The method according to claim 1, characterized in that, The video frame sequence is a video frame sequence of N consecutive frames including the standard image, where t*f - c < N < t*f + c, c is a preset constant, f is the frame rate of video encoding, and t is the preset duration.

4. The method according to claim 1, characterized in that According to the preset standard audio, performing voice recognition on the audio stream collected from the calibration device, and determining an audio frame sequence corresponding to the standard audio through the voice recognition specifically includes: Putting the audio stream collected from the calibration device into a preset first-in-first-out FIFO buffer queue after decoding, and the size of the FIFO buffer queue is equal to the data size of the standard audio; Matching the data in the FIFO buffer queue with the standard audio, and when the matching is successful, determining the audio frame sequence in the FIFO buffer queue as the audio frame sequence corresponding to the standard audio.

5. The method according to claim 1, wherein Receiving a user instruction through a user interface output to the user to implement the switching between the first working mode and the second working mode.

6. The method according to claim 1, wherein Extracting timestamps of the audio stream frame and / or the video stream frame; Moreover, calibrating the timestamp based on the timestamp calibration value determined by using a preset standard image and standard audio in the second working mode specifically includes: extracting the timestamp of the audio stream frame; using the sum of the timestamp and the timestamp calibration value as the new timestamp and updating it to the audio stream frame.

7. An electronic device, characterized in that, including: a memory for storing program instructions; a processor for calling the program instructions stored in the memory and executing the method according to any one of claims 1 to 6 according to the obtained program.

8. A computer program product for a computer, characterized in that, including a software code portion that, when the product runs on the computer, is used to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio and video synchronization test and correction method and device and electronic device

    CN112188259A

  • Audio and video synchronization method and device, electronic device and storage medium

    CN114339454A