An audio and video processing synchronization method and device, electronic equipment and storage medium

By separating and unifying the frame rate and sampling rate of audio and video data, the problem of poor synchronization in multi-channel audio and video processing is solved, achieving stable video and clear audio playback, and reducing resource waste.

CN115589498BActive Publication Date: 2025-11-04安徽文香科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211078332.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-11-04
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

When processing multiple audio and video signals, different video frame rates or different audio sampling rates can lead to poor synchronization, unstable video signal input, audio distortion, and serious waste of resources.

Method used

By acquiring multiple audio and video data, separating them into video and audio data, unifying the video frame rate and adjusting the audio playback duration, and performing synchronous processing under a reference clock, including frame extraction or frame interpolation to achieve frame rate uniformity, and adjusting the audio playback duration according to the data communication transmission volume, the data is then synthesized for output or display.

Benefits of technology

It improves the problems of unstable video and distorted audio, reduces resource consumption, and enhances the synchronization effect of multi-channel audio and video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115589498B_ABST
    Figure CN115589498B_ABST
Patent Text Reader

Abstract

The application provides an audio and video processing synchronization method and device, electronic equipment and a storage medium. The method comprises the following steps: separating at least two audio and video data to obtain at least two video data and at least two audio data; under a reference clock, unifying the frame rates of the at least two video data according to a reference frame rate and a current frame rate of any video data to obtain at least two video data after frame rate unification; adjusting the playing time of the at least two audio data, wherein the playing time of any audio at any time is equal to the current audio playing time plus the audio data length divided by the data communication transmission amount, to obtain at least two adjusted audio data; and outputting or sending the at least two video data after frame rate unification and the at least two adjusted audio data to other equipment or directly displaying. Through the application, the problem of poor synchronization effect in related art when multiple audio and video data are processed is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio and video processing, and more particularly to an audio and video processing synchronization method, apparatus, electronic device, and storage medium. Background Technology

[0002] When multiple different audio and video signal sources (which can be video signals captured by a device, recorded multimedia files, or network video streams) are processed in the same device or under the same clock signal, differences in video frame rates or audio sampling rates can lead to asynchrony between the multiple audio and video streams. Currently, to achieve synchronized playback of multiple audio and video streams, a selector can be used to switch which input is sent to the receiving terminal. After receiving the audio and video from each stream, the receiving terminal mixes and encodes them into a single audio and video data stream and sends it to each sending terminal. However, when processing audio and video from multiple different signal sources, video signal format issues can cause signal inaccessibility or instability, and different audio sampling rates can lead to sound distortion. Furthermore, if a particular stream is not needed for processing, it can consume resources. Therefore, existing technologies suffer from poor synchronization when processing multiple audio and video streams. Summary of the Invention

[0003] This invention provides an audio and video processing synchronization method, apparatus, electronic device, and storage medium to at least solve the problem of poor synchronization effect when processing multiple audio and video streams in related technologies.

[0004] According to a first aspect of the present invention, an audio-video processing synchronization method is provided. The method includes: acquiring at least two audio-video data; separating the at least two audio-video data to obtain at least two video data and at least two audio data, wherein the video data is raw video data; under a reference clock, unifying the frame rate of the at least two video data according to a reference frame rate and the current frame rate of any video data to obtain at least two video data with unified frame rate; adjusting the playback duration of the at least two audio data to obtain at least two audio data with adjusted frame rate, wherein the playback duration of any audio at any time is equal to the current audio playback duration plus the audio data length divided by the data communication transmission amount, wherein the data communication transmission amount is equal to the product of the sampling rate, the size of a single data item, and the number of channels; and synthesizing and outputting the at least two video data with unified frame rate and the at least two audio data with adjusted frame rate, or sending them to other devices or displaying them directly.

[0005] Optionally, the step of unifying the frame rates of the at least two video data under a reference clock based on a reference frame rate and the current frame rate of any video data to obtain at least two video data with unified frame rates includes: when the frame rates of the multiple video data are not the same, extracting frames when the current frame rate of any video data is greater than the reference frame rate; and supplementing frames when the current frame rate of any video data is less than the reference frame rate to obtain at least two video data with unified frame rates.

[0006] Optionally, the step of frame extraction when the current frame rate of any video data is greater than the reference frame rate includes: obtaining a frame extraction factor by dividing the current frame rate by the reference frame rate when the current frame rate of any video data is greater than the reference frame rate; if the frame extraction factor is an integer, extracting frames based on the number of frames and the frame extraction factor; if the frame extraction factor is not an integer, extracting frames based on the period under the reference clock and the frame extraction factor.

[0007] Optionally, the step of adding frames when the current frame rate of any video data is less than the reference frame rate includes: when the current frame rate of any video data is less than the reference frame rate, obtaining a frame addition factor by dividing the reference frame rate by the current frame rate; if the frame addition factor is an integer, adding frames according to the number of frames and the frame addition factor; if the frame addition factor is not an integer, adding frames according to the period under the reference clock and the frame addition factor.

[0008] Optionally, before separating the at least two audio and video data, the method includes: obtaining the at least two audio and video data by decoding when the signal source of the audio and video data is a network stream or a file; after separating the at least two audio and video data, the method includes: obtaining at least two video data, and controlling the at least two video data to be connected, disconnected, or dynamically shut down, wherein the connection is used to control the transmission and processing of the at least two video data, the disconnection is used to control the at least two video data to stop transmission and processing, and the dynamic shutdown is used to remotely control the signal source to shut down.

[0009] Optionally, the method further includes: converting the memory layout of at least two video data after frame rate unification according to a preset color space to obtain at least two video data after memory layout conversion; obtaining the data format corresponding to at least two video data; and outputting the processed at least two video data according to the at least two video data after memory layout conversion and the data format corresponding to the at least two video data.

[0010] According to a second aspect of the present invention, an audio-video processing synchronization device is also provided, the device comprising: a first obtaining module, configured to acquire at least two audio-video data, separate the at least two audio-video data to obtain at least two video data and at least two audio data, wherein the video data is raw video data; a second obtaining module, configured to, under a reference clock, unify the frame rate of the at least two video data according to a reference frame rate and the current frame rate of any video data, to obtain at least two video data with unified frame rate; a third obtaining module, configured to adjust the playback duration of the at least two audio data to obtain at least two adjusted audio data, wherein the playback duration of any audio at any time is equal to the current audio playback duration plus the audio data length divided by the data communication transmission amount, wherein the data communication transmission amount is equal to the product of the sampling rate, the size of a single data item, and the number of channels; and a synthesis module, configured to synthesize and output, or send to other devices, or directly display, the at least two video data with unified frame rate and the at least two adjusted audio data according to their correspondence.

[0011] Optionally, the second obtaining module includes: a frame extraction unit, used to extract frames when the current frame rate of any video data is greater than the reference frame rate when the frame rates of the plurality of video data are not the same; and a frame supplementation unit, used to supplement frames when the current frame rate of any video data is less than the reference frame rate, so as to obtain at least two video data with unified frame rates.

[0012] Optionally, the frame extraction unit includes: a obtaining submodule, configured to obtain a frame extraction factor by dividing the current frame rate by the reference frame rate when the current frame rate of any video data is greater than the reference frame rate; a first frame extraction submodule, configured to extract frames based on the number of frames and the frame extraction factor when the frame extraction factor is an integer; and a second frame extraction submodule, configured to extract frames based on the period under the reference clock and the frame extraction factor when the frame extraction factor is not an integer.

[0013] Optionally, the frame interpolation unit includes: a obtaining submodule, configured to obtain an interpolation factor by dividing the reference frame rate by the current frame rate when the current frame rate of any video data is less than the reference frame rate; a first interpolation submodule, configured to interpolate frames according to the number of frames and the interpolation factor when the interpolation factor is an integer; and a second interpolation submodule, configured to interpolate frames according to the period under the reference clock and the interpolation factor when the interpolation factor is a non-integer.

[0014] Optionally, the device further includes: a fourth obtaining module, used to obtain the at least two audio and video data by decoding when the signal source of the audio and video data is a network stream or a file; and a control module, used to obtain the at least two video data and control the at least two video data to be connected, disconnected, or dynamically shut down, wherein the connection is used to control the transmission and processing of the at least two video data, the disconnection is used to control the at least two video data to stop transmission and processing, and the dynamic shutdown is used to remotely control the signal source to shut down.

[0015] Optionally, the device further includes: a fifth obtaining module, configured to convert the memory layout of at least two video data after frame rate unification according to a preset color space, to obtain at least two video data after memory layout conversion; and an output module, configured to obtain the data format corresponding to at least two video data, and output the processed at least two video data according to the at least two video data after memory layout conversion and the data format corresponding to the at least two video data.

[0016] According to a third aspect of the present invention, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; wherein the memory is used to store a computer program; and the processor is used to execute the method steps of any of the above embodiments by running the computer program stored in the memory.

[0017] According to a fourth aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to execute the method steps of any of the above embodiments when running.

[0018] In this embodiment of the invention, at least two video data and at least two audio data are obtained by separating at least two audio and video data. Under a reference clock, the frame rates of the at least two video data are unified according to the reference frame rate and the current frame rate of any video data, resulting in at least two video data with unified frame rates. The playback duration of the at least two audio data is adjusted, with the playback duration of any audio at any given time equal to the current audio playback duration plus the audio data length divided by the data communication transmission volume, resulting in at least two adjusted audio data. The at least two video data with unified frame rates and the at least two adjusted audio data are then synthesized and output, either sent to other devices or directly displayed. Because the frame rates of multiple video streams and the sampling rates of multiple audio streams are synchronized under a reference clock, the problems of unstable video images and distorted audio sounds are improved, thereby solving the problem of poor synchronization effect in multi-channel audio and video processing in related technologies.

[0019] In this embodiment of the invention, by controlling connectivity, disconnection, or dynamic shutdown during data processing, the technical effect of reducing resource consumption is achieved, thus solving the problem of resource waste during multi-channel audio and video processing in related technologies. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the hardware environment for an optional audio and video processing synchronization method according to an embodiment of the present invention;

[0023] Figure 2 This is a flowchart illustrating an optional audio-visual processing synchronization method according to an embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of an optional video processing method according to an embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram of another optional video processing method according to an embodiment of the present invention;

[0026] Figure 5 This is a structural block diagram of an optional audio and video processing synchronization method according to an embodiment of the present invention;

[0027] Figure 6 This is a structural block diagram of an optional electronic device according to an embodiment of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] According to a first aspect of the present invention, an audio-video processing synchronization method is provided. Optionally, in this embodiment, the above-described audio-video processing synchronization method can be applied to, for example... Figure 1 In the hardware environment shown. For example... Figure 1 As shown, terminal 102 may include memory 104, processor 106, and display 108 (optional component). Terminal 102 can communicate with server 112 via network 110. Server 112 can provide services (such as application services) to the terminal or clients installed on the terminal. Database 114 can be set up on or independently of server 112 to provide data storage services to server 112. In addition, server 112 may run a processing engine 116, which can be used to execute the steps performed by server 112.

[0031] Optionally, terminal 102 may be, but is not limited to, a terminal capable of computing data, such as a mobile terminal (e.g., mobile phone, tablet computer), laptop computer, PC (Personal Computer), etc. The aforementioned network may include, but is not limited to, a wireless network or a wired network. The wireless network includes Bluetooth, Wi-Fi (Wireless Fidelity), and other networks that enable wireless communication. The aforementioned wired network may include, but is not limited to, a wide area network (WAN), a metropolitan area network (MAN), and a local area network (LAN). The aforementioned server 112 may include, but is not limited to, any hardware device capable of computing.

[0032] Furthermore, in this embodiment, the above-described audio and video processing synchronization method can also be applied to, but is not limited to, a powerful independent processing device without the need for data interaction. For example, the processing device can be, but is not limited to, a powerful terminal device; that is, the various operations in the above-described audio and video processing synchronization method can be integrated into a single independent processing device. The above is merely an example, and no limitation is made in this embodiment.

[0033] Optionally, in this embodiment, the above-described audio and video processing synchronization method can be executed by the server 112, by the terminal 102, or by both the server 112 and the terminal 102. The terminal 102 executing the audio and video processing synchronization method of this embodiment can also be executed by a client installed on it.

[0034] Taking the application of audio and video processing synchronization methods to the central processing unit as an example, Figure 2 This is a flowchart illustrating an optional audio-video processing synchronization method according to an embodiment of the present invention, as shown below. Figure 2 As shown, the process of this method may include the following steps:

[0035] Step S201: Acquire at least two audio and video data sets, separate the at least two audio and video data sets to obtain at least two video data sets and at least two audio data sets, wherein the video data sets are raw video data. Optionally, in this embodiment of the invention, acquiring at least two audio and video data sets from at least two different audio and video signal sources can be video signals acquired by a device, multimedia files generated from recording, or network video streams. Separating the at least two audio and video data sets to obtain at least two video data sets and at least two audio data sets, wherein the video data sets include RGB video pixel data or YUV video pixel data. Specifically, in the RGB space, any color can be generated by mixing the three primary colors of red, green, and blue in different proportions, representing brightness, hue, and saturation together; in the YUV space, each color has a brightness signal Y and two chromaticity signals U and V, where Y represents black and white information, and U and V represent color information. Similar to RGB, YUV is also a color encoding method in the video field. It separates luminance information (Y) from color information (UV). A complete image can be displayed even without UV information, but it will be black and white, thus achieving compatibility between color and black and white televisions.

[0036] Step S202: Under the reference clock, the frame rates of at least two video data are unified according to the reference frame rate and the current frame rate of any video data, resulting in at least two video data with unified frame rates. Optionally, the reference clock is a reference clock for audio and video synchronization. Time on the reference clock increases linearly, and both video data and audio processing are performed under the reference clock. Specifically, the video frame rate may be different values ​​such as 60 FPS (Frames Per Second), 30 FPS, or 25 FPS. In this case, a reference frame rate can be set, and the frame rates of at least two video data are unified according to the reference frame rate and the current frame rate of any video data, thereby obtaining at least two video data with unified frame rates.

[0037] Step S203: Adjust the playback duration of at least two audio data points to obtain adjusted at least two audio data points. The playback duration of any audio data point at any given time is equal to the current playback duration of the audio data plus the audio data length divided by the data communication transmission volume. The data communication transmission volume is equal to the product of the sampling rate, the size of a single data point, and the number of channels. Optionally, the playback duration of at least two audio data points is adjusted. Specifically, the playback duration of any audio data point at any given time is equal to the current playback duration of the audio data plus the audio data length divided by the data communication transmission volume. The data communication transmission volume is equal to the product of the sampling rate, the size of a single data point, and the number of channels. For example, the specific calculation formula for a sampling bit depth of 16 and a channel count of 2 is as follows:

[0038] double time = datalen / (sampling rate * 16 / 8 * 2);

[0039] audio->clock=audio->clock'+double time

[0040] Here, datalen represents the length of the audio data. It should be noted that the data communication transmission volume is in bytes per second. 1 byte equals 8 bits, so the size of a single data is 16 / 8. 2 indicates that it is a dual-channel system. audio->clock and audio->clock' represent the playback duration of the audio at any given moment and the current audio playback duration, respectively.

[0041] Step S204: Based on the matching of at least two video data points after frame rate unification and at least two audio data points after adjustment, synthesize and output the data, either sending it to other devices or displaying it directly. Optionally, specify the video output format for the processed video, synthesize it according to the audio playback duration, and output it. Then, allocate it to the display device as needed or let it enter the next processing signal. Alternatively, the input source can be directly echoed back to the display device without undergoing the above steps.

[0042] In this embodiment of the invention, at least two video data and at least two audio data are obtained by separating at least two audio and video data. Under a reference clock, the frame rates of the at least two video data are unified according to the reference frame rate and the current frame rate of any video data, resulting in at least two video data with unified frame rates. The playback duration of the at least two audio data is adjusted, with the playback duration of any audio at any given time equal to the current audio playback duration plus the audio data length divided by the data communication transmission volume, resulting in at least two adjusted audio data. The at least two video data with unified frame rates and the at least two adjusted audio data are then synthesized and output, either sent to other devices or directly displayed. Because the frame rates of multiple video streams and the sampling rates of multiple audio streams are synchronized under a reference clock, the problems of unstable video images and distorted audio sounds are improved, thereby solving the problem of poor synchronization effect in multi-channel audio and video processing in related technologies.

[0043] As an optional embodiment, under a reference clock, unifying the frame rates of at least two video data points based on a reference frame rate and the current frame rate of any video data to obtain at least two video data points with unified frame rates includes: when the frame rates of multiple video data points are different, skipping frames when the current frame rate of any video data point is greater than the reference frame rate; and padding frames when the current frame rate of any video data point is less than the reference frame rate, to obtain at least two video data points with unified frame rates. Optionally, when the frame rates of at least two video data points are different, for example, if the input source has video data at 60 FPS and 30 FPS, it is equivalent to different processing time periods, which will cause asynchrony in the input source frame rates. In this case, a reference frame rate can be set, skipping frames when the current frame rate of any video data point is greater than the reference frame rate, and padding frames when the current frame rate of any video data point is less than the reference frame rate. In this embodiment of the invention, at least two video data points with unified frame rates are obtained through frame skipping and frame padding.

[0044] As an optional embodiment, frame extraction when the current frame rate of any video data is greater than the reference frame rate includes: obtaining a frame extraction factor by dividing the current frame rate by the reference frame rate when the current frame rate of any video data is greater than the reference frame rate; if the frame extraction factor is an integer, extracting frames based on the number of frames and the frame extraction factor; if the frame extraction factor is not an integer, extracting frames based on the period under the reference clock and the frame extraction factor.

[0045] Optionally, when the input source has a frame rate of 60FPS, 30FPS, or other frame rates, it is equivalent to different processing time periods for frames. This can cause asynchrony in the frame rates of the input sources. To avoid this, a baseline frame rate F is specified. When the current frame rate F1 of an input source is greater than the baseline frame rate F, average frame subtraction is performed according to the subtraction factor F1 / F. For example, if the baseline frame rate is specified as 30FPS, and the subtraction factor of the 60FPS input source is an integer of 2, then frames 1, 2, 3, 4, 5, 6, 7, 8, 9… are subtracted by an average of 2 frames, i.e., 2, 3, 5, 6, 8, 9… are extracted, leaving 1, 4, 7… thus achieving a uniform frame rate. When the input source has video data with a frame rate of 30FPS, if the baseline frame rate is 25FPS and the subtraction factor is 1.2 (not an integer), then between the start time t1 and the end time t2 of video processing, the video data is processed according to the trigger time of each interval cycle under the baseline clock. Specifically, if the period is 33ms, then between t1 and t2, frames are extracted based on the average time point t = t1 + dt (dt can be 33 * frame extraction factor) within that period. In this embodiment of the invention, by calculating the frame extraction factor to process the video data, the effect of uniform frame rate is achieved.

[0046] As an optional embodiment, when the current frame rate of any video data is less than the reference frame rate, frame interpolation includes: when the current frame rate of any video data is less than the reference frame rate, obtaining an interpolation factor by dividing the reference frame rate by the current frame rate; if the interpolation factor is an integer, interpolating frames according to the number of frames and the interpolation factor; if the interpolation factor is not an integer, interpolating frames according to the period under the reference clock and the interpolation factor.

[0047] Optionally, when the current frame rate F2 of an input source is less than the reference frame rate F, frame interpolation is performed according to the interpolation factor F / F2. For example, if there are frames 1, 2, 3, 4, 5, 6, 7, 8, 9… interpolated by the interpolation factor 2, they can be represented as 1, 11, 12, 2, 21, 22, 3, 31, 32, 4… It should be noted that 11, 12, 21, 22, 31, 32 are all based on the previous frame image, with image copying compensation. When there is video data with a frame rate of 25 FPS, if the reference frame rate is 30 FPS, the interpolation factor for the 25 FPS video data is 1.2, which is not an integer. In this case, between the start time t1 and the end time t2 of video processing, the video data is processed according to the trigger time of each interval cycle under the reference clock. Specifically, if the period is 33ms, then between t1 and t2, frame interpolation is performed based on the average time point t = t1 + dt (dt can be 33 * interpolation factor) within that period. In this embodiment of the invention, by calculating the interpolation factor to process the video data, the effect of uniform frame rate is achieved.

[0048] As an optional embodiment, before separating at least two audio and video data, the process includes: obtaining at least two audio and video data by decoding when the signal source of the audio and video data is a network stream or a file; after separating the at least two audio and video data, the process includes: obtaining at least two video data, and controlling the at least two video data to be connected, disconnected, or dynamically shut down, wherein connection is used to control the transmission and processing of at least two video data, disconnection is used to control the cessation of transmission and processing of at least two video data, and dynamic shutdown is used to remotely control the shutdown of the signal source. Optionally, before separating the audio and video data, when the input signal source is a network stream or a file, decoding is required first, and then the decoded audio and video data is separated to obtain audio data and video data. During the processing of video data, necessary data transmission and processing can be controlled, unnecessary data transmission and processing can be stopped, data can be discarded to reduce resource consumption, or the signal source can be remotely controlled to be dynamically shut down.

[0049] In this embodiment of the invention, the data processing process is flexibly adjusted according to the current configuration, controlling whether it is connected, disconnected, or dynamically shut down, thereby achieving the technical effect of reducing resource consumption and solving the problem of resource waste during multi-channel audio and video processing in related technologies.

[0050] As an optional embodiment, the audio and video processing synchronization method further includes: converting the memory layout of at least two video data sets after frame rate unification according to a preset color space to obtain at least two video data sets with converted memory layouts; obtaining the data format corresponding to the at least two video data sets; and outputting the processed at least two video data sets according to the at least two video data sets with converted memory layouts and the corresponding data formats. Optionally, after unifying the frame rate of at least two video data sets, the memory layout of the at least two video data sets after frame rate unification can be converted according to a preset color space. Specifically, if the color space for predetermined negotiation processing is specified as RGB32, the input data is converted according to the RGB32 memory layout, and the current video data format is generated, including size, color space format, and bit depth. Further processing of multiple video frames is performed according to the memory layout, such as special effects transitions, video overlay, adding subtitles, and virtual keying. In this embodiment of the invention, the conversion of color space and memory layout achieves the effect of processing video according to actual needs.

[0051] As an optional embodiment, Figure 3 This is a schematic diagram of an optional video processing method according to an embodiment of the present invention. This method can be implemented using hardware or software. Figure 3As shown, the switch / pause receives video data (i.e., raw video data) and controls the downstream transmission of the video data, stops transmission, or remotely controls the dynamic shutdown of the signal source. It should be noted that when the input signal source is a network stream or file, it is first decoded using a decoder, such as... Figure 4 As shown, video data is decoded and then sent to a switcher / pause unit for processing. The switcher / pause unit sends the video data to a frame rate converter. The frame rate converter performs frame extraction or interpolation based on the relationship between the current frame rate and the reference frame rate of the video data to achieve frame rate unification. The color space converter receives the video data output by the frame rate converter, performs memory layout conversion on at least two video data sets after frame rate unification according to a preset color space, and generates the current video data format, which is then sent to a memory-based data processing converter. The memory-based data processing converter further processes the multiple video frames according to the memory layout. The processed output data can be displayed and rendered by the switcher / pause unit, sent to the encoder, or used to display and render the data from the input source. In this embodiment of the invention, through a series of processing steps on the video data, the conversion of frame rate, color space, and data format of at least two video data sets is achieved as needed.

[0052] According to another aspect of the present invention, an audio-video synchronization apparatus for implementing the above-described audio-video synchronization method is also provided. Figure 5 This is a structural block diagram of an optional audio-video synchronization device according to an embodiment of the present invention, such as... Figure 5 As shown, the device may include: a first obtaining module 501, used to acquire at least two audio and video data, separate the at least two audio and video data to obtain at least two video data and at least two audio data, wherein the video data is raw video data; a second obtaining module 502, used to unify the frame rate of the at least two video data according to the reference frame rate and the current frame rate of any video data under a reference clock, to obtain at least two video data with unified frame rate; a third obtaining module 503, used to adjust the playback duration of the at least two audio data to obtain at least two adjusted audio data, wherein the playback duration of any audio at any time is equal to the current audio playback duration plus the audio data length divided by the data communication transmission volume, and the data communication transmission volume is equal to the product of the sampling rate, the size of a single data, and the number of channels; and a synthesis module 504, used to synthesize and output the at least two video data with unified frame rate and the at least two adjusted audio data, or send them to other devices or display them directly.

[0053] It should be noted that the first obtaining module 501 in this embodiment can be used to execute the above step S201, the second obtaining module 502 in this embodiment can be used to execute the above step S202, the third obtaining module 503 in this embodiment can be used to execute the above step S203, and the synthesis module 504 in this embodiment can be used to execute the above step S204.

[0054] The above modules separate at least two audio and video data sets to obtain at least two video data sets and at least two audio data sets. Under a reference clock, the frame rates of the at least two video data sets are unified according to the reference frame rate and the current frame rate of any video data set, resulting in at least two video data sets with unified frame rates. The playback duration of the at least two audio data sets is adjusted, with the playback duration of any audio set at any given time equal to the current audio playback duration plus the audio data length divided by the data communication transmission volume, resulting in at least two adjusted audio data sets. The at least two video data sets with unified frame rates and the at least two adjusted audio data sets are then synthesized and output, either sent to other devices or directly displayed. Because the frame rates of multiple video streams and the sampling rates of multiple audio streams are synchronized under a reference clock, the problems of unstable video images and distorted audio sounds are improved, thus solving the problem of poor synchronization in multi-channel audio and video processing in related technologies.

[0055] As an optional embodiment, the second obtaining module includes: a frame extraction unit, used to extract frames when the current frame rate of any video data is greater than the reference frame rate when the frame rates of multiple video data are not the same; and a frame interpolation unit, used to interpolate frames when the current frame rate of any video data is less than the reference frame rate, so as to obtain at least two video data with unified frame rates.

[0056] As an optional embodiment, the frame extraction unit includes: a obtaining submodule, used to obtain a frame extraction factor by dividing the current frame rate by the reference frame rate when the current frame rate of any video data is greater than the reference frame rate; a first frame extraction submodule, used to extract frames according to the number of frames and the frame extraction factor when the frame extraction factor is an integer; and a second frame extraction submodule, used to extract frames according to the period under the reference clock and the frame extraction factor when the frame extraction factor is not an integer.

[0057] As an optional embodiment, the frame interpolation unit includes: a obtaining submodule, used to obtain an interpolation factor by dividing the reference frame rate by the current frame rate when the current frame rate of any video data is less than the reference frame rate; a first interpolation submodule, used to interpolate frames according to the number of frames and the interpolation factor when the interpolation factor is an integer; and a second interpolation submodule, used to interpolate frames according to the period under the reference clock and the interpolation factor when the interpolation factor is not an integer.

[0058] As an optional embodiment, the device further includes: a fourth obtaining module, used to obtain at least two audio and video data by decoding when the signal source of the audio and video data is a network stream or a file; and a control module, used to obtain at least two video data and control the at least two video data to be connected, disconnected, or dynamically shut down, wherein connecting is used to control the transmission and processing of at least two video data, disconnecting is used to control the at least two video data to stop transmission and processing, and dynamically shutting down is used to remotely control the signal source to be shut down.

[0059] As an optional embodiment, the device further includes: a fifth obtaining module, configured to convert the memory layout of at least two video data after frame rate unification according to a preset color space, to obtain at least two video data after memory layout conversion; and an output module, configured to obtain the data format corresponding to at least two video data, and output the processed at least two video data according to the at least two video data after memory layout conversion and the data format corresponding to at least two video data.

[0060] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should also be noted that the above modules, as part of a device, can operate in environments such as... Figure 1 The hardware environment shown can be implemented through software or hardware, and the hardware environment includes the network environment.

[0061] According to another aspect of the present invention, an electronic device for implementing the above-described audio and video processing synchronization method is also provided. The electronic device may be a server, a terminal, or a combination thereof.

[0062] Figure 6 This is a structural block diagram of an optional electronic device according to an embodiment of the present invention, such as... Figure 6As shown, the system includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. The processor 601, communication interface 602, and memory 603 communicate with each other via the communication bus 604. The memory 603 stores computer programs. When the processor 601 executes the computer program stored in the memory 603, it performs the following steps: acquiring at least two audio and video data sets; separating the at least two audio and video data sets to obtain at least two video data sets and at least two audio data sets, wherein the video data sets are raw video data; and under a reference clock, according to the base clock... The frame rate of at least two video data points is unified with the frame rate of any video data point and the current frame rate of any video data point to obtain at least two video data points with unified frame rates. The playback duration of at least two audio data points is adjusted to obtain at least two audio data points with adjusted frame rates. The playback duration of any audio data point at any given time is equal to the current audio playback duration plus the audio data length divided by the data communication transmission volume. The data communication transmission volume is equal to the product of the sampling rate, the size of a single data point, and the number of channels. The at least two video data points with unified frame rates and the at least two audio data points with adjusted frame rates are then synthesized and output, either sent to other devices or displayed directly.

[0063] Optionally, in this embodiment, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0064] The communication interface is used for communication between the aforementioned electronic device and other devices. The memory may include RAM, or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0065] As an example, such as Figure 6 As shown, the memory 603 may include, but is not limited to, the first obtaining module 501, the second obtaining module 502, the third obtaining module 503, and the synthesis module 504 in the audio-visual processing synchronization device. Furthermore, it may include, but is not limited to, other module units in the audio-visual processing synchronization device, which will not be elaborated upon in this example.

[0066] The aforementioned processor can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Furthermore, the aforementioned electronic device also includes: a display for displaying the audio and video processing synchronization results.

[0067] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0068] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. The device implementing the above audio and video processing synchronization method can be a terminal device, such as a smartphone (e.g., Android phone, iOS phone), tablet computer, PDA, mobile Internet Devices (MID), PAD, etc. Figure 6 This does not limit the structure of the aforementioned electronic devices. For example, the terminal device may also include components that are more advanced than those described above. Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.

[0069] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, ROM, RAM, disk or optical disk, etc.

[0070] According to another aspect of the present invention, a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to execute program code for an audio / video processing synchronization method.

[0071] Optionally, in this embodiment, the storage medium may be located on at least one of the network devices in the network shown in the above embodiment.

[0072] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: acquiring at least two audio and video data; separating the at least two audio and video data to obtain at least two video data and at least two audio data, wherein the video data is raw video data; under a reference clock, unifying the frame rate of the at least two video data according to the reference frame rate and the current frame rate of any video data to obtain at least two video data with unified frame rate; adjusting the playback duration of the at least two audio data to obtain at least two audio data with adjusted frame rate, wherein the playback duration of any audio at any time is equal to the current audio playback duration plus the audio data length divided by the data communication transmission volume, and the data communication transmission volume is equal to the product of the sampling rate, the size of a single data item, and the number of channels; and synthesizing and outputting the at least two video data with unified frame rate and the at least two audio data with adjusted frame rate according to the reference clock, or sending them to other devices or displaying them directly.

[0073] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated in this embodiment.

[0074] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.

[0075] According to another aspect of the present invention, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the audio and video processing synchronization method steps in any of the above embodiments.

[0076] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0077] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the audio and video processing synchronization method of the various embodiments of the present invention.

[0078] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0079] In the several embodiments provided by this invention, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0080] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the solution provided in this embodiment, depending on actual needs.

[0081] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0082] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for synchronizing audio and video processing, characterized in that, The method includes: Acquire at least two audio and video data sets, separate the at least two audio and video data sets to obtain at least two video data sets and at least two audio data sets, wherein the video data sets are raw video data sets; Under a reference clock, the frame rates of the at least two video data are unified according to the reference frame rate and the current frame rate of any video data to obtain at least two video data with unified frame rates. The playback duration of the at least two audio data is adjusted to obtain at least two adjusted audio data, wherein the playback duration of any audio at any time is equal to the current audio playback duration plus the audio data length divided by the data communication transmission volume, and the data communication transmission volume is equal to the product of the sampling rate, the size of a single data and the number of channels; Based on the correspondence between at least two video data points after frame rate unification and at least two audio data points after adjustment, the data is synthesized and output, either sent to other devices or displayed directly. The step of unifying the frame rates of at least two video data points under a reference clock, based on a reference frame rate and the current frame rate of any video data point, to obtain at least two video data points with unified frame rates, includes: When the frame rates of multiple video data are not the same, a frame is extracted when the current frame rate of any video data is greater than the reference frame rate. If the current frame rate of any video data is less than the reference frame rate, frames are padded to obtain at least two video data with unified frame rates.

2. The method according to claim 1, characterized in that, The step of frame extraction when the current frame rate of any video data is greater than the reference frame rate includes: When the current frame rate of any video data is greater than the reference frame rate, the frame skipping factor is obtained by dividing the current frame rate by the reference frame rate. If the frame extraction factor is an integer, frames are extracted according to the number of frames and the frame extraction factor. If the frame-skipping factor is not an integer, frames are skimmed according to the period under the reference clock and the frame-skipping factor.

3. The method according to claim 1, characterized in that, The step of adding frames when the current frame rate of any video data is less than the reference frame rate includes: When the current frame rate of any video data is less than the reference frame rate, the frame interpolation factor is obtained by dividing the reference frame rate by the current frame rate. If the frame interpolation factor is an integer, interpolate frames according to the number of frames and the frame interpolation factor; If the frame interpolation factor is not an integer, the frame is interpolated according to the period under the reference clock and the frame interpolation factor.

4. The method according to claim 1, characterized in that, Before separating the at least two audio and video data, the process includes: obtaining the at least two audio and video data by decoding when the signal source of the audio and video data is a network stream or a file; After separating the at least two audio and video data, the process includes: obtaining at least two video data, and controlling the at least two video data to be connected, disconnected, or dynamically shut down, wherein the connection is used to control the transmission and processing of the at least two video data, the disconnection is used to control the at least two video data to stop transmission and processing, and the dynamic shutdown is used to remotely control the signal source to shut down.

5. The method according to claim 1, characterized in that, The method further includes: The memory layout of at least two video data sets after frame rate unification is converted according to a preset color space to obtain at least two video data sets after memory layout conversion. Obtain the data format corresponding to at least two video data, and output the processed at least two video data after conversion according to the memory layout and the data format corresponding to the at least two video data.

6. An audio-visual processing synchronization device, characterized in that, The device includes: The first obtaining module is used to acquire at least two audio and video data, separate the at least two audio and video data to obtain at least two video data and at least two audio data, wherein the video data is raw video data; The second obtaining module is used to unify the frame rate of the at least two video data according to the reference frame rate and the current frame rate of any video data under the reference clock, so as to obtain at least two video data with unified frame rate. The third module is used to adjust the playback duration of the at least two audio data to obtain the adjusted at least two audio data, wherein the playback duration of any audio at any time is equal to the current audio playback duration plus the audio data length divided by the data communication transmission amount, and the data communication transmission amount is equal to the product of the sampling rate, the size of a single data and the number of channels. The compositing module is used to synthesize and output, or send to other devices or display directly, based on at least two video data points after frame rate unification and at least two audio data points after adjustment. The second module includes: A frame extraction unit is used to extract frames when the current frame rate of any video data is greater than the reference frame rate, when the frame rates of multiple video data are not the same. The frame interpolation unit is used to interpolate frames when the current frame rate of any video data is less than the reference frame rate, so as to obtain at least two video data with unified frame rates.

7. The apparatus according to claim 6, characterized in that, The frame extraction unit includes: A submodule is obtained, which is used to obtain a frame skipping factor by dividing the current frame rate by the reference frame rate when the current frame rate of any video data is greater than the reference frame rate; The first frame extraction submodule is used to extract frames based on the number of frames and the frame extraction factor when the frame extraction factor is an integer. The second frame-skipping submodule is used to skip frames according to the period under the reference clock and the frame-skipping factor when the frame-skipping factor is not an integer.

8. The apparatus according to claim 6, characterized in that, The frame interpolation unit includes: A submodule is obtained, which is used to obtain a frame interpolation factor by dividing the reference frame rate by the current frame rate when the current frame rate of any video data is less than the reference frame rate; The first frame interpolation submodule is used to interpolate frames according to the number of frames and the frame interpolation factor when the frame interpolation factor is an integer; The second frame interpolation submodule is used to perform frame interpolation based on the period under the reference clock and the frame interpolation factor when the frame interpolation factor is not an integer.

9. An electronic device comprising a processor, a communication interface, a memory, and a communication bus, wherein, The processor, the communication interface, and the memory communicate with each other via the communication bus, characterized in that... The memory is used to store computer programs; The processor is configured to perform the method steps of any one of claims 1 to 5 by running the computer program stored in the memory.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Bluetooth sound and picture synchronization method and device, display equipment and readable storage medium

    CN112261461A

  • Multichannel vocoder

    CN1495705A