Audio and video synchronization detection method and device, computer device and storage medium
By generating synchronized audio and video files and utilizing character recognition and speech recognition technologies to calculate the differences in audio and video time information, the problem of the inability to quantitatively evaluate audio and video synchronization in existing technologies is solved, and automatic and accurate audio and video synchronization detection is achieved.
Patent Information
- Application Number
- CN202111453995.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-12-01
AI Technical Summary
Existing technologies cannot quantitatively evaluate the audio-video synchronization effect in multimedia systems, relying mainly on subjective evaluation methods, which cannot accurately determine the synchronization delay difference between audio and video.
By acquiring the text of the test audio, a synchronized audio and video file is generated and input into a multimedia system for processing. Using character recognition and speech recognition technologies, the time information difference between the video and audio is calculated to quantify the audio and video synchronization detection results.
It enables automatic and quantitative detection of audio-visual synchronization effects without the need for subjective human evaluation, thus improving the accuracy and comprehensiveness of the detection.
Smart Images

Figure CN114339199B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the software technical field, in particular to an audio and video synchronization detection method and device, computer equipment and storage medium. BACKGROUND
[0002] With the development of multimedia technology, the application of multimedia system has penetrated into various fields of human life, such as communication, industry, medicine, teaching, etc., which brings great convenience to people's life. Moreover, multimedia system often involves the processing of audio and video, which leads to the problem of audio and video asynchronization. Therefore, measuring the synchronization effect of audio and video of multimedia system has become one of the evaluation indexes of multimedia system.
[0003] However, at present, subjective evaluation method is mostly used to judge whether the output audio and video of multimedia system are synchronized, which cannot quantify the synchronization delay difference of audio and video. SUMMARY
[0004] Therefore, it is necessary to provide a quantifiable audio and video synchronization detection method, device, computer equipment, storage medium and computer program product aiming at the above technical problems.
[0005] In a first aspect, the present application provides an audio and video synchronization detection method, which comprises:
[0006] obtaining a test audio and a text corresponding to the test audio; the text comprises at least one target text;
[0007] generating a synchronization audio and video file for synchronously playing the text and the test audio;
[0008] inputting first audio data and first video data corresponding to the synchronization audio and video file into a multimedia system, triggering the multimedia system to process the first audio data and the first video data in a preset audio and video processing flow, and outputting processed second audio data and second video data;
[0009] performing character recognition on the video corresponding to the second video data to obtain the target text and first time information of each target text appearing in the corresponding video;
[0010] performing speech recognition on the audio corresponding to the second audio data to obtain the target text and second time information of pronunciation of each target text appearing in the corresponding audio;
[0011] comparing the difference between the first time information and the second time information, and obtaining an audio and video synchronization detection result based on the difference.
[0012] In one of the embodiments, the obtaining the test audio and the corresponding text of the test audio comprises:
[0013] obtaining a test audio for audio-video synchronization test;
[0014] performing speech recognition on the test audio to obtain a corresponding text.
[0015] In one of the embodiments, the comparing the difference between the first time information and the second time information, and obtaining the audio-video synchronization test result based on the difference comprises:
[0016] confirming the difference between the first time information and the second time information corresponding to each target text in the text;
[0017] obtaining the audio-video synchronization test result based on the difference.
[0018] In one of the embodiments, the obtaining the audio-video synchronization test result based on the difference comprises:
[0019] if the difference corresponding to each target text is less than a preset difference threshold, confirming that the audio-video synchronization test result is audio-video synchronization;
[0020] if the difference corresponding to any target text is greater than the preset difference threshold, confirming that the audio-video synchronization test result is audio-video asynchronization.
[0021] In one of the embodiments, the multimedia system comprises a collection end and an output end; the inputting the first audio data and the first video data in the synchronized audio-video file into the multimedia system, triggering the multimedia system to process the first audio data and the first video data in a preset audio-video processing flow, and outputting the processed second audio data and second video data comprises:
[0022] inputting the first audio data and the first video data in the synchronized audio-video file into the collection end, triggering the collection end to simultaneously collect the input first audio data and the first video data respectively, and performing encoding respectively to obtain the encoded audio data and video data;
[0023] sending the encoded audio data and video data to the output end;
[0024] decoding the received encoded audio data and video data through the output end, and outputting the processed second audio data and second video data.
[0025] In one of the embodiments, the multimedia system is a video conference system, the collecting terminal is a first video conference terminal in the video conference system, and the output terminal is a second video conference terminal in the video conference system. The video conference system further includes a multipoint control unit. The sending of the encoded audio data and video data to the output terminal includes:
[0026] sending, by the first video conference terminal, the encoded audio data and video data to the multipoint control unit;
[0027] decoding, by the multipoint control unit, the received encoded audio data and video data, and synthesizing each of the decoded audio data into a same target audio data and each of the decoded video data into a same target video data;
[0028] encoding, by the multipoint control unit, the target audio data and the target video data to obtain intermediate audio data and intermediate video data;
[0029] sending, by the multipoint control unit, the intermediate audio data and the intermediate video data to the second video conference terminal.
[0030] In a second aspect, the present application further provides an audio-video synchronization detection device. The device includes:
[0031] a preparation module configured to acquire a test audio and a text corresponding to the test audio, the text including at least one target text, and generate a synchronized audio-video file for synchronously playing the text and the test audio;
[0032] an input module configured to input first audio data and first video data corresponding to the synchronized audio-video file to a multimedia system, trigger the multimedia system to process the first audio data and the first video data in a preset audio-video processing flow, and output second audio data and second video data processed by the multimedia system;
[0033] a calculation module configured to perform character recognition on a video corresponding to the second video data to obtain the target text and first time information of each of the target text appearing in the corresponding video, and perform speech recognition on an audio corresponding to the second audio data to obtain the target text and second time information of pronunciation of each of the target text appearing in the corresponding audio;
[0034] a determination module configured to compare the first time information and the second time information, and obtain an audio-video synchronization detection result based on a difference between the first time information and the second time information.
[0035] In a third aspect, the present application provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor executes the steps of the audio-video synchronization detection method.
[0036] In a fourth aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to perform the steps of the audio-video synchronization detection method.
[0037] In a fifth aspect, the present application provides a computer program product. The computer program product comprises a computer program, and the computer program is executed by a processor to perform the steps of the audio-video synchronization detection method.
[0038] The audio-video synchronization detection method, device, computer device, storage medium and computer program product described above, by obtaining a test audio and a text corresponding to the test audio, the text comprising at least one target text, generating a synchronized audio-video file for synchronously playing the text and the test audio. Inputting first audio data and first video data corresponding to the synchronized audio-video file into a multimedia system, triggering the multimedia system to process the first audio data and the first video data in a preset audio-video processing flow, outputting processed second audio data and second video data, and the second audio data and the second video data processed by the multimedia system may be out of synchronization relative to the first audio data and the first video data in the synchronized audio-video file. Performing character recognition on the video corresponding to the second video data to obtain the target text and first time information of each target text appearing in the corresponding video. Performing speech recognition on the audio corresponding to the second audio data to obtain the target text and second time information of pronunciation of each target text appearing in the corresponding audio. Comparing the difference between the first time information and the second time information, i.e., the delay difference between audio and video, and obtaining an audio-video synchronization detection result based on the difference. Therefore, without the need for manual subjective evaluation, the audio-video synchronization is automatically detected, and the audio-video synchronization detection result is quantitatively obtained. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 An application environment diagram of the audio-video synchronization detection method in an embodiment;
[0040] Figure 2 A flowchart of the audio-video synchronization detection method in an embodiment;
[0041] Figure 3 A flowchart of the audio-video synchronization detection method in an embodiment;
[0042] Figure 4 a flowchart of an audio-video synchronization detection method in an embodiment;
[0043] Figure 5 a flowchart of an audio-video synchronization detection method in an embodiment;
[0044] Figure 6 a block diagram of an audio-video synchronization detection device in an embodiment;
[0045] Figure 7 an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0046] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0047] The audio-video synchronization detection development method provided by the present application can be applied in an application environment as shown in the figure. Figure 1 The terminal 110 communicates with the multimedia system 120 through a network. The terminal 110 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices, and the multimedia system 120 can be implemented by a system composed of at least one terminal or / and at least one server.
[0048] The terminal 110 can obtain test audio and text corresponding to the test audio; the text includes at least one target word. The terminal 110 generates a synchronized audio-video file for synchronously playing the text and the test audio. The terminal 110 inputs first audio data and first video data corresponding to the synchronized audio-video file to the multimedia system 120, triggers the multimedia system 120 to process the first audio data and the first video data in a preset audio-video processing flow, and outputs processed second audio data and second video data to the terminal 110. The terminal 110 performs character recognition on video corresponding to the second video data, obtains the target word and first time information of each target word appearing in the corresponding video. The terminal 110 performs speech recognition on audio corresponding to the second audio data, obtains the target word and second time information of pronunciation of each target word appearing in the corresponding audio. The terminal 110 compares the difference between the first time information and the second time information, and obtains an audio-video synchronization detection result based on the difference. It can be understood that the terminal 110 can include a first terminal and a second terminal. The processing performed by the terminal 110 is performed by the first terminal and the second terminal together. Specifically, the first terminal generates a synchronized audio-video file, and inputs corresponding first audio data and first video data to the multimedia system 120. The multimedia system 120 outputs processed second audio data and second video data to the second terminal. The second terminal compares the difference between the first time information and the second time information after obtaining the first time information and the second time information of each target word, and obtains an audio-video synchronization detection result based on the difference.
[0049] In one embodiment, as shown in Figure 2 An audio-video synchronization detection method is provided. The method can be applied to a terminal, a server, a system including a terminal and a server, and a system including multiple terminals. In one embodiment, the method includes the following steps:
[0050] S202, obtaining test audio and text corresponding to the test audio; the text includes at least one target word; and generating a synchronized audio-video file for synchronously playing the text and the test audio.
[0051] The test audio is prepared for performing the audio-video synchronization detection method.
[0052] In one embodiment, the text is obtained based on speech recognition of the test audio.
[0053] In another embodiment, the text is obtained based on manual annotation of the test audio.
[0054] In one embodiment, a synchronized audio-video file that plays text and test audio synchronously means that the audio in the synchronized audio-video file is the test audio, and the video in the synchronized audio-video file displays the target text in the text, and the appearance time of the target text is synchronized with the appearance time of the corresponding pronunciation of the target text in the audio of the synchronized audio-video file.
[0055] Specifically, the terminal can acquire the test audio and its corresponding text. The text includes at least one target character. Based on the test audio and text, the terminal generates a synchronized audio / video file for synchronously playing the text and test audio.
[0056] In one embodiment, such as Figure 3 As shown, after the terminal acquires the test audio, its speech recognition module processes the audio, performs speech recognition, and generates corresponding text. The terminal then generates a video displaying the text in black on a white background. Finally, the terminal combines the black-and-white video and the test audio into a synchronized multimedia file, thus generating the multimedia file to be played—the synchronized audio-video file.
[0057] S204, the first audio data and the first video data corresponding to the synchronized audio and video files are input to the multimedia system, triggering the multimedia system to process the first audio data and the first video data according to the preset audio and video processing flow, and outputting the processed second audio data and the second video data.
[0058] Multimedia systems utilize computer technology and digital communication network technology to process and control multimedia information, such as video conferencing systems. The first audio data corresponding to a synchronized audio / video file refers to the audio data used to display the audio within the synchronized audio / video file; its format is not limited. The first video data corresponding to a synchronized audio / video file refers to the video data used to display the video within the synchronized audio / video file; its format is not limited.
[0059] In one embodiment, the terminal can transmit synchronized audio and video files to the multimedia system, so that the first audio data and the first video data corresponding to the synchronized audio and video files are input into the multimedia system.
[0060] In another embodiment, the terminal can transmit compressed synchronized audio and video files to the multimedia system, so that the first audio data and the first video data corresponding to the synchronized audio and video files are input into the multimedia system.
[0061] In another embodiment, the terminal can play synchronized audio and video files and input the first audio data and first video data corresponding to the synchronized audio and video files to the multimedia system through a physical interface. For example, when the terminal plays synchronized audio and video files, it transmits the first audio data and first video data corresponding to the synchronized audio and video files to the acquisition end of the multimedia system through an AV (Audio & Video) cable.
[0062] In one embodiment, the multimedia system encodes, transmits, and decodes the input first audio data and first video data respectively, and then outputs the processed second audio data and second video data.
[0063] In one embodiment, the multimedia system is a video conferencing system, including multiple video conferencing terminals and a multipoint control unit for combining at least one audio stream and at least one video stream into one audio stream and one video stream. A first video conferencing terminal performs acquisition and encoding, and sends the data to the multipoint control unit. The multipoint control unit performs decoding, audio-video synthesis processing, and encoding, and then sends the data to a second video conferencing terminal. The second video conferencing terminal decodes the data to generate processed second audio data and second video data.
[0064] In one embodiment, the multimedia system can output a media file that includes second audio data and second video data.
[0065] In another embodiment, the second video conferencing terminal of the multimedia system plays audio and video, and transmits the corresponding second audio data and second video data to the terminal simultaneously through a physical interface.
[0066] Specifically, the terminal can input the first audio data and the first video data corresponding to the synchronized audio and video files into the multimedia system. The multimedia system processes the first audio data and the first video data according to the preset audio and video processing flow, generates second audio data and second video data that may be out of sync and to be detected, and outputs the second audio data and second video data to the terminal.
[0067] S206, perform character recognition on the video corresponding to the second video data to obtain the target text and the first time information of each target text appearing in the corresponding video; perform speech recognition on the audio corresponding to the second audio data to obtain the target text and the second time information of the pronunciation of each target text appearing in the corresponding audio.
[0068] The video corresponding to the second video data refers to the video generated based on the second video data, and the audio corresponding to the second audio data refers to the audio generated based on the second audio data. The formats of the second video data and the second audio data are not limited.
[0069] Character recognition refers to identifying displayed text characters and storing the recognition results as text in a computer. For example, it can identify text characters on paper or in an image and extract the corresponding text. Time information includes multiple time points, which can be at the millisecond, microsecond, or second level.
[0070] In one embodiment, the terminal can perform character recognition on the video corresponding to the second video data. When the target text is recognized, the corresponding time point is recorded to generate first time information.
[0071] In one embodiment, the terminal can perform speech recognition on the audio corresponding to the second audio data, and when the pronunciation of the target text appears in the corresponding audio, record the corresponding time point to generate second time information.
[0072] It is understandable that the second video data, generated based on synchronized audio and video files, includes the target text in the corresponding video. Specifically, the terminal performs character recognition on the video corresponding to the second video data. When the target text is recognized, the corresponding time point is recorded, thus obtaining the target text and the first time information of each target text appearing in the corresponding video. The terminal performs speech recognition on the audio corresponding to the second audio data. When the text corresponding to the pronunciation of the audio is recognized as the target text, the corresponding time point is recorded, thus obtaining the target text and the second time information of the pronunciation of each target text appearing in the corresponding audio.
[0073] S208, compare the difference between the first time information and the second time information, and obtain the audio and video synchronization detection result based on the difference.
[0074] In one embodiment, the terminal can identify the difference between the first and second time information of each target character in the text and obtain the audio-visual synchronization detection result based on the difference value.
[0075] In one embodiment, the terminal can preset a difference threshold, and by comparing the preset difference threshold with the difference value of each target character, the audio and video synchronization detection result can be obtained as either audio and video playback synchronized or audio and video playback asynchronous.
[0076] Specifically, the terminal compares the differences between the first and second time information, judges the differences according to preset rules, and obtains the audio and video synchronization detection result as either audio and video playback is synchronized or audio and video playback is not synchronized.
[0077] The aforementioned audio-video synchronization detection method acquires test audio and its corresponding text; the text includes at least one target character; and generates a synchronized audio-video file for synchronously playing the text and test audio. The first audio and first video data from the synchronized audio-video file are input to a multimedia system, triggering the system to process them according to a preset audio-video processing flow, outputting processed second audio and second video data. The processed second audio and second video data may be out of sync with the first audio and first video data in the synchronized audio-video file. Character recognition is performed on the video corresponding to the second video data to obtain the target character and the first time information of each target character's appearance in the corresponding video. Speech recognition is performed on the audio corresponding to the second audio data to obtain the target character and the second time information of each target character's pronunciation appearing in the corresponding audio. The difference between the first and second time information, i.e., the delay difference between the audio and video, is compared, and the audio-video synchronization detection result is obtained based on this difference. Therefore, audio-video synchronization can be automatically detected without subjective human evaluation, and the results can be quantitatively derived.
[0078] In one embodiment, obtaining the test audio and the corresponding text includes obtaining the test audio used for audio-video synchronization testing; and performing speech recognition on the test audio to obtain the corresponding text.
[0079] Specifically, the terminal acquires the test audio used for audio-video synchronization testing, performs speech recognition on the test audio, and obtains the text corresponding to the test audio. It can be understood that the text obtained by performing speech recognition on the test audio is the same as the text obtained by performing speech recognition on the audio corresponding to the second audio data in step S206.
[0080] In this embodiment, test audio for audio-video synchronization testing is acquired; speech recognition is performed on the test audio to obtain corresponding text, ensuring that the generated text is identical to the text obtained from speech recognition of the audio corresponding to the second audio data in step S206. This ensures that the text obtained from speech recognition and character recognition in step S206 is identical without manual annotation, thus providing accurate data for step S208 and improving the accuracy of the audio-video synchronization detection results.
[0081] In one embodiment, comparing the difference between the first time information and the second time information, and obtaining the audio-video synchronization detection result based on the difference, includes, for each target text in the text, confirming the difference value between the first time information and the second time information corresponding to each target text; and obtaining the audio-video synchronization detection result based on the difference value.
[0082] Specifically, for each target character in the text, the terminal confirms the first time information and the second time information obtained in step S206 for each target character, and calculates the difference value between the first time information and the second time information. The terminal obtains the audio and video synchronization detection result based on the difference value.
[0083] In one implementation, the terminal can obtain the audio and video synchronization detection result based on the magnitude of the difference value.
[0084] In another embodiment, the terminal can obtain the stability of audio and video synchronization based on the magnitude of the difference value.
[0085] In another implementation, the terminal can obtain the average audio-video synchronization latency based on the average of the differences in all target text.
[0086] In this embodiment, for each target character in the text, the terminal confirms the difference between the first time information and the second time information corresponding to each target character; based on the difference value, the audio-video synchronization detection result is obtained. In this way, the terminal does not only analyze and obtain the audio-video synchronization detection result at a single point in time, but also analyzes and obtains the audio-video synchronization detection result at various points in the total audio-video playback time, thereby improving the accuracy and comprehensiveness of the audio-video synchronization detection result.
[0087] In one embodiment, obtaining the audio-video synchronization detection result based on the difference value includes confirming that the audio-video synchronization detection result is synchronized if the difference value corresponding to each target character is less than a preset difference threshold, and confirming that the audio-video synchronization detection result is out of sync if the difference value corresponding to any target character is greater than the preset difference threshold.
[0088] Specifically, the terminal obtains a preset difference threshold, acquires the difference value corresponding to each target character, and performs a comparison. If the difference value corresponding to each target character is less than the preset difference threshold, the audio and video synchronization detection result is confirmed as audio and video playback synchronized. If the difference value corresponding to any target character is greater than the preset difference threshold, the audio and video synchronization detection result is confirmed as audio and video playback desynchronized.
[0089] In this embodiment, the audio-video synchronization detection result is confirmed as either asynchronous or synchronous by judging the magnitude of the difference value corresponding to each target text. Thus, the terminal does not only determine the magnitude of the difference value at a single point in time to obtain the audio-video synchronization detection result, but also determines the magnitude of the difference value at various points throughout the total audio-video playback time, thereby improving the accuracy and comprehensiveness of the audio-video synchronization detection result.
[0090] In one embodiment, the multimedia system includes a acquisition end and an output end. The process involves inputting first audio data and first video data from a synchronized audio / video file into the multimedia system, triggering the multimedia system to process the first audio data and first video data according to a preset audio / video processing flow, and outputting processed second audio data and second video data. This includes inputting the first audio data and first video data from the synchronized audio / video file into the acquisition end, triggering the acquisition end to simultaneously acquire the input first audio data and first video data respectively, encoding them respectively to obtain encoded audio data and video data; sending the encoded audio data and video data to the output end; and decoding the received encoded audio data and video data through the output end to output processed second audio data and second video data.
[0091] In one embodiment, the multimedia system includes a capture end and an output end. The capture end can capture first audio data and first video data from a synchronized audio and video file. The input end can output processed second audio data and second video data.
[0092] In one embodiment, the terminal plays a synchronized audio and video file and inputs the first audio data and first video data generated during the playback of the synchronized audio and video file to the acquisition end through a physical interface.
[0093] Specifically, the terminal inputs the first audio data and the first video data from the synchronized audio and video file to the acquisition end. The acquisition end simultaneously acquires the input first audio data and the first video data, encodes them separately, and obtains encoded audio data and video data. The acquisition end sends the encoded audio data and video data to the output end. The output end decodes the received encoded audio data and video data and outputs processed second audio data and second video data.
[0094] In this embodiment, the multimedia system includes a acquisition end and an output end, and processes audio and video. This processing may cause audio and video to be out of sync, so that the output second audio data and second video data need to be processed in step S206 to obtain the audio and video synchronization detection result.
[0095] In one embodiment, the multimedia system is a video conferencing system, the acquisition end is a first video conferencing terminal in the video conferencing system, and the output end is a second video conferencing terminal in the video conferencing system. The video conferencing system also includes a multipoint control unit. Sending encoded audio and video data to the output end includes sending the encoded audio and video data to the multipoint control unit through the first video conferencing terminal; decoding the received encoded audio and video data through the multipoint control unit, and combining the decoded audio data into a single target audio data, and combining the decoded video data into a single target video data; encoding the target audio and video data through the multipoint control unit to obtain intermediate audio and video data; and sending the intermediate audio and video data to the second video conferencing terminal through the multipoint control unit.
[0096] Video conferencing systems refer to two-way, multi-point, real-time audio and video interaction systems that are not limited by geographical location and are built on network communication. A multipoint control unit is a device used to combine at least one audio stream into one audio stream and at least one video stream into one video stream.
[0097] Specifically, the first video conferencing terminal transmits encoded audio and video data to the multipoint control unit via network transmission. The multipoint control unit decodes the received encoded audio and video data, combining the decoded audio streams into a single target audio stream and the decoded video streams into a single target video stream. The multipoint control unit then encodes the target audio and video data to obtain intermediate audio and video data. Finally, the multipoint control unit transmits the intermediate audio and video data to the second video conferencing terminal via network transmission.
[0098] In this embodiment, the multimedia system is a video conferencing system, the acquisition end is the first video conferencing terminal in the video conferencing system, the output end is the second video conferencing terminal in the video conferencing system, and the video conferencing system also includes a multipoint control unit. While the first video conferencing terminal, the second video conferencing terminal, and the multipoint control unit each process the audio and video data to meet the business requirements of the multimedia system, this may result in asynchronous audio and video output. Therefore, in step S206, the output second audio data and second video data need to be processed to obtain the audio and video synchronization detection result.
[0099] In one embodiment, such as Figure 4As shown, the second video conferencing terminal of the multimedia system is connected to the terminal for audio-video synchronization detection via a physical interface. The terminal for audio-video synchronization detection receives the second audio data and the second video data transmitted by the second video conferencing terminal. The character recognition module in the terminal for audio-video synchronization detection processes the video corresponding to the second video data to generate text 1, which includes first time information. The speech recognition module in the terminal for audio-video synchronization detection processes the audio corresponding to the second audio data to generate text 2, which includes second time information. Text 1 and Text 2 are input to the audio-video synchronization judgment module of the terminal for audio-video synchronization detection. The audio-video judgment module obtains the audio-video synchronization detection result based on the difference between text 1 and text 2, thereby realizing a quantitative evaluation of the synchronization effect of the multimedia system.
[0100] In one embodiment, the multimedia system includes multiple video conferencing terminals and a multipoint control unit. Specifically, such as Figure 5As shown, the terminal generates corresponding text from the test audio using speech recognition technology, and generates a synchronized audio and video file based on the text and test audio. The terminal plays the synchronized audio and video file and transmits the first audio data and first video data corresponding to the synchronized audio and video file to the multimedia system through a physical interface. The first video conferencing terminal in the multimedia system collects the first video data and second video data corresponding to the synchronized audio and video file, encodes them, and transmits them to the multipoint control unit via the network. The multipoint control unit decodes the received encoded audio and video data, combines the decoded audio data into a single target audio data, and combines the decoded video data into a single target video data, and transmits it to the second video conferencing terminal via the network. The second video conferencing terminal decodes the received encoded audio and video data, outputs processed second audio data and second video data, and inputs the processed second audio data and second video data to the terminal used for audio and video synchronization detection through a physical interface. The terminal used for audio and video synchronization detection performs character recognition on the video corresponding to the second video data to obtain text 1 containing a timeline, and performs speech recognition on the audio corresponding to the second audio data to obtain text 2 containing a timeline. The terminal for audio-video synchronization detection compares Text 1 and Text 2 to see if the time difference is within an acceptable range, thereby obtaining the audio-video synchronization detection result. Specifically, for each target character in the text, the terminal confirms the difference value between the first time information and the second time information corresponding to that target character. If the difference value corresponding to each target character is less than a preset difference threshold, the terminal confirms the audio-video synchronization detection result as audio-video playback synchronized. If the difference value corresponding to any target character is greater than the preset difference threshold, the terminal confirms the audio-video synchronization detection result as audio-video playback asynchronized. It can be understood that the terminal in this embodiment can also be a module within the terminal for audio-video synchronization detection.
[0101] It should be understood that although the steps in the flowcharts of some embodiments of this application are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple steps or multiple stages, which are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0102] Based on the same inventive concept, this application also provides an audio-video synchronization detection device for implementing the audio-video synchronization detection method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more audio-video synchronization detection device embodiments provided below can be found in the limitations of the audio-video synchronization detection method described above, and will not be repeated here.
[0103] In one embodiment, such as Figure 6 As shown, an audio / video synchronization detection device 600 is provided, including: a preparation module 602, an input module 604, a calculation module 606, and a determination module 608, wherein:
[0104] The preparation module 602 is used to obtain the test audio and the corresponding text text; the text text includes at least one target text; and to generate a synchronized audio and video file for synchronously playing the text text and the test audio.
[0105] The input module 604 is used to input the first audio data and the first video data corresponding to the synchronized audio and video file into the multimedia system, trigger the multimedia system to process the first audio data and the first video data according to the preset audio and video processing flow, and output the processed second audio data and the second video data.
[0106] The calculation module 606 is used to perform character recognition on the video corresponding to the second video data to obtain the target text and the first time information of the appearance of each target text in the corresponding video; and to perform speech recognition on the audio corresponding to the second audio data to obtain the target text and the second time information of the appearance of the pronunciation of each target text in the corresponding audio.
[0107] The determination module 608 is used to compare the difference between the first time information and the second time information, and obtain the audio and video synchronization detection result based on the difference.
[0108] In one embodiment, the preparation module 602 is further configured to acquire test audio for audio-video synchronization testing; and to perform speech recognition on the test audio to obtain the corresponding text.
[0109] In one embodiment, the calculation module 606 is further configured to, for each target character in the text, confirm the difference between the first time information and the second time information corresponding to each target character;
[0110] Based on the difference values, the audio and video synchronization detection results are obtained.
[0111] In one embodiment, the calculation module 606 is further configured to confirm that the audio and video synchronization detection result is audio and video playback synchronization if the difference value corresponding to each target character is less than a preset difference threshold; and to confirm that the audio and video synchronization detection result is audio and video playback asynchrony if the difference value corresponding to any target character is greater than the preset difference threshold.
[0112] In one embodiment, the multimedia system includes a acquisition end and an output end; the input module 604 is further configured to input the first audio data and the first video data from the synchronized audio and video file into the acquisition end, triggering the acquisition end to simultaneously acquire the input first audio data and the first video data respectively, encode them respectively, and obtain encoded audio data and video data; send the encoded audio data and video data to the output end; and decode the received encoded audio data and video data through the output end to output processed second audio data and second video data.
[0113] In one embodiment, the multimedia system is a video conferencing system, the acquisition end is a first video conferencing terminal in the video conferencing system, and the output end is a second video conferencing terminal in the video conferencing system. The video conferencing system also includes a multipoint control unit. The input module 604 is further configured to send the encoded audio data and video data to the multipoint control unit through the first video conferencing terminal; the multipoint control unit decodes the received encoded audio data and video data, and combines the decoded audio data into a single target audio data, and combines the decoded video data into a single target video data; the multipoint control unit encodes the target audio data and target video data to obtain intermediate audio data and intermediate video data; and the multipoint control unit sends the intermediate audio data and intermediate video data to the second video conferencing terminal.
[0114] The aforementioned audio-video synchronization detection device acquires test audio and corresponding text text, where the text text includes at least one target character. It generates a synchronized audio-video file for synchronously playing the text text and test audio. The first audio and first video data from the synchronized audio-video file are input to a multimedia system, triggering the system to process them according to a preset audio-video processing flow. The system outputs processed second audio and second video data, which may be asynchronous with respect to the first audio and first video data in the synchronized audio-video file. Character recognition is performed on the video corresponding to the second video data to obtain the target character and the first time information of each target character's appearance in the corresponding video. Speech recognition is performed on the audio corresponding to the second audio data to obtain the target character and the second time information of each target character's pronunciation appearing in the corresponding audio. The difference between the first and second time information, i.e., the delay difference between the audio and video, is compared to obtain the audio-video synchronization detection result. Therefore, audio-video synchronization can be automatically detected without subjective human evaluation, and the detection result can be quantitatively obtained.
[0115] For specific limitations regarding the aforementioned audio-video synchronization detection device, please refer to the limitations of the aforementioned audio-video synchronization detection method above, which will not be repeated here. Each module in the aforementioned audio-video synchronization detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.
[0116] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an audio-visual synchronization detection method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0117] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0118] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0119] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0120] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0123] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for detecting audio and video synchronization, characterized in that, The method includes: Acquire test audio and perform speech recognition on the test audio to obtain corresponding text; the text includes at least one target character. A white-background, black-text video is generated based on the text. A synchronized audio-video file is generated based on the white-background, black-text video and the test audio to play the text and the test audio synchronously. The video in the synchronized audio-video file displays the target text in the text, and the appearance time of the target text and the appearance time of the corresponding pronunciation of the target text in the audio of the synchronized audio-video file are synchronized. Play the synchronized audio and video file, input the first audio data and the first video data of the synchronized audio and video file into the acquisition terminal of the multimedia system through the physical interface, trigger the acquisition terminal to simultaneously acquire the input first audio data and the first video data respectively, encode them respectively to obtain encoded audio data and video data, send the encoded audio data and video data to the output terminal of the multimedia system, decode the received encoded audio data and video data through the output terminal, and output the processed second audio data and second video data. The second audio data and the second video data are the data to be detected for audio and video synchronization. Character recognition is performed on the video corresponding to the second video data. When the target text is recognized, the corresponding time point is recorded to obtain the target text and the first time information of the appearance of each target text in the corresponding video. Speech recognition is performed on the audio corresponding to the second audio data. When the text corresponding to the pronunciation of the audio is identified as the target text, the corresponding time point is recorded to obtain the target text and the second time information of the appearance of the pronunciation of each target text in the corresponding audio. For each target character in the text, the difference between the first time information and the corresponding second time information of the target character is confirmed. Based on the average value of the difference value corresponding to each target character, the average audio-video synchronization delay is obtained. Based on the average audio-video synchronization delay, the audio-video synchronization detection result is obtained. The audio-video synchronization detection result is one of audio-video playback synchronization and audio-video playback asynchrony.
2. The method according to claim 1, characterized in that, The acquisition of the test audio and the corresponding text text includes: Obtain test audio for audio-video synchronization testing; The test audio is subjected to speech recognition to obtain the corresponding text.
3. The method according to claim 1, characterized in that, The audio-video synchronization detection results obtained based on the difference values include: If the difference value corresponding to each target character is less than the preset difference threshold, then the audio and video synchronization detection result is confirmed as audio and video playback synchronization. If the difference value corresponding to any target text is greater than the preset difference threshold, then the audio and video synchronization detection result is confirmed as audio and video playback being out of sync.
4. The method according to claim 1, characterized in that, The multimedia system is a video conferencing system, the acquisition end is a first video conferencing terminal in the video conferencing system; the output end is a second video conferencing terminal in the video conferencing system; the video conferencing system also includes a multipoint control unit; sending the encoded audio data and video data to the output end of the multimedia system includes: The encoded audio and video data are sent to the multipoint control unit via the first video conferencing terminal. The multi-point control unit decodes the received encoded audio and video data, and combines the decoded audio data into a single target audio data, and combines the decoded video data into a single target video data. The target audio data and target video data are encoded by the multi-point control unit to obtain intermediate audio data and intermediate video data; The intermediate audio data and intermediate video data are sent to the second video conferencing terminal through the multi-point control unit.
5. An audio-visual synchronization detection device, characterized in that, The device includes: A preparation module is used to acquire test audio and perform speech recognition on the test audio to obtain corresponding text; the text includes at least one target character; a white-background black-text video displaying the text is generated based on the text; a synchronized audio-video file is generated based on the white-background black-text video and the test audio to synchronize the playback of the text and the test audio; the video in the synchronized audio-video file displays the target character in the text, and the appearance time of the target character is synchronized with the appearance time of the pronunciation of the target character in the audio of the synchronized audio-video file; The input module is used to play the synchronized audio and video file. It inputs the first audio data and the first video data of the synchronized audio and video file into the acquisition terminal of the multimedia system through a physical interface, triggering the acquisition terminal to simultaneously acquire the input first audio data and the first video data, encode them respectively, and obtain encoded audio data and video data. The encoded audio data and video data are sent to the output terminal of the multimedia system. The output terminal decodes the received encoded audio data and video data and outputs the processed second audio data and second video data. The calculation module is used to perform character recognition on the video corresponding to the second video data, and when the target text is recognized, record the corresponding time point to obtain the target text and the first time information of the appearance of each target text in the corresponding video; and to perform speech recognition on the audio corresponding to the second audio data, and when the text corresponding to the pronunciation of the audio is recognized as the target text, record the corresponding time point to obtain the target text and the second time information of the appearance of the pronunciation of each target text in the corresponding audio. The determination module is used to determine the difference between the first time information and the corresponding second time information of each target character in the text, obtain the average audio-video synchronization delay based on the average difference value corresponding to each target character, and obtain the audio-video synchronization detection result based on the average audio-video synchronization delay. The audio-video synchronization detection result is one of audio-video playback synchronization and audio-video playback asynchrony.
6. The apparatus according to claim 5, characterized in that, The multimedia system is a video conferencing system; the acquisition end is the first video conferencing terminal in the video conferencing system; the output end is the second video conferencing terminal in the video conferencing system; the video conferencing system also includes a multipoint control unit. The input module is further configured to send the encoded audio data and video data to the multipoint control unit via the first video conferencing terminal; decode the received encoded audio data and video data via the multipoint control unit, and combine the decoded audio data into a single target audio data and the decoded video data into a single target video data; encode the target audio data and target video data via the multipoint control unit to obtain intermediate audio data and intermediate video data; and send the intermediate audio data and intermediate video data to the second video conferencing terminal via the multipoint control unit.
7. The apparatus according to claim 5, characterized in that, The preparation module is also used to acquire test audio for audio-video synchronization testing; and to perform speech recognition on the test audio to obtain the corresponding text.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method, device and system for testing audio and video synchronization in video instant messaging
CN105898505A
Method and system for video conference using mobile terminals
CN108933914A
Video playing quality detection method and device
CN112511818A
Closed caption signal processing apparatus and method
US20050060145A1