Audio and video lag detection method and device, electronic equipment, storage medium

By employing audio-video separation and image synthesis techniques, the low accuracy problem caused by neglecting the influence of audio in existing technologies has been solved, achieving higher accuracy in audio-video stutter detection.

CN119517099BActive Publication Date: 2025-11-18CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411616858.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-11-18
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing audio and video stuttering detection technologies ignore the impact of audio on stuttering detection, resulting in low detection accuracy.

Method used

By separating the target audio and video, processing the video and audio separately, extracting image frames and audio sampling points according to a preset period, generating a spectrogram, and synthesizing the images at the same time, the stuttering detection result is determined by the difference between the target test images.

Benefits of technology

It improves the accuracy of audio and video stutter detection, ensuring that the detection results are based on audio and video information at different times, resulting in richer information and improved detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119517099B_ABST
    Figure CN119517099B_ABST
Patent Text Reader

Abstract

The application provides a kind of audio and video lag detection method and device, electronic equipment, storage medium, related to artificial intelligence technical field, applicable to the field of financial technology.The method comprises: obtaining target audio and video;Audio and video separation is carried out on the target audio and video, to obtain target video and target audio;According to the preset period, frame extraction is carried out on the target video, to obtain image frame;According to the preset period, audio sampling is carried out on the target audio, to obtain audio sampling point;Spectrum conversion is carried out on the audio sampling point, to obtain spectrum diagram;When the extraction time of image frame is the same as the sampling time of audio sampling point, image synthesis is carried out on the image frame and the spectrum diagram, to obtain target test image;According to the difference between any two target test images, the lag detection result of the target audio and video is determined;Wherein, the lag detection result is used to represent that the target audio and video exists lag or does not exist lag.The application can improve the accuracy of audio and video lag detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and is applicable to the financial technology field, particularly to an audio / video stuttering detection method and device, electronic device, and storage medium. Background Technology

[0002] Audio and video refer to content created through audio and video recording. In the fintech field, during financial transactions (such as when banks or insurance companies introduce financial products), it is necessary to record the entire process to obtain audio and video, which facilitates future review and management. However, errors during the recording process may cause audio and video stuttering, thus requiring stuttering detection.

[0003] Existing stuttering detection techniques typically rely on differences between image frames in a video for stuttering detection. However, this approach ignores the impact of audio on audio-visual stuttering, resulting in low accuracy in detecting stuttering in both audio and video. Summary of the Invention

[0004] The main objective of this application is to provide a method, apparatus, electronic device, and storage medium for detecting audio and video stuttering, which can improve the accuracy of audio and video stuttering detection.

[0005] To achieve the above objectives, a first aspect of this application proposes an audio / video stuttering detection method, the method comprising:

[0006] Acquire target audio and video;

[0007] The target audio and video are separated to obtain the target video and the target audio.

[0008] The target video is frame-by-frame extracted according to a preset period to obtain image frames;

[0009] The target audio is sampled according to the preset period to obtain audio sampling points;

[0010] The audio sampling points are subjected to spectral conversion to obtain a spectrogram;

[0011] When the extraction time of the image frame is the same as the sampling time of the audio sampling point, the image frame and the spectrogram are combined to obtain the target test image.

[0012] Based on the difference between any two target test images, the stuttering detection result of the target audio and video is determined; wherein, the stuttering detection result is used to characterize whether the target audio and video has stuttering or not.

[0013] In some embodiments, determining the stuttering detection result of the target audio / video based on the difference between any two target test maps includes:

[0014] Obtain the target time for each target test image; wherein the target time, extraction time, and sampling time are the same for the same target test image.

[0015] Based on the target time, a first target test map and a second target test map are extracted from each of the target test maps; wherein, the target time of the second target test map is after the target time of the first target test map;

[0016] Image difference calculation is performed based on the first target test image and the second target test image to obtain the first target difference image;

[0017] The first test pixel sum is obtained by summing the pixels of each pixel in the first target difference map;

[0018] The stuttering detection result is determined based on the first test pixel.

[0019] In some embodiments, determining the stuttering detection result of the target audio / video based on the difference between any two target test maps further includes:

[0020] A third target test map is extracted from each of the target test maps based on the target time corresponding to the second target test map; wherein the target time of the third target test map is after the target time of the second target test map;

[0021] The image difference is calculated based on the first target test image and the third target test image to obtain the second target difference image;

[0022] The second test pixel sum is obtained by summing the pixels of each pixel in the second target difference map;

[0023] Image difference calculation is performed based on the second target test image and the third target test image to obtain the third target difference image;

[0024] The third test pixel sum is obtained by summing the pixels of each pixel in the third target difference map;

[0025] Pixel stuttering is detected based on the first test pixel sum, the second test pixel sum, and the third test pixel sum, and the stuttering detection result is obtained.

[0026] In some embodiments, the step of performing pixel stuttering detection based on the first test pixel sum, the second test pixel sum, and the third test pixel sum to obtain the stuttering detection result includes:

[0027] Construct a stuttering graph count marker, wherein the stuttering graph count marker is 0;

[0028] The first test pixel sum, the second test pixel sum, and the third test pixel sum are used alternately as the target test pixel sum;

[0029] If the target test pixel sum is less than or equal to the preset pixel sum threshold, then the number of stuttering images is incremented by 1;

[0030] Based on the number of stutters marked in the stuttering graph, a number of stutters are detected to obtain the video stuttering detection result.

[0031] In some embodiments, when the extraction time of the image frame is the same as the sampling time of the audio sampling point, image synthesis is performed on the image frame and the spectrogram to obtain the target test image, including:

[0032] The image frame is scaled according to a preset target scale, and the spectrogram is scaled according to the target scale, so that the scale of the image frame is the same as the scale of the spectrogram.

[0033] When the extraction time of the image frame is the same as the sampling time of the audio sampling point, the spectrogram is stitched to the left side of the image frame to obtain the target test image.

[0034] In some embodiments, after determining the stuttering detection result of the target audio / video based on the difference between any two target test maps, the method further includes:

[0035] If the stuttering detection result of the target audio and video indicates that the target audio and video does not stutter, then face detection is performed on the first target test image to obtain the location information of the first face region;

[0036] Based on the location information of the first face region, the first target difference map is extracted to obtain the face region of the first difference map.

[0037] The pixel sum of the first face region is obtained by summing the pixels of each pixel in the face region of the first difference image.

[0038] The stuttering detection result of the target audio and video is updated based on the pixel sum of the first face region.

[0039] In some embodiments, after performing audio-video separation on the target audio and video to obtain the target video and target audio, the method further includes:

[0040] Obtain the duration of the target video to get the target video duration;

[0041] Obtain the duration of the target audio;

[0042] The absolute duration difference is obtained by calculating the difference between the target video duration and the target audio duration.

[0043] If the absolute duration difference is greater than a preset duration difference threshold, a target prompt message is generated; wherein, the target prompt message is used to prompt the re-recording of the target audio and video.

[0044] To achieve the above objectives, a second aspect of this application provides an audio / video stuttering detection device, the device comprising:

[0045] The data acquisition module is used to acquire the target audio and video;

[0046] The data separation module is used to separate the target audio and video to obtain the target video and target audio.

[0047] The image extraction module is used to extract frames from the target video according to a preset period to obtain image frames;

[0048] An audio sampling module is used to sample the target audio according to the preset period to obtain audio sampling points;

[0049] The spectrum conversion module is used to perform spectrum conversion on the audio sampling points to obtain a spectrum diagram;

[0050] An image synthesis module is used to synthesize the image frame and the spectrogram to obtain a target test image when the extraction time of the image frame is the same as the sampling time of the audio sampling point.

[0051] The result determination module is used to determine the stuttering detection result of the target audio and video based on the difference between any two target test images; wherein, the stuttering detection result is used to characterize whether the target audio and video has stuttering or not.

[0052] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the audio / video stuttering detection method described in the first aspect.

[0053] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the audio / video stuttering detection method described in the first aspect.

[0054] The audio / video stuttering detection method, apparatus, electronic device, and storage medium proposed in this application address the problem of low accuracy in related technologies that rely solely on image information in video for video stuttering detection. In this application, when the target audio / video is acquired, it is first separated into target video and target audio, allowing for independent processing of the video and audio. Further, frames are extracted from the target video according to a preset period to obtain image frames. Thus, one image frame is extracted every preset period, and the extraction times of any two image frames are different. Further, in addition to extracting image frames according to the preset period, audio sampling is also performed on the target audio according to the preset period to obtain audio sampling points, thereby obtaining a spectrogram. Thus, one audio sampling point is extracted every preset period, and the sampling times of any two audio sampling points are different. Since both image frames and audio sampling points are acquired based on the preset period, each extraction time has one and only one sampling time identical to the extraction time. Furthermore, when the extraction time of the image frame coincides with the sampling time of the audio sampling point, the image frame and the spectrogram are synthesized to obtain the target test map. This target test map represents the audio and video information corresponding to the target audio and video at a specific moment. Finally, the stuttering detection result of the target audio and video is determined based on the difference between any two target test maps. By first separating the target audio and video, and then generating the target test map through a series of operations, and because the target test map represents the audio and video information corresponding to the target audio and video at a specific moment, determining the stuttering detection result based on the target test map is equivalent to determining the stuttering detection result based on the audio and video information of the target audio and video at different moments. This provides richer information and improves the accuracy of stuttering detection.

[0055] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0056] Figure 1 This is a flowchart of the audio / video stuttering detection method provided in the embodiments of this application;

[0057] Figure 2 yes Figure 1 The flowchart for step 106 in the document;

[0058] Figure 3 yes Figure 1 A flowchart of step 107 in the process;

[0059] Figure 4 yes Figure 1 Another flowchart for step 107 in the process;

[0060] Figure 5 yes Figure 4The flowchart for step 410 in the middle;

[0061] Figure 6 This is a block diagram of the module structure of the audio / video stuttering detection device provided in the embodiments of this application;

[0062] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0064] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0066] First, let's analyze some of the terms used in this application:

[0067] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0068] Audio and video recording, also known as dual-recording video, refers to the process of banks or insurance companies recording audio and video of the entire process when introducing financial products. Dual-recording video stuttering detection refers to the manual or intelligent detection of stuttering in the audio and video recordings. For example, the data generated from recording and videotaping of on-site personnel handling business is captured by dual-recording equipment and stored in a cloud storage system. This is used to monitor the operational standards and professional skills of on-site personnel, and to preserve evidence of business-related operations and behaviors. After the recorded audio and video are uploaded to the business quality inspection platform, they need to be inspected to ensure they meet the specified requirements. If they do, the recorded audio and video are saved; otherwise, they need to be re-recorded. Complete and clear video facilitates future review and management, and helps resolve future disputes; therefore, it is necessary to check the completeness of the video, whether there are any stutters, and to reduce the risk of missing key parts of the video.

[0069] This application provides a method, apparatus, electronic device, and storage medium for detecting audio and video stuttering. The core of this application lies in simultaneously considering the impact of audio and video on stuttering, thereby improving the accuracy of stuttering detection.

[0070] The audio / video stuttering detection method provided in this application can be applied to terminals and servers, or it can be software running on the server. The server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or it can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the audio / video stuttering detection method, etc., but is not limited to the above forms.

[0071] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0072] This application provides an audio / video stuttering detection method, an audio / video stuttering detection device, an electronic device, and a computer-readable storage medium. The specific details are illustrated in the following embodiments. First, the audio / video stuttering detection method in this application is described.

[0073] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0074] Reference Figure 1 , Figure 1 This is an optional flowchart of the audio / video stuttering detection method provided in the embodiments of this application, which may include, but is not limited to, steps 101 to 107.

[0075] Step 101: Obtain the target audio and video;

[0076] Step 102: Perform audio-video separation on the target audio and video to obtain the target video and target audio;

[0077] Step 103: Extract frames from the target video according to a preset period to obtain image frames;

[0078] Step 104: Sample the target audio according to a preset period to obtain audio sampling points;

[0079] Step 105: Perform spectral conversion on the audio sampling points to obtain a spectrogram;

[0080] Step 106: When the extraction time of the image frame is the same as the sampling time of the audio sampling point, the image frame and the spectrogram are synthesized to obtain the target test image.

[0081] Step 107: Determine the stuttering detection result of the target audio and video based on the difference between any two target test images; wherein, the stuttering detection result is used to characterize whether the target audio and video has stuttering or not.

[0082] Steps 101 to 107 shown in this embodiment address the problem of low accuracy in video stutter detection relying solely on image information in the video in related technologies. In this embodiment, when the target audio and video are acquired, they are first separated to obtain the target video and target audio, allowing for independent processing of the video and audio. Further, frame extraction is performed on the target video according to a preset period to obtain image frames. Thus, one image frame is extracted every preset period, and the extraction times of any two image frames are different. Further, in addition to extracting image frames according to the preset period, audio sampling is also performed on the target audio according to the preset period to obtain audio sampling points, thereby obtaining a spectrogram. Thus, one audio sampling point is extracted every preset period, and the sampling times of any two audio sampling points are different. Since both image frames and audio sampling points are acquired based on the preset period, each extraction time has one and only one sampling time identical to the extraction time. Furthermore, when the extraction time of the image frame coincides with the sampling time of the audio sampling point, the image frame and the spectrogram are synthesized to obtain the target test map. This target test map represents the audio and video information corresponding to the target audio and video at a specific moment. Finally, the stuttering detection result of the target audio and video is determined based on the difference between any two target test maps. By first separating the target audio and video, and then generating the target test map through a series of operations, and because the target test map represents the audio and video information corresponding to the target audio and video at a specific moment, determining the stuttering detection result based on the target test map is equivalent to determining the stuttering detection result based on the audio and video information of the target audio and video at different moments. This provides richer information and improves the accuracy of stuttering detection.

[0083] In step 101 of some embodiments, the target audio / video is acquired. The target audio / video refers to the audio / video to be tested for stuttering. For example, data information generated from recording audio and video of on-site personnel's business transactions can be captured by a terminal (such as a mobile phone or dual-recording device) to obtain the target audio / video. When the above-described audio / video stuttering detection method is applied to the server side, the server receives the target audio / video from the terminal (such as a mobile phone or dual-recording device).

[0084] In step 102 of some embodiments, the target audio and video are separated to obtain the target video and target audio. Separation tools (such as FFmpeg, Audacity, Adobe Premiere Pro, Final Cut Pro) can be used to perform the audio and video separation.

[0085] In one embodiment, after step 102, the audio / video stuttering detection method provided in this embodiment may further include:

[0086] Obtain the duration of the target video;

[0087] Obtain the duration of the target audio;

[0088] The absolute duration difference is obtained by calculating the difference between the target video duration and the target audio duration.

[0089] If the absolute duration difference is greater than the preset duration difference threshold, a target prompt message is generated; the target prompt message is used to prompt the re-recording of the target audio and video.

[0090] For example, the target audio / video is parsed into target audio and target video, resulting in the target audio duration of t1 seconds and the target video duration of t2 seconds. The absolute duration difference is then calculated as abs(t2–t1). The threshold for this duration difference is 30 seconds. If abs(t2–t1) is greater than 30 seconds, the target audio / video is considered incomplete, and a target prompt message needs to be generated.

[0091] The advantage of the above embodiments is that they can promptly prompt the user to re-record the target audio and video during the execution of stuttering detection tasks, thereby improving the feasibility of parallel tasks.

[0092] In one embodiment, step 102 may include: demultiplexing the target audio and video using a demultiplexer to obtain an audio packet queue and a video packet queue; decoding the audio packet queue using an audio decoder to obtain the target audio; and decoding the video packet queue using a video decoder to obtain the target video.

[0093] In step 103 of some embodiments, the target video is frame extracted according to a preset period to obtain image frames. The preset period is, for example, 1 second, 2 seconds, etc. For example, assuming the preset period is 1 second, step 103 includes: extracting frames from the target video every 1 second to obtain image frames for the 1st second, the 2nd second, the 3rd second, the 4th second, etc.

[0094] In step 104 of some embodiments, the target audio is sampled according to a preset period to obtain audio sampling points. For example, assuming the preset period is 1 second, step 104 includes: sampling the target audio every 1 second to obtain audio sampling points for the 1st second, the 2nd second, the 3rd second, the 4th second, etc.

[0095] It should be noted that the preset period for frame extraction is the same as the preset period for audio sampling.

[0096] In step 105 of some embodiments, the audio sampling points are subjected to spectral transformation to obtain a spectrogram. For example, for each audio sampling point, the audio sampling point is sampled twice using a preset sampling rate (e.g., 16kHz) to obtain audio secondary sampling points; a Fast Fourier Transform (FFT) is performed on the audio secondary sampling points to obtain a spectrogram. For example, the mel spectrogram is obtained according to the mel spectrum transform or the MFCC transform formula. For the audio sampling points at the 1st second, 2nd second, 3rd second, and 4th second in the example above, the spectrograms for the 1st second, 2nd second, 3rd second, and 4th second can be obtained.

[0097] In step 106 of some embodiments, when the extraction time of the image frame is the same as the sampling time of the audio sampling point, the image frame and the spectrogram are synthesized to obtain the target test image. For example, the image frame of the 1st second and the spectrogram of the 1st second are synthesized to obtain the target test image of the 1st second, the image frame of the 2nd second and the spectrogram of the 2nd second are synthesized to obtain the target test image of the 2nd second, the image frame of the 3rd second and the spectrogram of the 3rd second are synthesized to obtain the target test image of the 3rd second, the image frame of the 4th second and the spectrogram of the 4th second are synthesized to obtain the target test image of the 4th second, and so on.

[0098] In one embodiment, reference is made to Figure 2 Step 106 may include:

[0099] Step 201: Scale the image frame according to the preset target scale, and scale the spectrogram according to the target scale, so that the scale of the image frame is the same as the scale of the spectrogram.

[0100] Step 202: When the extraction time of the image frame is the same as the sampling time of the audio sampling point, the spectrogram is stitched onto the left side of the image frame to obtain the target test image.

[0101] In step 201, assuming the target scale is 320*240, the scale of the image frame is scaled to 320*240, and the scale of the spectrogram is scaled to 320*240.

[0102] In step 202, for example, the spectrogram of the first second is stitched to the left side of the image frame of the first second to obtain the target test image of the first second with a scale of 640*240; the spectrogram of the second second is stitched to the left side of the image frame of the second second to obtain the target test image of the second second with a scale of 640*240; the spectrogram of the third second is stitched to the left side of the image frame of the third second to obtain the target test image of the third second with a scale of 640*240; the spectrogram of the fourth second is stitched to the left side of the image frame of the fourth second to obtain the target test image of the fourth second with a scale of 640*240; and so on.

[0103] The advantage of the embodiments of steps 201 to 204 above is that they can standardize the scale of the target test map, which facilitates the subsequent determination of stuttering detection results based on the differences between the target test maps, thereby helping to improve stuttering detection efficiency.

[0104] In step 107 of some embodiments, the stuttering detection result of the target audio / video is determined based on the difference between any two target test images. The stuttering detection result is used to characterize whether the target audio / video has stuttering or not.

[0105] In one embodiment, reference is made to Figure 3 Step 107 may include:

[0106] Step 301: Obtain the target time for each target test image; wherein the target time, extraction time, and sampling time are the same for the same target test image.

[0107] Step 302: Extract the first target test map and the second target test map from each target test map according to the target time; wherein, the target time of the second target test map is after the target time of the first target test map;

[0108] Step 303: Calculate the image difference based on the first target test image and the second target test image to obtain the first target difference image;

[0109] Step 304: Sum the pixels in the first target difference map to obtain the first test pixel sum;

[0110] Step 305: Determine the stuttering detection result based on the first test pixel.

[0111] In step 301, referring to the example above, the target test map includes the target test map at the 1st second, the target test map at the 2nd second, the target test map at the 3rd second, and the target test map at the 4th second. Therefore, the target time is the 1st second, the 2nd second, the 3rd second, and the 4th second in sequence.

[0112] In step 302, for example, if the first target test map is the target test map at the 1st second, then the second target test map is the target test map at the 2nd second. Or, for example, if the first target test map is the target test map at the 2nd second, then the second target test map is the target test map at the 3rd second.

[0113] In step 303, specifically, the pixel difference between each first pixel in the first target test image and the corresponding second pixel in the second target test image is calculated to obtain a first target difference image. The second pixel corresponding to the first pixel specifically refers to the second pixel at the same position as the first pixel.

[0114] In step 304, the pixels of all points in the first target difference image are summed to obtain the first test pixel sum. For example, if the first target difference image has 3 pixels, and the pixel values ​​of the pixels are 0.1 pixels, 0.1 pixels, and 0.2 pixels respectively, then the first test pixel sum is 0.4 pixels.

[0115] Step 305: Determine the stuttering detection result based on the sum of the first test pixels. For example, if the sum of the first test pixels is less than or equal to a preset threshold (e.g., 0 pixels or 0.3 pixels), a stuttering detection result is generated to indicate that the target audio / video has stuttering; if the sum of the first test pixels is greater than the preset threshold (e.g., 0 pixels or 0.3 pixels), a stuttering detection result is generated to indicate that the target audio / video does not have stuttering. Alternatively, a stuttering detection result can be obtained by looking up a preset pixel stuttering mapping table based on the sum of the first test pixels. Specifically, firstly, the pixel stuttering mapping table is called, whereby the pixel stuttering mapping table indicates the stuttering detection result corresponding to each pixel interval. Next, the sum of the first test pixels is compared with each pixel interval in the pixel stuttering mapping table, and the detection result corresponding to the pixel interval containing the sum of the first test pixels is taken as the stuttering detection result.

[0116] The advantage of the embodiments of steps 301 to 305 described above is that the stuttering detection result can be determined based on the difference between two adjacent target test maps, and the operation is simple and flexible.

[0117] In one embodiment, reference is made to Figure 4 Step 107 may include:

[0118] Step 401: Obtain the target time for each target test image; wherein, the target time, extraction time, and sampling time are the same for the same target test image;

[0119] Step 402: Extract the first target test map and the second target test map from each target test map according to the target time; wherein the target time of the second target test map is after the target time of the first target test map;

[0120] Step 403: Calculate the image difference based on the first target test image and the second target test image to obtain the first target difference image;

[0121] Step 404: Sum the pixels in the first target difference map to obtain the first test pixel sum;

[0122] Step 405: Extract the third target test map from each target test map according to the target time corresponding to the second target test map; wherein, the target time of the third target test map is after the target time of the second target test map;

[0123] Step 406: Calculate the image difference based on the first target test image and the third target test image to obtain the second target difference image;

[0124] Step 407: Sum the pixels in the second target difference map to obtain the second test pixel sum;

[0125] Step 408: Calculate the image difference based on the second target test image and the third target test image to obtain the third target difference image;

[0126] Step 409: Sum the pixels in the third target difference map to obtain the third test pixel sum;

[0127] Step 410: Perform pixel stuttering detection based on the first test pixel sum, the second test pixel sum, and the third test pixel sum to obtain stuttering detection results.

[0128] Steps 401 to 404 are basically the same as steps 301 to 304 above, and will not be repeated here.

[0129] In step 405, for example, if the second target test map is the target test map at the 2nd second, then the third target test map is the target test map at the 3rd second. Or, for example, if the second target test map is the target test map at the 3rd second, then the third target test map is the target test map at the 4th second.

[0130] In step 406, specifically, the pixel difference between each first pixel in the first target test image and the corresponding third pixel in the third target test image is calculated to obtain a second target difference image. The corresponding third pixel specifically refers to the third pixel at the same position as the first pixel.

[0131] In step 407, the pixels of all points in the second target difference image are summed to obtain the second test pixel sum. For example, if the second target difference image has 3 pixels, and the pixel values ​​of the pixels are 0.1 pixels, 0.1 pixels, and 0.1 pixels respectively, then the second test pixel sum is 0.3 pixels.

[0132] In step 408, specifically, the pixel difference between each second pixel in the second target test image and the corresponding third pixel in the third target test image is calculated to obtain a third target difference image. The third pixel corresponding to the second pixel specifically refers to the third pixel at the same position as the second pixel.

[0133] In step 409, the sum of all pixels in the third target difference image is calculated to obtain the third test pixel sum. For example, if the third target difference image has 3 pixels, and the pixel values ​​of the pixels are 0.05 pixels, 0.05 pixels, and 0.1 pixels respectively, then the third test pixel sum is 0.2 pixels.

[0134] Step 410: Perform pixel stuttering detection based on the first test pixel sum, the second test pixel sum, and the third test pixel sum to obtain stuttering detection results. For example, if the first test pixel sum, the second test pixel sum, and the third test pixel sum are all less than or equal to a preset threshold (e.g., 0 or 0.5), a stuttering detection result is generated to characterize that the target audio / video has stuttering; if the first test pixel sum, the second test pixel sum, and the third test pixel sum are all greater than the preset threshold (e.g., 0 or 0.5), a stuttering detection result is generated to characterize that the target audio / video does not have stuttering.

[0135] In one embodiment, reference is made to Figure 5 Step 410 may include:

[0136] Step 501: Construct a stuttering graph count marker, with the stuttering graph count marker set to 0;

[0137] Step 502: Take the first test pixel sum, the second test pixel sum, and the third test pixel sum as the target test pixel sum in turn;

[0138] Step 503: If the target test pixel sum is less than or equal to the preset pixel sum threshold, then increment the number of stuttering images by 1.

[0139] Step 504: Perform quantity lag detection based on the lag graph quantity markers to obtain lag detection results.

[0140] In one example, if the number of stuttering graph markers is greater than a preset threshold, a stuttering detection result is generated to characterize that the target audio / video has stuttering; if the number of stuttering graph markers is less than or equal to the preset threshold, a stuttering detection result is generated to characterize that the target audio / video does not have stuttering.

[0141] The advantage of the embodiments of steps 501 to 504 described above is that the stuttering detection result can be dynamically determined based on the number of stuttering graph markers, resulting in high accuracy and flexibility.

[0142] In one example, assume there are 5 target test images, denoted as f1, f2, f3, f4, and f5. Then calculate the 5-second image difference, including: f 12 =abs(f1–f2),f 13 =abs(f1–f3),f 14 =abs(f1–f4),f 15=abs(f1-f5),f 23 =abs(f2–f3),f 24 =abs(f2–f4),f 25 =abs(f2–f5),f 34 =abs(f3–f4),f 35 =abs(f3–f5),f 45 =abs(f4–f5), resulting in 4+3+2+1=10 target difference maps, and the pixel sum of each target difference map is calculated. Judgment condition 1: If the pixel sum of each target difference map is equal to 0, then it is considered to be laggy. Judgment condition 2: If the pixel sum is less than a pixel sum threshold (e.g., 50), the number of target difference maps that meet the conditions is N. If the proportion of N to the total number of target difference maps is greater than a preset proportion (e.g., 0.5), then it is considered that there may be laggy behavior, requiring further confirmation; if the proportion of N to the total number of target difference maps is less than or equal to the preset proportion...

[0143] In one embodiment, after step 107, the audio / video stuttering detection method provided in this embodiment may further include:

[0144] If the stuttering detection result of the target audio and video indicates that there is no stuttering in the target audio and video, then face detection is performed on the first target test image to obtain the location information of the first face region;

[0145] Based on the location information of the first face region, the first target difference map is used to extract the face region of the first difference map.

[0146] The pixel sum of the first face region is obtained by summing the pixels of each pixel in the face region of the first difference image.

[0147] Based on the pixel sum of the first face region, update the stuttering detection results of the target audio and video.

[0148] The advantage of the above embodiments is that by using the face region in the target test image for stuttering detection, the accuracy of stuttering detection is further improved.

[0149] Specifically, face detection can be performed on the first target test image using existing face detection models. For example, performing Retinaface face detection on the first target test image yields the face bounding box and the coordinates of five landmark points P1, P2, P3, P4, and P5, representing the left eye, right eye, nose, left corner of mouth, and right corner of mouth, respectively. The eye region is calculated: the centers of the left and right eyes are P1 and P2, respectively. The side length L = (P2.x – P1.x) / 3 is calculated, so the left eye region R_lef = cvRect(P1.xL / 2, P1.yL / 2, L, L). The domain R_right = cvRect(P2.xL / 2,P2.yL / 2,L,L); finally, the mouth region is calculated: side length L1 = P5.x – P4.x, and the region is R_mouth = cvRect(P4.x-L1 / 2,P4.y-L1 / 2,L1,L1); in the first target difference map, the pixels of the face region in the first difference map including these 3 regions are obtained, and the sum of the pixels of the 3 regions in the face region of the first difference map is calculated to obtain the pixel sum of the first face region.

[0150] The step of updating the stuttering detection result of the target audio and video based on the sum of the pixels of the first face region may include: if the sum of the pixels of the first face region is less than or equal to a preset threshold (e.g., 0 pixels or 0.3 pixels), then the stuttering detection result of the target audio and video is updated to indicate that stuttering exists in the target audio and video; if the sum of the pixels of the first face region is greater than the preset threshold (e.g., 0 pixels or 0.3 pixels), then the stuttering detection result of the target audio and video is updated to indicate that stuttering does not exist in the target audio and video.

[0151] The step of updating the stuttering detection result of the target audio and video based on the sum of pixels in the first face region may further include: calling a pixel stuttering mapping table, wherein the pixel stuttering mapping table is used to indicate the stuttering detection result corresponding to each pixel interval. Then, the sum of pixels in the first face region is compared with each pixel interval in the pixel stuttering mapping table, and the stuttering detection result of the target audio and video is updated to the detection result corresponding to the pixel interval where the sum of pixels in the first face region is located.

[0152] In one embodiment, in addition to performing face detection on the first target test image, face detection can also be performed on the second target test image to obtain the location information of the second face region, and face detection can also be performed on the third target test image to obtain the location information of the third face region, etc., thereby updating the stuttering detection result together based on the face region pixels of each difference image. In this way, the accuracy of the stuttering detection result update can be further improved.

[0153] In one example, combining the previous example, there are 10 target difference maps. After extracting the face region, 10 difference map face regions can be obtained. Judgment condition 1: If the sum of pixels in each difference map face region is equal to 0, then it is considered to be lagging. Judgment condition 2: If the sum of pixels is less than a pixel threshold (e.g., 50), the number of target difference maps that meet the conditions is N. If the proportion of N to the total number of target difference maps is greater than a preset proportion (e.g., 0.5), then it is considered that there may be lagging, and further confirmation is needed; if the proportion of N to the total number of target difference maps is less than or equal to the preset proportion...

[0154] In summary, the present application can achieve the following technical effects: accurately detect the integrity of audio and video, ensuring that dual-recorded videos can be traced back; execute quickly, consume little time, respond promptly, and facilitate business completion; reduce manual quality inspection, further reduce costs and increase efficiency, and enhance value.

[0155] Please see Figure 6 This application also provides an audio / video stuttering detection device, which can implement the above-described audio / video stuttering detection method. Figure 6 The present invention provides a module structure block diagram of an audio / video stuttering detection device, which includes: a data acquisition module 601, a data separation module 602, an image extraction module 603, an audio sampling module 604, a spectrum conversion module 605, an image synthesis module 606, and a result determination module 607. The system comprises the following modules: a data acquisition module 601 for acquiring target audio and video; a data separation module 602 for separating the target audio and video to obtain target video and target audio; an image extraction module 603 for extracting frames from the target video according to a preset period to obtain image frames; an audio sampling module 604 for sampling the target audio according to a preset period to obtain audio sampling points; a spectrum conversion module 605 for performing spectrum conversion on the audio sampling points to obtain a spectrum diagram; an image synthesis module 606 for synthesizing the image frame and the spectrum diagram to obtain a target test image when the extraction time of the image frame is the same as the sampling time of the audio sampling point; and a result determination module 607 for determining the stuttering detection result of the target audio and video based on the difference between any two target test images. The stuttering detection result is used to characterize whether the target audio and video has stuttering or not.

[0156] In one embodiment, after separating the target audio and video to obtain the target video and target audio, the audio and video stuttering detection device further includes a prompting module, which is used to: obtain the duration of the target video; obtain the duration of the target audio; calculate the difference between the target video duration and the target audio duration to obtain an absolute duration difference; if the absolute duration difference is greater than a preset duration difference threshold, generate target prompt information; wherein, the target prompt information is used to prompt the re-recording of the target audio and video.

[0157] In one embodiment, after determining the stuttering detection result of the target audio / video based on the difference between any two target test images, the audio / video stuttering detection device further includes a result update module, configured to: if the stuttering detection result of the target audio / video indicates that the target audio / video does not have stuttering, then perform face detection on the first target test image to obtain the location information of the first face region; extract the region from the first target difference image based on the location information of the first face region to obtain the face region of the first difference image; sum the pixels of each pixel in the face region of the first difference image to obtain the pixel sum of the first face region; and update the stuttering detection result of the target audio / video based on the pixel sum of the first face region.

[0158] It should be noted that the specific implementation method of this audio and video stuttering detection device is basically the same as the specific implementation method of the above-mentioned audio and video stuttering detection method, and will not be repeated here.

[0159] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned audio / video stuttering detection method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0160] Please see Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0161] The processor 701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0162] The memory 702 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 702 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called and executed by the processor 701 using the audio / video stuttering detection method of the embodiments of this application.

[0163] The input / output interface 703 is used to implement information input and output;

[0164] The communication interface 704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0165] Bus 705 transmits information between various components of the device (e.g., processor 701, memory 702, input / output interface 703, and communication interface 704);

[0166] The processor 701, memory 702, input / output interface 703, and communication interface 704 are connected to each other within the device via bus 705.

[0167] This application also provides a computer-readable storage medium for computer-readable storage, wherein the storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described audio and video stuttering detection method.

[0168] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0169] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0170] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0171] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0172] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0173] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0174] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0175] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0176] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0177] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0178] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0179] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An audio-video stalling detection method, characterized in that, The method comprises: acquiring a target audio-video; performing audio-video separation on the target audio-video to obtain a target video and a target audio; extracting frames from the target video according to a preset period to obtain image frames; sampling the target audio according to the preset period to obtain audio sampling points; performing spectrum conversion on the audio sampling points to obtain a spectrum graph; when the extraction time of the image frames is the same as the sampling time of the audio sampling points, performing image synthesis on the image frames and the spectrum graph to obtain a target test image; acquiring a target time of each target test image; wherein the target time, the extraction time and the sampling time of the same target test image are the same; extracting a first target test image and a second target test image from each target test image according to the target time; wherein the target time of the second target test image is later than the target time of the first target test image; performing image difference calculation on the first target test image and the second target test image to obtain a first target difference image; performing pixel summation on each pixel point in the first target difference image to obtain a first test pixel sum; determining a stall detection result of the target audio-video according to the first test pixel sum; wherein the stall detection result is used to represent whether the target audio-video has a stall or not.

2. The method of claim 1, wherein, The method further comprises: extracting a third target test image from each target test image according to the target time corresponding to the second target test image; wherein the target time of the third target test image is later than the target time of the second target test image; performing image difference calculation on the first target test image and the third target test image to obtain a second target difference image; performing pixel summation on each pixel point in the second target difference image to obtain a second test pixel sum; performing image difference calculation on the second target test image and the third target test image to obtain a third target difference image; performing pixel summation on each pixel point in the third target difference image to obtain a third test pixel sum; performing pixel stall detection on the first test pixel sum, the second test pixel sum and the third test pixel sum to obtain the stall detection result.

3. The method of claim 2, wherein, The performing pixel stall detection on the first test pixel sum, the second test pixel sum and the third test pixel sum to obtain the stall detection result comprises: constructing a stall graph number mark, the stall graph number mark being 0; taking the first test pixel sum, the second test pixel sum and the third test pixel sum as a target test pixel sum in turn; if the target test pixel sum is less than or equal to a preset pixel sum threshold, increasing the stall graph number mark by 1; performing number stall detection according to the stall graph number mark to obtain the stall detection result.

4. The method according to any one of claims 1 to 3, characterized in that, The performing image synthesis on the image frames and the spectrum graph to obtain a target test image when the extraction time of the image frames is the same as the sampling time of the audio sampling points comprises: scaling the image frame according to a preset target scale, and scaling the spectrum graph according to the target scale, so that the scale of the image frame is the same as the scale of the spectrum graph; when the extraction time of the image frame is the same as the sampling time of the audio sampling point, splicing the spectrum graph on the left side of the image frame to obtain the target test image.

5. The method according to any one of claims 1 to 3, characterized in that, After the method according to the first test pixel sum determines the stutter detection result of the target audio video, the method further comprises: if the stutter detection result of the target audio video indicates that the target audio video does not exist stutter, performing face detection on the first target test image to obtain first face region position information; extracting the first target difference graph according to the first face region position information to obtain a first face region difference graph; performing pixel summation according to each pixel point in the first face region difference graph to obtain a first face region pixel sum; updating the stutter detection result of the target audio video according to the first face region pixel sum.

6. The method according to any one of claims 1 to 3, characterized in that, After the method of separating the target audio video into audio and video to obtain a target video and a target audio, the method further comprises: obtaining the length of the target video to obtain a target video length; obtaining the length of the target audio to obtain a target audio length; performing difference calculation according to the target video length and the target audio length to obtain an absolute time length difference value; if the absolute time length difference value is greater than a preset time length difference threshold value, generating a target prompt information; wherein the target prompt information is used to prompt to re-record the target audio video.

7. An audio-video stutter detection apparatus, comprising: The device comprises: a data acquisition module for acquiring a target audio video; a data separation module for separating the target audio video into audio and video to obtain a target video and a target audio; an image extraction module for extracting frames from the target video according to a preset period to obtain an image frame; an audio sampling module for sampling audio from the target audio according to the preset period to obtain an audio sampling point; a spectrum conversion module for converting the audio sampling point into a spectrum graph; an image synthesis module for synthesizing the image frame and the spectrum graph when the extraction time of the image frame is the same as the sampling time of the audio sampling point to obtain a target test image; A result determining module is configured to: acquire a target time point of each target test image; wherein the target time point, the extraction time point and the sampling time point of a same target test image are the same; extract a first target test image and a second target test image from each target test image according to the target time point; wherein the target time point of the second target test image is later than that of the first target test image; perform image difference calculation according to the first target test image and the second target test image to obtain a first target difference image; perform pixel summation according to each pixel point in the first target difference image to obtain a first test pixel sum; and determine a freezing detection result of the target audio and video according to the first test pixel sum; wherein the freezing detection result is used to represent whether the target audio and video has freezing or not.

8. An electronic device, comprising: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method of any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Game frame stagnation test method and device

    CN105761255A

  • Catton detection method and device

    CN109089137A