A Dual-Recording Video Freezing Detection Method Based on Frame Mean Residual

Through the method based on the frame mean residual, the absolute residual between video frames is calculated and the similarity index is constructed, which solves the efficiency and accuracy of traditional detection methods in complex scenarios, and is suitable for the quality inspection of double-recorded videos in the financial industry.

CN120186323BActive Publication Date: 2025-07-25GUANGDONG MICROPATTERN SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510641224.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-25
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

When facing complex scenarios, traditional video stutter detection methods have low detection efficiency and high false alarm rate, making it difficult to meet the needs of double-recorded video quality inspection in the financial industry.

Method used

Using a method based on the frame mean residual, the pixel mean of the head and tail frames is calculated as the reference frame for adjacent three frame groups, and the absolute residual matrix between the intermediate frame and the reference frame is calculated, a normalized similarity index is constructed, and the stuttering period is determined.

Benefits of technology

It significantly reduces the false alarm rate and improves detection efficiency. It is especially suitable for the quality inspection of double-recorded videos in the financial industry, improving the accuracy and robustness of the inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186323B_ABST
    Figure CN120186323B_ABST
Patent Text Reader

Abstract

The present invention discloses a dual-recording video freeze detection method based on frame mean residual, which relates to the technical field of video freeze detection. It preprocesses video frames, constructs reference frames, analyzes residuals, and determines freezes, achieving efficient detection of video freezes. First, video frames are extracted at fixed time intervals, and each frame is uniformly downsampled to a specific resolution to generate a standardized frame sequence. Second, for adjacent three-frame groups, the pixel mean of the first and last frames is calculated as the reference frame. Then, the residual between the middle frame and the reference frame is calculated, and the maximum residual value is extracted as a measure of the inter-frame difference. Finally, a normalized similarity index is constructed. When the similarity exceeds the threshold, it is determined that a freeze occurs in the corresponding time interval. By traversing all three-frame groups, freeze detection in the entire time domain of the video is achieved. The present invention avoids the disturbance of video compression distortion that traditional one-way prediction is vulnerable to through the residual analysis of three-frame groups, improving the accuracy of the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video stutter detection, and in particular to a method for detecting double-recording video stutter based on frame mean residual. Background Art

[0002] With the rapid development of digital video technology and its wide application in various fields, the importance of video stutter detection technology has become increasingly prominent. Video stutter not only affects the user experience but may also damage the integrity and accuracy of video content. In fields such as video surveillance, online education, and remote conferencing, the smoothness of video is one of the key indicators for measuring service quality.

[0003] However, traditional video stutter detection methods often rely on unidirectional prediction or simple inter-frame difference analysis, which are easily affected by noise interference and complex scenes, resulting in a high false alarm rate and low detection efficiency. Especially in the quality inspection of double-recording videos in the financial industry, the problem of video stutter is particularly important. Financial institutions such as banks, securities companies, and insurance companies are required by regulatory authorities to record audio and video (referred to as double-recording) when selling wealth management products, precious metals, and insurance products. As a digital medium for playback and display, double-recording videos first need to ensure that the internal audio and video streams can be normally displayed and played. What can be normally played generally includes clear audio and video content, smooth audio and video playback without stutter, and audio and video synchronization. Among them, video stutter is a common quality problem that may affect the compliance and usability of double-recording videos. However, traditional video stutter detection methods often have difficulty detecting stutter phenomena efficiently and accurately in the face of complex scenes of double-recording videos, resulting in low quality inspection efficiency and a high false alarm rate.

[0004] In response to this problem, although some improvement methods have been proposed in the prior art, there are still problems such as insufficient detection accuracy, high computational complexity, and difficulty in adapting to complex video scenes. Therefore, there is an urgent need for a more efficient, accurate, and robust video stutter detection method to meet the requirements in practical applications. Summary of the Invention

[0005] In order to solve the above technical problems of video stutter detection, the present invention provides a method for detecting double-recording video stutter based on frame mean residual. The following technical solutions are adopted:

[0006] A method for detecting double-recording video stutter based on frame mean residual, comprising the following steps:

[0007] Step 1, preprocessing the video to be detected to obtain a standardized sequence;

[0008] Step 2: In the standardized sequence obtained in Step 1, calculate the pixel mean of the first and last frames of adjacent three-frame groups as the reference frame;

[0009] Step 3: Calculate the absolute residual matrix between the intermediate frame and the reference frame, and extract the maximum residual value as a measure of the inter-frame difference;

[0010] Step 4: Construct a normalized similarity index. When the similarity exceeds the threshold, it is determined that stuttering occurs within the corresponding time interval;

[0011] Step 5: Traverse all three-frame groups to perform stuttering detection on the entire time domain of the video to be detected, and output the stuttering detection result of the video to be detected.

[0012] By adopting the above technical solution, first, extract the video frames to be detected at a fixed time interval, and uniformly downsample each video frame to a specific resolution to generate a standardized frame sequence.

[0013] Secondly, for adjacent three-frame groups, calculate the pixel mean of the first and last frames as the reference frame.

[0014] Then, calculate the residual between the intermediate frame and the reference frame, and extract the maximum residual value as a measure of the inter-frame difference.

[0015] Finally, construct a normalized similarity index. When the similarity exceeds the threshold, it is determined that stuttering occurs within the corresponding time interval. By traversing all three-frame groups, stuttering detection of the entire time domain of the video is achieved. This method avoids the disturbance of video compression distortion that traditional single-direction prediction is vulnerable to through the residual analysis of three-frame groups, and improves the accuracy of the algorithm.

[0016] This method effectively avoids the limitations of traditional methods by constructing a reference frame for three-frame groups and analyzing the inter-frame residuals, significantly reducing the false alarm rate and improving the detection efficiency. This method provides a new solution for video stuttering detection, is particularly suitable for the quality inspection scenario of dual-recording videos in the financial industry, and has important application value and promotion significance.

[0017] Optionally, the method for preprocessing in Step 1 is:

[0018] Extract video frames at a fixed time interval and obtain the luminance image of the video frame to generate a sampling sequence , uniformly downsample each video frame to resolution to obtain a standardized sequence , and the corresponding timestamps are .

[0019] By adopting the above technical solution, extract video frames at a fixed time interval (generally seconds), and obtain the luminance image of this frame to generate a sampling sequence . Uniformly downsample each frame to resolution (generally ), obtain the standardized sequence , and the corresponding timestamp is . By controlling the resolution downsampling, the interference of high-frequency signal noise on the video frame similarity evaluation can be effectively reduced. At the same time, this downsampling can complete the summary of a single pixel for a local small area of the video, which is convenient for subsequent measurement of the inter-frame similarity.

[0020] Optionally, the method for constructing the reference frame in step 2 is as follows:

[0021] In the standardized sequence, for adjacent three-frame groups , calculate the pixel mean value of the first and last frames as the reference frame:

[0022] ;

[0023] where , is the brightness pixel value of the th row and th column of the frame, is the th row and th column of the frame, and

[0024] is the pixel mean value of the first and last frames. By adopting the above technical solution, the pixels at the corresponding positions of the two frames are averaged to achieve the content averaging of the two image frames and obtain a reference value. The absolute value of the difference between the pixel values of the current frame and the reference frame is taken, and the size of the absolute value reflects the difference between the two frames, which is conducive to judging whether there is monotonic repetition in the video frames. The similarity evaluation of the past frame and the future frame of a certain video frame over a certain time interval avoids the incompleteness and residual error accumulation existing in the unidirectional prediction on the time axis, and can also avoid false alarms caused by the high similarity of adjacent frames.

[0025] Optionally, the method for constructing the normalized similarity index in step 4 is as follows:

[0026] ;

[0027] where , is the normalized similarity index, set the similarity threshold , if it is judged that , then it is judged that there is a freeze in the time interval where the three-frame group corresponding to is located.

[0028] By adopting the above technical solution, That is, the maximum value of the residual is transformed into the similarity of the frames. Using a single maximum pixel represents the difference between frames, which can effectively adapt to the problem that the moving picture area accounts for a small proportion in fixed-position camera (such as surveillance) devices, and avoid misjudgment of frame similarity caused by a large area of the main picture being static and only small targets moving.

[0029] Optionally, the method for the step 5 to output the card detection result of the video to be detected is as follows:

[0030] Number all three-frame groups of the video to be detected, repeat steps 2-4 to complete the traversal of all video three-frame groups, and summarize and output the card detection result of the video to be detected based on the card detection results of the three-frame groups corresponding to the numbers.

[0031] A dual-recording video card detection device based on frame mean residual is used to implement a dual-recording video card detection method based on frame mean residual. The dual-recording video card detection device includes a memory and a processor. The memory stores a dual-recording video card detection program designed by a dual-recording video card detection method based on frame mean residual, and stores the video to be detected. The processor is communicatively connected to the memory, inputs the video to be detected into the dual-recording video card detection program, and runs the dual-recording video card detection program to output the dual-recording video card detection result.

[0032] Optionally, it further includes a display, and the display is communicatively connected to the processor for displaying the dual-recording video card detection result output by the processor.

[0033] Memory, and the memory stores a dual-recording video card detection program designed by a dual-recording video card detection method based on frame mean residual.

[0034] In summary, the present invention includes at least the following beneficial technical effects:

[0035] The present invention can provide a dual-recording video card detection method based on frame mean residual, and realizes the efficient detection of video card through preprocessing of video frames, reference frame construction, residual analysis and card determination. First, video frames are extracted at fixed time intervals, and each frame is uniformly downsampled to a specific resolution to generate a standardized frame sequence. Secondly, for adjacent three-frame groups, the pixel mean of the first and last frames is calculated as the reference frame. Then, the residual between the middle frame and the reference frame is calculated, and the maximum residual value is extracted as a measure of the difference between frames. Finally, a normalized similarity index is constructed, and when the similarity exceeds the threshold, it is determined that a card occurs within the corresponding time interval. By traversing all three-frame groups, the card detection of the entire time domain of the video is realized. This method avoids the disturbance of video compression distortion that traditional one-way prediction is vulnerable to through the residual analysis of three-frame groups, and improves the accuracy of the algorithm.

[0036] It provides a brand-new solution for video stutter detection, which is especially suitable for the dual-recording video quality inspection scenario in the financial industry and has important application value and promotion significance. Brief Description of the Drawings

[0037] Figure 1 is a schematic flowchart of a method for detecting stutters in dual-recording videos based on frame mean residuals according to the present invention;

[0038] Figure 2 is a schematic diagram of a video frame in a specific embodiment of the present invention;

[0039] Figure 3 is a group of three adjacent frames in a specific embodiment of the present invention schematic diagram;

[0040] Figure 4 is a reference frame in a specific embodiment of the present invention effect schematic diagram;

[0041] Figure 5 is a residual image in a specific embodiment of the present invention effect schematic diagram. Detailed Embodiments

[0042] The present invention will be further described in detail below with reference to the accompanying drawings. An embodiment of the present invention discloses a method for detecting stutters in dual-recording videos based on frame mean residuals.

[0043] Referring to Figures 1 - 5 , a method for detecting stutters in dual-recording videos based on frame mean residuals includes the following steps:

[0044] Step 1: Preprocess the video to be detected to obtain a standardized sequence;

[0045] Step 2: In the standardized sequence obtained in Step 1, calculate the pixel mean of the first and last frames of each group of three adjacent frames as the reference frame;

[0046] Step 3: Calculate the absolute residual matrix between the middle frame and the reference frame, and extract the maximum residual value as a measure of the inter-frame difference;

[0047] Step 4: Construct a normalized similarity index, and when the similarity exceeds the threshold, it is determined that stuttering occurs within the corresponding time interval;

[0048] Step 5: Traverse all groups of three frames to perform stutter detection on the entire time domain of the video to be detected, and output the stutter detection result of the video to be detected.

[0049] First, extract the video frames to be detected at fixed time intervals, and uniformly downsample each video frame to a specific resolution to generate a standardized frame sequence.

[0050] Secondly, for adjacent three-frame groups, calculate the pixel mean of the first and last frames as the reference frame.

[0051] Then, calculate the residual between the middle frame and the reference frame, and extract the maximum residual value as a measure of the inter-frame difference.

[0052] Finally, construct a normalized similarity index. When the similarity exceeds the threshold, it is determined that stuttering occurs within the corresponding time interval. By traversing all three-frame groups, stuttering detection for the entire time domain of the video is achieved. This method avoids the disturbance of video compression distortion that traditional one-way prediction is vulnerable to through the residual analysis of three-frame groups, improving the accuracy of the algorithm.

[0053] This method effectively avoids the limitations of traditional methods by constructing a reference frame for three-frame groups and analyzing the inter-frame residuals, significantly reducing the false alarm rate and improving the detection efficiency. This method provides a new solution for video stuttering detection, especially suitable for the quality inspection scenario of dual-recording videos in the financial industry, and has important application value and promotion significance.

[0054] The method of step 1 preprocessing is as follows:

[0055] At a fixed time interval Extract video frames, obtain the luminance image of the video frames, and generate a sampling sequence , uniformly downsample each video frame to Resolution to obtain a standardized sequence , and the corresponding timestamps are .

[0056] At a fixed time interval (usually seconds) extract video frames, and obtain the luminance image of this frame to generate a sampling sequence . Uniformly downsample each frame to Resolution (usually ), to obtain a standardized sequence , and the corresponding timestamp is . By controlling the resolution downsampling, the interference of high-frequency signal noise on the video frame similarity evaluation can be effectively reduced. At the same time, this downsampling can complete the summary of a single pixel for a small local area of the video, facilitating the subsequent measurement of inter-frame similarity.

[0057] The method of step 2 constructing the reference frame is as follows:

[0058] In the standardized sequence, for adjacent three-frame groups , calculate the pixel mean of the first and last frames as the reference frame:

[0059] ;

[0060] Where , is the luminance pixel value at the th row and th column of the frame, is the luminance pixel value at the th row and th column of the

[0061] Take the average of the pixels at the corresponding positions of two frames of images to average the content of the two image frames and obtain a reference value. Subtract the corresponding pixel values of the current frame from the reference frame and take the absolute value. The magnitude of the absolute value reflects the difference between the two frames, which is conducive to determining whether there is monotonic repetition in the video frames. Evaluate the similarity of the past frame and the future frame of a certain video frame over a certain time interval, which avoids the incompleteness and cumulative residual error existing in the one-way prediction on the time axis, and can also avoid false alarms caused by the high similarity between adjacent frames.

[0062] The method for constructing the normalized similarity index in step 4 is:

[0063] ;

[0064] where and is the normalized similarity index. Set the similarity threshold . If it is judged that , then it is judged that there is a freeze in the time interval where the corresponding three-frame group is located.

[0065] Convert i.e., the maximum value of the residual, into the similarity of the frame. Using a single maximum pixel to represent the difference between frames can effectively adapt to the problem that the proportion of the moving picture area is small in fixed-position camera (such as surveillance) devices, and avoid misjudgment of frame similarity caused by a large area of the main picture being stationary and only small targets moving.

[0066] The method for the output of the freeze detection result of the video to be detected in step 5 is:

[0067] Number all the three-frame groups of the video to be detected, and repeat steps 2 - 4 to complete the traversal of all video three-frame groups. Based on the freeze detection results of the three-frame groups corresponding to the numbers, summarize and output the freeze detection result of the video to be detected.

[0068] A dual-recording video freeze detection device based on frame mean residual is used to implement a dual-recording video freeze detection method based on frame mean residual. The dual-recording video freeze detection device includes a memory and a processor. The memory stores a dual-recording video freeze detection program designed by a dual-recording video freeze detection method based on frame mean residual, and stores the video to be detected. The processor is communicatively connected to the memory, inputs the video to be detected into the dual-recording video freeze detection program, and runs the dual-recording video freeze detection program to output the dual-recording video freeze detection result.

[0069] It further includes a display, which is communicatively connected to the processor and is used to display the dual-recording video freeze detection result output by the processor.

[0070] The memory stores a dual-recording video freeze detection program designed by a dual-recording video freeze detection method based on frame mean residual.

[0071] The following illustrates the implementation principle of a dual-recording video freeze detection method based on frame mean residual based on an embodiment:

[0072] Step 1: Video preprocessing:

[0073] Take a 352×288 video with a duration of 20 seconds.

[0074] Extract video frames at a fixed time interval of seconds, and obtain the luminance image of this frame. As Figure 2 shown, generate a sampling sequence . Uniformly downsample each frame to a resolution of 100×100 to obtain a standardized sequence .

[0075] Step 2: Reference frame construction:

[0076] In the standardized sequence , for adjacent three-frame groups , as Figure 3 shown, calculate the pixel mean of the first and last frames as the reference frame , as Figure 4 shown, that is .

[0077]

[0078] Here , .

[0079] Step 3: Residual analysis:

[0080] Calculate the residual between the middle frame and the reference frame, that is.

[0081] ;

[0082] Here, is the residual corresponding to the second frame. As Figure 5 shown, is the luminance pixel value of the th row and th column of the frame. Then :

[0083] ;

[0084] The maximum value of the residual image ;

[0085] Step 4: Judging the freeze in the three-frame interval:

[0086] Construct a normalized similarity index:

[0087] ;

[0088] Since , it is determined that there is no freeze in the time interval .

[0089] Step 5: Judging the freeze of the entire video:

[0090] By repeating Steps 2 - 4 to complete the traversal of frames ( ), the freeze detection of all time intervals of the video can be covered. It can be known that there is a freeze in the interval where S8 and S9 are located, that is, there is a freeze in the entire video interval.

[0091]

[0092] The above are all the preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for detecting double-recorded video freezing based on frame mean residual, characterized in that, Including the following steps: Step 1: Preprocess the video to be detected to obtain a standardized sequence; Step 2: In the standardized sequence obtained in Step 1, calculate the pixel mean of the first and last frames of adjacent three-frame groups as the reference frame; Step 3: Calculate the absolute residual matrix between the middle frame and the reference frame, and extract the maximum residual value as a measure of the inter-frame difference; Step 4: Construct a normalized similarity index. When the similarity exceeds the threshold, it is determined that stuttering occurs within the corresponding time interval; Step 5: Traverse all three-frame groups, perform stuttering detection on the entire time domain of the video to be detected, and output the stuttering detection result of the video to be detected; The method for constructing the reference frame in Step 2 is: In the standardized sequence, for the adjacent three-frame group , calculate the pixel mean of the first and last frames as the reference frame: ; Among them , is the luminance pixel value of the th row and th column of the frame, is th row and th column of the frame; is the average pixel value of the first and last frames; The formula for calculating the absolute residual matrix in Step 3 is: ; Wherein is the residual corresponding to the th frame, and is the th row and th column luminance pixel value of the th frame. Then, the maximum residual value of is extracted as a measure of the inter-frame difference.

2. The dual-recording video freeze detection method based on frame mean residual according to claim 1, wherein The method for preprocessing in Step 1 is: At fixed time intervals Extract video frames, obtain the luminance images of the video frames, and generate a sampling sequence , uniformly downsample each video frame to The resolution to obtain a standardized sequence , and the corresponding timestamps are respectively .

3. The dual-recording video freeze detection method based on frame mean residual according to claim 1, characterized in that The method for constructing the normalized similarity index in Step 4 is: ; Among them , is a normalized similarity index, and a similarity threshold is set. If it is judged that , then it is judged that there is a freeze in the time interval where the corresponding three-frame group is located.

4. The dual-recording video freeze detection method based on frame mean residual according to claim 1, characterized in that The method for the output in Step 5 to output the stuttering detection result of the video to be detected is: Number all three-frame groups of the video to be detected, repeat Steps 2 - 4 to complete the traversal of all video three-frame groups, and summarize and output the stuttering detection result of the video to be detected based on the stuttering detection results of the three-frame groups corresponding to the numbers.

5. A dual-recording video freezing detection device based on frame mean residual, characterized in that, A dual-recording video stuttering detection device for implementing the dual-recording video stuttering detection method according to any one of claims 1 - 4 includes a memory and a processor. The memory stores a dual-recording video stuttering detection program designed by the dual-recording video stuttering detection method according to any one of claims 1 - 4, and stores the video to be detected. The processor is communicatively connected to the memory, inputs the video to be detected into the dual-recording video stuttering detection program, and runs the dual-recording video stuttering detection program to output the dual-recording video stuttering detection result.

6. The dual-recording video freeze detection device based on frame mean residual according to claim 5, wherein: It further includes a display, which is communicatively connected to the processor and is used to display the dual-recording video stuttering detection result output by the processor.

7. A memory, characterized in that, The memory stores a dual-recording video stuttering detection program designed by the dual-recording video stuttering detection method according to any one of claims 1 - 4.

Citation Information

Patent Citations

  • Method and device for detecting video lag, computer equipment and storage medium

    CN110996094A

  • Audio and video media transmission method and system

    CN119767009A