Video stability assessment method and device, electronic equipment and storage medium

Through Fourier transform and spatiotemporal mapping map generation methods, the sampling location of video stability evaluation is automatically determined, which solves the problem of manual intervention selection of regions in the prior art, and realizes efficient automatic video stability evaluation.

CN120298952APending Publication Date: 2025-07-11BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510453975.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing video stability assessment techniques require manual intervention to select areas of interest, with low automation and low evaluation efficiency.

Method used

The frequency domain video sequence is obtained through Fourier transform, the time domain sequence of frequency components is extracted, the target sampling position is determined, the space-time map is generated, and the video stability is evaluated. The sampling position is automatically determined using frequency domain feature analysis to realize an end-to-end automated evaluation process.

Benefits of technology

The degree of automation and evaluation efficiency of video stability evaluation has been improved, and efficient stability evaluation without manual intervention has been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298952A_ABST
    Figure CN120298952A_ABST
Patent Text Reader

Abstract

The invention provides a video stability evaluation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a to-be-evaluated video, and performing Fourier transform on each frame of original video image in the to-be-evaluated video to obtain a frequency domain video sequence; for each spatial position in the frequency domain video sequence, extracting a frequency component time domain sequence of the spatial position in the frequency domain video sequence; determining a target sampling position according to the frequency component time domain sequence corresponding to each spatial position; extracting pixel data from each frame of original video image according to the target sampling position; a space-time mapping method is applied to the pixel data to generate a space-time mapping graph, and the stability of the to-be-evaluated video is evaluated according to the space-time mapping graph. According to the method, the target sampling position is automatically determined through frequency domain feature analysis, and an end-to-end automatic evaluation process is realized, so that the evaluation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technologies, and in particular, to a method, apparatus, electronic device, and storage medium for video stability evaluation. Background Art

[0002] With the popularization of video acquisition devices, video stability evaluation has become a key technical requirement in fields such as security monitoring, film and television production, and motion analysis. An ideal video stability evaluation method should have characteristics such as real-time processing ability, strong scene adaptability, and objective and accurate evaluation results. Currently, the mainstream video stability evaluation technologies include: the method based on image registration: calculating the displacement amount by aligning adjacent frames; the method based on optical flow: analyzing the motion trajectories of pixels.

[0003] However, the current technical solutions often require manual intervention to select the region of interest (i.e., the sampling position), with a low degree of automation. At the same time, the evaluation efficiency is also very low. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method, apparatus, electronic device, and storage medium for video stability evaluation to solve the problem that the current technical solutions often require manual intervention to select the region of interest, with a low degree of automation, and at the same time, the evaluation efficiency is also very low. The specific technical solutions are as follows:

[0005] In a first aspect, this application provides a method for video stability evaluation, including:

[0006] Obtaining a video to be evaluated, and performing Fourier transform on each frame of the original video image in the video to be evaluated to obtain a frequency-domain video sequence;

[0007] For each spatial position in the frequency-domain video sequence, extracting the time-domain sequence of frequency components of the spatial position in the frequency-domain video sequence;

[0008] Determining a target sampling position according to the time-domain sequence of frequency components corresponding to each spatial position;

[0009] Extracting pixel data from each frame of the original video image according to the target sampling position;

[0010] Applying a spatio-temporal mapping method to the pixel data to generate a spatio-temporal mapping diagram, and evaluating the stability of the video to be evaluated according to the spatio-temporal mapping diagram.

[0011] In a possible implementation manner, the determining a target sampling position according to the time-domain sequence of frequency components corresponding to each spatial position includes:

[0012] For each spatial position, calculating high-frequency energy distribution data according to the time-domain sequence of frequency components of the spatial position;

[0013] Determine the spatial positions where the corresponding high-frequency energy distribution data exceeds a preset energy threshold as candidate sampling positions;

[0014] For each candidate sampling position, calculate motion saliency data according to the time-domain sequence of frequency components at the candidate sampling position;

[0015] Perform a weighted summation operation on the motion saliency data and high-frequency energy distribution data of the candidate sampling position to obtain position scoring data;

[0016] Sort all candidate sampling positions in descending order according to the corresponding position scoring data, and determine the top several candidate sampling positions as target sampling positions.

[0017] In a possible implementation manner, the method further includes:

[0018] For each candidate sampling position, determine the content category corresponding to the video content at the candidate sampling position;

[0019] In the case where the content category is a static scene, increase the weight of the high-frequency energy distribution data and decrease the weight of the motion saliency data;

[0020] In the case where the content category is a dynamic scene, increase the weight of the motion saliency data and decrease the weight of the high-frequency energy distribution data.

[0021] In a possible implementation manner, the extracting pixel data from each frame of the original video image according to the target sampling position includes:

[0022] Intercept a rectangular pixel region in each frame of the original video image with the target sampling position as the center;

[0023] Perform time alignment on the rectangular pixel regions corresponding to all the original video images to obtain a set of pixel regions;

[0024] Organize the set of pixel regions into a three-dimensional spatio-temporal data structure to obtain pixel data, where two dimensions in the three-dimensional spatio-temporal data structure represent spatial coordinates and the third dimension represents a time sequence.

[0025] In a possible implementation manner, the evaluating the stability of the video to be evaluated according to the spatio-temporal mapping graph includes:

[0026] Determine strip continuity data and detail retention rate data according to the spatio-temporal mapping graph, where the strip continuity data is used to characterize the smoothness of video temporal changes, and the detail retention rate data is used to characterize the retention degree of high-frequency spatial features;

[0027] Perform a weighted sum on the strip continuity data and the detail retention rate data to obtain the stability score of the video to be evaluated.

[0028] In a possible implementation manner, the determining the strip continuity data according to the spatio-temporal mapping graph includes:

[0029] Detect texture strips extending along the time dimension in the spatio-temporal mapping graph;

[0030] Extract local binary pattern feature vectors of the strip region in consecutive time units;

[0031] Calculate the similarity of the local binary pattern feature vectors within adjacent time units;

[0032] Normalize the temporal average value of the similarity to a preset interval to obtain the strip continuity data.

[0033] In a possible implementation manner, the determining the detail retention rate data according to the spatio-temporal mapping graph includes:

[0034] Extract the target gradient magnitude spectrum of the video to be evaluated according to the spatio-temporal mapping graph;

[0035] Calculate the energy ratio of the target gradient magnitude spectrum and a preset standard gradient magnitude spectrum within a preset high-frequency band, where the standard gradient magnitude spectrum is the gradient magnitude statistical benchmark of a stable video sequence under the same shooting conditions;

[0036] Use the energy ratio as the detail retention rate data.

[0037] In a second aspect, the present application provides a video stability evaluation device, including:

[0038] An acquisition module, configured to acquire a video to be evaluated and perform Fourier transform on each frame of the original video image in the video to be evaluated to obtain a frequency-domain video sequence;

[0039] A sequence extraction module, configured to extract a time-domain sequence of frequency components of each spatial position in the frequency-domain video sequence for each spatial position in the frequency-domain video sequence;

[0040] A determination module, configured to determine a target sampling position according to the time-domain sequence of frequency components corresponding to each spatial position;

[0041] A data extraction module, configured to extract pixel data from each frame of the original video image according to the target sampling position;

[0042] An evaluation module, configured to apply a spatio-temporal mapping method to the pixel data to generate a spatio-temporal mapping graph and evaluate the stability of the video to be evaluated according to the spatio-temporal mapping graph.

[0043] In a possible implementation, the determining module is specifically configured to:

[0044] For each spatial position, calculate high-frequency energy distribution data according to the time-domain sequence of frequency components of the spatial position;

[0045] Determine the spatial positions where the corresponding high-frequency energy distribution data exceeds a preset energy threshold as candidate sampling positions;

[0046] For each candidate sampling position, calculate motion saliency data according to the time-domain sequence of frequency components of the candidate sampling position;

[0047] Perform a weighted summation operation on the motion saliency data and the high-frequency energy distribution data of the candidate sampling position to obtain position scoring data;

[0048] Sort all candidate sampling positions in descending order according to the corresponding position scoring data, and determine the top several candidate sampling positions as target sampling positions.

[0049] In a possible implementation, the determining module is further configured to:

[0050] For each candidate sampling position, determine the content category corresponding to the video content at the candidate sampling position;

[0051] In the case where the content category is a static scene, increase the weight of the high-frequency energy distribution data and decrease the weight of the motion saliency data;

[0052] In the case where the content category is a dynamic scene, increase the weight of the motion saliency data and decrease the weight of the high-frequency energy distribution data.

[0053] In a possible implementation, the data extraction module is specifically configured to:

[0054] Intercept a rectangular pixel region in each frame of the original video image with the target sampling position as the center;

[0055] Perform time alignment on the rectangular pixel regions corresponding to all the original video images to obtain a set of pixel regions;

[0056] Organize the set of pixel regions into a three-dimensional spatio-temporal data structure to obtain pixel data, where two dimensions in the three-dimensional spatio-temporal data structure represent spatial coordinates and the third dimension represents a time series.

[0057] In a possible implementation, the evaluation module is specifically configured to:

[0058] Determine the stripe continuity data and the detail retention rate data according to the spatio-temporal mapping graph, where the stripe continuity data is used to characterize the smoothness of the video temporal variation, and the detail retention rate data is used to characterize the retention degree of high-frequency spatial features;

[0059] Perform weighted summation on the stripe continuity data and the detail retention rate data to obtain the stability score of the video to be evaluated.

[0060] In a possible implementation manner, the evaluation module is further configured to:

[0061] Detect texture stripes extending along the time dimension in the spatio-temporal mapping graph;

[0062] Extract local binary pattern feature vectors of the stripe region in consecutive time units;

[0063] Calculate the similarity of local binary pattern feature vectors within adjacent time units;

[0064] Normalize the temporal average value of the similarity to a preset interval to obtain the stripe continuity data.

[0065] In a possible implementation manner, the evaluation module is further configured to:

[0066] Extract the target gradient magnitude spectrum of the video to be evaluated according to the spatio-temporal mapping graph;

[0067] Calculate the energy ratio between the target gradient magnitude spectrum and a preset standard gradient magnitude spectrum within a preset high-frequency band, where the standard gradient magnitude spectrum is the gradient magnitude statistical benchmark of a stable video sequence under the same shooting conditions;

[0068] Use the energy ratio as the detail retention rate data.

[0069] In a third aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0070] The memory is used to store a computer program;

[0071] The processor is configured to implement the method steps of any one of the first aspect when executing the program stored on the memory.

[0072] In a fourth aspect, a computer-readable storage medium is provided, characterized in that the computer-readable storage medium stores a computer program, and the computer program realizes the method steps of any one of the first aspect when executed by a processor.

[0073] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-mentioned video stability assessment methods.

[0074] Beneficial effects of the embodiments of the present application:

[0075] The embodiments of the present application provide a method, device, electronic device and storage medium for evaluating video stability. In the embodiments of the present application, first, a video to be evaluated is obtained, and each frame of the original video image in the video to be evaluated is subjected to Fourier transform to obtain a frequency domain video sequence. Then, for each spatial position in the frequency domain video sequence, the time domain sequence of the frequency component of the spatial position in the frequency domain video sequence is extracted. Then, the target sampling position is determined according to the time domain sequence of the frequency component corresponding to each spatial position, and pixel data is extracted from each frame of the original video image according to the target sampling position. Finally, a space-time mapping method is applied to the pixel data to generate a space-time mapping diagram, and the stability of the video to be evaluated is evaluated according to the space-time mapping diagram. The present application automatically determines the target sampling position through frequency domain feature analysis to achieve an end-to-end automated evaluation process, thereby improving evaluation efficiency.

[0076] Of course, implementing any product or method of the present application does not necessarily require achieving all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0078] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0079] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0080] Figure 1 A flowchart of a video stability evaluation method provided in an embodiment of the present application;

[0081] Figure 2 A flowchart of another video stability evaluation method provided in an embodiment of the present application;

[0082] Figure 3A flowchart of another video stability evaluation method provided by an embodiment of the present application;

[0083] Figure 4 A schematic structural diagram of a video stability evaluation device provided by an embodiment of the present application;

[0084] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0086] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0087] Figure 1 A schematic flow diagram of a video stability evaluation method provided by an embodiment of the present application. This method can be applied to one or more electronic devices such as smart phones, laptop computers, desktop computers, portable computers, and servers. In addition, the execution subject of this method can be hardware or software. When the above execution subject is hardware, the execution subject can be one or more of the above electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above execution subject is software, this method can be implemented as multiple software or software modules, or can be implemented as a single software or software module. No specific limitation is made here.

[0088] As Figure 1 shown, the method specifically includes:

[0089] S101. Obtain a video to be evaluated, and perform Fourier transform on each frame of the original video image in the video to be evaluated to obtain a frequency-domain video sequence.

[0090] The video to be evaluated refers to the original video sequence that needs to be analyzed for stability, usually consisting of consecutive frame images (such as a 1080p video at 30fps). For example, a video of a bumpy road captured by a vehicle-mounted camera, a jittery image captured by a drone during aerial photography, etc.

[0091] In the embodiments of this application, the video to be evaluated is denoted as where F t represents the original video image of the t-th frame, and N is the number of original video images in the video to be evaluated. Apply the Fourier transform to each frame of the original video image to convert the spatial domain image into a frequency domain image, and obtain: where represents the Fourier transform. Finally, record all the frequency domain images to form a frequency domain video sequence, denoted as:

[0092] S102. For each spatial position in the frequency domain video sequence, extract the time-domain sequence of the frequency components at this spatial position in the frequency domain video sequence.

[0093] Spatial position refers to the spatial frequency coordinate point (u, v) in the frequency domain image, corresponding to a specific frequency component of the spatial domain image.

[0094] The time-domain sequence of frequency components refers to the time-series signal composed of the frequency values at a certain fixed (u, v) position on all frames.

[0095] In the embodiments of this application, the specific method for extracting the time-domain sequence of frequency components is as follows: for each spatial frequency coordinate point (u, v) in the frequency domain video sequence, extract the magnitude of its corresponding complex frequency component along the time axis to form a one-dimensional time-domain signal sequence with a length of N (the total number of video frames). Specifically, first, index the three-dimensional frequency domain data (width × height × number of frames) according to the (u, v) coordinates to obtain the complex Fourier coefficients at this position on all frames, then extract its magnitude (modulus) to form a sequence, and finally, perform band-pass filtering (such as 0.5 - 5Hz) on this sequence to enhance the signal in the frequency band related to jitter, and obtain the time-domain sequence of frequency components, which is the time-domain feature reflecting the variation characteristics of this frequency point over time.

[0096] S103. Determine the target sampling position according to the time-domain sequence of frequency components corresponding to each spatial position.

[0097] In the embodiments of this application, for each spatial position, the high-frequency energy distribution and motion significance of this spatial position can be calculated according to the time-domain sequence of frequency components corresponding to this spatial position, and the optimal sampling position, that is, the target sampling position, can be selected according to the high-frequency energy distribution and motion significance.

[0098] As for how to specifically screen the target sampling positions, it will be explained in detail through the following embodiments and will not be elaborated here for the time being.

[0099] S104. Extract pixel data from each frame of the original video image according to the target sampling positions.

[0100] Pixel data refers to a local image block (such as a 16×16 pixel area) extracted centered on the target sampling position, including spatial domain luminance / color information and its time series changes.

[0101] Specifically, S104 may include the following steps: intercept a rectangular pixel area centered on the target sampling position in each frame of the original video image; perform time alignment on the rectangular pixel areas corresponding to all the original video images to obtain a set of pixel areas; organize the set of pixel areas into a three-dimensional spatio-temporal data structure to obtain pixel data, where two dimensions in the three-dimensional spatio-temporal data structure represent spatial coordinates and the third dimension represents the time series.

[0102] The rectangular pixel area refers to a local image block (such as 16×16 pixels) intercepted from the original video image centered on the target sampling position, including the luminance / color information at this position and the features of its surrounding neighborhoods.

[0103] Time alignment refers to compensating for the inter-frame displacement through motion estimation (such as phase correlation, optical flow) to ensure that the pixel blocks corresponding to the same sampling position in different frames are strictly aligned on the time axis, avoiding spatial offset interference caused by jitter for analysis.

[0104] The set of pixel areas refers to the set of aligned pixel blocks corresponding to the same target sampling position in all video frames (i.e., the original video images), constituting the spatio-temporal analysis unit at this position.

[0105] The three-dimensional spatio-temporal data structure refers to a data cube (such as 16×16×N, N being the number of frames) formed by organizing the set of pixel areas according to the spatial (width×height) and time (frame sequence) dimensions for subsequent spatio-temporal mapping analysis.

[0106] The core of this step is to extract the local spatio-temporal information of the target sampling positions from the original video, and the specific process is as follows: Local interception: For each target sampling position (such as (x,y)), intercept a rectangular pixel block with a fixed size (such as 16×16) on each frame of the image, covering this position and its neighborhood; Time alignment: Calculate the sub-pixel level offset of adjacent frames through motion compensation (such as the phase correlation method), and register the pixel blocks of all frames to eliminate the inter-frame displacement caused by jitter or motion; Data integration: Stack the aligned pixel blocks in chronological order to form a three-dimensional spatio-temporal cube (width×height×time), where the spatial dimensions (x,y) describe the local image features and the time dimension (t) records the change law of this area over time.

[0107] In this solution, the key positions screened by frequency-domain analysis are converted into spatio-temporal data blocks that can be quantitatively analyzed to ensure the accuracy of subsequent stability evaluation.

[0108] S105. Apply a spatio-temporal mapping method to the pixel data to generate a spatio-temporal mapping graph, and evaluate the stability of the video to be evaluated according to the spatio-temporal mapping graph.

[0109] The spatio-temporal mapping graph refers to a two-dimensional image generated by the spatio-temporal mapping method. Its horizontal axis usually represents time, and the vertical axis represents spatial positions (such as the rows or columns of pixel blocks). The continuity of the texture in the graph directly reflects the stability of the video (for example, parallel stripes indicate stability, while breaks / distortions indicate jitter).

[0110] In the embodiments of the present application, first, the pixel data at the target sampling positions is converted into a visual spatio-temporal mapping graph, and then, the video stability is evaluated based on the texture features of the spatio-temporal mapping graph. The specific process is as follows:

[0111] Spatio-temporal mapping generation: For the three-dimensional pixel data (such as 16×16×N) at each target sampling position, stack the pixel values of the central row or column along the time axis frame by frame to generate a two-dimensional spatio-temporal graph (the horizontal axis is the frame number, and the vertical axis is the pixel value);

[0112] Texture analysis: Quantify the degree of jitter by calculating features such as the continuity of the texture (such as the pixel difference between adjacent frames) and the detail fidelity in the spatio-temporal graph;

[0113] Comprehensive scoring: Combine the spatio-temporal graph features of all sampling positions and calculate the overall stability score of the video by weighted calculation (for example: the continuity accounts for 60%, and the detail fidelity accounts for 40%).

[0114] As for how to specifically evaluate the stability of the video to be evaluated according to the spatio-temporal mapping graph, it will be described in detail in the following embodiments and will not be elaborated here first.

[0115] Through this solution, abstract spatio-temporal data can be converted into stability evidence that can be intuitively interpreted, and at the same time, it supports automated scoring and manual re-inspection.

[0116] In the embodiments of the present application, first, an input video to be evaluated is obtained, and Fourier transform is performed on each frame of the original video image in the input video to be evaluated to obtain a frequency-domain video sequence. Then, for each spatial position in the frequency-domain video sequence, a time-domain sequence of frequency components at the spatial position in the frequency-domain video sequence is extracted. Furthermore, a target sampling position is determined according to the time-domain sequence of frequency components corresponding to each spatial position, and pixel data is extracted from each frame of the original video image according to the target sampling position. Finally, a spatio-temporal mapping method is applied to the pixel data to generate a spatio-temporal mapping diagram, and the stability of the input video to be evaluated is evaluated according to the spatio-temporal mapping diagram. The present application automatically determines the target sampling position through frequency-domain feature analysis, realizes an end-to-end automated evaluation process, and thus improves the evaluation efficiency.

[0117] See Figure 2 , which is a flowchart of an embodiment of another video stability evaluation method provided by the embodiments of the present application. The Figure 2 process shown is based on the Figure 1 process shown, and describes how to determine the target sampling position according to the time-domain sequence of frequency components corresponding to each spatial position. As Figure 2 shown, the process may include the following steps:

[0118] S201. For each spatial position, calculate high-frequency energy distribution data according to the time-domain sequence of frequency components at the spatial position.

[0119] S202. Determine the spatial positions where the corresponding high-frequency energy distribution data exceeds a preset energy threshold as candidate sampling positions.

[0120] S203. For each candidate sampling position, calculate motion saliency data according to the time-domain sequence of frequency components at the candidate sampling position.

[0121] S204. Perform a weighted summation operation on the motion saliency data and the high-frequency energy distribution data of the candidate sampling position to obtain position score data.

[0122] S205. Sort all candidate sampling positions in descending order according to the corresponding position score data, and determine the top several candidate sampling positions as the target sampling positions.

[0123] For ease of understanding, S201 - S205 are described together as follows:

[0124] The high-frequency energy distribution data (E_hf) refers to the total energy of the high-frequency band (ω > ω_L) in the frequency domain calculated by integration, and characterizes the richness of high-frequency details (such as edges / textures) at the position.

[0125] The preset energy threshold (E th) which refers to a dynamically set threshold value, usually taken as 1.5 times the upper quartile (Q3) of the high-frequency energy distribution data, and is used to screen candidate regions with rich details.

[0126] The motion saliency data (M) refers to a motion intensity index calculated based on the optical flow field, reflecting the intensity of motion at the spatial position (x, y).

[0127] In the embodiments of the present application, first, according to the time-domain sequence of frequency components of each spatial position (x, y) in the entire video sequence, its high-frequency energy distribution is calculated, and according to the preset high-frequency energy threshold E th Candidate sampling positions are screened out, and the specific process is as follows:

[0128] 2.1.1: The time-domain sequence of frequency components of the (x, y) position in the frequency-domain video sequence is denoted as:

[0129]

[0130] 2.1.2: High-frequency energy distribution data:

[0131]

[0132] where f hf is the high-frequency lower limit, f max is the maximum frequency, and |·| represents the amplitude of the frequency component.

[0133] 2.1.3: According to the preset high-frequency energy threshold E th , select the spatial positions with high-frequency energy exceeding E th as candidate sampling positions: C = {x, y|E hf (x, y)>E th};

[0134] Then, according to the time-domain sequence of frequency components of each spatial position (x, y) in the entire video sequence, its motion saliency data is calculated, and a weighted sum operation is performed on the motion saliency data and the high-frequency energy distribution data to obtain position scoring data. The specific process is as follows:

[0135] 2.2.1: Obtain the motion saliency data through methods such as optical flow method, block matching method, and phase correlation method.

[0136] 2.2.2: Perform a weighted sum operation on the motion saliency data and the high-frequency energy distribution data to obtain position scoring data:

[0137] S optimal = {x, y∈C|w hf ·E hf (x, y)+w motion ·M(x, y)}

[0138] Among them, w hf represents the weight of the high-frequency energy distribution data, M(x, y) represents the motion saliency data at the position (x, y), and w motion represents the weight of the motion saliency data.

[0139] Finally, all candidate sampling positions are sorted in descending order according to the corresponding position scoring data, and several candidate sampling positions ranked at the front are determined as the target sampling positions. For example, the 10 positions ranked at the front are determined as the target sampling positions.

[0140] In addition, in another embodiment of the present application, the method may further include the following steps:

[0141] For each candidate sampling position, determine the content category corresponding to the video content at the candidate sampling position; in the case where the content category is a static scene, increase the weight of the high-frequency energy distribution data, and decrease the weight of the motion saliency data; in the case where the content category is a dynamic scene, increase the weight of the motion saliency data, and decrease the weight of the high-frequency energy distribution data.

[0142] The content category refers to the video content attribute at the candidate sampling position, and is divided into a static scene (such as a static background, a fixed object) and a dynamic scene (such as a moving object, a rapid movement of the camera).

[0143] In the application, for each candidate position, it can be determined whether it belongs to a static scene (low optical flow / variance) or a dynamic scene (high optical flow / variance) through the optical flow amplitude or the frequency domain variance (i.e., the frequency change variance).

[0144] Taking the determination of the content category through the frequency change variance as an example:

[0145] Frequency change variance

[0146] Among them, is the average value of the frequency components at this position.

[0147] Then, the frequency change variance data > the preset variance threshold, it is determined as a dynamic scene; the frequency change variance data ≤ the preset variance threshold, it is determined as a static scene.

[0148] In this solution, the weight of the high-frequency detail richness standard is adaptively adjusted according to the video content, and the optimal sampling position set is selected.

[0149] For a static scene, increase the weight of the high-frequency details: w hf,static = w hf +Δw;

[0150] For dynamic scenes, reduce the weight of high-frequency details and increase the weight of motion saliency: w hf,dynamic = w hf - Δw; w motion,dynamic = w motion + Δw; where Δw is an empirical value.

[0151] That is to say, static scenes focus on high-frequency energy (detail stability is more important), reduce the weight of motion saliency (to avoid noise interference); dynamic scenes increase the weight of motion saliency (to exclude the interference of object movement), and reduce the dependence on high-frequency energy (movement may cause detail blurring). In this way, the problem of misjudgment in complex scenes by traditional methods (such as misidentifying a moving object as jitter) is solved, and the evaluation accuracy is improved.

[0152] Figure 2 In the process shown, first, based on the high-frequency energy distribution data, quickly lock the candidate areas with rich details, greatly reducing the subsequent calculation amount; secondly, introduce the dynamic weighted fusion of motion saliency data and high-frequency energy to effectively distinguish real jitter from object movement interference; finally, through scoring and ranking, select the most representative target sampling positions to improve the accuracy of stability evaluation.

[0153] See Figure 3 , which is a flowchart of an embodiment of another video stability evaluation method provided by an embodiment of the present application. This Figure 3 The process shown is based on the process shown above Figure 1 to describe how to evaluate the stability of the video to be evaluated according to the spatio-temporal mapping diagram. As Figure 3 shown, this process may include the following steps:

[0154] S301. Determine the stripe continuity data and detail retention rate data according to the spatio-temporal mapping diagram, where the stripe continuity data is used to characterize the smoothness of the temporal change of the video, and the detail retention rate data is used to characterize the retention degree of high-frequency spatial features;

[0155] S302. Perform weighted summation on the stripe continuity data and the detail retention rate data to obtain the stability score of the video to be evaluated.

[0156] For easy understanding, the following provides a unified description of S301 and S302:

[0157] The stripe continuity data is used to measure the coherence and consistency of texture stripes in the time dimension in the spatio-temporal mapping diagram. It reflects whether the change of stripes between video frames is smooth and stable.

[0158] Detail retention rate data is used to measure the degree of retention of detail information in the output result of a video algorithm. It reflects whether the algorithm can accurately retain high-frequency details in the input video, such as edges, textures, etc.

[0159] In the embodiments of the present application, first, strip continuity data and detail retention rate data are determined according to the spatio-temporal mapping graph.

[0160] Specifically, determining the strip continuity data according to the spatio-temporal mapping graph may include the following steps: detecting texture strips extending along the time dimension in the spatio-temporal mapping graph; extracting local binary pattern feature vectors of the strip regions in consecutive time units; calculating the similarity of the local binary pattern feature vectors within adjacent time units; normalizing the temporal average value of the similarity to a preset interval to obtain the strip continuity data.

[0161] In this solution, video stability is accurately quantified through multi-stage processing: first, texture strips extending along the time axis in the spatio-temporal mapping graph are detected to capture the temporal continuity characteristics of the video sequence; then, local binary pattern (LBP) feature vectors of the strip regions are extracted, which have strong representational ability for texture changes; then, the similarity of the LBP features of adjacent frames is calculated, and the temporal consistency is evaluated through measurement methods such as cosine similarity or Hamming distance; finally, the similarity results of the entire video sequence are temporally averaged and normalized to the [0,1] interval (i.e., the preset interval) to generate the final strip continuity score, where a high score indicates smooth and stable temporal changes in the video, and a low score reflects obvious jitter or breakage.

[0162] This method combines LBP texture analysis with temporal similarity calculation, significantly improving the accuracy of jitter detection while ensuring computational efficiency.

[0163] Determining the detail retention rate data according to the spatio-temporal mapping graph may include the following steps: extracting the target gradient magnitude spectrum of the video to be evaluated according to the spatio-temporal mapping graph; calculating the energy ratio of the target gradient magnitude spectrum and a preset standard gradient magnitude spectrum within a preset high-frequency band, where the standard gradient magnitude spectrum is the gradient magnitude statistical benchmark of a stable video sequence under the same shooting conditions; using the energy ratio as the detail retention rate data.

[0164] In this solution, the degree of video detail retention is quantified through gradient analysis: first, the gradient magnitude spectrum of the target video is extracted from the spatio-temporal mapping graph, reflecting the distribution of high-frequency information such as image edges and textures; then, the spectrum is compared with a preset standard gradient magnitude spectrum (obtained by statistically analyzing stable videos) in key high-frequency bands, and the ratio of the two is calculated; finally, the ratio is used as the detail retention rate data, intuitively representing the degree of retention of high-frequency details in the processed video relative to an ideal stable video.

[0165] This method establishes an objective evaluation criterion based on the gradient spectrum. By comparing the energy of frequency bands, it effectively avoids the uncertainty of subjective evaluation. The closer the ratio is to 1, the more complete the details are retained. If it is much less than 1, it indicates obvious detail loss or blurring during the video processing.

[0166] Then, through adaptive weighted fusion (for example, emphasizing the detail retention rate in static scenes and continuity in dynamic scenes), the stability score of the video to be evaluated is obtained, making the stability score take into account both spatial and temporal characteristics. The specific formula is as follows:

[0167] S stability =α·C temporal +β·F detail

[0168] Where α and β are weight coefficients (if emphasizing the detail retention rate in static scenes, the value of β is relatively high; if emphasizing continuity in dynamic scenes, the value of α is relatively high). C temporal is the strip continuity data, and F detail is the detail retention rate data.

[0169] Figure 3 The process shown achieves significant beneficial effects through two-dimensional quantitative evaluation: First, the strip continuity data accurately captures the smoothness changes of the video sequence from the time dimension and can effectively identify inter-frame jitter and breakage phenomena. Second, the detail retention rate data objectively evaluates the loss degree of high-frequency features from the spatial dimension, avoiding misjudgment of blurring processing by traditional methods and greatly improving the reliability and applicability of the evaluation results.

[0170] Based on the same technical concept, the embodiment of the present application also provides a video stability evaluation device, as Figure 4 shown. The device includes:

[0171] An acquisition module 41, configured to acquire the video to be evaluated and perform Fourier transform on each frame of the original video image in the video to be evaluated to obtain a frequency-domain video sequence;

[0172] A sequence extraction module 42, configured to extract the time-domain sequence of frequency components at the spatial position in the frequency-domain video sequence for each spatial position in the frequency-domain video sequence;

[0173] A determination module 43, configured to determine the target sampling position according to the time-domain sequence of frequency components corresponding to each spatial position;

[0174] A data extraction module 44, configured to extract pixel data from each frame of the original video image according to the target sampling position;

[0175] An evaluation module 45, configured to apply a spatio-temporal mapping method to the pixel data to generate a spatio-temporal mapping graph, and evaluate the stability of the video to be evaluated according to the spatio-temporal mapping graph.

[0176] In a possible implementation manner, the determining module is specifically configured to:

[0177] For each spatial position, calculate high-frequency energy distribution data according to the time-domain sequence of frequency components at the spatial position;

[0178] Determine the spatial positions where the corresponding high-frequency energy distribution data exceeds a preset energy threshold as candidate sampling positions;

[0179] For each candidate sampling position, calculate motion saliency data according to the time-domain sequence of frequency components at the candidate sampling position;

[0180] Perform a weighted summation operation on the motion saliency data and the high-frequency energy distribution data at the candidate sampling position to obtain position scoring data;

[0181] Sort all candidate sampling positions in descending order according to the corresponding position scoring data, and determine the top several candidate sampling positions as target sampling positions.

[0182] In a possible implementation manner, the determining module is further configured to:

[0183] For each candidate sampling position, determine the content category corresponding to the video content at the candidate sampling position;

[0184] When the content category is a static scene, increase the weight of the high-frequency energy distribution data, and decrease the weight of the motion saliency data;

[0185] When the content category is a dynamic scene, increase the weight of the motion saliency data, and decrease the weight of the high-frequency energy distribution data.

[0186] In a possible implementation manner, the data extraction module is specifically configured to:

[0187] Intercept a rectangular pixel region in each frame of the original video image with the target sampling position as the center;

[0188] Perform time alignment on the rectangular pixel regions corresponding to all the original video images to obtain a set of pixel regions;

[0189] Organize the set of pixel regions into a three-dimensional spatio-temporal data structure to obtain pixel data, where two dimensions in the three-dimensional spatio-temporal data structure represent spatial coordinates, and the third dimension represents a time sequence.

[0190] In a possible implementation, the evaluation module is specifically configured to:

[0191] Determine stripe continuity data and detail retention rate data according to the spatio-temporal mapping graph, where the stripe continuity data is used to characterize the smoothness of video temporal changes, and the detail retention rate data is used to characterize the retention degree of high-frequency spatial features;

[0192] Perform weighted summation on the stripe continuity data and the detail retention rate data to obtain the stability score of the video to be evaluated.

[0193] In a possible implementation, the evaluation module is further configured to:

[0194] Detect texture stripes extending along the time dimension in the spatio-temporal mapping graph;

[0195] Extract local binary pattern feature vectors of the stripe region in consecutive time units;

[0196] Calculate the similarity of local binary pattern feature vectors within adjacent time units;

[0197] Normalize the temporal average value of the similarity to a preset interval to obtain the stripe continuity data.

[0198] In a possible implementation, the evaluation module is further configured to:

[0199] Extract the target gradient amplitude spectrum of the video to be evaluated according to the spatio-temporal mapping graph;

[0200] Calculate the energy ratio of the target gradient amplitude spectrum and a preset standard gradient amplitude spectrum within a preset high-frequency band, where the standard gradient amplitude spectrum is the gradient amplitude statistical benchmark of a stable video sequence under the same shooting conditions;

[0201] Use the energy ratio as the detail retention rate data.

[0202] Based on the same technical concept, an embodiment of the present application further provides an electronic device, as Figure 5 shown, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114,

[0203] The memory 113 is used to store a computer program;

[0204] When the processor 111 is used to execute the program stored on the memory 113, the following steps are implemented:

[0205] Obtain the video to be evaluated, and perform Fourier transform on each frame of the original video image in the video to be evaluated to obtain a frequency-domain video sequence;

[0206] For each spatial position in the frequency-domain video sequence, extract the time-domain sequence of frequency components of the spatial position in the frequency-domain video sequence;

[0207] Determine the target sampling position according to the time-domain sequence of frequency components corresponding to each spatial position;

[0208] Extract pixel data from each frame of the original video image according to the target sampling position;

[0209] Apply a spatio-temporal mapping method to the pixel data to generate a spatio-temporal mapping diagram, and evaluate the stability of the video to be evaluated according to the spatio-temporal mapping diagram.

[0210] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0211] The communication interface is used for communication between the above electronic device and other devices.

[0212] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0213] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0214] In another embodiment provided by the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of any of the above video stability evaluation methods are implemented.

[0215] In another embodiment provided by the present application, a computer program product containing instructions is further provided. When it runs on a computer, the computer is made to execute any of the video stability evaluation methods in the above embodiments.

[0216] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0217] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0218] It should be understood that the terms used in the text are only for the purpose of describing specific example embodiments and are not intended to be restrictive. Unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" used in the text may also include the plural form. The terms "include", "comprise", "contain" and "have" are inclusive and thus specify the presence of the stated features, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described in the text are not to be construed as necessarily requiring them to be executed in the specific order described or illustrated, unless the execution order is clearly stated. It should also be understood that additional or alternative steps can be used.

[0219] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for evaluating video stability, characterized in that, The method includes: Obtain the video to be evaluated, and perform Fourier transform on each frame of the original video image in the video to be evaluated to obtain a frequency-domain video sequence; For each spatial position in the frequency-domain video sequence, extract the time-domain sequence of frequency components of the spatial position in the frequency-domain video sequence; Determine the target sampling positions according to the time-domain sequence of frequency components corresponding to each spatial position; Extract pixel data from each frame of the original video image according to the target sampling positions; Apply a spatio-temporal mapping method to the pixel data to generate a spatio-temporal mapping graph, and evaluate the stability of the video to be evaluated according to the spatio-temporal mapping graph.

2. The method according to claim 1, wherein The determining the target sampling positions according to the time-domain sequence of frequency components corresponding to each spatial position includes: For each spatial position, calculate the high-frequency energy distribution data according to the time-domain sequence of frequency components of the spatial position; Determine the spatial positions where the corresponding high-frequency energy distribution data exceeds a preset energy threshold as candidate sampling positions; For each candidate sampling position, calculate the motion saliency data according to the time-domain sequence of frequency components of the candidate sampling position; Perform a weighted summation operation on the motion saliency data and the high-frequency energy distribution data of the candidate sampling positions to obtain position score data; Sort all candidate sampling positions in descending order according to the corresponding position score data, and determine the top several candidate sampling positions as the target sampling positions.

3. The method according to claim 2, wherein The method further includes: For each candidate sampling position, determine the content category corresponding to the video content at the candidate sampling position; In the case where the content category is a static scene, increase the weight of the high-frequency energy distribution data, and decrease the weight of the motion saliency data; In the case where the content category is a dynamic scene, increase the weight of the motion saliency data, and decrease the weight of the high-frequency energy distribution data.

4. The method according to claim 1, wherein The extracting pixel data from each frame of the original video image according to the target sampling positions includes: Intercept a rectangular pixel region in each frame of the original video image with the target sampling position as the center; Perform time alignment on the rectangular pixel regions corresponding to all the original video images to obtain a set of pixel regions; Organize the set of pixel regions into a three-dimensional spatio-temporal data structure to obtain pixel data, where two dimensions in the three-dimensional spatio-temporal data structure represent spatial coordinates, and the third dimension represents a time sequence.

5. The method according to claim 1, wherein The evaluating the stability of the video to be evaluated according to the spatio-temporal mapping graph includes: Determine streak continuity data and detail retention rate data according to the spatio-temporal mapping graph, where the streak continuity data is used to characterize the smoothness of video temporal changes, and the detail retention rate data is used to characterize the retention degree of high-frequency spatial features; Perform a weighted summation on the streak continuity data and the detail retention rate data to obtain the stability score of the video to be evaluated.

6. The method according to claim 5, characterized in that The determining the streak continuity data according to the spatio-temporal mapping graph includes: Detect texture streaks extending along the time dimension in the spatio-temporal mapping graph; Extract the local binary pattern feature vectors of the streak regions in consecutive time units; Calculate the similarity of the local binary pattern feature vectors within adjacent time units; Normalize the temporal average value of the similarity to a preset interval to obtain the stripe continuity data.

7. The method according to claim 5, characterized in that, The determining the detail retention rate data according to the spatio-temporal mapping graph includes: Extract the target gradient amplitude spectrum of the video to be evaluated from the spatio-temporal mapping graph; Calculate the energy ratio of the target gradient amplitude spectrum and a preset standard gradient amplitude spectrum within a preset high-frequency band, where the standard gradient amplitude spectrum is the gradient amplitude statistical benchmark of a stable video sequence under the same shooting conditions; Use the energy ratio as the detail retention rate data.

8. A video stability evaluation device, characterized in that, The device includes: An acquisition module, configured to acquire a video to be evaluated and perform Fourier transform on each frame of the original video image in the video to be evaluated to obtain a frequency-domain video sequence; A sequence extraction module, configured to extract the time-domain sequence of the frequency components of each spatial position in the frequency-domain video sequence for each spatial position in the frequency-domain video sequence; A determination module, configured to determine the target sampling position according to the time-domain sequence of the frequency components corresponding to each spatial position; A data extraction module, configured to extract pixel data from each frame of the original video image according to the target sampling position; An evaluation module, configured to apply a spatio-temporal mapping method to the pixel data to generate a spatio-temporal mapping graph and evaluate the stability of the video to be evaluated according to the spatio-temporal mapping graph.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is configured to implement the video stability evaluation method according to any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the video stability evaluation method according to any one of claims 1-7.