Video stability assessment method and device, electronic equipment and storage medium

By constructing a two-dimensional spatiotemporal matrix to analyze the time dimension jitter and spatial dimension distortion of videos, the problem of lack of unified space-time and space-time quantitative analysis in the existing video stability evaluation methods is solved, and a comprehensive evaluation of video stability is achieved.

CN120339914APending Publication Date: 2025-07-18BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510453972.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video stability evaluation methods lack unified quantitative analysis standards for space-time dimensions, resulting in one-sided evaluation results.

Method used

By obtaining the continuous frame sequence of the video, selecting pixel bands according to the preset sampling mode, building a two-dimensional space-time matrix, analyzing the time-dimensional jitter and spatial dimension distortion of the video, and generating stability evaluation.

Benefits of technology

A comprehensive quantitative evaluation of video stability is achieved, and the evaluation one-sided problem caused by analyzing only a single dimension in traditional methods is overcome, providing efficient and intuitive stability evaluation methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339914A_ABST
    Figure CN120339914A_ABST
Patent Text Reader

Abstract

The invention provides a video stability evaluation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a continuous frame sequence of a to-be-evaluated video; selecting at least one pixel band from each frame of image of the continuous frame sequence according to a preset sampling mode, wherein the pixel band is a continuous pixel set with a fixed spatial position; extracting pixel values of the pixel bands frame by frame according to a time sequence to form a pixel band time sequence; the pixel band time sequences are spliced in the time axis direction, a two-dimensional space-time matrix is generated, the first dimension of the matrix represents the spatial arrangement sequence of pixel bands, and the second dimension of the matrix represents the time sequence; and evaluating the stability of the video to be evaluated according to the two-dimensional space-time matrix. According to the method, the two-dimensional space-time matrix is constructed, time dimension jitter and space dimension distortion of the video are uniformly represented as texture features of the matrix, and the problem of one-sidedness of evaluation caused by analysis of only one dimension in a traditional method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technologies, and in particular, to a method, apparatus, electronic device, and storage medium for video stability evaluation. Background Art

[0002] With the popularization of video acquisition devices, video stability evaluation has become a key technical requirement in fields such as security monitoring, film and television production, and motion analysis. An ideal video stability evaluation method should have characteristics such as real-time processing ability, strong scene adaptability, and objective and accurate evaluation results.

[0003] Currently, the mainstream video stability evaluation technologies mainly include three categories: methods based on image registration: calculating the displacement amount by aligning adjacent frames; methods based on optical flow: analyzing the pixel motion trajectories; methods based on deep learning: relying on training with large-scale labeled data.

[0004] However, existing methods either only analyze the jitter in the time dimension (such as methods based on image registration) or only focus on spatial distortion (such as methods based on optical flow), lacking a unified quantitative analysis standard for the spatio-temporal dimension, resulting in one-sided evaluation results. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a method, apparatus, electronic device, and storage medium for video stability evaluation to solve the problem that existing methods lack a unified quantitative analysis standard for the spatio-temporal dimension, resulting in one-sided evaluation results. The specific technical solutions are as follows:

[0006] In a first aspect, this application provides a method for video stability evaluation, including:

[0007] Obtaining a continuous frame sequence of the video to be evaluated;

[0008] Selecting at least one pixel band from each frame image of the continuous frame sequence according to a preset sampling pattern, where the pixel band is a set of continuous pixels with fixed spatial positions;

[0009] Extracting the pixel values of the pixel band frame by frame in chronological order to form a pixel band time series;

[0010] Splicing the pixel band time series along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the chronological order;

[0011] Evaluating the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix.

[0012] In a possible implementation manner, the selecting at least one pixel band from each frame image of the continuous frame sequence according to a preset sampling pattern includes:

[0013] When the preset sampling mode is the row sampling mode, at least one row of continuous pixels is selected from each frame image of the continuous frame sequence as a pixel band;

[0014] When the preset sampling mode is the column sampling mode, at least one column of continuous pixels is selected from each frame image of the continuous frame sequence as a pixel band;

[0015] When the preset sampling mode is the multi-position sampling mode, at least one row of continuous pixels and at least one column of continuous pixels are selected from each frame image of the continuous frame sequence as a pixel band.

[0016] In a possible implementation manner, the selecting at least one pixel band from each frame image of the continuous frame sequence according to the preset sampling mode includes:

[0017] When the preset sampling mode is the first detection sampling mode, edge detection is performed on the first frame image of the continuous frame sequence to identify the high-frequency feature region in the first frame image;

[0018] Calculate the geometric center coordinates of the high-frequency feature region, and use the geometric center coordinates as a reference to extend a preset length along the horizontal direction or the vertical direction to obtain a straight line segment;

[0019] Select the continuous pixels with the same spatial coordinates as the straight line segment from each frame image of the continuous frame sequence as the pixel band.

[0020] In a possible implementation manner, the selecting at least one pixel band from each frame image of the continuous frame sequence according to the preset sampling mode includes:

[0021] When the preset sampling mode is the second detection sampling mode, for each frame image of the continuous frame sequence, edge detection is performed on the image to identify the high-frequency feature region in the image;

[0022] Calculate the centroid coordinates of the high-frequency feature region in the image, and generate candidate straight line segments with the centroid coordinates as reference points;

[0023] Calculate the spatial consistency scores of the candidate straight line segments and the pixel band of the previous frame image;

[0024] Select the continuous pixels at the position of the candidate straight line segment with the highest spatial consistency score as the pixel band of the image.

[0025] In a possible implementation manner, the evaluating the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix includes:

[0026] Convert the two-dimensional spatio-temporal matrix into a frequency-domain representation to obtain a two-dimensional frequency-domain distribution including spatial frequency and temporal frequency;

[0027] Detect the high-frequency components and periodic components in the two-dimensional frequency domain distribution;

[0028] When the energy proportion of the high-frequency components exceeds a first threshold, it is determined that there is high-frequency jitter in the video to be evaluated;

[0029] When the energy concentration of the periodic components in a preset frequency band exceeds a second threshold, it is determined that there is regular fluctuation in the video to be evaluated.

[0030] In a possible implementation manner, the evaluating the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix includes:

[0031] Perform element-by-element difference calculation on the pixel values of adjacent time points in the two-dimensional spatio-temporal matrix to obtain a difference matrix;

[0032] Calculate the change intensity data of the difference matrix, where the change intensity data is used to characterize the cumulative effect of pixel value mutations in the difference matrix;

[0033] When the change intensity data exceeds a preset threshold, determine the video frame corresponding to the time point as an unstable frame.

[0034] In a possible implementation manner, the method further includes:

[0035] For each unstable frame, extract the difference matrix over-standard region corresponding to the unstable frame;

[0036] Merge the over-standard regions with adjacent spatio-temporal positions into an abnormal event set;

[0037] Calculate the spatio-temporal distribution density of the abnormal event set, where the density is jointly characterized by the number of abnormal events per unit time and the spatial area covered by the abnormal events;

[0038] Generate a stability rating for the video to be evaluated according to the spatio-temporal distribution density.

[0039] In a second aspect, the present application provides a video stability evaluation device, including:

[0040] An acquisition module, configured to acquire a continuous frame sequence of a video to be evaluated;

[0041] A selection module, configured to select at least one pixel band from each frame image of the continuous frame sequence according to a preset sampling mode, where the pixel band is a set of continuous pixels with fixed spatial positions;

[0042] An extraction module, configured to extract the pixel values of the pixel band frame by frame in chronological order to form a pixel band time series;

[0043] A splicing module, which is used to splice the pixel band time series along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the time order;

[0044] An evaluation module, which is used to evaluate the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix.

[0045] In a possible implementation manner, the selection module is specifically used for:

[0046] In the case where the preset sampling mode is the row sampling mode, at least one row of continuous pixels is selected from each frame of the continuous frame sequence as the pixel band;

[0047] In the case where the preset sampling mode is the column sampling mode, at least one column of continuous pixels is selected from each frame of the continuous frame sequence as the pixel band;

[0048] In the case where the preset sampling mode is the multi-position sampling mode, at least one row of continuous pixels and at least one column of continuous pixels are selected from each frame of the continuous frame sequence as the pixel band.

[0049] In a possible implementation manner, the selection module is further used for:

[0050] In the case where the preset sampling mode is the first detection sampling mode, edge detection is performed on the first frame image of the continuous frame sequence to identify the high-frequency feature region in the first frame image;

[0051] Calculate the geometric center coordinates of the high-frequency feature region, and take the geometric center coordinates as a reference, and extend a preset length along the horizontal direction or the vertical direction to obtain a straight line segment;

[0052] Select the continuous pixels with the same spatial coordinates as the straight line segment from each frame of the continuous frame sequence as the pixel band.

[0053] In a possible implementation manner, the selection module is further used for:

[0054] In the case where the preset sampling mode is the second detection sampling mode, for each frame image of the continuous frame sequence, edge detection is performed on the image to identify the high-frequency feature region in the image;

[0055] Calculate the centroid coordinates of the high-frequency feature region in the image, and generate a candidate straight line segment with the centroid coordinates as the reference point;

[0056] Calculate the spatial consistency score between each candidate straight line segment and the pixel band of the previous frame image;

[0057] Select the continuous pixels at the position of the candidate straight line segment with the highest spatial consistency score as the pixel band of the image.

[0058] In a possible implementation, the evaluation module is specifically configured to:

[0059] Convert the two-dimensional spatio-temporal matrix into a frequency-domain representation to obtain a two-dimensional frequency-domain distribution including spatial frequency and temporal frequency;

[0060] Detect high-frequency components and periodic components in the two-dimensional frequency-domain distribution;

[0061] When the energy ratio of the high-frequency components exceeds a first threshold, it is determined that there is high-frequency jitter in the video to be evaluated;

[0062] When the energy concentration of the periodic components in a preset frequency band exceeds a second threshold, it is determined that there is regular fluctuation in the video to be evaluated.

[0063] In a possible implementation, the evaluation module is further configured to:

[0064] Perform element-by-element difference calculation on the pixel values at adjacent time points in the two-dimensional spatio-temporal matrix to obtain a difference matrix;

[0065] Calculate the change intensity data of the difference matrix, where the change intensity data is used to characterize the cumulative effect of pixel value mutations in the difference matrix;

[0066] When the change intensity data exceeds a preset threshold, determine the video frame corresponding to the corresponding time point as an unstable frame.

[0067] In a possible implementation, the evaluation module is further configured to:

[0068] For each unstable frame, extract the difference matrix exceeding-standard region corresponding to the unstable frame;

[0069] Merge the exceeding-standard regions with adjacent spatio-temporal positions into an abnormal event set;

[0070] Calculate the spatio-temporal distribution density of the abnormal event set, where the density is jointly characterized by the number of abnormal events per unit time and the spatial area covered by the abnormal events;

[0071] Generate a stability rating for the video to be evaluated according to the spatio-temporal distribution density.

[0072] In a third aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0073] The memory is used to store a computer program;

[0074] A processor, when executing a program stored in a memory, implements the method steps described in any one of the first aspects.

[0075] In a fourth aspect, a computer-readable storage medium is provided, characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method steps described in any one of the first aspects are implemented.

[0076] In a fifth aspect, a computer program product containing instructions is provided, which when running on a computer causes the computer to execute the video stability evaluation method described above.

[0077] Advantages of the embodiments of the present application:

[0078] The embodiments of the present application provide a video stability evaluation method, apparatus, electronic device and storage medium. In the embodiments of the present application, first, a continuous frame sequence of a video to be evaluated is obtained. Then, at least one pixel band is selected from each frame image of the continuous frame sequence according to a preset sampling pattern, and the pixel band is a set of continuous pixels with fixed spatial positions. Furthermore, pixel values of the pixel band are extracted frame by frame in chronological order to form a pixel band time series, and the pixel band time series is spliced along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the time order. Finally, the stability of the video to be evaluated is evaluated according to the two-dimensional spatio-temporal matrix. By constructing the two-dimensional spatio-temporal matrix, the present application uniformly characterizes the time dimension jitter (such as camera shake) and spatial dimension distortion (such as motion blur) of the video as the texture features of the matrix, overcoming the problem of one-sided evaluation caused by traditional methods that only analyze a single dimension (time or space).

[0079] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. Description of the Drawings

[0080] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present invention and used together with the description to explain the principles of the present invention.

[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0082] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a scale limitation.

[0083] Figure 1 It is a flowchart of a video stability evaluation method provided by an embodiment of the present application;

[0084] Figure 2 It is a flowchart of another video stability evaluation method provided by an embodiment of the present application;

[0085] Figure 3 It is a flowchart of yet another video stability evaluation method provided by an embodiment of the present application;

[0086] Figure 4 It is a schematic structural diagram of a video stability evaluation device provided by an embodiment of the present application;

[0087] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0088] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0089] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0090] Figure 1A schematic flowchart of a video stability evaluation method provided by an embodiment of this application. This method can be applied to one or more electronic devices such as smart phones, laptop computers, desktop computers, portable computers, servers, etc. In addition, the execution subject of this method can be hardware or software. When the above execution subject is hardware, the execution subject can be one or more of the above electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above execution subject is software, this method can be implemented as multiple software or software modules, or can be implemented as a single software or software module. No specific limitation is made here.

[0091] As Figure 1 shown, the method specifically includes:

[0092] S101. Obtain a continuous frame sequence of the video to be evaluated.

[0093] The video to be evaluated refers to a video that needs to be evaluated. For example, a movie video, a TV drama video, a surveillance video, etc.

[0094] The continuous frame sequence is a set of video images arranged in chronological order, and the time interval between adjacent frames is fixed.

[0095] In one embodiment, the video to be evaluated can be obtained by receiving a video input by a user.

[0096] In another embodiment, each video can be sequentially obtained from a preset video set as a video to be evaluated.

[0097] S102. Select at least one pixel band from each frame image of the continuous frame sequence according to a preset sampling mode, where the pixel band is a set of continuous pixels with fixed spatial positions.

[0098] The pixel band refers to a linear set of pixels with unchanged spatial coordinates in a video frame.

[0099] The preset sampling mode includes a row sampling mode, a column sampling mode, a multi-position sampling mode, a first detection sampling mode, and a second detection sampling mode. In applications, users can set specific sampling modes according to actual needs.

[0100] In one embodiment, S102 can specifically include the following steps: When the preset sampling mode is the row sampling mode, select at least one row of continuous pixels from each frame image of the continuous frame sequence as the pixel band; when the preset sampling mode is the column sampling mode, select at least one column of continuous pixels from each frame image of the continuous frame sequence as the pixel band; when the preset sampling mode is the multi-position sampling mode, select at least one row of continuous pixels and at least one column of continuous pixels from each frame image of the continuous frame sequence as the pixel band.

[0101] Among them, the execution process of the row sampling mode is as follows: Determine the sampling row index r, the default value is the middle row of the video frame (i.e., floor(H / 2), in applications, it can also be other rows selected by the user), traverse each column index w from 1 to W, and add the coordinates (r, w) to the set P: P = P ∪ {(r, w) | w = 1, 2, …, W}, that is, the pixel band.

[0102] The execution process of the column sampling mode is as follows: Determine the sampling column index c, the default value is the middle column of the video frame (i.e., floor(W / 2), in applications, it can also be other columns selected by the user), traverse each row index h from 1 to H, and add the coordinates (h, c) to the set P: P = P ∪ {(h, c) | h = 1, 2, …, H}, that is, the pixel band.

[0103] The execution process of the multi-position sampling mode is as follows: Determine the row sampling set R = {r1, r2, …, r m and the column sampling set C = {c1, c2, …, c n , for each row index r in the set R, traverse each column index w from 1 to W, and add the coordinates (r, w) to the set P: P = P ∪ {(r, w) | r ∈ R, w = 1, 2, …, W}; for each column index c in the set C, traverse each row index h from 1 to H, and add the coordinates (h, c) to the set P: P = P ∪ {(h, c) | c ∈ C, h = 1, 2, …, H}, that is, the pixel band.

[0104] Through this solution, the sampling mode can be flexibly set according to user requirements to extract the pixel band according to user needs.

[0105] In another embodiment, S102 may further include the following steps: When the preset sampling mode is the first detection sampling mode, perform edge detection on the first frame image of the continuous frame sequence to identify the high-frequency feature region in the first frame image; calculate the geometric center coordinates of the high-frequency feature region, and extend a preset length along the horizontal direction or the vertical direction with the geometric center coordinates as the reference to obtain a straight line segment; select continuous pixels with exactly the same spatial coordinates as the straight line segment from each frame image of the continuous frame sequence as the pixel band.

[0106] In this embodiment, the first detection sampling mode determines the optimal sampling position by analyzing the feature distribution of the first frame image of the video: First, perform edge detection on the first frame to identify key regions with rich textures, and calculate the geometric center of this region as the reference point; taking this reference point as the center, extend a fixed length along the horizontal or vertical direction to generate a sampling line segment, and the coordinate parameters of this line segment will be strictly fixed in all subsequent frames. This mode adaptively determines the sampling position based on the features of the first frame, which can not only specifically capture the key motion features in the video, but also ensure the spatio-temporal consistency of subsequent analysis through coordinate locking, and is particularly suitable for scenarios that require long-term stable observation (such as fixed surveillance cameras). When the scene changes significantly, the system can trigger re-initialization and perform a new round of feature detection and sampling position calculation.

[0107] In yet another embodiment, S102 may further include the following steps: When the preset sampling mode is the second detection sampling mode, for each frame image of the continuous frame sequence, perform edge detection on the image to identify the high-frequency feature region in the image; calculate the centroid coordinates of the high-frequency feature region in the image, and generate a candidate line segment with the centroid coordinates as the reference point; calculate the spatial consistency score between each candidate line segment and the pixel band of the previous frame image; select the continuous pixels at the position of the candidate line segment with the highest spatial consistency score as the pixel band of the image.

[0108] In this embodiment, the second detection sampling mode realizes adaptive sampling by adopting a dynamic tracking strategy: perform edge detection on each frame in real time and locate the high-frequency feature region, calculate the centroid of the feature region as the dynamic reference point, and generate multiple candidate sampling paths (horizontal / vertical / diagonal) based on this point; by evaluating the geometric consistency (including position offset and angle change) between the candidate paths and the sampling position of the previous frame, select the path with the highest matching degree as the pixel band of the current frame. This mode adapts to the change of scene content while maintaining sampling continuity through inter-frame correlation constraints and dynamic path optimization, and is particularly suitable for scenarios with complex background motion such as action cameras and in-vehicle videos, and can effectively avoid the problem of feature loss caused by fixed sampling positions.

[0109] S103. Extract the pixel values of the pixel band frame by frame in chronological order to form a pixel band time series.

[0110] In the embodiment of the present application, the specific process of extracting the pixel band for analysis from the continuous frame sequence is as follows:

[0111] Input: Video frame sequence: Video data composed of multiple frames I1, I2, I3,..., I T Composition: Fixed position coordinate set: Fixed position coordinate set P obtained through step S102.

[0112] Execution process:

[0113] (1) Initialize an empty set of pixel bands {B1, B2, B3, ..., B T};

[0114] (2) For each frame I at time point t t Extract the pixel band B at a fixed position t , which contains all the pixels on the coordinate P, B t = I t (i, j) ∣ (i, j) ∈ P.

[0115] Output: The set of pixel bands B1, B2, B3, ..., B T , that is, the time series of pixel bands.

[0116] S104. Concatenate the time series of pixel bands along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the time order.

[0117] In the embodiments of the present application, the specific process of concatenating the time series of pixel bands along the time axis direction to generate a two-dimensional spatio-temporal matrix is as follows:

[0118] Input: The set of pixel bands: containing the pixel bands B1, B2, B3, ..., B of each frame T .

[0119] Process:

[0120] (1) Initialize an empty two-dimensional matrix M with a size of H x T or W x T;

[0121] (2) Concatenate the pixel bands B in each frame of video data t into the matrix M in chronological order:

[0122] Vertical sampling situation: [M(i, t) = I t (i, k) i = 1, 2, …, H; t = 1, 2, …, T];

[0123] Horizontal sampling situation: [M(y, t) = I t (k, y) y = 1, 2, …, W; t = 1, 2, …, T];

[0124] Output: The time-space mapping matrix M, that is, the two-dimensional spatio-temporal matrix.

[0125] S105. Evaluate the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix.

[0126] In the embodiments of the present application, the quantitative evaluation of the video stability is realized by analyzing the texture features of the two-dimensional spatio-temporal matrix.

[0127] As for how to specifically evaluate the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix, it will be described in detail through the following embodiments and will not be elaborated here first.

[0128] Through the above steps, the present application provides a Fixed-Position Temporal Stacking (FPTS) method, which provides an efficient and intuitive means for evaluating stability by analyzing the spatio-temporal changes at fixed positions in the video.

[0129] In the embodiment of the present application, first, a continuous frame sequence of the video to be evaluated is obtained. Then, at least one pixel band is selected from each frame image of the continuous frame sequence according to a preset sampling pattern. The pixel band is a set of continuous pixels with fixed spatial positions. Furthermore, the pixel values of the pixel band are extracted frame by frame in time sequence to form a pixel band time series, and the pixel band time series is spliced along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the time order. Finally, the stability of the video to be evaluated is evaluated according to the two-dimensional spatio-temporal matrix. By constructing the two-dimensional spatio-temporal matrix, the present application uniformly represents the time dimension jitter (such as camera shake) and the spatial dimension distortion (such as motion blur) of the video as the texture features of the matrix, overcoming the problem of one-sided evaluation caused by the traditional method of only analyzing a single dimension (time or space).

[0130] See Figure 2 , which is a flowchart of an embodiment of another method for evaluating video stability provided by the embodiment of the present application. The Figure 2 shown process is based on the process shown above Figure 1 and describes how to evaluate the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix. As Figure 2 shown, the process may include the following steps:

[0131] S201. Convert the two-dimensional spatio-temporal matrix into a frequency domain representation to obtain a two-dimensional frequency domain distribution including spatial frequency and time frequency;

[0132] S202. Detect the high-frequency components and periodic components in the two-dimensional frequency domain distribution;

[0133] S203. When the energy ratio of the high-frequency components exceeds a first threshold, it is determined that there is high-frequency jitter in the video to be evaluated;

[0134] S204. When the energy concentration of the periodic components in a preset frequency band exceeds a second threshold, it is determined that there is regular fluctuation in the video to be evaluated.

[0135] For ease of understanding, the following provides a unified description of steps S201 - S204:

[0136] In the embodiments of the present application, first, a two - dimensional Fourier transform is performed on the two - dimensional spatio - temporal matrix to generate an energy spectrum distribution (i.e., two - dimensional frequency - domain distribution) that simultaneously contains spatial and temporal frequencies; then, the high - frequency components (reflecting sudden jitter) and periodic energy aggregations in specific frequency bands (reflecting regular fluctuations) in the energy spectrum are detected respectively; when the high - frequency energy exceeds the first threshold, it is determined that there is high - frequency vibration of the device (such as camera jitter), and when the energy concentration in the periodic frequency band exceeds the standard, it is determined that there is environmental regular perturbation (such as engine vibration of a vehicle - mounted camera).

[0137] The specific execution process is as follows:

[0138] Input: Time - space mapping matrix M.

[0139] Process:

[0140] (1) Perform Fourier transform on the time - space mapping matrix M to analyze the frequency - domain information:

[0141]

[0142] (2) Calculate the spectral amplitude:

[0143]

[0144] (3) By calculating the spectral amplitude, it can be determined which frequency components are significant, that is, the frequency components with larger amplitudes Threshold = α·max(|F trans (u,v)|), where α is an empirical value (such as 0.1) used to determine the threshold of significant frequency components.

[0145] Output: Video stability analysis result.

[0146] Through this solution, different types of unstable factors can be effectively distinguished, providing a basis for targeted image stabilization processing, and is particularly applicable to industrial inspection and motion analysis scenarios that require identifying the source of jitter.

[0147] See Figure 3 , which is a flowchart of an embodiment of another video stability evaluation method provided by the embodiments of the present application. The Figure 3 shown process, based on the process shown above Figure 1 describes how to evaluate the stability of the video to be evaluated according to the two - dimensional spatio - temporal matrix. As Figure 3 shown, the process may include the following steps:

[0148] S301. Perform element - by - element difference calculation on the pixel values of adjacent time points in the two - dimensional spatio - temporal matrix to obtain a difference matrix;

[0149] S302. Calculate the change intensity data of the difference matrix, where the change intensity data is used to characterize the cumulative effect of pixel value mutations in the difference matrix;

[0150] S303. When the change intensity data exceeds a preset threshold, determine the video frame corresponding to the time point as an unstable frame.

[0151] For ease of understanding, the following provides a unified description of steps S301 - S303:

[0152] This embodiment realizes real-time detection of video stability through temporal difference analysis. Specifically: First, calculate the pixel-level difference values of adjacent time points in the two-dimensional spatio-temporal matrix to generate a difference matrix reflecting inter-frame mutations; then, perform statistical quantization on the difference matrix (such as calculating the absolute value mean or the proportion of out-of-limit pixels) to obtain the change intensity index data characterizing the overall disturbance degree; when the change intensity index data exceeds the preset threshold, determine that the corresponding video frame has significant jitter and mark it as an unstable frame.

[0153] Through this solution, sudden unstable frames can be captured in real time, which is especially suitable for embedded devices with limited computing power or monitoring scenarios that require low-latency processing.

[0154] In addition, in another embodiment of the present application, the method may further include the following steps: For each unstable frame, extract the out-of-limit region of the difference matrix corresponding to the unstable frame; merge the out-of-limit regions with adjacent spatio-temporal positions into an abnormal event set; calculate the spatio-temporal distribution density of the abnormal event set, where the density is jointly characterized by the number of abnormal events per unit time and the spatial area covered by the abnormal events; generate the stability rating of the video to be evaluated according to the spatio-temporal distribution density.

[0155] This embodiment realizes refined evaluation of video stability through abnormal event aggregation analysis. Specifically: First, extract the pixel regions exceeding the threshold from the difference matrix of each unstable frame as abnormal regions, and then cluster these discrete abnormal regions into a complete abnormal event set according to spatio-temporal continuity (spatial overlap or temporal proximity); calculate the comprehensive spatio-temporal distribution density index by statistically analyzing the occurrence frequency of abnormal events per unit time and their spatial coverage area; finally, divide the stability rating of the video (such as A / B / C level) based on the intensity interval of this density index.

[0156] In this solution, by aggregating analysis, the frame-level jitter detection is upgraded to event-level stability evaluation, which can not only identify short-term sudden jitters but also capture persistent unstable phenomena. It is especially suitable for film and television post-production and video monitoring quality control scenarios that require quantitative evaluation of the overall video quality, and its density index can intuitively reflect the severity and spatial influence range of instability.

[0157] Based on the same inventive concept, an embodiment of the present application further provides a video stability evaluation device, as Figure 4 shown, the device includes:

[0158] An acquisition module 41, configured to acquire a continuous frame sequence of a video to be evaluated;

[0159] A selection module 42, configured to select at least one pixel band from each frame image of the continuous frame sequence according to a preset sampling mode, where the pixel band is a set of continuous pixels with fixed spatial positions;

[0160] An extraction module 43, configured to sequentially extract pixel values of the pixel band frame by frame in time order to form a pixel band time series;

[0161] A splicing module 44, configured to splice the pixel band time series along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the time order;

[0162] An evaluation module 45, configured to evaluate the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix.

[0163] In a possible implementation manner, the selection module is specifically configured to:

[0164] When the preset sampling mode is a row sampling mode, select at least one row of continuous pixels from each frame image of the continuous frame sequence as the pixel band;

[0165] When the preset sampling mode is a column sampling mode, select at least one column of continuous pixels from each frame image of the continuous frame sequence as the pixel band;

[0166] When the preset sampling mode is a multi-position sampling mode, select at least one row of continuous pixels and at least one column of continuous pixels from each frame image of the continuous frame sequence as the pixel band.

[0167] In a possible implementation manner, the selection module is further configured to:

[0168] When the preset sampling mode is a first detection sampling mode, perform edge detection on the first frame image of the continuous frame sequence to identify a high-frequency feature region in the first frame image;

[0169] Calculate the geometric center coordinates of the high-frequency feature region, and use the geometric center coordinates as a reference to extend a preset length along the horizontal direction or the vertical direction to obtain a straight line segment;

[0170] Select continuous pixels with exactly the same spatial coordinates as the straight line segment from each frame image of the continuous frame sequence as the pixel band.

[0171] In a possible implementation, the selection module is further configured to:

[0172] When the preset sampling mode is the second detection sampling mode, for each frame image of the continuous frame sequence, perform edge detection on the image to identify the high-frequency feature region in the image;

[0173] Calculate the centroid coordinates of the high-frequency feature region in the image, and generate candidate line segments based on the centroid coordinates;

[0174] Calculate the spatial consistency score between each candidate line segment and the pixel band of the previous frame image;

[0175] Select the continuous pixels at the position of the candidate line segment with the highest spatial consistency score as the pixel band of the image.

[0176] In a possible implementation, the evaluation module is specifically configured to:

[0177] Convert the two-dimensional spatio-temporal matrix into a frequency-domain representation to obtain a two-dimensional frequency-domain distribution including spatial frequency and temporal frequency;

[0178] Detect the high-frequency components and periodic components in the two-dimensional frequency-domain distribution;

[0179] When the energy proportion of the high-frequency components exceeds the first threshold, determine that the video to be evaluated has high-frequency jitter;

[0180] When the energy concentration of the periodic components in the preset frequency band exceeds the second threshold, determine that the video to be evaluated has regular fluctuations.

[0181] In a possible implementation, the evaluation module is further configured to:

[0182] Perform element-by-element difference calculation on the pixel values at adjacent time points in the two-dimensional spatio-temporal matrix to obtain a difference matrix;

[0183] Calculate the change intensity data of the difference matrix, where the change intensity data is used to characterize the cumulative effect of pixel value mutations in the difference matrix;

[0184] When the change intensity data exceeds the preset threshold, determine the video frame corresponding to the corresponding time point as an unstable frame.

[0185] In a possible implementation, the evaluation module is further configured to:

[0186] For each unstable frame, extract the difference matrix exceeding-standard region corresponding to the unstable frame;

[0187] Merge the exceeding-standard regions with adjacent spatio-temporal positions into an abnormal event set;

[0188] Calculate the spatio-temporal distribution density of the set of abnormal events, where the density is jointly characterized by the number of abnormal events per unit time and the spatial area covered by the abnormal events;

[0189] Generate a stability rating for the video to be evaluated based on the spatio-temporal distribution density.

[0190] Based on the same technical concept, an embodiment of the present application also provides an electronic device, as Figure 5 shown, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114.

[0191] The memory 113 is used to store computer programs;

[0192] When the processor 111 is used to execute the program stored on the memory 113, the following steps are implemented:

[0193] Obtain a continuous frame sequence of the video to be evaluated;

[0194] Select at least one pixel band from each frame image of the continuous frame sequence according to a preset sampling mode, where the pixel band is a set of continuous pixels with fixed spatial positions;

[0195] Extract the pixel values of the pixel band frame by frame in chronological order to form a pixel band time series;

[0196] Concatenate the pixel band time series along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the chronological order;

[0197] Evaluate the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix.

[0198] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.

[0199] The communication interface is used for communication between the above electronic device and other devices.

[0200] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0201] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0202] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above video stability evaluation methods are implemented.

[0203] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute any of the video stability evaluation methods in the above embodiments.

[0204] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0206] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an", and "the" may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or their combinations. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0207] The above description is only the specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for evaluating video stability, characterized in that, The method includes: Obtaining a sequence of consecutive frames of the video to be evaluated; Selecting at least one pixel band from each frame image of the consecutive frame sequence according to a preset sampling pattern, where the pixel band is a set of consecutive pixels with fixed spatial positions; Sequentially extracting pixel values of the pixel band frame by frame in time order to form a pixel band time series; Stitching the pixel band time series along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the time order; Evaluating the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix.

2. The method according to claim 1, characterized in that The step of selecting at least one pixel band from each frame image of the consecutive frame sequence according to a preset sampling pattern includes: When the preset sampling pattern is a row sampling pattern, selecting at least one row of consecutive pixels from each frame image of the consecutive frame sequence as the pixel band; When the preset sampling pattern is a column sampling pattern, selecting at least one column of consecutive pixels from each frame image of the consecutive frame sequence as the pixel band; When the preset sampling pattern is a multi-position sampling pattern, selecting at least one row of consecutive pixels and at least one column of consecutive pixels from each frame image of the consecutive frame sequence as the pixel band.

3. The method according to claim 1, wherein The step of selecting at least one pixel band from each frame image of the consecutive frame sequence according to a preset sampling pattern includes: When the preset sampling pattern is a first detection sampling pattern, performing edge detection on the first frame image of the consecutive frame sequence to identify the high-frequency feature region in the first frame image; Calculating the geometric center coordinates of the high-frequency feature region, and extending a preset length along the horizontal or vertical direction with the geometric center coordinates as the reference to obtain a straight line segment; Selecting the consecutive pixels with exactly the same spatial coordinates as the straight line segment from each frame image of the consecutive frame sequence as the pixel band.

4. The method according to claim 1, characterized in that The step of selecting at least one pixel band from each frame image of the consecutive frame sequence according to a preset sampling pattern includes: When the preset sampling pattern is a second detection sampling pattern, for each frame image of the consecutive frame sequence, performing edge detection on the image to identify the high-frequency feature region in the image; Calculating the centroid coordinates of the high-frequency feature region in the image, and generating candidate straight line segments with the centroid coordinates as the reference points; Calculating the spatial consistency scores between each candidate straight line segment and the pixel band of the previous frame image; Selecting the consecutive pixels at the position of the candidate straight line segment with the highest spatial consistency score as the pixel band of the image.

5. The method according to claim 1, characterized in that The step of evaluating the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix includes: Converting the two-dimensional spatio-temporal matrix into a frequency domain representation to obtain a two-dimensional frequency domain distribution including spatial frequency and time frequency; Detecting high-frequency components and periodic components in the two-dimensional frequency domain distribution; When the energy proportion of the high-frequency components exceeds a first threshold, determining that there is high-frequency jitter in the video to be evaluated; When the energy concentration of the periodic components in a preset frequency band exceeds a second threshold, determining that there is regular fluctuation in the video to be evaluated.

6. The method according to claim 1, wherein The step of evaluating the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix includes: Perform element-by-element difference calculation on the pixel values of adjacent time points in the two-dimensional spatio-temporal matrix to obtain a difference matrix; Calculate the change intensity data of the difference matrix, where the change intensity data is used to characterize the cumulative effect of pixel value mutations in the difference matrix; When the change intensity data exceeds a preset threshold, determine the video frame corresponding to the time point as an unstable frame.

7. The method according to claim 6, wherein The method further includes: For each unstable frame, extract the difference matrix exceeding-standard region corresponding to the unstable frame; Merge the exceeding-standard regions with adjacent spatio-temporal positions into an abnormal event set; Calculate the spatio-temporal distribution density of the abnormal event set, where the density is jointly characterized by the number of abnormal events per unit time and the spatial area covered by the abnormal events; Generate a stability rating for the video to be evaluated according to the spatio-temporal distribution density.

8. A video stability evaluation device, characterized in that, The device includes: An acquisition module, configured to acquire a continuous frame sequence of the video to be evaluated; A selection module, configured to select at least one pixel band from each frame image of the continuous frame sequence according to a preset sampling pattern, where the pixel band is a set of continuous pixels with fixed spatial positions; An extraction module, configured to extract the pixel values of the pixel band frame by frame in time order to form a pixel band time series; A splicing module, configured to splice the pixel band time series along the time axis direction to generate a two-dimensional spatio-temporal matrix, where the first dimension of the matrix represents the spatial arrangement order of the pixel bands, and the second dimension represents the time order; An evaluation module, configured to evaluate the stability of the video to be evaluated according to the two-dimensional spatio-temporal matrix.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; When the processor is configured to execute the program stored on the memory, it implements the video stability evaluation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the video stability evaluation method according to any one of claims 1-7.

Citation Information

Cited By

  • Video fingerprint processing method and device, electronic equipment and storage medium

    CN121214294A