A video cropping method and a quality evaluation method for cropped video
Through the improved automatic video cropping model and quality evaluation method, the existing video cropping methods are solved, and efficient and automatic video cropping and quality evaluation are achieved, improving the effect and quality of video cropping.
Patent Information
- Application Number
- CN202411886091.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The existing video cropping methods are inefficient, difficult to meet the needs of large-scale video data processing, and difficult to ensure the content integrity and timing stability of the cropped video, affecting the user's viewing experience.
Using an improved video automatic cropping model, an intermediate result significant graph is generated through the significance area prediction module, the center position and size of the cropping box are predicted, and quality evaluation is performed through content integrity, content consistency and timing stability.
It realizes automatic determination of the optimal crop position, improves the efficiency and quality of video cropping, ensures the content integrity and timing stability of the cropped video, and improves the user's viewing experience.
Smart Images

Figure CN119342207B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and video processing, and in particular to a video cropping method and a quality evaluation method of cropped videos. Background Art
[0002] Current video cropping methods mainly rely on manual operations or fixed cropping areas, which is not only inefficient but also difficult to meet the needs of large-scale video data processing. Especially when processing a large video library, manual cropping is particularly time-consuming and labor-intensive, and cannot complete the task efficiently. In addition, traditional methods are difficult to ensure the content integrity of the cropped video. Manual cropping often results in the loss of important content or incomplete coverage of significant areas, which ultimately affects the viewing experience of the video and the accuracy of information transmission.
[0003] The existing technology has difficulty in effectively evaluating and maintaining the temporal coherence between video frames during the video cropping process, resulting in screen jumps and frequent target switching in the cropped video, which seriously affects the user's viewing experience. At the same time, the lack of a unified and effective evaluation standard to evaluate the quality of the cropped video leads to the strong subjectivity and lack of comparability of the evaluation results, which further restricts the development of video processing technology. In addition, the existing technology has a low degree of automation in video cropping and quality assessment, and still relies on manual setting of cropping parameters and manual quality assessment. This not only increases the complexity of the operation, but also makes it difficult to ensure the objectivity and consistency of the evaluation results. In the face of these problems, there is an urgent need for new methods that can automatically determine the optimal cropping position and comprehensively evaluate the quality of the cropped video through quantitative evaluation methods, thereby providing a comprehensive solution to significantly improve the effect and quality of video cropping. Summary of the invention
[0004] The purpose of the present invention is to provide a video cropping method and a method for evaluating the quality of cropped videos. First, the original video data and the aspect ratio of the video to be cropped are used as input parameters, and a video automatic cropping model is used to obtain the cropped video, and the intermediate result saliency map and cropping frame information are output at the same time; secondly, the quality of the cropped video is evaluated by three aspects: content integrity, content consistency, and temporal stability; finally, the weighted summation method is used to comprehensively evaluate the content integrity, temporal stability, and content consistency scores to obtain an overall quality evaluation of the cropped video. The effect and quality of video cropping are improved. The present invention is achieved through the following technical solutions.
[0005] In a first aspect, the present invention provides a video cropping method, comprising:
[0006] The salient region prediction module of the video automatic cropping model is used to predict the salient region of the original video and generate an intermediate salient map. ;
[0007] Saliency map of intermediate results Mapping is performed to predict the center position of the cropping box ;
[0008] According to the aspect ratio of the original video Calculate the size of the crop box ; and The width and height of the original video respectively;
[0009] According to the center position of the cropping frame and the size of the crop box Calculate the coordinates of the lower left corner of the cropping box and the upper right corner coordinates ;
[0010] According to the center position of the cropping frame , the size of the cropping frame , the lower left corner coordinates of the cropping box and the upper right corner coordinates Output the cropped video;
[0011] Wherein, the video automatic cropping model includes a salient region prediction module and an output module connected in sequence;
[0012] The salient region prediction module in the automatic video cropping model is a module obtained by improving the prior knowledge modeling module in the original UAVSa model. The specific improvement is to remove the observation prior module and the environmental semantic prior module from the original UAVSa model and only retain the Gaussian prior module. The Gaussian prior module's modeling ability for the center bias phenomenon is used to effectively extract salient regions and complete the prediction of salient regions.
[0013] In practical applications, since traditional video cropping methods mainly rely on manual operations or fixed cropping areas, they are inefficient and difficult to meet the needs of large-scale video data processing. The present invention provides a new video cropping method, which improves the original UAVSa model and designs a video automatic cropping model. The model is used to predict saliency regions and cropping frames, and the original video data and the aspect ratio of the video to be cropped are used as input parameters. The cropping direction and size are determined according to the cropping aspect ratio, and the intermediate result saliency map and cropping frame are output. The new video cropping method proposed in the present invention not only solves the problems of low efficiency and incomplete content in the existing video cropping technology, but also the automatically determined optimal cropping position provides a basis for subsequent video quality evaluation methods.
[0014] Optionally, the intermediate result saliency map Mapping is performed to predict the center position of the cropping box include:
[0015] The intermediate result saliency map Mapped to a value through multiple convolutional layers in the output module of the video automatic cropping model , and use the sigmoid function to convert the value Mapping to clipping position coefficients , the clipping position coefficient Calculated by the following formula:
[0016] ,
[0017] The crop position coefficient is set according to the crop direction Multiply it by the width w or height h of the original video to get the center position of the cropping frame , the center position of the cropping frame Calculated by the following formula:
[0018] ,
[0019] In the formula, The aspect ratio of the preset video cropping frame is based on the aspect ratio of the preset video cropping frame. Determine the crop direction. If the original video has an aspect ratio of Aspect ratio larger than the preset video cropping frame , the cropping box is cropped in the horizontal direction and the vertical height remains unchanged, otherwise the cropping box is cropped in the vertical direction and the horizontal width remains unchanged.
[0020] Optionally, the size of the cropping frame is the crop length of the original video in the horizontal or vertical direction, the size of the crop frame Calculated by the following formula:
[0021] .
[0022] Optionally, the lower left corner coordinates of the cropping box and the upper right corner coordinates They are calculated by the following formulas:
[0023] ,
[0024] .
[0025] In a second aspect, based on the video cropping method described in the first aspect, the present invention provides a method for evaluating the quality of cropped video, comprising the following steps:
[0026] Perform content integrity assessment based on the cropped video to obtain a content integrity score ;
[0027] Perform content consistency assessment based on the cropped video to obtain a content consistency score ;
[0028] Perform timing stability evaluation based on the cropped video to obtain a timing stability score ;
[0029] Score content completeness , content consistency score and timing stability scores Perform weighted summation to obtain the final quality evaluation score.
[0030] In practical applications, due to the lack of effective evaluation criteria to evaluate the quality of the cropped video, and still relying on manually set cropping parameters and manual quality assessment, the evaluation results are highly subjective and lack comparability, which further restricts the development of video processing technology. The present invention evaluates content integrity by calculating the average value of the significance ratio in the cropping frame through content integrity; by calculating the Hamming distance between the perceived hash values of adjacent frame images based on the perceptual hash algorithm through content consistency, and evaluating the content changes between video frames for content consistency evaluation; by calculating the position coordinates of each frame cropping frame through temporal stability, and calculating the second-order derivative and its standard deviation based on these coordinates to evaluate temporal stability. Through the above three aspects, a comprehensive solution is provided for the video cropped by automatically determining the optimal cropping position, which helps to improve the effect and quality of the cropped video.
[0031] Optionally, the content integrity assessment is performed based on the cropped video to obtain a content integrity score. include:
[0032] Calculation of the percentage of significant values: by calculating the significant map of the intermediate results The saliency value of each frame image in is accumulated to obtain the The overall saliency value of the frame image ; By accumulating the saliency value of each frame image in the cropping frame, we get The crop box saliency value of the frame image ; Wherein, the saliency value is used to indicate the strength of the saliency of the position of each pixel value in the video;
[0033] Calculate the first The saliency value of the image within the frame cropping box The significant value of the total The proportion of :
[0034] ;
[0035] In the formula, express from Get , express from Get , express OK The significant value of the column position; is the total number of frames of the cropped video; is the length coordinate of the pixel point in the video image; is the width coordinate of the pixel point in the video image;
[0036] Average saliency ratio calculation: calculate the saliency ratio of all image frames Add them up and divide by the total number of image frames Get the average significance ratio as the content completeness score , the content completeness score Calculated by the following formula:
[0037]
[0038] Optionally, the content consistency is evaluated based on the cropped video to obtain a content consistency score. include:
[0039] Calculate the perceptual hash value for each frame: Scale each frame of the cropped video to size , and then convert the scaled image into a grayscale image , for grayscale images Perform a two-dimensional discrete cosine transform to obtain the discrete cosine transform DCT coefficient matrix , the discrete cosine transform DCT coefficient matrix Calculated by the following formula:
[0040]
[0041] in,
[0042] In the above formula, c(u) and c(v) are normalization coefficients. The input grayscale image Middle position The pixel value of is the size of the image, and All are frequency indices after discrete cosine transform;
[0043] Let the matrix , extract the DCT coefficient matrix Top left 8×8 sub-block , sub-block Contains the main visual features of the cropped video image, the sub-block The expression is as follows:
[0044] ,
[0045] The sub-block is calculated by the following formula The mean :
[0046] ,
[0047] In the formula, represents the mean function;
[0048] Sub-block Middle Row m column value With the mean Compare and get Binary hash value of the frame image at row n and column m , the binary hash value Calculated by the following formula:
[0049]
[0050] The first All binary hash values of the frame image are concatenated to form a binary hash value with a total length of 64 bits, and then the first Frame image length The binary hash value at ;
[0051] Calculate the Hamming distance between adjacent frames: For each pair of adjacent frames, calculate the Hamming distance between the hash values of adjacent frames using the following formula:
[0052] ,
[0053] Calculate similarity: Calculate the similarity by the following formula based on the Hamming distance between the hash values of adjacent frames. Frame similarity :
[0054]
[0055] Calculate the content consistency score: Frame similarity The content consistency score is calculated by the following formula :
[0056]
[0057] Optionally, the timing stability evaluation is performed based on the cropped video to obtain a timing stability score. Includes: Get the center position of the cropping frame The horizontal axis : Calculate the center position of the cropping frame for each frame The horizontal axis Get the center position of the cropping frame The horizontal coordinate array , ,in Indicates The center position of the crop box of the frame The horizontal axis of
[0058] Calculate the second-order derivative: Calculate the center position of the cropping box for each frame by discrete difference method The second derivative of the horizontal axis , the formula is:
[0059] ,
[0060] In the above formula, i=1, 2, ... I;
[0061] By calculating the center position of the cropping box for each frame The second derivative of the horizontal axis , forming a second-order derivative array , the second-order derivative array The expression is as follows:
[0062] ,
[0063] Calculate the standard deviation of the second-order derivative array: According to the second-order derivative array The second-order derivative array is calculated by the following formula Standard Deviation :
[0064]
[0065] In the formula, is the second-order derivative array The average value of
[0066] Calculating the Timing Stability Score: Standard Deviation Use the sigmoid function for normalization and calculate the timing stability score using the following formula :
[0067]
[0068] In the formula, is a constant.
[0069] Optionally, the content integrity score , content consistency score and timing stability scores The final quality evaluation score is obtained by weighted summation:
[0070] Score the completeness of the content obtained , content consistency score and timing stability scores Assign weights to each , and , the final quality evaluation score is obtained by weighted summation using the following formula :
[0071]
[0072] The above evaluation process can improve the evaluation efficiency by evaluating the cropped video in a manner of sequentially evaluating content integrity, content consistency and temporal stability. In other embodiments, the above evaluation order can also be adjusted, for example, evaluating content consistency, content integrity and temporal stability in sequence, or evaluating content integrity, temporal stability and content consistency in sequence, or evaluating content consistency, temporal stability and content integrity in sequence, or evaluating temporal stability, content integrity and content consistency in sequence, or evaluating temporal stability, content consistency and content integrity in sequence.
[0073] Beneficial Effects
[0074] (1) The present invention improves the salient region prediction module obtained by the UAVSa model. The module can quickly and effectively extract the salient region in the video, accurately predict the center position and size of the cropping frame, and thus achieve precise cropping. By adjusting the center position and size of the cropping frame, it can flexibly cope with videos with different aspect ratios, and dynamically adjust the cropping frame according to different application requirements, so that it is suitable for cropping of various types of video content. This method is particularly suitable for video editing tasks that require high-precision cropping, such as movie editing, advertising production, etc.
[0075] (2) The cropped video quality evaluation method proposed in the present invention comprehensively evaluates the cropped video through content integrity score, content consistency score and temporal stability score. It not only takes into account the integrity of the video content, but also evaluates the content consistency and temporal stability after cropping, thereby providing a comprehensive quality evaluation. In practical applications, users can adjust the weights of each score according to specific needs to achieve the best cropping effect. For example, in some applications, content integrity may be more important, while in other applications, temporal stability may be the key indicator. By adjusting the weights, users can flexibly adjust the evaluation criteria to achieve the optimal cropping quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 The figure is a flow chart of a video cropping method and a method for evaluating the quality of cropped video in an embodiment of the present invention;
[0077] Figure 2 The figure is a schematic diagram of a significant result and a cropping frame position determination process in an embodiment of the present invention;
[0078] Figure 3 The figure is a flow chart of content integrity assessment in an embodiment of the present invention;
[0079] Figure 4 Shown is a flow chart of content consistency evaluation in an embodiment of the present invention. DETAILED DESCRIPTION
[0080] The following is a further description in conjunction with the accompanying drawings and specific embodiments. In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features.
[0081] Example 1
[0082] This embodiment provides a video cropping method, including:
[0083] The salient region prediction module of the video automatic cropping model is used to predict the salient region of the original video and generate an intermediate salient map. ;
[0084] Saliency map of intermediate results Mapping is performed to predict the center position of the cropping box ;
[0085] According to the aspect ratio of the original video Calculate the size of the crop box ; and The width and height of the original video respectively;
[0086] According to the center position of the cropping frame and the size of the crop box Calculate the coordinates of the lower left corner of the cropping box and the upper right corner coordinates ;
[0087] According to the center position of the cropping frame , the size of the cropping frame , the lower left corner coordinates of the cropping box and the upper right corner coordinates Output the cropped video;
[0088] Wherein, the video automatic cropping model includes a salient region prediction module and an output module connected in sequence;
[0089] The salient region prediction module in the automatic video cropping model is a module obtained by improving the prior knowledge modeling module in the original UAVSa model (the full name in English is An EfficientUnmanned Aerial Vehicle Video Saliency Prediction Model, and the Chinese name is Fast UAV Video Saliency Prediction Model). The specific improvement is to remove the observation prior module and the environmental semantic prior module from the original UAVSa model, retain the Gaussian prior module, and use the Gaussian prior module's modeling ability for the center bias phenomenon to extract the salient region and complete the prediction of the salient region.
[0090] In practical applications, since traditional video cropping methods mainly rely on manual operations or fixed cropping areas, they are inefficient and difficult to meet the needs of large-scale video data processing. The present invention provides a new video cropping method, which improves the original UAVSa model and designs a video automatic cropping model. The model is used to predict saliency regions and cropping frames, and the original video data and the aspect ratio of the video to be cropped are used as input parameters. The cropping direction and size are determined according to the cropping aspect ratio, and the intermediate result saliency map and cropping frame are output. The new video cropping method proposed in the present invention not only solves the problems of low efficiency and incomplete content in the existing video cropping technology, but also the automatically determined optimal cropping position provides a basis for subsequent video quality evaluation methods.
[0091] like Figure 2As shown in FIG. 1 , it is a schematic diagram of the process of determining the saliency result and the position of the cropping frame of the present invention. In this embodiment, the i-th frame image of the original video selected shows a scene with two people standing under a tree and a sun in the background. The intermediate result saliency map is obtained by the saliency region prediction module of the automatic video cropping. The scene position in the saliency map is: the cloud-like occlusion in the image is the saliency region of the image, and the image after saliency processing is displayed, highlighting the two people and the tree. According to the output module of the automatic video cropping, the position of the cropping frame is obtained, and videos with different cropping ratios can be obtained according to user requirements.
[0092] Through these cropping boxes of different proportions, we can see the different cropping effects of the salient area and how to choose the appropriate cropping box size in different scenarios to retain the most important content.
[0093] Example 2
[0094] Based on Example 1, this example introduces a specific implementation process of a video cropping method, which specifically includes the following contents:
[0095] When people observe a scene, the visual system receives a large amount of visual signal data. The visual system does not pay equal attention to every area, but selectively and quickly detects the salient areas in the scene, thereby quickly acquiring valuable visual information. This ability is called the visual attention mechanism, and these partial areas that can quickly attract people's attention are called salient areas.
[0096] The study of visual saliency is mainly to locate the areas that may attract visual attention in images and videos. For video cropping tasks, it can assist in determining the cropping area and be used for the evaluation of cropping results. To achieve this goal, the present invention performs saliency region detection on the original video based on the existing UAVSa model, wherein the prior knowledge modeling module in the model is improved, the observation prior module and the environmental semantic prior module are removed, and only the Gaussian prior is retained. By retaining only the Gaussian prior, its modeling ability for the center bias phenomenon can be utilized to effectively extract the saliency region, thereby ensuring the efficiency and accuracy of the prediction, and the video can be processed more efficiently while generating accurate saliency predictions.
[0097] Step 1: Salient Region Detection
[0098] Take a video sequence, i.e. the original video, as input, perform visual salient region detection on each frame of the video image, and generate an intermediate result salient map : , , They are the width and height of the original video respectively.
[0099] Step 2: Crop box prediction
[0100] The intermediate result saliency map Mapped to a value through multiple convolutional layers in the output module of the improved UAVSa model , and use the sigmoid function to convert the value Mapping to clipping position coefficients , the clipping position coefficient Calculated by the following formula:
[0101] ,
[0102] The crop position coefficient is set according to the crop direction The width of the original video or height Multiply to get the center position of the cropping frame , the center position of the cropping frame Calculated by the following formula:
[0103] ,
[0104] In the formula, The aspect ratio of the preset video cropping frame is based on the aspect ratio of the preset video cropping frame. Determine the crop direction. If the original video has an aspect ratio of Aspect ratio larger than the preset video cropping frame , the cropping box is cropped in the horizontal direction and the vertical height remains unchanged, otherwise the cropping box is cropped in the vertical direction and the horizontal width remains unchanged.
[0105] Crop box size is the crop length of the original video in the horizontal or vertical direction, the size of the crop frame Calculated by the following formula:
[0106] .
[0107] The lower left corner coordinates of the crop box and the upper right corner coordinates They are calculated by the following formulas:
[0108] ,
[0109] .
[0110] Finally, according to the center position of the cropping box , the size of the cropping frame , the lower left corner coordinates of the cropping box and the upper right corner coordinates Output the cropped video.
[0111] Example 3
[0112] Based on Example 2, this example introduces a method for evaluating the quality of a cropped video, which specifically includes the following contents:
[0113] Perform content integrity assessment based on the cropped video to obtain a content integrity score ;
[0114] Perform content consistency assessment based on the cropped video to obtain a content consistency score ;
[0115] Perform timing stability evaluation based on the cropped video to obtain a timing stability score ;
[0116] Score content completeness , content consistency score and timing stability scores Perform weighted summation to obtain the final quality evaluation score.
[0117] In practical applications, due to the lack of effective evaluation criteria to evaluate the quality of the cropped video, and still relying on manually set cropping parameters and manual quality assessment, the evaluation results are highly subjective and lack comparability, which further restricts the development of video processing technology. The present invention evaluates content integrity by calculating the average value of the significance ratio in the cropping frame through content integrity; by calculating the Hamming distance between the perceived hash values of adjacent frame images based on the perceptual hash algorithm through content consistency, and evaluating the content changes between video frames for content consistency evaluation; by calculating the position coordinates of each frame cropping frame through temporal stability, and calculating the second-order derivative and its standard deviation based on these coordinates to evaluate temporal stability. Through the above three aspects, a comprehensive solution is provided for the video cropped by automatically determining the optimal cropping position, which helps to improve the effect and quality of the cropped video.
[0118] Example 4
[0119] Based on Example 3, this example introduces a specific implementation process of a method for evaluating the quality of a cropped video, which specifically includes the following contents:
[0120] 1. Content Completeness Assessment
[0121] Content integrity mainly evaluates whether the cropped video can fully express the main content of the original video, ensuring that the cropped video can still fully express the core content of the original video. The main targets or areas in the original video should continue to exist in the cropped video to avoid important information being omitted or weakened due to cropping. The cropped video should contain all the valid information in the original video as much as possible to reduce the information loss caused by cropping. For example, the key dialogues, actions and plots in the original video should be completely retained in the cropped video.
[0122] In order to effectively evaluate the content integrity of the cropped video, the intermediate result saliency map is used As a basis for calculating the ratio of the saliency values of the video before and after cropping, a quantitative method is used to evaluate the ratio of the saliency information in each cropped frame to the total saliency information. The saliency map can highlight the most important area for visual attention in the video, that is, the main target object of the video content. Through this quantitative evaluation, it can be objectively measured whether the cropped video retains enough key content. Figure 3 As shown in the figure, it is a flow chart of content integrity assessment of the present invention. First, the input video frame number, salient value ratio and content integrity score are initialized, and it is determined whether the input video frame number is less than the total video frame number. If it is less, the score is calculated directly; if it is not less, the calculation is performed according to the following steps. The specific steps are as follows:
[0123] Step 1: Calculation of significant value ratio
[0124] By analyzing the intermediate result saliency map The saliency value of each frame image in is accumulated to obtain the The overall saliency value of the frame image ; By accumulating the saliency value of each frame image in the cropping frame, we get The crop box saliency value of the frame image ; Wherein, the saliency value is used to indicate the strength of the saliency of the position of each pixel value in the video;
[0125] Calculate the first The saliency value of the image within the frame cropping box The overall significant value The proportion of :
[0126] ;
[0127] In the formula, express from Get , express from Get , express OK The significant value of the column position; is the total number of frames of the cropped video; is the length coordinate of the pixel point in the video image; is the width coordinate of the pixel point in the video image;
[0128] Step 2: Calculation of average significant value ratio
[0129] The proportion of significant values of all image frames Add them up and divide by the total number of image frames Get the average significance ratio as the content completeness score , the content completeness score Calculated by the following formula:
[0130]
[0131] 2. Content consistency assessment
[0132] In the field of video processing, content consistency is one of the important criteria for evaluating the quality of cropped videos. The core of content consistency evaluation is to ensure that the cropped video can continuously and completely contain the same salient target in time sequence, and to minimize frequent cropping target switching. When evaluating the content consistency of cropped videos, the main purpose is to measure whether the cropped video maintains temporal coherence, especially the continuity of the same salient target in the time series.
[0133] To achieve this goal, the present invention introduces similarity, which is used to compare the similarity between different video frames and can accurately evaluate the continuity of salient targets in time sequence. If the similarity is high, it means that the content of the cropped video frame has strong consistency in time sequence, that is, the salient targets are continuously and completely preserved in time sequence.
[0134] At the same time, the present invention designs a perceptual hash algorithm for calculating similarity. The algorithm can effectively extract the visual features of the image and simplify the complex image information into a binary hash value of a fixed length, so that the similarity comparison between different images becomes more concise and efficient. It can quickly generate the hash value of the image and has strong robustness to slight changes in the image. The generated hash value has a fixed length, which facilitates the similarity comparison between different frames and improves the accuracy of content consistency assessment.
[0135] like Figure 4As shown, it is the content consistency evaluation flow chart of the present invention. First, the number of video frames, similarity and content consistency score are initialized; it is determined whether the input number of video frames is less than the total number of video frames. If it is less, the score is directly calculated; if it is not less, the calculation is performed according to the following steps: the perceptual hash value of each frame is calculated, and the main visual features are extracted by converting each frame image into a grayscale image and performing a two-dimensional discrete cosine transform to generate a 64-bit binary hash value. Then, for each pair of adjacent frames, the Hamming distance between adjacent frames is calculated, and the similarity between frames is measured by this distance. Finally, the content consistency score is calculated based on the similarity, so as to ensure that the cropped video maintains the coherence of the salient target in time sequence. The calculation steps are as follows:
[0136] Step 1: Calculate the perceptual hash value for each frame
[0137] Scale each frame of the cropped video to size , and then convert the scaled image into a grayscale image , for grayscale images Perform a two-dimensional discrete cosine transform to obtain the discrete cosine transform DCT coefficient matrix , the discrete cosine transform DCT coefficient matrix Calculated by the following formula:
[0138]
[0139] in,
[0140] In the above formula, c(u) and c(v) are normalization coefficients. The input grayscale image Middle position The pixel value of is the size of the image, and All are frequency indices after discrete cosine transform;
[0141] Let the matrix , extract the DCT coefficient matrix Top left 8×8 sub-block , sub-block Contains the main visual features of the cropped video image, the sub-block The expression is as follows:
[0142] ,
[0143] The sub-block is calculated by the following formula The mean :
[0144] ,
[0145] In the formula, represents the mean function;
[0146] Sub-block Middle The value of the mth column of the row With the mean Compare and get Frame image Binary hash value of row m column , the binary hash value Calculated by the following formula:
[0147]
[0148] The first All binary hash values of the frame image are concatenated to form a binary hash value with a total length of 64 bits, and then the first Frame image length The binary hash value at ;in, is a binary hash value length The binary number at The value range of is 1-64; in this embodiment, each frame of the image is divided into 8 rows and 8 columns, and there is a binary number in each row and column, that is, each frame has 8*8=64 binary numbers, and these 64 binary numbers are converted from an 8*8 matrix into a row with a length of 64. It's the length The binary number at the position, that is, the binary hash value, can also be considered as a 1*64 matrix: [1,0,1,1,0,0…,0]. The binary hash value at the length 2 is 0, and the binary hash value at the length j is 0 or 1.
[0149] Step 2: Calculate the Hamming distance between adjacent frames
[0150] For each pair of adjacent frames, the Hamming distance between the hash values of adjacent frames is calculated using the following formula:
[0151] ,
[0152] Step 3: Calculate similarity
[0153] The similarity of the i-th frame is calculated based on the Hamming distance between the hash values of adjacent frames using the following formula: :
[0154]
[0155] Step 4: Calculate content consistency score
[0156] According to Frame similarity The content consistency score is calculated by the following formula :
[0157]
[0158] 3. Timing Stability Evaluation
[0159] In the field of video processing, timing stability refers to whether the time intervals between frames remain consistent during video playback, which is directly related to the smoothness and naturalness of the viewing experience. If the timing of the video is unstable, the picture may jitter, skip frames, or switch unnaturally, thus affecting the overall visual quality of the video and user experience.
[0160] In order to evaluate the timing stability of the cropped video, the present invention considers the following key points: whether the time interval between frames is consistent, that is, whether the time difference between two frames is stable; whether the switching part of the video after editing is natural to avoid abrupt picture jumps, and whether the transition is smooth; whether there is jitter in the video during playback, and jitter refers to irregular jumps of video frames during playback.
[0161] Based on the above, the evaluation method proposed by the present invention effectively quantifies the smoothness or stability of the video in terms of time sequence, and avoids sudden changes or incoherence caused by cropping. This method uses the center position of the cropping frame and the size of the cropping frame to calculate the center position of each frame cropping frame. The horizontal axis , and based on the horizontal axis Calculate its second-order derivative and its standard deviation to measure the jitter of the cropping position. The specific steps are as follows:
[0162] Step 1: Get the center position of the cropping frame The horizontal axis
[0163] Calculate the center position of the cropping frame for each frame The horizontal axis Get the center position of the cropping frame The horizontal coordinate array , ,in Indicates the center position of the cropping frame of the i-th frame The horizontal axis of
[0164] Step 2: Calculate the second derivative
[0165] Calculate the center position of the cropping frame of each frame by discrete difference method The second derivative of the horizontal axis , the formula is:
[0166] ,
[0167] In the above formula, i=1, 2, ... I;
[0168] By calculating the center position of the cropping box for each frame The second derivative of the horizontal axis , forming a second-order derivative array , the second-order derivative array The expression is as follows:
[0169] ,
[0170] Step 3: Calculate the standard deviation of the second derivative array
[0171] According to the second-order derivative array The second-order derivative array is calculated by the following formula Standard Deviation :
[0172]
[0173] In the formula, is the second-order derivative array The average value of
[0174] Step 4: Calculate the Timing Stability Score
[0175] Standard Deviation Use the sigmoid function for normalization and calculate the timing stability score using the following formula :
[0176]
[0177] In the formula, is a constant.
[0178] Represents the standard deviation The result obtained after standardization. When normalizing the scoring criteria, the sigmod function usually converts the standard deviation Compressing to a fixed range (0 to 1) helps standardize the ratings, but it also causes the existing differences to be further compressed, making the rating differences between different videos less obvious. To restore these differences, the standard deviation can be adjusted before the sigmoid function is applied. Multiply by a constant This constant The amplification operation can preserve the original differences and more accurately reflect them in the normalized results.
[0179] Optionally, the differences between different videos are reduced after sigmoid compression, by calculating the standard deviation Multiply by a constant , which can effectively restore these differences and ensure that the scoring results are more discriminative. This adjustment will not destroy the smooth continuity of the sigmoid function and can still provide stable and interpretable scoring results.
[0180] 4. Calculate the weighted total score
[0181] When evaluating the quality of the cropped video, the content integrity score obtained by the above method is calculated by weighted summation. , content consistency score and timing stability scores The scores of the above three categories are combined to get the final quality evaluation score. The specific contents are as follows:
[0182] Step 1: Determine the weights
[0183] Assign a weight to each evaluation method for content integrity, content consistency and temporal stability, and the weights are , and The weights can be adjusted according to actual needs and importance. In this embodiment, the weights obtained through experiments are 0.4, 0.2, and 0.4 respectively.
[0184] Step 2: Weighted summation
[0185] Multiply each score by the corresponding weight and then sum them up to get the final quality evaluation score , which is used to measure the overall quality of the cropped video. The final quality evaluation score is obtained by weighted summation using the following formula :
[0186]
[0187] Figure 1 The flowchart of a video cropping method and a method for evaluating the quality of cropped video of the present invention only shows the logical sequence of the method described in this embodiment. In other possible embodiments of the present invention, different methods may be used without conflict. Figure 1 The steps shown or described are completed in the order shown, that is, the scoring method of the three aspects of the present invention, the content integrity scoring , content consistency score and timing stability scores The calculations can be performed in parallel.
[0188] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.
Claims
1. A video cropping method, characterized in that: include: The salient region prediction module of the video automatic cropping model is used to predict the salient region of the original video and generate an intermediate salient map. ; The output module of the video automatic cropping model is used to crop the intermediate result saliency map Mapping is performed to predict the center position of the cropping box ; According to the aspect ratio of the original video Calculate the size of the crop box ; and The width and height of the original video respectively; According to the center position of the cropping frame and the size of the crop box Calculate the coordinates of the lower left corner of the cropping box and the upper right corner coordinates ; According to the center position of the cropping frame , the size of the cropping frame , the lower left corner coordinates of the cropping box and the upper right corner coordinates Output the cropped video; Wherein, the video automatic cropping model includes a salient region prediction module and an output module connected in sequence; The salient region prediction module in the automatic video cropping model is a module obtained by improving the prior knowledge modeling module in the original UAVSa model. Specifically, the improvement is to remove the observation prior module and the environmental semantic prior module from the original UAVSa model, retain the Gaussian prior module, and use the Gaussian prior module's modeling ability for the center bias phenomenon to extract the salient region and complete the prediction of the salient region. The intermediate result saliency map Mapping is performed to predict the center position of the cropping box include: The intermediate result saliency map Mapped to a value through multiple convolutional layers in the output module of the video automatic cropping model , and use the sigmoid function to convert the value Mapping to clipping position coefficients , the clipping position coefficient Calculated by the following formula: , The crop position coefficient is set according to the crop direction Multiply it by the width w or height h of the original video to get the center position of the cropping frame , the center position of the cropping frame Calculated by the following formula: , In the formula, The aspect ratio of the preset video cropping frame is based on the aspect ratio of the preset video cropping frame. Determine the crop direction. If the original video has an aspect ratio of Aspect ratio larger than the preset video cropping frame , the cropping frame is cropped in the horizontal direction and the vertical height remains unchanged, otherwise the cropping frame is cropped in the vertical direction and the horizontal width remains unchanged; The size of the crop box is the crop length of the original video in the horizontal or vertical direction, the size of the crop frame Calculated by the following formula: ; The lower left corner coordinates of the cropping box and the upper right corner coordinates They are calculated by the following formulas: , 。 2. A method for evaluating the quality of a cropped video, characterized in that: The cropped video is obtained by the video cropping method according to claim 1, and the quality evaluation of the cropped video comprises the following steps: Perform content integrity assessment based on the cropped video to obtain a content integrity score ; Perform content consistency assessment based on the cropped video to obtain a content consistency score ; Perform timing stability evaluation based on the cropped video to obtain a timing stability score ; Score content completeness , content consistency score and timing stability scores Perform weighted summation to obtain the final quality evaluation score; The content integrity assessment is performed based on the cropped video to obtain a content integrity score include: By analyzing the intermediate result saliency map The saliency value of each frame image in is accumulated to obtain the The overall saliency value of the frame image ; By accumulating the saliency value of each frame image in the cropping frame, we get The crop box saliency value of the frame image ; Wherein, the saliency value is used to indicate the strength of the saliency of the position of each pixel value in the video; Calculate the first The saliency value of the image within the frame cropping box The overall significant value The proportion of : ; In the formula, express from Get , express from Get , express OK The significant value of the column position; is the total number of frames of the cropped video; is the length coordinate of the pixel point in the video image; is the width coordinate of the pixel point in the video image; The proportion of significant values of all image frames Add them up and divide by the total number of image frames Get the average significance ratio as the content completeness score , the content completeness score Calculated by the following formula: ; The content consistency evaluation is performed based on the cropped video to obtain a content consistency score include: Scale each frame of the cropped video to size , and then convert the scaled image into a grayscale image , for grayscale images Perform a two-dimensional discrete cosine transform to obtain the discrete cosine transform DCT coefficient matrix , the discrete cosine transform DCT coefficient matrix Calculated by the following formula: ; in, ; In the above formula, c(u) and c(v) are normalization coefficients. The input grayscale image Middle position The pixel value of is the size of the image, and They are all frequency indices after discrete cosine transform; Let the matrix , extract the DCT coefficient matrix Top left 8×8 sub-block , sub-block Contains the main visual features of the cropped video image, the sub-block The expression is as follows: , The sub-block is calculated by the following formula The mean : , In the formula, represents the mean function; Sub-block Middle The value of the mth column of the row With the mean Compare and get Frame image Binary hash value of row m column , the binary hash value Calculated by the following formula: ; The first All binary hash values of the frame image are concatenated to form a binary hash value with a total length of 64 bits, and then the first Frame image length The binary hash value at ; For each pair of adjacent frames, the Hamming distance between the hash values of adjacent frames is calculated using the following formula: , The Hamming distance between the hash values of adjacent frames is calculated by the following formula Frame similarity : ; According to Frame similarity The content consistency score is calculated by the following formula : ; The timing stability evaluation is performed based on the cropped video to obtain a timing stability score. include: Calculate the center position of the cropping frame for each frame The horizontal axis Get the center position of the cropping frame The horizontal coordinate array , ,in Indicates The center position of the crop box of the frame The horizontal axis of Calculate the center position of the cropping frame of each frame by discrete difference method The second derivative of the horizontal axis , the formula is: , In the above formula, i=1, 2, ... I; By calculating the center position of the cropping box for each frame The second derivative of the horizontal axis , forming a second-order derivative array , the second-order derivative array The expression is as follows: , According to the second-order derivative array The second-order derivative array is calculated by the following formula Standard Deviation : ; In the formula, is the second-order derivative array The average value of Standard Deviation Use the sigmoid function for normalization and calculate the timing stability score using the following formula : ; In the formula, is a constant; The content completeness score , content consistency score and timing stability scores The final quality evaluation score is obtained by weighted summation: Score the completeness of the content obtained , content consistency score and timing stability scores Assign weights to each , and , the final quality evaluation score is obtained by weighted summation using the following formula : 。
Citation Information
Patent Citations
Image cutting method and device thereof, computer equipment and storage medium
CN113516666A