Video shot segmentation method
Through the video shot segmentation method, key frames are used to segment shots with high similarity, and combined with intra-frame and inter-frame compression, the problem of unsatisfactory compression efficiency when video scenes change frequently is solved, and higher compression efficiency and ratio are achieved.
Patent Information
- Application Number
- CN202510843246.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When processing video scenes with frequent changes, the existing technology has unsatisfactory compression ratio and compression efficiency. The existing algorithm is highly complex and it is difficult to effectively improve the compression efficiency.
Through the video shot segmentation method, key frames are used to segment shots with high similarity. A combination of intra-frame and inter-frame compression is adopted to determine the shot type based on the RGB segmentation value, accurately find the segmentation point, reduce the entire frame processing volume, and improve compression efficiency.
It effectively improves the ratio and efficiency of video compression, reduces the amount of whole frame processing, and solves the problem of insufficient compression efficiency in the existing technology.
Smart Images

Figure CN120602669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video compression, and in particular to a video shot segmentation method. Background Art
[0002] The core goal of video compression is to reduce the amount of compressed data and shorten the compression time, that is, to improve the compression ratio and compression efficiency while maintaining visual quality as much as possible.
[0003] Frequent scene changes in a video (such as rapid camera switching, dynamic backgrounds, flash effects, etc.) will significantly affect the compression ratio, compression efficiency, bit rate distribution, and reconstruction quality. The size of the compression ratio directly affects the size of the storage space, and the length of the compression time directly affects the compression efficiency.
[0004] In the existing technology, there are improvements by optimizing the frame compression algorithm, but the effect is not obvious and the algorithm is highly complex. How to improve the compression ratio and compression efficiency when the video scene changes frequently is an urgent problem to be solved. Summary of the Invention
[0005] In response to the shortcomings of existing methods, the present invention finds that by accurately finding the key frames in the video, using the key frames to segment shots with high similarity, and performing inter-frame compression on the shots with high similarity, both the compression ratio and the compression efficiency can be improved.
[0006] The technical solution adopted by the present invention is: a video shot segmentation method comprising the following steps:
[0007] Step 1: Capture video footage and perform frame processing;
[0008] Step 2: cropping the framed image;
[0009] As a preferred embodiment of the present invention, cropping is performed with the aspect ratio being the same as the height ratio.
[0010] Step 3: Set four corner sampling areas at the four corners of the cropped image and randomly select several sampling points;
[0011] Step 4: Set a first regular octagonal sampling side with the same center point in the cropped image, set a second regular octagonal sampling side with the same center point in the first regular octagonal sampling side, and set a first rectangular sampling area with the same center point in the second regular octagonal sampling side; and randomly select a number of sampling points within the first and second regular octagonal sampling sides and the first rectangular sampling area;
[0012] Step 5: Separate and weight the RGB values of all sampling points to obtain a weighted pixel vector, and calculate the pixel mean and pixel variance of the weighted pixel vector;
[0013] Step 6: Calculate the pixel mean difference using the weighted pixel vector and pixel mean of all sampling points;
[0014] Step 7: Calculate the RGB segmentation value using pixel mean, pixel variance, and pixel mean difference;
[0015] As a preferred embodiment of the present invention, the formula of RGB segmentation value is:
[0016] BOU R =[((2*R 1 ave,k *R 1 ave,k-1 )+δ1)*(2*R 1 MD,k +δ2)] / [(R 1 ave,k ^2 +R 1 ave,k-1 ^2 +δ1)*(R 1 var,k +R 1 va r,k-1 +δ2)];
[0017] BOU G =[((2*G 1 ave,k *G 1 ave,k-1 )+δ1)*(2*G 1 MD,k +δ2)] / [(G 1 ave,k ^2 +G 1 ave,k-1 ^2 +δ1)*(G 1 var,k +G 1 var,k-1 +δ2)];
[0018] BOU B =[((2*B 1 ave,k *B 1 ave,k-1 )+δ1)*(2*B 1 MD,k +δ2)] / [(B 1 ave,k ^2 +B 1 ave,k-1 ^2 +δ1)*(B 1 var,k+B 1 va r,k-1 +δ2)];
[0019] BOU RGB =(BOU R +BOU G +BOU B ) / 3;
[0020] Among them, δ1 and δ2 are the first and second correction coefficients respectively; R 1 ave,k , G 1 ave,k 、B 1 ave,k are the mean values of R, G, and B pixels of the k-th frame sampling point; R 1 MD,k , G 1 MD,k 、B 1 MD,k are the mean differences of R, G, and B pixels of the k-th frame sampling point; R 1 var,k , G 1 var,k 、B 1 var,k are the R, G, and B pixel variances of the k-th frame sampling point respectively.
[0021] Step 8: Determine whether a hard cut shot or a gradient shot is used based on the RGB segmentation value;
[0022] As a preferred embodiment of the present invention, the hard-cut lens includes:
[0023] When BOU RGB,k <=ε1, the kth frame is the cutting point; ε1 is the first threshold, BOU RGB,k is the RGB segmentation value of the k-th frame.
[0024] As a preferred embodiment of the present invention, the hard-cut lens further comprises:
[0025] When ε1 <BOU RGB,k <=ε2 and BOU RGB,k-1 When >ε3, the kth frame is the cutting point, and ε2 and ε3 are the second and third thresholds respectively.
[0026] As a preferred embodiment of the present invention, the gradient lens includes:
[0027] When both the peak cut point and the valley cut point meet the conditions or only the valley cut point meets the conditions, the valley cut point is used as the split point;
[0028] When only the peak cut point meets the conditions, the peak cut point is used as the split point.
[0029] As a preferred embodiment of the present invention, the peak cut point satisfies the following conditions:
[0030] When BOU RGB,λ >=ε3 and BOU RGB,max >BOU RGB,max+λj And BOU RGB,max >BOU RGB,max-λi When , it is determined that there is a peak cut point;
[0031] And when ε4 <BOU RGB,max <ε5 and BOU RGB,max+1 >ε4 and BOU RGB,max-1 When >ε4, the peak frame BOU RGB,max is the split point;
[0032] Among them, ε3, ε4, and ε5 are the third, fourth, and fifth thresholds respectively; BOU RGB,max-λi , BOU RGB,max+λj BOU RGB,max The front lambda i Frame and post-λ j Frame; BOU RGB,max is the peak frame; BOU RGB,λ is the BOU of consecutive threshold frames λ RGB .
[0033] As a preferred embodiment of the present invention, the valley cutting point satisfies the following conditions:
[0034] When BOU RGB,λ >=ε3 and BOU RGB,min <BOU RGB,min+λj And BOU RGB,min <BOU RGB,min-λi When , it is judged that there is a valley cutting point;
[0035] And when ε3 <BOU RGB,min <ε6, the valley frame BOU RGB,min is the split point;
[0036] Among them, ε6 is the sixth threshold; BOU RGB,min-λi , BOU RGB,min+λj BOU RGB,min The front λ i Frame and post-λ j Frame; BOU RGB,min is the valley frame.
[0037] As a preferred embodiment of the present invention, a video shot segmentation system includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement a video shot segmentation method.
[0038] As a preferred embodiment of the present invention, a computer-readable medium stores computer program code, and the computer program code implements a video shot segmentation method when executed by a processor.
[0039] Beneficial effects of the present invention:
[0040] 1. The present invention samples each frame of image by dividing the area, obtains the key parts of the image frame, and reduces the processing load of the entire frame;
[0041] 2. Use sampling points to build a segmentation value algorithm and divide the shots reasonably according to the segmentation value;
[0042] 3. According to the characteristics of different lenses, use the corresponding threshold judgment method to accurately find the segmentation point;
[0043] 4. Solve the problem of insufficient compression ratio and compression efficiency when existing key frames are divided by custom equal intervals. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a flow chart of the video shot segmentation method of the present invention;
[0045] Figure 2 The original image P and the cropped image P of the present invention are 1 Schematic diagram;
[0046] Figure 3 Schematic diagram of the four-corner sampling area image of the present invention;
[0047] Figure 4 It is a schematic diagram of the octagonal and rectangular sampling area images of the present invention;
[0048] Figure 5 This is a schematic diagram of the present invention that only satisfies the peak cutting point;
[0049] Figure 6 This is a schematic diagram of the present invention that only satisfies the valley value cutting point;
[0050] Figure 7 This is a schematic diagram of the present invention that satisfies the peak-valley value cutting points simultaneously. DETAILED DESCRIPTION
[0051] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, and therefore only shows the components related to the present invention.
[0052] Frequent scene changes in a video will affect the compression ratio and compression efficiency. Existing compression algorithms such as H.264 / H.265 use intra-frame compression for keyframe I by setting it by default, and inter-frame compression for the image frames between keyframes.
[0053] For example, in a 60-frame video, if the number of key frames I is set to 3, frames 1, 21, and 41 are key frames I. Frame 1 uses intra-frame compression, while frames 2-20 use inter-frame compression. Similarly, frames 21-40 and 41-60 undergo significant image changes between frames 2-20, 21-40, or 41-60, resulting in suboptimal compression efficiency and ratio. This is because H.264 / H.265 relies on inter-frame prediction (motion estimation and inter-frame differencing), leveraging the similarity (temporal redundancy) between adjacent frames to reduce data size. However, when scenes change frequently and significantly, adjacent frames differ significantly, motion compensation fails, and inter-frame prediction efficiency decreases. Furthermore, bitrate fluctuations are significant. Static scenes require low bitrate, allowing more bits to be allocated to key frames I. However, dynamic scenes experience a surge in instantaneous bitrate demand (requiring frequent I-frame refreshes). If the rate control strategy (such as CBR / VBR) is not adapted, insufficient bitrate, blocking artifacts, and blurring (quantization distortion) may result.
[0054] The present invention analyzes the video frame by frame to find the demarcation value between dynamic and static scenes, and uses the demarcation value as the key frame I to divide the video into several sub-shots; and performs targeted compression on continuous frames with high similarity in the sub-shots to improve the compression ratio and compression efficiency.
[0055] like Figure 1 As shown, a video shot segmentation method includes the following steps:
[0056] Step 1: Capture video footage and perform frame processing;
[0057] For example, a 60-second video shot with a frame rate of 60fps will have 3600 frames of images;
[0058] Step 2: crop the original image of each frame with the same aspect ratio coefficient based on the center point of the original image;
[0059] like Figure 2 As shown, the width and height of the original image P are W*H, and the cropped image P 1 The width and height are W 1 *H 1 , the aspect ratio coefficient is ɑ, W 1 =W*ɑ,H 1 =H*ɑ; where 256 / (W*H)=<ɑ<=1 / 2;
[0060] The center point of image P is O, P 1 The center point is O ’, O and O 1 concentric;
[0061] For example, if the original image P is 1280*1280 pixels and the aspect ratio coefficient ɑ=1 / 2, then the cropped image P 1 The size is 640*640, which is 1 / 4 of the original image P.
[0062] Step 3: After cropping the image P 1 Set up several four-corner sampling area images P at the four corners i 2 , randomly select m in the four corner sampling areas 1 i sampling points;
[0063] Four corner sampling area image P i 2 The width and height are W i 2 *H i 2 Among them, W i 2 =W 1 / β i , H i 2 =H 1 / β i , 2<=β i <=W 1 / 6 or 2<=β i <=H 1 / 6; i is the index of the four corner sampling area, i<=4 corresponds to the four corners in the clockwise direction, that is, the upper left is 1, the upper right is 2, the lower left is 3, and the lower right is 4;
[0064] The preferred β1=β2=β3=β4;
[0065] Among them, 36= <m 1 i <=W i 2 *H i 2 ;
[0066] m 1 The index of the four corner sampling areas;
[0067] like Figure 3 This is a schematic diagram of the width and height of the sampling area corresponding to index 1 in the upper left corner.
[0068] Step 4: Use the cropped image P 1 The center point of the cropped image P is the center. 1Set the first regular octagon inside, set the second regular octagon inside the first regular octagon, set the first rectangular area inside the second regular octagon; randomly select m on the edge of the first regular octagon 2 sampling points, randomly select m on the edge of the second regular octagon 3 sampling points; select m sampling points in the first rectangular area 4 sampling points;
[0069] Among them, m 2 Indicates the index of the first regular octagon, m 3 Indicates the index of the second regular octagon, m 4 Indicates the index of the first rectangular area.
[0070] Among them, the side length of the first regular octagon is x1,16 <x1<=w 1 / 2;
[0071] The side length of the second regular octagon is x2, 16 <= x2 <x1;
[0072] The width W of the first rectangle 3 , high H 3 , 8<=W 3 、H 3 <=(1+1 / sqrt(2))*x2;
[0073] Step 5: Separate the RGB values of the sampling points in each sampling area to obtain the pixel vector R j , G j 、B j value; R j , G j 、B j The value is weighted γ j Get the weight pixel vector R 1 j , G 1 j 、B 1 j , calculate the pixel mean R of all sampling points 1 ave , G 1 ave 、B 1 ave and pixel variance R 1 var , G 1 var 、B 1 var ;
[0074] For example, 1 1.m 1 2.m 1 3.m 14.m 2 、m 3 、m 4 There are seven sampling points in the sampling area (the maximum value of the four corner sampling areas is 4), that is, when j = 1, it corresponds to the sampling area m 1 1, j = 2 corresponds to the sampling area m 1 2...j=7 corresponds to sampling area m 4 ;
[0075] R j The value is used as an example to illustrate;
[0076] R 1 j =R j *γ j ; 1=<γ j <1.5, preferably, γ1=γ2=γ3=γ4=1, γ5=1.2, γ6=1, γ7=1.33;
[0077] Similarly, we get G 1 j 、B 1 j .
[0078] Step 6: Use the pixel vector R of all sampling points 1 j The pixel values in R 1 ave The difference between two consecutive frames of image P is calculated 1 The pixel mean difference R 1 MD,k ;
[0079] R 1 MD,k =Mean Difference[(R 1 k -R 1 ave,k ),(R 1 k-1 -R 1 ave,k-1 )];
[0080] Among them, Mean Difference is the mean difference function, k is the kth frame, R 1 ave,k is the pixel mean of all sampling points in the kth frame;
[0081] Similarly, we get G 1 MD,k 、B 1 MD,k ;
[0082] Step 7: Calculate the RGB segmentation value BOU using the pixel mean, pixel variance and pixel mean difference of all sampling points RGB ;
[0083] Split Value BOU RGB The formula is:
[0084] BOU R =[((2*R 1 ave,k *R 1 ave,k-1 )+δ1)*(2*R 1 MD,k +δ2)] / [(R 1 ave,k ^2 +R 1 ave,k-1 ^2 +δ1)*(R 1 var,k +R 1 va r,k-1 +δ2)];
[0085] Similarly, we get:
[0086] BOU G =[((2*G 1 ave,k *G 1 ave,k-1 )+δ1)*(2*G 1 MD,k +δ2)] / [(G 1 ave,k ^2 +G 1 ave,k-1 ^2 +δ1)*(G 1 var,k +G 1 var,k-1 +δ2)];
[0087] BOU B =[((2*B 1 ave,k *B 1 ave,k-1 )+δ1)*(2*B 1 MD,k +δ2)] / [(B 1 ave,k ^2 +B 1 ave,k-1 ^2 +δ1)*(B 1 var,k +B1 va r,k-1 +δ2)];
[0088] calculate:
[0089] BOU RGB =(BOU R +BOU G +BOU B ) / 3;
[0090] Among them, δ1 is the first correction coefficient, which is an arbitrary number not equal to 0, ensuring that the denominator is not equal to 0;
[0091] δ2 is the second correction coefficient, which is an arbitrary number not equal to 0, ensuring that the denominator is not equal to 0.
[0092] Step 8. According to BOU RGB Identify hard cuts and fades;
[0093] Step 81: Determining a hard cut shot includes:
[0094] Step 811: When BOU RGB,k <=ε1, set the frame as the cutting point; preferably, ε1=0.24;
[0095] Indicates that the frame is a hard cut shot and the cut point needs to be recorded, indicating that the background of the two frames is quite different;
[0096] For example, 60 frames of image, BOU RGB,1 is the segmentation value between the first and second frames, BOU RGB,1 <0.24; BOU RGB,59 is the segmentation value between the 59th frame and the 60th frame, BOU RGB,59 <0.24;
[0097] Find two cut values BOU RGB,1 , BOU RGB,59 , the 60-frame video is divided into 3 shots, and intra-frame compression is used for the 1st and 60th frames with large background differences; for the 2nd to 59th frames with small background differences, a combination of intra-frame compression and inter-frame compression is used, that is, intra-frame compression is used for the 2nd frame, and inter-frame compression is used for frames 3-59. Since the background differences of frames 3-59 are small and the similarity is high, the compression efficiency and compression ratio can be effectively improved.
[0098] Step 812: When ε1 <BOU RGB,k <=ε2, and BOU RGB,k-1 >ε3, set the frame as the cutting point;
[0099] Among them, ε1, ε2, and ε3 are custom parameters, ε3>ε2>ε1;
[0100] Preferably, ε2=0.56, ε3=0.685.
[0101] Step 82: determining that the lens is a gradient lens includes:
[0102] Step 821: When BOU RGB,λ >=ε3 and the average demarcation peak frame BOU within consecutive threshold frames λ RGB,max Greater than the previous λ of the peak frame i Frame and post-λ j When the value of the frame is reached, it is determined that there is a peak cutting point;
[0103] When BOU RGB,λ >=ε3 and the average demarcation valley frame BOU within consecutive threshold frames RGB,min Less than the previous λ of the peak frame i Frame and post-λ j When the value of the frame is 0, it is determined that there is a valley cutting point;
[0104] In this embodiment, i =λ j =3;
[0105] The threshold frame λ is a custom parameter;
[0106] Step 8211: When only the peak cut point exists and ε4 <BOU RGB,max <ε5 and BOU RGB,max+1 >ε4 and BOU RGB,max-1 When >ε4, the peak frame is used as the segmentation point;
[0107] Among them, ε4 and ε5 are custom parameters, ε5>ε4; BOU RGB,max-1 , BOU RGB,max+1 Indicates BOU RGB,max 1 frame before and after;
[0108] Preferably, ε4=0.9, ε5=0.938;
[0109] Step 8212: When only the valley segmentation point exists and ε3 <BOU RGB,min When <ε6, the valley frame is used as the segmentation point;
[0110] Among them, ε6 is a custom parameter, ε6>ε3;
[0111] Preferably, ε6=0.9;
[0112] Step 8213: When both the peak and valley cut points are satisfied, first determine whether step 8212 is satisfied. If step 8212 is not satisfied, execute step 8211.
[0113] Among them, step 8211 and step 8212 are parallel logical relationships, not front-to-back logical relationships;
[0114] ε1 to ε6 are the first to sixth thresholds respectively;
[0115] like Figure 5 This is an example of only meeting the peak cut point;
[0116] First, the threshold frame λ = 9, the 4th frame is BOU RGB,max ; Among them, the BOU of the 4th frame RGB Greater than frames 1, 2, 3, 5, 6, and 7;
[0117] Although the 7th frame is BOU RGB,min , but does not meet the last 3 frame conditions;
[0118] Secondly, determine the BOU of the 4th frame RGB =0.93, satisfying: 0.9<0.93<0.938; and BOU RGB,3 =0.912>0.9 and BOU RGB,4 =0.913>0.9; so the 4th frame is used as the cutting point of the gradient lens.
[0119] like Figure 6 This is an example of a point that only satisfies the valley cut;
[0120] First, the threshold frame λ = 11, the 4th frame is BOU RGB,min , BOU of the 4th frame RGB Smaller than the 1st, 2nd, 3rd, 5th, 6th, and 7th frames;
[0121] Although the 11th frame is BOU RGB,max , but does not meet the last 3 frame conditions;
[0122] Secondly, determine the BOU of the 4th frame RGB =0.84, satisfying: 0.84<0.9; the 4th frame is the gradual lens cutting point.
[0123] like Figure 7 This is an example of meeting both peak and valley cut points;
[0124] First, the threshold frame λ = 12, the first frame is BOU RGB,max ; But the last 3 frames conditions are not met;
[0125] Secondly, the 4th frame is BOU RGB,min , satisfying the conditions of the three frames before and after;
[0126] Secondly, the 8th frame is BOU RGB,max , satisfying the conditions of the three frames before and after;
[0127] Secondly, determine the BOU of the 4th frameRGB =0.87, which satisfies: 0.87<0.9; the 4th frame is used as the gradient lens cutting point, and there is no need to judge the 8th frame.
[0128] Step 73: When the cut point condition is met, the next frame of the cut point frame is used as the starting frame to perform the threshold frame judgment of the next cycle;
[0129] For example, for a 60-frame image, λ=12, the first cycle determines frames 1-12. When the fourth frame is detected as a cutting point, the next cycle is frames 5-16, and so on.
[0130] Since the shot similarity is low when different shots are switched, the lower the similarity, the smaller the compression ratio, and the compression effect is poor; however, the present invention accurately finds the segmentation points (i.e., key frames) and compresses similar frames in the shots between the cutting points, effectively improving the compression performance.
[0131] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification and must be determined according to the scope of the claims.
Claims
1. A video shot segmentation method, characterized in that: The following steps are involved: Step 1: Capture video footage and perform frame processing; Step 2: cropping the framed image; Step 3: Set four corner sampling areas at the four corners of the cropped image and randomly select several sampling points; Step 4: Set a first regular octagonal sampling side with the same center point in the cropped image, set a second regular octagonal sampling side with the same center point in the first regular octagonal sampling side, and set a first rectangular sampling area with the same center point in the second regular octagonal sampling side; and randomly select a number of sampling points within the first and second regular octagonal sampling sides and the first rectangular sampling area; Step 5: Separate and weight the RGB values of all sampling points to obtain a weighted pixel vector, and calculate the pixel mean and pixel variance of the weighted pixel vector; Step 6: Calculate the pixel mean difference using the weighted pixel vector and pixel mean of all sampling points; Step 7: Calculate the RGB segmentation value using pixel mean, pixel variance, and pixel mean difference; Step 8. Determine hard-cut shots and gradient shots based on RGB segmentation values.
2. The video shot segmentation method according to claim 1, wherein: The formula for RGB segmentation value is: BOU R =[((2*R 1 ave,k *R 1 ave,k-1 )+δ1)*(2*R 1 MD,k +δ2)] / [(R 1 ave,k ^2 +R 1 ave,k-1 ^2 +δ1)*(R 1 var,k +R 1 var,k-1 +δ2)]; BOU G =[((2*G 1 ave,k *G 1 ave,k-1 )+δ1)*(2*G 1 MD,k +δ2)] / [(G 1 ave,k ^2 +G 1 ave,k-1 ^2 +δ1*(G 1 var,k +G 1 var,k-1 +δ2)]; BOU B =[((2*B 1 ave,k *B 1 ave,k-1 )+δ1)*(2*B 1 MD,k +δ2)] / [(B 1 ave,k ^2 +B 1 ave,k-1 ^2 +δ1)*(B 1 var,k +B 1 var,k-1 +δ2)]; BOU RGB =(BOU R +BOU G +BOU B ) / 3; Among them, δ1 and δ2 are the first and second correction coefficients respectively; R 1 ave,k , G 1 ave,k 、B 1 ave,k are the mean values of R, G, and B pixels of the k-th frame sampling point; R 1 MD,k , G 1 MD,k 、B 1 MD,k are the mean differences of R, G, and B pixels of the k-th frame sampling point; R 1 var,k , G 1 var,k 、B 1 var,k are the R, G, and B pixel variances of the k-th frame sampling point respectively.
3. The video shot segmentation method according to claim 2, wherein: Cut shots include: When BOU RGB,k <=ε1, the kth frame is the cutting point; ε1 is the first threshold, BOU RGB,k is the RGB segmentation value of the k-th frame.
4. The video shot segmentation method according to claim 2, wherein: Hard cut shots also include: When ε1 <BOU RGB,k <=ε2 and BOU RGB,k-1 When >ε3, the kth frame is the cutting point, and ε2 and ε3 are the second and third thresholds respectively.
5. The video shot segmentation method according to claim 2, wherein: Gradient lenses include: When both the peak cut point and the valley cut point meet the conditions or only the valley cut point meets the conditions, the valley cut point is used as the split point; When only the peak cut point meets the conditions, the peak cut point is used as the split point.
6. The video shot segmentation method according to claim 5, wherein: The peak cut point meets the conditions: When BOU RGB,λ >=ε3 and BOU RGB,max >BOU RGB,max+λj And BOU RGB,max >BOU RGB,max-λi When , it is determined that there is a peak cut point; And when ε4 <BOU RGB,max <ε5 and BOU RGB,max+1 >ε4 and BOU RGB,max-1 When >ε4, the peak frame BOU RGB,max is the split point; Among them, ε3, ε4, and ε5 are the third, fourth, and fifth thresholds respectively; BOU RGB,max-λi , BOU RGB,max+λj BOU RGB,max The front λ i Frame and post-λ j Frame; BOU RGB,max is the peak frame; BOU RGB,λ is the BOU of consecutive threshold frames λ RGB .
7. The video shot segmentation method according to claim 5, wherein: The valley cutting point meets the following conditions: When BOU RGB,λ >=ε3 and BOU RGB,min <BOU RGB,min+λj And BOU RGB,min <BOU RGB,min-λi When , it is judged that there is a valley cutting point; And when ε3 <BOU RGB,min <ε6, the valley frame BOU RGB,min is the split point; Among them, ε6 is the sixth threshold; BOU RGB,min-λi , BOU RGB,min+λj BOU RGB,min The front λ i Frame and post-λ j Frame; BOU RGB,min is the valley frame.
8. The video shot segmentation method according to claim 1, wherein: Cropping is done with the same ratio of aspect ratio to height ratio.
9. The video shot segmentation system is characterized by: include: a memory for storing instructions executable by the processor; A processor, configured to execute instructions to implement the video shot segmentation method according to any one of claims 1 to 8.
10. A computer-readable medium storing computer program code, characterized in that When the computer program code is executed by a processor, the video shot segmentation method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
News video scene generating method
CN102685398A
Shot segmentation method based on X264 compressed video
CN104869403A
Method and device for recognizing shot cut
CN106331524A
Video shot segmentation method, system and device and storage medium
CN110766711A
Video scene recognition method and device, computer equipment and storage medium
CN114187558A