A Key Frame Extraction Method for Team Sports Videos Based on Global Motion Statistical Features

By using global motion statistical features for fine-grained slicing and keyframe extraction in team sports videos, the problem of difficulty in extracting related keyframes in the prior art is solved, and more efficient and accurate video summary generation is achieved.

CN113032631BActive Publication Date: 2025-06-10BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110204179.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-24
Publication Date
2025-06-10
Estimated Expiration
2041-02-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract keyframes related to team sports video events, especially when the color characteristics between frames are similar, and are susceptible to redundant local motion characteristics.

Method used

The video fine-grained segmentation is performed using a method based on global motion statistical features, candidate keyframes are extracted through global motion statistical features, and redundant frames are removed in combination with spatiotemporal consistency and hierarchical clustering, and finally a representative keyframe set is obtained.

Benefits of technology

Improve the accuracy of keyframe extraction, make the extracted keyframes more relevant to the competition event, reduce the impact of redundant information, and significantly improve the quality of the video digest.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113032631B_ABST
    Figure CN113032631B_ABST
Patent Text Reader

Abstract

A key frame extraction method for team sports videos based on global motion statistical features belongs to the field of video analysis. First, preprocess the video, segment the video in units of shots, extract the game shots and structure them; then estimate the global motion corresponding to the video frames and calculate the global motion statistical features corresponding to the video frames; further, based on the horizontal translation amount of the shots, further segment the shot video segments into fine-grained video segments; then extract candidate key frames according to the global motion statistical features, and finally combine spatio-temporal consistency and hierarchical clustering to extract representative key frames from the candidate key frame set to remove redundant frames to obtain the final key frame set. The present invention makes full use of the global motion information of team sports videos, improves the key frame extraction performance, reduces the interference of redundant optical flow information, and is beneficial to enhancing the correlation between the extracted key frames and the key events in the game.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video analysis. Background Art

[0002] Videos are ubiquitous in many fields such as movies, surveillance, sports, news, etc. The huge number of video files makes the workload of analyzing and understanding video content quite heavy. In addition, browsing videos in a video database consumes a large amount of time. Therefore, video summarization that can assist users in quick browsing is an important and evolving research field. Key frame extraction is an important way of video summarization, which can visually represent the summary of a video through a series of video frames.

[0003] Sports videos are a type of broadcast television video. Team sports such as football and basketball are deeply loved by people and have a wide mass base. Therefore, the research on team sports videos has broad application prospects. Accurate key frame extraction can contribute to applications such as the generation of video highlights, quick browsing of game content, and extraction of game technical statistics information.

[0004] The general framework of existing video key frame extraction methods is as follows: First, detect shot boundaries, and then segment the video into shot video segments. After that, use features such as color or optical flow to extract key frames through frame clustering. Since video frames corresponding to key events in the same video often have relatively similar visual features, these methods are difficult to effectively extract key frames related to game events.

[0005] Vennila et al. used Markov chains, transition probability matrices and other means to enhance the clustering results of visual features to obtain a central clustering with higher confidence. Mendi et al. used hybrid optical flow to calculate the motion vectors in video frames, then found local minima according to the motion feature time curve, and selected the corresponding frames as key frames.

[0006] The key frames extracted by the algorithm of Vennila et al. can accurately describe the main content of the video in most cases, but the clustering effect for video frames with similar color features is not ideal. The algorithm of Mendi et al. can effectively depict the shot motion of the video and dynamically extract key frames, but this algorithm is vulnerable to some redundant local motion features.

[0007] For sports videos, the video frames of game shots often have relatively similar visual features, so it is difficult for existing methods to effectively extract key frames related to game events. And only performing basic shot segmentation on the original video is not enough for sports videos because a single shot often contains multiple events, so fine-grained video segmentation is required. In addition, the motion features commonly used in algorithms for sports videos are extracted from optical flow, which includes both the global motion from the camera and the local motion of players on the field, and even involves changes in the game logo or scoreboard. Mixed optical flow is a mixture of the motion optical flows from the above different subjects, some of which is irrelevant to events and cannot help in the extraction of key frames. Redundant information may instead affect the performance of algorithms based on optical flow. Summary of the Invention

[0008] The object of the present invention is to provide a method for extracting key frames of team sports videos based on global motion statistical features. As is well known, most events in team sports videos are the transitions between offense and defense, which are closely related to global motion transitions. Therefore, by using fine-grained segmentation based on global motion, the key frames extracted from fine-grained segments can be more relevant to key events. Thus, after extracting game shots, the present invention performs fine-grained video segmentation based on the horizontal translation amount of the global motion at the shot level; then extracts candidate key frames according to global motion statistical features, and finally extracts representative key frames from the set of candidate key frames to remove redundant frames to obtain the final set of key frames. In addition, a dataset SportKF is constructed, which includes 25 videos from four team sports (basketball, football, American football, and hockey), totaling 112 minutes, containing 197,878 frames, among which there are 764 key frames.

[0009] Figure 1 is the main flowchart of the technical solution of the present invention. As Figure 1 shown, a method for extracting key frames of team sports videos based on global motion statistical features proposed by the present invention includes the following steps:

[0010] 1. Extract the game shot video segments through a method for extracting game shots of team sports videos based on color features.

[0011] 2. Extract the mixed optical flow of the video frames in each shot video segment and estimate the global motion.

[0012] 3. Calculate the global motion statistical features corresponding to the video frames, including the shot scaling amount and translation amount.

[0013] 4. Further segment the shot video segments into fine-grained video segments based on the horizontal translation amount of the shots.

[0014] 5. Draw the curves of the lens' translation amount and zoom amount, and extract the candidate key frames in each fine-grained video segment based on the maximum values of the curves.

[0015] 6. Extract representative key frames from the candidate key frame set by combining spatio-temporal consistency and hierarchical clustering,

[0016] to obtain the final key frame set.

[0017] Step 1:

[0018] Input: The original team sports video

[0019] Output: The set of game shot video segments.

[0020] Step 2:

[0021] Input: The video frames of the game shot video segments

[0022] Output: The global optical flow field corresponding to the video frames

[0023] Step 3:

[0024] Input: The global optical flow values of the four corner points of the video frames

[0025] Output: The global motion statistical features: including the zoom amount and translation amount of the lens

[0026] As Figure 2 shown in the schematic diagram of lens translation and zoom (this schematic diagram shows the displacement of the points corresponding to the conversion of the lens from the original lens range to the zoomed and translated lens range. Among them, the rectangle P tl P bl P br P tr represents the original lens range, and the rectangle P' tl P' bl P' br P' tr represents the transformed lens range. To more intuitively illustrate this process, assume that the motion is divided into two steps: translation and then zoom, then the dashed rectangle represents the lens range after translation in the intermediate step.).

[0027]

[0028] In the formula, X zoom , Y zoom are the zoom vectors of the lens in the X and Y directions respectively, and MAG zoom is the combined zoom amount of the lens. (x tl , y tl ), (x br , y br ) are the global motion vectors of the upper left and lower right corner points of the video frame respectively.

[0029] X trans = x tl -X zoom

[0030] Y trans = y tl -Y zoom

[0031]

[0032] Wherein X trans , Y trans are respectively the translation vectors of the lens in the X and Y directions, and DIS trans is the combined translation amount of the lens.

[0033] Step 4:

[0034] Input: Set of competition shot video clips

[0035] Output: Set of fine-grained video clips

[0036] Trans_con = SGN(X trans (Frame i )·X trans (Frame i+1 ))

[0037] Where Trans_con represents the conversion of the global horizontal translation amount direction, the SGN(x) function returns the sign of x, if x>0, then SGN(x) = 1; if x = 0, then SGN(x) = 0, indicating that the horizontal translation amount direction of the lens has not changed; if x<0, then SGN(x) = -1, indicating that the horizontal translation amount direction of the lens has changed; X trans (Frame i ), X trans (Frame i+1 ) are respectively the translation vectors of the corresponding lenses in the X direction for the i-th frame and the (i + 1)-th frame.

[0038] Specific operation: In each shot video clip, calculate Trans_con corresponding to the i-th frame and the (i + 1)-th frame one by one. If Trans_con = -1, then separate the i-th frame and the (i + 1)-th frame until the second-to-last video frame of each shot video clip is calculated; then store the frames before the first interval as the first fine-grained video clip, then store the frames between each two intervals as the corresponding fine-grained video clips, and the last fine-grained video clip consists of all the frames after the last interval;

[0039] Step 5:

[0040] Input: All video frames of the set of fine-grained video clips

[0041] Output: Candidate key frame set

[0042] Step 1: For each fine-grained video segment, draw the change curve MAG of the combined scaling amount of the video frame sequence zoom (f) and the change curve DIS of the combined translation amount trans (f) (f is the frame number corresponding to the video frame sequence).

[0043] Step 2: Simultaneously scan the change curves of the combined translation amount and the combined scaling amount. Taking the change curve of the combined scaling amount as an example, find f 1 , f 3 , f 2 as three consecutive video frames. If MAG zoom (f 3 ) > MAG zoom (f 1 ) and MAG zoom (f 3 ) > MAG zoom (f 2 ), then f 3 is the frame number corresponding to a maximum value of the MAG zoom (f) curve. The DIS trans (f) curve is the same. If the same frame number f 3 is both the frame number corresponding to the maximum value of the MAG zoom (f) curve and the frame number corresponding to the maximum value of the DIS trans (f) curve, then select the frame corresponding to f 3 as the key frame; repeat Step 2 until the search is complete;

[0044] Step 3: Traverse the initially extracted key frame set, find two frames k 1 , k 2 . If k 2 - k 1 ≥ med (the med value is a set threshold. The larger the med, the fewer adjacent key frames that meet the condition for extracting intermediate frames, and thus the fewer intermediate frames are extracted, and vice versa), then extract the video frames with the frame numbers ( represents rounding x downwards) in the frame set of the original fine-grained video segment, and store them in the candidate key frame set arranged in ascending order of frame numbers; repeat Step 3 until the search is complete;

[0045] An example of candidate key frame extraction is shown in Figure 3 .

[0046] Step 6:

[0047] Input: Candidate key frame set

[0048] Output: Final Key Frame Set

[0049] Extraction of Representative Key Frames Based on Spatiotemporal Consistency:

[0050] Step 1: Color-code the global optical flow fields corresponding to all candidate key frames (output of Step 2) to obtain global optical flow images;

[0051] Step 2: Convert all global optical flow images from the RGB color space to the HSI color space;

[0052] Step 3: Sort all global optical flow images in the HSI color space according to the time series of the corresponding key frames;

[0053] Step 4: Calculate the two-dimensional histogram of the HI channel corresponding to all global optical flow images in the HSI color space;

[0054] Step 5: Use the compareHist function in the OpenCV library based on Python 3.7, set the method parameter to cv2.HISTCMP_BHATTACHARYYA, and calculate the Bhattacharyya distance B between the two-dimensional histograms of the HI channels of adjacent images;

[0055] The Bhattacharyya distance is used to measure the probability distributions of two discrete data. In histogram similarity calculation, the Bhattacharyya distance has good results, and its formula for calculating histogram similarity is as follows:

[0056]

[0057] where h 1 and h 2 are two histograms, n is the number of histogram groups, and h i [m] represents the frequency corresponding to the m-th group; the smaller the value, the higher the correlation, 0 for perfect match, and 1 for complete mismatch;

[0058] Step 6: If B > TH B , detect whether the key frames corresponding to these two adjacent images have been extracted. If not, extract the unextracted key frames;

[0059] Extraction of Representative Key Frames Based on Hierarchical Clustering:

[0060] Step 7: Group all global optical flow images according to the fine-grained video segments where the corresponding key frames were previously located, and sort them according to the time series;

[0061] Step 8: Calculate the three-channel histogram of each global optical flow image and horizontally splice them into an array containing 768 feature values;

[0062] Step 9: The cluster.hierarchy module in the scipy library based on Python 3.7 can be used to perform hierarchical clustering on the arrays corresponding to all global optical flow images in each group. Set the method to Average Linkage and set the inter-cluster distance threshold to TH HC ;

[0063] Step 10: Extract the key frames corresponding to the arrays with the smallest distance to each array in each cluster after clustering;

[0064] Step 11: Detect whether the first and last candidate key frames of each fine-grained video segment are extracted. If not, extract the unextracted key frames;

[0065] Step 12: Take the union of the above-extracted key frames and remove the duplicate key frames. Obtain the final key frame set.

[0066] The present invention proposes a method for extracting key frames of team sports videos based on global motion statistical features. Introducing global motion can perform fine-grained segmentation on videos and effectively extract video key frames. Experimental results on the SportKF dataset show that the proposed KEGMS can achieve the best performance. This method can effectively extract key frames from team sports videos with similar inter-frame color features, and at the same time effectively avoids the influence of some redundant local motion features on the algorithm for extracting key frames based on optical flow. In addition, fine-grained video segmentation can better make the extracted key frames correspond to the key events in the game.

[0067] To comprehensively evaluate the performance of the present invention, the accuracy, recall rate, and F-value metrics of the generated video key frame set are used as objective criteria to evaluate the quality of the abstract. The calculation formulas for the three objective evaluation criteria for video key frames, accuracy (P), recall rate (R), and F-value (F), are as follows:

[0068]

[0069] In the formula, N matched represents the matching degree between the key frames automatically extracted by the algorithm and the manually extracted key frames in the dataset; N AS represents the number of key frames automatically extracted; N US represents the number of key frames manually extracted.

[0070] Precision can reflect the ability of the algorithm to automatically extract key-frame matching data in the dataset of manually extracted key frames. Recall can reflect the proportion of the number of matched key frames to the number of manually extracted key frames. At the same time, these two criteria restrict each other. Generally, a higher recall rate corresponds to a lower precision, and vice versa. Therefore, to comprehensively measure precision and recall, the F-value is selected as an indicator to evaluate the comprehensive quality of video summaries.

[0071] In addition, a method for extracting competition shots of team sports videos based on color features is proposed. The specific implementation steps are as follows:

[0072] Input: Original team sports video

[0073] Output: Set of video clip segments of competition shots

[0074] Definition:

[0075] S = {S 1 , S 2 , ……, S S}: Set of video clip segments after video segmentation;

[0076] F: Set of video frames of video clip segments;

[0077] AM R : Arithmetic mean of the R-channel values corresponding to all pixel points in the video frame;

[0078] AM G : Arithmetic mean of the G-channel values corresponding to all pixel points in the video frame;

[0079] AM B : Arithmetic mean of the B-channel values corresponding to all pixel points in the video frame;

[0080] R R : Range of AM R in the video clip segment;

[0081] R G : Range of AM G in the video clip segment;

[0082] R B : Range of AM B in the video clip segment;

[0083] Arithmetic mean of AM R corresponding to all video frames in the video clip segment;

[0084] Arithmetic mean of AM G corresponding to all video frames in the video clip segment;

[0085] The arithmetic mean value corresponding to all video frames in the lens video clip; B ;

[0086] A RGB : The array stacked in row order;

[0087] C max : The class with the most lens video clips after K-means clustering;

[0088] TH BS : The threshold for distinguishing candidate reference lenses from all other lenses;

[0089] CBS: The set of candidate reference lens video clips;

[0090] S FM : The lens video clip with the most video frames in CBS;

[0091] The lens video clip S FM corresponding

[0092] The lens video clip S FM corresponding

[0093] The lens video clip S FM corresponding

[0094] with ;

[0095] with ;

[0096] with ;

[0097] TH GS : The threshold for distinguishing game lenses from non-game lenses;

[0098] S’: The set of game lens video clips.

[0099] Step 1: Segment the video into lens video clips and save the video frame by frame in picture form to obtain S

[0100] Step 2: Calculate the AM of all video frames R , AM G,AM B

[0101] Step 3: Calculate R of all shot video clips R ,R G ,R B ; and A RGB

[0102] Step 4: Perform K-means clustering on A corresponding to all shot video clips and find C RGB and find C max

[0103] Step 5: Perform the following operations on all shot video clips in C: If at least one of R max in C is less than or equal to TH R ,R G ,R B , then store this shot video clip into CBS BS , then store this shot video clip into CBS

[0104] Step 6: Find

[0105] Step 7: Calculate of all shot video clips and perform the following operations on all shot video clips: If are all less than or equal to TH GS , then store it into S'.

[0106] Value: TH BS = 100, TH GS = 15%.

[0107] Using the mean property in the ImageStat module of the Python image processing library PIL, the average value AM of the pixel values of the three channels in each frame within a single shot video clip can be calculated R ,AM G ,AM B . The formula is as follows:

[0108]

[0109] In the formula, r, g, and b are the RGB color pixel values corresponding to the pixel points in the video frame, and n is the number of pixel points

[0110]

[0111] In the formula, N is the number of video frames contained in a single shot video clip

[0112]

[0113] The present invention uses color features to extract the game shots of team sports videos, and uses the main color difference of the entire video frame to eliminate close - up and off - field shots, while retaining the main game shots in the video. In the existing related work, the classification of shots often focuses on the venue itself. This method does not stick to this conventional idea and proposes a new idea. First of all, this method is not limited by venue factors and can be applied to multiple team sports. At the same time, this method can also well make up for some misdetections in the shot boundary detection method: if there is an incorrect segmentation in a non - game shot, the entire shot will also be recognized and removed, without affecting the subsequent research and analysis of the video. Therefore, this method can extract the most important game shot part in the entire team sports video, laying a good structured foundation for subsequent semantic event detection and summary generation.

[0114] To test the effectiveness of this method, this method is applied to the SportKF dataset, and the recall rate and precision rate of 5 football videos are compared with the method proposed by Yu Junqing et al. The calculation formulas for the recall rate and precision rate are as follows:

[0115] Recall rate=(number of actually detected - number of misdetected) / number of should - be - detected

[0116] Precision rate=(number of actually detected - number of misdetected) / number of actually detected

[0117] For the extracted game shots, the recall rate of the four team sports is 100%, and the precision rate is 86.67%. For football videos, the recall rate of the method proposed by Yu Junqing et al. is 93.2%, while the recall rate of this method reaches 100%; the precision rate of the method proposed by Yu Junqing et al. is 88.7%, while the recall rate of this method reaches 93.4%.

[0118] The experimental results show that this method has achieved good results in the detection and recognition of game shots, and has improved in both the recall rate and precision rate compared with the relatively representative method proposed by Yu Junqing et al. Description of the Drawings

[0119] Figure 1 Main Flow Chart

[0120] Figure 2 Schematic Diagram of Shot Translation and Zoom

[0121] Figure 3 Schematic Diagram of Candidate Key Frame Extraction

[0122] Figure 4 Hierarchical Clustering Dendrogram

[0123] Figure 5 Flow Chart of Representative Key Frame Extraction

[0124] Figure 6(a) Frame No. 12126 (b) Frame No. 12127

[0125] Figure 7 (a) Frame No. 11263 (b) Frame No. 11274

[0126] Figure 8 (a) Global optical flow map corresponding to Frame No. 11263

[0127] (b) Global optical flow map corresponding to Frame No. 11274

[0128] Figure 9 (a) Global optical flow map in HSI color space corresponding to Frame No. 11263

[0129] (b) Global optical flow map in HSI color space corresponding to Frame No. 11274

[0130] Figure 10 (a) 2D histogram of the HI channel of the global optical flow map corresponding to Frame No. 11263

[0131] (b) 2D histogram of the HI channel of the global optical flow map corresponding to Frame No. 11274

[0132] Figure 11 Flowchart

[0133] Figure 12 Video cover of the SportKF dataset

[0134] Figure 13 Long shot

[0135] Figure 14 LOGO shot Detailed implementation method

[0136] Taking an NBA game video with a duration of 7 minutes, a frame width of 640, a frame height of 360, and a frame rate of 30 frames per second as an example (the following specific values are all rounded to 2 decimal places).

[0137] Step 1:

[0138] Input: Original team sports video (this example contains a total of 33 shots and 12,593 frames)

[0139] Output: A set of video clips of game shots without relatively redundant content such as advertisements and close-up shots (including 11 shots numbered 2, 4, 6, 9, 13, 15, 17, 21, 25, 27, 31, all of which are game shots and there is no missed detection).

[0140] Step 2:

[0141] Input: Video frames of the video clips of game shots

[0142] Output: Global optical flow field of video frames

[0143] Optical flow is the instantaneous velocity of the pixel motion of a spatial moving object on the observation imaging plane. It uses the change of pixels in the time domain in the image sequence and the correlation between adjacent frames to find the corresponding relationship between the previous frame and the current frame, and thus calculates the motion information of the object between adjacent frames. For example, if two consecutive frames are input, Figure 6 The left figure is Frame No. 12126, Figure 6 The right figure is Frame No. 12127. Then, by comparison, the global optical flow field corresponding to Frame No. 12127 can be obtained.

[0144] In this example, the global optical flow field of Frame No. 12127 is a two-dimensional array of 640×360, where each array corresponds to the global optical flow value of the pixel value at that point. Taking the global optical flow values of the four corner points of this optical flow field as an example: upper left (2.38, 1.73), lower left (2.38, -2.07), lower right (-1.58, -2.07), upper right (-1.58, 1.73).

[0145] Step 3:

[0146] Input: Global optical flow values of four corner points of video frames

[0147] Output: Global motion statistical features (including the zoom amount and translation amount of the lens)

[0148] Taking Frame No. 12127 as an example, the global optical flow values of its four corner points are: upper left (2.38, 1.73), lower left (2.38, -2.07), lower right (-1.58, -2.07), upper right (-1.58, 1.73). That is, x tl = 2.38, y tl = 1.73, x br = -1.58, y br = -2.07,

[0149] According to formula (1), X can be calculated zoom ≈ 1.98

[0150] According to formula (2), Y can be calculated zoom ≈ 1.90

[0151] According to formula (3), MAG can be calculated zoom ≈ 2.74

[0152] According to formula (4), X can be calculated trans ≈ 0.40

[0153] According to formula (5), Y can be calculated tran ≈ -0.17

[0154] DIS can be calculated according to formula (6). trans ≈0.43

[0155] Step 4:

[0156] Input: Set of game shot video clips

[0157] Output: Set of fine-grained video clips

[0158] Trans_con = SGN(X trans (Frame i )·X trans (Frame i+1 ))

[0159] Taking the 12126th frame and the 12127th frame as an example, the X corresponding to the 12126th frame trans = -0.18, while the X corresponding to the 12127th frame trans = 0.40, then the product of x trans (12126)·x trans (12127) is negative. Therefore, Trans_con = -1 (the SGN(x) function returns the sign of x. If x > 0, then SGN(x) = 1; if x = 0, then SGN(x) = 0; if x < 0, then SGN(x) = -1). Then, the 12126th frame and the 12127th frame need to be separated. The 12126th frame belongs to the 17th fine-grained segment of shot No. 30, while the 12127th frame belongs to the 18th fine-grained segment of shot No. 30. The corresponding event is a fast break by the opponent after a mistake, with an offense-defense transition.

[0160] Step 5:

[0161] Input: All video frames of the set of fine-grained video clips

[0162] Output: Set of candidate key frames

[0163] Taking Figure 3 the curve from the 350th frame to the 480th frame corresponding to Figure 3 (the complete curve shown) as an example, the 360th frame corresponding to the "jump ball" event is extracted. Find f 1 = 359, f 3 = 360, f 2 = 361. At this time, MAG zoom (f 1 ) = 3.78, MAG zoom (f 3 ) = 6.10, MAG zoom (f 2 ) = 3.54, then f 3= 360 is MAG zoom (f) The frame number corresponding to a maximum value of the curve; DIS trans (f 1 ) = 3.59, DIS trans (f 3 ) = 5.69, DIS trans (f 2 ) = 3.31, then f 3 = 360 is DIS trans (f) A maximum value of the curve, so f 3 = 360 is both the frame number corresponding to the maximum value of the MAG zoom (f) curve and the frame number corresponding to the maximum value of the DIS trans (f) curve, then extract the 360th frame and store it in the candidate key frame set.

[0164] In this example, med is set to 20. When traversing the first fine-grained video segment of the shot with serial number 20, find k 1 = 7107, k 2 = 7191, then k 2 - k 1 ≥ med, so extract f m = (7107 + 7191 | 2 + 1 = 7150's video frame (corresponding to the "free throw" event) and store it in the candidate key frame set according to the frame number sequence.

[0165] Step 6:

[0166] Input: Candidate key frame set

[0167] Output: Final key frame set

[0168] Extraction of representative key frames based on spatio-temporal consistency:

[0169] In the candidate key frame set, take the 11263rd frame and the 11274th frame in the 9th fine-grained segment of the video segment of the shot with serial number 30 as an example. These two frames are adjacent candidate key frames.

[0170] Then the color-coded image is as Figure 8 shown.

[0171] The image after conversion to the HSI color space is as Figure 9 shown.

[0172] The two-dimensional histogram of the HI channel corresponding to the global optical flow image in the HSI color space is as Figure 10 shown.

[0173] Calculate the Bhattacharyya distance B = 0.38 between the two-dimensional histograms of the HI channels of the global optical flow maps of these two adjacent candidate key frames;

[0174] In this example, TH B is set to 0.22, and since 0.38 > 0.22, the 11263rd frame (corresponding to the end of the two-point layup in the event) and the 11274th frame (corresponding to the start of the rebound in the event) are extracted.

[0175] Extraction of representative key frames based on hierarchical clustering:

[0176] Taking the 9th fine-grained segment of the video clip of shot No. 30 as an example, its clustering process is as Figure 4 shown. In this example, TH HC is set to 4.00, that is, the inter-cluster distance threshold is set to 4.00. Figure 4 As shown by the dashed line in the figure, it can be divided into 5 clustering clusters. From left to right, the first cluster contains the corresponding arrays of 8 global optical flow maps with corresponding frame numbers 11111, 11116, 11121, 11124, 11138, 11260, 11263, 11274. The second cluster contains the corresponding array of 1 global optical flow map with the corresponding frame number 11150. The third cluster contains the corresponding arrays of 2 global optical flow maps with corresponding frame numbers 11221 and 11241. The fourth cluster contains the corresponding array of 1 global optical flow map with the corresponding frame number 11173. The fifth cluster contains the corresponding arrays of 4 global optical flow maps with corresponding frame numbers 11155, 11197, 11189, 11210.

[0177] In the first cluster, the key frame number corresponding to the array with the smallest distance to each array is 11260. In the second cluster, the corresponding frame number is 11150. In the third cluster, the corresponding frame number is 11221. In the fourth cluster, the corresponding frame number is 11173. In the fifth cluster, the corresponding frame number is 11155.

[0178] After extraction, it is detected whether the first and last key frames of the 9th fine-grained segment of the video clip of shot No. 30 are extracted, that is, the corresponding frame numbers are 11111 and 11274. These two frames are not extracted, so they are extracted.

[0179] Taking the union of the above-extracted key frames, such as the 11274th frame, which is extracted in the extraction of representative key frames based on spatio-temporal consistency and is also extracted in the extraction of representative key frames based on hierarchical clustering. Therefore, the duplicate frames are removed, and the final key frame set can be obtained, which contains 523 frames, accounting for 4.15% of the original 12593 frames, that is, 95.85% of the redundant frames are removed.

[0180] The present invention (KEGMS) is compared with two representative key frame extraction schemes, including the algorithm based on motion features proposed by Mendi et al. and the algorithm based on color features (SFKE) proposed by Vennila et al. The present invention provides two sets of results, one set with a high F value (TH B = 0.26, THHC = 3.80); Another group has a high recall rate (TH B = 0.22, TH HC = 4.00), and weakens the F value. The comparison results are shown in Table 1.

[0181] Table 1 Comparative Experiments

[0182] P R F Mendi 0.07 0.85 0.129 SFKE 0.30 0.60 0.398 KEGMS (High F-value) 0.31 0.76 0.441 KEGMS (High recall) 0.10 0.98 0.174

[0183] It can be found that the KEGMS with a high recall rate obtains the highest recall rate of 0.98, which is at least 0.13 higher than the compared algorithms. The KEGMS with the optimal threshold obtains the maximum F value of 0.441, which is at least 0.043 higher than the existing algorithms. This experiment shows that the introduction of motion information helps in the extraction of key frames. Using global motion can further improve the performance.

[0184] In addition, a method for extracting competition shots of team sports videos based on color features is also proposed. The specific implementation steps are as follows:

[0185] Taking an NBA game video with an input duration of 7 minutes, a frame width of 640, a frame height of 360, and a frame rate of 30 frames per second as an example. Among them, the long shots are as Figure 13 shown, and the LOGO shots are as Figure 14 shown (the following specific values are all taken to 2 decimal places, rounded).

[0186] Step 1: Segment the video into video shot segments to obtain S = {S 1 , S 2 , ……, S 33}. (Including a total of 33 shots)

[0187] Step 2: Calculate the AM R , AM G , AM B . For example, the corresponding AM Figure 13 for R = 100.50, AM G = 80.34, AM B = 74.88;

[0188] Step 3: Calculate the R R , R G , R B ; and A RGB

[0189] Figure 13 The R R corresponding to the shot where G = 37.22, RB = 40.65; A RGB = (110.05, 90.39, 79.10)

[0190] Figure 14 The R corresponding to the shot (serial number 12) R = 13.84, R G = 12.63, R B = 1.80; A RGB = (105.70, 90.43, 74.45)

[0191] Step Four: Perform K-means clustering on all shot video segments corresponding to A RGB and find C max In this example, C max contains 9 shots: 4, 6, 9, 15, 17, 20, 21, 27, 31

[0192] Step Five: Perform the following operations on all shot video segments in C max : If at least one of R R , R G , R B is less than or equal to TH BS , then store this shot video segment in CBS

[0193] In this example, TH BS is set to 100. It can be seen that for the shot with serial number 17 in C max , the corresponding R R = 40.13, R G = 28.46, R B = 19.81 satisfies the condition "If at least one of R R , R G , R B is less than or equal to TH BS ", so it is stored in CBS; on the contrary, for the shot with serial number 20 in C max , the corresponding R R = 212.30, R G = 216.18, R B = 196.36 does not satisfy the condition "If at least one of R R , R G , R B is less than or equal to TH BS ", so it is not stored in CBS; in this example, the result of Step Five, CBS, contains 8 shots: 4, 6, 9, 15, 17, 21, 27, 31

[0194] Step Six: Locate

[0195] In this example, among the 8 shots included in CBS, the shot with the largest number of frames is the shot numbered 31, which contains 1,895 frames. Therefore, the corresponding

[0196] Step Seven: Calculate the for all video segments of the shots and perform the following operations on all video segments of the shots: If are all less than or equal to TH GS , then store it in S’

[0197] In this example, TH GS is set to 15%. For example, for the shot where Figure 13 is located (shot number 17), the corresponding Δ satisfies the condition of " are all less than or equal to TH GS ", so it is stored in S’. In this example, the final result, the set S’ of the competition shot video segments, contains 11 shots numbered 2, 4, 6, 9, 13, 15, 17, 21, 25, 27, and 31. All of them are competition shots, and there is no missed detection.

Claims

1. A method for extracting key frames from team sports videos based on global motion statistical features, characterized in that, it includes the following steps: Step 1: Extract the game shot video segments through a method for extracting game shot videos of team sports videos based on color features; Input: Original team sports video Output: Set of game shot video segments; Step 2: Extract the mixed optical flow of the video frames in each shot video segment and estimate the global motion; Input: Video frames of the game shot video segments Output: Global optical flow field corresponding to the video frames; Step 3: Calculate the global motion statistical features corresponding to the video frames, including the lens scaling amount and the translation amount; Input: Global optical flow values of the four corner points of the video frame Output: Global motion statistical features, including the lens scaling amount and the translation amount; Among them, rectangle P tl P bl P br P tr represents the original lens range, and rectangle P’ t1 P’ b1 P’ br P′ tr represents the transformed lens range; where X zoom , Y zoom are respectively the scaling vectors of the lens in the X and Y directions, and MAG zoom is the combined scaling amount of the lens; (x tl , y tl ), (x br , y br ) are respectively the global motion vectors of the upper left and lower right corner points of the video frame; X trans = x tl -X zoom Y trans = y tl -Y zoom where X trans , Y trans are respectively the translation vectors of the lens in the X and Y directions, and DIS trans is the combined translation amount of the lens; Step 4: Further divide the shot video segments into fine-grained video segments based on the horizontal translation amount of the lens; Input: Set of game shot video segments Output: Set of fine-grained video segments Trans_con = SGN(X trans (Frame i )·X trans (Frame i+1 )) Where Trans_con represents the conversion in the direction of the global horizontal translation amount, the SGN(x) function returns the sign of x. If x > 0, then SGN(x) = 1; if x = 0, then SGN(x) = 0, indicating that the direction of the horizontal translation amount of the lens has not changed; if x < 0, then SGN(x) = -1, indicating that the direction of the horizontal translation amount of the lens has changed; X trans (Frame i ),X trans (Frame i+1 ) are the translation vectors of the corresponding lenses in the X direction for the i-th frame and the (i + 1)-th frame respectively; Specific operation: In each shot video segment, calculate Trans_con corresponding to the i-th frame and the (i + 1)-th frame one by one. If Trans_con = -1, then separate the i-th frame and the (i + 1)-th frame until the second-to-last video frame of each shot video segment is calculated; then save the frames before the first interval as the first fine-grained video segment, then save the frames between every two intervals as the corresponding fine-grained video segments, and the last fine-grained video segment consists of all the frames after the last interval; Step 5: Draw the curves of the combined translation amount and the combined scaling amount of the lens, and extract the candidate key frames in each fine-grained video segment according to the maximum value of the curve; Input: All video frames of the set of fine-grained video segments Output: Set of candidate key frames Step 1): For each fine-grained video segment, plot the curve MAG of the change in the scaling factor of the video frame sequence zoom (f) and the curve DIS of the change in the translation amount trans (f), where f is the frame number corresponding to the video frame sequence; Step 2): Simultaneously scan the curves of the combined translation amount and the combined scaling amount; taking the curve of the combined scaling amount as an example, find f 1 , f 3 , f 2 as three consecutive video frames. If MAG zoom (f 3 ) > MAG zoom (f 1 ) and MAG zoom (f 3 ) > MAG zoom (f 2 ), then f 3 is the frame number corresponding to a maximum value of the MAG zoom (f) curve; the DIS trans (f) curve is the same; if the f 3 with the same frame number is both the frame number corresponding to the maximum value of the MAG zoom (f) curve and the frame number corresponding to the maximum value of the DIS trans (f) curve, then the frame corresponding to f 3 is selected as the key frame; Step three): Repeat step two) until the search is completed; Step 4): Traverse the initially extracted set of key frames to find two frames k 1 , k 2 . If k 2 -k 1 ≥med, then extract the video frames with frame numbers ( denotes rounding x downwards) from the original fine-grained video clip frame set, and store them in the candidate key frame set arranged in ascending order of frame numbers; Value of med: Basketball: 35 Football: 88 Rugby: 53 Hockey: 160 Step five): Repeat step four) until the search is completed; Step 6: Extract representative key frames from the set of candidate key frames by combining spatio-temporal consistency and hierarchical clustering to obtain the final set of key frames; Input: Set of candidate key frames Output: Final set of key frames Extraction of representative key frames based on spatio-temporal consistency: Step one: Color-code the global optical flow fields corresponding to all candidate key frames to obtain global optical flow images; Step two: Convert all global optical flow images from the RGB color space to the HSI color space; Step three: Sort all global optical flow images in the HSI color space according to the time series of the corresponding key frames; Step four: Calculate the two-dimensional histogram of the HI channel corresponding to all global optical flow images in the HSI color space; Step five: Calculate the Bhattacharyya distance B between the two-dimensional histograms of the HI channels of adjacent images; The Bhattacharyya distance is used to measure the probability distributions of two discrete data; in the calculation of histogram similarity, the Bhattacharyya distance has good effects, and its formula for calculating histogram similarity is as follows: where h 1 and h 2 are two histograms, n is the number of histogram groups, and h i [m] represents the frequency corresponding to the m-th group; the smaller the value, the higher the correlation, 0 for a perfect match, and 1 for a complete mismatch; Step 6: If B > TH B , check whether the key frames corresponding to these two adjacent images have been extracted. If not, extract the unextracted key frames; Extraction of representative key frames based on hierarchical clustering: Step seven: Group all global optical flow images according to the fine-grained video segments where the corresponding key frames were previously located, and sort them according to the time series; Step 8: Calculate the RGB three-channel histograms corresponding to each global optical flow image, and horizontally splice them into an array containing 768 eigenvalues; Step Nine: Perform hierarchical clustering on the arrays corresponding to all global optical flow images in each group, and set the inter-cluster distance threshold to TH HC ; Step 10: Extract the key frames corresponding to the arrays with the smallest distance to each array in each cluster after clustering; Step 11: Detect whether the first and last candidate key frames of each fine-grained video segment are extracted. If not, extract the unextracted key frames; Step 12: Take the union of the above-extracted key frames and remove the duplicate key frames; obtain the final set of key frames; TH B ,TH HC Recommended value: High recall rate: Basketball: TH B = 0.22, TH HC = 4.00 Football: TH B = 0.36, TH HC = 8.00 Rugby: TH B = 0.17, TH HC = 9.00 Hockey: TH B = 0.14, TH HC = 3.50 High F value: Basketball: TH B = 0.26, TH HC = 3.80 Football: TH B = 0.40, TH HC = 7.00 Rugby: TH B = 0.20, TH Hc = 6.50 Hockey: TH B = 0.24, TH Hc = 3.00 When counting the technical data of the competition, the high recall rate mode is adopted to ensure the data accuracy; while when quickly browsing the competition content and storing the video, the high F value mode is adopted.

2. A method for extracting key frames of team sports videos based on global motion statistical features according to claim 1, characterized in that, Step 1 is specifically as follows: S = {S 1 , S 2 ,......, S S}: The set of shot video segments after video segmentation; F: The set of video frames of the lens video segment; AM R : The arithmetic mean of the R-channel values corresponding to all pixel points in the video frame; AM G : The arithmetic mean of the G-channel values corresponding to all pixel points in the video frame; AM B : The arithmetic mean of the B-channel values corresponding to all the pixels in the video frame; R R : Range of AM in the lens video segment R ; R G : Range of AM in the lens video segment G ; R B : Range of AM in the lens video segment B ; The arithmetic mean corresponding to all video frames in the lens video clip for AM R ; The arithmetic mean corresponding to all video frames in the lens video segment for AM G ; The arithmetic mean corresponding to all video frames in the lens video clip for AM B ; A RGB : An array stacked in row order; ​ C max : The class with the most video clips of shots after K-means clustering; TH BS : A threshold value used to distinguish candidate reference lenses from all other lenses; CBS: The set of candidate reference lens video segments; S FM : The video clip with the largest number of video frames in the CBS; Lens video clip S FM corresponding Lens video segment S FM corresponding Lens video segment S FM corresponding with percentage error; and percentage error; with percentage error; TH GS : Threshold for distinguishing between game shots and non-game shots; S’: The set of competition lens video segments; Step 1: Split the video into video lens segments, and save the video frame by frame in the form of pictures to obtain S Step 2: Calculate the AM of all video frames R , AM G , AM B Step 3: Calculate the R of all video segments of the shots R , R G , R B; and A RGB Step 4: Perform K-means clustering on all A corresponding to the video segments of the shots RGB and find C max Step Five: Perform the following operations on all the video clips of the shots in C max : If at least one of R R, R G, R B is less than or equal to TH BS , then store the video clip of this shot into CBS Step Six: Locate Step Seven: Calculate for all the video clips of the shots and perform the following operations on all the video clips of the shots: If are all less than or equal to TH GS , then store them in S'; Value: TH BS = 100, TH GS = 15%.

Citation Information

Patent Citations

  • Video key frame extraction method based on multi-characteristic fusion shot clustering

    CN107220585A

  • Video abstract extraction method, video abstract extraction system, video abstract extraction device and storage medium

    CN110381392A