Failure sensing method for player tracking in football video and program product

By introducing technical means such as feature extraction, pose estimation and template library update in player tracking in football videos, the problem of insufficient accuracy of failure perception in the existing technology is solved, and the accuracy of player tracking is improved.

CN120182884APending Publication Date: 2025-06-20HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510184571.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The player tracking method in existing football videos cannot fully consider the sports characteristics of football players in actual games, resulting in insufficient accuracy of failure perception, affecting the accuracy of player tracking.

Method used

By introducing technical means such as feature extraction, pose estimation and template library update in player tracking, combined with attention mechanism and Kalman filtering algorithm, the accuracy of failure perception is improved.

Benefits of technology

It improves the accuracy of failure perception during player tracking in football videos, reduces tracking drift caused by occlusion and disappearance, and improves the overall accuracy of player tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182884A_ABST
    Figure CN120182884A_ABST
Patent Text Reader

Abstract

The invention discloses a failure sensing method for player tracking in a football video and a program product, and belongs to the field of video analysis, and the method comprises the steps: obtaining a target frame set corresponding to a search region in a current frame; selecting target frames from the target frame set as positioning frames in sequence according to the descending order of the confidence degrees, removing the target frames which are repeated with the selected positioning frames and have the confidence degrees lower than that of the positioning frames in the target frame set, and performing attitude estimation on the remaining target frames to obtain human body skeleton models corresponding to the target frames; calculating the number Nsum of key points contained in the human body skeleton model, if Nsum is not greater than Nsum, if Nsum is greater than Nsum; if so, judging that the current frame is shielded; 0 lt; eta < lt >; 1; n represents the number of target frames; nper represents the number of standard key points of the human skeleton model. According to the method, the motion characteristics of the football player in the actual match can be fully considered, the accuracy of failure perception in the football player tracking process in the football video is improved, and then the accuracy of football player tracking is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video analysis, and more specifically, relates to a failure perception method and program product for player tracking in football videos. Background Art

[0002] With the continuous development of network media platforms, more and more viewers choose to watch football games through video live broadcasts. In order to meet the different viewing needs of different viewers, football game videos can be assisted and analyzed by target tracking technology, so as to provide more diverse content. Ordinary viewers pay more attention to the critical moments during the game, such as the performance of players when dribbling, running positions, and shooting. This requires tracking the target, and relying on the tracking results, more advanced semantic analysis tasks can be carried out, such as highlight shots, player action recognition, and video summarization. For referees, the ball possession and fouls during the game require additional auxiliary information to ensure the fairness of the game. Player tracking technology can help referees track players suspected of fouls and conduct a detailed analysis of their behaviors. For professional personnel such as players and coaches, analyzing the data generated during the game can help improve the players' sports skills and also help them make more informed tactical decisions during the game. Therefore, player tracking is a basic task for many practical applications in football videos, and it is of great significance for enhancing the broadcast and entertainment value of football games.

[0003] Currently, the most effective technology for tracking target players is video analysis technology. Mainstream tracking algorithms can already perform efficient tracking in general scenarios, but target tracking in sports videos, especially player tracking in football game scenarios, faces some difficulties and challenges, including:

[0004] (1) Frequent player occlusion and crossing. During the game, players move violently and interact frequently, and occlusion occurs from time to time, making it difficult for the tracker to locate the target position. In addition, when players cross, the tracker also needs to handle the problem of target confusion in the crossed part.

[0005] (2) Interference from similar players. Players on the field all wear similar jerseys, and each player, especially players on the same team, has a similar appearance to the target player. Moreover, most of the shots during the game are long shots, and the proportion of players in the picture is small, making it difficult to capture the characteristics of players, causing difficulties in distinguishing similar players and easily resulting in tracking drift.

[0006] (3) Difficult repositioning. When the target is occluded and moves out of the picture, the tracker completely loses the target, and at this time, the target needs to be repositioned. The movement of the camera and the interference of other players all pose obstacles to repositioning.

[0007] (4) High algorithm complexity requirements. Live video broadcast has the characteristics of real-time, requiring the tracker to obtain accurate results in a timely manner, which puts higher requirements on the running speed and accuracy of the tracking algorithm.

[0008] The above difficulties and challenges make it particularly important to effectively perceive failures such as player occlusion and disappearance during player tracking. To this end, some studies have deeply analyzed the motion characteristics of players in football game videos and proposed a tracking algorithm that can perceive and correct failure states. In failure perception, the encoder based on the attention mechanism and the improved decoder are combined to enhance the feature representation ability. The drift is reduced by limiting the target scale and adding a pre-window. A dual search domain detection method is designed. At the same time, the position, confidence and movement direction information are integrated to perceive the occurrence of player occlusion and disappearance with almost no additional calculation. This method effectively improves the accuracy of player tracking in football videos. However, when the proposed dual search domain algorithm is used for occlusion perception, it mainly judges whether occlusion occurs based on the change in the number of players in the two defined search domains. In some scenarios, the players are concentrated in a certain area of ​​the stadium, the interaction between the players increases, and a large number of gatherings and occlusions occur. The detection of the number of players is difficult and the detection accuracy cannot be guaranteed. On the other hand, this method mainly perceives disappearance based on the confidence of targets at edge positions and the direction of movement of targets between two adjacent frames. However, in actual games, there is great uncertainty in the direction of movement of the football on the field in a short period of time. Correspondingly, there will also be great uncertainty in the position and direction of movement of the players. This disappearance perception method cannot fully consider the actual movement characteristics of the players, and there will be certain errors in the judgment results of the disappearance state.

[0009] In general, existing player tracking methods in football videos cannot fully consider the movement characteristics of football players in actual games, and the accuracy of failure perception needs to be further improved. Summary of the invention

[0010] In view of the defects of the prior art and the need for improvement, the present invention provides a method and program product for failure perception of player tracking in football videos, the purpose of which is to fully consider the movement characteristics of football players in actual games, improve the accuracy of failure perception in the process of player tracking in football videos, and thus improve the accuracy of player tracking.

[0011] To achieve the above object, according to one aspect of the present invention, a method for detecting player tracking failure in a football video is provided, comprising:

[0012] S1: For the current frame, extract features from the search region therein to obtain search region features, and use the template feature as a convolution kernel to perform a convolution operation on the search region features to obtain a response map; the template feature is the feature of the tracking result of the target to be tracked in the historical frame.

[0013] S2: Regress the response map to obtain target bounding boxes and the confidence levels corresponding to each target bounding box, and remove the target bounding boxes with confidence levels lower than a preset first threshold to obtain a set of target bounding boxes.

[0014] S3: Select the target bounding box with the highest confidence level from the set of target bounding boxes as the localization bounding box.

[0015] S4: For each target bounding box in the current set of target bounding boxes with a confidence level lower than that of the localization bounding box, calculate the intersection over union (IoU) with the localization bounding box respectively. If the IoU is greater than a preset second threshold, then determine the corresponding target bounding box as a duplicate target bounding box and remove it from the set of target bounding boxes.

[0016] S5: If there are still target bounding boxes in the set of target bounding boxes that have not been selected as the localization bounding box, then select the next target bounding box from the set of target bounding boxes in descending order of confidence level as the localization bounding box, and go to S4; otherwise, use the current set of target bounding boxes as the set of candidate target bounding boxes, and go to S6.

[0017] S6: Perform pose estimation on each target bounding box in the set of candidate target bounding boxes to obtain a human skeleton model corresponding to each target bounding box, and count the number of key points N included in the human skeleton model. sum , if N sum < η × n × N per , then determine that occlusion occurs in the current frame.

[0018] Among them, 0 < η < 1; n represents the number of target bounding boxes in the set of candidate target bounding boxes; N per represents the standard number of key points of the human skeleton model.

[0019] Further, S6 also includes:

[0020] For each target bounding box in the set of candidate target bounding boxes, respectively determine whether it is located at the edge of the current frame; for each target bounding box located at the edge of the current frame, perform the following operations:

[0021] T1: Denote the target position in the target bounding box as Pos i , and denote the positions of the template feature in the previous F frames as Pos i-1 , Pos i-2 , …… Pos i-F , and calculate the minimum distance from each position to the frame border of the current frame min(|bbox - Pos i |), min(|bbox - Posi-1 |), min(|bbox - Pos i-2 |), …… min(|bbox - Pos i-F |);

[0022] T2: If min(|bbox - Pos i |) < min(|bbox - Pos i - 1|) < min(|bbox - Pos i-2 |) < …… < min(|bbox - Pos i-F |), then it is determined that disappearance occurs in the current frame;

[0023] Where F is an integer greater than or equal to 2.

[0024] Further, the method for determining whether the target box is located at the edge of the frame includes:

[0025] Calculate the minimum distance dist between the target position (x, y) in the target box and the frame border according to dist = min(|x|, |y|, |W - x|, |H - y|). If dist < γ, then it is determined that the target box is located at the edge of the frame;

[0026] Where W and H respectively represent the width and height of the frame; γ represents a preset critical value, and γ > 0.

[0027] Further, after S6, it further includes: If disappearance does not occur in the current frame but occurs in the previous frame, then perform disappearance correction; The disappearance correction includes:

[0028] Perform target detection on the players in the current frame to obtain the detection positions of each player in the current frame;

[0029] Use the Kalman filter algorithm to predict the positions of each player in the previous frame in the current frame as the predicted positions of the corresponding players;

[0030] Taking the Euclidean distances between each detection position and each predicted position as elements, obtain the cost matrix D;

[0031] Perform the Hungarian algorithm on the cost matrix D to obtain the row index list row_i and column index list col_i of the successfully matched players in the previous frame and the current frame;

[0032] If the index corresponding to the player in the current frame does not appear in col_i, then put the player into the set unmatched_det of reappearing players;

[0033] Traverse the set unmatched_det of players. For each player traversed, if the player meets the preset conditions, then determine the player as a target that reappears after disappearance;

[0034] Among them, the preset conditions include: the target box corresponding to the player is located in the search area of the current frame and at the edge of the current frame, and at the same time, the similarity between the player target box and the previously disappeared player target box is greater than the preset similarity threshold.

[0035] Furthermore, S4 also includes: maintaining a set of interference boxes for the current positioning box; if the intersection over union between the positioning box and the target boxes in the target box set with a confidence level lower than that of the positioning box is not greater than the second threshold, then determine the target box as an interference box and add it to the set of interference boxes of the positioning box;

[0036] Moreover, the method for perceiving the failure of player tracking in a football video further includes: if there is no occlusion in the current frame while there is occlusion in the previous frame, then perform occlusion correction; the occlusion correction includes:

[0037] Locate the last frame before the current frame where there was no occlusion, use it as the occlusion start frame, and obtain the set of interference boxes B corresponding to the tracking result of the template feature in the occlusion start frame;

[0038] For the current frame, denote its set of candidate target boxes as the set of candidate target boxes U. For each target box in the set of candidate target boxes U, calculate the average feature similarity between it and each target box in the set of interference boxes B as the interference index of the corresponding target box;

[0039] Determine the target box with the minimum interference index in the set of candidate target boxes as the tracking result of the template feature in the current frame.

[0040] Furthermore, in S2, before performing regression on the response map, it also includes:

[0041] Weight the response map with a Gaussian window function; the Gaussian window function is centered at the geometric position center of the response map.

[0042] Furthermore, the method for perceiving the failure of player tracking in a football video provided by the present invention further includes: maintaining a template library with a size not exceeding L for storing the features of the tracking results of the target to be tracked in one or more frames; and, when tracking the current frame, the template feature is the feature constructed using the features in the template library;

[0043] Moreover, the method for perceiving the failure of player tracking in a football video further includes:

[0044] If the sequence number i of the current frame differs from the frame sequence number corresponding to the last update of the template library by M, then perform template library update; the template library update includes:

[0045] Set the length to m = floor(Mδ / (δ max -δ min)) and make it cover the current frame and its previous m - 1 frames;

[0046] Among the tracking results of the template feature in each frame within the window, add the target bounding boxes with confidence greater than the preset third threshold to the positive tracking result set;

[0047] The weights of each frame within the initial window j represents the frame sequence number within the window, 1 ≤ j ≤ m; the closer the frame is to the current frame, the higher its initial weight;

[0048] For each frame within the window, calculate the average feature difference value D between the tracking result of the template feature in this frame and the tracking results in the remaining frames within the window cur-j , and add the frames with feature difference values less than the preset fourth threshold to the reliable result set;

[0049] According to Update the weights of each frame in the reliable result set; is the updated weight of the j - th frame within the window;

[0050] Add the tracking results of the frames in the reliable result set with weights greater than the preset fifth threshold to the reverse tracking result set;

[0051] Sequentially add the features of the target bounding boxes located in the intersection of the positive tracking result set and the reverse tracking result set to the template library;

[0052] Each time a feature is added to the template library, perform the following operations:

[0053] If the size of the template library is less than L, directly add the feature to the template library;

[0054] If the size of the template library is equal to L, after adding the feature to the template library, calculate the similarity between each feature and the remaining features to obtain the similarity matrix X; locate the position (x', y') of the largest element in the similarity matrix X, and calculate the average similarity S between the x'-th feature in the template library and the remaining features x' , and the average similarity S between the y'-th feature in the template library and the remaining features y' ; if S x' > S y' , then delete the x'-th feature in the template library; otherwise, delete the y'-th feature in the template library;

[0055] Among them, both L and M are preset positive integers; δ represents the confidence variance of the first M frames including the current frame, δ max and δ min respectively represent the maximum and minimum values of the variances accumulated during the tracking process; floor() represents rounding down; α is the learning rate, α > 0.

[0056] Further, when tracking the current frame, template features are constructed using the features in the template library, including:

[0057] Selecting one or more features from the template library that have the highest average similarity to the remaining features;

[0058] According to the time of the corresponding frame, using exponential moving to assign corresponding feature weights to the selected features, and performing weighted fusion on the features to obtain template features.

[0059] According to another aspect of the present invention, there is provided a computer program product, including a computer program; when the computer program is executed by a processor, it implements the above-mentioned failure perception method for player tracking in a football video provided by the present invention.

[0060] According to another aspect of the present invention, there is provided a computer-readable storage medium, including a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the above-mentioned failure perception method for player tracking in a football video provided by the present invention.

[0061] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0062] (1) During the process of tracking players in the current frame, after obtaining the target boxes and the confidence levels corresponding to each target box, the present invention will first eliminate the target boxes with low confidence levels and duplicate target boxes, thereby being able to exclude target boxes that are clearly unlikely to be tracking results as early as possible, improving the subsequent analysis efficiency and ensuring the real-time nature of tracking; on this basis, occlusion perception is realized based on the number of key points in the human skeleton model obtained by pose detection. Specifically, when the actually detected number of key points is significantly less than the standard number of key points, it is determined that occlusion occurs in the current frame, thereby realizing occlusion perception. This occlusion perception method makes full use of the pose information of the players, can achieve a finer-grained determination, and is not restricted by the actual standing position scenario, and can effectively improve the accuracy of occlusion perception.

[0063] (2) For the target boxes located at the edge of the current frame, the present invention will comprehensively consider the historical movement trajectories of the corresponding targets. When the movement of the target gets closer and closer to the edge in more than 3 consecutive frames, it is only then determined that the target will disappear. This determination method can more accurately judge the movement trend of the target and improve the accuracy of disappearance perception.

[0064] (3) Before performing regression on the response map, the present invention will first weight the response map with a Gaussian window function, and the selected Gaussian window function is centered on the geometric position center of the response map, thereby being able to focus more on the central region of the response map and effectively enhance the target.

[0065] (4) When updating the template library based on the results of bidirectional tracking (i.e., forward tracking and backward tracking), the present invention sets a sliding window with variable length and only performs backward tracking on the frames within the window. During backward tracking, the current frame is located at the end of the window, and each frame is assigned an initial weight. The closer a frame is to the current frame, the higher its initial weight. Subsequently, the frame weights are updated based on the average feature difference value between the tracking result of each frame and the tracking results of the remaining frames within the window, and reliable results in backward tracking are determined based on the frame weights. Finally, the template library is updated according to the results that are reliable in both forward tracking and backward tracking. This method for updating the template library based on a sliding window can ensure that the target does not deviate too much within each tracking result, and the initialization and update methods of the corresponding frame weights can accurately reflect the importance of frames in backward tracking, ultimately improving the quality of the template features in the template library and the accuracy of subsequent player tracking.

[0066] (5) In a preferred embodiment of the present invention, when performing player tracking on the current frame, multiple features are selected from the template library, and weights are assigned to each feature using exponential moving according to the time of the corresponding frame. The final template feature is obtained through weighted fusion. This method for generating template features can comprehensively consider the similarity between features and the temporal correlation between frames, further improving the quality of the template features and thus the accuracy of player tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a flowchart of a failure perception method for player tracking in a football video provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0069] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0070] Like other tracking tasks, the initial data for the player tracking task in football videos, in addition to the video to be tracked, also requires the position information of the tracking target in the historical frames. This information is generally a rectangular box represented by a quadruple (x, y, w, h). The target is boxed, where x and y represent the upper left coordinates of the positioning box, with the upper left corner of the given image as the origin, and w and h represent the width and height of the initial box, both in pixels. When predicting the position of the target in the next frame, a search area needs to be selected and it is ensured that the target is within the search area. Since the movement of the target between adjacent frames is small, generally a square box of appropriate size centered on the target position in the previous frame is selected as the search area. This rectangular box that frames the target is the template image in player tracking.

[0071] Using the features of the template image as the convolution kernel, after performing a convolution operation on the features of the search area image, a response map is obtained; the value of each point in the response map represents the confidence that the corresponding position in the search area image is the target. When there are many suspected players in the search area, multiple peaks will appear in the response map, and these peaks can be used to locate the suspected targets.

[0072] In order to improve the accuracy of failure perception during player tracking in football videos and thus improve the accuracy of player tracking, the present invention provides a method and program product for failure perception in player tracking in football videos. The overall idea is to fully consider the movement characteristics of football players in actual games and improve the judgment basis for failure perception.

[0073] The following are examples.

[0074] Example 1:

[0075] A method for failure perception in player tracking in football videos, as Figure 1 shown, includes:

[0076] S1: For the current frame, extract the features of the search area therein to obtain the search area features, and use the template features as the convolution kernel to perform a convolution operation on the search area features to obtain a response map; the template features are the features of the tracking results of the target to be tracked in the historical frames;

[0077] S2: Perform regression on the response map to obtain the target boxes and the confidence levels corresponding to each target box, and eliminate the target boxes with confidence levels lower than a preset first threshold to obtain a set of target boxes;

[0078] The target boxes with low confidence levels are obviously not the tracking results. Directly eliminating these target boxes can reduce the subsequent calculation amount without affecting the tracking accuracy, improve the calculation efficiency, and ensure real-time performance;

[0079] Accurately calculating the confidence of the target box is beneficial to improving the accuracy of player tracking and failure perception. To improve the calculation accuracy of the target box confidence, as a preferred implementation, before regressing the response map in this embodiment, the response map will be enhanced for the target first. Specifically, the response map is weighted with a Gaussian window function; the Gaussian window function is centered on the geometric position center of the response map, so that the central region of the response map can be more concerned. Since the center of the search area in the current frame usually coincides with the center of the target position in the previous frame, through such an operation, the target can be effectively enhanced;

[0080] S3: Select the target box with the highest confidence from the target box set as the positioning box;

[0081] S4: For each target box in the current target box set with a confidence lower than that of the positioning box, calculate the intersection-over-union ratio with the positioning box respectively. If the intersection-over-union ratio is greater than the preset second threshold, the corresponding target box is determined as a duplicate target box and removed from the target box set;

[0082] By identifying and removing duplicate target boxes, the amount of duplicate calculations can be reduced and the situation of misjudgment can be avoided;

[0083] S5: If there are still target boxes in the target box set that have not been selected as the positioning box, select the next target box from the target box set in descending order of confidence as the positioning box, and transfer to S4; otherwise, use the current target box set as the candidate target box set and transfer to S6;

[0084] S6: Perform pose estimation on each target box in the candidate target box set to obtain the human skeleton model corresponding to each target box, and count the number of key points N included in the human skeleton model sum , if N sum <η×n×N per , it is determined that occlusion occurs in the current frame;

[0085] When there is no occlusion, the human skeleton point model obtained through pose estimation will have a standard number of key points. When occlusion occurs, the occlusion will cause the number of key points in the actually detected human skeleton point model to be less than the standard number of key points due to occlusion. Based on this, when the number of actually detected key points is significantly less than the standard number of key points, it can be determined that occlusion has occurred;

[0086] Among them, 0<η<1. Optionally, in this embodiment, η = 0.7; n represents the number of target boxes in the candidate target box set; N per represents the standard number of key points of the human skeleton model. Pose estimation can be completed through the corresponding human pose estimation model. When the model is trained, the corresponding standard number of key points may be different depending on the dataset used. Optionally, in this embodiment, Nper = 17.

[0087] This embodiment realizes occlusion perception based on the number of key points in the human skeleton model obtained by pose detection, makes full use of the pose information of the players, can achieve a finer-grained determination, and is not restricted by the actual standing position scenario, which can effectively improve the accuracy of occlusion perception.

[0088] This embodiment further includes disappearance perception. Correspondingly, S6 also includes:

[0089] For each target box in the candidate target box set, respectively determine whether it is located at the edge of the current frame; for each target box located at the edge of the current frame, perform the following operations:

[0090] T1: Denote the target position in the target box as Pos i , and denote the positions of the template feature in the previous F frames as Pos i-1 , Pos i-2 , …… Pos i-F , and calculate the minimum distances min(|bbox - Pos i |), min(|bbox - Pos i - 1|), min(|bbox - Pos i-2 |), …… min(|bbox - Pos i-F |) between each position and the picture border of the current frame;

[0091] T2: If min(|bbox - Pos i |) < min(|bbox - Pos i-1 |) < min(|bbox - Pos i-2 |) < …… < min(|bbox - Pos i-F |), that is, the target gets closer and closer to the edge during the continuous F + 1 frame movement, then it is determined that disappearance occurs in the current frame;

[0092] where F is an integer greater than or equal to 2.

[0093] Optionally, in this embodiment, the method for determining whether the target box is located at the edge of the frame includes:

[0094] Calculate the minimum distance dist between the target position (x, y) in the target box and the frame border according to dist = min(|x|, |y|, |W - x|, |H - y|). If dist < γ, then determine that the target box is located at the edge of the frame;

[0095] where W and H respectively represent the width and height of the frame; γ represents a preset critical value, and γ > 0.

[0096] Based on the above steps, when performing disappearance perception in this embodiment, the historical movement trajectories of the corresponding targets will be comprehensively considered. This judgment method can more accurately judge the movement trend of the targets and improve the accuracy of disappearance perception.

[0097] Based on the above occlusion perception and disappearance perception, this embodiment further proposes methods for occlusion correction and disappearance correction.

[0098] The purpose of occlusion correction is to accurately track the target player after the occlusion ends. To achieve occlusion correction, in this embodiment, S4 further includes: maintaining a set of interference boxes for the current positioning box; if the intersection over union between the positioning box and the target boxes in the target box set with a confidence level lower than that of the positioning box is not greater than the second threshold, then determine the target box as an interference box and add it to the interference box set of the positioning box;

[0099] Moreover, the method for failure perception of player tracking in a football video further includes: if there is no occlusion in the current frame while there was occlusion in the previous frame, indicating that the occlusion has ended, then perform occlusion correction; the occlusion correction specifically includes:

[0100] Locate the last frame before the current frame where there was no occlusion, use it as the occlusion start frame, and obtain the set of interference boxes B corresponding to the tracking result of the template feature in the occlusion start frame;

[0101] For the current frame, denote its set of candidate target boxes as the set of candidate target boxes U. For each target box in the set of candidate target boxes U, calculate the average feature similarity between it and each target box in the set of interference boxes B as the interference index of the corresponding target box;

[0102] Determine the target box with the minimum interference index in the set of candidate target boxes as the tracking result of the template feature in the current frame.

[0103] The purpose of disappearance correction is to accurately identify the reappearing player after the disappearing player reappears in the video. To achieve disappearance correction, in this embodiment, after S6, it further includes: if there is no disappearance in the current frame while there was disappearance in the previous frame, then perform disappearance correction; the disappearance correction includes:

[0104] Perform target detection on the players in the current frame to obtain the detection positions of each player in the current frame;

[0105] Use the Kalman filter algorithm to predict the positions of each player in the current frame from the previous frame as the predicted positions of the corresponding players;

[0106] Taking the Euclidean distances between each detection position and each predicted position as elements, obtain the cost matrix D;

[0107] Execute the Hungarian algorithm on the cost matrix D to obtain the list of row indices row_i and the list of column indices col_i of the successfully matched players in the previous frame and the current frame;

[0108] Based on the list of row indices and the list of column indices of the successfully matched players in the previous frame and the current frame, the players can be classified as follows: If the index corresponding to the player in the previous frame does not appear in row_i, put it into the set of disappeared players unmatched_track; if the index corresponding to the player in the current frame does not appear in col_i, put the player into the set of reappeared players unmatched_det; if the index corresponding to the player in the previous frame appears in row_i, it means that the player is successfully matched with the player corresponding to the index in col_i, and put the successfully matched players into the set matched_pair;

[0109] Traverse the set of players unmatched_det, and for each player traversed, if the player meets the preset conditions, determine the player as a target that reappears after disappearing;

[0110] Among them, the preset conditions include: the target box corresponding to the player is located in the search area of the current frame and at the edge of the current frame, and at the same time, the similarity between the player target box and the target box of the player that disappeared previously is greater than the preset similarity threshold.

[0111] To alleviate the problem of tracking failure caused by the change of players during the tracking process, this embodiment will maintain a template library with a size not exceeding L, which is used to store the features of the tracking results of the target to be tracked in one or more frames; when tracking the current frame, the template feature is the feature constructed using the features in the template library. This embodiment also proposes a two-way tracking method based on an adaptive weight sliding window to screen reliable intermediate tracking results and update the template library so that the tracker can adapt to the appearance changes of players. Specifically,

[0112] If the sequence number i of the current frame differs from the frame sequence number corresponding to the last update of the template library by M, then perform the template library update; the template library update includes:

[0113] Set a window with a length of m = floor(Mδ / (δ max -δ min )) and make it cover the current frame and the previous m - 1 frames;

[0114] Add the target boxes with confidence greater than the preset third threshold in the tracking results of the template features in each frame within the window to the forward tracking result set;

[0115] The weights of each frame within the initial window Let \(j\) represent the frame sequence number within the window, where \(1\leq j\leq m\); the closer the frame is to the current frame, the higher its initial weight.

[0116] For each frame within the window, calculate the average feature difference value \(D\) between the tracking result of the template feature in this frame and the tracking results in the remaining frames within the window. cur-j And add the frames with feature difference values less than the preset fourth threshold to the reliable result set.

[0117] According to Update the weights of each frame in the reliable result set. Is the updated weight for the \(j\)-th frame within the window.

[0118] Add the tracking results of the frames with weights greater than the preset fifth threshold in the reliable result set to the reverse tracking result set.

[0119] Successively add the features of the target boxes located in the intersection of the forward tracking result set and the reverse tracking result set to the template library.

[0120] When there are too many features in the template library, similar features will also increase. The increase in similar features will not only not improve the model accuracy, but also increase the computational amount and reduce the running speed. To avoid this situation, in this embodiment, the size of the template library is limited, and it is expected to retain features with a large enough gap in the template library. Specifically, each time a feature is added to the template library in this embodiment, the following operations are performed:

[0121] If the size of the template library is less than \(L\), directly add the feature to the template library.

[0122] If the size of the template library is equal to \(L\), after adding the feature to the template library, calculate the similarity between each feature and the remaining features to obtain the similarity matrix \(X\); locate the position \((x',y')\) of the largest element in the similarity matrix \(X\), and calculate the average similarity \(S\) between the \(x'\)-th feature in the template library and the remaining features x' And the average similarity \(S\) between the \(y'\)-th feature in the template library and the remaining features y' ; if \(S\) x' \(>\) \(S\) y' , then delete the \(x'\)-th feature in the template library; otherwise, delete the \(y'\)-th feature in the template library.

[0123] Where \(L\) and \(M\) are both preset positive integers; \(\delta\) represents the confidence variance of the first \(M\) frames including the current frame, \(\delta\) max and \(\delta\) min respectively represent the maximum and minimum values of the variances accumulated during the tracking process; \(floor()\) represents rounding down; \(\alpha\) is the learning rate, \(\alpha>0\).

[0124] When tracking the current frame based on the maintained template library, one or more features in the template library can be used to construct template features, specifically including:

[0125] Select one or more features from the template library that have the highest average similarity to the remaining features;

[0126] According to the time of the corresponding frame, use exponential moving to assign corresponding feature weights to the selected features, and perform weighted fusion on the features to obtain template features.

[0127] In this embodiment, when constructing template features, weights are assigned to each feature according to the time of the corresponding frame using exponential moving, so that the newer the feature, the higher the weight. Based on this weight, multiple features are weighted and fused to obtain the final template features, which can comprehensively consider the similarity between features and the temporal correlation between frames, further improve the quality of template features, and thus improve the accuracy of player tracking.

[0128] It is easy to understand that when only one feature is selected from the template library, the selected feature is the one with the highest average similarity to the remaining features, and this feature will be used as the final template feature.

[0129] Embodiment 2:

[0130] A computer program product includes a computer program; when the computer program is executed by a processor, it implements the method for failure perception of player tracking in a football video provided in the above Embodiment 1.

[0131] Embodiment 3:

[0132] A computer-readable storage medium includes a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the method for failure perception of player tracking in a football video provided in the above Embodiment 1.

[0133] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for detecting player tracking failure in a soccer video, characterized in that: include: S1: For the current frame, extract features from the search area therein to obtain search area features, and use template features as convolution kernels to perform convolution operations on the search area features to obtain a response map; the template features are features of the tracking results of the target to be tracked in the historical frames; S2: regressing the response map to obtain a target frame and a confidence level corresponding to each target frame, removing target frames whose confidence levels are lower than a preset first threshold, and obtaining a target frame set; S3: Selecting a target frame with the highest confidence from the target frame set as a positioning frame; S4: for each target frame in the current target frame set whose confidence is lower than that of the positioning frame, respectively calculate the intersection-over-union ratio between the target frame and the positioning frame, and if the intersection-over-union ratio is greater than a preset second threshold, determine the corresponding target frame as a duplicate target frame and remove it from the target frame set; S5: If there are still target frames in the target frame set that have not been selected as the positioning frame, the next target frame is selected from the target frame set as the positioning frame in descending order of confidence, and the process goes to S4; otherwise, the current target frame set is used as the candidate target frame set, and the process goes to S6; S6: Perform posture estimation on each target frame in the candidate target frame set, obtain the human skeleton model corresponding to each target frame, and count the number of key points N contained in the human skeleton model sum , if N sum <η×n×N per , it is determined that occlusion occurs in the current frame; Where 0<η<1; n represents the number of target frames in the candidate target frame set; N per Represents the standard number of key points of the human skeleton model.

2. The method for detecting player tracking failure in a soccer video according to claim 1, wherein: The S6 also includes: For each target frame in the candidate target frame set, determine whether it is located at the edge of the current frame; for each target frame located at the edge of the current frame, perform the following operations: T1: Record the target position in the target box as Pos i , the positions of the template features in the first F frames are recorded as Pos i-1 、Pos i-2 , ...Pos i-F , and calculate the minimum distance min(|bbox-Pos i |)、min(|bbox-Pos i-1 |)、min(|bbox-Pos i-2 |)、……min(|bbox-Pos i-F |); T2: If min(|bbox-Pos i |) <min(|bbox-Pos i-1 |) <min(|bbox- Pos i-2 |)<…… <min(|bbox-Pos i-F |), it is determined that disappearance occurs in the current frame; Wherein, F is an integer greater than or equal to 2.

3. The method for detecting player tracking failure in a soccer video according to claim 2, wherein: Methods for determining whether the target box is at the edge of the frame include: According to dist=min(x|,|y|,|Wx|,|Hy), the minimum distance dist between the target position (x, y) in the target frame and the frame border is calculated. If dist<γ, the target frame is judged to be at the edge of the frame. Wherein, W and H represent the width and height of the frame respectively; γ represents a preset critical value, and γ>0.

4. The method for detecting player tracking failure in a soccer video according to claim 2, wherein: After S6, the method further includes: if disappearance does not occur in the current frame but disappearance occurs in the previous frame, performing disappearance correction; the disappearance correction includes: Perform target detection on the players in the current frame to obtain the detection position of each player in the current frame; Use the Kalman filter algorithm to predict the position of each player in the previous frame in the current frame as the predicted position of the corresponding player; The cost matrix D is obtained by taking the Euclidean distance between each detection position and each predicted position as an element; Execute the Hungarian algorithm on the cost matrix D to obtain the row index list row_i and column index list col_i of the players that successfully matched in the previous frame and the current frame; If the index corresponding to the player in the current frame does not appear in col_i, the player is put into the reappearing player set unmatched_det; Traversing the player set unmatched_det, for each traversed player, if the player meets the preset condition, the player is determined as a target that reappears after disappearing; The preset conditions include: the target frame corresponding to the player is located in the search area of ​​the current frame and at the edge of the current frame, and the similarity between the player target frame and the previously disappeared player target frame is greater than a preset similarity threshold.

5. The method for detecting player tracking failure in a soccer video according to claim 1, wherein: S4 also includes: maintaining an interference frame set for the current positioning frame; if the intersection-and-union ratio between the positioning frame and a target frame in the target frame set whose confidence is lower than that of the positioning frame is not greater than a second threshold, determining the target frame as an interference frame and adding it to the interference frame set of the positioning frame; Furthermore, the method for detecting player tracking failure in a football video further includes: if no occlusion occurs in the current frame but occlusion occurs in the previous frame, performing occlusion correction; the occlusion correction includes: Locate the last frame before the current frame that is not occluded, use it as the occlusion start frame, and obtain the interference frame set B corresponding to the tracking result of the template feature in the occlusion start frame; For the current frame, its candidate target frame set is recorded as candidate target frame set U. For each target frame in candidate target frame set U, the average feature similarity between it and each target frame in interference frame set B is calculated as the interference index of the corresponding target frame. The target frame with the smallest interference index in the candidate target frame set is determined as the tracking result of the template feature in the current frame.

6. The method for detecting player tracking failure in a soccer video according to claim 1, wherein: In S2, before regressing the response graph, it also includes: The response map is weighted with a Gaussian window function; the Gaussian window function is centered at the geometric center of the response map.

7. The method for detecting player tracking failure in a soccer video according to any one of claims 1 to 6, characterized in that: Also includes: Maintaining a template library whose size does not exceed L, for storing features of tracking results of the target to be tracked in one or more frames; and when tracking the current frame, the template features are features constructed using features in the template library; Furthermore, the method for detecting failure of player tracking in a football video further includes: If the sequence number i of the current frame differs by M from the frame sequence number corresponding to the last time the template library was updated, the template library is updated; the template library update includes: Set the length to m = floor(Mδ / (δ max -δ min ))'s window, and make it cover the current frame and its previous m-1 frames; Adding the target frame whose confidence level is greater than a preset third threshold in the tracking results of the template feature in each frame within the window to the forward tracking result set; The weight of each frame in the initial window j represents the frame number in the window, 1≤j≤m; the closer the frame is to the current frame, the higher its initial weight; For each frame in the window, calculate the average feature difference value D between the tracking result of the template feature in the frame and the tracking result in the remaining frames in the window. cur-j , and adding frames whose feature difference values ​​are less than a preset fourth threshold to a reliable result set; according to Update the weight of each frame in the reliable result set; is the updated weight of the jth frame in the window; Adding the tracking results in the frames whose weights are greater than a preset fifth threshold in the reliable result set to the reverse tracking result set; Sequentially adding features of the target box located in the intersection of the forward tracking result set and the backward tracking result set to the template library; Each time a feature is added to the template library, the following operations are performed: If the size of the template library is smaller than L, the feature is directly added to the template library; If the size of the template library is equal to L, after adding the features to the template library, calculate the similarity between each feature and the remaining features to obtain a similarity matrix X; locate the position (x', y') of the largest element in the similarity matrix X, and calculate the average similarity S between the x'th feature in the template library and the remaining features x' , and the average similarity S between the y'th feature and the remaining features in the template library y' If S x' >S y' , then delete the x'th feature in the template library; otherwise, delete the y'th feature in the template library; Where L and M are both preset positive integers; δ represents the confidence variance of the previous M frames including the current frame, δ max and δ min They represent the maximum and minimum values ​​of the variance accumulated during the tracking process respectively; floor() means rounding down; α is the learning rate, α>0.

8. The method for detecting player tracking failure in a soccer video according to claim 7, wherein: When tracking the current frame, the features in the template library are used to construct template features, including: Select one or more features from the template library that have the highest average similarity with the remaining features; According to the time of the corresponding frame, the selected features are assigned corresponding feature weights using exponential shift, and the features are weighted fused to obtain the template features.

9. A computer program product, characterized in that It comprises a computer program; when the computer program is executed by a processor, it implements the method for sensing player tracking failure in a football video as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The invention comprises a stored computer program; when the computer program is executed by a processor, the device where the computer-readable storage medium is located is controlled to execute the method for sensing failure of player tracking in a football video according to any one of claims 1 to 8.

Citation Information

Cited By

  • Adaptive sliding window type incremental adjustment variable-scale motion body posture refinement method

    CN121026054A

  • Method for refining pose of moving body by adaptive sliding window incremental adjustment with variable scale

    CN121026054B

  • Target tracking detection method and device based on dynamic scoring and time window adjustment

    CN121582297A

  • Target tracking detection method and device based on dynamic score and time window adjustment

    CN121582297B