Method and system for detecting ball and acquiring movement information in football image
The method employs AI models to preprocess and analyze soccer videos, effectively addressing the challenge of detecting balls and tracking movement in dynamic sports environments, resulting in improved accuracy and user experience.
Patent Information
- Application Number
- PCT/KR2024/018317
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-11-20
- Publication Date
- 2025-06-12
AI Technical Summary
Existing technologies face challenges in accurately detecting a ball in soccer videos and generating reliable movement information, especially in dynamic and cluttered environments where objects are frequently occluded.
A method and system utilizing artificial intelligence models, specifically CNN-based models like YOLO, to search for and track objects of interest in soccer videos. This involves preprocessing images to enhance accuracy, searching for objects in current frames, and using temporal correlation to determine the location of objects across frames.
The proposed solution enables efficient and accurate detection of balls and generation of movement information, even in complex scenarios with occlusions, thereby improving the quality of sports video analysis and user experience.
Smart Images

Figure KR2024018317_12062025_PF_FP_ABST
Abstract
Description
Method and system for detecting a ball in a soccer video and obtaining movement information
[0001] The present invention relates to a method and system for detecting a ball in a soccer video and obtaining movement information.
[0002] The application of ICT to sports footage is increasing. In particular, technology that automatically extracts objects of interest from sports footage can provide a variety of value-added services. Objects of interest are objects of interest to users in sports footage, such as the ball, specific players, players from specific teams, and scores. By tracking the ball with a camera used to film a sporting event or zooming in and editing only the ball from the entire stadium footage, a video of key events can be naturally generated. Alternatively, by tracking a specific player with a camera used to film a sporting event or zooming in and editing only a specific player from the entire stadium footage, a video of a specific player can be naturally generated. Furthermore, by automatically extracting movement information from specific players, a heat map can be created for that player, or by automatically extracting movement information from all players, a comprehensive analysis of the game can be performed.
[0003] In particular, sports videos such as soccer or futsal videos feature large stadiums and wide coverage areas. Therefore, extracting only the portions of interest desired by the user from the sports video can create various additional services. For example, while a user would have to manually follow a specific player and edit the video to extract footage of him or her in action in a soccer video, if a technology for extracting objects of interest exists, the technology can be applied to the original sports video source to automatically extract footage centered on that specific player. Alternatively, while a user would have to manually follow the ball while filming a soccer video to ensure it doesn't leave the camera frame, technology for extracting objects of interest can allow the camera to automatically follow the ball, resulting in a desired sports broadcast footage.
[0004] There are various methods for extracting objects of interest from images. Representative methods include methods that do not use machine learning and methods that do. Methods that do not use machine learning predefine the characteristics of the object of interest and then search for objects within the image based on these defined characteristics. For example, if the object of interest is a ball, typical characteristics of the ball (e.g., area occupied by the ball (number of pixels), curvature, color) can be defined in advance and the ball can be extracted based on these characteristics. While this method can be utilized in images with minimal noise and relatively clear images, it is prone to errors and can increase in video, especially in sports videos, where objects move frequently and are not static images.
[0005] One way to utilize machine learning is to learn the characteristics of objects of interest in advance and then use the learned model to locate them in images. A representative example is YOLO (You Only Look Once), a technique that extracts features based on a convolutional neural network (CNN) deep learning model and then uses these to predict the type and location of an object. While this machine learning-based object extraction method demonstrates excellent performance, sports videos often feature a large number of objects, often obscuring the object of interest from other objects. Consequently, its performance remains incomplete.
[0006] A method for detecting a ball in a soccer video and generating movement information of the detected ball may be provided.
[0007] The technical task to be achieved by this embodiment is not limited to the technical task described above, and other technical tasks can be inferred from the following embodiments.
[0008] A method for detecting a ball in a soccer video and generating movement information of the detected ball may be provided.
[0009] The technical task to be achieved by this embodiment is not limited to the technical task described above, and other technical tasks can be inferred from the following embodiments.
[0010] A method for detecting a ball or player in a soccer video and generating motion information of the detected object may be provided.
[0011] FIG. 1 illustrates a flowchart of a method for generating motion information of an object of interest in a sports video, according to one embodiment.
[0012] FIG. 2 illustrates an original of a sports video, according to one embodiment.
[0013] FIG. 3 illustrates a flowchart of a method for preprocessing a sports image, according to one embodiment.
[0014] FIG. 4 illustrates a state in which an object of interest is searched in a sports video according to one embodiment.
[0015] FIG. 5 illustrates an original of a sports video according to one embodiment.
[0016] FIG. 6 illustrates a two-dimensional sports stadium according to one embodiment.
[0017] FIG. 7 illustrates an analysis result generated based on location information of an object of interest, according to one embodiment.
[0018] FIG. 8 illustrates a heatmap generated based on location information of an object of interest, according to one embodiment.
[0019] FIG. 9 illustrates a flowchart of a method for obtaining location information of an object of interest, according to one embodiment.
[0020] FIG. 10 illustrates a state in which an object of interest is searched in a sports video according to one embodiment.
[0021] FIG. 11 illustrates a state in which a ball is not detected in a sports video according to one embodiment.
[0022] FIG. 12 illustrates a current frame and surrounding frames in a sports video according to one embodiment.
[0023] FIG. 13 illustrates a current frame and surrounding frames in a sports video according to one embodiment.
[0024] FIG. 14 illustrates the movement of an object of interest in a sports video, according to one embodiment.
[0025] FIG. 15 illustrates the movement of an object of interest in a sports video, according to one embodiment.
[0026] FIG. 16 shows a state in which multiple balls are searched in a sports video according to one embodiment.
[0027] FIG. 17 illustrates a candidate region of an object of interest in a sports video, according to one embodiment.
[0028] Figures 18a and 18b illustrate labeled objects of interest, according to one embodiment.
[0029] FIG. 19 illustrates a flowchart of a method for performing labeling on an object of interest, according to one embodiment.
[0030] FIG. 20 illustrates a state in which tracking of a labeled object of interest fails, according to one embodiment.
[0031] FIG. 21 illustrates a flowchart of a method for obtaining ball position information according to one embodiment.
[0032] FIG. 22 illustrates a current frame and surrounding frames in a sports video according to one embodiment.
[0033] FIG. 23 illustrates a flowchart of a method for obtaining location information of an object of interest, according to one embodiment.
[0034] A method for collecting movement information of a ball in a sports video performed in a computing device may include the steps of sequentially receiving each frame of the sports video as an input video, searching for at least one object classified as a ball in the input video using an artificial intelligence model, classifying the input video as a key frame and determining the object as a ball if there is one object classified as a ball by the artificial intelligence model in the input video and the probability that the object is a ball is greater than or equal to a reference value, determining an alternative region where a ball is expected to be based on a location of a ball in a surrounding key frame if there is no object classified as a ball by the artificial intelligence model in the input video, determining the plurality of objects as candidate regions and classifying the input video as a candidate frame if there are multiple objects classified as balls by the artificial intelligence model in the input video, and determining one of the ball candidate regions of the candidate frame as a ball.
[0035] The above-determining step may determine the candidate region with the highest temporal correlation with the surrounding key frame as the ball. The surrounding key frame may be the frame closest to the current frame among the key frames located before and after the current frame.
[0036] A method for searching for an object of interest in a sports video performed by a computing device may include the steps of: receiving a current frame of the sports video; tracking the object of interest in the current frame based on a previous search result if there is a previous search result for the object of interest in the previous frame; searching for the object of interest using an artificial intelligence model if there is no previous search result or tracking the object of interest fails; and determining an alternative region in which the object of interest is estimated to be located using a frame having a search result among previous frames if searching for the object of interest fails.
[0037] A method for collecting position information of an object of interest in a sports video performed on a computing device includes the steps of receiving a current frame of the sports video as an input video, searching for an object of interest in the input video using an artificial intelligence model, assigning label information to at least one of the searched objects of interest, and obtaining two-dimensional position information of the object of interest by perspective transforming the position information of the object of interest to which the label information has been assigned, wherein the object of interest may include a player, a ball, or a referee.
[0038] A method for preprocessing a sports image performed on a computing device includes the steps of converting the sports image into a frequency domain to remove a background area having a frequency component greater than a reference value, generating a color histogram of a remaining area from which the background area has been removed and determining a color having the largest number based on the color histogram, and extracting a bounding box for pixels having the color, wherein the color histogram can be generated by performing K-means clustering on the remaining area to classify pixels constituting the remaining area into a predetermined number of clusters and counting the number of pixels for each of the clusters.
[0039] A method for searching for an object of interest in a sports video performed in a computing device may include the steps of searching for an object of interest in a current frame using an artificial intelligence model, determining the current frame as a key frame if the searched object of interest is one, designating each of the searched objects of interest as a first candidate region and a second candidate region and determining the current frame as a candidate frame, if the searched object of interest is plural, determining an alternative region in which the object of interest is expected to be based on a location of the object of interest in each of the surrounding key frames if the searched object of interest is not found, and determining a candidate region having a greater temporal correlation with a location of the object of interest in each of the surrounding key frames as a location of the object of interest.
[0040] A method for generating a heat map of a soccer game performed on a computing device may include a step of searching for a ball for a current frame using an artificial intelligence model, a step of recording a position of the searched ball, a step of converting the recorded position of the ball into a two-dimensional plane, and a step of generating a heat map based on the position of the ball converted into the two-dimensional plane.
[0041] Below, several embodiments will be described clearly and in detail with reference to the attached drawings so that those skilled in the art (hereinafter, “ordinary technicians”) can easily practice the present invention.
[0042] Hereinafter, objects of interest may include types of objects that constitute a match in a sports video and may be of interest to the user. For example, objects of interest may include a ball, players, players of a specific team, a specific player, a referee, etc.
[0043] Hereinafter, the object to be tracked is an object of interest that the user is directly interested in. For example, the user may be interested in the movement of the ball or the movement of a specific player (e.g., player number 7 or the goalkeeper). Alternatively, the user may be interested in the movement of all objects of interest. Therefore, the object to be tracked may be at least one of the objects of interest.
[0044] The following methods can be performed on a computing device.
[0045] FIG. 1 illustrates a flowchart of a method for generating motion information of an object of interest in a sports video, according to one embodiment.
[0046] Referring to FIG. 1, a method for generating motion information of an object of interest in a sports video may include a step of receiving a sports video to be analyzed (S12), a step of searching for an object of interest (S14), a step of converting it into a position on a two-dimensional stadium (S16), and a step of generating an analysis result (S18).
[0047] In step S12, the computing device can input an original sports video to be analyzed.
[0048] Sports videos may include sports broadcast videos, videos taken with personal cameras, etc. Sports videos may include, but are not limited to, soccer videos, futsal videos, basketball videos, rugby videos, etc. Sports videos may be captured by a single camera or videos captured by multiple cameras. Videos captured by a single camera may be, but are not limited to, videos of a sports event captured by a single camera in a fixed location. Videos captured by multiple cameras may be, but are not limited to, videos captured by cameras located at different locations and combined. Sports videos may include a stadium, player(s), and a ball. In step S12, the sports videos input to the computing device may be dynamic videos or one of multiple static videos (e.g., frames) that constitute a dynamic video. For example, the computing device may receive the Nth frame (N is an integer) of the sports videos. FIG. 2 illustrates an original video input to a computing device according to an embodiment.
[0049] A computing device can preprocess an input image. Preprocessing of the input image is a step of enhancing the image to improve the accuracy of object-of-interest detection. The computing device can perform preprocessing on the current frame. Image enhancement may include removing the background of a sports video or extracting a region of interest (e.g., a soccer stadium). Since soccer videos are typically shot from a wide field and from a distance, image enhancement before detecting an object of interest can improve detection accuracy. For example, background removal can be performed by removing high-frequency regions (e.g., BG in Figure 2). Region-of-interest extraction can be performed by extracting green or green-like regions based on pixel values (e.g., green). Alternatively, a flat area whose area exceeds a threshold value can be extracted as the region of interest. This is because a stadium typically consists of a large surface composed of similar colors, such as grass, dirt, mats, and courts. In one embodiment, a region slightly enlarged from the region of interest (e.g., an area expanded by a threshold value from the region-of-interest bounding box) can be determined as the actual region of interest. For example, practical areas of interest may include referees outside the line, substitutes, etc.
[0050] According to one embodiment, the primary color constituting the region of interest (i.e., field color) may be determined based on the color histogram of the current image. The computing device may perform color-based clustering on the current input image to classify the current input image into multiple color clusters. Similar colors may be classified into one cluster. The clustering method may use, but is not limited to, a K-means clustering technique. The color belonging to the cluster with the largest number of pixels may be determined as the color of the field area. For example, in a soccer image, the color of the grass may be determined as the primary color, in a basketball image, the color of the court (e.g., ochre, yellow) may be determined as the primary color, and in a tennis image played on a blue court, blue may be determined as the primary color. The bounding box of pixels belonging to the field color (e.g., the color belonging to the cluster with the largest number of pixels as a result of the clustering) may be extracted as the field area. FIG. 3 illustrates a flowchart of a preprocessing method performed on an input image, but is not limited thereto.
[0051] Referring back to FIG. 1, at step S14, the computing device can search for an object of interest in the input image.
[0052] Objects of interest may include specific players, players belonging to a specific team, players, a ball, a referee, etc. A computing device may search for objects of interest in an input image using an artificial intelligence model. The artificial intelligence model is designed to identify objects of interest from sports images and may be generated by performing learning based on training data. In one embodiment, the artificial intelligence model is based on a neural network such as a convolutional neural network (CNN) and may be trained to identify specific objects from a plurality of training images. In one embodiment, the artificial intelligence model may be trained to identify players, a ball, etc. from the current image. For example, the artificial intelligence model may be YOLO (You Only Look Once) or R-CNN (Regions with Convolutional Neural Networks).
[0053] In one embodiment, an AI model can detect players, referees, and / or balls in a sports video. FIG. 4 illustrates an example of an AI model detecting objects of interest in a sports video.
[0054] A computing device can search for an object of interest for each frame of the input image or at predetermined intervals. The computing device can determine location information of the searched object of interest. For example, a bounding box containing the searched object of interest may be drawn, and one of the four corners of the bounding box or the center point of the bounding box may be determined as the location of the object of interest.
[0055] Referring again to FIG. 1, in steps S16 to S18, the computing device may generate motion analysis information of the object of interest based on location information of the object of interest.
[0056] In step S16, the computing device can convert the position information of the objects of interest acquired in step S14 into position information on a two-dimensional stadium. The position information of the objects of interest acquired in step S14 is usually a three-dimensional stadium image photographed from the front of the stadium (e.g., the sideline position), so if the position information acquired in step S14 is converted into a two-dimensional stadium image (i.e., a form of the stadium viewed from above, see FIG. 6), more accurate object position information can be generated. The computing device can project the position information of the objects of interest acquired in step S14 onto the two-dimensional stadium. Based on the position information of the two-dimensional image, the positions of the players and the positions of the ball can be known at a glance, making it efficient for analyzing the results of the game or the movements of the players. The two-dimensional conversion of the position information can be performed using perspective transformation. Perspective transformation is a method of selecting four vertices in a two-dimensional stadium, selecting four points representing the corresponding vertices in a captured video, representing how far apart each vertex is in a matrix, and mapping it by multiplying the matrix to the image.
[0057] In step S18, the computing device may generate various analysis results based on the location information of step S16. The various analysis results may include distances traveled by players, the movement of a specific player, a heat map representing the overall movement of a specific player during a game, ball movement, etc. FIG. 7 illustrates, according to one embodiment, information on passes made during a game, and FIG. 8 illustrates, according to one embodiment, a heat map representing information on the movement of at least one player during a game.
[0058] FIG. 9 is a flowchart of a method for a computing device to search for an object of interest, according to one embodiment.
[0059] The method of FIG. 9 may correspond to, but is not limited to, the detailed steps of step S14 of FIG. 1. The method of FIG. 9 may be a flowchart of a method for a computing device to obtain movement information of a specific object of interest from a sports video.
[0060] In step S21, the computing device may receive a sports video. The sports video may be one of the frames constituting a sports video.
[0061] In step S22, the computing device can search for an object of interest in the input image.
[0062] A computing device can use an artificial intelligence model to detect objects of interest. In one embodiment, the artificial intelligence model may be trained to detect players, goalkeepers, referees, and balls. The artificial intelligence model can determine the type of each detected object and output a probability that each object is of the determined type. Figure 10 illustrates objects of interest detected by the artificial intelligence model, according to one embodiment. Each object of interest is indicated by a bounding box and a probability value.
[0063] In step S23, the computing device can determine whether the object of interest has been explored.
[0064] The object of interest in step S23 may refer to a tracked object. For example, a user may wish to observe the movement of a ball during a game. Alternatively, the user may wish to observe the movement of a specific player (e.g., player number 7 or a goalkeeper) during a game. Alternatively, the user may wish to observe the movement of a specific player and the ball during a game. In other words, an object whose movement information is desired may be labeled, and the labeled object may be referred to as a tracked object.
[0065] Successful detection of a target object can mean that an object classified as the target object type exists in the current image and that the probability of the target object being the type is greater than or equal to a threshold value. For example, if the target object is a ball, the ball was detected in the image of FIG. 10 (when the probability threshold value is 0.7). Failure to detect an object of interest can mean that a supposedly existing object of interest has not been detected or that multiple objects of interest have been detected. In the image of FIG. 11, the ball was not detected. Cases where a ball was not detected may include not only cases where the ball itself was not detected, as in FIG. 11, but also cases where an object classified as a ball exists but has a probability value less than a threshold value. FIG. 16 illustrates a state in which multiple target objects of interest have been detected, according to one embodiment. If the target object of interest is a ball or a specific player, it is desirable to detect only one object in the image. However, if two or more balls or specific players are detected, only one of the detected objects is the desired target object of tracking.
[0066] Referring back to FIG. 9, if the search for the object of interest is unsuccessful (No), in step S24, the computing device can determine the type of failure in searching for the object of interest. First, if the object to be tracked is not currently being searched (No), in step S25, the computing device can determine an alternative area where the object to be tracked is expected to be. For example, if the original image source is a panoramic image capturing the entire soccer field, the ball must be searched. However, the ball may not be searched if it is obscured by players, overlapped by other objects, or due to image quality issues, etc. (see FIG. 11). In such cases, the computing device can determine the location of the object to be tracked using the search results of the surrounding frames. Referring to FIG. 12, the surrounding frames are frames located before or after the current image (CF), and preferably include frames within 10 frames of the current image.
[0067] For example, at least one of the previous frames in which the ball was detected in the previous frame may be used to determine a replacement area of the target object to be tracked in the current frame. In one embodiment, the replacement area of the ball in the current frame may be determined based on the area of the ball in the frame closest to the current frame among the previous frames in which the ball was detected. The area in which the area of the ball in the previous frame has moved by an estimated motion vector may be determined as the replacement area of the ball. The estimated motion vector refers to the direction and degree in which the target object to be tracked is expected to have moved between the last frame in which the ball was detected and the current frame. The estimated motion vector may be determined based on at least one previous frame or based on at least one subsequent frame. The estimated motion vector may be determined based on the previous frame and the subsequent frame. There are various methods for determining the estimated motion vector based on surrounding frames. The estimated motion vector may be determined by averaging the motion of the ball in the previous frames. This is because the time interval between frames is very short (e.g., 1 / 60 to 1 / 30 second), so it is unlikely to deviate significantly from the direction or magnitude of the motion of the surrounding frame. Figure 13 shows the previous frames of the current frame (CF) for calculating the estimated motion vector.
[0068] [Mathematical Formula 1]
[0069] Location of replacement area = (x a + MV x , y a + MV y )
[0070] In mathematical expression 1, x a, y a is the position where the tracked object was searched in the previous frame in which the tracked object was last searched. MV x and MV yis an estimated motion vector in which the tracked object is expected to have moved in the last frame in which the tracked object was detected. The estimated motion vector can be determined based on the movements of the tracked object in several past frames. For example, if the last frame in which the tracked object was detected is the previous frame, the estimated motion vector can be derived by averaging the movements of several tracked object in the previous frames.
[0071] In one embodiment, a replacement area of a ball in a current frame may be determined based on at least one of the previous frames in which the ball was searched and at least one of the subsequent frames in which the ball was searched. The area between the area in which the ball was searched in the previous frame and the area in which the ball was searched in the subsequent frame may be determined as the replacement area. Referring to FIG. 14, B n-1 is the area where the ball was explored in the previous frame, and B n+1 The area where the ball is explored in the subsequent frames, B n This is the alternate area where the ball is expected to be in the current frame.
[0072] MV in Equation 2 x and MV y is an estimated motion vector in which the tracking target object is expected to have moved in the last frame in which the tracking target object was searched, and the estimated motion vector can be determined based on the frame closest to the current frame among the previous frames in which the tracking target object was searched and the frame closest to the current frame among the frames after the tracking target object was searched. FIG. 15 illustrates a tracking target object (MO) in a previous frame, a current frame, and a subsequent frame according to an embodiment. The previous frame is N frames before the current frame, and the subsequent frame is M frames after the current frame. In Equation 2, x a, y a is the position where the tracked object was searched in the previous frame in which the tracked object was last searched.
[0073] [Equation 2]
[0074] MV x = (x b - x a ) XN / (N+M), MV y = (y b - y a ) XN / (N+M))
[0075] There are numerous ways to determine the replacement region of a tracked object based on the frames before and after the tracked object was actually explored, so a detailed description is omitted.
[0076] Referring back to FIG. 9, if multiple objects to be tracked are detected (Yes), in step S26, the computer device may select one of the multiple objects as the target object to be tracked by referring to the surrounding frames. For example, if the target object to be tracked is a ball, this means a situation where multiple objects of a shape similar to a ball are detected within one frame. FIG. 16 illustrates a situation where two balls are detected, according to one embodiment. For example, an artificial intelligence model for detecting balls may detect multiple objects determined to be balls in the current frame. For example, object A has a probability of being a ball of 0.67 and object B has a probability of being a ball of 0.55. In this embodiment, if the reference value is 0.5, both objects A and B may be detected as balls. In this case, one of the multiple objects A and B needs to be finally determined as the target object to be tracked.
[0077] According to one embodiment, among the searched multiple objects, the object closest to the searched position of the estimated target object in the surrounding frame may be finally determined as the estimated target object. According to one embodiment, among the estimated target objects searched in the current frame, an object belonging to the surrounding area of the searched position of the estimated target object in the surrounding frame may be determined as the final estimated target object. For example, referring to FIG. 17, if the objects searched as balls in the current frame are P12_1 and P12_2, and the area of the estimated target object in the previous frame is P12_P, P12_1 located in the surrounding area TA of P12_P may be determined as the final estimated target object. The surrounding area TA is not necessarily determined to include the area of the tracking target object in the previous frame. The surrounding area TA may also be determined as an area moved by the estimated motion vector determined from the surrounding frame. Since the estimated motion vector is the same as the above-described method for determining the replacement area, a detailed description thereof will be omitted. Because it is unlikely that a particular object's position will suddenly change significantly from the surrounding frame, objects that are far from the surrounding area are more likely to be misdetected.
[0078] Referring back to FIG. 9, the computing device can perform labeling on the object to be tracked determined in step S27. Labeling refers to assigning identification information to the object to be tracked, and based on the identification information displayed on the screen, the user can effectively visually view the movement of the object to be tracked. In addition, there is an advantage in that the computing device can determine whether the search for the object of interest in the sports video is being performed well by performing labeling. FIG. 18A illustrates a state in which a ball is labeled according to an embodiment, and FIG. 18B illustrates a frame following the frame of FIG. 18A according to an embodiment. The ball is also labeled in FIG. 18B. Labeling can be performed directly by the user initially (for example, the first frame in which no labeling information is assigned), but after the labeling of the object to be tracked is assigned, it can be assigned automatically according to object tracking.
[0079] FIG. 19 illustrates a flowchart of a method by which a computing device performs labeling on a searched object, according to one embodiment.
[0080] In step S191, if the currently input image is the first frame of the original image or there is no other labeling information to refer to (No), the computing device can assign labeling information in step S192. According to one embodiment, labeling can be performed on objects to be tracked. For example, labeling can be performed on the ball and / or player 7. Labeling information can be performed by the artificial intelligence model or can be directly input by the user for a specific object. The user can check the labeling information assigned by the artificial intelligence model and modify the labeling information. For example, if the artificial intelligence model is trained to detect player 7 and the ball within the field, the objects detected by the artificial intelligence model can be assigned '7P' and 'ball', respectively.
[0081] In step S191, if the current frame is not the first frame or there is previous labeling information that can be referenced (Yes), the computing device can perform object tracking in step S193. Object tracking is a method of determining whether an object identical to an object assigned a label in the previous frame also exists in the current frame, and a known algorithm (e.g., ByteTrack) can be used for object tracking. If object tracking is successful (Yes), the labeling information of the previous frame can be assigned to the object to be tracked in the current frame in the same manner.
[0082] If object tracking fails (No), in step S194, the computing device may provide a recommended labeling list of objects determined not to be in the current frame among the objects assigned labeling information in the previous frame. Failure in object tracking may mean that the currently searched object did not exist in the previous frame but has newly appeared in the current frame, or that an object in the previous frame corresponding to the currently searched object actually exists, but the corresponding object has not been searched for algorithmic reasons. FIG. 20 illustrates a state in which tracking of a labeled object of interest has failed, according to one embodiment. For example, a ball was assigned labeling information as 'ball' among the objects of interest in the previous frame, and the same ball that existed in the previous frame was searched in the current frame, but for some reason it was not determined to be the same object. For example, this may occur due to lighting reasons, when the shape, shape, or color of an object changes as it moves or rotates, or when it is covered by or overlaps with another object. For example, if there were objects labeled as '7P' and 'ball' in the previous frame, but there are objects labeled as '7P' in the current frame, but no objects labeled as 'ball', then the 'ball' labeling information can be recommended as the labeling information for the searched and tracked object in the current frame. Alternatively, if there were objects labeled as '7P', 'referee', and 'ball' in the previous frame, but there are objects labeled as '7P' in the current frame, but no objects labeled as 'referee' and 'ball', then the 'referee' and 'ball' labeling information can be recommended as the labeling information for the searched and tracked object in the current frame.
[0083] In step S195, if the user selects one of the recommended lists, labeling information from the previous frame may be assigned. Otherwise (NO), in step S197, the computing device may assign new labeling information to the object, treating it as a newly appeared object in the current frame. The new labeling information may be automatically assigned by the computing device based on the type of the current object of interest, or may be input by the user.
[0084] Referring again to FIG. 9, in step S28, the computing device can determine whether the current input image is the final frame. If the current input image is not the final frame (No), the computing device can input the next frame, and if the current input image is the final frame (Yes), the computing device can terminate the object of interest search.
[0085] FIG. 21 is a flowchart of a method for a computing device to obtain position information of a ball in a sports video, according to one embodiment.
[0086] In step S2010, the computing device may receive a sports video. The sports video may be one of the frames constituting the sports video. The computing device may sequentially receive the frames constituting the sports video.
[0087] In step S2020, the computing device may use an artificial intelligence model to detect a ball in an input image. The artificial intelligence model may be trained to detect balls. The artificial intelligence model may detect an object determined to be a ball in the input image and output a probability that the detected object is a ball.
[0088] If the computing device successfully searches (Yes), it can classify the current input image as a key frame and determine the searched object as a ball (step S2070). In step S2070, the computing device can obtain the location of the object determined to be a ball. Successful search may mean that there is one object classified as a ball in the input image and that the probability of the object being a ball is greater than or equal to a reference value. The key frame becomes a reference frame for determining the ball area in other frames where ball search fails.
[0089] If the computing device fails to search (No), the computing device can determine whether there are multiple objects classified as balls or no objects classified as balls (step S2040). If no objects are classified as balls (No), the computing device can determine an alternative area where the ball is expected to be based on the location of the ball in the surrounding key frames (step S2050). The computing device can obtain the location of the alternative area. The method for determining the alternative area has been described above with reference to FIGS. 9 to 20.
[0090] In step S2060, if there are multiple objects classified as balls, the objects classified as balls can be determined as candidate regions, and the current input image can be classified as candidate frames. The situation where there are multiple objects classified as balls may be due to an object being falsely detected as a ball and / or the presence of multiple balls (a "new ball" situation, where a new ball enters the field). The computing device can acquire the location of each candidate region.
[0091] In step S2080, the computing device can determine whether the current input image is the final frame. If the current input image is not the final frame (No), the computing device can input the next frame. If the current input image is the final frame (Yes), the computing device can terminate the object of interest search.
[0092] In step S2090, the computing device may determine one of the candidate regions as a ball for each of the candidate frames. The candidate regions may be determined based on the temporal correlation with the object detected as a ball in surrounding key frames. Since it is unlikely that a specific object in each frame image constituting the video will suddenly and significantly change in the direction or size of movement, even if the frame changes, there is a high probability that the object is actually a ball if it is within the range where the specific object is expected to be. This is called temporal correlation, and a high temporal correlation may mean that, when considering the movement of the ball over time in surrounding key frames, the object is likely to move according to the movement of the actual ball in surrounding key frames over time and be at the location of the object. For example, candidate regions that are incorrectly detected as a ball often suddenly appear at a location far from the location of the ball in a previous or subsequent key frame. Among the candidate regions, the candidate region with the highest temporal correlation with the key frame can be determined as the final object to be searched (e.g., a ball). This method is also useful in cases where the original object being searched is a ball, but a new ball is introduced into the field, requiring the search to be performed using the new ball instead of the original.
[0093] According to one embodiment, the surrounding key frame for determining the temporal correlation may be determined as the key frame that is temporally closer to the current candidate frame among the immediately preceding key frame and the immediately following key frame. Referring to FIG. 22, the immediately preceding key frame (BF) is the closest frame among the previous key frames of the candidate frame (CF), and the immediately following key frame (FF) is the closest frame among the subsequent key frames of the candidate frame (CF). The immediately preceding key frame (BF) is 6 frames away from the candidate frame, and the immediately following key frame (FF) is 1 frame away from the candidate frame. In this embodiment, the immediately following key frame (FF) may be selected as the key frame for calculating the temporal correlation with the candidate regions of the candidate frame (CF). This is because a frame that is relatively closer to the candidate frame is more accurate for calculating the temporal correlation.
[0094] In one embodiment, a first value, which is a smaller value among the distance between a first candidate region of a candidate frame and a ball position in the preceding key frame and a distance between the first candidate region and a ball position in the following key frame, and a second value, which is a smaller value among the distance between a second candidate region and a ball position in the preceding key frame and a distance between the second candidate region and a ball position in the following key frame, may be compared, and if the first value is smaller, the temporal correlation of the first candidate region may be determined to be higher. This embodiment places more weight on the physical location of the ball than on the temporal distance between a reference frame and a current frame.
[0095] FIG. 23 is a flowchart of a method for a computing device to search for an object of interest, according to one embodiment.
[0096] In step S2020, original sports footage can be input.
[0097] In step S2040, an object of interest can be searched using an artificial intelligence model.
[0098] In step S2060, labeling can be performed on the object to be tracked.
[0099] In step S2070, the computing device can determine whether the current input image is the final frame. If the current input image is not the final frame (No), the computing device can input the next frame, and if the current input image is the final frame (Yes), the computing device can terminate the object of interest search.
[0100] The descriptions are intended to provide exemplary configurations and operations for implementing the present invention. The technical concept of the present invention encompasses not only the embodiments described above, but also implementations that can be achieved by simply modifying or altering the above embodiments. Furthermore, the technical concept of the present invention encompasses implementations that can be easily achieved by modifying or altering the above embodiments in the future.
Claims
1. A method for collecting ball movement information from a sports video performed on a computing device, A step of sequentially receiving each frame of the above sports video as an input video; A step of searching for at least one object classified as a ball in the input image using an artificial intelligence model; A step of classifying the input image into a key frame and determining the object as a ball if there is one object classified as a ball by the artificial intelligence model in the input image and the probability of the object being a ball is greater than or equal to a reference value; If there is no object classified as a ball by the artificial intelligence model in the input image, a step of determining an alternative area where a ball is expected to be based on the location of the ball in surrounding key frames; If there are multiple objects classified as balls by the artificial intelligence model in the input image, a step of determining the multiple objects as candidate regions and classifying the input image as candidate frames; and comprising a step of determining one of the above candidate regions of the above candidate frame as a ball; A method for collecting movement information of a ball in a sports video, wherein the step of determining determines, as a ball, the candidate region having the highest temporal correlation with surrounding key frames.
Citation Information
Patent Citations
System and method for tracking multiple objects
KR1020180093402A
Multiple nozzle system for architectural 3D printer
KR1020220141668A
Lane-changing vehicles detect method of CCTV
KR1020230065554A
Apparatus and method for analyzing basketball match based on multiple neural networks
KR102211135B1
Computer-implemented method for automated detection of a moving area of interest in a video stream of field sports with a common object of interest
US20230092516A1