Competition video editing method and system based on tracking holder data, and storage device
By tracking PTZ data and analyzing user activity, and selecting highlights from event videos, this technology solves the problems of high computing power consumption and misidentification in existing technologies, achieving efficient and low-consumption event video editing.
Patent Information
- Application Number
- CN202511442556.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies require significant computing power and time for automatic editing of sports videos, and often misidentify exciting segments, making it difficult to efficiently filter out the highlights.
By utilizing tracking gimbal data, video clips with gimbal yaw angles exceeding a threshold are selected. Combined with crowd activity and specific scene recognition, potential highlights are marked, and edited clips are generated.
It achieves efficient filtering of highlights from sports videos without using artificial intelligence algorithms, significantly reducing computing power consumption and time. The false recognition rate has increased, but it can be further reduced through subsequent large-scale model filtering.
Smart Images

Figure CN120916033A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of video clips, and particularly to a sports video clipping method based on tracking gimbal data, a system and a storage device. BACKGROUND
[0002] Automatic clipping of sports videos to obtain highlights or high points in the sports is a hot demand in the current sports shooting market. In the prior art, automatic clipping of sports videos usually uses AI analysis methods, such as Chinese patents “CN108900896A Video clipping method and device” and “CN116095496A Method and system for automatically extracting basketball game clipping segments”, which both identify various key targets and key features in the entire game video through artificial intelligence model algorithms.
[0003] Whether a small model or a large model is used to identify highlights, a large number of algorithm calculations are required. A sports event lasts about 1.5-2 hours, and there are more than 100,000 video frames at a rate of 20 frames per second; analyzing, identifying or reasoning on 100,000 video frames through artificial intelligence algorithms requires a lot of computing power and takes a long time. SUMMARY
[0004] In the prior art, there is a technical solution for tracking and shooting live events using AI analysis, such as CN120378750A, CN120321492A or CN120302158A, which can track a ball target, track a specified human target or track the fastest several people among multiple human targets in the picture. Based on this, to solve the above technical problems, the present application uses tracking and shooting logs associated with sports videos to efficiently screen possible highlights in sports videos without using machine learning or AI algorithms to obtain all possible clipping segments that may be highlights.
[0005] In a first aspect of the present application, a sports video clipping method based on tracking gimbal data is provided, the sports video comprising a tracking log, the tracking log comprising a gimbal yaw angle, the method comprising:
[0006] S100: screening all video segments when the absolute value of the gimbal yaw angle exceeds a first threshold from the sports video as a first video segment set;
[0007] S200: for each video segment in the first video segment set, marking any one time point in the continuous frame segment where the crowd activity continuously exceeds a second threshold to obtain a first time point set;
[0008] S300: save a video clip segment with a preset length before and after each first time point in the first time point set as a clip segment.
[0009] Preferably, the step S200 comprises:
[0010] S210: for each video clip segment in the first video clip segment set, exclude video data with crowd motion trend opposite to the yaw direction of the gimbal, to obtain a second video clip segment set;
[0011] S220: for each video clip segment in the second video clip segment set, mark any one time point in a continuous frame segment with crowd activity lasting beyond a second threshold, to obtain a first time point set.
[0012] In one possible embodiment, the event is football, and the method further comprises:
[0013] S101: screen all video clip segments with absolute value of the yaw angle of the gimbal not greater than a third threshold from the event video, as a third video clip segment set;
[0014] S201: for each video clip segment in the third video clip segment set, identify a kick-off scene and mark, to obtain a second time point set;
[0015] S301: based on each second time point in the second time point set, determine the last clip segment before the second time point as a goal video clip segment.
[0016] The identifying a kick-off scene and marking comprises:
[0017] marking any one time point in a continuous frame segment with crowd activity lasting below a fourth threshold.
[0018] In one possible embodiment, the event is football, and the method further comprises:
[0019] marking any one time point in a continuous frame segment with crowd gathering degree lasting beyond a fifth threshold from the event video, to obtain a third time point set;
[0020] based on each third time point in the third time point set, determining the last clip segment before the third time point as a corner kick video clip segment.
[0021] In one possible embodiment, the event is basketball, and the method further comprises:
[0022] S102: for each video clip segment in the first video clip segment set, screen video data with crowd motion trend opposite to the yaw direction of the gimbal, to obtain a fourth video clip segment set;
[0023] S202: For each video clip in the fourth video clip set, identify a goal-line outside-scene;
[0024] S302: Label all goal-line outside-scenes to obtain a fourth time point set;
[0025] S402: Based on each fourth time point in the fourth time point set, determine the last clip segment before it as a goal video clip.
[0026] The tracking log further includes a rate of a preset number of human body targets, and the crowd activity level is a mean value of the rate of the preset number of human body targets in the video frame.
[0027] In a second aspect of the present application, a sports video processing system is provided, comprising a shooting module and a clip module, the shooting module is adapted to track and shoot a sports event, generate a sports video and a tracking log, the tracking log includes a pan-tilt yaw angle; the clip module is adapted to obtain the sports video and the tracking log generated by the shooting module, and execute the method of the first aspect of the present application to generate video clips.
[0028] In a third aspect of the present application, a computing device is provided, comprising a memory and a processor, the memory has a computer program stored thereon, and the processor executes the computer program to implement the method of the first aspect of the present application.
[0029] In a fourth aspect of the present application, a computer readable storage medium is provided, having a computer program stored thereon, and the computer program is executed by a processor to implement the method of the first aspect of the present application.
[0030] In a fifth aspect of the present application, a computer program product is provided, having a computer program stored therein, and the computer program is executed by a processor to implement the method of the first aspect of the present application.
[0031] Through the clip method of the embodiments of the present application, the clip of the sports video is implemented. Compared with the method of analyzing by artificial intelligence in the prior art, although there may be relatively more misidentified wonderful clips, since the method is based on the clip of the video shot by AI intelligent tracking, the data generated in the shooting process can be fully utilized in the clip process, and the entire clip process can be completely free of operation of the artificial intelligence model, therefore, the consumption of computing power is extremely small, and the time is relatively short, and the technical effect is very obvious. BRIEF DESCRIPTION OF DRAWINGS
[0032] The above and other features, advantages, and aspects of the embodiments of the present application will become more apparent by describing in detail the following detailed description in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals indicate the same or similar elements, in which:
[0033] Figure 1 A flowchart of a method for clipping a sports video according to an embodiment of the present application;
[0034] Figure 2 A schematic diagram of a standard size football field according to an embodiment of the present application;
[0035] Figure 3 A schematic diagram of a structure of a terminal device or a server suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the purposes, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0037] In the field of computer vision analysis, it is generally assumed that the origin of the coordinate system is located at the upper left corner of the picture. In each of the embodiments of the present application, the upper left corner is used as the origin of the picture coordinate system unless otherwise stated. Those skilled in the art should know that such a coordinate system setting is not absolutely fixed. When the origin of the coordinate system is set at any position within or outside the picture, a corresponding technical solution can be obtained by simple adjustment of the present solution without creative work, which is within the scope of protection of the present application.
[0038] In addition, the term "and / or" in this document is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the front and rear associated objects.
[0039] The sports video clipping method provided by the embodiments of the present application can be run in various computing devices, including notebook computers, smart phones, tablet computers, or other smart devices with data acquisition capabilities, such as smart cameras, wearable devices, etc., and can also be run in the cloud to provide clipping services through network communication.
[0040] According to one embodiment of the present application, a method for clipping a sports video is provided. In the embodiment of the present application, the sports video to be clipped is captured by a tracking camera. The tracking camera can be implemented as a smart camera holder or a smart camera, which is placed at one side of the middle line of the court. The tracking camera can calculate the current tracking point position by analyzing the real-time captured image, and adjust the shooting angle based on the tracking point position. The specific scheme of the tracking camera is known in the art, which has been described in the summary of the present application and is not the content to be protected by the present application. While the tracking camera is capturing the sports video, a tracking log synchronized with the video needs to be output. The tracking log is a record of the tracking situation of each video frame, which includes the timestamp and the current holder angle of each video frame. The angle of the holder includes at least the pitch angle and the yaw angle, and in the present application, only the yaw angle of the holder is used. Preferably, a complete tracking log example is as follows:
[0041] 2025-06-17 19:54:52.040172
[0042] Frame2 fps16 SZ35 Yaw32.0 Pit12.0
[0043] FFilterP11 mot4 Deque10(V19,x19)(V13,x-12)(V8,x-5)
[0044] BallSz12 BallSuccess
[0045] Frame2 FIN:x1280 y720 V0
[0046] In the above example:
[0047] The first behavior log date and timestamp;
[0048] The Frame field in the second row is the frame number, which is the second frame of the current video in the above log; fps is the processing speed of the current video, which is 16 frames per second in the above log; Yaw and Pit are the current pitch angle and yaw angle;
[0049] There are 10 elements in the motion tracking Deque queue of the third row, but since the tracking algorithm of this test only tracks the preset number of personal target boxes with the fastest speed, the Deque only displays the speed of the three fastest targets and the displacement of the targets on the x-axis within 1 frame. The speed is always positive, and the horizontal and vertical coordinate offsets of the target in the current frame and the previous frame (or several previous frames, which can be freely defined according to the video frame rate) are calculated according to the Pythagorean theorem; preferably, if the perspective situation produces near large and far small, the obtained value can also be divided by the preset multiple of the target box height to obtain the optimized target speed. The displacement of the target on the x-axis is an integer, and the sign represents the moving direction. The displacement and speed are both in length units of pixels.
[0050] The BallSuccess in the fourth row indicates that the ball target is successfully detected, and BallSz12 indicates that the area of the ball target box is 12;
[0051] The FIN:x1280 y720 in the fifth row is the tracking point position of the current frame according to the tracking strategy, and v0 is the rotation speed in the rotation instruction calculated according to the tracking point position.
[0052] As shown in Figure 1 , the clip method of the event video in an embodiment of the present application includes:
[0053] S100: Select all video clips when the gimbal yaw angle absolute value exceeds the first threshold value from the event video as the first video clip set.
[0054] In basketball, football and other games, since the scoring area is on both sides of the court, exciting clips such as goals, steals, fast attacks, blocks, saves, etc. may occur at the positions of the goals on both sides of the court, and the probability of occurring in the middle area of the court is very small. At the same time, in the tracking shooting of the event, since the shooting device is placed on one side of the center line and will automatically adjust the shooting angle to track the crowd or the ball on the court, when the tracking target is located on both sides of the court, the shooting angle of the gimbal, i.e. the yaw angle, will also be located on the left or right side. In this embodiment, all video clips when the gimbal yaw angle absolute value exceeds the first threshold value are selected as the first video clip set, which is the first step of filtering the event video. In this step, a large number of non-active pictures with the gimbal angle of view located in the center or biased to the center can be filtered out, thereby reducing the amount of data to be analyzed. The rotation of the gimbal is continuous, and the gimbal yaw angle is a gradually changing value, so the multiple video frames when the gimbal yaw angle falls within an interval range must be continuous.
[0055] In one embodiment, according to the analysis of a large number of basketball and football data, almost all the highlights are located in the pictures when the yaw angle of the gimbal is greater than 30° or less than -30°. Therefore, the first threshold can be set to 30°; if a small amount of missed detection rate is acceptable to further reduce the amount of data to be analyzed, the first threshold can also be moderately increased, for example, set to 32° or 35°. Those skilled in the art should know that the shooting field of view will be very different according to the field of view angle of the shooting device; the larger the field of view angle of the device, the smaller the angle the gimbal needs to rotate when shooting the same content, therefore, the larger the field of view angle of the device, the smaller the first threshold can be set. Those skilled in the art can obtain the appropriate range of the first threshold suitable for the specified shooting device through a small amount of testing.
[0056] S200: For each video clip in the first video clip set, mark any one time point in the continuous frame clip whose crowd activity level continuously exceeds the second threshold to obtain a first time point set.
[0057] The crowd activity level refers to the activity of human targets in a video frame, which can be represented by the average speed of human targets in the picture. In the intense stage of the game, the competition for the ball and position will be more intense, so it can be considered that the crowd activity level is higher, and therefore the crowd activity level can be used to filter the highlights. In one possible embodiment, the displacement of the same human target in the x-axis and y-axis between the current frame and the previous frame or several previous frames is calculated to obtain the instantaneous speed of all human targets in each frame; the average speed of all human targets in the current frame, or the average speed of the fastest several human targets, is calculated as the crowd activity level in the current frame.
[0058] In another possible embodiment, the tracking strategy of the fastest several human targets has been used when the event video is shot, that is, the speed of all human targets, or at least the speed of the fastest several human targets, is included in the tracking log of each frame. At this time, only the speed of all corresponding human targets included in the tracking log needs to be directly read to obtain the average value, that is, the crowd activity level of the current frame.
[0059] The crowd activity level continuously exceeding the second threshold means that the activity level of the crowd remains above the second threshold within a preset time interval. The preset time interval can be set to 80-120 frames, corresponding to a time period of 4-6 seconds. The specific setting of the second threshold is related to the frame rate of the video acquisition device. When a 2K screen is used for shooting, through experiments, it can be set to between 100-150 pixels / frame; with the increase of the frame rate of the acquisition device, the specific value of the second threshold can be increased accordingly.
[0060] The value of the instantaneous speed is continuously gradual, so the change of the crowd activity is also gradual, and thus the frames of the crowd activity in a certain interval are also continuous. Any one time point in the continuous frame segment in which the crowd activity exceeds the second threshold value is marked to obtain a first time point set. Each video segment in the first video segment set can have one or more clip segments, that is, there can be only one first time point in each video segment, or there can be multiple first time points. Specifically, for each continuous frame segment in which the crowd activity exceeds the second threshold value, it can be considered to belong to one highlight moment, and any time point in the continuous frame segment is recorded as a mark. Specifically, the middle point of the continuous frame segment can be selected, or the starting point of the continuous frame segment can be selected. Considering that the crowd activity can be advanced or delayed, in order to comprehensively average the error, one random time point can also be selected, and the preset length before and after the saved video segment is also increased.
[0061] In a preferred embodiment, the S200 step can be divided into two steps, and the first video segment set is filtered again by the crowd movement trend to further exclude non-active segments. Specifically:
[0062] S210: For each video segment in the first video segment set, video data in which the crowd movement trend is opposite to the yaw direction of the holder is excluded to obtain a second video segment set.
[0063] The first video segment processed by the S100 step is a video segment in which the absolute value of the yaw angle of the holder exceeds the first threshold value, that is, a video segment in which the real-time picture faces the backcourt of both sides. In the backcourt picture, although all highlight moment videos are included, a certain length of non-active stage can also be included. For example, after the attack of the attacking side is terminated, the ball holding movement of the defending side can have high activity but no highlight moment. Therefore, preferably, the video in which the movement trend is opposite to the yaw direction of the holder is excluded by performing secondary screening in the first video segment according to the personnel movement trend.
[0064] The crowd motion trend refers to the overall motion direction of the players in the picture, which can be determined by the average value of the motion speed of a plurality of human body targets on the horizontal axis. In a possible embodiment, the displacement of the same human body target in the current frame and the previous frame is calculated to obtain the speed vector of all human body targets in each frame, and the speed component on the x-axis is decomposed. The average value of the x-axis speed component of all human body targets in the current frame, or the average value of the x-axis speed component of the fastest human body targets, is calculated, and the positive or negative nature is taken as the crowd motion trend in the current frame. In another possible embodiment, when the event video is shot, the strategy of tracking the fastest several people has been used, and the x-axis displacement of the fastest several human body targets has been provided in the tracking log. The motion trend can be determined by the positive or negative nature of the average value of the x-axis displacement in the tracking log. Taking the example of the example log in the foregoing embodiment, in this frame, the x-axis displacement of the three players with the fastest motion speed in the log is 19, -12, and -5, respectively, and the average value is 2 / 3. It is determined that the crowd motion trend in the current frame is the positive direction of the x-axis.
[0065] The yaw direction of the gimbal is determined by the gimbal yaw angle, which can be positive, zero or negative. When the gimbal yaw angle is positive, it means that the gimbal is deviated to the positive direction of the x-axis on the horizontal axis. When the gimbal yaw angle is negative, it means that the gimbal is deviated to the negative direction of the x-axis on the horizontal axis. When the gimbal yaw angle is zero, it means that the gimbal is not yawed.
[0066] If the gimbal yaw angle is positive and the crowd motion trend is the negative direction of the x-axis, or the gimbal yaw angle is negative and the crowd motion trend is the positive direction of the x-axis, it means that the crowd motion trend is opposite to the yaw direction of the gimbal.
[0067] S220: For each video segment in the second video segment set, mark the frames with the crowd activity exceeding the second threshold value to obtain a first time point set.
[0068] For each video segment in the second video segment set filtered twice in S210, mark the frames with the crowd activity exceeding the second threshold value, and the specific implementation manner is the same as the corresponding part in S200, which will not be described here.
[0069] S300: Based on each first time point in the first time point set, save the video segments with a preset length before and after the first time point as a clip segment.
[0070] Each of the first time points in the first time point set represents a possible highlight moment around the time point. Video clips of a preset length before and after each first time point, i.e., a possible highlight moment, are saved. Specifically, the preset length can be defined by the user, for example, 5-8 seconds. In order not to miss more shots, the value of the preset length can be increased. In order to make the finished video more concise, the value of the preset length can be reduced.
[0071] In one possible embodiment, when the event is football, the following method is preferably used to mark the types of the clip set and determine the goal video clip. In a football match, after one team scores a goal, according to the rules, both teams should send one player to the position on both sides of the center of the center circle to perform the kick-off action, at this time, other players must be outside the center circle. For a large model, this is a very easy scene to learn and identify; from the perspective of the log, at this time, almost all players are in a static state, and the player speed recorded in the log can be used for screening, and the significance is very high. Therefore, in the embodiment of the application, the previous clip segment of the center circle kick-off scene is determined as the goal clip segment by identifying the center circle kick-off scene. Specifically, the method includes:
[0072] S101: Screen all video clips in which the absolute value of the yaw angle of the gimbal is not greater than a third threshold value from the event video as a third video clip set;
[0073] When the absolute value of the yaw angle of the gimbal is small, the field of view of the gimbal is near the center line, and the complete center circle can be observed. The center circle kick-off scene is unlikely to occur in a scene with a large yaw angle of the gimbal, so in this embodiment, first, the video clips in which the yaw angle of the gimbal is small, i.e., the absolute value is not greater than a third threshold value, are screened from the event video as a third clip set, which is used for subsequent kick-off scene judgment, which can further reduce the calculation steps and time.
[0074] The third threshold value can be set as the possible range in which the center circle does not leave the field of view of the gimbal. Taking a standard football field (105 meters long, 68 meters wide, and a center circle radius of 9.15 meters) as an example, as shown in FIG. 5, the field of view angle of the gimbal is the angle between the rays f and g. When the view of the gimbal (located at point G) is tilted to the right, the limit yaw angle at which the entire center circle can be displayed is the angle between the ray h (the angle bisector of the angle between the rays f and g) and the positive direction of the y-axis. Since the center circle is 9.15 meters in radius, the limit yaw angle is 9.15 / 68=0.134 radians. Figure 2 Figure 2 In the limit case shown, the ray f is tangent to the circle E, and the sine of the angle between the straight line f and the positive direction of the y-axis can be obtained by the trigonometric function, which is the radius of the middle circle divided by half of the width of the football field. The aforementioned angle is about 15.6°. With the default field of view angle of 60° of a general smart phone, the limit yaw angle is 14.4°. When the view angle of the holder is deviated to the left side, the calculation result is the same because the football field is axisymmetric based on the center line (i.e., the y-axis). Therefore, in a general scenario, the third threshold value can be set to 15°. According to the size of the football field and the field of view angle of the device, the third threshold value will change accordingly. After reading the embodiments of the present application and Figure 2 the description, those skilled in the art can quickly calculate the actual required third threshold value without any creative thinking.
[0075] S201: For each video segment in the third video segment set, identify and mark the kick-off scene to obtain a second time point set;
[0076] Specifically, in a possible embodiment, the kick-off scene can be identified by training and analyzing a large model because the kick-off scene has obvious features. Preferably, in another possible embodiment, the crowd activity is used for screening because the crowd is generally static for a period of time during the kick-off scene. The continuous frame segment with a crowd activity continuously lower than a fourth threshold value is regarded as a kick-off scene. The calculation method of the crowd activity has been described in detail in the previous embodiments, which will not be repeated here. The specific setting of the fourth threshold value is related to the frame rate of the video collection device. When a 2K screen is used for shooting, the fourth threshold value can be set to between 50 and 100 pixels after experiments. With the improvement of the frame rate of the collection device, the specific value of the second threshold value can be improved accordingly.
[0077] The identification and marking of the kick-off scene are not intended to cut out the kick-off scene, and therefore, the marking points are not limited. The second time point can be marked at any time point of a kick-off scene, and the previous cut segment is unchanged regardless of the location of the second time point in the kick-off scene.
[0078] S301: Based on each second time point in the second time point set, the last cut segment before the second time point is determined as a goal video segment.
[0079] Specifically, through the aligned unique time axis, the last cut segment before each second time point can be found from all the cut segments as a goal video segment.
[0080] In this embodiment, steps S101-S301 are further marking of special events in the clip segment obtained in S300. Therefore, S101-S301 can be performed after S300, or can be performed in parallel with S100-S300, for example, S100-S101-S200-S201-S300-S301, or in other order, which are all within the protection scope of the present application.
[0081] In another possible embodiment, when the event is football, the type of the clip segment can be further marked to determine a set piece segment. Set pieces include corner kicks, free kicks and penalty kicks. When a set piece is executed in a football match, there will be an abnormal gathering of people, which will not occur when the match is proceeding normally. For example, when a corner kick is executed, most players will gather in the penalty area in front of the goal; when a free kick is executed, a certain number of players of the other team will gather in front of the goal; when a penalty kick is executed, other players will gather behind the player executing the kick. Therefore, the gathering degree of people in the picture can be calculated using this scene, and the scene where the gathering degree of people exceeds the fifth threshold value continuously is marked, and the clip segment where the marking is located or the nearest clip segment after the marking is the set piece segment.
[0082] In a preferred embodiment, the K-means clustering algorithm is used to cluster the spatial coordinates of the players extracted from the video picture with 4K resolution and 20 frames per second, so as to determine the gathering degree of people. Specifically, the two-dimensional coordinate points of all players are first extracted from each picture by a target detection model, and then the data of 20 frames per second are integrated to form a point set which is input into the K-means algorithm. The number of clusters K is preset to 3 to ensure that the distribution of the defense team, the attack team and the scattered individuals can be reflected. The average point distance of each cluster is calculated, and when the average point distance of a cluster is less than 1.2 meters and the number of people in the cluster is more than 6, it is determined that there is a high gathering area. The continuous frames are further counted, and when the gathering degree index exceeds the fifth threshold value in at least 15 consecutive frames (about 0.75 seconds), i.e. the average point distance is less than 1.2 meters and the number of people is more than 6, the time period is marked as a people gathering scene. If the marked scene corresponds to a clip segment or the nearest clip segment after the clip segment, the set piece segment is determined.
[0083] Taking a corner kick as an example, during the experiment, when 8 to 10 players are concentrated in an area with a radius of 5 meters in the penalty area, the K-means algorithm will automatically divide them into a single cluster, and the average distance between points in the cluster is about 0.9 meters, which is significantly lower than the threshold of 1.2 meters. In this scenario, the system continuously detects the aggregation state for more than 20 frames, meeting the threshold and duration requirements, and therefore automatically labels the segment as a corner kick and kick segment. Through the above method, the system can effectively identify corner kicks, free kicks, and penalty kicks and other kick events in a football match based on the calculation of the crowd aggregation degree.
[0084] In another possible embodiment, when the event is basketball, the preferred method for refining the label of the clip segment to determine the scoring video segment is as follows. In a basketball game, after one side scores a goal, the defending side needs to throw a boundary ball by one player leaving the baseline according to the rules. This is a very easy scene for the large model to learn and identify. To reduce the video length analyzed by the large model, this embodiment filters out the video data that may be the post-throwing segment for the large model analysis through the judgment of crowd movement trends. Specifically, the method includes:
[0085] S102: For each video segment in the first video segment set, filter out the video data with the crowd movement trend opposite to the yaw direction of the pan-tilt, to obtain a fourth video segment set;
[0086] When the defending side throws a boundary ball, all players will move to the center of the court, while the pan-tilt view is still on the side of the defending team at this time. Therefore, in this special scenario, the crowd movement trend is opposite to the yaw direction of the pan-tilt, and the video data with the crowd movement trend opposite to the yaw direction of the pan-tilt is filtered out to obtain the fourth video segment set. The calculation method of the crowd movement trend has been described in detail in the previous embodiments, which will not be repeated here.
[0087] S202: For each video segment in the fourth video segment set, identify the boundary ball throwing scene.
[0088] The analysis of the bottom line out-of-bounds action can be performed by a large model training and analysis method. In a preferred embodiment, the goal line is identified by a target recognition algorithm, and the relative position of the player to the goal line is analyzed. When an event of a player leaving the goal line is identified, the current video segment is identified as a bottom line out-of-bounds scene. Specifically, first, the target recognition algorithm is used to identify the person target box and the goal line in the current frame; second, it is determined whether the foot point of the person crosses the goal line, i.e., whether the bottom edge midpoint of each person target box moves from one side of the goal line to the other side. Let the identified goal line equation be ax+by+c=0; the foot point coordinates of the human target box fed back by the target recognition algorithm are substituted into the goal line function f(x,y)=ax+by+c to obtain the instantaneous value of the goal line function; the goal line function value of each person's foot point is monitored, and when the value changes in sign (i.e., from positive to negative, or from negative to positive), it means that the foot point has crossed the goal line. At this time, we determine that a player has left the goal line.
[0089] S302: Label all bottom line out-of-bounds scenes to obtain a fourth time point set;
[0090] Specifically, when the bottom line out-of-bounds scene is identified in S202, a time point in the video segment of the scene is selected for labeling to obtain a fourth time point set. The identification and labeling of the bottom line out-of-bounds scene is not intended to cut out the scene, so the labeling point is not limited and can be labeled at any time point in the bottom line out-of-bounds scene. The first clip segment before the second time point does not change regardless of the location of the second time point in the bottom line out-of-bounds scene.
[0091] S402: Based on each fourth time point in the fourth time point set, determine the last video segment before it as a goal video segment.
[0092] Specifically, through the aligned unique time axis, the last clip segment before each fourth time point can be found from all clip segments as a goal video segment.
[0093] The steps S102-S402 in this embodiment are further labeling of special events in the clip segments obtained in S300, so S102-S402 can be performed after S300, or can be run in parallel with S100-S300, for example, S100-S102-S200-S202-S302-S300-S402, or other orders.
[0094] In a possible embodiment, the sound information in the event video is analyzed to identify the referee whistle in the game. The referee whistle has a symbolic meaning in ball games and is a signal of interruption and the starting signal of an important event. The scenes where the referee whistle appears include the following categories:
[0095] 1) If the whistle scene is immediately followed by a crowd gathering scene, the video segment can be identified as a corner kick scene.
[0096] 2) If the whistle scene is immediately followed by a kick-off scene, the segment before the whistle scene can be identified as a goal scene. This identification scheme can be used as a supplementary step to the aforementioned S301, i.e., determining whether there is a whistle scene between each second time point and the last clip segment before it, and if so, determining the last clip segment before it as a goal segment.
[0097] Compared with the prior art method of analyzing using artificial intelligence, although the number of misidentified highlight segments may be relatively increased by the clipping method of the present application, the present method is based on AI intelligent tracking of the video for clipping, and in the clipping process, the data generated during the shooting process can be fully utilized, and the entire clipping process can completely or rarely call the artificial intelligence model for operation, so that the consumption of computing power is extremely small, the time is relatively short, and the technical effect is very obvious.
[0098] When it is necessary to further reduce misidentification, in a possible embodiment, further based on the clipping method of the present application, the video segment obtained by the clipping method in the present application can be further judged by the large model in the prior art, so as to further screen out the game highlight segments. Compared with the method of directly using the large model for analysis, the number of video frames analyzed by the large model in the present embodiment will be greatly reduced, thereby effectively reducing the demand for computing power.
[0099] Another embodiment of the present application provides a kind of event video processing system, comprising:
[0100] The shooting module is suitable for tracking shooting of the event, generating event video and tracking log, and the tracking log includes the yaw angle of the pan-tilt head; in a possible embodiment, the tracking log further includes the speed of a preset number of human body targets; the speed of the preset number of human body targets can be all human body targets or the fastest preset number of human body targets. The shooting module can be implemented as an automatic tracking camera, or as an automatic tracking pan-tilt head with a detachable installed camera.
[0101] The clipping module is adapted to acquire the event video and the tracking log generated by the shooting module, and perform the clipping method of the event video in the foregoing embodiments, which will not be described herein again. The clipping module can be implemented as an executable program or code in an intelligent terminal or a server.
[0102] In one possible implementation, the event video and the tracking log generated by the shooting module are automatically uploaded to the storage module in a designated cloud server, or are uploaded to the storage module in the cloud server in response to the behavior of a user, the cloud server is deployed with the clipping module, and when receiving a clipping requirement, the clipping module processes the event video in the storage module to generate a clipping segment.
[0103] Figure 3 A structural diagram of a computing device adapted to be used to implement the embodiments of the present application is shown. The computing device can be implemented as a terminal device or a server.
[0104] As shown in Figure 3 , the terminal device or the server includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded from a storage portion 508 to a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the terminal device or the server are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0105] The following components are connected to the I / O interface 505: an input portion 506 including a keyboard, a mouse, and the like; an output portion 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage portion 508 including a hard disk, and the like; and a communication portion 509 including a network interface card such as a LAN card, a modem, and the like. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 510 as needed, so that a computer program read therefrom is installed in the storage portion 508 as needed.
[0106] In particular, the above method flow steps can be implemented as a computer software program in accordance with embodiments of the present application. For example, embodiments of the present application include a computer program product which includes a computer program tangibly embodied on a machine readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable media 511. When the computer program is executed by the central processing unit (CPU) 501, the above-described functions defined in the system of the present application are executed.
[0107] It should be noted that the computer readable medium shown in the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or instrument, or any combination of the above. More specific examples of computer readable storage media can include, but are not limited to, electrical connections with one or more conductive wires, portable computer disks, hard disks, random access memories (RAM), read only memories (ROM), erasable programmable read only memories (EPROM or flash memory), optical fibers, portable compact disk read only memories (CD-ROM), optical storage devices, magnetic storage devices or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, device or instrument. In the present application, the computer readable signal medium can include a data signal propagating in a baseband or as a carrier wave part of a carrier wave, in which a computer readable program code is carried. Such a propagating data signal can take various forms, including but not limited to electromagnetic signals, optical signals or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in combination with an instruction execution system, device or instrument. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0108] In addition, the present application may, in embodiments, refer to the identification of some specific scenes or poses by large models. According to a certain scene or pose identification requirement, a commercial large model is selected to directly call its SDK for analysis, or an open source large model is selected for deployment, training and calling, wherein the training, parameter adjustment and calling are all common technical means for those skilled in the art, and are not within the scope to be protected by the present application.
[0109] The computer program product of the present application can be a computer program including a plurality of program instructions. When the computer program instructions are executed by one or more processors, the one or more processors perform a function according to the computer program instructions. The computer program product can be a computer program that is executed on one computer or a computer program that is executed on a plurality of computers.
[0110] The units or modules described in the embodiments of the present application can be implemented by software, or can be implemented by hardware. The described units or modules can also be implemented in a processor. In some cases, the name of the unit or module does not constitute a limitation on the unit or module itself.
[0111] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately and not be assembled into the electronic device. The computer readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the methods described in the present application.
[0112] As yet another aspect, the embodiments of the present application also provide a computer program product, and the computer program / instructions are executed by a processor to implement the method of any of the above embodiments.
[0113] The above description is merely preferred embodiments of the present application and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application disclosed in the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above application concept. For example, the above features can be replaced with similar technical features disclosed in the present application (but not limited to) to form a technical solution.
Claims
1. A method for event video clip based on tracking gimbal data, the event video comprising a tracking log, the tracking log comprising a gimbal yaw angle, characterized in that, The method comprises: S100: filtering out video clips in which the absolute value of the yaw angle of the gimbal exceeds a first threshold from the event video as a first video clip set; S200: for each video clip in the first video clip set, marking any one time point in the continuous frame clip in which the crowd activity level continuously exceeds a second threshold to obtain a first time point set; S300: based on each first time point in the first time point set, saving video clips of a preset length before and after the first time point as a clip segment.
2. The method of claim 1, wherein, The S200 step comprises: S210: for each video clip in the first video clip set, excluding video data in which the crowd movement trend is opposite to the yaw direction of the gimbal to obtain a second video clip set; S220: for each video clip in the second video clip set, marking any one time point in the continuous frame clip in which the crowd activity level continuously exceeds a second threshold to obtain a first time point set.
3. The method of claim 1, wherein, The event is football, and the method further comprises: S101: filtering out video clips in which the absolute value of the yaw angle of the gimbal is not greater than a third threshold from the event video as a third video clip set; S201: for each video clip in the third video clip set, identifying a kick-off scene and marking to obtain a second time point set; S301: based on each second time point in the second time point set, determining the last clip segment before the second time point as a goal video clip.
4. The method of claim 3, wherein, The identification of the kick-off scene and the marking comprise: Marking any one time point in the continuous frame clip in which the crowd activity level continuously falls below a fourth threshold.
5. The method of claim 1, wherein, The event is football, and the method further comprises: Marking any one time point in the continuous frame clip in which the crowd gathering degree continuously exceeds a fifth threshold in the event video to obtain a third time point set; Based on each third time point in the third time point set, determining the clip segment in which the third time point is located as a corner kick segment.
6. The method of any one of claims 1-5, wherein, The tracking log further comprises the speed of a preset number of human body targets, and the crowd activity level is the average speed of the preset number of human body targets in the video frame.
7. An event video processing system comprising a shooting module and an editing module, characterized in that, The shooting module is adapted to track and shoot the event to generate an event video and a tracking log, and the tracking log comprises a yaw angle of a gimbal; the clip module is adapted to obtain the event video and the tracking log generated by the shooting module and execute the method according to any one of claims 1-6 to generate a video clip.
8. A computing device comprising a memory and a processor, said memory having stored thereon a computer program, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method according to any one of claims 1-6.
10. A computer program product in which a computer program is stored, characterized in that the computer program comprises program code which, when executed by a computer, causes the computer to perform the method according to any one of the preceding claims. The computer program is executed by the processor to implement the method according to any one of claims 1-6. The computer program is executed by the processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Video editing method and apparatus
CN108900896A
Method and system for automatically extracting wonderful video of basketball match
CN116095496A
Intelligent tracking shooting method, shooting holder, equipment and storage medium
CN120302158A
Automatic tracking shooting method and device for motion in fixed score area, and storage medium
CN120321492A
Tracking shooting method, device and equipment for ball games and storage medium
CN120378750A