Site element marking method, tracking shooting method based on site boundary filtering

Through the method of linking the control end with the gimbal, the user clicks on the site elements and converts coordinates on the control end to mark the site edges, solving the problem of difficult filtering of interference targets outside the site edges in the prior art, and improving the multi-objective tracking accuracy of sports intelligent shooting and the stability of video analysis.

CN120182898BActive Publication Date: 2025-08-01SUZHOU DEEPSIGHT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510645876.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-01
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively filter out interfering targets outside the sideline of the field in sports intelligent shooting, resulting in a decrease in the accuracy of multi-objective tracking.

Method used

By linking the control end with the gimbal, the user clicks on the site elements to record the coordinates and gimbal angle on the control end, and performs coordinate conversion and mapping on the analysis end, marks the site edges, and filters out the human targets located outside the edges.

Benefits of technology

It improves the accuracy of multi-objective tracking, reduces the interference of off-site personnel to the tracking algorithm, and improves the stability and robustness of video analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182898B_ABST
    Figure CN120182898B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for marking site elements and a tracking shooting method based on filtering of site sidelines. Among them, the method for marking site elements includes: in response to a user's click on a site element in the control terminal screen, recording the first coordinate of the user's click and the pan-tilt angle at the time of clicking; based on the first coordinate, obtaining the second coordinate of the site element in the analysis terminal screen; based on the pan-tilt angle, converting the second coordinate into the third coordinate in the virtual three-dimensional coordinate system; when it is necessary to mark the site element in the analysis terminal screen, based on the third coordinate and the second pan-tilt angle at the time of the video frame, mapping the third coordinate to the position coordinate of the site element in the video frame. This embodiment provides a solution for manually marking site elements by the user and mapping the coordinates of the elements to the real-time screen of the analysis terminal, which can solve the problems of tight computing power, poor recognition effect, and easy loss of position caused by simply using video analysis technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical fields of computer vision and intelligent sports shooting, and particularly to a method for filtering field sidelines and a method for marking field elements. Background Art

[0002] In the field of intelligent sports shooting technology, the prior art generally performs object recognition and tracking on the human targets of athletes, including human body shapes, postures, face data, etc. Especially when it comes to multi-object tracking, the accuracy of identifying the targets and positions of athletes has a great impact on the tracking effect. In the prior art, the factor that has the greatest interference on misidentification is off-field personnel. In basketball, football and other games, there are usually non-sporting targets such as onlookers, coaches, substitute players, etc. standing outside the field, and these targets are usually located relatively close to the field sidelines. Therefore, in the object tracking algorithm, these off-field personnel are easily identified as athletes and tracked, which will significantly affect the tracking accuracy of sporting targets. In the prior art, there are some ideas to identify the sidelines in the real-time video image and filter out the personnel whose foot points are outside the sidelines based on this. CN116416648A solves the personnel filtering problem by the above method. However, in fact, the on-site situation of the stadium is very complex and the difficulty of sideline recognition is great, resulting in poor filtering effect based on automatic sideline recognition. Summary of the Invention

[0003] In view of this, the present application provides a tracking and shooting solution based on field sideline filtering, which can effectively filter out interference targets outside the field sidelines through the linkage between the control end and the pan-tilt head, and improve the effect of multi-object tracking.

[0004] In the first aspect of the present application, a method for marking field elements is provided, and the method includes:

[0005] Responding to a user's click on a field element in the first real-time image of the control end, recording the first coordinate of the user's click and the first pan-tilt head angle at the time of the click;

[0006] Based on the first coordinate, obtaining the second coordinate of the field element in the second real-time image of the analysis end;

[0007] Based on the first pan-tilt head angle, converting the second coordinate into a third coordinate in a virtual three-dimensional coordinate system;

[0008] When it is necessary to mark a field element in the video frame of the second real-time image, based on the third coordinate and the second pan-tilt head angle at the time when the video frame is located, mapping the third coordinate into a fourth coordinate, and the fourth coordinate is the position coordinate of the field element in the video frame.

[0009] In a possible implementation, the first real-time image is acquired by the shooting end and synchronously transmitted to the control end, and is displayed in the image of the control end; based on the first coordinate, the second coordinate of the field element in the second real-time image of the analysis end is obtained, including:

[0010] Based on the position difference and / or field of view angle difference between the camera at the shooting end and the camera at the analysis end, perform displacement and / or distortion processing on the first coordinate.

[0011] Optionally, convert the second coordinate to a third coordinate in a virtual three-dimensional coordinate system, including:

[0012] Convert the second coordinate to a first intermediate coordinate in a coordinate system with the center of the second real-time image as the origin;

[0013] Add a third-dimensional coordinate to the first intermediate coordinate to obtain a first three-dimensional vector, and the modulus length of the first three-dimensional vector is the focal length value of the camera at the analysis end;

[0014] Based on the first pan-tilt angle, rotate the first three-dimensional vector reversely back to the initial angle to obtain the third coordinate.

[0015] In a possible implementation, based on the third coordinate and the second pan-tilt angle at the time when the video frame is located, map the third coordinate to a fourth coordinate in the video frame, including:

[0016] Rotate the third coordinate according to the second pan-tilt angle to obtain a second three-dimensional vector;

[0017] Extract the first two dimensions in the second three-dimensional vector to obtain a second intermediate coordinate;

[0018] Convert the second intermediate coordinate to a fourth coordinate in the default coordinate system of the second real-time image.

[0019] In the second aspect of the present application, a tracking shooting method based on filtering of the field boundary is provided, and the method includes:

[0020] Mark a plurality of boundary points of the field in the real-time video frame, and the plurality of boundary points at least include each folding point of the boundary of the field. The boundary points are marked by the method of the first aspect of the present application, and each boundary point corresponds to a fifth coordinate;

[0021] Based on all the fifth coordinates, enclose to obtain a field polygon in the real-time video frame;

[0022] Identify all human targets in the real-time video frame to obtain a first target set, and calculate the foot points of each human target in the first target set;

[0023] Filter out all human targets whose foot points are located outside the field polygon from the first target set to obtain a second target set;

[0024] Track and photograph the human targets in the second target set.

[0025] In the third aspect of the present application, a site element marking system is provided, including:

[0026] A control terminal, which is adapted to control the pan-tilt angle in real time, obtain and display the first real-time image of the shooting terminal; the control terminal is adapted to respond to the user's click on the site element in the first real-time image, record the first coordinate of the user's click and the first pan-tilt angle at the time of clicking and send them to the analysis terminal.

[0027] A pan-tilt, including:

[0028] A shooting terminal, which is adapted to obtain the first real-time image;

[0029] An analysis terminal, which is adapted to obtain the second coordinate of the site element in the second real-time image of the analysis terminal based on the first coordinate; based on the first pan-tilt angle, convert the second coordinate into the third coordinate in the virtual three-dimensional coordinate system; when it is necessary to mark the site element in the video frame of the second real-time image, map the third coordinate to the fourth coordinate based on the third coordinate and the second pan-tilt angle at the time when the video frame is located, and the fourth coordinate is the position coordinate of the site element in the video frame.

[0030] Preferably, the analysis terminal is further adapted to perform displacement and / or distortion processing on the first coordinate based on the position difference and / or field of view angle difference between the camera of the shooting terminal and the camera of the analysis terminal.

[0031] In the fourth aspect of the present application, a tracking and shooting system is provided, including:

[0032] A control terminal, which is adapted to control the pan-tilt angle in real time, obtain and display the real-time image of the shooting terminal; the control terminal is adapted to respond to the user's click on the site boundary point in the real-time image, and the boundary point at least includes each folding point of the boundary of the site, record the coordinates of each boundary point clicked by the user and the pan-tilt angle at the time of clicking and send them to the analysis terminal.

[0033] A pan-tilt, including:

[0034] A shooting terminal, which is adapted to obtain the first real-time image;

[0035] An analysis terminal, which is adapted to execute the method of the second aspect of the present application to obtain a tracking and shooting strategy and send it to the moving terminal;

[0036] A moving terminal, which is adapted to adjust the angle of the pan-tilt based on the tracking and shooting strategy.

[0037] In a fifth aspect of the present application, there is provided an electronic device, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the method according to the first aspect or the second aspect of the present application is implemented.

[0038] In a sixth aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method according to the first aspect or the second aspect of the present application is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present application will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0040] Figure 1 FIG. is a flowchart of a method for marking site elements according to an embodiment of the present application;

[0041] Figure 2 FIG. is a schematic diagram of a pan-tilt head including a shooting end according to an embodiment of the present application;

[0042] Figure 3 FIG. is an architecture diagram of a site element marking system according to an embodiment of the present application;

[0043] Figure 4 FIG. is an architecture diagram of a tracking shooting system according to an embodiment of the present application;

[0044] Figure 5 FIG. is a schematic structural diagram of a terminal device or a server suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0046] In the field of computer vision analysis, it is usually default that the origin of the coordinate system is located at the upper left corner of the screen. In each embodiment of the present application, unless otherwise stated, the upper left corner is used as the origin of the screen coordinate system; those skilled in the art should be aware that such a coordinate system setting is not absolutely fixed. When the origin of the coordinate system is set at any position inside or outside the screen, the corresponding technical solutions obtained by making simple adjustments to the present solution without creative efforts all fall within the scope of protection of the present application.

[0047] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0048] Figure 1 Flowchart of a method for marking site elements according to an embodiment of the present application. Figure 1 , the method comprising:

[0049] First, in response to the user clicking on a site element in the first real-time image of the control terminal 31 , the first coordinate of the user's click and the first pan / tilt angle at the time of the click are recorded.

[0050] Among them, the first coordinate clicked by the user, that is, the coordinate of the pixel clicked by the user on the first real-time screen; the first gimbal angle when the user clicks, that is, the real-time angle of the gimbal 22 when the user clicks, including the pitch angle (rotation angle on the pitch axis) and yaw angle (rotation angle on the yaw axis) of the gimbal 22.

[0051] In one implementation, the control terminal 31 is a smart terminal in communication with the capture terminal 21. The control terminal 31 and the capture terminal 21's image on the pan / tilt head 22 are synchronized in real time. Specifically, the first real-time image from the capture terminal 21 is synchronized to the control terminal 31. Furthermore, the control terminal 31 can also transmit control signals through the capture terminal 21 to control the rotation of the pan / tilt head 22. Therefore, in order to simultaneously obtain the first coordinates in the first real-time image and the first pan / tilt head angle at the time of the click, data collection must be performed using the control terminal 31; direct marking on the capture terminal 21 is not possible.

[0052] In another implementation, the shooting end 21 and the analysis end 23 are both fixed devices on the pan-tilt head 22, the control end 31 is in communication with the pan-tilt head 22, and can choose to obtain the real-time image of the shooting end 21 or the real-time image of the analysis end 23.

[0053] Before accepting a user click, the user can preferably be prompted to select or be prompted with the name of the field element to be marked, so that the analysis terminal 23 can accurately identify the element information in subsequent marking. Field elements are generally used to identify specific events in a game. For example, points on the sidelines of a court can be used to identify the court boundary, thereby identifying out-of-bounds events in various games; points on the basket can be used to assist in identifying goal events; points on the free throw line in basketball can be used to identify free throw events in basketball; points on the left and right goal posts can be used to identify the location of the goal, thereby identifying goal events in soccer games; points on both sides of the net can be used to identify the location of the net, thereby identifying hit events in competitive sports, etc.

[0054] In a possible implementation, the elements that the user needs to mark may not be in the real-time image of the initial state. Therefore, the user may control the pan-tilt 22 to rotate to mark specific elements. At this time, the user remotely controls the pan-tilt 22 to rotate through the control terminal 31 until the specific element is displayed in the real-time image. Preferably, since the distortion caused by the difference in the field of view angle between the shooting-end camera 211 and the analysis-end camera 221 may not be completely eliminated, and the distortion is smaller at positions closer to the center of the image, it can be required that the user controls the pan-tilt 22 to rotate until the specific element is displayed in the center of the real-time image before starting to mark.

[0055] Secondly, based on the first coordinate, obtain the second coordinate of the site element in the second real-time image of the analysis end 23.

[0056] Wherein, the second real-time image is the real-time image collected by the analysis-end camera 221.

[0057] In a possible implementation, the analysis-end camera 221 is also suitable for collecting and recording real-time video. At this time, the shooting end 21 and the analysis end 23 are the same device, and the first real-time image in the control terminal 31 is the same as the second real-time image in the analysis end 23. At this time, there is no difference between the first coordinate and the second coordinate.

[0058] In another possible implementation, although the shooting end 21 and the analysis end 23 are not the same device, the control terminal 31 can directly obtain the real-time image of the analysis end 23. At this time, the first real-time image in the control terminal 31 is the same as the second real-time image in the analysis end 23, and there is no difference between the first coordinate and the second coordinate.

[0059] In yet another possible implementation, the analysis end 23 and the shooting end 21 are different devices installed on the pan-tilt 22. At this time, there must be differences in the position and / or field of view angle between the analysis-end camera 221 and the shooting-end camera 211. Since the analysis end 23 needs to collect a large amount of visual data, preferably, the analysis end 23 uses a wide-angle camera; since the shooting end 21 needs to have a better view, preferably, the position of the shooting end 21 is usually higher. Therefore, preferably, based on the position difference and / or field of view angle difference between the shooting-end camera 211 and the analysis-end camera 221, the first coordinate is subjected to displacement and / or distortion processing. Combining the description in this application, those skilled in the art should understand that the aforementioned position difference and field of view angle difference may not exist, may only have a position difference, or may have both a position difference and a field of view angle difference. Correspondingly, the process of processing the first coordinate may not require any processing, may only perform displacement processing, or may perform both displacement and distortion processing.

[0060] Figure 2 Schematic diagram of a pan-tilt 22 including a shooting end 21 according to an embodiment of the present application, showing the positional relationship between the analysis end camera 221 and the shooting end camera 211 in one implementation. As Figure 2 shown, the shooting end 21 is detachably fixed on the pan-tilt 22, and the shooting end camera 211 is located in the upper right corner of the analysis end camera 221. Since Figure 2 is the rear view of the shooting end 21, actually, the shooting end camera 211 is located in the upper left corner of the analysis end camera 221.

[0061] Based on the positional difference between the shooting end camera 211 and the analysis end camera 221, the first coordinate is displaced, that is, the first coordinate is displaced by corresponding pixels in the opposite direction of the shooting end 21 relative to the analysis end 23. In the embodiment as Figure 2 shown, since the shooting end camera 211 is located in the upper left corner of the analysis end camera 221, the real-time image of the shooting end camera 211 is in the upper left corner of the real-time image of the analysis end camera 221. Then, the position of the same target element in the real-time image of the analysis end camera 221 is more towards the upper left corner compared to its position in the real-time image of the shooting end camera 211. Therefore, in the embodiment as Figure 2 shown, the first coordinate should be moved towards the upper left corner. The specific pixel value to be moved is determined by the distance between the two cameras and the DPI (Dots Per Inch, the number of pixels per inch of the physical screen) of the shooting end camera 211. The specific formula is: movement pixel value = distance between the two cameras * DPI / 2.54. Among them, 2.54 is the conversion coefficient between inches and centimeters.

[0062] Based on the field of view angle difference between the shooting end camera 211 and the analysis end camera 221, the first coordinate is distorted, that is, the first coordinate is converted according to the distortion coefficient using an existing distortion model. Through pre-calibrated distortion parameters (such as radial distortion coefficients k1, k2, tangential distortion coefficients p1, p2), the corrected coordinates are calculated according to the distortion model. The present application does not specifically limit the distortion model and the parameter calibration method. In one implementation, the Zhang Zhengyou calibration method can be used to extract the corner coordinates as control points by shooting checkerboard images from multiple angles, construct a non-linear model to solve the parameters, and the following process can be adopted: image acquisition → corner detection → establishment of distortion equation → least squares optimization (LM algorithm) → output of parameters such as k1, k2, p1, p2. An example of a specific distortion model is:

[0063]

[0064] where, x real and y realis the coordinate of the point to be measured before distortion correction, i.e., the first coordinate; x corrected and y corrected is the coordinate of the point to be measured after distortion correction, i.e., the second coordinate; r is the distance of the point to be measured from the center of the screen before distortion correction; c x and c y are the coordinates of the center point of the screen.

[0065] In actual engineering, to avoid iterative calculations, a mapping table is directly constructed and the look-up table method is used.

[0066] Again, based on the first pan-tilt angle, the second coordinate is converted into a third coordinate in the virtual three-dimensional coordinate system.

[0067] In a possible implementation, the specific process of converting the second coordinate into the third coordinate is divided into three steps:

[0068] The first step: Since operations such as rotation and 3D projection of the virtual coordinate composition algorithm need to be based on the optical axis (image center), first the second coordinate is converted into a first intermediate coordinate in the coordinate system with the center point of the second real-time image as the origin. Specifically, if the second coordinate is (x, y), then after translating the coordinate system to the original image center (x c , y c ) as the origin, the first intermediate coordinate is (x - x c , y - y c ).

[0069] The second step: A third-dimensional coordinate is added to the first intermediate coordinate to obtain a first three-dimensional vector, and the modulus of the first three-dimensional vector is the focal length value of the analysis end camera 221. To solve the mapping distortion problem from a two-dimensional image to a three-dimensional space and achieve coordinate unity for multi-view and large-range scenes through the geometric characteristics of the spherical model, therefore, spherical transformation is adopted, and a third-dimensional coordinate is added to the two-dimensional first intermediate coordinate to be converted into a virtual three-dimensional coordinate.

[0070] Specifically, a Z-direction component is added to the first intermediate coordinate (x - x c , y - y c ) so that the vector falls on the sphere with a fixed radius of the focal length value f (in pixel units), then the expression of the first three-dimensional vector is obtained as . In a specific implementation, the right-hand coordinate system is used to add the z component, and the right-hand coordinate system is the coordinate system in which the X axis is to the right, the Y axis is downward, and the Z axis is forward in the 2D plane.

[0071] In the third step, based on the first pan-tilt angle, reverse-rotate the first three-dimensional vector back to the initial angle to obtain the third coordinate. By reversely compensating for the perspective change caused by the movement of the pan-tilt 22, the 2D image points captured in different poses are uniformly mapped to the initial coordinate system, thereby ensuring the consistency of the global coordinates.

[0072] For example, if the pan-tilt 22 rotates by an angle around the Y-axis, the image points it captures are equivalent to the result of the original coordinate system rotating clockwise around the Y-axis. If the pan-tilt 22 rotates by an angle around the X-axis, the image points it captures are equivalent to the result of the original coordinate system rotating clockwise around the X-axis. To map the current image points to the initial coordinate system (yaw = 0, pitch = 0), a reverse rotation needs to be applied to the first three-dimensional vector (i.e., rotate - around the Y-axis and - around the X-axis), thereby canceling the influence of the movement of the pan-tilt 22. Specifically, the rotation matrix for the Y-axis is , and the rotation matrix for the X-axis is , and the total rotation matrix is .

[0073] Among them, is the angle of the reverse rotation of the first three-dimensional vector around the Y-axis; is the angle of the reverse rotation of the first three-dimensional vector around the X-axis; R y (- ) is the rotation matrix of the first three-dimensional vector around the Y-axis; R x (- ) is the rotation matrix of the first three-dimensional vector around the X-axis; R is the total rotation matrix of the first three-dimensional vector in the three-dimensional coordinate system.

[0074] Multiply the first three-dimensional vector by the total rotation matrix to obtain the third coordinate, which is the coordinate of the site element in the three-dimensional coordinate system of the virtual spherical surface.

[0075] Figure 1The above three steps are the pre - processing of the behavior data of the user - marked site elements obtained by the control end 31. Furthermore, as the time axis and the perspective change, when it is necessary to mark the site elements in the video frames of the second real - time screen, based on the third coordinate and the second pan - tilt angle at the time of the video frame, the third coordinate is mapped to a fourth coordinate, and the fourth coordinate is the position coordinate of the site element in the video frame. The main function of "marking" the site element is to assist the AI algorithm in position calculation. Therefore, generally, it is not necessary to actually display the position of the site element in the screen; "marking" the site element actually means obtaining the position coordinate of the site element in the video frame, that is, obtaining the fourth coordinate of the site element in the video frame of the second real - time screen. The time of the video frame refers to the real - time time corresponding to the position of the video frame on the time axis. In one embodiment, there is a one - to - one correspondence among the second real - time screen video, the time axis, and the pan - tilt log. Each second real - time screen frame corresponds to a time point on the time axis and a pan - tilt log, and the pan - tilt log contains the pan - tilt angle.

[0076] This step is to inversely map the third coordinate in the virtual spherical three - dimensional coordinate system back to the fourth coordinate in the plane coordinate system of the second real - time screen. Specifically:

[0077] Rotate the third coordinate by the second pan - tilt angle to obtain a second three - dimensional vector.

[0078] For example, if the second pan - tilt angle is a rotation around the X - axis and a rotation around the Y - axis , then the rotation matrix around the X - axis is , and the rotation matrix around the Y - axis is .

[0079] Among them, is the angle of the second pan - tilt angle rotating around the X - axis; is the angle of the second pan - tilt angle rotating around the Y - axis; R x ( ) is the rotation matrix of the second pan - tilt angle around the X - axis; R y ( ) is the rotation matrix of the second pan - tilt angle around the X - axis.

[0080] Multiply the third coordinate by R x ( ) and then multiply by R y ( ), and the second three - dimensional vector is obtained;

[0081] Extract the first two dimensions of the second three-dimensional vector to obtain a second intermediate coordinate, and convert the second intermediate coordinate into a fourth coordinate in the default coordinate system of the second real-time screen.

[0082] For example, if the second three-dimensional vector is (x', y', z'), then the first two dimensions are extracted to obtain the second intermediate coordinate (x', y'). At this time, the second intermediate coordinate is the coordinate of the site element in the coordinate system with the center point as the origin in the second real-time screen, and it needs to be converted into the default coordinate system, that is, the coordinate in the coordinate system with the upper left corner of the screen as the origin in the embodiments of the present application. Suppose the center point coordinate is (x' c , y' c ), then the converted fourth coordinate is (x' + x' c , y' + y' c ). The fourth coordinate is the real position coordinate of the site element in the current video frame of the second real-time screen.

[0083] The above embodiments provide a solution in which the user manually marks the site elements and maps the coordinates of the elements to the real-time screen of the analysis end 23, which can solve the problems of tight computing power, poor recognition effect, and easy loss of position caused by simply using video analysis technology. For example, in the real-time shooting and analysis of a basketball game scene, it is necessary to locate the position of the basket to identify scoring events. If a pure video analysis algorithm is used, it needs to be sent frame by frame into the inference model for recognition, and when the basket is blocked by a player or the ball, it cannot be located. If the technical solution of the above embodiments is adopted, before the user starts shooting, the user can be guided to control the rotation of the pan-tilt 22 through the control end 31 to manually mark the positions where the two baskets are located; after the analysis end 23 obtains the marked positions and the pan-tilt angles at the time of marking, a virtual spherical three-dimensional coordinate system can be generated, and the positions of the two baskets are recorded in this virtual coordinate system. When the user starts shooting, if the pan-tilt 22 rotates to the position where the basket is visible, the analysis end 23 can immediately deduce the coordinates of the basket position in the real-time screen, and it will not disappear due to occlusion, so that events related to the basket can be judged more optimally, and the stability and robustness of the video analysis results can be improved.

[0084] Another embodiment of the present application provides a tracking shooting method based on filtering by the site sidelines. Through the site element marking method of the above embodiments, several points on the site sidelines are marked by the user in advance, and then the analysis end 23 connects them into the site sidelines, and filters the personnel outside the site based on this sideline. The solution of this embodiment can significantly reduce the interference of off-site personnel on target recognition in the multi-target tracking algorithm.

[0085] First, multiple edge points of the venue are marked in the real-time video frame, wherein the multiple edge points include at least each inflection point of the edge of the venue. The edge points are marked using the venue element marking method in the embodiment of the present application. Each edge point corresponds to a fifth coordinate, and based on all the fifth coordinates, a venue polygon in the real-time video frame is enclosed.

[0086] Competition venues are typically rectangular, though polygonal venues like octagonal cages are also possible. Curved edges are rarely present in sports venues. Therefore, to completely map the venue's edges on the analysis terminal 23, the user must at least mark each inflection point (i.e., corner point) along the edges. For example, using a rectangular venue, the user must mark at least the four corners of the venue as venue elements, with each point corresponding to a fifth coordinate. The specific marking method is the same as the venue element marking method described in the previous embodiments of this application and will not be further elaborated here.

[0087] Camera lenses always cause distortion. Therefore, straight lines in the image will not appear perfectly straight, but rather curved. The longer the line, the more pronounced the distortion. Therefore, it's best to mark not only the corners of the field's edges, but also other points along the edges, such as the midpoints. This will create a polygon that more closely resembles the actual field's shape.

[0088] Next, all human targets in the real-time video frame are identified to obtain a first target set, and the foot points of each human target in the first target set are calculated. The foot points of the human targets are the coordinates of the human targets' feet. In one embodiment, all human targets in the real-time video frame are identified using a target recognition algorithm such as YOLO (You Only Look Once) or SSD (Single Shot MultiBox Detector), with each human target corresponding to a target box. In one embodiment, the midpoint of the bottom edge of the target box can be used as the target foot point. Preferably, in edge filtering scenarios, it is necessary to minimize misidentification caused by target box errors. Therefore, the foot point can be moved up from the midpoint of the bottom edge of the target box by a predetermined ratio. This ratio is related to the error rate of the target recognition algorithm used, and this ratio is not limited in this application. For example, in the YOLOv8 algorithm, this ratio can be set to 5%, that is, the foot point is moved up from the midpoint of the bottom edge of the target box by 5% of the target box height.

[0089] Thirdly, all human targets whose feet are outside the field polygon are filtered out from the first target set to obtain a second target set.

[0090] Specifically, there are various simple plane geometry calculation methods for determining whether a point on a plane is located inside a polygon, and the present application does not limit this comparison. For example, the coordinates of the point can be substituted into the polygon equation to determine the positive or negative value. If the value is positive, the point is outside the polygon; if the value is negative, the point is inside the polygon.

[0091] For all human targets whose foot points are located outside the field polygon, it can be considered that they are actually outside the field boundary and do not belong to the athletes competing in the field. Therefore, they are excluded from the set of targets to be tracked to obtain a second target set.

[0092] Finally, track and photograph the human targets in the second target set. Specifically, the tracking and photographing can adopt general multi-target tracking algorithms, such as Kalman filters, SORT (Simple Online and Realtime Tracking), etc. The present application does not limit this. Any situation where any publicly available or proprietary multi-target tracking algorithm is used to implement this solution falls within the protection scope of the present application.

[0093] The method of the embodiment of the present application has been described above. Next, the system of the embodiment of the present application is provided.

[0094] Figure 3 It is an architecture diagram of a site element marking system 30 according to an embodiment of the present application. As Figure 3 shown:

[0095] The site element marking system 30 includes a control end 31 and a pan-tilt head 22. The control end 31 is communicatively connected to the pan-tilt head 22, wherein:

[0096] The control end 31 is adapted to perform real-time control on the angle of the pan-tilt head 22, obtain and display the first real-time image of the shooting end 21; the control end 31 is adapted to respond to the user's click on the site element in the first real-time image, record the first coordinate clicked by the user and the first pan-tilt head angle at the time of clicking and send them to the analysis end 23.

[0097] The pan-tilt head 22 includes:

[0098] The shooting end 21 is adapted to obtain the first real-time image;

[0099] The analysis terminal 23 is adapted to obtain a second coordinate of the site element in a second real-time image of the analysis terminal 23 based on the first coordinate; convert the second coordinate into a third coordinate in a virtual three-dimensional coordinate system based on the first pan-tilt angle; when it is necessary to mark the site element in a video frame of the second real-time image, map the third coordinate into a fourth coordinate based on the third coordinate and the second pan-tilt angle at the time when the video frame is located, and the fourth coordinate is the position coordinate of the site element in the video frame. The analysis terminal 23 is further adapted to perform displacement and / or distortion processing on the first coordinate based on the position difference and / or the field of view angle difference between the camera 211 of the shooting terminal and the camera 221 of the analysis terminal.

[0100] Figure 4 It is an architecture diagram of a tracking shooting system 40 according to an embodiment of the present application. As Figure 4 shown:

[0101] The tracking shooting system 40 includes a control terminal 31 and a pan-tilt 22. The control terminal 31 is communicatively connected to the pan-tilt 22, wherein:

[0102] The control terminal 31 is adapted to perform real-time control on the angle of the pan-tilt 22, obtain and display a real-time image of the shooting terminal 21; the control terminal 31 is adapted to respond to a user's click on a site boundary point on the real-time image, where the boundary point at least includes each folding point of the boundary of the site, record the coordinates of each boundary point clicked by the user and the pan-tilt angle at the time of clicking, and send them to the analysis terminal 23.

[0103] The pan-tilt 22 includes:

[0104] The shooting terminal 21 is adapted to obtain the first real-time image;

[0105] The analysis terminal 23 is adapted to execute the tracking shooting method based on site boundary filtering as described in the embodiments of the present application, obtain a tracking shooting strategy, and send it to the moving terminal 24;

[0106] The moving terminal 24 is adapted to adjust the angle of the pan-tilt 22 based on the tracking shooting strategy.

[0107] Wherein, after the moving terminal 24 adjusts the angle of the pan-tilt 22 based on the tracking shooting strategy, the angle of the shooting terminal 21 fixedly connected to the pan-tilt 22 will be adjusted accordingly, so as to achieve the effect of tracking shooting. The tracking shooting strategy is obtained by a multi-object tracking algorithm, and the multi-object tracking algorithm is a prior art known to those skilled in the art and has been introduced above, so it will not be elaborated here.

[0108] In addition, those skilled in the art should be aware that in the venue element marking system 30 and the tracking and shooting system 40 in the embodiments of the present application, the mechanism in the pan-tilt 22 is decomposed into a shooting end 21, an analysis end 23, and a motion end 24 according to functions. The shooting end 21 is a camera device that actually collects event videos in response to user requirements, and the analysis end 23 is an intelligent analysis device. Generally, the videos collected by the analysis end 23 are not stored, but in another embodiment, the real-time video of the analysis end 23 can be directly used as the actually collected event video. At this time, the shooting end 21 and the analysis end 23 are the same device component in terms of entity.

[0109] For the convenience and simplicity of description, the specific working processes of the venue element marking system 30 and the tracking and shooting system 40 in the embodiments of the present application can refer to the specific implementation manners of the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0110] Figure 5 The structural schematic diagram of the terminal device or server suitable for implementing the embodiments of the present application is shown.

[0111] As Figure 5 shown, the terminal device or server includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the terminal device or server are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0112] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that the computer program read from it can be installed into the storage section 508 as needed.

[0113] In particular, according to an embodiment of the present application, the above method flow steps can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present application are executed.

[0114] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the foregoing module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0116] The units or modules involved in the embodiments described in the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.

[0117] As another aspect, the present application also provides a computer-readable storage medium, which can be included in the electronic device described in the foregoing embodiments; or can exist separately without being assembled into the electronic device. The foregoing computer-readable storage medium stores one or more programs, and when the foregoing programs are executed by one or more processors, they implement the methods described in the present application.

[0118] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the application involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions in the present application.

Claims

1. A method for marking site elements, characterized in that, Including: In response to the user clicking on the site element in the first real-time image of the control terminal, record the first coordinate of the user's click and the first pan-tilt angle at the time of the click; Based on the first coordinate, obtain the second coordinate of the site element in the second real-time image of the analysis terminal; Convert the second coordinate into a first intermediate coordinate in a coordinate system with the center of the second real-time image as the origin; Add a third-dimensional coordinate to the first intermediate coordinate to obtain a first three-dimensional vector, and the modulus of the first three-dimensional vector is the camera focal length value of the analysis terminal; Based on the first pan-tilt angle, reverse-rotate the first three-dimensional vector back to the initial angle to obtain a third coordinate; When it is necessary to mark the site element in the video frame of the second real-time image: Rotate the third coordinate according to the second pan-tilt angle at the time of the video frame to obtain a second three-dimensional vector; Extract the first two dimensions in the second three-dimensional vector to obtain a second intermediate coordinate; Convert the second intermediate coordinate into a fourth coordinate in the default coordinate system of the second real-time image, and the fourth coordinate is the position coordinate of the site element in the video frame.

2. The site element marking method according to claim 1, characterized in that, The first real-time image is obtained by the shooting terminal and synchronously transmitted to the control terminal and displayed in the control terminal; The obtaining the second coordinate of the site element in the second real-time image of the analysis terminal based on the first coordinate includes: Based on the position difference and / or field-of-view angle difference between the camera of the shooting terminal and the camera of the analysis terminal, perform displacement and / or distortion processing on the first coordinate.

3. A tracking shooting method based on sideline filtering of a venue, characterized in that, Including: Mark a plurality of side points of the site in the real-time video frame, the plurality of side points at least include each folding point of the side of the site, the side points are marked by the method according to any one of claims 1-2, and each side point corresponds to a fifth coordinate; Enclose a site polygon in the real-time video frame based on all the fifth coordinates; Identify all human targets in the real-time video frame to obtain a first target set, and calculate the foot points of each human target in the first target set; Filter out all human targets in the first target set whose foot points are outside the site polygon to obtain a second target set; Track and shoot the human targets in the second target set.

4. A site element marking system, characterized in that, Including: A control terminal, which is suitable for performing real-time control of the pan-tilt angle, obtaining and displaying the first real-time image of the shooting terminal; The control terminal is suitable for responding to the user's click on the site element in the first real-time image, recording the first coordinate of the user's click and the first pan-tilt angle at the time of the click and sending them to the analysis terminal; A pan-tilt, including: A shooting terminal, which is suitable for obtaining the first real-time image; An analysis terminal, where the analysis terminal is adapted to obtain a second coordinate of the site element in a second real-time image at the analysis terminal based on the first coordinate; convert the second coordinate into a first intermediate coordinate in a coordinate system with the center of the second real-time image as the origin; add a third-dimensional coordinate to the first intermediate coordinate to obtain a first three-dimensional vector, and the modulus of the first three-dimensional vector is the focal length value of the camera at the analysis terminal; based on the first pan-tilt angle, rotate the first three-dimensional vector back to the initial angle in the reverse direction to obtain a third coordinate; when it is necessary to mark the site element in the video frame of the second real-time image: rotate the third coordinate according to the second pan-tilt angle at the time when the video frame is located to obtain a second three-dimensional vector; extract the first two dimensions in the second three-dimensional vector to obtain a second intermediate coordinate; convert the second intermediate coordinate into a fourth coordinate in the default coordinate system of the second real-time image, and the fourth coordinate is the position coordinate of the site element in the video frame.

5. A site element marking system as claimed in claim 4, characterized in that, The analysis terminal is further adapted to perform displacement and / or distortion processing on the first coordinate based on the position difference and / or field of view angle difference between the camera at the shooting terminal and the camera at the analysis terminal.

6. A tracking shooting system, characterized in that, Including: A control terminal, where the control terminal is adapted to perform real-time control on the pan-tilt angle, obtain and display the real-time image at the shooting terminal; the control terminal is adapted to respond to the user's click on the site boundary point on the real-time image, where the boundary point at least includes each folding point of the boundary of the site, record the coordinates of each boundary point clicked by the user and the pan-tilt angle at the time of clicking and send them to the analysis terminal; A pan-tilt, including: A shooting terminal, adapted to obtain a first real-time image; An analysis terminal, where the analysis terminal is adapted to execute a tracking shooting method based on site boundary filtering as described in claim 3, obtain a tracking shooting strategy, and send it to the moving terminal; A moving terminal, adapted to adjust the angle of the pan-tilt based on the tracking shooting strategy.

7. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method and device for filtering staffs in court

    CN116416648A

  • Video image real-time rotation method and system based on FPGA

    CN119722466A

  • Marking device for building decoration

    CN218238842U