Player feature extraction method, real-time tracking shooting method, system and storage medium
The method uses ball jersey number and facial feature extraction to improve athlete tracking accuracy and reduce costs in sports event video capture by employing CNN models for intelligent, real-time adjustments.
Patent Information
- Application Number
- CN202510437492.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
AI Technical Summary
The existing video tracking and shooting methods cannot achieve independent shooting decisions in dynamic scenes, it is difficult to accurately distinguish individual athletes from specific teams, and the hardware storage cost is high, making it difficult to meet the real-time needs of live events.
By obtaining the player's face feature vector and jersey number, combining the target detection model and classification model, players' feature extraction and real-time tracking and shooting are realized, and the shooting gimbal angle is automatically adjusted to track the target.
It realizes efficient single-player intelligent tracking and shooting, improves tracking accuracy in dynamic scenes, reduces shooting costs, and meets the real-time needs of live events.
Smart Images

Figure CN120318276A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image recognition and tracking shooting, and particularly to a method for extracting player features, a real-time tracking shooting method, a device, and a storage medium. Background Art
[0002] Tracking shooting technology is an important means for collecting images of sports events and is widely used in professional scenarios such as athlete motion capture and event live broadcast. However, there are still significant technical bottlenecks in the intelligent tracking and multi-angle collection of dynamic scenes in currently commonly used video tracking shooting methods. First, traditional tracking shooting pan-tilt systems mostly use fixed-angle shooting or manual control shooting, and cannot make autonomous shooting decisions in dynamic scenes, easily resulting in the lack of multi-dimensional image information. Second, tracking algorithms generally adopt the principle of giving priority to the largest target or feature recognition based on the color of human clothing, and can only track the overall group on the field, unable to accurately distinguish individual athletes of a specific team, lacking a facial feature extraction and multi-modal verification mechanism, and being extremely prone to target misjudgment or loss in scenarios where multiple people have similar clothing, resulting in insufficient dynamic tracking accuracy. In addition, existing automatic shooting methods for sports events generally adopt panoramic stitching technology, and construct a panoramic image through multi-wide-angle lens distortion correction and image stitching, with a relatively large hardware storage cost and being difficult to meet the real-time requirements of event live broadcast.
[0003] Therefore, there is an urgent need for a real-time intelligent tracking shooting method to meet the intelligent, low-cost, and high-efficiency shooting requirements of sports events. Summary of the Invention
[0004] According to an embodiment of the present disclosure, a method for extracting player features and an intelligent tracking shooting solution based on jersey information are provided, which can achieve efficient single-person intelligent tracking shooting and greatly reduce the difficulty of shooting specific personnel in sports events.
[0005] In a first aspect of the present disclosure, a method for extracting player features is provided. The method includes:
[0006] Obtain a picture containing a human target wearing a jersey, where the picture contains the front of one or more human targets;
[0007] Obtain the face feature vector of the human target;
[0008] Input the picture into a first target detection model to identify the number area in the picture and crop it to obtain a picture of the jersey number;
[0009] Input the picture of the jersey number into a number classification model to obtain the jersey number of the human target.
[0010] In a possible implementation, there are multiple front-facing images of human targets, so multiple face feature vectors of the human targets are stored in a feature vector queue.
[0011] In a second aspect of the present disclosure, a real-time tracking and shooting method based on player features is provided. The method includes:
[0012] S101: Obtain a first human body image containing the target to be tracked;
[0013] S102: Execute the method of the first aspect of the present disclosure to obtain one or more first face feature vectors and a first jersey number of the target to be tracked;
[0014] S103: Input the real-time shooting image into a second target detection model to obtain all second human body images containing human targets in the real-time shooting image;
[0015] S104: Execute the method of the first aspect of the present disclosure to obtain the second jersey numbers corresponding to all the second human body images;
[0016] S105: Compare the first jersey number with all the second jersey numbers in sequence:
[0017] If the comparison is successful, extract the face feature vectors of the second human body image corresponding to the successfully compared second jersey number, and compare the similarity with one or more first face feature vectors:
[0018] If they are similar, perform real-time tracking and shooting on the second human target in the second human body image corresponding to the successfully matched second jersey number.
[0019] Preferably, repeat S103 - S105 for subsequent frames of the real-time shooting image until the real-time shooting image is interrupted.
[0020] In a possible implementation, if the first jersey number fails to match all the second jersey numbers, continue to execute S103 - S105 for the next frame of the real-time shooting image.
[0021] In a possible implementation, after the first jersey number matches one of the second jersey numbers, first determine whether the second human body image corresponding to the second jersey number overlaps with other human body images; if not, then perform the extraction and similarity comparison of one or more first face feature vectors.
[0022] In another possible implementation, the real-time tracking and shooting of the second human target can be performed in the following manner:
[0023] Predict the movement of the second human target through a target tracking algorithm to obtain the next predicted position of the second human target;
[0024] Automatically adjust the angle of the shooting gimbal according to the next predicted position, and continuously obtain a second real-time shooting image containing the second human target.
[0025] In a third aspect of the present disclosure, a method for extracting jersey information is provided. The method includes:
[0026] Obtain an image containing a human target wearing a jersey;
[0027] Input the image into a first target detection model, identify the number area in the image and crop it to obtain a jersey number image;
[0028] Detect the background color in the jersey number image to obtain the jersey color of the human target;
[0029] Input the jersey number image into a number classification model to obtain the jersey number of the human target.
[0030] In a possible implementation, the background color of the jersey number image is determined by the color at the edge of the jersey number image.
[0031] In a fourth aspect of the present disclosure, a real-time tracking shooting method based on jersey information is provided. The method includes:
[0032] S201: Obtain a first human image containing the target to be tracked;
[0033] S202: Execute the method according to the third aspect of the present disclosure to obtain the first jersey color and the first jersey number of the target to be tracked;
[0034] S203: Input the real-time shooting image into a second target detection model to obtain all second human images containing human targets in the real-time shooting image;
[0035] S204: Execute the method according to the third aspect of the present disclosure to obtain the second jersey color and the second jersey number of all second human images, and each second human image corresponds to a second jersey number and a second jersey color;
[0036] S205: Compare the first jersey number with all the second jersey numbers in sequence:
[0037] If the comparison is successful, determine whether the first jersey color is the same as the second jersey color corresponding to the successfully matched second jersey number;
[0038] If they are the same, perform real-time tracking shooting on the second human target in the second human image corresponding to the successfully matched second jersey number.
[0039] Preferably, repeat the execution of S203 - S205 until the real-time shooting image is interrupted.
[0040] In a possible implementation, when determining whether the first jersey color is the same as the second jersey color, the cosine value of the vector angle between the first jersey color and the second jersey color corresponding to the successfully matched second jersey number can be calculated; when the cosine value is greater than a preset value, it is determined that the first jersey color is the same as the second jersey color corresponding to the successfully matched second jersey number. The preset value can be set to 0.95.
[0041] In another possible implementation, the following method can be used to perform real-time tracking and shooting of the second human target:
[0042] Predict the movement of the second human target through a target tracking algorithm to obtain the next predicted position of the second human target;
[0043] According to the next predicted position, automatically adjust the angle of the shooting pan-tilt head to continuously obtain a second real-time shooting image containing the second human target.
[0044] In the fifth aspect of the present disclosure, a real-time tracking and shooting system is provided, including:
[0045] A shooting module for collecting real-time shooting images;
[0046] An extraction module for performing the player feature extraction method of the first aspect of the present disclosure or the jersey information extraction method of the third aspect;
[0047] A comparison module for performing the real-time tracking and shooting method of the second aspect or the fourth aspect of the present disclosure;
[0048] A movement module for receiving a movement instruction sent by the comparison module and changing the angle of the shooting module according to the instruction.
[0049] In the sixth aspect of the present disclosure, an electronic device is provided. The electronic device includes: a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the methods of the first to fourth aspects of the present disclosure are implemented.
[0050] In the seventh aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the methods of the first to fourth aspects of the present disclosure are implemented.
[0051] The real-time tracking and shooting method provided by the embodiments of the present disclosure determines the position of the person to be tracked and shot and makes a prediction by performing a two-step comparison between the pre-extracted feature information of the person to be tracked and the feature information of the person in the real-time shooting image, and adjusts the pan-tilt head angle based on the prediction result for continuous tracking, realizing efficient single-person intelligent tracking and shooting, improving the difficulty and efficiency of event shooting, and reducing the shooting cost.
[0052] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0054] Figure 1 is a flowchart of a player feature extraction method according to an embodiment of the present disclosure;
[0055] Figure 2 is a schematic diagram of a player feature extraction method according to an embodiment of the present disclosure;
[0056] Figure 3 is a flowchart of a real-time tracking shooting method according to an embodiment of the present disclosure;
[0057] Figure 4 is a flowchart of a jersey information extraction method according to an embodiment of the present disclosure;
[0058] Figure 5 is a schematic diagram of a jersey information extraction method according to an embodiment of the present disclosure;
[0059] Figure 6 is a flowchart of a real-time tracking shooting method according to another embodiment of the present disclosure;
[0060] Figure 7 is an architecture diagram of a real-time tracking shooting system 700 according to an embodiment of the present disclosure;
[0061] Figure 8 is a schematic structural diagram of a terminal device or a server suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0063] In addition, the term "and / or" in the present disclosure is merely a relational term describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally represents an "or" relationship between the associated objects before and after.
[0064] Embodiment 1:
[0065] Figure 1 It is a flowchart of a player feature extraction method according to an embodiment of the present disclosure. Refer to Figure 1 The method includes:
[0066] First, obtain a picture of a human target containing a jersey. Since face feature vectors need to be extracted, it must contain one or more pictures of the front of the human target. Regarding the source of the picture, the present disclosure does not make any limitations. The picture can be from the user's active upload, or from the user's specified actions such as clicking or bounding in the captured picture. Generally, there are jersey numbers on both the front and back of the jersey. Therefore, preferably, it is required that the user specify at least two pictures of the human target, one of the front and one of the back.
[0067] Second, obtain the face feature vector of the human target. Among them, the face feature vector is extracted from the picture of the front of the human target.
[0068] There are various technical solutions for the specific process of face feature extraction to obtain feature vectors from the face picture. The present disclosure does not limit the specific extraction method. Generally, open-source CNN (Convolutional Neural Network) algorithms such as the FaceNet algorithm, ResNet / DenseNet algorithm, and ArcFace algorithm can be used to extract the feature information of the face to obtain a multi-feature multi-dimensional vector.
[0069] In one embodiment, when the user specifies multiple pictures of the front of the human target, the method will extract the face features of each picture of the front of the human target to obtain multiple face feature vectors and store them in a feature vector queue.
[0070] Third, input the picture of the human target containing the jersey into the first target detection model to identify and crop the number area in the picture to obtain a picture of the jersey number, and input the picture of the jersey number into the number classification model to obtain the jersey number of the human target. Among them, the number classification model is used to identify and classify the numbers in the picture of the jersey number.
[0071] In the prior art, the recognition of jersey numbers is completed through OCR technology. However, the effect of OCR technology in dynamic images is not good, and when the jersey is wrinkled, accurate numbers cannot be recognized. Considering that in the use scenario of jersey number recognition, the possible text content may only be the ten digits from 0 to 9, using OCR technology not only wastes computing power but also may recognize results other than numbers, thus further reducing the accuracy. Therefore, in this embodiment, the process of jersey number recognition is divided into two steps to achieve number region recognition and number classification successively, which can improve the accuracy of jersey number recognition.
[0072] In the present disclosure, there is no limitation on the specific selection of the first object detection model and the number classification model for number region recognition: The YOLOv8 (You Only Look Once version 8) model can be used to integrally achieve number region recognition and number classification; alternatively, a combination of a recognition model and a classification model can be used. For example, the first object detection model can adopt a convolutional neural network, ViT (Vision Transformer), etc.; the classification model can adopt large models such as MobileNet, EfficientNet-Lite, or traditional machine learning models such as traditional vector machines and random forests.
[0073] The training method of the model is a commonly used technical means for those skilled in the art, and the present disclosure does not limit the specific model training method. In one embodiment, the first object detection model and the classification model can be trained through the following steps: Collect a large number of pictures containing jersey numbers under different teams, different angles, and different lighting conditions, label the number regions and numbers in each picture (0 - 99 or numbers within a specific range), and use methods such as rotation, scaling, and brightness adjustment to increase data diversity; use pre-trained models such as ResNet and EfficientNet for fine-tuning; set the output layer to the number of number categories. Then, identify the number of digits in the first stage and identify each digit separately in the second stage.
[0074] Figure 2 It is a schematic diagram of the player feature extraction method according to an embodiment of the present disclosure. In this embodiment, the panoramic view captured in the picture includes three players and several spectators.
[0075] First, a player is recognized through the second object detection model, framed with a bounding box, and preliminarily cropped based on the bounding box.
[0076] In one embodiment, the second object detection model directly calls the YOLO (You Only Look Once, a commonly used object detection model) model to implement.
[0077] Secondly, the number area in the picture is identified through the first target detection model, that is, Figure 2 the small picture containing the number 17 in Figure 2 is used as the picture of the jersey number;
[0078] Thirdly, the picture of the jersey number is input into the classification model, and the jersey number is identified as 17.
[0079] Meanwhile, the face information in the picture is extracted through the face feature extraction algorithm to obtain the face feature vector.
[0080] It should be emphasized that although the descriptions such as "secondly", "thirdly", "meanwhile" are used in the description of the embodiments above, in different embodiments, after obtaining the picture containing the human target wearing the jersey, the two steps of face feature vector extraction and jersey number recognition can be carried out simultaneously, or the jersey number recognition can be carried out first and then the face feature vector extraction. The present disclosure does not limit the sequence of these two steps.
[0081] Figure 3 is a flowchart of the real-time tracking shooting method according to the embodiment of the present disclosure. Refer to Figure 3 and the method includes:
[0082] S101: Obtain the first human body picture containing the target to be tracked.
[0083] Among them, in order to ensure the recognition quality and accuracy, the first human body picture in S101 must contain at least one front picture of the target to be tracked. For the acquisition method of the first human body picture, the present disclosure does not make any limitation. It can be actively uploaded by the user, or it can be a partial picture specified by the user through methods such as point selection and frame selection in the real-time shooting picture.
[0084] S102: Obtain one or more first face feature vectors and the first jersey number of the target to be tracked.
[0085] In S102, the method for extracting the player features in the foregoing Figure 1 embodiment is used to obtain one or more first face feature vectors and the first jersey number of the target to be tracked, and details are not described here again. In one embodiment, when there are multiple first face feature vectors, the multiple first face feature vectors are stored in the first feature vector queue for subsequent vector similarity comparison.
[0086] S101 and S102 are the preparation stages of the real-time tracking and shooting method based on jersey information in this embodiment, which are used to determine the information of the target to be tracked. Since in a game, usually both competing sides have independent numbering rules, it is very likely that there are two players with the same jersey number on the field. If tracking is only based on the jersey number, it is easy to have the situation of incorrect target recognition. Therefore, in order to distinguish the situation of the same jersey number, the recognition of players' facial features is introduced in the embodiments of the present disclosure.
[0087] S103: Input the real-time shooting image into the second target detection model to obtain all second human images containing human targets in the real-time shooting image.
[0088] Among them, the second target detection model is used to identify and locate the position information of human targets in the image in the image, which is a common target detection algorithm and can be implemented using algorithms such as YOLO. In one embodiment, the target detection model includes an input layer, a Backbone network layer, a Neck feature fusion layer, and a prediction layer. Among them, the input layer is used to preprocess the input image to make it meet the requirements of model training and prediction. The Backbone network layer is used for feature extraction, and it uses the deeply optimized C2f module as the basic unit, improving the performance while reducing the network size. The Neck feature fusion layer is used to fuse the feature maps from different stages of the Backbone network layer to enhance the feature representation ability. The prediction layer includes an SPPF (Spatial Pyramid Pooling Fast) module, a PAA (Probabilistic Anchor Assignment) module, a PAN (Path Aggregation Network) module, and a Head module. The SPPF module is used to splice feature maps of different scales together to improve the model's detection ability for image targets of different sizes. The PAA module is used to intelligently allocate anchor boxes to optimize the selection of positive and negative samples and improve the training effect of the model. The PAN module is used to perform path aggregation on features at different levels to enhance the expression ability of the feature map through bottom-up and top-down paths. The Head module is used for the final target detection prediction, marking all human images in the current video frame in the form of target boxes, that is, all second human images.
[0089] S104: Obtain the second jersey number of each second human image.
[0090] Among them, using the player feature extraction method disclosed in the embodiments of the present disclosure as Figure 1 shown, extract the second jersey number from each second human image identified in step S103. Obviously, each second human image corresponds one-to-one to the second jersey number identified from it.
[0091] S105: Compare each second jersey number with the first jersey number:
[0092] If the comparison is successful, it indicates that there is a jersey number in it that is the same as the jersey number of the target to be tracked. Therefore, extract the face feature vector of the second human body image corresponding to the successfully compared second jersey number, and compare its similarity with the first face feature vector, or compare its similarity with multiple first face feature vectors in the first feature vector queue respectively. If the comparison fails, it means that there is no jersey number in the current real-time video image that is the same as the jersey number of the target to be tracked, that is, the current frame matching fails. Repeat the steps of S103 - S105 for the subsequent frames of the real-time video image. In different embodiments, the subsequent frame can be the next real-time frame or the next frame after frame extraction.
[0093] Preferably, before performing the similarity comparison, it can be determined whether the second human body image overlaps with other human body images in the real-time captured image. When multiple target recognition frames overlap, it is impossible to determine whether the corresponding relationship between the current second human body image and its corresponding second jersey number is correct, and there is likely to be a situation of misidentification. Therefore, the steps of feature vector extraction and similarity comparison can be performed only when it is determined that the second human body image does not overlap with other human body images, so as to save computing power and reduce the possibility of tracking errors.
[0094] The extraction of the feature vector is the same as the feature vector extraction step in the player feature extraction method shown in this disclosure Figure 1 and will not be elaborated here.
[0095] In one embodiment, the vector Euclidean distance is used to compare the similarity of the feature vectors. The Euclidean distance is the straight-line distance between two points in an n-dimensional space, and is the arithmetic square root of the sum of the squares of the differences between the corresponding dimension values of two identical feature vectors. Its calculation formula is:
[0096]
[0097] where A and B are two n-dimensional vectors, A i and B i respectively represent the coordinate values of the vector in the i-th dimension.
[0098] In one embodiment, when the Euclidean distance between two face feature vectors is less than 0.6, it can be considered that these two face feature vectors are similar.
[0099] In another embodiment, the cosine value of the vector angle is used to compare the similarity of the feature vectors. The cosine value of the vector angle is the ratio of the dot product of two feature vectors to the product of their magnitudes, i.e., r(f,g)=(f·g) / (||f||·||g||). The closer the value of the cosine of the angle is to 1, the more similar the two feature vectors are. In one embodiment, when the cosine value of the angle is greater than 0.85, it can be considered that the two face feature vectors are similar.
[0100] When the face feature vector of the second human body image corresponding to the successfully matched second jersey number is similar to the first face feature vector, or similar to at least one of the first face feature vectors in the first feature vector queue, it is determined that the second human body target in the second human body image at this time is the target to be tracked.
[0101] After determining the target to be tracked, the movement of the second human body target is predicted through a target tracking algorithm, and the next predicted position of the second human body target is obtained; according to the next predicted position, the angle of the shooting pan-tilt is automatically adjusted to continuously obtain a second real-time shooting image containing the second human body target.
[0102] Among them, the specific selection of the target tracking algorithm is not specifically limited in the present invention. Generally, mature target tracking algorithms such as the MHT algorithm, the MOT algorithm, and the SORT algorithm can all implement the technical solution of the present invention. In a possible implementation manner, the SORT algorithm is used to implement target tracking. SORT (Simple Online and Realtime Tracking) is an efficient multi-target tracking algorithm based on motion modeling and data association. Its core idea is to achieve continuous tracking of cross-frame target identities by fusing the target motion prediction in the time dimension and the detection box position information in the space dimension. First, the boundary box position of the first human body target in the next frame is predicted through Kalman filtering according to the historical trajectory of the first human body target. Then, the Hungarian algorithm is introduced to establish the association relationship between the current frame containing the first human body target and the previous frame containing the first human body target, and the first human body target in the current frame and the first human body target in the previous frame are matched by minimizing the association cost, ensuring that the boundary box of each first human body target is associated with at most one predicted boundary box, and then the next predicted position of the first human body target is obtained.
[0103] In a possible implementation manner, the initial position of the first human body target is (x1,y 1) and the next predicted position of the first human body target predicted by the multi-target tracking algorithm is (x2,y 2) then the yaw angle that the shooting pan-tilt needs to adjust is:
[0104]
[0105] Among them, D is the target level parameter. The pitch angle that the shooting pan-tilt needs to adjust is:
[0106]
[0107] In addition, the rotation speed of the shooting pan-tilt can be controlled by a PID controller, and a smooth rotation speed command is generated according to the pitch angle difference or yaw angle difference to avoid the jitter of the shooting pan-tilt and smooth the rotation speed The calculation formula is as follows:
[0108]
[0109] Among them, is the pitch angle difference or yaw angle difference, is a hyperparameter with a set value of 0.8, is a hyperparameter with a set value of 0.2, is a hyperparameter with a set value of 0.1.
[0110] In one embodiment, since the jersey number recognition may be incorrect and the target may be lost during the tracking process, therefore, regardless of whether the target to be tracked is determined in step S105, preferably, the steps of S103-S105 are continued to be executed for the subsequent frames of the live video. In different embodiments, the subsequent frame may be the next real-time frame or the next frame after frame extraction.
[0111] In this embodiment, the intelligent tracking and shooting of the first human target is realized through the multi-target tracking algorithm, which improves the robustness in the complex event shooting scenario.
[0112] In the foregoing embodiment 1 of the present disclosure, it is intended to solve the problem that multiple people may correspond to the same number in simple jersey number recognition, and the secondary matching of player features is introduced, so that the target to be tracked can be accurately recognized. The player feature matching in embodiment 1 uses the feature vector of the face, and there may also be a situation where the face information cannot be collected for matching due to the player's back to the screen. To solve this problem, the following embodiment 2 of the present disclosure is introduced, which uses the jersey color as the player feature.
[0113] Embodiment 2:
[0114] Figure 4 is a flowchart of the jersey information extraction method according to the embodiment of the present disclosure. Refer to Figure 4 This method includes:
[0115] First, obtain an image of a human target wearing a jersey. Regarding the source of the image, the present disclosure places no restrictions. The image can be sourced from a user's active upload or from a user's specified actions such as clicking or bounding box selection within the captured image. Generally, the jersey number is present on both the front and back of the jersey. Therefore, preferably, the user is required to specify two images of the human target, one for the front and one for the back.
[0116] Secondly, input the image of the human target wearing the jersey into the first target detection model to identify the number area in the image and perform cropping to obtain the jersey number image.
[0117] Thirdly, determine the jersey color by identifying the background color of the jersey number image. In a ball game scenario, the main part of the jersey is usually a solid color, and the jersey number is black or white. Therefore, after obtaining the cropped jersey number image, the image should be a pattern with the background color being the jersey color and black or white numbers in the middle. Therefore, preferably, the edge color of the cropped jersey number image is the jersey color. The jersey color is stored in RGB format.
[0118] In one embodiment, the background color of the jersey number image is identified by calling the built-in color picker function of the system (getColorAtPosition in the Android system and Core Graphics in the iOS system). The specific identification point can use the edge point at the lower left corner of the image, that is, the point with both x / y axis pixel coordinate values being extremely small.
[0119] At the same time, input the jersey number image into the number classification model to obtain the jersey number of the human target.
[0120] Among them, the number classification model is used to identify and classify the numbers in the jersey number image.
[0121] Usually, the recognition of jersey numbers is completed through OCR technology. However, the effect of OCR technology in dynamic images is not good, and when the jersey is wrinkled, accurate numbers cannot be recognized either. Considering that in the usage scenario of jersey number recognition, the possible text content may only be the ten digits from 0 to 9, using OCR technology not only wastes computing power but may also recognize results other than numbers, further reducing the accuracy. Therefore, in this embodiment, the process of jersey number recognition is divided into two steps to successively achieve number area recognition and number classification, which can improve the accuracy of jersey number recognition.
[0122] In the present disclosure, there are no restrictions on the specific selection of the first target detection model, number classification model, and the model training method for number area recognition, and they are consistent with the first target detection model, number classification model, and the model training method mentioned in Embodiment 1, and will not be elaborated here.
[0123] Figure 5 Schematic diagram of a method for extracting jersey information according to an embodiment of the present disclosure. In this embodiment, the panoramic view captured in the picture includes three players and several spectators.
[0124] First, a player is identified through a second object detection model, framed with a bounding box, and preliminarily cropped based on the bounding box.
[0125] In one embodiment, the second object detection model is directly implemented by calling the YOLO (You Only Look Once, a commonly used object detection model) model.
[0126] Secondly, the number area in the picture is identified through a first object detection model, that is, Figure 5 the small picture containing the number 17 is used as the picture of the jersey number.
[0127] Thirdly, the edge color of the picture of the jersey number is extracted and identified as black.
[0128] At the same time, the picture of the jersey number is input into a classification model, and the jersey number is identified as 17.
[0129] It should be emphasized that although the description of "at the same time" is used in the foregoing description of the embodiment, in different embodiments, after obtaining the picture of the jersey number, the color of the jersey can be identified first, and then the jersey number can be identified; or the jersey number can be identified first, and then the color of the jersey can be identified. The present disclosure does not limit whether these two steps occur successively or their order.
[0130] Figure 6 Flowchart of a real-time tracking shooting method according to another embodiment of the present disclosure. Refer to Figure 6 , the method includes:
[0131] S201: Obtain a first human body picture containing the target to be tracked.
[0132] Among them, in order to ensure the recognition quality and accuracy, preferably, the first human body picture is a front and / or back photo of the target to be tracked containing the jersey number. For the acquisition method of the first human body picture, the present disclosure does not make any restrictions, and it can be actively uploaded by the user, or it can be a local picture specified by the user through methods such as point selection and box selection in the real-time shooting picture.
[0133] S202: Obtain the first jersey color and the first jersey number of the target to be tracked.
[0134] In S202, the method for obtaining the first jersey color and the first jersey number of the target to be tracked is the jersey information extraction method in the foregoing Figure 4 embodiment, which will not be elaborated here.
[0135] S201 and S202 are the preparation stages of the real-time tracking and shooting method based on jersey information in this embodiment, which are used to determine the information of the target to be tracked. Since in a game, usually both competing sides have independent numbering rules, it is very likely that there are two players with the same jersey number on the field. If tracking is only based on the jersey number, it is easy to have the situation of incorrect target recognition. Considering that the situation of the same jersey number can only occur in two opposing teams, and the jersey colors of the opposing teams must be different. Therefore, in order to distinguish the situation of the same jersey number, the recognition of jersey colors is introduced in the embodiments of the present disclosure.
[0136] S203: Input the real-time shooting image into the second target detection model to obtain all second human images containing human targets in the real-time shooting image.
[0137] Among them, the second target detection model is used to identify and locate the position information of human targets in the image from the image, which is a common target detection algorithm and can be implemented using algorithms such as YOLO. In one embodiment, the target detection model includes an input layer, a Backbone network layer, a Neck feature fusion layer, and a prediction layer. Among them, the input layer is used to preprocess the input image to meet the requirements of model training and prediction. The Backbone network layer is used for feature extraction, and it uses the deeply optimized C2f module as the basic unit, improving the performance while reducing the network size. The Neck feature fusion layer is used to fuse the feature maps from different stages of the Backbone network layer to enhance the feature representation ability. The prediction layer includes an SPPF (Spatial Pyramid Pooling Fast) module, a PAA (Probabilistic Anchor Assignment) module, a PAN (Path Aggregation Network) module, and a Head module. The SPPF module is used to splice feature maps of different scales together to improve the model's detection ability for image targets of different sizes. The PAA module is used to intelligently allocate anchor boxes to optimize the selection of positive and negative samples and improve the training effect of the model. The PAN module is used to perform path aggregation on features at different levels to enhance the expression ability of the feature map through bottom-up and top-down paths. The Head module is used for the final target detection prediction, marking all human images in the current video frame in the form of target boxes, that is, all second human images.
[0138] S204: Obtain the second jersey number and the second jersey color of each second human image.
[0139] Among them, using the jersey information recognition method disclosed in the embodiments of the present disclosure, the second jersey number and the second jersey color are extracted from each second human body image recognized in step S203. Obviously, each second human body image corresponds one-to-one to the second jersey number and the second jersey color recognized therefrom.
[0140] S205: Compare the second jersey number and the second jersey color of each second human body image with the first jersey number and the first jersey color recognized in S202. If the comparison is successful, it indicates that the human body in this second human body image is the target to be tracked set in step S201, and call the target tracking algorithm to track the target to be tracked.
[0141] Among them, first compare with the numbers, that is, compare the second jersey number of each second human body image with the first jersey number. Since not all athletes may be included in the real-time video image, even if the result of the number comparison is that only one jersey number comparison is successful, it is still impossible to determine that the athlete corresponding to this jersey number must be the target to be tracked. Therefore, in the solution disclosed in this embodiment, instead of waiting until all jersey numbers are compared and then comparing the jersey colors, when a certain second jersey number comparison is successful, immediately compare the jersey colors. If the jersey color comparison is also successful, it is determined that the athlete corresponding to this jersey number is the target to be tracked. If the jersey color comparison fails, it indicates that the athlete corresponding to this jersey number is not the target to be tracked, and continue to compare the remaining second jersey numbers. If all second jersey numbers are compared and failed, it indicates that there is no target to be tracked in the current image.
[0142] Due to reasons such as lighting or shooting angle, it cannot be guaranteed that the second jersey color in the real-time video image is exactly the same as the first jersey color, and there may be a certain color difference. Therefore, the matching of the jersey colors does not require the RGB three channels to have exactly the same values, and only the error needs to be less than a preset value. In one embodiment, the similarity of hue and saturation is measured by calculating the cosine of the angle between color vectors. The cosine of the angle is the ratio of the dot product of two color vectors to the product of the moduli, that is, r(f,g)=(f·g) / (||f||·||g||). The closer the value of the cosine of the angle is to 1, the closer the hues and saturations of the two colors are. For example, if Color1=(216,8,24) and Color2=(240,16,16), then f·g = 216*240 + 8*16 + 24*16 = 52352, ||f||·||g||≈52,475.85, and the cosine of the angle is 0.999, indicating that their hues are highly similar.
[0143] When the cosine of the angle is greater than a preset value, it is determined that the two colors are close, and the jersey color matching is considered successful. In a possible embodiment, the preset value can be 0.95.
[0144] In one embodiment, since there may be errors in the identification of the jersey number and the target may be lost during the tracking process, preferably, steps S203 - S205 are continued to be executed for the subsequent frames of the live video regardless of whether the target to be tracked is determined in step S205. In different embodiments, the subsequent frame may be the next real-time frame or the next frame after frame extraction.
[0145] After determining the target to be tracked, the movement of the second human target is predicted through a target tracking algorithm to obtain the next predicted position of the second human target; according to the next predicted position, the angle of the shooting pan-tilt is automatically adjusted to continuously obtain a second live shooting image containing the second human target.
[0146] Among them, the target tracking algorithm is the same as the target tracking algorithm in Embodiment 1 of the present disclosure and will not be elaborated here.
[0147] It should be noted that Embodiment 2 of the present disclosure can be independently used as a jersey information extraction and real-time tracking shooting scheme based on jersey information; it can also be used as a supplement to Embodiment 1: when the face feature vector of the second human image cannot be successfully obtained in Embodiment 1, the jersey color matching scheme of Embodiment 2 is used as a supplement; similarly, Embodiment 1 can also be used as a supplement to Embodiment 2: when the jersey color cannot be extracted in Embodiment 2, or the jersey colors of the two competing sides are too close, the face feature vector comparison scheme of Embodiment 1 is used as a supplement.
[0148] Figure 7 FIG. 700 is an architecture diagram of a real-time tracking shooting system according to an embodiment of the present disclosure.
[0149] In the real-time tracking shooting system, a shooting module 701, an extraction module 702, a comparison module 703, and a movement module 704 are included. Among them:
[0150] The shooting module 701 is suitable for collecting real-time video images and can generally be implemented as a camera or an intelligent electronic device with a video recording function;
[0151] The extraction module 702 is suitable for extracting the human image in the real-time video image and identifying the jersey number and player features of the human image; the player features can be the face feature vector in Embodiment 1 of the present disclosure or the jersey color in Embodiment 2 of the present disclosure;
[0152] The comparison module 703 is suitable for comparing the jersey number and player features identified by the extraction module 702, identifying the target to be tracked, and sending a movement instruction to the movement module 704 based on the tracking algorithm;
[0153] In one embodiment, both the extraction module 702 and the comparison module 703 are located in the smart device.
[0154] The motion module 704 is adapted to receive the motion instruction sent by the comparison module 703 and change the angle of the shooting module 701 according to the motion instruction. In one embodiment, the motion module 704 is implemented as a gimbal.
[0155] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0156] Figure 8 The structural schematic diagram of the terminal device or server suitable for implementing the embodiments of the present disclosure is shown.
[0157] As Figure 8 shown, the terminal device or server includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the terminal device or server are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0158] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as required. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as required so that the computer program read from it can be installed into the storage section 808 as required.
[0159] In particular, according to embodiments of the present disclosure, the above method flow steps may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a machine-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through a communication section 809, and / or installed from a removable medium 811. When the computer program is executed by a central processing unit (CPU) 801, the above functions defined in the system of the present disclosure are performed.
[0160] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the foregoing module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0162] The units or modules involved in the embodiments described in the present disclosure can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases.
[0163] As another aspect, the present disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the foregoing embodiments; or may exist separately without being assembled into the electronic device. The foregoing computer-readable storage medium stores one or more programs, and when the foregoing programs are executed by one or more processors, they implement the methods described in the present disclosure.
[0164] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the application involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing application concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions applied in the present disclosure.
Claims
1. A method for extracting player characteristics, characterized in that, Comprising: Obtain a picture of a human target wearing a jersey, wherein the picture of the human target includes a picture of one or more fronts of the human target; Obtain the face feature vector of the human target; Input the picture into a first target detection model, identify the number area in the picture and crop it to obtain a picture of the jersey number; Input the picture of the jersey number into a number classification model to obtain the jersey number of the human target.
2. The player feature extraction method according to claim 1, wherein When there are pictures of multiple fronts of the human target, after the step of obtaining the face feature vector of the human target, it further includes: Store the multiple face feature vectors corresponding to the pictures of the multiple fronts of the human target in a feature vector queue.
3. A real-time tracking and shooting method based on jersey information, characterized in that, Comprising: S101: Obtain a first human picture including a target to be tracked; S102: Execute the method according to claim 1 or 2 to obtain one or more first face feature vectors and a first jersey number of the target to be tracked; S103: Input the real-time captured picture into a second target detection model to obtain all second human pictures including human targets in the real-time captured picture; S104: Execute the method according to claim 1 to obtain the second jersey numbers corresponding to all the second human pictures; S105: Compare the first jersey number with all the second jersey numbers in sequence: If the comparison is successful, extract the face feature vector of the second human picture corresponding to the second jersey number and compare it with the one or more first face feature vectors for similarity; If they are similar, perform real-time tracking and shooting on the second human target in the second human picture corresponding to the successfully matched second jersey number.
4. The real-time tracking and shooting method based on jersey information according to claim 3, wherein It further includes: Repeat S103 - S105 for subsequent frames of the real-time captured picture until the real-time captured picture is interrupted.
5. The real-time tracking and shooting method according to claim 3, characterized in that The step of comparing the first jersey number with all the second jersey numbers in sequence further includes: If the comparison fails, continue to execute S103 - S105 for the next frame of the real-time captured picture.
6. The real-time tracking and shooting method according to claim 3, wherein The step of if the comparison is successful, extract the face feature vector of the second human picture corresponding to the second jersey number and compare it with the one or more first face feature vectors for similarity includes: If the comparison is successful, determine whether the second human picture corresponding to the second jersey number overlaps with other human pictures; If there is no overlap, extract the face feature vector of the second human picture corresponding to the second jersey number and compare it with the one or more first face feature vectors for similarity.
7. The real-time tracking and shooting method according to any one of claims 3-6, characterized in that, The step of performing real-time tracking and shooting on the second human target in the second human picture includes; Predict the movement of the second human target through a target tracking algorithm to obtain the next predicted position of the second human target; According to the next predicted position, automatically adjust the angle of the shooting pan-tilt to continuously obtain a second real-time captured picture including the second human target.
8. A real-time tracking and shooting system, characterized in that, Comprising: A shooting module for collecting real-time captured pictures; An extraction module for executing the player feature extraction method according to any one of claims 1 or 2; A comparison module, configured to execute the real-time tracking shooting method according to any one of claims 3-7; A motion module, configured to receive a motion instruction sent by the comparison module and change the angle of the shooting module according to the instruction.
9. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that When the processor executes the computer program, the method according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Sportsman photograph sorting method based on number plate identification and face identification
CN107609108A
Method and device, terminal device and storage medium for identifying targets
CN108875667A
A player identification method based on deep learning
CN109766768A
Method, device and system for identifying identity of moving person in real time
CN110852161A
Object close-up switching method and device in specific scene and electronic equipment
CN115529468A