Video clip classification method and device for motion event, equipment and storage medium

By identifying the action execution object and target sphere positions in the video screen, key action recognition and motion event determination are performed, the problem of low efficiency in traditional video clip processing is solved, and more efficient video clip classification processing is achieved.

CN120014501APending Publication Date: 2025-05-16ARASHI VISION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311537199.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

During the traditional video clip processing, after determining the highlight moment of the motion event triggered by the sphere, the subsequent processing efficiency is low and the amount of data is large, resulting in a long processing time.

Method used

By identifying the action execution object and the position of the target sphere in the video screen, key action recognition is performed, motion events are determined, and video clips are classified and processed according to the key action execution object.

Benefits of technology

The efficiency of key action recognition is improved, the amount of data involved in video processing is reduced, and the video clips corresponding to motion events need to be classified, thereby improving the efficiency of classification processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014501A_ABST
    Figure CN120014501A_ABST
Patent Text Reader

Abstract

The invention relates to a video clip classification method and device of a motion event, a handheld holder, computer equipment, a storage medium and a computer program product. The method comprises the following steps: identifying an action execution object in a video picture, and detecting the position of a target sphere in the video picture; performing key action recognition on the action execution object according to the position of the target sphere to obtain a key action execution object; the key action execution object is an action execution object for executing a key action on the target sphere; determining a motion event triggered by the target ball; and classifying video clips corresponding to the motion event according to the key action execution object. By adopting the method, the video clip classification efficiency of the motion event can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, device, computer equipment, storage medium and computer program product for classifying video clips of motion events. Background Art

[0002] With the development of image processing technology, more and more scenes are worthy of forming corresponding highlights through image editing. In the process of video production, video clip processing can be performed based on sports events in ball sports.

[0003] In the process of traditional video clip processing, the highlight moment of the motion event triggered by the ball is determined, and then subsequent processing is performed through the video clips before and after the highlight moment. This subsequent processing process takes a lot of time and has low processing efficiency. Summary of the invention

[0004] Based on this, it is necessary to provide a method, device, pan / tilt head, computer equipment, computer-readable storage medium and computer program product for classifying video clips of motion events in response to the above technical problems, which can improve the efficiency of classification processing.

[0005] In a first aspect, the present application provides a method for classifying video clips of sports events. The method comprises:

[0006] Identify the action execution object in the video picture, and detect the position of the target sphere in the video picture;

[0007] According to the position of the target sphere, key action recognition is performed on the action execution object to obtain a key action execution object; the key action execution object is an action execution object that performs the key action on the target sphere;

[0008] Determining a motion event triggered by the target sphere;

[0009] The video clips corresponding to the motion events are classified according to the key action execution objects.

[0010] In one embodiment, the step of identifying an action execution object in a video screen includes:

[0011] Determine the biometric template and the object to be identified in the video image;

[0012] Determining a target biometric template that matches the object to be identified based on the attribute similarity between the object to be identified and the biometric template;

[0013] Among the objects to be identified, the action execution object is determined according to the identity information corresponding to the target biometric template.

[0014] In one embodiment, determining a target biometric template matching the object to be identified based on the attribute similarity between the object to be identified and the biometric template includes:

[0015] Determining the attribute similarity between the object to be identified and the biometric feature template; the feature template has corresponding identity information;

[0016] According to the attribute similarity, determining whether there is a biometric template that meets the matching condition;

[0017] If yes, determining a target biometric template that matches the object to be identified based on the biometric template that meets the matching condition;

[0018] If not, a target biometric template matching the object to be identified and identity information corresponding to the target biometric template are created according to the attribute information of the object to be identified.

[0019] In one embodiment, performing key action detection on the action execution object according to the position of the target sphere to obtain the key action execution object includes:

[0020] Determine a detection range corresponding to the position of the target sphere;

[0021] Determining an object to be detected located in the detection range from among the action execution objects;

[0022] Based on the action type of the object to be detected in the video picture, a key action execution object in the object to be detected is identified.

[0023] In one embodiment, the step of identifying a key action execution object in the object to be detected based on the action type of the object to be detected in the video picture includes:

[0024] Performing human body key point detection on the video screen for the objects to be detected respectively to obtain skeleton information of multiple frames;

[0025] Performing action recognition based on the skeleton information of the multiple frames in time sequence to obtain the action type of the object to be detected;

[0026] According to the action type that matches the pitching action, a key action execution object is determined from the action execution objects.

[0027] In one embodiment, determining the motion event triggered by the target sphere includes:

[0028] Determining a sphere trajectory of the target sphere in the video image according to the position of the target sphere;

[0029] Determining a reference object of the target sphere, and determining a relative position between the target sphere and the reference object;

[0030] According to the ball trajectory and the relative position, a motion event in the video picture is detected.

[0031] In one embodiment, the classifying the video clips corresponding to the motion event according to the key action execution object includes:

[0032] Determine whether there is a classification storage address of the identifier of the key action execution object;

[0033] If yes, then classify and store the video clips corresponding to the motion event according to the existing classification storage addresses;

[0034] If not, create a classified storage address of the identifier to obtain a created classified storage address; and classify and store the video clips corresponding to the motion event according to the created classified storage address.

[0035] In a second aspect, the present application also provides a video clip classification device for sports events. The device comprises:

[0036] A picture detection module, used to identify the action execution object in the video picture and detect the position of the target sphere in the video picture;

[0037] An object detection module is used to perform key action detection on the action execution object according to the position of the target sphere to obtain a key action execution object; the key action execution object is an action execution object that performs a key action on the target sphere;

[0038] An event detection module, used to determine the motion event triggered by the target sphere;

[0039] The segment classification module is used to classify the video segments corresponding to the motion events according to the key action execution objects.

[0040] In a third aspect, the present application also provides a handheld gimbal, comprising a motor and a processor, wherein the motor is used to control the rotation of the gimbal, and the processor implements the steps of classifying video clips of motion events in any of the above embodiments when executing the computer program.

[0041] In a fourth aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of classifying video clips of motion events in any of the above embodiments are implemented.

[0042] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of classifying video clips of motion events in any of the above embodiments are implemented.

[0043] In a sixth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of classifying video clips of motion events in any of the above embodiments are implemented.

[0044] The video clip classification method, device, handheld pan / tilt head, computer equipment, storage medium and computer program product of the above-mentioned motion event can identify the action execution object in the video screen and detect the position of the target sphere in the video screen, so as to determine the identity information of the action execution object, and perform key action recognition on the action execution object according to the position of the target sphere to obtain the key action execution object. Since the position of the target sphere changes in real time, the key action recognition can be applied to the action execution object around the target sphere, thereby reducing the amount of data involved in video processing and improving the detection efficiency of the key action execution object. The motion event triggered by the target sphere can be determined, so that it can be clear that the video clip corresponding to the motion event needs to be classified, so as to classify the video clip corresponding to the motion event according to the key action execution object, and improve the efficiency of the classification processing through the correlation between the key action execution object and the video clip corresponding to the motion event. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A diagram showing an application environment of a method for classifying video clips of motion events in one embodiment;

[0046] Figure 2 is a schematic flow chart of a method for classifying video clips of motion events in one embodiment;

[0047] Figure 3 A schematic diagram of a process for identifying an action execution object in a video screen in one embodiment;

[0048] Figure 4 A schematic diagram of the structure of skeleton information of each frame in an embodiment;

[0049] Figure 5 A schematic diagram of performing action recognition on skeleton information of multiple frames in one embodiment;

[0050] Figure 6 is a structural block diagram of a video clip classification device for motion events in one embodiment;

[0051] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The video clip classification method of the motion event provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 can be, but is not limited to, various cameras, video cameras, panoramic cameras, sports cameras, personal computers, laptops, smart phones, tablet computers, pan / tilt bodies and portable wearable devices, and the portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The terminal 102 can be fixed to the pan / tilt body by welding or the like, and can also be detachably connected or rotatably connected to the pan / tilt body.

[0054] In one embodiment, Figure 2 As shown, a method for classifying video clips of motion events is provided, and the method is applied to Figure 1 The terminal 102 in the example is used as an example to illustrate, and the following steps are included:

[0055] Step 202: Identify the action execution object in the video picture, and detect the position of the target sphere in the video picture.

[0056] The video screen is a multi-frame screen in the video. The video screen includes the action execution object and the target sphere, and may also include objects such as obstacle objects. Obstacle objects include non-target sphere objects, and non-target sphere objects include objects such as human heads, spheres drawn on posters, and spheres outside the stadium.

[0057] The action execution object is an object with identity information. Optionally, after identifying the action object in the video screen, a motion trajectory of the action execution object can be generated for the identity information of a certain action execution object, and the identity information of the action execution object in different video screens can be maintained by the motion trajectory. Optionally, there can be multiple action execution objects in the video screen, and each action execution object is identified by its own identity information. Optionally, the action execution object is an object for action recognition. Optionally, the action execution object is an athlete in the video screen.

[0058] The target sphere is a sphere used to control video acquisition during the video acquisition process. Optionally, the direction of video acquisition, zoom factor and other image acquisition parameters can be determined according to the position of the target sphere in the video screen, so as to control the video acquisition process to be performed according to the target sphere. In the video screen, the position of the target sphere changes with the timestamp.

[0059] Optionally, the process of step 202 may contain at least one of target detection and target tracking techniques; target detection refers to finding all the action execution objects and target spheres of interest in the image, and determining their types and positions. Unlike image classification, which only focuses on whether there is a target in the image, target detection also requires determining the position and bounding box. Target tracking refers to predicting the position of the action recognition object and the target sphere in a frame of image according to a tracking algorithm to obtain their positions in subsequent frames of video. Optionally, the method also includes: creating a database for maintaining athlete information to identify the action execution objects in the video screen based on the database.

[0060] Optionally, the method also includes: detecting the action execution objects in the video screen and the number of action execution objects; if the number of action execution objects detected is unique, the action execution objects specified according to the number are the key action execution objects of the key action points; if the number of action execution objects detected is not unique, executing step 204.

[0061] Step 204 , performing key action recognition on the action execution object according to the position of the target sphere to obtain the key action execution object; the key action execution object is an action execution object that executes the key action on the target sphere.

[0062] The key action is the result of action recognition belonging to a preset action or preset behavior. Optionally, the key action is set for a preset ball. Optionally, there is a certain correspondence between the key action and the motion event triggered by the target ball, and based on this correspondence, when the motion event is triggered by the target ball, the key action execution object corresponding to the motion event can be determined. Optionally, in the case where the preset ball is basketball, the preset action refers to at least one action type determined based on the pattern and features, and the action type can be preset. Exemplarily, in the case where the action type to which the key action belongs is a shooting action, the motion event triggered by the target ball is a goal event; in the case where the action type to which the key action belongs is a rebound action, the motion event triggered by the target ball is a non-goal event; and so on, for some other interception actions, there may also be corresponding motion events.

[0063] Optionally, the key action and the motion event triggered by the target ball have a temporal sequence, and steps 202, 204 and 206 may be performed in chronological order; or step 206 may be performed first, and then steps 202 and 204 are performed through the cached video screen. Optionally, if step 206 is performed first, and then steps 202 and 204 are performed through the cached video screen, the number of executions of the key action detection is relatively small, which helps to improve processing efficiency.

[0064] The key action execution object is an action execution object that performs a key action on the target sphere in the video screen. Key action recognition refers to the process of identifying the preset action or preset behavior of an object by analyzing the patterns and characteristics of human actions. According to the position of the target sphere, the key action recognition of the action execution object can make the detection range of the key action recognition dynamically change with the position of the target sphere, so as to detect the action execution objects around the target sphere, thereby reducing the data resources required for the key action recognition process.

[0065] Optionally, identifying the action execution object in the video picture and detecting the position of the target sphere in the video picture includes: using a detector to detect the action execution object and the target sphere in each frame of the video picture to obtain the respective positions and bounding boxes of the action execution object and the target sphere.

[0066] Correspondingly, according to the position of the target sphere, key action recognition is performed on the action execution object to obtain the key action execution object, including: according to the position of the target sphere in the current frame video screen, the action execution object associated with the target sphere is determined to obtain the current action execution object in the current frame video screen; the current action execution object is a multi-target tracking result; key action recognition is performed on the current action execution object to obtain the key action execution object.

[0067] Specifically, a detector is used to detect the action execution object and the target sphere in each frame of the video to obtain the respective positions and bounding boxes of the action execution object and the target sphere, including: starting from an initial frame set by a user, calling a detection algorithm for detection at a certain time interval, where the detector can be a single-task detector or a multi-task detector; if it is a single-task detector, the time intervals for calling each single-task detector are different.

[0068] Among them, detection algorithms such as the Viola-Jones object detection framework (Viola-Jones), histogram of oriented gradients (HOG), the features in the histogram of oriented gradients are a feature descriptor used for object detection in computer vision and image processing, scale-invariant feature transform algorithm (SIFT), edge detection algorithm, template matching, color features and other methods, or based on deep learning methods such as the two-stage target detection algorithm (Faster Region-based Convolutional Neural Networks, Faster R-CNN), a single-stage dense box detector (Single Shot MultiBox (Detector) SSD), a single-stage dense box detector (RetinaNet, retinal network), YOLOX network, YOLOX network is in the Exceeding YOLO series in 2021.

[0069] Specifically, according to the position of the target sphere, the action execution object associated with the target sphere is determined, and the multi-target tracking result of the current frame is obtained, including the multi-target tracking algorithm such as SORT and Deep SORT to track the target, and obtain the multi-target tracking result of the current frame; assuming that the current frame tracking result contains n targets, it can be expressed as: T = {T_1, T_2, ..., T_ n Assuming that the current frame is the tth frame, the nth face target tracking sequence contains the target rectangle from the target tracking start frame to the current frame, which can be expressed as: T_ n = {B n_t ,B n_t-1 ,...,B n_t0}, where t0 is the frame number in which the target first appears in the picture, t is the frame number in which the target (t+1)th appears in the picture, B n_t Represents the rectangular box of the target in the (t+1)th frame.

[0070] In a feasible implementation, according to the position of the target sphere, a key action detection is performed on the action execution object to obtain the key action execution object, including: determining the detection range corresponding to the position of the target sphere; detecting the behavior pattern and characteristics of the action execution object in the detection range; and identifying the key action execution object that performs the key action according to the behavior pattern and characteristics. Thus, through the behavior pattern of the action execution object, action recognition is performed in combination with its own characteristics to more accurately identify the key action execution object that performs the key action.

[0071] Step 206, determining the motion event triggered by the target sphere.

[0072] A motion event is a certain ball event triggered by a target ball. Optionally, the motion event is determined based on a reference area path of the target ball relative to a reference object to improve the accuracy of motion event recognition. Optionally, the motion event triggered by the target ball includes a goal event. Optionally, the motion event is set for a certain moment and is used to determine the motion event at this moment.

[0073] In one embodiment, determining the motion event triggered by the target ball includes: after obtaining the key action execution object, determining the motion event triggered by the target ball; thereby, once a shooting action is identified, determining the motion event triggered by the target ball to more accurately determine the moment when the motion event is triggered.

[0074] In one embodiment, determining a motion event triggered by a target ball includes: in a video screen arranged in chronological order, judging based on computer vision technology whether the target ball triggers a goal event; if so, determining that the motion event triggered by the target ball is a goal event; if not, determining that the motion event triggered by the target ball is a rebound event.

[0075] Step 208: classify the video clips corresponding to the motion events according to the key action execution objects.

[0076] The video clip corresponding to the sports event includes the video clip when the sports event is triggered. The video clip can be a video clip during the execution of the sports event, and each sports event has its own key action execution object. Optionally, in a basketball scene, the video clip corresponding to the sports event is a goal clip.

[0077] In one embodiment, the video clips corresponding to the sports events are classified according to the key action execution objects, including: the video clips corresponding to the sports events are classified and processed according to the key action execution objects, so that the video clips corresponding to each sports event are processed differently for different key action execution objects. Thus, in the case where there are multiple action execution objects on the basketball court, the video clips corresponding to each key action execution object can be personalized according to the needs of the key action execution object itself, and their respective highlights can be obtained, thereby realizing efficient processing of the video clips. Optionally, the classification processing includes, but is not limited to, classified storage according to the key action execution objects, classified addition of special effects, and classified input into certain neural network models.

[0078] In one embodiment, video clips corresponding to motion events are classified according to key action execution objects, including: according to identity information of key action execution objects, video clips corresponding to motion events are classified and stored, so that the identity information of the key action execution objects can be used as an index to find the video clips corresponding to each motion event.

[0079] In the above-mentioned video clip classification method of motion events, the action execution object in the video screen is identified, and the position of the target sphere in the video screen is detected, so that the identity information of the action execution object can be determined, and the key action recognition of the action execution object can be performed according to the position of the target sphere to obtain the key action execution object. Since the position of the target sphere changes in real time, the key action recognition can be applied to the action execution object around the target sphere, which reduces the amount of data involved in video processing and improves the detection efficiency of the key action execution object. The motion event triggered by the target sphere can be determined, so that it can be clear that the video clip corresponding to the motion event needs to be classified, so that the video clip corresponding to the motion event is classified according to the key action execution object, and the efficiency of the classification processing is improved by the correlation between the key action execution object and the video clip corresponding to the motion event.

[0080] In one embodiment, Figure 3 As shown, identifying the action execution object in the video screen includes:

[0081] Step 302: Determine the biometric template and the object to be identified in the video image.

[0082] A biometric template is a set of features set for an object, used to characterize the identity information of the object. Optionally, the biometric information of an object can be collected, and then the features of the biometric information can be extracted and modeled to obtain a biometric template; wherein, a sensor or device can be used to collect the biometric information of an individual, and the biometric information includes but is not limited to fingerprints, faces, irises, voiceprints or palms, appearance and other information; the collected biometric information can be converted into a digital model or feature vector for subsequent comparison and verification. This step involves technologies such as image processing, pattern recognition and feature extraction algorithms. Optionally, a biometric template is a preset feature set for a user, used to characterize the user's identity information.

[0083] The object to be identified is an object in the video screen whose identity information needs to be identified. Optionally, the human body object in the video screen can be determined by classification identification, and the human body object with unknown identity information is the object to be identified.

[0084] Step 304: Determine a target biometric template that matches the object to be identified based on the attribute similarity between the object to be identified and the biometric template.

[0085] Attribute similarity refers to the similarity between the biometric template and the object to be identified in terms of attributes. Optionally, when the attribute similarity between the object to be identified and a certain biometric template exceeds a certain threshold, the biometric template can be determined as the target biometric template of the object to be identified.

[0086] In one embodiment, step 304 is a re-identification technology (Re-IDentification, ReID), which refers to a technology for re-identifying the same biometric template in multiple cameras or time periods.

[0087] In an optional implementation, person re-identification (Person ReID) is used, and the person re-identification process is based on the appearance features of the human body. Specifically, according to the attribute similarity between the object to be identified and the biometric template, a target biometric template matching the object to be identified is determined, including: in the video screen, the human body area framed by the human detector is cut off and scaled to a specified size; the human body area of ​​the specified size is input into the feature extraction network to obtain a feature vector; then the feature vector is calculated one by one with the feature vectors of all athletes in the database; if the similarity with a certain athlete biometric template in the database is greater than a preset similarity threshold, it is determined that the object feature template with a similarity greater than the preset similarity threshold is the target biometric template matching the object to be identified.

[0088] In an optional implementation, face recognition is used, which is a biometric recognition technology that automatically performs identity recognition based on human facial features (such as statistical, geometric or depth features, etc.). In the video screen, face detection is performed in the human body area, and the face area is cut out and scaled to a specified face size such as 96x112; the face area of ​​the specified face size is input into the feature extraction network to obtain a feature vector; then the feature vector is calculated one by one with the feature vectors of all athletes in the database; if the similarity with a certain athlete biometric template in the database is greater than a preset similarity threshold, it is determined that the object feature template with a similarity greater than the preset similarity threshold is the target biometric template that matches the object to be identified.

[0089] In an optional implementation, face re-identification may be used first, and if the target biometric template that matches the object to be identified is not determined by face re-identification, pedestrian re-identification may be used. Thus, face recognition is used first to focus on facial features, which ensures higher recognition accuracy with a smaller amount of data; while pedestrian re-identification is performed on the features of various parts of the body, and can determine the behavior pattern of the object to be identified during movement, thereby determining the target biometric template that matches the object to be identified.

[0090] In one embodiment, determining a target biometric template that matches the object to be identified based on the attribute similarity between the object to be identified and the biometric template includes: determining the attribute similarity between the object to be identified and the biometric template; the feature template has corresponding identity information; judging whether there is a biometric template that meets the matching conditions based on the attribute similarity; if so, determining a target biometric template that matches the object to be identified based on the biometric template that meets the matching conditions; if not, creating a target biometric template that matches the object to be identified and the identity information corresponding to the target biometric template based on the attribute information of the object to be identified.

[0091] The attribute information of the object to be identified is the biometric information collected from the video screen, which is used to characterize a certain object to be identified. The attribute information of the object to be identified is a feature set of the object to be identified, which can be used to characterize the identity information of the object to be identified. The attribute information of the object to be identified includes but is not limited to the facial information and body information of the object to be identified.

[0092] The target biometric template may be a pre-set biometric template or may be attribute information created based on the attribute information of the object to be identified. Optionally, when creating the target biometric template based on the attribute information of the object to be identified, the attribute information of the object to be identified may be used as the target biometric template, and an identity identifier may be generated according to certain rules as the identity information corresponding to the target biometric template; optionally, the attribute information of the object to be identified may be classified and reorganized according to storage rules such as the storage format and storage data volume of the biometric template to obtain the target biometric template; optionally, an identity identifier may be created according to date and time information, and the identity identifier may be used as the identity information corresponding to the target biometric template.

[0093] The matching condition is an evaluation index of attribute similarity. When the attribute similarity satisfies a preset threshold or is greater than or equal to the preset threshold, the evaluation index is achieved, and the biometric template can be determined as the target biometric template that matches the object to be identified.

[0094] In a feasible implementation, determining the attribute similarity between the object to be identified and the biometric template includes: determining at least one of face attribute similarity and body attribute similarity between the facial features of the object to be identified and the biometric template.

[0095] Correspondingly, judging whether there is a biometric template that meets the matching conditions based on the attribute similarity includes: judging whether there is a biometric template that meets the matching conditions based on at least one attribute similarity of face attribute similarity and body attribute similarity.

[0096] In a feasible implementation, face re-recognition is first used. If face re-recognition is passed, it is determined whether there is a target biometric template that matches the object to be identified. Then, pedestrian re-recognition is used to determine whether there is a target biometric template that matches the object to be identified. Specifically, according to the attribute similarity, it is determined whether there is a biometric template that meets the matching conditions, including: if there is a biometric template with a facial attribute similarity greater than a preset threshold, then the biometric template with a facial attribute similarity greater than the preset threshold is a biometric template that meets the matching conditions; if there is no biometric template with a facial attribute similarity greater than the preset threshold, it is determined whether there is a biometric template with a human body attribute similarity greater than the preset threshold, and the biometric template with a human body attribute similarity greater than the preset threshold is a biometric template that meets the matching conditions; if there is no biometric template with a facial attribute similarity greater than the preset threshold, and there is no biometric template with a human body attribute similarity greater than the preset threshold, then there is no biometric template that meets the matching conditions.

[0097] Therefore, the target biometric template can be pre-recorded or dynamically created during the game. The user can freely choose the method of determining the target biometric template. When there is no biometric template that meets the matching conditions, the target biometric template and the identity information corresponding to the target biometric template can still be created to more efficiently classify the video clips corresponding to the sports event.

[0098] Step 306: Determine the action execution object among the objects to be identified according to the identity information corresponding to the target biometric template.

[0099] In one embodiment, in the object to be identified, determining the action execution object according to the identity information corresponding to the target biometric template includes: determining the identity information of the object to be identified according to the identity information corresponding to the target biometric template, so that the object to be identified serves as the action execution object.

[0100] Optionally, when there are multiple objects to be identified in the video image, a target biometric template of each object to be identified may be determined to determine the identity information of each object to be identified, thereby determining each action execution object.

[0101] In this embodiment, the biometric template is compared with the subsequent user data capture to determine the attribute similarity, so as to use the attribute similarity to reflect the unique physical or behavioral attributes of an individual and realize identity authentication. On this basis, since in this method, the key action of the action execution object is identified according to the position of the target sphere, the identity information of the same key action execution object may be re-identified, and the consistency of the identity information can be better guaranteed by the biometric template, and the object obtained by re-identification will not change the identity information, so that the video clips corresponding to the motion event can be classified, avoiding too many classifications, so as to improve the efficiency of subsequent classification processing.

[0102] In one embodiment, key action identification is performed on the action execution object according to the position of the target sphere to obtain the key action execution object, including: determining the detection range corresponding to the position of the target sphere; determining the object to be detected located in the detection range from the action execution object; and identifying the key action execution object in the object to be detected based on the action type of the object to be detected in the video screen.

[0103] The detection range is determined according to the position of the target sphere. At different time points, the position of the target sphere in the video image changes, and the detection range changes accordingly, thereby causing the action execution object to be detected to change. Optionally, the detection range is the distance interval between each action execution object and the target sphere.

[0104] In one embodiment, an object to be detected located in a detection range is determined from the action execution objects, including: in the current frame video picture, respectively calculating the distances between the positions of multiple action execution objects and the target sphere to obtain each distance to be detected in the current frame video picture; and determining the object to be detected based on the action execution objects whose distance to be detected is less than or equal to the critical value of the detection range.

[0105] In another embodiment, an object to be detected located in a detection range is determined from the action execution objects, including: in a current frame video image, respectively calculating the distances between a plurality of action execution objects and the positions of a target sphere to obtain each distance to be detected in the current frame video image; and determining the object to be detected based on at least one action execution object with the smallest distance to be detected.

[0106] Optionally, before step 204, the method further includes: the multi-target tracker creates a track for each target and assigns an ID to the track, and maintains the ID during the life cycle of the entire track, and each ID is unique in the entire video. For the action execution object, when starting a track life cycle, the action execution object is first identified by biometric technology. If the athlete already exists in the database, then the athlete's ID is taken out from the database and assigned to the track, otherwise a new ID is created in the data, and then the ID is assigned to the track, so as to determine the object to be detected located in the detection range from the action execution object according to the track.

[0107] In this embodiment, the detection range corresponding to the position of the target sphere is determined, so that the detection range can be changed based on the position of the target sphere, so as to determine the objects to be detected located in the detection range from the action execution objects, so that the number of action execution objects to be detected is reduced; and the action type of the objects to be detected in the video picture can accurately identify the key action execution objects among the objects to be detected.

[0108] In one embodiment, based on the action type of the object to be detected in the video picture, the key action execution object in the object to be detected is identified, including: performing human body key point detection on the object to be detected in the video picture to obtain multi-frame skeleton information; performing action recognition based on the multi-frame skeleton information in time sequence to obtain the action type of the object to be detected; and determining the key action execution object from the action execution objects according to the action type that conforms to the pitching action.

[0109] The key points of the human body are set for each object to be detected and are used to reflect the skeleton information of the object to be detected. Optionally, the key points of the human body are positions that can produce curvature changes, including but not limited to key points of the human body such as the neck, shoulder, elbow, wrist, waist, knee, ankle and other positions.

[0110] By arranging the skeleton key points of each object to be detected in time sequence, multi-frame skeleton information of each object to be detected can be formed, so as to represent the action of the corresponding object to be detected through the multi-frame skeleton information. The multi-frame skeleton information is connected in sequence according to the human body key points of each object to be detected in the video screen. The multi-frame skeleton information is composed of the skeleton information of each frame. The skeleton information of each frame is as follows: Figure 4 As shown in (1), (2) and (3) in FIG, a schematic diagram of action recognition based on the skeleton information of multiple frames of images based on the temporal sequence is shown in FIG. Figure 5 As shown; among them, Figure 4 The a-th frame, the b-th frame, and the c-th frame are the second video frame, the third video frame, and the fourth video frame, respectively.

[0111] In a feasible implementation, action recognition is performed based on skeleton information of multiple frames in time sequence to obtain the action type of the object to be detected, including: based on the time sequence, action recognition is performed on the skeleton information of multiple frames to obtain the action characteristics of the object to be detected; and the action type of the object to be detected is determined according to the action type that the action characteristics of the object to be detected conform to.

[0112] In another feasible implementation, action recognition is performed based on skeleton information of multiple frames in time sequence to obtain the action type of the object to be detected, including: standardizing the spatiotemporal graph composed of key points of the human body in the video image to obtain skeleton information of multiple frames in time sequence; substituting the skeleton information of the multiple frames in time sequence into a certain action recognition algorithm for encoding to obtain action features of the object to be detected; inputting the action features of the object to be detected into a classification network for action recognition to realize action recognition, and outputting the defined action type.

[0113] Optionally, since the size of the human body in the picture is variable, the detected human body key points are at different scales, so the scales of the key points of different frames of video need to be unified. In addition, due to the viewing angle, the human body may rotate, so the connection lines of some human body key points need to be consistent with the direction of the reference line. Position standardization generally involves stringing together human body frames of different frames according to a point (such as the center point of the spine) for standardization on the time axis after completing scale normalization and angle normalization.

[0114] In a feasible implementation, according to the action type that matches the pitching action, the key action execution object is determined from the action execution objects, including: determining the action execution object whose action type is the pitching action as the key action execution object; or, filtering out the key action execution object from the action execution objects whose action type is the pitching action.

[0115] For example, the model input dimension of the classification network for action recognition is usually (N, C, T, V, M), where:

[0116] N represents the number of videos. Usually a batch has 256 videos. The number of videos should preferably be a power of 2.

[0117] C represents the characteristics of the joint. Usually a joint contains three features: x, y, and acc (if it is a three-dimensional skeleton, it is four). x and y are the position coordinates of the node joint, and acc is the confidence.

[0118] T represents the number of key frames in the video. Generally, a video has 150 frames.

[0119] V represents the number of joints, and usually one person labels 18 joints.

[0120] M represents the number of people in a frame, and generally the two people with the highest average confidence are selected

[0121] Optionally, skeleton information may be acquired using a deep learning algorithm; the deep learning algorithm used to acquire skeleton information includes a real-time human posture recognition algorithm and a skeleton-based behavior recognition algorithm.

[0122] Among them, the real-time human posture recognition algorithms include but are not limited to OpenPose, DensePose and EfficientPose. OpenPose is a human posture recognition project developed by Carnegie Mellon University (CMU) in the United States based on convolutional neural networks and supervised learning and developed with Caffe as the framework. DensePose is a human posture recognition project developed by Facebook researchers Natalia Neverova, Iasonas Kokkinos and INRIA in France. An amazing real-time human pose recognition system developed by Alp Guler; EfficientPose, which is an efficient, accurate and scalable end-to-end 6D multi object pose estimation approach, is a new method for 6D target pose estimation. The input of the network is an RGB image, and the output of the network is the 2D BoundingBox of all objects to be detected in the image and the 6D Pose (x, y, z, roll, yaw, pitch) in three-dimensional space. For example, OpenPose extracts 18 body joints and 17 lines connecting the joints to obtain the skeleton. These skeleton information will be stored in chronological order for input features for subsequent action recognition.

[0123] Skeleton-based behavior recognition algorithms include but are not limited to Spatial Temporal Graph Convolutional Networks (ST-GCN), Adaptive Graph Convolutional Neural Networks (AGCN), and Channel-wise Topology Refinement Graph Convolution (CTR-GCN).

[0124] In one embodiment, determining a motion event triggered by a target sphere includes: determining a sphere trajectory of the target sphere in a video screen according to the position of the target sphere; determining a reference object of the target sphere and determining a relative position between the target sphere and the reference object; and determining a motion event in the video screen according to the sphere trajectory and the relative position.

[0125] The sphere trajectory is a set of positions of the target sphere in multiple frames of the video, which is obtained by arranging the positions of the target sphere in chronological order. Through the sphere trajectory, the multiple frames of the video can complement each other's information to ensure more accurate identification of the position of the target sphere.

[0126] The relative position of the target sphere and the reference object can reduce the position deviation of the target sphere caused by the camera movement process, so that the position change of the target sphere can be clearly displayed. Therefore, when the focal length changes and causes the video picture to change, the change of the video picture will also cause the reference system to send this change, which in turn causes the position of the target sphere and the reference object to change synchronously, making the change in relative position smaller, thereby ensuring accuracy.

[0127] Optionally, the target sphere is a moving sphere, and the sphere trajectory is a sphere area trajectory of the sphere in the video picture; according to the sphere trajectory and relative position, detecting motion events in the video picture, including: identifying the moving sphere based on the sphere area trajectory in multiple frames of video pictures; determining the target sphere from each moving sphere according to the confidence that each moving sphere belongs to a preset ball type; according to the position information of the target sphere, collecting multiple frames of moving pictures of the target sphere; in each frame of the moving picture of the target sphere, determining the motion event triggered by the target sphere according to the relative position of the target sphere and a reference object.

[0128] Identify the sphere area in the picture. This process can include all spheres that can be classified and identified, or it can be inferred through the trajectory of the identified sphere at a certain stage. The sphere area trajectory is the area where each sphere is located in multiple frames of video. The areas that belong to the same sphere and change sequentially in multiple frames are correlated between frames to obtain the sphere area trajectory of the sphere.

[0129] Among them, the moving sphere occupies a small proportion in the picture, and may be in a semi-occluded state, or in a motion blur state in the air. Therefore, the sphere area in the multi-frame video picture is inter-frame associated to form a sphere area trajectory, so that the moving sphere can be identified from the spheres detected stably in continuous multi-frames. The moving sphere is a sphere that moves in the multi-frame picture, that is, a non-stationary sphere. Therefore, although the characteristics of the sphere are not very obvious, it is not easy to misdetect the spherical logo (spherical LOGO), human head, human face, etc. on the poster as a moving sphere to avoid false alarms.

[0130] The preset ball is a preset round ball, and the size of the preset round ball can be set for a certain specification or for multiple specifications; the preset ball includes but is not limited to basketball, football, volleyball, and table tennis. Optionally, the preset ball is set according to the sports scene; when the sports scene is a basketball court, the preset ball is basketball; when the sports scene is a football field, the preset ball is football, and so on, the relationship between the sports scene and the preset ball. The confidence that the sports ball belongs to the preset ball is calculated based on whether the features of the sports ball match the features of the preset ball. Optionally, the confidence of the sports ball in a frame of video can be determined based on image features, semantic features, association features or other features, and then the confidence of the sports ball in each frame of video is comprehensively calculated.

[0131] In one embodiment, in each frame of the target sphere motion picture, the motion event triggered by the target sphere is determined according to the relative position of the target sphere and the reference object, including: in each frame of the target sphere motion picture, according to the relative position of the target sphere and the reference object, the reference area path of the target sphere relative to the reference object is determined; according to the reference area path, the candidate event picture in each frame of the motion picture is determined; in the candidate event picture, the motion event of the target sphere is determined according to the occlusion relationship between the target sphere and the reference object. Thus, the reference area path can be generated by the relative position of the target sphere and the reference object such as the basket, and the motion event of the target sphere can be more accurately identified during the camera movement, and the spatial misjudgment problem during the camera movement can be solved through the occlusion relationship between the target sphere and the reference object such as the basket.

[0132] Among them, the candidate event screen is the screen of the ball event determined according to the reference area path. Since the motion screen of the ball event may occur after the state judgment is performed according to the reference area path, and the reference area path represents the positional relationship of the screen in two-dimensional space, it may encounter the problem of spatial dislocation. Therefore, it is judged whether the target sphere generates the corresponding candidate event screen through the occlusion relationship between the target sphere and the basket. The occlusion relationship is obtained by analyzing the characteristics of the target sphere and the reference object when there is a mutual occlusion relationship between the target sphere and the reference object. Optionally, the occlusion relationship can be represented by the characteristics of the target sphere occluding the reference object, or by the characteristics of the reference object occluding the target sphere.

[0133] In this embodiment, motion events in the video image are determined based on the trajectory and relative position of the ball, so that motion events occurring in the target ball can be more accurately identified. The occlusion relationship between the target ball and reference objects such as the basket can be used to solve the problem of spatial misjudgment in the camera movement process, thereby making the execution process of the video clip classification method more efficient.

[0134] In one embodiment, the video clips corresponding to the motion events are classified according to the key action execution objects, including: determining whether there is a classification storage address that identifies the key action execution object; if so, classifying and storing the video clips corresponding to the motion events according to the existing classification storage address; if not, creating an identified classification storage address to obtain a created classification storage address; and classifying and storing the video clips corresponding to the motion events according to the created classification storage address.

[0135] The identifier of the key action execution object is used to characterize the identity information of the key action execution object. Optionally, the identifier of the key action execution object is the ID of the key action execution object. Optionally, the identifier of the key action execution object corresponds to the key action execution object, and each identifier of the key action execution object corresponds to a key action execution object. Exemplarily, the identifier of the key action execution object is the ID of the athlete.

[0136] The classification storage address is an address generated based on the identification of the key action execution object; optionally, the classification storage address is a storage path, which is a folder name, and this folder name contains the identification of the key action execution object, so that the video clip corresponding to the motion event is stored in the folder belonging to this folder name.

[0137] The created classification storage address refers to a classification storage address created by the terminal 102 when the identification of the key action execution object does not exist. After the creation of the classification storage address, if the video clips corresponding to the motion event are classified again according to the same key action execution object, the classification storage address of the identification of the key action execution object exists.

[0138] In an exemplary embodiment, after a goal-scoring event is identified, a goal-scoring video clip is captured and stored in a folder named after the ID of the player who generated the shooting action; if this folder does not exist, one is created and then saved.

[0139] When a key action execution object has a classified storage address, the video clips corresponding to the motion events can be classified and stored according to the classified storage address of this key action execution object, so that after the same key action execution object executes the key action, the video clips corresponding to the motion events caused by it are stored centrally to form a highlight collection, thereby improving the efficiency of classified storage; when a key action execution object does not have a classified storage address, the video clips corresponding to the motion events are classified and stored according to the classified storage address after creation, without the need to manually create the classified storage address of this key action execution object, thereby improving the efficiency of classified storage.

[0140] In an exemplary embodiment, when using a basketball mode of a handheld gimbal, the mode will automatically identify highlight moments during the recording of the game, and save the clips before and after the highlight moments for subsequent editing. However, the highlight moments are not tied to the athletes, which means that the program only knows when a goal is scored but does not know who scored the goal. There will be many highlight moments in a game, and if you want to find all the videos of a certain athlete, you have to look through them one by one. This method is time-consuming and laborious, and to some extent, it will kill the enthusiasm of the athletes for post-editing.

[0141] After the handheld pan-tilt is integrated with any embodiment of the technical solution, the highlight moment is automatically identified during the recording of the game, and the highlight moment is intercepted and stored in the shooter's folder. After the game is over, the classified highlight clips will be obtained, so that each athlete can quickly obtain his own highlights. The present invention proposes a method for automatically associating a goal clip with an athlete ID (the identifier of the key action execution object). Each goal highlight clip will be classified according to the athlete's ID and stored in different folders. Athletes can quickly obtain their own highlights, thereby stimulating their enthusiasm for editing.

[0142] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0143] Based on the same inventive concept, the embodiment of the present application also provides a motion event video segment classification device for implementing the above-mentioned motion event video segment classification method. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above-mentioned method, so the specific limitations of the one or more motion event video segment classification device embodiments provided below can refer to the above-mentioned limitations on the motion event video segment classification method, and will not be repeated here.

[0144] In an exemplary embodiment, Figure 6 As shown, a video clip classification device for motion events is provided, comprising:

[0145] A picture detection module 602 is used to identify the action execution object in the video picture and detect the position of the target ball in the video picture;

[0146] The object detection module 604 is used to perform key action detection on the action execution object according to the position of the target sphere to obtain a key action execution object; the key action execution object is an action execution object that performs a key action on the target sphere;

[0147] An event detection module 606 is used to determine a motion event triggered by the target ball;

[0148] The segment classification module 608 is used to classify the video segments corresponding to the motion events according to the key action execution objects.

[0149] In one embodiment, the picture detection module 602 is used to:

[0150] Determine the biometric template and the object to be identified in the video image;

[0151] Determining a target biometric template that matches the object to be identified based on the attribute similarity between the object to be identified and the biometric template;

[0152] Among the objects to be identified, the action execution object is determined according to the identity information corresponding to the target biometric template.

[0153] In one embodiment, the picture detection module 602 is used to:

[0154] Determining the attribute similarity between the object to be identified and the biometric feature template; the feature template has corresponding identity information;

[0155] According to the attribute similarity, determining whether there is a biometric template that meets the matching condition;

[0156] If yes, determining a target biometric template that matches the object to be identified based on the biometric template that meets the matching condition;

[0157] If not, a target biometric template matching the object to be identified and identity information corresponding to the target biometric template are created according to the attribute information of the object to be identified.

[0158] In one embodiment, the object detection module 604 is used to:

[0159] Determine a detection range corresponding to the position of the target sphere;

[0160] Determining an object to be detected located in the detection range from among the action execution objects;

[0161] Based on the action type of the object to be detected in the video picture, a key action execution object in the object to be detected is identified.

[0162] In one embodiment, the object detection module 604 is used to:

[0163] Performing human body key point detection on the video screen for the objects to be detected respectively to obtain skeleton information of multiple frames;

[0164] Performing action recognition based on the skeleton information of the multiple frames in time sequence to obtain the action type of the object to be detected;

[0165] According to the action type that matches the pitching action, a key action execution object is determined from the action execution objects.

[0166] In one embodiment, the event detection module 606 is used to:

[0167] Determining a sphere trajectory of the target sphere in the video image according to the position of the target sphere;

[0168] Determining a reference object of the target sphere, and determining a relative position between the target sphere and the reference object;

[0169] According to the ball trajectory and the relative position, a motion event in the video picture is detected.

[0170] In one embodiment, the segment classification module 608 is used to:

[0171] Determine whether there is a classification storage address of the identifier of the key action execution object;

[0172] If yes, then classify and store the video clips corresponding to the motion event according to the existing classification storage addresses;

[0173] If not, create a classified storage address of the identifier to obtain a created classified storage address; and classify and store the video clips corresponding to the motion event according to the created classified storage address.

[0174] Each module in the video clip classification device for motion events can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0175] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a video clip classification method for a motion event is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0176] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0177] The present application also provides a handheld gimbal, comprising a motor and a processor, wherein the motor is used to control the rotation of the gimbal, and the processor implements the steps of classifying video clips of motion events in any of the above embodiments when executing the computer program.

[0178] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0179] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0180] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0181] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0182] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0183] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0184] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for classifying video clips of motion events, characterized in that: The method comprises: Identify the action execution object in the video picture, and detect the position of the target sphere in the video picture; According to the position of the target sphere, key action recognition is performed on the action execution object to obtain a key action execution object; the key action execution object is an action execution object that performs the key action on the target sphere; Determining a motion event triggered by the target sphere; The video clips corresponding to the motion events are classified according to the key action execution objects.

2. The method according to claim 1, characterized in that The step of identifying the action execution object in the video screen includes: Determine the biometric template and the object to be identified in the video image; Determining a target biometric template that matches the object to be identified based on the attribute similarity between the object to be identified and the biometric template; Among the objects to be identified, the action execution object is determined according to the identity information corresponding to the target biometric template.

3. The method according to claim 2, characterized in that The step of determining a target biometric template that matches the object to be identified based on the attribute similarity between the object to be identified and the biometric template includes: Determining the attribute similarity between the object to be identified and the biometric feature template; the feature template has corresponding identity information; According to the attribute similarity, determining whether there is a biometric template that meets the matching condition; If yes, determining a target biometric template that matches the object to be identified based on the biometric template that meets the matching condition; If not, a target biometric template matching the object to be identified and identity information corresponding to the target biometric template are created according to the attribute information of the object to be identified.

4. The method according to claim 1, characterized in that: The step of performing key action detection on the action execution object according to the position of the target sphere to obtain the key action execution object comprises: Determine a detection range corresponding to the position of the target sphere; Determining an object to be detected located in the detection range from among the action execution objects; Based on the action type of the object to be detected in the video picture, a key action execution object in the object to be detected is identified.

5. The method according to claim 4, characterized in that The step of identifying a key action execution object in the object to be detected based on the action type of the object to be detected in the video picture includes: Performing human body key point detection on the video screen for the objects to be detected respectively to obtain skeleton information of multiple frames; Performing action recognition based on the skeleton information of the multiple frames in time sequence to obtain the action type of the object to be detected; According to the action type that matches the pitching action, a key action execution object is determined from the action execution objects.

6. The method according to claim 1, characterized in that The step of determining the motion event triggered by the target sphere comprises: Determining a sphere trajectory of the target sphere in the video image according to the position of the target sphere; Determining a reference object of the target sphere, and determining a relative position between the target sphere and the reference object; According to the ball trajectory and the relative position, a motion event in the video picture is detected.

7. The method according to claim 1, characterized in that The classifying the video clips corresponding to the motion events according to the key action execution objects includes: Determine whether there is a classification storage address of the identifier of the key action execution object; If yes, then classify and store the video clips corresponding to the motion event according to the existing classification storage addresses; If not, create a classified storage address of the identifier to obtain a created classified storage address; and classify and store the video clips corresponding to the motion event according to the created classified storage address.

8. A video clip classification device for sports events, characterized in that: The device comprises: A picture detection module, used to identify the action execution object in the video picture and detect the position of the target sphere in the video picture; An object detection module is used to perform key action detection on the action execution object according to the position of the target sphere to obtain a key action execution object; the key action execution object is an action execution object that performs a key action on the target sphere; An event detection module, used to determine the motion event triggered by the target sphere; The segment classification module is used to classify the video segments corresponding to the motion events according to the key action execution objects.

9. A handheld gimbal, characterized in that: It comprises a motor and a processor, wherein the motor is used to control the rotation of the pan / tilt head, and the processor is used to implement the steps of the method described in any one of claims 1 to 7.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.