Shooting counting method and device, electronic equipment and storage medium

By processing basketball shooting video streams using computer vision and intelligent analysis algorithms, the system identifies and tracks human and basketball targets, solving the problem of low efficiency in manual counting and achieving automation and high efficiency in basketball shooting counting.

CN121747006APending Publication Date: 2026-03-27绍兴市北大信息技术科创中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, shot counting mainly relies on manual counting, which is inefficient and prone to omissions or miscounts.

Method used

By acquiring basketball shooting video streams, computer vision and intelligent analysis algorithms are used to identify and track human targets, key points of the human body, basketball targets, nets and hoops in the video frames, and match the movement trajectories of human targets and basketball targets to achieve automatic and real-time shooting counting.

Benefits of technology

It automates, enables real-time, and increases the efficiency of shot counting, improving the accuracy and efficiency of counting while reducing deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747006A_ABST
    Figure CN121747006A_ABST
Patent Text Reader

Abstract

The invention discloses a shooting counting method and device, electronic equipment and a storage medium. Comprising the following steps: for each video frame in an obtained video frame shooting video stream, positioning a human body target in the video frame and positioning human body key points to obtain all human body areas and human body key point positions; positioning a basketball target in the video frame to obtain all basketball areas; positioning baskets and nets in the video frame to obtain all basket areas and net areas; tracking the human body target in each video frame to obtain at least one human body motion track; tracking a basketball target in each video frame to obtain at least one basketball movement track; and for a human body target in each human body motion trail, matching a basketball target for the human body target, and obtaining a shooting counting result of the human body target based on the basketball motion trail of the matched basketball target, and the basket area and the net area on each video frame.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application generally relates to the technical field of computer vision. More particularly, the present application relates to a shooting counting method and device, an electronic device and a storage medium. BACKGROUND

[0002] Basketball is a popular sport, which is not only the core content of school sports teaching and various levels of competition, but also an important way to improve students' physical quality and motor coordination. Shooting practice, as one of the basic techniques of basketball training, is often used to measure the training intensity, ball control proficiency and technical stability of the training personnel, so the counting of shooting times is particularly important.

[0003] Currently, the counting of shooting times mainly relies on manual counting. However, manual counting is low in efficiency and prone to missing or miscounting. Therefore, it is urgent to provide a shooting counting method, device, electronic device and storage medium to improve the efficiency and accuracy of shooting counting. SUMMARY

[0004] In order to at least solve one or more technical problems mentioned above, the present application provides a shooting counting method, device, electronic device and storage medium in multiple aspects.

[0005] In a first aspect, the present application provides a shooting counting method, comprising: obtaining a shooting video stream; for each video frame in the shooting video stream, performing region positioning on a human target in the video frame and positioning a human key point to obtain all human regions in the video frame and the positions of human key points in each human region; performing region positioning on a basketball target in the video frame to obtain all basketball regions in the video frame; performing region positioning on a basket target and a net target in the video frame to obtain all basket regions and net regions in the video frame; tracking the human target in each video frame based on the human regions in the video frame to obtain at least one human motion trajectory; tracking the basketball target in each video frame based on the basketball regions in the video frame to obtain at least one basketball motion trajectory; for the human target in each human motion trajectory, matching a basketball target for the human target, and obtaining a shooting counting result of the human target based on the basketball motion trajectory of the matched basketball target, the basket regions and the net regions on each video frame.

[0006] In some embodiments, the method satisfies at least one of the following: the region positioning of the human target and the positioning of the human key points in the video frame to obtain all human regions in the video frame and the human key point positions in each human region includes: positioning the human target and the human key points in the video frame by using a trained human key point detection model to obtain all human regions in the video frame and the human key point positions in each human region; the region positioning of the basketball target in the video frame to obtain all basketball regions in the video frame includes: positioning the basketball target in the video frame by using a trained basketball detection model to obtain all basketball regions in the video frame; and the region positioning of the basket target and the net target in the video frame to obtain all basket regions and net regions in the video frame includes: positioning the basket target and the net target in the video frame by using a trained net and basket detection model to obtain all basket regions and net regions in the video frame.

[0007] In some embodiments, the method satisfies at least one of the following: the human key point detection model adopts a YOLOv8-pose model; the basketball detection model adopts a YOLOv8 model, and the YOLOv8 model adopted by the basketball detection model adopts a Dy Sample dynamic up-sampling; and the net and basket detection model adopts a YOLOv8 model.

[0008] In some embodiments, before matching a basketball target for the human target in each human motion trajectory, the method further includes: selecting a to-be-identified video frame from all video frames for the human target in each human motion trajectory; wherein the face of the human target in the to-be-identified video frame is complete and clear; cropping a face region from the human region in the to-be-identified video frame; extracting features of the face region, and comparing the extracted target features with each face feature stored in a first database to obtain identity information of the human target.

[0009] In some embodiments, the human key points include wrist key points, and the matching of a basketball target for the human target in each human motion trajectory includes: acquiring a to-be-detected video frame when the human target stands with a ball in the shooting video stream for the human target in each human motion trajectory; and determining the basketball target matched with the human target according to the wrist key point position of the human target and each basketball region in the to-be-detected video frame.

[0010] In some embodiments, the method further comprises: based on the basketball movement trajectory of the matched basketball target, determining the motion state of the basketball target in each basketball region of the basketball movement trajectory of the matched basketball target based on the relative position of the basketball region and the basket region and the net region on the video frame where the basketball region is located; the motion state comprises: a pre-entry state, an entry state and a post-entry state; when the state transition process of one valid shot is determined based on the motion state of each basketball target, one shot counting is performed to obtain the shot counting result of the human target; wherein the state transition process of one valid shot refers to the state transition process of the basketball target from the pre-entry state, the entry state to the post-entry state.

[0011] In some embodiments, the method further comprises: based on the basketball movement trajectory of the matched basketball target, performing difference calculation on the longitudinal coordinates of the set points of the basketball regions in the continuous video frames to obtain a difference sequence comprising a plurality of difference values; when the state transition process of one valid shot is determined based on the motion state of each basketball target, it is determined whether the motion direction of the basketball target in the state transition process of the current valid shot is from top to bottom and whether the basketball movement trajectory of the matched basketball target passes through the basket region first and then passes through the net region based on the difference values involved in the state transition process of the current valid shot in the difference sequence; the performing one shot counting comprises: if the motion direction of the basketball target in the state transition process of the current valid shot is from top to bottom and the basketball movement trajectory of the matched basketball target passes through the basket region first and then passes through the net region, one shot counting is performed.

[0012] In some embodiments, the method further comprises: when the state transition process of one valid shot is determined based on the motion state of each basketball target, it is determined whether the net shakes in the state transition process of the current valid shot based on the net region in the video frame involved in the state transition process of the current valid shot; the performing one shot counting comprises: if the net shakes in the state transition process of the current valid shot, one shot counting is performed.

[0013] In some embodiments, the method further comprises: after a set time window after the one shot counting is performed, shot detection is performed again.

[0014] In some embodiments, the human body key point positions include: a wrist key point, an elbow key point, a shoulder key point, a hip joint key point, a knee key point, and an ankle key point; the method further includes: for each valid shot, calculating an elbow angle and a knee angle of the human body target based on the human body key point positions in each video frame involved in the valid shot; wherein the elbow angle refers to the angle of the included angle formed by the shoulder key point, the elbow key point, and the wrist key point; the knee angle refers to the angle of the included angle formed by the hip joint key point, the knee key point, and the ankle key point; based on the elbow angle, the knee angle, the wrist key point position, and the shoulder key point obtained in each video frame, the shooting stage of the human body target in the video frame is determined; the shooting stage includes: a preparation stage and a release stage; for the video frames in which the human body target is in the preparation stage, the index values of at least one first evaluation index are calculated based on the human body key point positions; for the video frames in which the human body target is in the release stage, the index values of at least one second evaluation index are calculated based on the human body key point positions; based on the index values of the at least one first evaluation index of the human body target in each video frame in the preparation stage, a first target video frame is selected from the video frames in which the human body target is in the preparation stage; and based on the index values of the at least one second evaluation index of the human body target in each video frame in the release stage, a second target video frame is selected from the video frames in which the human body target is in the release stage; based on the index values of the first evaluation index in the first target video frame and the corresponding set threshold range, and the index values of the second evaluation index in the second target video frame and the corresponding set threshold range, a shooting posture score result of the human body target in the shot is calculated.

[0015] In some embodiments, the method further includes: inputting the shooting posture score result of the human body target, the index values of the first evaluation index, the index values of the second evaluation index, and the basic information of the human body target into the trained large language model based on the prompt word template, matching N retrieval results closest to the input of the large language model from the second database set by the large language model, and outputting the specific training suggestions of the human body target based on the retrieval results and the shooting posture score result of the human body target, the index values of the first evaluation index, the index values of the second evaluation index, and the basic information of the human body target; wherein the basic information of the human body target at least includes: the identity information, age, and ball age of the human body target.

[0016] In a second aspect, this application provides a basketball shot counting device, comprising: a basketball shot video stream acquisition module for acquiring a basketball shot video stream; a first positioning module for locating human targets and key points in each video frame of the basketball shot video stream, thereby obtaining all human body regions and the positions of key points within each human body region in the video frame; a second positioning module for locating basketball targets in the video frame, thereby obtaining all basketball regions in the video frame; and a third positioning module for locating the hoop and net targets in the video frame, thereby obtaining... The video frame contains all the basket and net areas; a first tracking module is used to track human targets in each video frame based on the human body area in each video frame to obtain at least one human motion trajectory; and a second tracking module is used to track basketball targets in each video frame based on the basketball area in each video frame to obtain at least one basketball motion trajectory; and a shot counting module is used to match a basketball target for each human target in each human motion trajectory, and obtain the shot count result of the human target based on the basketball motion trajectory of the matched basketball target, the basket area and the net area in each video frame.

[0017] In a third aspect, this application provides an electronic device comprising: a processor configured to execute program instructions; and a memory configured to store the program instructions, which, when loaded and executed by the processor, cause the processor to perform a shot counting method according to the first aspect or any optional embodiment of the first aspect.

[0018] In a third aspect, this application provides a computer-readable storage medium storing program instructions, characterized in that, when the program instructions are loaded and executed by a processor, the processor performs the shot counting method according to the first aspect or any optional embodiment of the first aspect.

[0019] By means of the shot counting method, device, electronic equipment and storage medium provided as above, the shot counting method provided by the embodiments of the present application realizes shot counting of training personnel by acquiring a shot video stream and then by computer vision and intelligent analysis algorithm, that is, the human body target, human body key point, basketball target, basketball net and basketball hoop in each video frame are all identified, then the human body target and the basketball target are tracked and matched, and then the shot counting result of the human body target can be obtained according to the basketball movement track of the basketball target matched by each human body target, the basketball hoop region and the basketball net region on each video frame, so as to realize automatic, real-time and non-contact counting of the number of shots, without manual intervention, and the efficiency and accuracy of shot counting are improved; and by matching the human body target and the basketball target, the shots of multiple training personnel can be counted at the same time, and the efficiency is high; and the method only uses a camera device, and the deployment cost is low. BRIEF DESCRIPTION OF DRAWINGS

[0020] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description read in conjunction with the accompanying drawings, in which several embodiments of the present application are shown by way of example, and wherein the same reference numerals identify similar or corresponding elements throughout the several drawings, in which: Figure 1 An exemplary flowchart of a shot counting method of some embodiments of the present application is shown; Figure 2 An exemplary structural block diagram of a shot counting device of some embodiments of the present application is shown; Figure 3 An exemplary structural block diagram of an electronic device of some embodiments of the present application is shown. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0022] It should be understood that the terms "comprise" and "include" used in the specification and claims of the present application indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or sets thereof.

[0023] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and in the claims, the terms "a," "an," and "the" include both singular and plural referents unless the context clearly dictates otherwise. It is further to be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items, and that the term "at least one of" as used herein means one or more.

[0024] As used in this specification and claims, the terms "if' can be construed to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be construed to mean "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]," depending on the context.

[0025] The specific embodiments of the present application will now be described in detail with reference to the accompanying drawings.

[0026] Figure 1 An exemplary flowchart of a shot counting method 100 of some embodiments of the present application is shown. It can be appreciated that the above shot counting method 100 can be performed by any suitable device with data processing capability, such as, but not limited to, a terminal device, a processor, a server, etc.

[0027] In embodiments of the present application, the above shot counting method can be applied to a basketball training and teaching scenario, in which the training court can include a plurality of basketball stands, each including a net and a basket. A camera device can be provided on the basketball training court, specifically, the camera device (e.g., a camera) can be disposed above the training court, at a height of 2m~3m from the ground, and its lens can be disposed at the side of the trainee to ensure that the whole process of the trainee's shot practice (including holding the ball, jumping, shooting, and scoring) can be captured.

[0028] In embodiments of the present application, each camera can count the shots of a plurality of trainees within its scanning range.

[0029] As Figure 1As shown, the shot counting method 100 includes: step S110, acquiring a shot video stream; step S120, for each video frame in the shot video stream, performing region positioning on a human target in the video frame and positioning a human key point, to obtain all human regions in the video frame and the human key point positions in each human region; step S130, performing region positioning on a basketball target in the video frame, to obtain all basketball regions in the video frame; step S140, performing region positioning on a basket target and a net target in the video frame, to obtain all basket regions and net regions in the video frame; step S150, tracking the human target in each video frame based on the human regions in the video frame, to obtain at least one human motion trajectory; step S160, tracking the basketball target in each video frame based on the basketball regions in the video frame, to obtain at least one basketball motion trajectory; and step S170, for the human target in each human motion trajectory, matching a basketball target for the human target, and obtaining a shot counting result of the human target based on the basketball motion trajectory of the matched basketball target, the basket regions and the net regions on each video frame.

[0030] Exemplarily, the shot video stream in step S110 is dynamic video data of a trainee or a player performing a shot practice, which is collected by a camera device. After the camera device collects the shot video stream, the shot video stream can be sent to the device with data processing capability for processing. Specifically, the camera and the device with data processing capability can be connected to the same local area network, and the video stream can be transmitted in real time through the RTSP protocol. In addition, in the embodiment of the present application, the frame rate of the camera device can be set according to the actual processing capability of the device with data processing capability.

[0031] Exemplarily, the video frame in step S120 is all video frames in the shot video stream (for example, when the frame rate of the camera device is 30 fps, 30 target video frames are included in one second), or part of the video frames in the shot video stream (for example, the shot video stream is processed by frame extraction (for example, one frame is extracted every other frame), to obtain the video frame). The embodiment of the present application does not make specific limitation on this. Each video frame includes at least one trainee and at least one basketball.

[0032] The human body target in the step S120 refers to a natural human individual appearing in the target video frame, and the number thereof is at least one. The human body key point refers to an anatomical position that is a landmark of the human body, and specifically can include 21 key points, such as a nose key point, a left eye key point, a right eye key point, a left ear key point, a right ear key point, a left shoulder key point, a right shoulder key point, a left elbow key point, a right elbow key point, a left wrist key point, a right wrist key point, a left hip joint key point, a right hip joint key point, a left knee key point, a right knee key point, a left ankle key point, a right ankle key point, a left toe key point, a right toe key point, a left heel key point, and a right heel key point.

[0033] In the embodiment of the present application, the human body region is a rectangular bounding box region containing the human body target obtained through region positioning, which can be represented as (x1, y1, x2, y2), where x1 represents the horizontal coordinate of the upper left corner point of the human body region in the image coordinate system, y1 represents the vertical coordinate of the upper left corner point of the human body region in the image coordinate system, x2 represents the horizontal coordinate of the lower right corner point of the human body region in the image coordinate system, and y2 represents the vertical coordinate of the lower right corner point of the human body region in the image coordinate system.

[0034] In the embodiment of the present application, the human body key point position can be the center position or the barycenter position of each key point, and the like, which is not specifically limited in the embodiment of the present application and can only represent the human body key point. Specifically, each human body key point position can be represented as (x, y), where x represents the horizontal coordinate of the human body key point in the image coordinate system, and y represents the vertical coordinate of the human body key point in the image coordinate system.

[0035] Here, the direction of the image coordinate system can be defined in advance, for example, the right direction and the downward direction are the positive directions of the two coordinate axes of the image coordinate system.

[0036] In the embodiment of the present application, for each video frame, the human body target in the video frame is subjected to region positioning and the human body key point is positioned, and there can be many methods to obtain all the human body regions in the video frame and the human body key point positions in each human body region, for example, through a trained human body key point detection model. Specifically, the video frame is input into the human body key point detection model, and the human body key point detection model outputs all the human body regions in the video frame and the human body key point positions in each human body region.

[0037] The human body key point detection model is a neural network model trained in advance through a large number of human body images, which can be, for example, a yolov8-pose model. As for the training method, the following embodiments are exemplarily described, which is not described herein.

[0038] Exemplarily, the basketball target in the step S130 refers to the basketball appearing in the shooting video stream, and the number thereof is at least one. The basketball region refers to a rectangular bounding box region containing the basketball target obtained through region positioning, which can be represented as (x3, y3, x4, y4), wherein x3 represents the horizontal coordinate of the top-left corner point of the basketball region in the image coordinate system, y3 represents the vertical coordinate of the top-left corner point of the basketball region in the image coordinate system, x4 represents the horizontal coordinate of the bottom-right corner point of the basketball region in the image coordinate system, and y4 represents the vertical coordinate of the bottom-right corner point of the basketball region in the image coordinate system.

[0039] In the embodiment of the present application, there are many methods for region positioning of the basketball target in the video frame to obtain all the basketball regions in the video frame, for example, through a trained basketball detection model. Specifically, the video frame is input into the above-mentioned basketball detection model, and all the basketball regions in the video frame are output by the basketball detection model.

[0040] The above-mentioned basketball detection model is a neural network model trained in advance through a large number of basketball images, which can be, for example, a YOLOv8 model. As for the training method, the following embodiments are exemplarily described, which will not be described here.

[0041] As a preferred embodiment of the present application, since the proportion of the basketball target in the picture of the video frame is small, usually only a few tens or even a few dozen pixel points, the resolution is low when the YOLOv8 model is used to detect the basketball target, and the feature information of the basketball target is easy to be lost. Based on this, the YOLOv8 model can be improved to improve the detection accuracy of the basketball target.

[0042] Specifically, Dy Sample dynamic upsampling can be introduced into the YOLOv8 model used in the above-mentioned basketball detection model. The Dy Sample dynamic upsampling here is a content-aware dynamic upsampling method, the core idea of which is to generate a sampling offset according to the content of the input feature, which is different from the traditional fixed sampling kernel, so that the target feature can be accurately positioned.

[0043] Specifically, the network structure of the YOLOv8 model can include Backbone (feature extraction), Neck (feature fusion), and Head (detection head) three parts, and the Dy Sample dynamic upsampling is set in the Neck part of the YOLOv8 model to replace the fixed upsampling of the Neck part. Specifically, the input feature is first predicted through a small convolutional layer to obtain a sampling offset, and then the offset is combined with a sampling grid to form a dynamic sampling point set. Compared with the traditional sampling method, Dy Sample can better extract the texture details and edge features of the basketball target.

[0044] The embodiment of the application introduces Dy Sample dynamic up-sampling into the YOLOv8 model. Through perception of input features, the sampling position of small targets such as basketball targets can be dynamically adjusted, the contour information and subtle features of the basketball target are better preserved, and the detailed features of the basketball are restored, thereby improving the basketball detection recall rate.

[0045] Exemplarily, the basketball net target in the step S140 is the basketball net appearing in the shooting video stream, and the basket target is the basket appearing in the shooting video stream. In the embodiment of the application, the number of the basketball net and the number of the basket are both at least one.

[0046] The basket region refers to a rectangular boundary region containing the basket target obtained through region positioning, which can be represented as (x5, y5, x6, y6), wherein x5 represents the horizontal coordinate of the upper left corner point of the basket region in the image coordinate system, y5 represents the vertical coordinate of the upper left corner point of the basket region in the image coordinate system, x6 represents the horizontal coordinate of the lower right corner point of the basket region in the image coordinate system, and y6 represents the vertical coordinate of the lower right corner point of the basket region in the image coordinate system.

[0047] The basketball net region refers to a rectangular boundary region containing the basketball net target obtained through region positioning, which can be represented as (x7, y7, x8, y8), wherein x7 represents the horizontal coordinate of the upper left corner point of the basketball net region in the image coordinate system, y7 represents the vertical coordinate of the upper left corner point of the basketball net region in the image coordinate system, x8 represents the horizontal coordinate of the lower right corner point of the basketball net region in the image coordinate system, and y8 represents the vertical coordinate of the lower right corner point of the basketball net region in the image coordinate system.

[0048] In the embodiment of the application, the basket target and the basketball net target in the video frame are region positioned to obtain all the basket regions and the basketball net regions in the video frame. There can be many methods, for example, through a trained basketball net and basket detection model, specifically, the video frame is input into the above-mentioned basketball net and basket detection model, and all the basket regions and the basketball net regions in the video frame are output by the basketball net and basket detection model.

[0049] The above-mentioned basketball net and basket detection model is a neural network model trained in advance through a large number of basket and basketball net images, which can be a YOLOv8 model, for example. As for the training method, the following embodiments are exemplarily described, which will not be described herein.

[0050] It should be emphasized that the execution order of the above-mentioned steps S120, S130 and S140 is not fixed and can be executed in any order. Of course, it can also be executed synchronously, and the embodiment of the application does not make specific limitation thereon.

[0051] Exemplarily, the human motion trajectory in the step S150 refers to a sequence of position changes of a same human target in continuous video frames, which can be represented by the first identification information and coordinates of the human region in each video frame.

[0052] The basketball motion trajectory in the step S160 refers to a sequence of position changes of a same basketball target in continuous video frames, which can be used for subsequent shot counting analysis. The basketball motion trajectory can be represented by the second identification information and coordinates of the basketball region in each video frame.

[0053] The first identification information refers to information for uniquely identifying a human target, which can have many implementation manners, for example, a digital number, a letter number, a word, a color, and the like. The embodiment of the present application does not make a specific limitation on the first identification information.

[0054] The second identification information refers to information for uniquely identifying a basketball target, which can also have many implementation manners, for example, a digital number, a letter number, a word, a color, and the like. The embodiment of the present application does not make a specific limitation on the second identification information.

[0055] It should be noted that the first identification information of the human region belonging to a same human target is the same, and the first identification information of the human region belonging to different human targets is different. Similarly, the second identification information of the basketball region belonging to a same basketball target is the same, and the second identification information of the basketball region belonging to different basketball targets is different. Moreover, the first identification information and the second identification information can be of the same kind (for example, both are letter numbers), or can be of different kinds (for example, the first identification information is a digital number, and the second identification information is a letter number). Of course, when the first identification information and the second identification information are of the same kind, the first identification information and the second identification information are different.

[0056] In the embodiment of the present application, the method for tracking the human target in each video frame based on the human region in each video frame and the method for tracking the basketball target in each video frame based on the basketball region in each video frame are the same, and can be many, for example, a ByteTrack multi-target tracking algorithm.

[0057] Taking tracking of the basketball target in each video frame as an example, specifically, a Kalman filtering algorithm is used to predict the position of the basketball target in the next video frame according to the position and speed of the basketball target in the current video frame, and a Hungarian algorithm is used to match the position predicted by the last video frame and the basketball region detected in the current video frame. Specifically, the basketball regions detected in the current video frame are sorted in descending order of confidence, the basketball region with high confidence (for example, the confidence is greater than or equal to 85%) is matched first, that is, the basketball region is taken as the position of the basketball target in the current video frame corresponding to the predicted position closest to the basketball region, and is marked as the second identification information; then the basketball region with low confidence (for example, the confidence is less than 85%) is matched, that is, the intersection over union of the basketball region with low confidence and each un-matched predicted position can be calculated, if the intersection over union is greater than a set threshold (for example, 90%), the basketball region is matched and is marked as the second identification information, and the basketball region with low confidence (for example, the confidence is less than 60%) caused by occlusion and the like is removed and is not matched.

[0058] In the embodiment of the present application, when the position predicted by the last video frame does not appear in the current video frame, the predicted position is taken as the position of the basketball target in the current video frame, that is, the trajectory prediction is still maintained, thereby effectively avoiding the problem of tracking error caused by occlusion and the like.

[0059] Exemplarily, after obtaining the multiple human motion trajectories and the multiple basketball motion trajectories in step S170, a basketball target can be matched for each human target in each human motion trajectory, and a corresponding relationship between the human target and the basketball target is established. Taking the first identification information as a letter number and the second identification information as a number as an example, human A-basketball 1, human B-basketball 2, … can be obtained, so as to count the shots of each human target.

[0060] As for the specific matching method, the following embodiments are exemplarily described, which are not described herein.

[0061] In the embodiment of the present application, after obtaining the corresponding relationship between the human target and the basketball target, a shot counter is initialized for each corresponding relationship, and the initial value is 0. When an effective shot is detected, the shot counter is increased by 1. Here, an effective shot refers to the process that the basketball target passes through the basket and the net in turn from above the basket.

[0062] Based on the above description of the effective shot, the relative position relationship between the basketball target and the basket and the net in each video frame can be determined through the basketball motion trajectory and the basket region and the net region in each video frame, and then the effective shot is determined through the relative position relationship in each video frame, and finally the shot count result of the human target is obtained after the motion ends.

[0063] The embodiment of the application realizes counting of the shooting of the training personnel by acquiring a shooting video stream and then through computer vision and intelligent analysis algorithm, that is, the human body target, human body key point, basketball target, basketball net and basket in each video frame are identified, then the human body target and the basketball target are tracked and matched, and then the shooting counting result of the human body target can be obtained according to the basketball movement track of the basketball target matched by each human body target, the basket area and the basketball net area on each video frame, the shooting quantity is counted automatically, in real time and in a non-contact manner, without manual intervention, and the efficiency and accuracy of the shooting counting are improved; and through the matching of the human body target and the basketball target, the shooting of multiple training personnel can be counted at the same time, and the efficiency is high; and only a camera device is used, and the deployment cost is low.

[0064] The training process of the human body key point detection model, the basketball detection model and the basketball net and basket detection model is described as follows: (1) Training data acquisition: The human body image of the training personnel in the shooting scene is acquired by the camera device, and the human body images in multiple scenes are downloaded from the search engine through the web crawler to obtain a first image set, the first image set covers different collection angles, different light conditions and different shooting actions as much as possible; the first image set is used to train the human body key point model; The basketball image of the training personnel in the shooting scene is acquired by the camera device, and the basketball images in multiple scenes are downloaded from the search engine through the web crawler to obtain a second image set, the second image set covers different collection angles, different light conditions and different shooting actions as much as possible; the second image set is used to train the basketball detection model; The basketball net and basket image of the training personnel in the shooting scene is acquired by the camera device, and the basketball net and basket images in multiple scenes are downloaded from the search engine through the web crawler to obtain a third image set, the third image set covers different collection angles, different light conditions and different shooting actions as much as possible; the third image set is used to train the basketball net and basket detection model.

[0065] (2) Data labeling: The human body region and 21 human body key points on each image in the first image set are labeled by using a labeling tool to obtain a first labeled image set; The basketball region on each image in the second image set is labeled by using a labeling tool to obtain a second labeled image set; The basketball net region and the basket region on each image in the third image set are labeled by using a labeling tool to obtain a third labeled image set.

[0066] (3) Model training: The first image set and the first labeled image set are divided into a first training set, a first validation set and a first test set in a ratio of 7:2:1, which are respectively used for model training, model validation and model testing; then the first training set is used to train the yolov8-pose model, and the performance in the model training process is verified based on the first validation set, so that the model is adjusted when the performance does not meet the requirements, and finally a trained human key point detection model is obtained; the generalization performance of the trained human key point detection model is tested based on the first test set.

[0067] The second image set and the second labeled image set are also divided into a second training set, a second validation set and a second test set in a ratio of 7:2:1, which are respectively used for model training, model validation and model testing; then the second training set is used to train the yolov8 model with Dy Sample dynamic up-sampling, and the performance in the model training process is verified based on the second validation set, so that the model is adjusted when the performance does not meet the requirements, and finally a trained basketball detection model is obtained; the generalization performance of the trained basketball detection model is tested based on the second test set.

[0068] The third image set and the third labeled image set are also divided into a third training set, a third validation set and a third test set in a ratio of 7:2:1, which are respectively used for model training, model validation and model testing; then the third training set is used to train the yolov8 model, and the performance in the model training process is verified based on the third validation set, so that the model is adjusted when the performance does not meet the requirements, and finally a trained basketball net and basket detection model is obtained; the generalization performance of the trained basketball net and basket detection model is tested based on the third test set.

[0069] At this point, the training of the human key point detection model, the basketball detection model and the basketball net and basket detection model is completed.

[0070] As an optional embodiment of the present application, before matching a basketball target for each human target in each human motion trajectory, the above shot counting method 100 further comprises: selecting a to-be-identified video frame from all video frames for each human target in each human motion trajectory; wherein the face of the human target in the to-be-identified video frame is complete and clear; cropping the face region from the human region in the to-be-identified video frame; performing feature extraction on the face region, and comparing the extracted target features with each face feature stored in the first database to obtain the identity information of the human target.

[0071] Exemplarily, the face of the human target in the to-be-identified video frame is complete and clear, that is, the face of the human target is not occluded, and the image clarity satisfies a set clarity condition. For the human target in each human motion trajectory, one video frame in which the face of the human target is clear and complete is selected from all the video frames, so as to avoid the failure of human target identity recognition due to the occlusion of the face, image blur, and the like.

[0072] Here, the completeness of the face in each video frame can be determined by a face detection algorithm, and then the clarity of each video frame can be calculated by a clarity detection algorithm. Then, the face completeness and the image clarity of each video frame are weighted and averaged, and the video frame with the highest weighted average value is taken as the to-be-identified video frame.

[0073] Exemplarily, the first database is a database in which the basic information of the students and the corresponding face features are stored in advance. Specifically, in the training and teaching scene, the trainer manually inputs the basic information of all the students on the front-end interface (i.e., the display interface of the electronic device) in advance. The basic information can include but is not limited to identity information (i.e., name), age, ball age, height, weight, and contact information, and the like. A student's front face photo (with clear features, no occlusion, and uniform illumination) or a face photo captured from a camera device is uploaded.

[0074] After obtaining the student's front face photo, feature extraction can be performed on the front face photo. For example, a trained face recognition model is used for feature extraction, and the face features are stored in the above-mentioned database together, which can be used for subsequent identity recognition.

[0075] In the embodiment of the present application, after obtaining the to-be-identified video frame, the human region in the to-be-identified video frame can be input into the face recognition model, and then the face region is located in the human region and cropped out. Then, feature extraction is performed on the face region, and the extracted target feature is compared with each face feature stored in the first database. Specifically, the cosine similarity of the target feature and each face feature in the first database can be calculated, and the identity information of the human target corresponding to the face feature with the highest cosine similarity is taken as the identity information of the human target.

[0076] The embodiment of the present application combines the first database to perform identity recognition on the human target in each human motion trajectory, which can accurately and independently count the number of shots of each human target. Moreover, the video frame with clear and unoccluded face is processed, which improves the accuracy of identity recognition.

[0077] As an optional embodiment of the present application, the human body key point comprises a wrist key point; for each human body target in the human body motion track, a basketball target is matched for the human body target, comprising: for each human body target in the human body motion track, a to-be-detected video frame of the human body target standing with a basketball is acquired in the shooting video stream; and the wrist key point position of the human body target in the to-be-detected video frame and each basketball region are used to determine the basketball target matched for the human body target.

[0078] Exemplarily, the wrist key point can be a left wrist key point of the human body target or a right wrist key point of the human body target, and the embodiments of the present application are not limited in this regard. The embodiments of the present application are described by taking the right wrist key point of the human body target as an example.

[0079] In the embodiments of the present application, the to-be-detected video frame refers to one of the static video frames in which the human body target is in a standing state and holds a basketball, and is a key input frame for basketball target matching.

[0080] In the embodiments of the present application, for each human body target in the human body motion track, the to-be-detected video frame of the human body target standing with a basketball can be acquired from the shooting video stream, and then the wrist key point position in the to-be-detected video frame and the position relationship of each basketball region are used to determine the basketball target matched for the human body target. Specifically, the distance (for example, the position of the center point of each basketball region) from the wrist key point position to each basketball region can be calculated, and then the basketball target of the basketball region closest to the wrist key point position is taken as the basketball target matched for the human body target.

[0081] The embodiments of the present application can match one basketball target for each human body target by using the close position relationship between the wrist key point of the training personnel in the ball holding stage and the basketball, and through distance screening, the association between each human body target and the corresponding basketball can be accurately established in the scene in which multiple training personnel hold balls at the same time, and matching confusion can be avoided.

[0082] The embodiments of the present application can use various verification methods to count the number of shots, including but not limited to state machine verification, trajectory analysis verification, and basketball net shaking auxiliary verification, etc. The various verification methods are exemplarily described in the following embodiments, and the specific embodiments are as follows: As an optional embodiment of the present application, the shooting count result of the human target is obtained based on the matched basketball movement track of the basketball target, the basket area on each video frame, and the net area, and includes: for each basketball area in the matched basketball movement track of the basketball target, determining the movement state of the basketball target in the basketball area based on the relative position of the basketball area and the basket area and the net area on the video frame where the basketball area is located; the movement state includes: a pre-entry state, an entry state, and a post-entry state; when the state transition process of a valid shot is determined based on the movement state of each basketball target, a shot count is performed to obtain the shooting count result of the human target; wherein the state transition process of a valid shot refers to the state transition process of the basketball target from the pre-entry state, the entry state to the post-entry state.

[0083] Exemplarily, in the embodiment of the present application, a state machine can be established based on the position of the basketball relative to the net and the position of the net. Specifically, the above-mentioned movement state can include: a pre-entry state, an entry state, and a post-entry state. Among them, the pre-entry state is the initial state of the basketball target; when the basketball target appears above the basket area below a set pixel height (for example, 6% of the height of the video frame), the entry state is entered; when the basketball target appears in the net or below the basket, the post-entry state is entered.

[0084] Based on the above description, the state transition process of a valid shot can be the state transition process of the basketball target from the pre-entry state, the entry state to the post-entry state, that is, the process of the basketball target sequentially passing through the basket and the net from above the basket.

[0085] In the embodiment of the present application, for each basketball area in the matched basketball movement track of the basketball target, the movement state of the basketball target in the basketball area is determined based on the relative position of the basketball area and the basket area and the net area on the video frame where the basketball area is located. Specifically, the vertical coordinates of the first set point (for example, the center point of the bottom of the basketball area) of the basketball target area, the vertical coordinates of the second set point (for example, the center point of the top of the basket area) of the basket area, and the vertical coordinates of the third set point (for example, the center point of the top of the net area) of the net area can be calculated to determine the movement state of the basketball target in the basketball area.

[0086] The embodiment of the present application also sets a frame number timeout anti-stuck mechanism, that is, if the basketball target does not enter the next movement state for a certain number of frames, the movement state of the basketball target is reset to the pre-entry state. The certain number of frames can be a reasonable value set in advance, for example, 15 frames, 10 frames, etc., which is not limited in the embodiment of the present application.

[0087] The embodiment of the present application sets a frame number timeout anti-stuck mechanism to avoid the state from being stuck due to the basketball target bouncing after hitting the basket.

[0088] After determining the movement state of each basketball target on each video frame, a shot count can be performed when determining the state transition process (i.e., pre-entry state-entry state-post-entry state) of an effective shot based on the movement state of each basketball target, and after the shot count is performed, the state of the basketball target is reset to the pre-entry state, until the training is completed, and the shot count result of the human target can be obtained.

[0089] By judging the movement state of the basketball target and then performing shot counting based on the movement state of the basketball target, the application embodiment can effectively exclude invalid conditions such as rebound after the basketball hits the basket, does not enter the net, and passes the basket, and reduce the misjudgment rate of counting.

[0090] As an optional embodiment of the application, the shot counting method 100 further includes: performing difference calculation on the vertical coordinates of the set points of the basketball region in the continuous multiple video frames based on the basketball movement trajectory of the matched basketball target, to obtain a difference sequence including multiple difference values; when determining the state transition process of an effective shot based on the movement state of each basketball target, determining whether the movement direction of the basketball target in the state transition process of the effective shot is from top to bottom and whether the basketball movement trajectory of the matched basketball target passes through the basket region first and then passes through the net region based on the difference values involved in the state transition process of the effective shot in the video frame difference sequence; performing a shot count, including: if the movement direction of the basketball target in the state transition process of the effective shot is from top to bottom and the basketball movement trajectory of the matched basketball target passes through the basket region first and then passes through the net region, performing a shot count.

[0091] Exemplarily, the set point of the basketball region in the video frame refers to the first set point described above. In the application embodiment, the difference calculation on the vertical coordinates of the set points of the basketball region in the continuous multiple video frames based on the basketball movement trajectory of the matched basketball target to obtain a difference sequence including multiple difference values can be specifically calculated by the following formula: Formula 1 Wherein, represents the tth vertical coordinate; represents the (t-1)th vertical coordinate; represents the difference value of the tth vertical coordinate and the (t-1)th vertical coordinate.

[0092] In the application embodiment, , it indicates that the basketball target is falling, When the basketball target is rising, the movement direction of the basketball target is from top to bottom. Therefore, when the state transition process of the current valid shot is determined based on the movement state of each basketball target, whether the movement direction of the basketball target in the state transition process of the current valid shot is from top to bottom can be determined based on the difference values involved in the state transition process of the current valid shot in the video frame difference sequence. That is, when all the difference values involved in the state transition process of the current valid shot are greater than 0, it can be determined that the movement direction of the basketball target in the state transition process of the current valid shot is from top to bottom.

[0093] In the embodiment of the present application, when the state transition process of the current valid shot is determined based on the movement state of each basketball target, whether the basketball movement trajectory of the matched basketball target first passes through the basket area and then passes through the net area. Here, when determining the passing order of the basketball movement trajectory, the frame in which the basketball movement trajectory first overlaps with the basket area and the net area can be extracted. The frame in which the basketball trajectory first overlaps with the basket area is recorded as the first frame, and the frame in which the basketball trajectory first overlaps with the net area is recorded as the second frame. If the first frame is earlier than the second frame, it is determined that the basketball movement trajectory first passes through the basket area and then passes through the net area. If the first frame is later than the second frame, it is determined that the basketball movement trajectory does not first pass through the basket area and then pass through the net area.

[0094] In the embodiment of the present application, when the state transition process of the current valid shot is determined based on the movement state of each basketball target, if it is further determined that the movement direction of the basketball target in the state transition process of each valid shot is from top to bottom, and the basketball movement trajectory of the matched basketball target first passes through the basket area and then passes through the net area, then the shot count is performed.

[0095] When the shot count is performed, the movement direction of the basketball target and the passing order of the basketball movement trajectory are also verified to exclude non-shot actions such as pass from bottom to top and rebound, so that the shot count result is more accurate.

[0096] As an optional embodiment of the present application, the shot counting method 100 further includes: when the state transition process of the current valid shot is determined based on the movement state of each basketball target, determining whether the net shakes in the state transition process of the current valid shot based on the net area in the video frame involved in the state transition process of the current valid shot; and performing a shot count, including: if the net shakes in the state transition process of the current valid shot, performing a shot count.

[0097] Exemplarily, the embodiment of the present application can also determine whether the net shakes during the state transition process of the current valid shot based on the motion state of each basketball goal. In the embodiment of the present application, when determining whether the net shakes during the state transition process of the current valid shot, the change of the center point of the net area can be used to determine whether the net shakes during the state transition process of the current valid shot. Specifically, the distance between the center point of the net area in the last video frame of the state transition process of the current valid shot and the average center point of the net area in N (for example, 10) historical video frames can be calculated. If the distance is greater than or equal to a set distance (for example, 5), it is considered that the net shakes during the state transition process of the current valid shot. If the distance is less than the set distance, it is considered that the net does not shake during the state transition process of the current valid shot.

[0098] In the embodiment of the present application, if the net shakes during the state transition process of the current valid shot, a shot count is performed.

[0099] The embodiment of the present application can further improve the accuracy of shot counting by using the net shaking to assist in verifying whether to perform shot counting.

[0100] As an optional embodiment of the present application, the shot counting method 100 further includes: after a set time window after the shot count is performed, shot detection is performed again.

[0101] Exemplarily, the set time window can be a time period set in advance, for example, 1s, 2s. In the embodiment of the present application, no shot counting is performed in the set time window, and shot detection is performed again after the set time window after the shot count is performed, which can prevent repeated counting caused by the bouncing of the basketball in the basket and the net, and has strong anti-interference ability.

[0102] At this point, the description of the shot counting is completed. The embodiment of the present application verifies a valid shot by using multiple methods such as state machine verification, trajectory analysis verification, and net shaking assisted verification. Compared with a single detection method, the shot recognition accuracy can be effectively improved, and the method can better adapt to scenes such as light changes and occlusions.

[0103] In addition to counting the number of shots, the embodiment of the present application can also score the shooting posture of the training personnel, so that the training of the training personnel can be optimized based on the shooting posture score.

[0104] As an optional embodiment of the present application, the human body key point position includes a wrist key point, an elbow key point, a shoulder key point, a hip joint key point, a knee key point, and an ankle key point; the above-mentioned shooting counting method 100 further includes: for each valid shooting, respectively calculating the elbow angle and the knee angle of the human body target based on the human body key point position in each video frame involved in the valid shooting; wherein the elbow angle refers to the angle of the included angle formed by the shoulder key point, the elbow key point, and the wrist key point; the knee angle refers to the angle of the included angle formed by the hip joint key point, the knee key point, and the ankle key point; based on the elbow angle, the knee angle, the wrist key point position, and the shoulder key point obtained from each video frame, the shooting stage of the human body target in the video frame is determined; the shooting stage includes a preparation stage and a release stage; for the video frame in which the human body target is in the preparation stage, the index value of at least one first evaluation index is calculated based on the human body key point position; for the video frame in which the human body target is in the release stage, the index value of at least one second evaluation index is calculated based on the human body key point position; based on the index value of at least one first evaluation index of the human body target on each video frame in the preparation stage, the first target video frame is selected from the video frames in which the human body target is in the preparation stage; and based on the index value of at least one second evaluation index of the human body target on each video frame in the release stage, the second target video frame is selected from the video frames in which the human body target is in the release stage; based on the index value of each first evaluation index in the first target video frame and the corresponding set threshold range, and the index value of each second evaluation index in the second target video frame and the corresponding set threshold range, the shooting posture score result of the human body target in the shooting is calculated.

[0105] Exemplarily, the elbow angle of the human body target refers to the angle of the included angle formed by the shoulder key point, the elbow key point, and the wrist key point, for example, the angle of the included angle formed by the left shoulder key point, the left elbow key point, and the left wrist key point, for another example, the angle of the included angle formed by the right shoulder key point, the right elbow key point, and the right wrist key point, the present application embodiment does not make specific limitation, and the present application embodiment only takes the elbow angle as an example to describe that it refers to the angle of the included angle formed by the right shoulder key point, the right elbow key point, and the right wrist key point.

[0106] The above-mentioned knee angle refers to the angle of the included angle formed by the hip joint key point, the knee key point, and the ankle key point, for example, the angle of the included angle formed by the left hip joint key point, the left knee key point, and the left ankle key point, for another example, the angle of the included angle formed by the right hip joint key point, the right knee key point, and the right ankle key point; the present application embodiment does not make specific limitation, and the present application embodiment only takes the knee angle as an example to describe that it refers to the angle of the included angle formed by the right hip joint key point, the right knee key point, and the right ankle key point.

[0107] In the embodiments of the present application, for each video frame involved in each valid shot, the elbow angle and the knee angle of the human target are calculated based on the human key point positions.

[0108] Exemplarily, the shot stage includes a preparation stage and a release stage. The preparation stage refers to the power accumulation and aiming stage of the shot, and the release stage refers to the power release stage.

[0109] In the embodiments of the present application, after the elbow angle and the knee angle of the human target in each video frame are calculated, the shot stage of the human target in the video frame can be determined based on the elbow angle, the knee angle, the wrist key point (for example, the right wrist key point) and the shoulder key point (for example, the right shoulder key point). Specifically, the elbow angle threshold (for example, 160 o ) and the knee angle threshold (for example, 170 o ) can be set in advance, and then the elbow angle and the elbow angle threshold and the knee angle and the knee angle threshold are compared respectively. When the elbow angle is less than 160 o and the knee angle is less than 170 o , it is determined that the shot stage of the human target in the video frame is the preparation stage. When the elbow angle is greater than 160 o and the vertical coordinate of the right wrist key point is less than the vertical coordinate of the right shoulder key point, it is determined that the shot stage of the human target in the video frame is the release stage. Other cases do not belong to the preparation stage or the release stage.

[0110] After determining the shot stage of the human target in each video frame, different processing is performed on the video frames in different shot stages, and one video frame (denoted as a first target video frame (with the optimal elbow angle) and a second target video frame (with the optimal knee angle)) is selected for shot posture scoring of the human target.

[0111] Exemplarily, the first scoring indicator can be at least one of the knee angle, the angle between the upper arm and the body, the elbow position, the elbow angle, the trunk alignment and the footstep occupation standardization. The second scoring indicator can be at least one of the elbow angle and the shooting point position.

[0112] It should be noted in advance that all the parts mentioned below to distinguish between left and right key points are described by taking the right key point as an example, and for some scoring indicators, a set threshold range can be set to evaluate the standard degree of the shot posture. Based on this, the index values of the at least one first scoring indicator and the at least one second scoring indicator based on the human key point positions can be calculated by the following method: Knee angle: calculate the angle of the included angle formed by the right hip joint key point, the right knee key point and the right ankle key point; the threshold range is set to be 130 o ~160 o , and the closer to 145 o , the higher the score of this item; Angle between upper arm and body: calculate the angle of the included angle formed by the right elbow key point, the right shoulder key point and the right hip key point; the threshold range is set to be 80 o ~100 o , and the closer to 90 o , the higher the score of this item, which ensures the coordination between the shooting arm and the trunk; Elbow position: calculate the midpoint between the right shoulder key point and the right hip joint key point, and calculate the difference between the horizontal coordinates of the midpoint and the right elbow key point; the threshold range can be 0-30 pixels, and the closer to 0, the higher the score of this item; Elbow angle: calculate the angle of the included angle formed by the right shoulder key point, the right elbow key point and the right wrist key point; the threshold range can be 80 o ~100 o , and the closer to 90 o , the higher the score of this item; Trunk alignment: calculate the difference between the horizontal coordinates of the right wrist key point and the right elbow key point, the difference between the horizontal coordinates of the right elbow key point and the right knee key point, and the difference between the horizontal coordinates of the right knee key point and the right toe key point, and the threshold range is less than 50 pixels. If all the above differences are less than 50 pixels, the trunk alignment can be scored, which ensures the stability of the body center of gravity; Foot standing specification: the left foot key point (left toe key point) is in front, the right foot key point (for example, the right toe key point) is in back, and the vertical coordinate of the left foot key point should be less than or equal to the vertical coordinate of the right foot key point + 20 pixels (i.e. the threshold range), and the horizontal distance between the left and right toe key points should be between 0.5 times and 1.5 times the horizontal distance between the left and right shoulder key points (i.e. the threshold range).

[0113] Shooting point position: compare the relative height of the right wrist key point and the nose key point. If the vertical coordinate of the right wrist key point is higher than the vertical coordinate of the nose key point, it can be scored, which ensures the rationality of the shooting arc and the shooting speed.

[0114] In the embodiments of the present application, after the index values of at least one first score index of the human target in each video frame in the preparation stage are calculated, the deviations of each first score index can be determined based on the set threshold range corresponding to each first score index, and then the video frame with the smallest sum of the first deviations of all first score indexes is determined as the first target video frame.

[0115] Similarly, after the index values of the at least one second score indicator of the human target in each video frame of the release stage are calculated, the deviations of each second score indicator can be determined based on the set threshold range corresponding to each second score indicator, and then the video frame with the minimum sum of second deviations of all second score indicators is determined as the second target video frame.

[0116] In this way, the shooting posture score result of the human target in the current shooting can be calculated based on the index values of the first score indicators in the first target video frame and the set threshold range corresponding thereto (i.e., the minimum first deviation sum), and the index values of the second score indicators in the second target video frame and the set threshold range corresponding thereto (i.e., the minimum second deviation sum), for example, the sum of the minimum first deviation sum and the minimum second deviation sum.

[0117] The embodiments of the present application score the shooting posture of the human target by multiple indicators, making the scoring more comprehensive, and subdividing the shooting stage into the preparation stage and the release stage for evaluation, which can help the trainee better locate the problem.

[0118] The embodiments of the present application can also store (e.g., store in the first database) the shooting count result, the hit rate, the posture score result, and the key frame (i.e., the first target video frame and the second target video frame) of each training of each training personnel, so as to facilitate subsequent analysis and viewing.

[0119] As an optional embodiment of the present application, the shooting counting method 100 further comprises: inputting the shooting posture score result of the human target, the index values of the first score indicators, the index values of the second score indicators, and the basic information of the human target into the trained large language model based on the prompt word template, matching N retrieval results closest to the input of the video frame large language model from the set second database by the video frame large language model, and outputting the targeted training suggestion of the human target based on the video frame retrieval result and the shooting posture score result of the human target, the index values of the first score indicators, the index values of the second score indicators, and the basic information of the human target; wherein the basic information of the human target at least includes: the identity information, age, and ball age of the human target.

[0120] Exemplarily, the prompt word template is a template set in advance for input into the large language model to obtain the training suggestion, which may, for example, include but is not limited to: the shooting posture score result of the human target, the index values of the first score indicators of the video frame, the index values of the second score indicators of the video frame, and the basic information of the human target, etc. The basic information of the human target here may include but is not limited to the identity information, age, and ball age of the human target.

[0121] As one specific embodiment of the present application, the above prompt word template can be, for example: Student {identity information}, age {X} years old, ball age {Y} years, this time the shooting posture score is {Z} points. First scoring index: knee bending angle {K1}°, arm and body angle {K2}°, elbow position {K3}°, elbow angle {K4}°, trunk alignment pixel difference {K5} pixels, foot standing level distance {K6} pixels / front and back distance {K7} pixels; second scoring index: elbow angle {K8}°, shooting point position difference {K9} pixels. Please analyze the shooting posture problems of this student in combination with the professional knowledge of basketball teaching, and give personalized training suggestions.

[0122] The above-mentioned second database is a vector database that stores professional knowledge of basketball (including technical essentials (i.e., standard motion descriptions of 8 scoring indexes), common error analysis (for example, insufficient knee bending, too low shooting point, etc.), improvement methods (for example, half squat jumping exercise, weight-bearing knee bending exercise, etc.), etc.) set in advance. The content in the second database can come from the experience summary of coaches, literature related to basketball sports, and authoritative basketball teaching theory, etc., facilitating subsequent retrieval matching.

[0123] In the embodiments of the present application, retrieval generation technology can be used for matching in the second database. The retrieval generation technology here refers to a technical framework that combines information retrieval and large language model (LLM), and the core idea is "let the large language model retrieve first and then generate". By deeply integrating external knowledge base with LLM, the static knowledge boundary of the model is broken through, and dynamic knowledge updating and accurate content generation are realized.

[0124] Specifically, the shooting posture score result of the human body target, the index value of the first scoring index of the video frame, the index value of the second scoring index of the video frame, and the basic information of the human body target can be input into the trained large language model, and the large language model can match N retrieval results closest to the input of the large language model from the second database. Specifically, the similarity between the input vector and the vector in the database can be calculated by cosine similarity to determine the N retrieval results closest to the input of the large language model. Then the large language model outputs the targeted training suggestions for the human body target based on the video frame retrieval result and the shooting posture score result of the human body target, the index value of the first scoring index, the index value of the second scoring index, and the basic information of the human body target.

[0125] The embodiments of the present application generate training suggestions suitable for the technical level and physical condition of the training personnel by combining the scoring index, age, ball age, etc. of the training personnel, and ensure the authority and scientificity of the suggestions based on the professional basketball teaching corpus through retrieval enhancement generation technology, avoiding non-professional guidance.

[0126] Of course, for each trainer, a comprehensive training report can also be generated containing technical analysis, progress curve and suggestions for improvement.

[0127] Figure 2 An exemplary structural block diagram of the shot counting device is shown.

[0128] As Figure 2 shown, the shot counting device 200 includes: a shot video stream acquisition module 210, configured to acquire a shot video stream; a first positioning module 220, configured to, for each video frame in the shot video stream, perform region positioning on a human target in the video frame and perform positioning on human key points to obtain all human regions in the video frame and human key point positions in each human region; a second positioning module 230, configured to perform region positioning on a basketball target in the video frame to obtain all basketball regions in the video frame; and a third positioning module 240, configured to perform region positioning on a basket target and a net target in the video frame to obtain all basket regions and net regions in the video frame; a first tracking module 250, configured to track the human target in each video frame based on the human regions in the video frame to obtain at least one human motion trajectory; and a second tracking module 260, configured to track the basketball target in each video frame based on the basketball regions in the video frame to obtain at least one basketball motion trajectory; and a shot counting module 270, configured to, for the human target in each human motion trajectory, match one basketball target for the human target, and obtain a shot counting result of the human target based on the basketball motion trajectory of the matched basketball target, the basket regions and the net regions on each video frame.

[0129] As an optional embodiment of the present application, the above-mentioned shot counting device 200 satisfies at least one of the following: the first positioning module 220 is specifically configured to perform region positioning on the human target in the video frame and perform positioning on human key points by using a human key point detection model that has been trained to obtain all human regions in the video frame and human key point positions in each human region; the second positioning module 230 is specifically configured to perform region positioning on the basketball target in the video frame to obtain all basketball regions in the video frame, including: performing region positioning on the basketball target in the video frame by using a basketball detection model that has been trained to obtain all basketball regions in the video frame; and the third positioning module 240 is specifically configured to perform region positioning on the basket target and the net target in the video frame by using a basket-net detection model that has been trained to obtain all basket regions and net regions in the video frame.

[0130] As an optional embodiment of the present application, the shot counting device 200 satisfies at least one of the following: the human key point detection model adopts a YOLOv8-pose model; the basketball detection model adopts a YOLOv8 model, and the video frame basketball detection model adopts a YOLOv8 model using Dy Sample dynamic up-sampling; and the basket and net detection model adopts a YOLOv8 model.

[0131] As an optional embodiment of the present application, the shot counting device 200 further includes: a to-be-identified video frame selection module configured to select a to-be-identified video frame from all video frames for each human target in a human motion trajectory; wherein the face of the human target in the to-be-identified video frame is complete and clear; a cropping module configured to crop a face region from a human region in the to-be-identified video frame; and an identity recognition module configured to extract features of the face region and compare the extracted target features with each face feature stored in a first database to obtain identity information of the human target.

[0132] As an optional embodiment of the present application, the human key points include wrist key points, and the shot counting module 270 is specifically configured to: for each human target in a human motion trajectory, obtain a to-be-detected video frame of the human target standing with a ball in a shot video stream; and determine a basketball target matched with the human target according to the position of the wrist key points of the human target in the to-be-detected video frame and each basketball region.

[0133] As an optional embodiment of the present application, the shot counting module 270 is further configured to: for each basketball region in a basketball motion trajectory of the matched basketball target, determine a motion state of the basketball target in the basketball region based on the relative positions of the basketball region, a basket region and a net region in a video frame in which the basketball region is located; the video frame motion state includes a pre-entry state, an entry state and a post-entry state; when a state transition process of a valid shot is determined based on the motion state of each basketball target, a shot count is performed to obtain a shot count result of the human target; wherein the state transition process of a valid shot refers to a state transition process of the basketball target from the pre-entry state, the entry state to the post-entry state.

[0134] As an optional embodiment of the present application, the shot counting device 200 further comprises a difference calculation module configured to calculate the vertical coordinates of the set points of the basketball region in the continuous video frames based on the basketball movement track of the matched basketball target, to obtain a difference sequence comprising a plurality of difference values; a judgment module configured to determine whether the movement direction of the basketball target in the state transition process of the current valid shot is from top to bottom and whether the basketball movement track of the matched basketball target passes through the basket region first and then the net region based on the difference values involved in the state transition process of the current valid shot in the video frame difference sequence when the state transition process of the one valid shot is determined based on the movement state of each basketball target; and the shot counting module 270 is further configured to count one shot if the movement direction of the basketball target in the state transition process of the current valid shot is from top to bottom and the basketball movement track of the matched basketball target passes through the basket region first and then the net region.

[0135] As an optional embodiment of the present application, the shot counting device 200 further comprises a shaking determination module configured to determine whether the net shakes in the state transition process of the current valid shot based on the net region in the video frames involved in the state transition process of the current valid shot when the state transition process of the one valid shot is determined based on the movement state of each basketball target; and the shot counting module 270 is further configured to count one shot if the net shakes in the state transition process of the current valid shot.

[0136] As an optional embodiment of the present application, the shot counting device 200 further comprises a cooling module configured to perform shot detection again after a set time window after one shot is counted.

[0137] As an optional embodiment of the present application, the human body key point position includes: a wrist key point, an elbow key point, a shoulder key point, a hip joint key point, a knee key point, and an ankle key point; the shot counting device 200 further includes: an angle calculation module, configured to calculate the elbow angle and the knee angle of the human body target based on the human body key point position in each video frame involved in each valid shot; wherein the elbow angle refers to the angle of the included angle formed by the shoulder key point, the elbow key point, and the wrist key point; the knee angle refers to the angle of the included angle formed by the hip joint key point, the knee key point, and the ankle key point; a shot phase determination module, configured to determine the shot phase of the human body target in each video frame based on the elbow angle, the knee angle, the wrist key point position, and the shoulder key point obtained in the video frame; the shot phase includes: a preparation phase and a release phase; a first index value calculation module, configured to calculate the index value of at least one first scoring index based on the human body key point position for the video frame in which the human body target is in the preparation phase; a second index value calculation module, configured to calculate the index value of at least one second scoring index based on the human body key point position for the video frame in which the human body target is in the release phase; a first target video frame selection module, configured to select a first target video frame from the video frames in which the human body target is in the preparation phase based on the index value of at least one first scoring index of the human body target in each video frame in the preparation phase; and a second target video frame selection module, configured to select a second target video frame from the video frames in which the human body target is in the release phase based on the index value of at least one second scoring index of the human body target in each video frame in the release phase; a shot scoring module, configured to calculate the shot posture score result of the human body target in the current shot based on the index value of each first scoring index and the corresponding set threshold range in the first target video frame, and the index value of each second scoring index and the corresponding set threshold range in the second target video frame.

[0138] As an optional embodiment of the present application, the shot counting device 200 further includes: a training suggestion generation module, configured to input the shot posture score result of the human body target, the index value of the first scoring index, the index value of the second scoring index, and the basic information of the human body target into the trained large language model based on the prompt word template, match N retrieval results closest to the input of the video frame large language model from the second database set based on the large language model, and output the targeted training suggestion of the human body target based on the video frame retrieval result and the shot posture score result of the human body target, the index value of the first scoring index, the index value of the second scoring index, and the basic information of the human body target; wherein the basic information of the human body target at least includes: the identity information, age, and ball age of the human body target.

[0139] Thus, the description of the shot counting device 200 shown in the figure is completed. Figure 2 The description of the shot counting device 200 shown in the figure is completed.

[0140] Correspondingly, the embodiments of the present application also provide a device for implementing the above-mentioned shot counting method 100. Figure 2 The hardware structure diagram of the device is shown in FIG. 3, and the device is specifically as follows. Figure 3 As shown in FIG. 3, the electronic device 300 can be a device for implementing the above-mentioned shot counting method 100. As shown in FIG. 3, the electronic device 300 includes a processor 310 and a memory 320. Wherein the memory 320 is configured to store program instructions; the processor 310 is configured to load and execute the program instructions stored in the memory 320, so as to implement the embodiments of the corresponding shot counting method 100 as shown above. Figure 3 As shown in FIG. 3, the electronic device 300 includes a processor 310 and a memory 320. Wherein the memory 320 is configured to store program instructions; the processor 310 is configured to load and execute the program instructions stored in the memory 320, so as to implement the embodiments of the corresponding shot counting method 100 as shown above.

[0141] As an embodiment, the memory 320 can be any electronic, magnetic, optical or other physical storage device, and can contain or store information such as program instructions, data, etc. For example, the memory 320 can be a volatile memory, a non-volatile memory or a similar storage medium. Specifically, the memory 320 can be a RAM (Radom Access Memory, Random Access Memory), a flash memory, a storage drive (such as a hard disk drive), a solid state disk, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof.

[0142] So far, the description of the electronic device is completed. Figure 3 So far, the description of the electronic device is completed.

[0143] Although the embodiments of the present application have been shown and described herein, it is obvious to those skilled in the art that such embodiments are provided only in an exemplary manner. Those skilled in the art can think of many changes, changes and alternative ways without departing from the idea and spirit of the present application. It should be understood that various alternatives to the embodiments of the present application described herein can be employed in practicing the present application. The appended claims are intended to define the scope of protection of the present application and thus cover the equivalents or alternatives within the scope of the claims.

Claims

1. A method for counting basketball shots, characterized in that, include: Obtain the shooting video stream; For each video frame in the basketball shooting video stream, the human target in the video frame is located in the region and the key points of the human body are located to obtain all human body regions in the video frame and the positions of the key points of the human body in each human body region. as well as Localize the basketball targets in the video frame to obtain all basketball regions in the video frame; and Region localization is performed on the basket and net targets in the video frame to obtain all basket and net regions in the video frame; Based on the human body region in each video frame, the human target in each video frame is tracked to obtain at least one human motion trajectory. as well as Based on the basketball region in each video frame, the basketball target in each video frame is tracked to obtain at least one basketball trajectory. For each human target in the human motion trajectory, a basketball target is matched for that human target, and the shot count result of that human target is obtained based on the basketball motion trajectory of the matched basketball target, the basket area and net area on each video frame.

2. The method according to claim 1, characterized in that, The method satisfies at least one of the following: The step of locating human targets in the video frame and locating human key points to obtain all human regions in the video frame and the positions of human key points in each human region includes: locating human targets in the video frame and locating human key points using a pre-trained human key point detection model to obtain all human regions in the video frame and the positions of human key points in each human region. The step of locating the basketball target in the video frame to obtain all basketball regions in the video frame includes: locating the basketball target in the video frame using a trained basketball detection model to obtain all basketball regions in the video frame. The step of locating the basket and net targets in the video frame to obtain all basket and net regions in the video frame includes: using a trained basket and net detection model to locate the basket and net targets in the video frame to obtain all basket and net regions in the video frame.

3. The method according to claim 1 or 2, characterized in that, The method satisfies at least one of the following: The human body key point detection model adopts the YOLOv8-pose model; The basketball detection model uses the YOLOv8 model, and the YOLOv8 model used in the basketball detection model uses DySample dynamic upsampling. The basketball net and basket detection model uses the YOLOv8 model.

4. The method according to claim 1, characterized in that, Before matching a basketball target to a human target in each human motion trajectory, the method further includes: For each human target in the human motion trajectory, one video frame to be identified is selected from all video frames; wherein the face of the human target in the video frame to be identified is complete and clear; The face region is cropped out from the human body region in the video frame to be identified; Feature extraction is performed on the face region, and the extracted target features are compared with the face features stored in the first database to obtain the identity information of the human target.

5. The method according to claim 1, characterized in that, The key points of the human body include: wrist key points; the matching of a basketball target to a human target in each human movement trajectory includes: For each human target in the human motion trajectory, obtain the video frame to be detected when the human target is standing with the ball in the basketball shooting video stream; Based on the location of the key points of the human target's wrist and the basketball areas in the video frame to be detected, a basketball target matching the human target is determined.

6. The method according to claim 1, characterized in that, The method of obtaining the shot count result of the human target based on the basketball trajectory of the matched basketball target, the basket area and the net area on each video frame includes: For each basketball region in the basketball trajectory of the matched basketball target, the motion state of the basketball target in the basketball region is determined based on the relative position of the basketball region with the basket region and the net region on the video frame; the motion state includes: pre-entry state, entry state and post-entry state; When determining the state transition process of an effective shot based on the motion state of each basketball target, a shot count is performed to obtain the shot count result of the human target; where, the state transition process of an effective shot refers to the state transition process of the basketball target from the pre-entry state, the entry state to the post-entry state.

7. The method according to claim 6, characterized in that, The method further includes: Based on the basketball trajectory of the matched basketball target, the ordinate of the set point in the basketball area in multiple consecutive video frames is differentially calculated to obtain a differential sequence including multiple differential values. When determining the state transition process of an effective shot based on the motion state of each basketball target, the system determines whether the direction of the basketball target's motion during the state transition process of the effective shot is from top to bottom based on the difference value involved in the state transition process of the effective shot in the difference sequence, and whether the basketball trajectory of the matched basketball target passes through the basket area before passing through the net area. The process of counting a shot includes: If the basketball target moves from top to bottom during the state transition of this valid shot, and the trajectory of the matched basketball target passes through the basket area and then through the net area, then a shot count is performed.

8. The method according to claim 6 or 7, characterized in that, The method further includes: When determining the state transition process of a valid shot based on the motion state of each basketball target, the net area in the video frame involved in the state transition process of this valid shot is used to determine whether the net shakes during the state transition process of this valid shot; The process of counting a shot includes: If the net shakes during the transition of a valid shot, a shot count will be recorded.

9. The method according to claim 8, characterized in that, The method further includes: After a set time window following a shot count, another shot detection is performed.

10. The method according to claim 1, characterized in that, Key points on the human body include: wrist key points, elbow key points, shoulder key points, hip joint key points, knee key points, and ankle key points; the method also includes: For each valid shot, the elbow angle and knee angle of the human target are calculated based on the positions of the human key points in each video frame involved in the valid shot; wherein, the elbow angle refers to the angle formed by the shoulder key point, the elbow key point and the wrist key point; the knee angle refers to the angle formed by the hip joint key point, the knee key point and the ankle key point. The shooting phase of the human target in each video frame is determined based on the elbow angle, knee angle, wrist key point position, and shoulder key point obtained in each video frame; the shooting phase includes: preparation phase and release phase; For video frames where the human target is in the preparation stage, calculate the index value of at least one first scoring index based on the position of the human key points. For video frames where the human target is in the release phase, the index value of at least one second scoring index is calculated based on the position of the human key points. The first target video frame is selected from each video frame in the preparation phase based on the index value of at least one first scoring index of the human target in each video frame of the preparation phase; and The second target video frame is selected from each video frame in which the human target is in the release phase based on the index value of at least one second scoring index of the human target in each video frame of the release phase. The shooting posture score of the human target in this shot is calculated based on the index value of each first scoring indicator in the first target video frame and its corresponding set threshold range, and the index value of each second scoring indicator in the second target video frame and its corresponding set threshold range.

11. The method according to claim 10, characterized in that, The method further includes: Based on the prompt word template, the shooting posture score of the human target, the index value of the first scoring indicator, the index value of the second scoring indicator, and the basic information of the human target are input into a pre-trained large language model. The large language model then matches the N search results that are closest to the input of the large language model from a pre-defined second database. Based on the search results, the shooting posture score of the human target, the index value of the first scoring indicator, the index value of the second scoring indicator, and the basic information of the human target, targeted training suggestions for the human target are output. The basic information of the human target includes at least the following: the human target's identity information, age, and years of playing basketball.

12. A basketball shot counting device, characterized in that, include: The shooting video stream acquisition module is used to acquire shooting video streams; The first positioning module is used to perform region positioning and human key point positioning for each video frame in the basketball shooting video stream, so as to obtain all human regions in the video frame and the positions of human key points in each human region. as well as The second positioning module is used to locate the basketball target in the video frame and obtain all basketball areas in the video frame. as well as The third positioning module is used to perform region positioning of the basket and net targets in the video frame, and obtain all basket and net regions in the video frame. The first tracking module is used to track human targets in each video frame based on the human body region in each video frame, and obtain at least one human motion trajectory. as well as The second tracking module is used to track the basketball target in each video frame based on the basketball area in each video frame, and obtain at least one basketball trajectory. The shot counting module is used to match a basketball target to each human target in the human motion trajectory, and obtain the shot count result of the human target based on the basketball motion trajectory of the matched basketball target, the basket area and net area on each video frame.

13. An electronic device, characterized in that, include: A processor, configured to execute program instructions; as well as A memory configured to store the program instructions, which, when loaded and executed by the processor, cause the processor to perform the shot counting method according to any one of claims 1-11.

14. A computer-readable storage medium storing program instructions, characterized in that, When the program instructions are loaded and executed by the processor, the processor performs the shot counting method according to any one of claims 1-11.