System and method for shot improvement through machine learning
Through machine learning and computer vision technology, real-time analysis of shooters’ video data, track the movement of body logos, generate shooting motion data, and provide improvement suggestions, solving the problem that shooters have difficulty quantifying and improving their skills, achieving improved shooting accuracy and consistency.
Patent Information
- Application Number
- CN202380066092.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-13
- Filing Date
- 2023-09-13
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art is difficult to effectively quantify and improve shooter skills, especially in the absence of personalized guidance, which makes it difficult for shooters to improve shooting accuracy and consistency.
Through machine learning and computer vision technology, real-time analysis of shooter video data, track the movement of body logos, generate shooting motion data, and provide improvement suggestions, including using machine learning models to identify shooting errors and posture problems, providing real-time feedback and training suggestions.
Real-time feedback and personalized guidance for shooters are achieved, shooting accuracy and consistency are improved, and shooters can identify and correct problems in posture and movement.
Smart Images

Figure CN120569769A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 406,245, filed on September 13, 2022, entitled “SYSTEMS AND METHODS FOR MARKSMANSHIP IMPROVEMENT THROUGH MACHINE LEARNING,” and U.S. Provisional Patent Application No. 63 / 406,208, filed on September 13, 2022, entitled “SYSTEMS AND METHODS FOR AUTOMATED TARGET IDENTIFICATION, CLASSIFICATION, AND SCORING,” and U.S. Provisional Patent Application No. 63 / 406,241, filed on September 13, 2022, entitled “SYSTEMS AND METHODS FOR MARKSMANSHIP DIGITIZING AND ANALYZING,” the contents of which are incorporated herein by reference in their entirety.
[0003] background
[0004] According to many reports, millions of people enjoy target shooting every year, and the number of people who regularly shoot targets has increased over the past decade and continues to increase. In the United States alone, it is estimated that more than 52 million people regularly shoot targets. There are many different types of recreational shooting activities, from simple casual shooting (plinking) with a pistol or rifle pointed at a paper or steel target, to skilled long-range rifle shooting competitions that require a high degree of self-discipline and skill, to fun and fast-paced pistol shooting at pop-up or fixed targets, or shotgun shooting at skeet, trap, sporting clays, etc. In addition to recreational shooting, more and more target shooters practice as part of their profession, such as law enforcement officers, military personnel, and security personnel.
[0005] While many participants are satisfied with their current skill level, a growing number of shooters are looking to improve their skills. However, in some cases, many participants don't know how to improve. Shooting range attendance continues to grow, and these participants, whether for sport, recreation, personal defense, or public defense, are eager to improve their skills. However, without personalized shooting instruction, or even with regular instruction, it's difficult to quantify improvements and any issues affecting accuracy. Furthermore, many participants don't know how to get better.
[0006] Therefore, there is a need for a system and method that can analyze participant habits and provide feedback and recommendations for improvement. There is also a need for a system that can provide the above benefits in near real time using consumer-grade devices. These and other benefits will become apparent from the following disclosure.
[0007] Overview
[0008] A system having one or more computers can be configured to perform specific operations or actions by installing software, firmware, hardware, or a combination thereof on the system, wherein the software, firmware, hardware, or a combination thereof causes the system to perform the actions. One or more computer programs can be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform the actions. One general aspect includes a method for improving shooting performance. The method also includes receiving video data of a shooter; determining one or more body markers of the shooter, tracking the one or more body markers during the shooting to generate shooting motion data, determining a score for the shot, associating the shooting motion data with the score, and generating a recommendation for changing the motion data for a subsequent shot. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0009] Implementations may include one or more of the following features. In the method, determining one or more body landmarks of the shooter may include generating a wireframe model by connecting the body landmarks. Associating the shooting motion data with a score may include executing a classification and regression tree machine learning model to identify a causal relationship between the shooting motion data and the score. The method may include determining the shooter's grip by image analysis of video data of the shooter. The method may include analyzing the shooter's grip and providing grip recommendations on a display screen to change the grip. Determining the score of the shot may include: receiving target video data; performing image analysis on the received target video data; determining a hit on the target; and determining a score for the hit. Receiving the video data may include capturing the video data via a mobile phone. Determining the one or more body landmarks may include determining 17 body landmarks. Tracking the one or more body landmarks may include generating a bounding box around each of the one or more body landmarks. The method may include executing a machine learning model to associate the shooting motion data with a score. The machine learning model is configured to determine motion data that results in an off-center target hit. Implementations of the described technology may include hardware, methods or processes, or computer software on a computer-accessible medium.
[0010] One general aspect includes a method for improving the causal outcomes of body movement. The method also includes receiving video data of body movement; determining one or more body landmarks viewable in the video data of body movement; tracking the one or more body landmarks during the movement; generating movement data based at least in part on the tracking of the one or more body landmarks; determining a score associated with the movement data; associating the movement data with the score; and generating a recommendation for changing the movement data for a subsequent movement. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0011] Implementations may include one or more of the following features. In the method, receiving video data includes capturing the video data via a mobile computing device. The method may include executing a machine learning model to associate motion data with a score. The machine learning model is configured to determine the motion data that resulted in a decreased score. The method may include predicting, by the machine learning model, a predicted score based on the motion data. The method may include comparing the predicted score to the score. Implementations of the described technology may include hardware, methods or processes, or computer software on a computer-accessible medium. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are part of the present disclosure and are incorporated into this specification. The accompanying drawings illustrate examples of embodiments of the present disclosure and, in conjunction with the description and claims, serve to at least partially explain various principles, features, or aspects of the present disclosure. Certain embodiments of the present disclosure are described more fully below with reference to the accompanying drawings. However, various aspects of the present disclosure may be implemented in many different forms and should not be construed as limited to the implementations set forth herein. The same numbers always refer to similar, but not necessarily identical or identical, elements.
[0014] Figure 1A An image of a shooter captured by an image capture device is shown in accordance with some embodiments.
[0015] Figure 1B According to some embodiments Figure 1A Wireframe model of a shooter.
[0016] Figure 2 Motion data associated with multiple body landmarks is shown in accordance with some embodiments.
[0017] Figure 3 Motion data associated with multiple body landmarks and automatic shot detection based on the motion data are shown in accordance with some embodiments.
[0018] Figure 4Analysis of motion data associated with multiple body landmarks and identifying shooting errors based on the motion data are shown in accordance with some embodiments.
[0019] Figure 5 Analysis of motion data associated with multiple body landmarks and identifying shooting errors based on the motion data are shown in accordance with some embodiments.
[0020] Figure 6 Analysis of motion data associated with multiple body landmarks is shown in accordance with some embodiments.
[0021] Figure 7 Analysis of motion data associated with multiple body landmarks and identification of performance-enhancing posture changes are shown in accordance with some embodiments.
[0022] Figure 8A An image of a shooter captured with an imaging device is shown in accordance with some embodiments.
[0023] Figure 8B Shown is a diagram of a body with recognized body landmarks according to some embodiments. Figure 8A Computer-generated wireframe model of a shooter.
[0024] Figure 9 Motion data associated with a plurality of body markers and identifying body motions that result in decreased performance are shown in accordance with some embodiments.
[0025] Figure 10A and Figure 10B A diagnostic target for identifying shooter behavior based on shooting patterns is shown in accordance with some embodiments.
[0026] Figure 11 An image of a shooter's hand and a computer-generated wireframe model associated with body landmarks that identifies and analyzes the shooter's grip is shown in accordance with some embodiments.
[0027] Figure 12A is an illustration of a user interface for a computer software program that captures video images, tracks the movement of body landmarks, and analyzes a shooter's posture and movement data, according to some embodiments.
[0028] Figure 12B is an illustration of a user interface for a computer software program that analyzes a shooter's performance and automatically scores the performance, according to some embodiments.
[0029] Figure 13A and Figure 13B A system for marksmanship digitization and analysis is shown in accordance with some embodiments.
[0030] Figure 13C A system for marksmanship digitization and analysis is shown in accordance with some embodiments.
[0031] Figure 14 A system for marksmanship digitization and analysis is shown in accordance with some embodiments.
[0032] Figure 15 The logic of a machine learning algorithm for quantizing shot samples is shown in accordance with some embodiments.
[0033] Figure 16 An example decision tree machine learning model is shown that correlates motion data of body markers with shooting performance in accordance with some embodiments.
[0034] Figure 17 An annotated decision tree machine learning model is shown that correlates motion data of body landmarks with shooting performance in accordance with some embodiments.
[0035] Figure 18 A pruned decision tree machine learning model is shown that correlates motion data of body markers with shooting performance in accordance with some embodiments.
[0036] Figure 19A A system for capturing video data of a shooter and a target is shown in accordance with some embodiments.
[0037] Figure 19B and Figure 19C Image data of a shooter's left and right views captured by an imaging device is shown in accordance with some embodiments.
[0038] Figure 19D Image data of a top view of a shooter captured by an imaging device is shown in accordance with some embodiments.
[0039] Figure 19E Image data of a target captured by an imaging device is shown in accordance with some embodiments.
[0040] Figure 20 An example user interface of a computer application configured to receive, analyze, and score marksmanship performances is shown in accordance with some embodiments.
[0041] Figure 21 An example user interface of a computer application configured to receive, analyze, and score marksmanship performance is shown, displaying shot placement and automatic scoring, in accordance with some embodiments.
[0042] Figure 22An example user interface of a computer application configured to receive, analyze, and score marksmanship performance is shown that allows for selection of a body marker and displays athletic data associated with the selected body marker in accordance with some embodiments.
[0043] Figure 23 Several key technologies enabled in the embodiments described herein and the features facilitated by the key technologies are shown.
[0044] Figure 24 An example user interface is shown for a system and a computer application configured to receive, analyze, and score marksmanship performance according to some embodiments, the system using multi-source collaboration and machine learning to detect fired shots, analyze shot placement and provide recommendations for improvement, and analyze grip posture and provide recommendations for improvement.
[0045] Figure 25 is a process flow for capturing athletic data, analyzing the athletic data, and generating recommendations for improvement, according to some embodiments.
[0046] Figure 26 is a process flow for associating a firearm with an individual shooter, according to some embodiments.
[0047] Figure 27 An example system for determining an acoustic signature according to some embodiments is shown;
[0048] Figure 28 shows an example flow chart for pre-training a DNN according to some embodiments;
[0049] Figure 29 shows an example flow chart for deriving spectrogram weights according to some embodiments; and
[0050] Figure 30 An example flow chart for performing classification and determining an acoustic signature is shown.
[0051] Figure 31 A system configured for automatically scoring shooting targets according to some embodiments is provided.
[0052] Figure 32 shows an example process flow for classifying and scoring targets according to some embodiments;
[0053] Figure 33 shows an example process flow for identifying and classifying objects according to some embodiments;
[0054] Figure 34 shows an example process flow for registering targets and determining scoring hits according to some embodiments;
[0055] Figure 35A 、 Figure 35B and Figure 35C A method for initializing a target scoring system to identify a target according to some embodiments is shown;
[0056] Figure 36 shows an example process flow for detecting an impact on a target according to some embodiments; and
[0057] Figure 37 An example process flow for scoring impacts on a target is shown in accordance with some embodiments.
[0058] Figure 38 An example user interface for automatic goal scoring in a software application is shown, according to some embodiments.
[0059] Detailed description
[0060] Body landmark and pose estimation
[0061] According to some embodiments, a system is described that uses computer vision and machine learning to significantly reduce the manual and laborious nature of current training techniques, thereby providing a system that continuously learns and adapts. According to some embodiments, the system includes a machine vision and machine learning system that can track and estimate the posture of the participant and determine positive and negative factors that affect the quality of participation, such as posture, movement, anticipation, recoil, grip, posture, etc. This can be largely performed by a computer vision system that can simultaneously track several body landmarks and, in some cases, correlate the movement of one or more body landmarks with marksmanship accuracy. As an example, the system can identify and track any number of body landmarks, such as 3, 5, 11, 17, 21, 25, or 30 or more body landmarks. When the system tracks the position of the landmarks (which can be two-dimensional or three-dimensional), the system can correlate the movement of the landmarks with the rounds sent down range and the scores of the individual rounds. Detection of a bullet fired at a target area may be determined by movement of one or more suitable markers, such as movement of a participant's hand or wrist (or other body marker) in response to recoil, the sound of a firearm, a pressure wave associated with a muzzle blast, a target hit, or some other marker.
[0062] The system may also monitor the participant from one, two, three, or more perspectives and analyze the movement of each body landmark, and may also monitor shooting accuracy and correlate body landmark movement with accuracy. Based on the accuracy, the system may further provide an analysis of body movements that resulted in less than perfect accuracy, and may further suggest ways to improve body movement to increase accuracy.
[0063] Motion capture can be performed by one or more cameras generally aimed at the participant and one or more cameras aimed at the target. In some cases, one or more of the cameras is associated with a mobile computing device, such as, for example, a smartphone, tablet, laptop, digital personal assistant, and a wearable device (e.g., a watch, glasses, body camera, smart hat, etc.). In some cases, the wearable device may include sensors such as accelerometers, vibration sensors, motion sensors, or other sensors that provide motion data to the system. In some embodiments, the system tracks body marker positions over time and generates a motion map.
[0064] refer to Figure 1A , one or more cameras can capture one or more views of the shooter 100. The cameras can capture video data of the shooter as the shooter draws his gun, aims, fires, reloads, and / or adjusts his position. The computer system can receive the video data, analyze the video data, and create a video such as Figure 1B 10. The model associated with the shooter is shown in FIG. 10. In some cases, the computer system identifies body landmarks and connects them into a wireframe model 102 that tracks the pose and movement of the shooter's body landmarks. In some cases, the body landmarks may include one or more of the nose, left ear, left eye, left hip, left knee 103, right ear, right eye, right hip, left ankle, left elbow, left wrist, right knee 104, right ankle, right elbow 106, right wrist 108, left shoulder, and right shoulder 110. Of course, other body landmarks are possible, but for the sake of efficiency throughout this disclosure, we will focus on these seventeen body landmarks. In some embodiments, a single camera may capture two-dimensional motion data associated with one or more of the body landmarks. In some examples, two or more cameras may be used to capture three-dimensional motion data for one or more of the body landmarks.
[0065] The body markers can be tracked over time, such as during a shooting string (e.g., a shooting session), and the movement of one or more of the body markers can be tracked during that time. In some cases, two-dimensional movement is tracked in the x and y directions corresponding to lateral and vertical movement. In some cases, three-dimensional movement of the body markers is tracked in the x, y, and z directions.
[0066] refer to Figure 2 , shows a graph of selected body markers 200. The body markers can be user-selectable to focus on individual or combinations of body markers for viewing. For example, the topmost line 202 represents the right wrist of a right-handed shooter in a horizontal direction, while the third downward line 203 represents the vertical movement of the right wrist over time. As can be seen, the movement of the wrist moves up and down during the motion capture process. In some cases, the system can associate the movement of one or more body markers with events or phases during a shot sequence.
[0067] For example, during the first stage 204, the right wrist is relatively low, and the system can associate this position and movement with the start of arming. The second stage 206, showing the right wrist moving upward in very short intervals, can be associated with drawing the pistol from the holster. The third stage 208 shows the right wrist remaining substantially stable in the vertical plane; however, there are sharp peaks in movement 207a-207c, which may be associated with firing a shot from the pistol.
[0068] During the fourth phase 210, the right wrist moves down and back to the firing position. The system can associate this motion with the reloading operation. In some cases, the system is trained on training data, which can be supervised learning, to associate similar motions with the various phases.
[0069] During the fifth phase 212, as the shooter aims, the shooter's right wrist initially moves upward and then stabilizes onto the target, followed by peaks in movement 213a-213c, which may be associated with a shot being fired at the target.
[0070] In the sixth stage 214, the right wrist again moves downward to the initial position, which may be associated with holstering the pistol.
[0071] While this example focuses on the shooter's right wrist in a vertical orientation, it should be apparent that any body landmark can be viewed, analyzed, and movement or combination of movements can be associated with an event or action of the shooter, including groups of body landmarks.
[0072] refer to Figure 3 , depicting a close-up view of a firing phase 300 showing a sharp peak 302 of right wrist movement and a sharp peak 304 of right elbow movement in the vertical direction. The system can analyze the motion data and automatically determine when to fire a shot. The system can be configured to associate a sharp peak in the vertical motion of the wrist and / or elbow with a fired shot. Figure 3As shown, each of arrows 306 can coincide with a fired shot. Additionally, audio data can be correlated with motion data to provide additional clues about when the shot was fired. In some cases, audio data can be combined and / or synchronized with motion data to provide additional details about the fired shot. Motion data can be used to infer additional information about the shooter, their habits, their posture, and other clues that can be brought to the shooter's attention in an effort to improve the shooter's accuracy.
[0073] refer to Figure 4 , showing motion data 400 of a shooter's right wrist 402 and right elbow 404. In some cases, the system may apply a trend curve 406 to the motion data, which may represent normalized motion data. Additionally, the system may make inferences and / or determinations based on the motion data 400. For example, Figure 4 As shown, at 408 places, once the shooter has aimed the firearm at the target and attempts to keep the firearm stable, the shooter will lower the wrist and elbow at 410 places, and then launch the shot at 412 places. The system can identify this pattern and determine that the movement of lowering the wrist and / or elbow immediately before shooting shows that the shooter attempts to prejudge the recoil of the firearm and attempts to resist recoil. In many cases, prejudging recoil can significantly reduce the accuracy of shooting because the shooter moves the firearm away from the target when prejudging the recoil that occurs when launching the shot. Similar examples of motion data that may reduce the shooter's accuracy include flinching, pre-ignition push, trigger jerk, eye closure, etc.
[0074] In some cases, the system can provide information to the shooter regarding recoil anticipation and provide information that may include one or more drill or practice sessions in an effort to improve the shooter's motion data regarding recoil anticipation. For example, the system can identify a practice regimen that may include dry firing, skip loading (e.g., mixing live rounds in a magazine of dummy rounds), or other skill-building exercises.
[0075] Figure 4 It also shows that the shooter experiences a drift in posture. For example, before the first shot 412, the shooter's right wrist and right arm are positioned at a first vertical height 414, and before the second shot 416, the shooter's right wrist and right elbow are positioned at a second vertical height 418, which is higher than the first vertical height. This data indicates that the shooter does not return to the same position from the first shot to the second shot, so the shooter's sight picture will be slightly different, which may reduce the accuracy of consecutive shots.
[0076] refer to Figure 5 The motion data 500 further illustrates drift. The motion data associated with the right wrist 502 and right elbow 504 of a right-handed shooter not only illustrates recoil anticipation, where the motion data shows a decrease in that body part just before the shot is fired, but also shows a trend line 506 that demonstrates the shooter's wrist and elbow continuing to drift upward during the string of shots. The shooter's inability to return to the same position between shots can significantly reduce the accuracy and precision of the shots fired within the string of shots.
[0077] In some cases, the system will identify drift of one or more body landmarks and may provide this information to the shooter. In some cases, the system will provide information on a display screen associated with a mobile computing device. For example, the system may be implemented on a mobile computing device associated with the shooter, and the display screen on the mobile computing device may provide the shooter with information, instructions, or exercises to improve drift and the resulting accuracy issues.
[0078] Similarly, the system can correlate body marker motion with other correctable flaws in the shooter's position or posture. For example, referring to the figure showing body marker motion data 600, Figure 6 and Figure 7 , motion data 600 shows the movement of multiple body landmarks during a shot string. The bottom line graph depicts right ankle motion data 602 that proves the shooter changed the position of the right foot. Changing position during a shot string may affect the aiming diagram, accuracy, precision, and other metrics associated with shooting. The system can determine whether the change in foot position is positive or negative in terms of shooting scoring and can provide recommendations to the shooter based on the change in posture. The system can review the shooting accuracy and / or precision of shots 604a-604e fired before and after the repositioning of the foot and determine whether the movement of the foot has a positive or negative impact on shooting performance.
[0079] In some cases, the system correlates a target's score (e.g., shooting accuracy and / or precision) with body landmarks and movement, and can indicate to the shooter which locations of various body landmarks affect their shooting performance, either for better or worse.
[0080] The system can be configured to correlate certain postures, movements, and combinations with shooting performance through machine learning. In some cases, body marker movements and combinations of movements can be associated with improved shooting performance, while others may be associated with decreased shooting performance.
[0081] The system can normalize the motion data to generate normalized coordinates for the position of each body part throughout the session. The scores and their moving averages can be represented by a signal, such as by displaying them in a user interface. The motion data and / or scores can be stored in a data file that can be analyzed in near real time or saved for later analysis.
[0082] The motion data can be associated with the shooting pattern data, such as the x and y coordinates of each shot, and the shot coordinates can be temporally associated with the motion data occurring at the time the shot was fired. Additionally, a score can be assigned to the shot and saved with the shot data.
[0083] One or more machine learning methods can be applied to the motion data and the shooting data to generate a correlation between motion and shooting accuracy. For example, in some embodiments, a convolutional deep neural network (CNN) can be designed to locate features in a collection of motion and shooting data. Other deep learning models for classification can also be used to correlate motion data and shooting data to identify patterns that lead to increased or decreased accuracy. Transformations that correlate any set of attributes with any other set of attributes can also be used.
[0084] Figure 8A An example camera angle of shooter 800 is shown, and Figure 8B Shown is shooter's gained wireframe model 802, which allows system tracking to comprise the motion of the shooter's body of selected body marker.As can be seen, this system can determine shooter's posture, attitude and the motion in whole shooting string.Wireframe model can comprise the representation of each major joint and body artifact (body artifact), can comprise one or more of shooter's nose or chin 804, right shoulder 806, right elbow 808, right wrist 810, hip 812, left femur 814, left knee 816, left shank 818, left ankle 820, right femur 822, right knee 824, right shank 826 and right ankle 828 etc. In some cases, this system can be trained according to the historical performances of those shooters relevant with the motion data from various shooters.In this case, system can determine the influence of specific posture and / or motion on shooting performance.In some cases, system can rely on the motion data from single shooter to help this shooter to adjust and / or further training exercise to improve shooter's performance.
[0085] Figure 9Shown is x-axis motion data 900 associated with a shooter's head 902 and nose 904. As can be seen, during the shot string, after each shot 906a-906d, the shooter's head moved backward, which can be correlated with shooting performance. That is, the system can determine that the shooter's head moved after each shot, for example, to see the target, and may have drifted excessively during the shot string without returning to the exact same point, which causes the shooter's performance to be below a threshold.
[0086] In some cases, the system may be configured with logic to determine probable cause and effect based on the shooter's movement and / or the scoring of the target. Figure 10A and Figure 10B , Figure 10A and Figure 10B Diagnostic targets are shown separately for a left-hand shooter and a right-hand shooter 1000. The right-hand target and the left-hand target are mirror images of each other, so only one target will be described below.
[0087] Assuming the shooter is able to keep the pistol on target, motion associated with the shooter may be the cause of off-center hits. In cases where off-center hits are regularly scattered to one side of the target, there may be some issues that can lead to off-center hits. For example, if a grouping of shots falls at the 12:00 position, the system can determine that the shooter bent their wrist 1002 upward (e.g., riding the recoil), which typically occurs in anticipation of recoil.
[0088] If the shot spread falls at the 1:30 position 1004, this may indicate a grip error known as heeling, where the heel of the hand pushes the butt of the pistol to the left in anticipation of the shot, forcing the muzzle to the right.
[0089] If the shot spread falls at the 3:00 position 1006, this indicates that the thumb of the shooting hand is applying too much pressure and pushing one side of the pistol to the right, forcing the muzzle to the right.
[0090] In the case where the shot spread falls at the 4:30 position 1008, this may indicate that the shooter tightened their grip (e.g., lobstering) when shooting. This would cause the front sight to depress when the trigger is pulled, resulting in a shot low and to the right for a right-handed shooter.
[0091] When the shot spread falls into the 6:00 position 1010, this can signal the wrist to press down and forward. This is usually an unconscious effort to control recoil and prevent the muzzle from rising.
[0092] In the case where the shot spread falls in the 7:00 position 1012, this may indicate a trigger pull or slapping. This may indicate that the shooter attempted to pull the trigger at the instant the sights were aligned with the target.
[0093] When the shot spread falls at the 8:00 position 1014, this may indicate to the shooter to tighten their grip during the shot.
[0094] In the event that the shot spread falls at the 9:00 position 1016, this may indicate a too little finger on the trigger. This often results in the shooter pulling the trigger at an angle during the final rearward movement of the trigger, which tends to push the muzzle to the left.
[0095] Where the shot spread falls at the 10:30 position 1018, this may indicate that the shooter pushed in anticipation of recoil, but lacked sufficient follow through.
[0096] The system can be programmed to look at target hits during a shooting string and, in combination with body landmark movements, determine if the shooter is causing recoil anticipation, trigger control errors, and / or grip errors. The system can make this determination for an individual shooter and can recommend practice drills and practice strings to address specific shooting issues.
[0097] In some cases, the system may access target prior engagement data (DOPE) associated with the shooter, which may include previous shooting sessions, records, scores, and analysis.
[0098] In addition to real-time scoring, complete shooting sessions can be recorded and stored so that users can review them along with determinations from the machine learning system, identifying and / or highlighting flaws in the shooter's form or technique, errors, and changes that improve or degrade the shooter's performance.
[0099] Synergy comes from the synchronization of all available information sources as described herein.
[0100] Continue to refer Figure 10A and Figure 10B , where the diagnostic goal is shown, the basic conventional ML model can map the target shooting pattern φ∈Φ to the shooter behavior β∈B,f T Φ→B. The behavior can be a four-dimensional (4D) phenomenon consisting of three-dimensional (3D) motion data plus time. The firing pattern Φ is a two-dimensional (2D) projection of the evolving 3-D phenomenon (2D plus time). This model can be based on causal assumptions.
[0101] Experts may study a larger collection of complex shooter behaviors and maps it to a larger set of firing patterns Φ * , f XP :B * →Φ. Behavior β∈B * can be a 4D phenomenon, and the shooting pattern φ∈Φ * It is a 2D phenomenon. One can speculate that this is in principle a causal model because the expert fits the model based on his or her understanding of the shooter's physiokinematics and psychology, as well as the physics of shooting.
[0102] A relatively simple iterative model for expert learning could consist of predicting shooting patterns based on observed shooter behavior: f XP :B * →Φ. The model can then evaluate the difference between the predicted firing pattern and the actual firing pattern: δ(φ′, φ). The model can then adjust the prediction model f of the causal information based on this difference δ(φ′, φ) XP This can be repeated for each n-shot training session.
[0103] This type of learning model considers the shooter behavior β∈B * is considered as a 3D phenomenon and the shooting pattern φ∈Φ * is considered a 2D phenomenon. In some cases, the model is enhanced to include B * Considered as a 4D phenomenon, Φ * According to some embodiments, the posture analysis method is performed by B ** The 3D projection (2D plus time) of represents the 4D shooter behavior B * .
[0104] The model can be implemented by executing the decision function f G : Φ→D, objectively identifying good shooting patterns (e.g., tightly scattered around the bullseye) from the rest of the “bad” shooting patterns, where D is a binary variable. The model can also be configured to factor in the shooter behavior E * Related to how (and why) the shooting mode φ is "bad".
[0105] In some embodiments, the sampled time signal of the shooter behavior can be reduced to a single row of data in the dataset for a single training. A bounding box can be derived around the set of XY coordinates of each body landmark in the training. The bounding box can be denoted as shooter behavior B. Therefore, during a time period, there can be a shooter behavior for each of the body landmarks. In some cases, the shooter behavior B and the shooting pattern Φ are represented in polar coordinates as a radius and an angle. The shooter behavior B (17 in some examples) is 2D with (x, y) coordinates or (r, θ) representation. The shooting pattern φ∈Φ is a causal consequence of the shooter behavior β∈B, and the method can regard the shooting pattern Φ as a proxy for the unobservable third dimension of the shooter behavior B. The method can regard the shooter behavior β∈B and the shooting pattern φ∈Φ as features, and the score s∈S as the target f of the ML model ML :B xΦ→S.
[0106] The model can be used for the pattern (β, φ, s)∈Ψ L ,Ψ H The analysis is performed to produce low scores and high scores. Implicit or explicit clustering techniques can be used for this purpose. In some cases, the model f ML : B xΦ→S can be constructed for multiple trainings of a specific shooter and for training of multiple shooters. In some cases, given a specific training (β, φ, s) and a model f ML : BxΦ→S, model quality can be evaluated by comparing the actual scores s∈S with the predicted scores s′∈S. Some models can be updated sequentially based on these comparisons.
[0107] In some embodiments, low score training (β, φ, s) can be evaluated for Ψ to determine which of the 17 components of observable shooter behavior β and shooting pattern φ, which is a proxy for unobservable shooter behavior, may be the cause of the low score. This may be easiest for an interpretable model. Unobservable behavior can be described by common labels in diagnostic targets (such as those described above). In some cases, the shooter diagnostic application can also distinguish low scores caused by aiming misalignment, which may be easiest to understand because the shots are closely scattered around the center of mass rather than the target bull's eye. In some cases, the system determines that a reduced score has been obtained and generates recommendations for improving the score. As used herein, the term "reduced score" is used to represent a score that is less than the target score. The target score can be a perfect score, the highest historical score of a participant, the average score of a participant, or some other metric. For example, in a shooting event, the highest score for a given shot is typically "10". In this case, a reduced score is any score below "10". Similarly, in other sports, a score less than the expected score can be assigned, such as, for example, a penalty kick in soccer can be a binary outcome, with a missed shot being a reduced score compared to a scored shot. Regarding golf shots, the average golfer's driving distance is 275 yards. A reduced score could result from a 250-yard drive, which is less than the average golfer's driving distance, and the system can observe behavior and determine which behaviors lead to a reduced score.
[0108] According to some embodiments, the system relies on artificial intelligence (AI) and / or machine learning (ML) to make two determinations: detecting shooter movement and shot placement, and analyzing shooter behavior to understand how these behaviors lead to shot placement. Having described detecting movement, there are multiple approaches to analyzing how shooter behavior leads to shot placement. First, an expert-based approach utilizes a comparison of an individual shooter's behavior in each training session with an expert's assessment of what should lead to a good shot. Second, a data-based approach is performed by building a model from repetitions ("training") of the shooting behavior of a single shooter or multiple shooters. This can be considered an AI / ML-based discovery strategy that determines which behaviors are associated with good shooting.
[0109] In some cases, AI / ML can be used to automate and augment expert-based analysis, meaning that if prototypical “ideal” behaviors are known a priori, models of these a priori known behaviors can be fit to the data of a single shooter or multiple shooters.
[0110] Viewed in this light, Transformer ML models can emulate automated and augmented expert-based analysis by combining a pre-trained data-based model for inferring relationships between language segments (similar to inferring relationships between shooter behavior and shot placement) with an additional stage of data-based adaptation.
[0111] refer to Figure 11 , the system determines the key points of the shooter's hand 1100. The key points can coincide with each movable joint of the wrist, hand, and fingers, as well as their respective positions and positioning relative to each other. The system can connect the key points into a wireframe model 1102 of the shooter's hand, which allows for accurate monitoring of posture and movement. Figure 11 A plurality of key points associated with the shooter's hand are shown, which can be used to determine the grip used by the shooter. For example, by referencing the key points of the shooter's hand, the system can determine whether the shooter uses a thumb-forward, thumb-up, cup and saucer grip, wrist grip, trigger guard gamer, or another style of grip. Different grips and pressures applied by the hands can impart motion to the pistol, and the system can determine that a different or modified grip will result in better performance.
[0112] The system receives signals associated with the shooter's position and movement and processes these signals to detect errors in the shooter's posture and grip. In addition, by combining single-frame analysis (e.g., finding insights from the relative positions of different body parts at a specific moment, which is an analysis across time), it can identify errors caused by changes in the shooter's position.
[0113] The system can also focus on the position of the hands on the pistol. A hand tracking machine learning model can be used to track the position of each bone in each finger in the video frame. The system will receive signals associated with the position of each finger over time.
[0114] refer to Figure 12A and Figure 12B , shows an application that can be executed on a mobile computing device and receives image data from an imaging sensor associated with the mobile computing device. As with any of the embodiments described herein, the methods and processes can be programmed into an application or instruction set that can be executed on a computing device. In some cases, the application can be executed on a mobile computing device and audio / video capture, analysis, scoring, recommendations, training exercises, and other feedback can be performed using the mobile computing device. In some cases, the system receives video and / or audio from multiple video capture devices that capture different views of the same shooter. The system can use multiple different views of the shooter in analysis and provide feedback to the shooter to improve performance.
[0115] like Figure 12A Shown, as the system that is applied on mobile computing device 1200 and runs, captures the video frame of shooter 1202, sets up body mark to track over time, creates shooter's wireframe model 1203, tracks the shooting 1204 that is launched and provides feedback to shooter 1202.As shown in the figure, system can identify shooter's posture, is " Weaver (Weaver) " posture in the example shown, tracks the quantity 1204 of shooting, and provides feedback 1206 about each shooting.For example, the shooting display that carries out after 7 seconds of shooting string starts has hit mass center.System recognition carries out the reloading operation 1208 that continues 1.2 seconds at 8 seconds, is subsequently at the shooting of 11 seconds, shows shooter's prejudgment recoil, and shooting deviates from center.In some cases, this function of system can be regarded as tracking shooting, provides feedback and provides the training instructor that is used to improve the advice of posture, grip, posture, trigger pull etc. to improve shooter's performance.
[0116] Figure 12B An additional screen 1210 is shown that can be displayed by the system in spotter mode, where the mobile computing device can point its video capture device at a target. In some cases, the mobile computing device can use internal or external lenses to obtain a better view of the target, and the mobile computing device can be coupled to a spotting scope or other type of optical or electronic telephoto zoom lens to obtain a better image of the target.
[0117] The system can switch between different modes, such as observer, drill instructor, DOPE, lock (in which information about different firearms owned by the shooter can be stored) and program settings. The observer mode can display an additional screen 1210 with a target view, on which target hits can be seen. The system can also display other information, such as the combination of firearm 1212a and the ammunition 1212b launched, the distance 1214 to the target, the type of target 1216, the time 1218 of the shooting string, the number of shots launched, the number 1220 and score 1222 of target hits, etc. It can also display an image 1224 of the target, and can further highlight the hits 1226 on the image 1224 of the target. It should be appreciated that the observer mode can display information in real time, including the time, hits, scores, and number of shots that have passed. In addition, the system can also store data associated with the shooting string for later playback, viewing, and analysis. For example, a training instructor mode may review the DOPE associated with a shooter and a combination of firearm and ammunition and provide a detailed analysis of the mistakes the shooter made with a specific firearm and offer exercises or suggestions to improve the mistakes and thereby enhance the shooter's performance.
[0118] refer to Figure 13A 、 Figure 13B and Figure 13C , shows an example architecture 1300 for data capture. According to some embodiments, a shooting range 1301 may be equipped with one or more sensors, such as image sensors, projectors, and computing devices. Figure 13A Shooting range 1301 is shown from the perspective of shooter 1302 looking toward a target area. Figure 13B The shooting range 1301 is shown from a side view, and Figure 13C Shooting range 1301 is shown from a top plan view.
[0119] One or more cameras 1304 may be mounted within the shooting range to capture video of the shooter, which may be from multiple angles. Cameras 1304 may be any suitable type of video capture device, including but not limited to CCD cameras, thermal cameras, dual thermal cameras 1305 (e.g., a capture device with both thermal and optical capture capabilities), etc. Cameras 1304 may be mounted at any suitable location within the shooting range, such as, for example, to the sides of shooters 1304a, 1304b, facing in front of shooter 1304c, overhead 1304d, etc.
[0120] In some cases, the cameras 1304 may be mounted to a rack system that provides a portable structure for supporting one or more cameras 1304 to capture multiple angles of the shooter.
[0121] In some embodiments, a projector 1306 may be provided to project an image onto a screen 1308. In some cases, the projector can project any target image onto a screen or shooting wall, and the shooter can practice dry-firing (e.g., without firing live ammunition), and the system can record hits on the projected image. For example, an image may display a target on a screen, and the shooter can practice dry-firing at the target. The system can record the aiming point when the trigger is pulled, and display the hit on the projected target image. A compatible device may be provided that can emit a laser through the barrel of the firearm when the trigger is pulled to indicate where the shot will hit the target. A thermal imaging camera 1304 can detect the laser hit location, and the system can be configured to display the hit on the projected target, and the system can further record and score the hit. In this way, shooters can practice using a shooting simulator with their own firearms without having to go to a dedicated shooting range. The system can recommend dry-firing practice to improve bad shooting habits.
[0122] Figure 14A system configured for multi-source synchronization is shown. In some embodiments, multiple video capture devices 1404 are used to capture a shooter 1402 from different angles. The system obtains characteristics from the shooter 1406, such as posture, gesture, and movement. The system can analyze captured audio signals 1408 to supplement video shot detection. That is, the system can analyze both video and audio to detect shots fired by the shooter, such as by correlating spikes in the audio waveform with sudden muzzle lifts of the firearm.
[0123] The system can be configured to detect firing phases 1410, such as a first shot sequence, reloading, and a second shot sequence, including shot release detection. The system can also identify and determine firing errors 1412, and identify errors and training and exercises to address firing errors. Error detection can be iterated, and further error corrections can be proposed 1414. The system can also incorporate shot detection 1416 as described herein, including correlation of shot detection with audio data.
[0124] The system can also capture an image 1418 of the target (such as a video image) and record hits on the target, which can be correlated with motion data before and during the trigger pull. As described herein, the system can determine the shot location 1420 and ultimately determine a score 1422 for the one or more shots fired. Thus, some embodiments of the system provide an intelligent coaching / adaptive training solution that transforms a time-intensive and non-scalable live training environment into an automated and adaptive virtual scenario-based training solution.
[0125] According to some embodiments, the model is based on instructions executed by one or more processors, which cause the processors to perform various actions. For example, the method may include collecting (as many as possible) gesture analysis files, which may be stored as comma separated value (CSV) files.
[0126] Figure 15The shot bounding box 1500 determined by the system is shown. The "drillparser.py" script can be configured to assemble the gesture analysis files. The system can then find the shot in the CSV file, for example, for each of the 17 gesture landmarks, and export an oriented bounding box 1500 around the xy coordinates for the shot for a 1.0 second time window up to and including the shot. The angle φ 1502 and dimensions l 1504 and w 1506 of the bounding box 1500 are then used as factors in the decision tree model for the score. The "all shots" method processes the data based on the CSV file, thus drawing a bounding box 1500 around all shots as a row in the resulting dataset. The "oneshot" and "oneshotd" methods process each shot individually, so the shot bounding box 1500 can be an infinitesimal box around a single shot, so each row in the dataset is a shot. The "oneshotd" method additionally orients the bounding box around the 17 gesture landmarks in chronological order, while the "single shot" method ignores time.
[0127] Thus, the system is able to derive an oriented bounding box around each shot, which can be oriented around the body landmarks.
[0128] 3) These datasets are processed in BigML as follows:
[0129] 1. Upload to BigML as source object
[0130] 2. The source object is converted into a dataset object
[0131] 3. The dataset objects are divided into training (80%) / testing (20%) datasets.
[0132] 4. Build a tree model using factors selected from the training dataset
[0133] 5. Create batch predictions using the appropriate model and test dataset
[0134] 6. Download the model as a JSON PML object for subsequent steps
[0135] 4) The "modelannotator.py" script annotates the model PML file as needed for interpretation operations. In some cases, this annotation simply involves adding a "targets" attribute to each node, which is a list of all target (score) values reachable from that node in the tree.
[0136] 5) The “itemexplainer.py” script uses the (virtual) pruned annotation model to “explain” which elements of the newly trained poses led to unsatisfactory score predictions.
[0137] 6) Of course, instances where the predicted scores differ significantly from the actual scores can be used to adjust the model by repeating the model training through machine learning.
[0138] The factors derived from the shot sample may include one or more of the following factors: shot_phi with bounding box rotation (-π≤φ≤π); shot_l—bounding box length; shot_w—bounding box width; shot_lw—bounding box area; shotθ—bounding box center rotation (-π≤φ≤π); shot_r—bounding box center radius. A bounding box is drawn around the shot sample 1500.
[0139] A very similar approach can be used to derive factors from a sample of body landmarks. As the body landmarks move over time, each body landmark can also be used to create a bounding box with length, width, area, and rotation, and associated with a shot factor.
[0140] In some cases, the system is configured to use one or more cameras for machine vision of the shooter, determine one or more body landmarks (in some cases, up to 17 or more body landmarks), and track the movement of each of these landmarks over time during shooting training. The system can correlate time-limited body landmark movement with scored shots and determine whether the shot was good or poor. A good shot can be related to the combination of the shooter's DOPE, firearm, and ammunition. For example, if a single shot is closer to the aiming point than the shooter's average shot, this can be classified as a good shot. Furthermore, by observing several shots over a period of time, the system can correlate shooter behavior with good or poor shots. Furthermore, by analyzing shooter behavior (e.g., body landmark movement), the system can predict whether a shot was a good or poor shot, even without seeing the target results. For example, a good shot can be considered a shot within the 9th or 10th ring of the target, while a poor shot can be any shot outside the 9th ring. The definition of a good shot and a poor shot can vary depending on the shooter's expertise. For example, for a very skilled shooter, any shot outside the 10th ring can be considered a poor shot.
[0141] Finally, by correlating shooter behavior with target accuracy, the system can provide the shooter with output regarding the behavior that led to reduced accuracy. Additionally, the system can recommend one or more specific trainings to the shooter to address the behavior that led to reduced accuracy.
[0142] Figure 16Some embodiments of a decision tree model 1600 for associating body markers with shot samples are shown. A decision tree algorithm is a machine learning algorithm that uses decision trees for prediction. It follows a tree-like model of decisions and their possible outcomes. In some cases, the algorithm works by recursively splitting the data into subsets based on the most important features at each node of the tree.
[0143] From the body marker data, various features are extracted to represent the kinematic properties of the body's movement. These features include the positions of the body markers over time, bounding boxes describing the movement of the body markers during the time preceding and including the fired shot, relative distances between the markers, and time derivatives that capture the dynamic aspects of the movement. Feature scaling is performed to facilitate model convergence.
[0144] In some embodiments, Figure 16 The decision tree depicted in Figure 1 is used as the core of the model. It is chosen for its ability to elucidate complex nonlinear relationships within the data while maintaining interpretability. The depth of the tree is optimized to minimize overfitting through techniques such as cross-validation or a predetermined maximum depth. The choice of splitting criterion, whether based on Gini impurity or information gain, depends on the specific problem. Pruning methods (such as enforcing a minimum number of samples per leaf node) are applied to control excessive branching.
[0145] Training the model can involve the recursive construction of a decision tree. The tree has nodes associated with body landmarks. For example, the left_ankle_phi node 1602 can branch into the right_wrist_lw node 1604 and the right_knee_phi node 1606, and these branches can show how the bounding box rotation of the right ankle motion data, combined with the bounding box area of the right wrist or the rotation of the bounding box associated with the right knee, affects the resulting shot placement. Similarly, the decision tree can associate shot placement with a combination of features. At each node, the training data is partitioned based on the feature values, and the selected loss function, such as mean squared error or cross entropy, is optimized. Hyperparameters can be fine-tuned via cross-validation, and the performance of the model is evaluated using various metrics.
[0146] During real-time operation, the decision tree model is continuously updated as new body landmark data becomes available for each shot fired. For each input feature vector derived from the real-time landmark data, the model traverses the decision tree, reaching a leaf node. The label assigned to the leaf node (indicating a positive or negative outcome) is used as the model's prediction. In the example shown, right shoulder bounding box lengths between values of 0.80 and 0.82 are consistent with the predicted results for the shooter's performance shown at leaf nodes 1608 and 1610.
[0147] In the context of this model, a positive outcome means the successful execution of a specific move or the achievement of a desired shot placement. Conversely, a negative outcome indicates an incorrect move, a deviation from the desired posture, or an off-center shot placement.
[0148] The performance of the model is rigorously evaluated using various metrics including, but not limited to, accuracy, precision, recall, F1 score, and area under the receiver operating characteristic (ROC) curve. A confusion matrix is used to quantify how proficient the model is at classifying positive and negative outcomes.
[0149] Figure 17 An annotation model 1700 is shown, which displays various values at several nodes. For example, during a shot sequence, the area of the bounding box for the right wrist 1702 indicates values of 0.47, 0.51, 0.57, 0.59, and 0.61. This indicates that during the shot sequence, the shooter moved her right wrist within the area defined by the displayed values. Following the branch below the right wrist 1702, for values below the average of these values, the system identifies that the left knee 1704 moves in a certain pattern, which can be associated with good or bad shots, to find a causal relationship between the movement of the right wrist and the movement of the left knee. Other features can be annotated similarly within the model. The annotated features can be stored in a feature vector and associated with each shot fired, used to analyze how different features combine to produce a specific shooting performance.
[0150] Figure 18 Pruning model 1800 is shown, which allows to explore various features and their mutual relationships more deeply. For example, when viewing the features of right_knee_phi 1802 (e.g., the rotation of the bounding box associated with the movement of the right knee during the shooting string), we can see values 0.8, 0.82, and 0.9. Given these values, we can see that the right_knee_phi value of 0.9 provides a result at the first result node 1806. When right_knee_phi 1802 is 0.8 or 0.82 and right_shoulder_1 bounding box length 1804 is 0.09, we see the second result at the second result node 1808 and the third result node 1810. In some cases, the second result node 1808 is associated with poor shooting performance, while the third node 1810 is associated with good shooting performance. These nonlinear causal relationships can be determined by the machine learning model so that the system can determine which motions lead to better shooting results on poor shooting results by executing one or more machine learning algorithms.
[0151] In some cases, the Isolation Forest encodes the dataset of trees such that the leaf instances for each leaf node are the same. Node splits can be randomly selected instead of reducing the impurity of the sub-objective value. In some cases, a node is a leaf node when the training instances at the node have the same target value. In some cases, rare instances in the training dataset reach leaf nodes earlier than less rare instances. According to some embodiments, new anomalous instances have a short path relative to the tree depth of all the trees in the forest, which makes anomaly identification more efficient.
[0152] In some cases, a Classification and Regression Tree (CART) is a predictive model that explains how a result variable can be predicted based on other values. A CART-style decision tree is a tree where each split is a split of a predictor variable, and each end node contains a prediction of the result variable. It constructs a binary tree to partition the feature space into segments that are homogeneous with respect to the target variable, and is a recursive algorithm that makes binary splits of the input features based on a specific criterion, thereby creating a tree-like structure. A CART-style decision tree can be used to predict the outcome of a shot based on motion data of one or more body markers. A CART-style decision tree is a multi-class classifier (e.g., N>1), while an Isolation Tree is a single-class classifier (e.g., "anomaly or not"). In some cases, a CART-style decision tree can be reduced to an Isolation Tree by designating M < N classes as "non-anomalous", designating N - M classes as "anomalous", and pruning the branches leading to leaf nodes with N - M "anomalous" to their root nodes. This can be analogous to training an Isolation Tree only on M < N "non-anomalous" instances, which allows impurity-reducing splits and impure "non-anomalous" leaf nodes.
[0153] Figure 19A A rack system 1900 is shown, which can provide a portable mounting structure for accommodating one or more imaging devices, including one or more cameras. In some cases, the structure includes one or more uprights 1902 and one or more crossbars 1904. The rack structure 1900 can be placed around a shooter, and in some cases, in the shooter's target area. For example, the rack 1900 can position the crossbar 1904 at a location that is approximately one foot (≈0.3 m) to 8 feet (≈2.4 m) above the shooter and between one foot (≈0.3 m) and 15 feet (≈4.5 m) in front of the shooter. In some examples, cameras are positioned on each upright and on the crossbar. Thus, in some embodiments, two, three, or more cameras are positioned on the rack, some of which are aimed at the shooter, and one or more cameras can additionally be aimed at target area targets.
[0154] Figure 19B 、 Figure 19C and Figure 19D Various views are shown captured by cameras mounted to the frame 1900. A first camera 1908 may be mounted to the crossbar 1904 or the upright 1902 and positioned to capture a left side view of the shooter ( Figure 19B ). The camera may be configured to capture the entire shooter's body, or may be configured to capture the shooter's upper body and head.
[0155] The second camera 1910 can be mounted on the crossbar 1904 or the column 1902 and is configured to capture a right side view of the shooter ( Figure 19C ). The camera can be configured to capture the entire body of the shooter, or can be configured to capture the upper body and head of the shooter. In some cases, the first camera 1908 and the second camera 1910 utilize different fields of view, so that one of the cameras captures the entire body of the shooter, while the other camera captures only a portion of the shooter's body.
[0156] A third camera 1912 may be positioned on the crossbar 1904 and configured to capture a top view 1914 of the shooter ( Figure 19D By positioning the cameras to capture the shooter's motion from various angles, the system can correlate the video data from each camera and determine three-dimensional motion data for selected body landmarks.
[0157] The fourth camera may be positioned to capture target 1916 ( Figure 19E )'s video data and target impacts 1918. The video data from each camera can be synchronized and analyzed to determine when a shot was fired and to correlate the fired shots with recorded target hits or misses.
[0158] Figure 20 A computer program user interface 2000 is shown that can be used with the systems and methods described herein. For example, the computer program can be configured to receive video data from one or more cameras, synchronize the video data, determine when a shot is fired, and record and score hits or misses on targets. User interface 2000 can include a start recording button 2002 that allows the shooter to begin video capture. In some cases, the shot sequence can be timed, and the start recording button can also start a timer. The user interface can also include a timer 2004 associated with the session. The user interface can be presented on a mobile computing device associated with the user, or on a mobile computing device associated with a facility. For example, a rack system can be located at the facility, and a computing device associated with the facility can be connected to the rack system and configured to receive video data and provide performance feedback for the embodiments described herein.
[0159] The user interface may be provided on any suitable display, such as a television, a touch screen display, a tablet screen, a smartphone screen, or any other visual computer interface. Figure 21 , the user interface 2000 can display indicia associated with the shot sequence, such as, for example, a video of the shot sequence and the motion during the shot sequence 2102. The video can be displayed in a playback window, and controls 2104 for playing, pausing, adjusting the volume, and navigating the video can be provided. In addition, there can be controls for selecting different views 2106, i.e., controls that allow the viewer to select video clips captured by different cameras during the shot sequence, and can allow the viewer to view different views of the shot sequence individually or in a combined view (such as a side-by-side view). The views can be synchronized so that the viewer can see different views of the same event simultaneously.
[0160] The user interface 2000 may additionally display a target 2110 and may identify target hits recorded by the system 2112. The user interface 2000 may additionally display a score for the most recent shot 2114 and an average score for the shot string 2116. Of course, other information may be displayed as desired by the user, which may include training tips, motion data that resulted in off-center shots, or other errors during a shot string.
[0161] Figure 22 An additional view of user interface 2000 is shown, in which a user can specify body marker selections 2202. Additionally, user interface 2000 provides a selection of signal settings 2204, which allows the user to specify details of either Y-coordinate motion or X-coordinate motion. In response to the user selection, user interface 2000 may display motion data 2206 associated with the selection. This type of review and analysis allows the shooter to view the movement of individual body markers during a shot string in great detail, and may also specifically view horizontal movement, vertical movement, or both for review.
[0162] Figure 23 Some features and techniques employed by embodiments of the described systems and methods are shown. In many embodiments, the disclosed system utilizes machine vision (e.g., computer vision) 2302 to track the body landmarks of a participant (e.g., a shooter), and also tracks changes in the target to provide automatic target scoring 2304. The systems and methods may also utilize audio processing 2306 to implement shot time detection 2308, which may also be combined with computer vision techniques. The disclosed systems and methods may also utilize signal processing 2310 to provide position correction 2312 of the shooter, including posture, gesture, movement, grip, trigger pull, etc. The systems and methods described herein also apply machine learning 2314 to determine the cause and effect of off-center shots, which may include analysis of the shooter's posture, grip 2316, trigger pull, gesture, body landmark movement, etc.
[0163] Figure 24 A method 2400 according to an embodiment described herein is shown. The system may receive video data and optionally audio data. The system may be configured to detect shots 2402 and score the target by image processing of the video data. The system may process the video data to determine the shot fired and the shooter's posture analysis 2404. Additionally, the system may analyze the audio data for shot detection 2406.
[0164] Shot detection 2402 provides the coordinates of the shot within the target (e.g., x, y coordinates) 2408. Posture analysis of the shooter 2404 provides the coordinates of body landmarks (e.g., x, y coordinates) 2410 during the shooting process. Audio shot detection 2406 provides the exact time of the shot 2412. Shot analysis may include one or more machine learning algorithms that receive shot and body data and detect and predict, through the machine learning algorithm, causal relationships between the shooter's shooting performance and the firearm and ammunition combination, which may be referred to as performing shot analysis 2414. For example, embodiments of the system may generate one or more of a session score 2416, a shot position recommendation 2418, a shot error and grip analysis 2420, and the like.
[0165] Figure 25 An example process flow 2500 for using machine vision and machine learning to track a participant's body landmarks and generate recommendations for improvement is shown. As discussed herein, the systems and methods described herein can be used for any event that benefits from repeatability and accurate body kinematics, such as a sports event. In addition to shooting, some such events include archery, golf, bowling, darts, running, swimming, pole vaulting, football, baseball, basketball, hockey, and many other types of sports. In any event, the system is configured to track the participant's body landmarks and determine how to change body motion to improve performance.
[0166] At block 2502, the system receives video data of a participant. This can come from a single image capture device, or two image capture devices, or three or more image capture devices. The image capture device can be any suitable imaging device configured to capture sequential images of a participant, and can include any consumer or professional-grade video camera, including cameras that are routinely incorporated into mobile computing devices.
[0167] At block 2504, the system determines one or more body landmarks of the participant. The body landmarks can be associated with any joint, body part, limb, or location associated with a joint, limb, or body part. In some cases, the system generates a wireframe based on the one or more body landmarks and may use fewer than all body landmarks when generating the wireframe model.
[0168] At block 2506, body markers are tracked during the performance of the activity to generate motion data. As non-limiting examples, body markers can be created for a golfer's hands, wrists, arms, head, shoulders, torso, waist, knees, ankles, and feet, which can be tracked during a golf club swing.
[0169] At block 2508, a score is determined and associated with the performance. As described, in activities involving missiles, the score can be associated with the path or destination of the missile. For example, in golf, the score can be based on distance, direction, proximity to the target, or a metric associated with the participant's average or past performance. In short, any metric can be used to evaluate the quality of the performance results.
[0170] At block 2510, the score is associated with the performance. That is, a relationship is established between the performance and the determined score, which can be stored for later analysis and to determine trends in the performance over time, or to compare one performance to another.
[0171] At block 2512, the system generates recommendations for changing the motion of the subsequent performance to improve the result. In some cases, the recommendations may involve the hands, including gripping of a firearm, golf club, bat, club, etc. The recommendations may also include changes in weight distribution or shifting. The recommendations may include motions of the hands, head, shoulders, torso, legs, feet, or other body parts. In many cases, the recommendations include suggestions for changing the motion of one or more body markers in an effort to improve the performance score in the subsequent attempt.
[0172] Acoustic Firearm Signature
[0173] According to some embodiments, a system is described that can use consumer-grade audio recording devices (e.g., iPhone, tablet computer, phone, camera) to quickly calculate the signature of the acoustic sound of a shooter's firearm firing a specific ammunition in a given environment. In some cases, the system can receive input audio or audio / video files and calculate the signature of the firearm based on the sound in the captured recording. The recording can be captured by any suitable audio and / or video capture device, such as, but not limited to, security cameras, traffic cameras, video cameras, television cameras, mobile device recorders such as smart phones or tablet computers, and other capture devices. In some cases, the capture device is an easily available consumer-grade recording device. The acoustic signature can also include determining other characteristics of the brand, model, silencer, ammunition type, ammunition manufacturer, and firearm explosion (firearm blast) based on the acoustic signature from the discharged firearm (discharged firearm).
[0174] This capability could be used to detect and separate a shooter's shots from those of other shooters on a typical shooting range, which could be useful, for example, for automated scoring. A version of this capability could be used to identify specific firearms and ammunition from sound recordings.
[0175] In some cases, the described solutions operate in near real time on a single consumer recording device, such as a mobile phone, using only a moderate amount of training data. As used herein, the terms "real time" or "near real time" are broad terms, and in the context of the present disclosure, relate to receiving input data, processing the input data, and outputting the results of the data analysis, with almost no human-perceptible delay. In other words, a system as described herein that outputs analysis data in less than one second is considered to be near real time. Systems that operate in real time or near real time can limit the amount of computation required to machine learn, or at least train a model, which can be used to characterize a particular firearm that fires the appropriate ammunition. Additionally, there are some methods that do not use ML techniques, meaning that they do not adjust ("train") the model on a large number of samples. Some of these other methods can work well, but can also be computationally intensive, or shift the computational load to shot recognition time, for example, searching a large dictionary of signatures to obtain a "match."
[0176] Existing methods that use artificial intelligence (AI) to analyze gunshot sounds have proposed a direct method using deep neural networks (DNNs) to classify shooting sounds according to firearm and ammunition type. In existing methods, DNNs must be pre-trained on a large number of shooting sounds for each firearm and appropriate ammunition type. These sounds may not be captured on the same recording equipment or under the same shooting conditions. This potentially improves the generalization ability of the trained DNN instance to classify firearm and ammunition types regardless of the environment and recording equipment. However, it reduces its ability to identify shooting sounds of specific firearms and ammunition types using any recording equipment in any environment.
[0177] Existing methods describe a two-step approach to classification. For example, some existing methods may use a pre-trained instance of a relatively general pre-trained DNN instance as an approximate classifier (predictor) and use an additional trainable step on the classifier output to improve the DNN prediction. In some cases, the DNN treats a finite length segment of the time-varying spectrogram of the shooting sound as an image and distinguishes the pooled image for each firearm and ammunition pair from the pooled images of other firearm and ammunition pairs. This approach has several disadvantages, including using a general DNN instance that is not particularly good at classifying sounds from different environments or different ammunition.
[0178] According to some embodiments, the described system overlays a trainable weighted mask onto an image, created based on time periods of a time-varying spectrogram and a direct training method for adjusting the mask using only a few instances of the sounds of a shooter's firearm and ammunition, recorded with specific equipment in a specific environment. Training only the weighted mask is significantly less computationally intensive than training the DNN itself. Thus, according to some embodiments, the DNN can be trained without training, or to a much smaller degree than existing methods, while the weighted mask receives training. This approach has several benefits.
[0179] For example, in abstract terms, the adjusted mask can be viewed as a signature for the combination of firearm, ammunition, environment, and recording device. Rather than searching a catalog of signatures for that combination of factors, this signature is used to condition the input spectrogram to enhance the DNN's classification performance for firearms and ammunition in a specific environment using a given recording device. The enhanced DNN classification can then be used to improve the detection of the shooter's shots and distinguish between the shooter's shots and those of other shooters.
[0180] This approach demonstrates significant advantages in terms of computational intensity, resulting in faster analysis and can be applied to many firearms across many different environments. For example, because the training dataset for a specific shooter is small, the weighted mask can be updated online using stochastic gradient descent or offline using gradient descent. In other words, the training data can be tailored to a specific shooter using a specific firearm and ammunition combination in a given environment, resulting in a much smaller dataset than if data points were aggregated from numerous shooters across different environments.
[0181] Additionally, both algorithms estimate the gradient of the function computed by the DNN at each iteration of the weight matrix update. While this does not result in excessive computational overhead, it may result in only the sign or a coarsely quantized version of the gradient being sufficient. This simplification depends heavily on whether the DNN computes a monotonically increasing or decreasing function for each input. There is a large literature studying the properties of the functions computed by DNNs. However, finding a short answer to this question is not easy. In some cases, DNNs do this if all internal layers use linear or affine activation functions and the output layer uses a monotonically non-decreasing activation function.
[0182] Thus, in some embodiments, the disclosed system is capable of rapidly fingerprinting firearm and ammunition pairs used in a specific environment using a specific recording device. In other words, the systems and methods described herein can very quickly determine the acoustic signature of a firearm and ammunition combination in an environment. This allows the system to distinguish the analyzed acoustic signature from other firearm and ammunition combinations. This is particularly useful in crowded shooting ranges, where distinguishing one shooter from another is important, for example, when combined with a system that performs automated target scoring. By being able to distinguish between firearm and ammunition pairs, the scoring system will have fewer false positives and missed shots, as the system can determine that a specific firearm and ammunition combination was used, which can be temporally matched to a hit on the target. In some examples, the system can be executed on a mobile computing device and used at a shooting range. If the mobile computing device has a microphone pointed in the general direction of a shooter of interest, shots fired by the shooter of interest will typically have an audio file dominated by shots fired by the shooter of interest, which can help determine whether the shot fired originated from the shooter of interest. In some cases, the recording device (e.g., a mobile computing device) has a microphone pointed at the target area, such as where the mobile computing device has a camera pointed at a target, such as for automatic target scoring, in which case the audio from the shots fired by the shooter of interest will have a volume that is more difficult to distinguish from other shooters at the range. In these cases, the described embodiments can use training of a machine learning algorithm or training of a weighted mask to quickly distinguish shots fired by the intended shooter from shots fired by all other shooters at the range.
[0183] In some cases, classifiers rely on feature extraction in the form of time-spectrograms. Mel-frequency cepstral coefficient (MFCC) vectors can be used as an alternative to raw power spectral density vectors (PSD) for frequency representation of spectrograms. Initially, Mel-frequency cepstral coefficients (MFC) were developed for speech processing and information was considered to be a way of encoding in speech waveforms. In some cases, MFC can be applied to firearm fingerprinting purposes. As used throughout this disclosure, "firearm fingerprint" is used to refer to an acoustic signature, or a specific sound or sound image that identifies a particular firearm and ammunition combination.
[0184] In some cases, MFCCs can be generated by performing a series of steps, including but not limited to: i) windowing a signal segment and computing a fast Fourier transform (FFT) of the signal segment, ii) combining the linear FFT coefficients into MEL frequency filter bank coefficients, iii) taking the logarithm of those coefficients, and iv) computing a discrete cosine transform (DCT) of the logarithmic MEL filter bank coefficients. The FFT of a signal segment typically results in a peak at the applied frequency and other peaks called side lobes, which are typically on either side of the peak frequency. In some cases, the DCT expresses a finite sequence of data points as a sum of cosine functions oscillating at different frequencies. In some cases, fewer or more steps than those disclosed can be implemented to derive a firearm fingerprint based on the sound of a gunshot. For example, in some cases, only steps i), ii), and iii) described above can be used for the sound of a gunshot.
[0185] According to some embodiments, MFCC or PSD coefficients may be used as input to a DNN that is trained to classify the time-frequency spectrogram into a certain number of independent classes. For example, MFCC and / or PSD coefficients may be input to two classes as proposed, or a certain number of classes corresponding to (firearm, ammunition) pairs, or some larger number of classes corresponding to tuples of k>2 attributes.
[0186] In some cases, the DNN can place the new gunshot sound "near" any category represented by the maximum classifier output among all classifier outputs. For example, the gunshot sound can be initially classified using a nearest neighbor method (such as the k-NN algorithm). Subsequent analysis can further classify the gunshot sound.
[0187] The inner layers of a DNN can represent different sets of attributes of the input. These can also be used to classify and train the sound of gunfire.
[0188] From an optimization theory perspective, training a DNN essentially defines a surface with multiple local optima, and the DNN can be thought of as guiding new inputs to the most suitable local optimum. In some embodiments, following a pre-trained DNN with one or more trainable layers can be thought of as extracting a different set of more optimal properties for the sound of a gunshot.
[0189] In some cases, before a pre-trained DNN has a trainable layer, weighting coefficients can be included in some cases, which can be thought of as tuning the pre-trained DNN so that the shot of interest to the shooter is the most positive example of all shots placed “near” any class represented by the maximum classifier output among all classifier outputs placed by the pre-trained DNN. In some cases, the decision threshold can be adjusted to optimize the confusion matrix for the training dataset used to tune the trainable input layer or another test dataset for some useful criterion.
[0190] According to some embodiments, the FFT of a windowed segment of the shot audio can be computed in O(n log n) time. Computing MFCCs has the same order of time, but involves computing two O(n log n) operations. In some cases, MFCCs may not be significantly better than FFTs, and they can be omitted.
[0191] According to some embodiments, the MEL frequency log spectrum (omitting the discrete cosine transform (DCT) that produces the cepstrum) can be an improvement over the original FFT, with only an additional computational cost of O(n). Therefore, in some examples, the MEL frequency log spectrum is used instead of the DCT that produces the cepstrum.
[0192] Typically, there is no fast form like the FFT for computing functions computed by a DNN (composed of convolutional neural network (CNN) layers). The FFT exploits regularities in the FFT kernel that are not typically present in the composition of an essentially arbitrary CNN. However, the examples described herein have an FFT of the input time signal, and a pre-trained DNN composed of a CNN can be implemented in the frequency domain.
[0193] As a non-limiting example, the following DNN-based, rapidly trainable classification recognizer may be implemented to quickly determine the acoustic signature of a firearm and ammunition combination.
[0194] Let f:RM→RK denote the analog transformation of a trained deep neural network from a real-valued M-dimensional input vector of class probabilities to a K-dimensional vector. The final discrete output mapping Ψ:RK→N≤K selects the most likely K classes.
[0195] The system can be configured to expand the trained DNN into an enhanced binary classifier for class k. In some cases, the system can add a rapidly trainable input stage to the DNN, which is implemented as a Hadamard product "⊙" of a weight matrix W and an input vector x. In some cases, the Hadamard product is a binary operation that receives two matrices, such as the weight matrix W and the input vector x, and returns a matrix of corresponding elements multiplied together. The system can then follow the DNN class probability vector output with a selector function:
[0196] Γ: RK×N≤K→R, which provides the single-category probability of category k. The enhanced DNN can implement the following function:
[0197] y(n)=Γ(f(W⊙x(n));k).
[0198] Suppose we have the set of additional input vectors All instances of the same class k. The system can be tuned to better recognize similar instances of class k by any of a number of suitable means. For example, one approach is online learning, which is commonly used with large training datasets. For online learning, we initially set W = 1 and then sequentially update the vectors in ζ using stochastic gradient descent as follows:
[0199]
[0200] where 0 < α < 1 is an adaptation constant that weights the relative contribution of ζ to W. For each vector in ζ, the target value d(n) = 1.0. Iterate until convergence for each z in ζ. Here, relu[...] represents the element-wise relu() of the parameter matrix.
[0201] The discharge of a firearm produces multiple acoustic events, such as the muzzle blast caused by the expansion of gases within the chamber and their expulsion through the barrel, and the ballistic shock wave generated by the projectile, which is in most cases supersonic, but in some cases can also be subsonic. The acoustic events are the result of variables that generate the firearm signature and can include firearm type, make, model, barrel length, ammunition type, powder quantity and identity, projectile weight, and projectile shape, among others.
[0202] Figure 26 Shown is an example process 2600 for associating a firearm with a specific shooter according to some embodiments. When a shooter visits a shooting range for practice, the shooting range may have other shooters who also fire their firearms at the shooting range. In some cases, a busy shooting range may have 10, 20 or 30 or more shooters who may all be shooting at the same time. For an acoustic system, it is quite difficult to record the shooting that a shooter of interest fires by the cacophony of firearm firing. In some cases, the system is configured to distinguish the firearm of a shooter of interest from the firearms of other shooters at the shooting range.
[0203] For example, at block 2602, the system receives video data of the shooter, which also includes audio data. This can be received, for example, via a multi-camera system, a dedicated microphone, or a consumer-grade audio / video capture device (such as a mobile computing device).
[0204] At block 2604, the system may determine the shooter's body landmarks.
[0205] At block 2606, the system may track one or more body markers as the shooter fires a shot and generate motion data associated with the body markers during the shot.
[0206] At block 2608, the system associates audio data with motion data to determine that the shooter of interest has launched shooting. In some cases, the system can receive the audio data indicating that the firearm is launched, which can be represented as a peak in an audio wave file. This can be associated with motion data (such as the shooter's wrist), and the recoil from the firearm indicates that the shooter's hand is displaced, thereby indicating that shooting has been launched. In some cases, the system is trained to distinguish the shooter's firearm from other firearms at the shooting range. In this case, the system can be trained to distinguish the shooting launched from the shooter of interest and other shooters at the shooting range.
[0207] At block 2610, the system may determine a score for the shot and associate the score with the athletic data that resulted in the score.
[0208] At block 2612, the system can analyze the motion data in conjunction with the score and determine any errors made by the shooter and provide suggestions to identify the errors and / or how to address them in the future. The system can also provide training exercises to allow the shooter to address the errors and improve his shooting performance.
[0209] refer to Figure 27, which illustrates an example system 2700 using online learning, receives audio 2702 and converts it into a delta spectrogram 2704. The spectrogram 2704 is framed 2706, for example, by time windowing, and used to determine MFCCs, for example, by determining the FFT of the windowed signal, combining the linear FFT coefficients into MEL frequency filterbank coefficients, determining the logarithm of the coefficients, and determining the DCT of the logarithmic MEL filterbank coefficients. The determined MFCCs can be input into a rapidly trainable input stage 2708 and then passed to a DNN multi-class classifier 2710. A selector 2734 determines the classifier output k with the highest class probability to classify the shot. The rapidly trainable input stage 2708 can be referred to as a trainable weighted mask, or simply a mask. The mask can be adjusted using examples of shooting sounds, such as those recorded from a shooter's firearm and ammunition in a specific environment. In some cases, training only the weighted mask is significantly less computationally intensive than training a DNN. The trained (e.g., adjusted) weighted mask can represent a signature for the combination of firearm, ammunition, environment, and recording device. This signature is not necessarily used to search the signature catalog, but is used to adjust the input spectrogram to enhance the DNN's classification performance for firearms and ammunition in a specific environment using a given recording device. In some cases, the training data set for a specific shooter (e.g., a specific firearm and ammunition combination) is small, so the weighted mask can be updated using an online method (such as by using stochastic gradient descent) or offline (such as by using gradient descent techniques).
[0210] The inner layers of the DNN classifier 2710 can represent different sets of attributes of the input. In some cases, the DNN is trained to define a surface with multiple local optima, and the DNN can be used to guide new inputs to the most appropriate local optimum. The enhanced DNN classification can then be used to improve the detection of shots from a firearm and the differentiation of shots from a firearm from shots from other shooters.
[0211] The weighted mask W of the input stage 2708 to the pre-trained DNN classifier 2710 can be adapted by an online learning loop 2712 that optimizes the weighted mask W.
[0212] The online learning loop 2712 includes copies 2714, 2716, and 2718 of the shot classifier input stages 2708, 2710, and 2734. When a new shot n is detected, a subtractor 2720 in the online learning loop 2712 compares the result of the selected output k from the copies of the shot classifiers 2714, 2716, and 2718 to the expected classification d(n) = 1.0 to calculate the classification error term.
[0213] Block 2730 of the adaptation loop computes the vector sign of the gradient of the DNN output relative to the spectrogram input of the current output spectrogram x(n) from the framing block 2706.
[0214] The raw incremental adjustment to the current weight vector Wi(n) is then calculated by Hadamard multiplier 2732 as the product of the current shot spectrogram from framer 2706 and the sign vector of classifier gradients 2730. Vector multiplier 2722 then scales the raw incremental adjustment to the weight vector Wi(n) by an arbitrary value α.
[0215] Finally, vector adder 2724 calculates preliminary updated weights by adding the current delta adjustment to the current Wi(n). The preliminary weights are converted to updated weight vector Wi+1(n) via vector relu[] operation 126.
[0216] The weight adaptation loop 2712 just described is iterated, as represented by the delta operator 2728, until the weight vector converges, e.g., |Wi(n)-Wi+1(n)|≤δ. The result Wi+1(n) is then selected as the weight vector W(n+1) for classifying the next shot.
[0217] Offline learning can be used, for example when ζ is small. In some cases, offline learning can be initiated by setting W = I and updating it using gradient descent:
[0218] Wl+1=relu[Wl+α·1 / M, for i=1 to M:
[0219]
[0220] In some cases, offline learning can require roughly the same amount of computation as online learning. However, online learning has an advantage over offline learning in that stochastic gradient descent does not require access to the entire training dataset ζ in each iteration of the update function, while offline learning does. Offline learning offers the advantage of finding a local optimum, while online learning may only approximate that local optimum.
[0221] Although the gradient evaluated at the current parameter Wl⊙zi The computational cost of is not too high, but if the DNN normalizes the input, pre-computing or It may be sufficient to have , where 1 is the unit vector. In some cases, this can significantly simplify offline learning in this context, at the expense of approximating the local optimum as in online learning.
[0222] While some systems aim to enhance DNNs for shot classification, they do so by adding layers to the output of a pretrained network to tailor it to a specific task. The pretrained network extracts many levels of increasingly abstract features, such as from an image, and additional trainable layers are used to focus on the problem of interest. In contrast, many of the systems and methods described herein operate in a very different manner, resulting in systems that are more efficient, faster, and more accurate. In many cases, the systems described herein only add a single layer to the input of the pretrained network. In some cases, the camera is pointed at the target rather than focused on the shooter, and the camera and microphone receive other shots, making it impossible to locally determine whether the shot originated from the shooter of interest. The acoustic signature is generated quickly and, in many cases, executed on a mobile device (e.g., a smartphone) that includes a camera and microphone. The mobile device can execute instructions (e.g., an application) that include the components and systems described herein, enabling classification and acoustic signature determination to be performed on the mobile device. In some cases, the pretrained network can be trained on a relatively comprehensive set of shots. The trainable input layer can pre-distort the input data to enable the pretrained network to recognize with high probability the specific firearm and ammunition and environment. This can then increase the probability of detecting shots of interest and rejecting all others.
[0223] According to some embodiments, the systems and methods described herein will converge to a useful weighted mask where the gradient of the function computed by the DNN is monotonically non-decreasing or monotonically non-increasing in each input.
[0224] refer to Figure 28 , which shows a pre-trained DNN, process 2800 begins at 2802 and opens an audio file at block 2804, which can be the first audio file, or the next or subsequent audio file. Using the audio file, and at step 2806, the system captures a block of samples that frame the first and / or next shot. In other words, each shot is windowed into time-limited samples. At block 2808, a spectrogram associated with the sample is generated and labeled with attributes.
[0225] At block 2810, the system determines whether the most recent sample is associated with the last shot, and if not, the system returns to block 2806 to capture a block of samples associated with the subsequent shot. If so, the system proceeds to block 2812 and determines whether the labeled spectrogram created at block 2808 is the last file. If not, the system returns to block 2804 to open or capture the next audio file. If the system determines that the most recent file is the last file, the system proceeds to block 2814, where the labeled spectrograms are aggregated. At block 2816, a DNN is trained on the labeled spectrograms. The system stops at block 2818 with the trained DNN.
[0226] refer to Figure 29 , which illustrates deriving spectrogram weights, process 2900 begins at 2902 and captures a collection of shots at block 2904. The shots may be captured by an audio and / or video recording device, or may include opening a file associated with one or more shots. At block 2906, the system captures a block of samples framing the first and / or next shot. In other words, each shot is windowed into time-limited samples. At block 2908, a spectrogram associated with the sample is created and labeled with attributes.
[0227] At block 2910, the system determines whether the most recent spectrogram is associated with the last shot, and if not, the system returns to block 2906 to capture a block of samples associated with the subsequent shot. If so, the system proceeds to block 2912 and determines whether the tagged spectrogram created at block 2908 is the last set. If not, the system returns to block 2904 to capture the next set of shots. If the system determines that the most recent file is the last file, the system proceeds to block 2914, where the tagged spectrograms are aggregated. At block 2916, the system updates W (weight) until the value converges. The system stops at block 2918 with the spectrogram weights.
[0228] refer to Figure 30 , shows a process 3000 for classifying a shot. The process begins at block 3002, and at block 3004, the system captures a block of samples framing the shot. At block 3006, the system determines the spectrogram associated with the block of samples. At block 3008, the system determines the Hadamard product of the spectrogram and the weight matrix. At block 3010, the Hadamard product of the spectrogram and the weight matrix is applied to the DNN. At block 3012, the system makes a binary decision, such as whether the shot is associated with or not associated with the signature of the firearm in question. In some cases, the system is able to identify the type of firearm and the ammunition fired by the firearm. For example, when receiving an audio sample, the system, without any prior knowledge of the firearm in question, can determine that the acoustic signature of the firearm in the audio sample corresponds to a 230-grain round-nosed projectile fired from a Beretta .45ACP. The process stops at block 3014.
[0229] The system may include one or more processors and one or more computer-readable media that may store various modules, applications, programs, or other data. The computer-readable media may include instructions that, when executed by one or more processors, cause the processors to perform the operations described herein for the system.
[0230] In some embodiments, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), both the CPU and GPU, a microprocessor, a digital signal processor, or other processing units or components known in the art. Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), etc. In addition, each of the processors may have its own local memory, which may also store program modules, program data, and / or one or more operating systems. One or more control systems, computer controllers, and remote controllers may include one or more cores.
[0231] Automatic target scoring
[0232] In some cases, the described systems operate in near real time on a single consumer device of record, such as a mobile phone, using only a modest amount of training data. As used herein, the terms "real time" or "near real time" are broad terms and, in the context of the present disclosure, relate to receiving input data, processing the input data, and outputting the results of the data analysis with little to no human-perceptible delay. In other words, a system as described herein that outputs analysis data in less than one second is considered to be near real time. Systems that operate in real time or near real time can limit the amount of computation used for machine learning, or at least training models, which the method can use to characterize specific target acquisition, classification, and scoring.
[0233] Existing methods for automatic target scoring rely on acoustic triangulation, optical triangulation, and piezoelectric sensor triangulation. Acoustic triangulation has been attempted using an acoustic chamber target, which uses the Mach waves of a missile to determine the position of a missile as it passes through a target. An acoustic triangulation automatic scoring system operates by using microphones to measure the sound waves of a missile as it passes through a target. Thus, the sound of a missile passing through a target from multiple audio sensors (e.g., microphones) can be used to determine the position of a missile passing through a target.
[0234] Optical triangulation automated scoring systems use three or more lasers, such as infrared lasers, to triangulate the position of a projectile as it passes through a target. Piezoelectric sensor triangulation systems rely on an array of piezoelectric sensors on a board that sense the vibrations caused by the projectile striking the target.
[0235] Figure 31System 3100 is shown, configured to automatically identify targets, classify targets, determine impacts on targets, and score shooting strings. System 3100 may include a computing resource 3102, which may be a mobile computing device associated with a participant at a shooting range. The computing resource may include any one or more of a variety of mobile computing devices, such as a smartphone, tablet, laptop, or other suitable computing device. Computing resource 3102 generally includes one or more processors 3104 and a memory 3106 storing one or more modules 3108. Modules 3108 may store instructions that, when executed, cause one or more processors 3104 to perform various actions. Computing resource 3102 may also include a data storage device, which may be a remote storage device (such as a remote server or cloud-based storage system), a local storage device, or a combination thereof. The data storage device may store data on past target hits (DOPE) (which may allow for data tracking over time) as well as comparative data between different shooters, different firearms, different ammunition, different targets, different environments, and the like.
[0236] The storage system may also allow for historical trend analysis, which can be used to show a shooter's performance over time, including tracking improvements. The data store may also be analyzed to provide performance predictions, rankings, social features, and other benefits.
[0237] The system may include one or more imaging sensors 3110, such as any suitable camera. In some cases, the imaging sensor 3110 may be associated with the computing resource 3102. For example, in some embodiments, the computing resource 3102 may be a smartphone with a built-in camera 3110.
[0238] The camera 3110 can be pointed to capture an image of a target 3112. The target can be located at any distance from the shooter, and the camera 3110 can be aimed and / or zoomed to capture an image of the target. In some embodiments, the camera can be coupled to a lens (such as a spotting scope or camera lens) to allow the camera to obtain a closer view of the target through optical zoom or digital zoom.
[0239] The computing resources 3102 may include instructions (eg, module 3108 ) that allow the computing device to initialize a target 3114 , detect an impact on the target 3116 , and score the impact on the target 3118 .
[0240] Figure 32A decision tree 3200 is shown that is configured to detect, identify, and classify targets. According to some embodiments, before the system begins searching for a target, the system does not know the type of target. For example, in some existing systems, the scoring system can be pre-programmed with the target that the shooter will aim at. This makes it easy for the system to understand the size and shape of the target, as well as the location and boundaries of each scoring ring or area. In the illustrated embodiment, the system is configured to automatically determine and determine the scoring rings and areas without prior data about the target type. For example, the system can have one or more video capture devices that can be integrated into one or more mobile computing devices. As used herein, a mobile computing device can be one or more of the following: a mobile phone, a smartphone, a tablet computer, a laptop computer, a personal digital assistant, smart glasses, a body cam, a wearable computing device, or some other computing device that a user can carry to a shooting range.
[0241] The mobile device may activate its camera and capture one or more frames of a target 3202. The computing device may be instructed to analyze the one or more frames using any suitable image analysis algorithm to identify the target in the one or more frames. If the target is detected and classified at block 3204, the target is registered with the system 3206, and scoring rings and regions are determined. The system may capture additional image frames containing the target and look for differences from one frame to the next that may be associated with an impact on the target. These frames may be compared, and at block 3208, a moving average may be generated. Moving averages are a fundamental mathematical and statistical technique used in image analysis and machine learning for various purposes, including noise reduction, feature extraction, and trend analysis. They involve calculating the average of pixel intensities or other data points within a moving window or kernel across an image or dataset. Moving averages can be used to extract meaningful features from an image. For example, by sliding a small window across an image and calculating the average pixel value within that window, important information can be highlighted. For example, in edge detection, a moving average can emphasize areas of sudden changes in pixel intensity, helping to identify the edges or boundaries of targets and scoring regions. Edge detection can also be used to identify impacts on a target.
[0242] In some examples, moving averages are used in time series data analysis. For example, to detect anomalies, moving averages can be used to establish a baseline behavior for a system. Any data point that deviates significantly from this baseline can be flagged as an anomaly or outlier. These anomalies can then be further analyzed to determine their impact on the target.
[0243] As the sequence of moving averages is generated, they can be combined into a long term moving average. At block 3210, the moving average image can be compared to the long term moving average to determine a difference from one frame to a subsequent frame that indicates a change in the target that is most likely associated with an impact on the target.
[0244] At block 3212, the impacts are selected and classified. For example, the system determines the boundaries of the scoring ring and determines the location of each impact and associates the location of each impact with a score for that impact.
[0245] Returning to block 3204, in the event that a target has not been previously detected and classified (such as when the shooter initialized the system or changed targets), the system determines whether a target has been detected at block 3214. If a target has not been detected, the system continues to detect targets at block 3216. If a target has been detected, the system classifies the target at block 3218, for example, by identifying the boundaries of the target, the boundaries of the scoring ring, and the values of the scoring ring.
[0246] In the event that the system has not yet detected the target, the system can capture one or more additional image frames and analyze the one or more additional image frames to determine that the target is within the field of view of the imaging device. Once the target is detected, the system can classify the target to determine the size of the target and the relative position and size of the scoring ring or area.
[0247] Figure 33 Further described are the initial steps that the system can take to identify and classify a target 3300 by analyzing one or more image frames. Object detection is a computer vision technique that involves identifying and locating multiple objects in an image or video stream. Unlike image classification, which determines whether a single object category is present in an entire image, object detection provides a more refined understanding by not only identifying the object but also specifying its location via a bounding box. In some embodiments, the object detection algorithm typically outputs a bounding box that surrounds the detected object. These bounding boxes consist of the coordinates (x, y) of the upper left corner of the object and the dimensions (width and height) that define the spatial extent of the object in the image.
[0248] At block 3302, the system applies object detection to one or more images of a target and searches for the target. In some embodiments, the object detection model is generic for the target, which allows the system to detect any target regardless of size or shape. At block 3304, if the target is found in the same location in subsequent images (e.g., 2 or more images, 3 or more images, 4 or more images, etc.), the system assumes that the system has located the target and defines a bounding box around the target. In some cases, finding the target in the same location in subsequent images includes determining a moving average of the images to determine the target location, size, and shape.
[0249] At block 3306, the objects are optionally classified. In addition to locating objects, the system can be configured to detect objects and classify each detected object into a predefined class or category. This allows the system to distinguish between different object types, such as circular objects, oval objects, rectangular objects, silhouette objects, or other types of objects.
[0250] The object classifier can be applied to the image within the bounding box. Thus, the system determines which reference object image to apply.
[0251] At box 3308, the system aligns the target with the reference target image. In some cases, this involves applying a contrast adjustment to the image. This can also involve iteratively modifying the initial bounding box, such as by adjusting the corners of the bounding box and then projecting the adjusted bounding box onto the reference target image. The difference between the two can be applied as a score, and a hill climbing technique can be applied to find the best corner, which can be associated with the initial position of the image. Hill climbing is an optimization algorithm for finding the local maximum (or minimum) of a given objective function. By iteratively taking small steps in the direction to produce higher values, the algorithm determines the highest value, the lowest value, and therefore can be used to determine the boundary of the target. In some cases, object detection is combined with semantic segmentation to provide pixel-level object masks. This allows a more accurate understanding of object boundaries within the image, such as target boundaries and scoring ring boundaries.
[0252] At box 3310, the system has been initialized and registered with the target and begins looking for shocks in the subsequent moving average.
[0253] Figure 34A process 3400 of aligning a target 3300 and determining an impact on the target is described. At block 3402, the target can be re-aligned, such as by performing a hill climbing technique to search for a new optimal set of corner points for the target. In some cases, the hill climbing search uses mean squared distance in a perceptual space technique. For example, mean squared distance (also known as mean squared perceptual error) is a metric used to measure the similarity or dissimilarity between two data points involving perceptual data (e.g., target corners). In some cases, for each data point, relevant perceptual features are extracted, in this case, the features are target corners, edges, scoring rings, etc. The features can be visual descriptors, which can be represented as vectors of perceptual features. These feature vectors capture the relevant information for each data point in a simplified and more informative form.
[0254] The mean squared distance between two points (represented by their respective feature vectors) is generated by determining the squared differences between corresponding features and calculating the mean of these squared differences. The resulting mean squared distance provides a quantitative measure of the dissimilarity between two data points in the perceptual space.
[0255] At block 3404, the system can apply a transformation matrix that can be used to map the set of corner points to an image. In some cases, the image onto which the coordinates are mapped has a size of 160 pixels, and in some cases less than 160 pixels.
[0256] At block 3406, the moving average image is updated, and in some cases, the long-term moving average is about 10 seconds or longer, while the short-term moving average is about 0.1 seconds. In some cases, the camera can capture more than 30 frames per second. For the short-term moving average, this is equivalent to averaging about 3 frames to determine the short-term moving average.
[0257] At block 3408, the system determines the difference between the moving averages. For example, a long-term moving average would be associated with a static target that has not changed for about 10 seconds and can be compared to a short-term moving average that reflects changes in the image. Thus, the difference between the short-term and long-term moving averages will highlight changes in the image, such as impacts on the target. The system can convolve any difference image with a simple impact kernel (in some cases, a 5×5, uniformly weighted square kernel) and find the maximal block-wise locations in the difference image. A kernel generally refers to a convolution filter that can be used to process and modify pixel values, such as for feature extraction. The square kernel can be convolved (or moved) across the image, and at each location, the kernel value can be multiplied by the pixel value in the corresponding neighborhood, and the results can be summed to produce a new pixel value in the output image. Of course, the kernel size can be varied to adjust the range of the neighborhood considered during the convolution and can include any of a set of aperture kernels with non-uniform weights and can have any suitable size.
[0258] The convolution returns a set of potential impacts on the target. This set of potential impacts can be further filtered, such as by using simple statistics on a window around the marked difference values. In some cases, the window is chosen to be 16×16, with the difference value in the middle of the window. Of course, other window sizes are entirely possible, and the pixel values described herein are merely illustrative of some embodiments. The system can also apply business rules to the windowed difference values, for example, the system should not detect multiple impacts at the exact same location.
[0259] At block 3410, an impact on the target is determined. In some cases, this is accomplished by passing the filtered set of differences to an impact classifier for scoring. If the difference score is above a threshold, the location is marked as an impact, and another window can be placed around the impact. In some cases, a 10×10 window is placed around the impact location in a long-term average, with the short-term average being for 5 to 10 subsequent frames. This ensures that the same impact is not detected again. Thus, the difference is windowed with a first window, and if the difference exceeds a threshold score, the difference is windowed with a second window that is smaller than the first window. The windowed difference associated with the short-term moving average can be added to a long-term moving average for at least 5 frames, or at least 6 frames, or at least 10 frames, or at least 12 frames, or at least 15 frames, or more. In some cases, when the difference score is below a threshold, the difference is marked as a false impact, and the system will not need to evaluate and classify it again.
[0260] According to some embodiments, the system may receive audio data associated with a shot fired and determine that a shot has been fired based on the audio data. In some cases, the audio data is associated with a target image, and the system may convolve a difference target image in response to the audio data indicating that a shot has been fired. In some cases, the system may not need to continuously convolve the difference image. In such cases, the system may determine that a shot has been fired based on the audio data, then update a short-term moving average and convolve the difference image to look for the shot. In some cases, the system is configured to distinguish shots fired by a user aiming at a target from other shooters at the shooting range. In this way, the system can know when a shooter of interest has fired a shot, even if there are other active shooters at the shooting range.
[0261] In some cases, audio data may be used for impact detection, such as by correlating the audio of a fired shot with an impact appearing on a target image.
[0262] Figures 35A-35C Initializing the scoring system by identifying and classifying an object is shown and described. In some cases, the system can automatically determine the boundaries of the object, while in some embodiments, user input can define the boundaries of the object. For example, using a human-computer interface (e.g., a touch screen, mouse, stylus, touchpad, etc.), a human can draw a boundary around the object to help the system identify the object. However, in many embodiments, the system uses machine vision to identify the object and its boundaries. Figure 35A An image 3500 captured by a camera associated with the system is shown. The image may include a target holder 3502, a target 3504, a target retaining clip 3506, and other features within the field of view. The system may determine an initial bounding box 3508 around the identified target, such as by using a trained target detection model. In some embodiments, a user may define the initial bounding box, for example, by drawing on a computer display using a human-machine interface. The human-machine interface may be any suitable interface, and in some cases may be a touch screen, pen, mouse, trackball, etc. The initial bounding box may not exactly conform to the edges and corners of the target, particularly in those cases where the bounding box is defined by the user. The initial bounding box and target image may be referred to as an initialization frame. The initialization frame may be converted to a Lab color space comprising the following components: luminance, a green to red axis, and a blue to yellow axis to generate perceptual uniformity. In some cases, the luminance channel is equalized using contrast-limited adaptive histogram equalization (CLAHE).
[0263] Figure 35BThe target is shown, where the coordinates of the target are determined as described above, and the coordinates in many cases imply a quadrilateral, which can be projected onto the reference target image 3510. The reference target image 3510 can also be converted to Lab color space, and the squared difference can be generated (in Lab space) between the projection and the target. A hill climbing algorithm can be applied to the coordinates, where the possible increments are small changes in the coordinates, and a better solution can be determined by squaring the perceptual difference.
[0264] Figure 35C The optimal coordinates, determined, for example, by minimizing the difference from several random restarts of a hill climbing algorithm, are shown. The coordinates can then be used to apply an updated bounding box 3512. Thus, even in cases where the target image is skewed, such as where the target appears to be a parallelogram rather than a rectangle from the camera's perspective, the initial bounding box can be modified to conform to the shape of the target as presented in the image captured by the camera.
[0265] In some embodiments, the system can define the edges of the target through image analysis; however, in some cases, the edges of the target are irrelevant and only the scoring rings are important. Therefore, in some cases, the system is configured to identify the scoring rings and not focus on the target boundaries. Furthermore, the system may not need to classify the target, but only need to identify the scoring rings. For example, the system may determine, through one or more machine learning models, that the target represents a central bullseye target with sequential scoring rings. The system can assign a score value to each ring, such as a bullseye with a score of 10, the next larger ring with a score of 9, and so on. Similarly, the system can identify a target with five bullseye-sized circles spaced throughout the target and assign a value of 10 to each of these scoring rings. One or more of the multiple bullseye-sized rings may have radially spaced larger scoring rings, which may be assigned a smaller value than the bullseye-sized rings. Thus, the system can omit the target classification step and focus solely on the size and position of the scoring rings.
[0266] Figure 36 Impact detection and scoring of detected impacts are shown and described. A machine learning model can be executed to determine whether a difference between image frames is likely a missile impact on target 3504. Target 3504 can be re-registered, for example, by applying a hill climbing technique to possible corner coordinates as in the target initialization step. The short-term moving average difference is compared to the long-term moving average, and the difference is windowed 3602a, 3602b, 3602c, generating a difference image 3612 between the current target and the long-term target average.
[0267] The difference image can be convolved with a shock kernel (e.g., a windowed kernel that is scanned over the image differences). Any point that exceeds a long-term exponential moving average (EMA) of the maximum convolution value by several standard deviations is marked as a possible shock 3604a, 3604b, 3604c.
[0268] The possible shocks 3604a-3604c are fed into a machine learning model (e.g., a classifier) 3606, which determines whether the difference is likely to be an actual shock. If the difference is above a threshold, the system labels the difference as an actual shock 3608. However, if the difference is below a threshold, the system labels the difference as a false shock 3610.
[0269] Figure 37 Scoring of impacts on target 3504 is shown and described. Different scoring zones can be determined by the system in the following ways: based on computer vision, by referencing registered targets from previous shooting activities, by retrieving stored target models from a known target database, or in some other way. Scoring zone 3702 on the target can be represented as a joint area of one or more simple shapes (e.g., ellipse, rectangle, circle, triangle, etc.). The coordinates of the detected impact 3704 can be normalized and converted to the axes implied by the reference image. In other words, the impact can be superimposed on the reference image, and the reference image can be used for the coordinates of the impact. The coordinates can be Cartesian coordinates represented by x and y values. In some cases, the coordinates can be radial coordinates, which represent the impact as an angle and distance (e.g., the distance from the center of the target). The system can then determine whether the impact is completely within a single scoring zone or intrudes into the scoring zone boundary, which allows the system to accurately score the impact. The system can use simple geometric shapes to determine whether any significant portion of a given impact is within any simple shape in each target zone.
[0270] In certain embodiments, the coordinates of the impact are used for further analysis by the system. For example, by generating and storing the coordinates of a given shooting string, the spread (grouping) can be quantified, which can be used as a measure of improvement over time. Similarly, the shooter's angular momentum (moment of angle, MOA) can be determined, which is a measure of the spread size from center to center and edge to edge in inches and minutes of anangle. Spread can also be used to define posture, grip or motion errors during the shooting string. Spread can be quantified, including the size of the spread, the rotation of the spread or other metrics.
[0271] Scoring can be quantified using any suitable metric. In some cases, scoring is based on points, with each zone of the target receiving points added or subtracted from the initial amount. In some cases, missed shots or extra shots fired are scored as negative or higher values, depending on the type of scoring. In some cases, timed scoring is used, where the total time is reflected in the score, and misses can be penalized by increasing the time. In some cases, spread size is used to determine the score, and extra shots or missed shots may penalize spread size. Of course, other metrics and combinations of metrics can be determined by the system used to score a particular shot string.
[0272] The system can be configured to return, for each shot in a string, the coordinates of the impact and the time at which the shot occurred in such a way that multiple metrics combining position and time can be used to range the string. For example, the time between shots can be measured, or the shots following a buzzer or other activation signal can be tracked and stored with a measure of accuracy.
[0273] In some cases, when a hit is identified, the system can draw a bounding box around one or more hits. When the shot string is complete, the system can draw a bounding box containing each shot within the spread and determine a metric based on the bounding box to determine a score.
[0274] like Figure 38 As shown, Figure 38 A user interface 3800 of a system developed and operable according to some embodiments described herein is shown, which can be configured to determine a bounding box 3802, which can pass through the center of the outermost impact or along the edge of the impact. The system can determine any of a number of metrics, such as, but not limited to, spread size 3804, total spread width 3806, spread height 3808, bounding box rotation angle, MOA, elevation offset 3810, windage offset 3812, and can further determine a shot distance 3814, which can be manually entered or determined based on the flight time of the detected projectile.
[0275] For example, the system can be configured to record the sound of a gunshot, the shockwave of a projectile or powder detonation, the movement of the firearm or shooter, or some other indication that a shot has been fired. The system can then detect when the impact on the target occurs and determine the time of flight of the ammunition and, based on the firearm, ammunition, and / or powder load, the target distance. This process can be accomplished in near real time using a simple consumer-grade mobile computing device. In some cases, the mobile computing device can utilize the zoom feature of a built-in image capture device. In some cases, an external zoom lens can be used by the mobile computing device to acquire image frames. For example, a mobile phone can be coupled to a scope that provides optical zoom through the scope to allow the mobile computing device to capture clearer images of targets that may be within the target area. Some mobile computing devices can rely on digital zoom to capture one or more images of targets positioned within the target area.
[0276] The system can further determine and display the number of shots fired in the current shooting string 3816, the average split time 3818 of each shot, which can be helpful for timing shooting events. The system can also show the score 3820 associated with each shot and the cumulative score 3822 of the shooting string.
[0277] Some embodiments also provide an automated and automatic scoring system that can quickly identify targets, classify targets, including identifying a scoring ring for a target, and score impact hits on targets at a shooting range. In some cases, the system is stored on a consumer-grade mobile computing device (e.g., an iPhone, tablet, phone, camera) and executed on these mobile computing devices. In some cases, the system includes a camera device pointed at a target of interest, and the system is configured to identify the target, classify the target, determine the impact on the target, and score the impact on the target. In some cases, the system is configured to prompt the shooter about the shooting phase. For example, the system can be configured for use during a CMP high-powered rifle competition, and the system can prompt the user that the current phase requires firing 20 shots from an off-hand position at the target within a 20-minute window. In some cases, the system knows how many shots are expected during the shooting phase (called a string of fire) and can prompt the user with information associated with the current shooting phase, such as the number of shots, timeframe, and shooting posture. In some cases, the shooter may input information associated with the shooting phase, such as the number of shots the system should expect, the firearm used, the distance to the target, etc. In some cases, the system is manually activated and deactivated, and shots on target are recognized only during the time the system has been activated.
[0278] Comprehensive training system
[0279] As described above, in some embodiments, the system utilizes a rack-type arrangement in which multiple recording devices can be used to capture audio and video of participants and / or targets. In some cases, the imaging device can be pointed toward the target and can have a zoom lens, a digital zoom feature, or rely on an external lens, such as a camera mounted to a scope, for capturing images of the target within the target area. In some cases, the system is configured to synchronize multiple sources, such as one or more video frames and / or audio data from one or more audio capture devices. In some cases, the system is configured to synchronize multiple video frames and audio data from one or more audio / video capture devices.
[0280] The described system thus provides a comprehensive firearms training solution that allows a shooter to track body movements even at a busy shooting range, identify errors in technique, associate scores with errors in technique, automatically score targets, and distinguish between shots fired by a participant's firearm and shots fired by other firearms.
[0281] This system can utilize one or more machine learning models for synchronization, prediction, verification, and can be further trained to analyze scoring data and scoring data is associated with motion data, to determine the correlation between specific motion data (for example, behavior) and the scoring trend. As an example, the system can be associated with the shooter's wrist rotation downwards and typically score outside 10 rings and below 10 rings before shooting, and determine that the shooter prejudges firearm recoil before shooting. The system can then provide feedback to the user, not only has the information relevant to motion / score correlation, but also can provide one or more exercises or training, to allow the user to identify and solve the behavior that causes score reduction. The system can be further trained to distinguish the launch of the firearm associated with the shooter of interest, even in the busy shooting range with many shooters. As described elsewhere herein, a similar process can be used together with any motion data from any activity or motion.
[0282] In some embodiments utilizing multiple video capture devices, the system can track body motion in two or three dimensions and from multiple angles. The two-dimensional or three-dimensional body motion data can be correlated, synchronized, and analyzed to determine two-dimensional or three-dimensional motion data, which can be further correlated with the resulting score.
[0283] While the embodiments of the described systems are described with respect to a shooter firing a burst of shots, it should be understood that the systems and methods described herein are applicable to capturing any type of body motion and to other sports where body motion can produce performance metrics. For example, embodiments of the systems described herein can be used to track, review, and improve body motion such as basketball free throw shooting, golf swings, figure skating elements, archery, soccer, baseball swings, or any other sport or activity in which the movement of a set of observable body markers can be recorded in time and there is some observed causal consequence of that movement.
[0284] The system may include one or more processors and one or more computer-readable media that may store various modules, applications, programs, or other data. The computer-readable media may include instructions that, when executed by one or more processors, cause the processors to perform the operations described herein for the system.
[0285] In some implementations, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), both a CPU and a GPU, a microprocessor, a digital signal processor, or other processing units or components known in the art. Alternatively or additionally, the functions described herein may be performed, at least in part, by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc. In addition, each of the processors may have its own local memory, which may also store program modules, program data, and / or one or more operating systems. One or more control systems, computer controllers, and remote controls may include one or more cores.
[0286] Embodiments may be provided as a computer program product comprising a non-transitory machine-readable storage medium having stored thereon instructions (in compressed or uncompressed form) that can be used to program a computer (or other electronic device) to perform the processes or methods described herein. Computer-readable media may include volatile and / or non-volatile memory, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Machine-readable storage media may include, but are not limited to, hard drives, floppy disks, optical disks, CD-ROMs, DVDs, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, flash memories, magnetic or optical cards, solid-state memory devices, or other types of media / machine-readable media suitable for storing electronic instructions. In addition, embodiments may also be provided as a computer program product comprising a transient machine-readable signal (in compressed or uncompressed form). Whether or not carrier modulation is used, examples of machine-readable signals include, but are not limited to, signals that a computer system or a machine hosting or running a computer program can be configured to access, including signals downloaded via the Internet or other networks.
[0287] Those skilled in the art will recognize that any process or method disclosed herein may be modified in a variety of ways. The process parameters and step sequences described and / or illustrated herein are provided as examples only and may be changed as needed. For example, although the steps illustrated and / or described herein may be shown or discussed in a particular order, the steps do not necessarily need to be performed in the order shown or discussed.
[0288] The various exemplary methods described and / or shown herein may also omit one or more steps described or shown herein, or include additional steps in addition to these disclosed steps. In addition, the steps of any method as disclosed herein may be combined with any one or more steps of any other method as disclosed herein.
[0289] The present disclosure sets forth exemplary embodiments and, therefore, is not intended to limit the scope of the embodiments of the present disclosure and the appended claims in any way. The embodiments have been described above with the help of functional building blocks that illustrate the implementation of specific functions and their relationships. For ease of description, the boundaries of these functional building blocks have been arbitrarily defined herein. Alternative boundaries may be defined to the extent that the specified functions and their relationships are appropriately performed.
[0290] The foregoing description of specific embodiments will fully reveal the general nature of the embodiments of the present disclosure so that others can, by applying the knowledge of one of ordinary skill in the art, readily modify and / or adapt various applications of such specific embodiments without departing from the general concepts of the embodiments of the present disclosure, without undue experimentation. Therefore, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments, based on the teachings and guidance presented herein. The wording or terminology herein is for the purpose of description and not for the purpose of limitation, so that the terms or wording of this specification will be interpreted by one of ordinary skill in the relevant art in light of the teachings and guidance presented herein.
[0291] The breadth and scope of embodiments of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
[0292] Conditional language, such as "can," "could," "might," or "may," among others, unless specifically stated otherwise or understood otherwise in the context as used, is generally intended to convey that certain embodiments may include certain features, elements, and / or operations, while other embodiments do not. Thus, such conditional language is generally not intended to imply that features, elements, and / or operations are in any way required for one or more embodiments, or that one or more embodiments must include logic for determining, with or without user input or prompting, whether such features, elements, and / or operations are included in or to be performed in any particular embodiment.
[0293] Unless otherwise indicated, the terms "connected to" and "coupled to" (and their derivatives) as used in the specification should be interpreted as allowing both direct and indirect connections (i.e., via other elements or components). Furthermore, the terms "a" or "an" as used in the specification should be interpreted as meaning "at least one of." Finally, for ease of use, the terms "including" and "having" (and their derivatives) as used in the specification do not exclude additional components and should be interpreted as open-ended.
[0294] The specification and drawings disclose examples of systems, devices, apparatus, and techniques that can provide systems and methods for determining the acoustic signature of a discharged firearm. Of course, it is not possible to describe every conceivable combination of elements and / or methods for the purpose of describing the multiple features of the present disclosure, but one of ordinary skill in the art recognizes that many other combinations and arrangements of the disclosed features are possible. Therefore, various modifications may be made to the present disclosure without departing from the scope or spirit of the present disclosure. In addition, other embodiments of the present disclosure may be apparent from consideration of the specification and drawings, as well as from the practice of the disclosed embodiments as presented herein. The examples set forth in the specification and drawings should be considered in all respects to be illustrative and not restrictive. Although specific terms are employed herein, these terms are used in a general and descriptive sense only and are not intended to be limiting.
[0295] Those skilled in the art will recognize that, in some implementations, the functionality provided by the above-described processes and systems can be provided in an alternative manner, such as by splitting or merging more software programs or routines into fewer programs or routines. Similarly, in some implementations, the processes and systems shown can provide more or less functionality than described, such as when other illustrated processes lack or include such functionality, or when the amount of functionality provided is changed. In addition, although various operations can be illustrated as being performed in a particular manner (e.g., serially or in parallel) and / or in a particular order, those skilled in the art will recognize that, in other implementations, operations can be performed in other orders and in other ways. Those skilled in the art will also recognize that the data structures discussed above can be constructed in different ways, such as by splitting a single data structure into multiple data structures or by merging multiple data structures into a single data structure. Similarly, in some implementations, the data structures shown can store more or less information than described, such as when other illustrated data structures lack or include such information, or when the amount or type of stored information is changed. Various methods and systems shown in the accompanying drawings and described herein represent example implementations. In other implementations, the methods and systems may be implemented in software, hardware, or a combination thereof. Similarly, in other implementations, the order of any method may be changed, and various elements may be added, reordered, combined, omitted, modified, etc.
[0296] In light of the foregoing, it should be understood that although specific embodiments have been described herein for illustrative purposes, various modifications may be made without departing from the spirit and scope of the appended claims and the elements described therein. In addition, although certain aspects are presented above in the form of certain claims, the inventors envision multiple aspects of any available claim form. For example, although only some aspects may be described as embodied in a specific configuration at present, other aspects may also be embodied in this manner. As will be apparent to those skilled in the art who benefit from this disclosure, various modifications and changes may be made. It is intended to include all such modifications and changes, and therefore, the above description should be considered illustrative rather than restrictive.
Claims
1. A method for improving shooting performance, comprising: receiving video data of the shooter; determining one or more physical landmarks of the shooter; tracking the one or more body landmarks during shooting to generate shooting motion data; determining a score for the shot; associating the shooting sports data with the score; and A recommendation is generated for altering the motion data for a subsequent shot.
2. The method according to claim 1, wherein Determining one or more body landmarks of the shooter includes generating a wireframe model by connecting the body landmarks.
3. The method according to claim 1, wherein Associating the shooting motion data with the score includes executing a classification and regression tree machine learning model to identify a causal relationship between the shooting motion data and the score.
4. The method of claim 1 , further comprising determining the shooter's grip through image analysis of the video data of the shooter.
5. The method of claim 4, further comprising analyzing the shooter's grip and providing grip recommendations on a display screen to change the grip.
6. The method according to claim 1, wherein Determining the score for the shot includes: receiving target video data; performing image analysis on the received target video data; Confirming hits on targets; and A score for the hit is determined.
7. The method according to claim 1, wherein Receiving the video data includes capturing the video data via a mobile phone.
8. The method according to claim 1, wherein Determining one or more body landmarks includes determining 17 body landmarks.
9. The method according to claim 1, wherein: Tracking the one or more body landmarks includes generating a bounding box around each of the one or more body landmarks.
10. The method of claim 1, further comprising executing a machine learning model to associate the shooting sports data with the score.
11. The method according to claim 10, wherein: The machine learning model is configured to determine motion data that resulted in off-center target hits.
12. A method for improving the causal consequences of physical movement, comprising: receiving video data of body movement; determining one or more body landmarks viewable in the video data of the body movement; tracking the one or more body landmarks during the action; generating athletic data based at least in part on tracking the one or more body landmarks; determining a score associated with the athletic data; associating the athletic data with the score; and Generate a recommendation for changing the motion data for a subsequent action.
13. The method according to claim 12, wherein: Receiving the video data includes capturing the video data via a mobile computing device.
14. The method of claim 12, further comprising executing a machine learning model to associate the athletic data with the score.
15. The method according to claim 14, wherein The machine learning model is configured to determine athletic data that resulted in a decreased score.
16. The method of claim 15, further comprising predicting, by the machine learning model, a prediction score based on the motion data. The method of claim 16 , further comprising comparing the predicted score to the score.