Sports training systems using computer vision and artificial intelligence for digitally supervised training and evaluation

The sports training mat with computer vision and mobile device technology provides a cost-effective, engaging platform for children to improve their skills through precise performance analysis and adaptive training, addressing the need for interactive and evaluative sports training systems.

WO2026097182A1PCT designated stage Publication Date: 2026-05-15U-PRO SOCCER TECHNOLOGIES INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
U-PRO SOCCER TECHNOLOGIES INC
Filing Date
2025-11-12
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing sports training technologies lack a robust, cost-effective system that provides comprehensive analytics, is fun and entertaining, and appeals to children, particularly for interactive and evaluative sports training starting from a young age without expensive hardware.

Method used

A sports training mat with machine-detectable markers and a camera-equipped mobile device using computer vision to define a three-dimensional training space, track body movements, and provide instructional guidance and evaluation, incorporating adaptive difficulty scaling and a game system for skill development.

Benefits of technology

Enables precise, real-time feedback and comprehensive analysis of athletic performance, promoting skill development and physical activity in children through an engaging and affordable system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025051506_15052026_PF_FP_ABST
    Figure CA2025051506_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A system for digitally supervised sports training features a training mat with a plurality of machine-detectable markers, a camera-equipped mobile device having a monoscopic camera aimable to encompass the training mat and an overlying space. Executable software embodied at least in part by the mobile device is operable to capture digital imagery of the mat and its markers using the camera, detect the markers within the captured imagery using computer vision, and then use at least a subset of the detected markers as reference points to digitally define virtual boundaries of a three-dimensional training space. Through a GUI of the mobile device, instructional guidance is given on a sports training exercise to be performed by the user, performance of which is digitally captured by the camera. The software performs automated tracking of body movements relative to the virtually bound training space, and uses same in a fully automated evaluation of the performance.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SPORTS TRAINING SYSTEMS USING COMPUTER VISION AND ARTIFICIAL INTELLIGENCE FOR DIGITALLY SUPERVISED TRAINING AND EVALUATION

[0002] CROSS- REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority benefit of United States Provisional Patent Application No. 63 / 718,772, filed November 11 , 2024, the entirety of which is incorporated herein by reference.

[0004] FIELD OF THE INVENTION

[0005] The present invention relates generally to the field of sports training, and more particularly to methods and systems for providing interactive virtual coaching of athletes, with particular emphasis on children to promote physical activity and other beneficial aspects of sports participation.

[0006] BACKGROUND

[0007] Integration of modem computer technology, and most recently artificial intelligence and machine learning, into the field of sports training, sports analysis and athletic performance evaluation has been on the rise, of which the following patent literature are some examples of prior efforts with some at least nominal relevance to the present invention.

[0008] Issued U.S. Patent US10010778 is by Pillar Vision Inc. discloses use of two depth sensing cameras are used to capture imagery of players performing training sequences, and to evaluate dribbling and passing performance based on various characteristics of the performance, among which posture (stance, body motion) is included as one example. The measured performance characteristics are evaluated against stored “proper” characteristics denoted by a set of evaluation data). The detailed embodiment focusses on basketball, but mention is made of applying the technology to soccer and hockey. Issued U.S. Patents US10600334 & US11642047 by Nex Team Inc. discloses a sports training and analysis tool with particular focus on body-eye coordination and reaction time. In this example, a smartphone is used to display the user guidance for the training exercises being performed, and to perform artificial intelligence (Al) based computer vision motion capture of the user’s performance for evaluation and scoring purposes.

[0009] Issued U.S. Patent 9358426 by Nike, Inc. discloses a system and method for guiding and monitoring physical exercises using machine vision to evaluate same against an “ideal” form of that exercise, which ideal form is represented onscreen by a virtual shadow over which an avatar of the user is overlaid for visual feedback on whether the user is following proper form for the given exercise. The performance is evaluated based on relative conformance / overlap between the user’s form and the shadowed ideal form, and can include location scores according to body part or grouped body parts. Sharing of scores via social media is described for ranking of multiple users.

[0010] Published U.S. Patent Application US2023 / 0145451 by the Sony Group Corporation teaches a system for monitoring and evaluation performance of strength training exercises in a gym environment, with time series 3D motion capture data of the user’s performance using computer vision, which data is then evaluated against a nominal movement pattern. Outputted performance indicators may be a quantified metric or identified deviation from the proper form.

[0011] Published U.S. Patent Application US2024 / 0261660 by Houlihan discloses hockey and soccer training systems using combination of a smartphone / tablet app for guiding performance of training exercises and a multi-colored game floor atop which those exercises are performed. Rather than motion capture of the player movements, use of an electronically equipped puck or ball is made to enable monitoring of the puck / ball movements. Published U.S. Patent Application US2023 / 0173366 by Cook discloses use of a floor mat in digitally aided performance of sports training exercises, with basketball dribbling as the illustrated example, but with soccer referenced as another example of a useful context for the solutions disclosed. Use of a smartphone application is contemplated for displaying video demonstration of the training exercise that the player is to perform, but the reference lacks implementation of motion capture or evaluation.

[0012] Published U.S. Patent Application US2007 / 0219024 by Allegre discloses a soccer training mat usable for performing various training exercises, and briefly mentions the notion of a computerized means for measuring and evaluating the player’s performance of the exercises, but with no enabling disclosure relating to such computerized monitoring and evaluation.

[0013] Published US Patent Application US2010 / 0137079 teaches a sports training setup by Skillz System Inc. A computer-connected camera is used to monitor the position of a ball or puck, for the purpose of reproducing a virtual representation of the ball or puck on a display screen, for the purpose of helping the user train to stick handle a ball or puck “by feel” without looking down. Though the detailed embodiment is in the context of hockey, use of a similar setup for soccer training is also mentioned. Practice drills are described using on-screen interactive content, particularly by display of obstacles / targets on screen to be navigated. There is reference to quantification and scoring of skills, and competitive scoring among multiple players, and optional networking among multiple players. An online demo video of the product shows inclusion of playable “video games” that use the ball / puck interaction as input to an alternative mode of gameplay, to offer added entertainment value to the intended child audience for the product.

[0014] Despite these prior innovations, there remains much room for improvement for a particularly robust system and methodology for such interactive and evaluative sports training technology, particularly one that can be implemented at reasonable cost without expensive hardware, with highly comprehensive analytics, an operating environment that is fun, entertaining and easy to navigate, and particularly appealing to children to encourage physical activity and skills development starting from a young age.

[0015] SUMMARY OF THE INVENTION

[0016] According to one aspect of the invention, there is provided system for digitally supervised sports training, said system comprising: a sports training mat comprising a plurality of machine-detectable markers on a topside of said sports training mat; a camera-equipped mobile device having a monoscopic camera whose field of view (FOV) is aimed or aimable in a manner encompassing said sports training mat and an overlying space thereabove; embodied at least partially by said camera-equipped mobile device, one or more processors and one or more non-transitory computer readable medium connected thereto and having stored therein statements and instructions executable by said one more processors to cause, when executed, performance of numerous steps comprising at least:

[0017] (a) with the topside of said sports training mat situated within a field of view (FOV) of a singular monoscopic camera of said camera-equipped mobile device, capture digital imagery of the topside of said sports training mat using said singular monoscopic camera;

[0018] (b) within said captured digital imagery of the topside of said sports training mat, detecting said machine-detectable markers using computervision;

[0019] (c) using at least a subset of the detected markers as reference points, digitally defining outer virtual boundaries of a three-dimensional training space situated overtop of the mat, the topside of which denotes a bottom extremity of said training space; (d) communicating, to a user, instructional guidance regarding performance of a sports training exercise to be performed by said user;

[0020] (e) digital capturing, by said singular monoscopic camera, of a userperformance of said sports training exercise; and

[0021] (f) one of either:

[0022] (i) automated tracking of body movements in captured digital imagery of said user-performance; or

[0023] (ii) outputting of said captured digital imagery to another computing resource configured to perform said automated tracking of the body movements in said captured digital imagery of said userperformance.

[0024] Preferably said machine-detectable markers comprise outer markers situated proximate an outer perimeter of the sports training mat, and inner markers residing further inward from the outer perimeter of the sports training mat, and step (c) comprises using said outer markers to define said outer virtual boundaries.

[0025] Preferably said outer markers comprise corner markers residing respectively proximate to outer perimeter comers of the sports training mat.

[0026] Preferably step (c) further comprises, using said inner markers as further references, subdividing the training space into a plurality of smaller zones.

[0027] Preferably step (e) further comprises tracking said body movements, at least in part, by detected presence or movement within, among or between different zones of the training space.

[0028] In some embodiments, the smaller zones may be geofenced from one another, and geofenced virtual boundaries between said smaller zones are used in the tracking of said body movements. Preferably the sports training mat comprises a plurality of human discernible target indica on the topside thereof, and said instructional guidance comprises identification of one or more of said target indicia and identification of a user action to be taken relative thereto.

[0029] Preferably at least a subset of said human discernible target indicia and at least a subset of the machine-readable markers of are of coincident location to on another on the topside of the mat.

[0030] The human discernible target indica may be defined by, or be integral parts of, the machine-readable markers.

[0031] Preferably the human discernible target indicia include comprise numerals, alphabetic characters, symbols, or any combination thereof.

[0032] Preferably the instructional guidance comprises visual demonstration of the sports training exercise with a prescribed pattern or form of body movement. Such visual demonstration may be presented as live-action video or animated video, and may optionally be accompanied by textual and / or auditory content. In other implementations, the instructional guidance may be as textual or auditory instruction, without visual demonstration.

[0033] Preferably a virtual reference model of the prescribed pattern or form of body movement is digitally stored, and the steps further comprises performing a computer automated evaluation of the recorded user-performance against said virtual reference model. The virtual reference model is preferably derived from a recorded performance of the training exercise by a professional or semi-professional athlete, whose performance denotes an optimal or ideal performance. The same recorded performance may be used as the visual demonstration. Preferably the steps further comprise logging the tracked body movements from the recorded user-performance as historical data set, and in subsequent repetition of the same sports training exercise, or variations thereof, at a later time, performing a computer automated evaluation thereof against the historical data to gauge physical development and / or skill level progression over time.

[0034] Preferably the steps further comprise outputting, from said computer automated evaluation, a plurality of evaluation scores, among which at least a subset are bodypart or body-region scores attributed to body movement evaluation of different parts or regions of the user’s body.

[0035] The steps may also comprise outputting, in aggregate, a plurality of user-specific evaluation scores respectively assigned to a plurality of different users, based on equivalent evaluation of respective performances of the sports training activity by said plurality of different users.

[0036] Preferably the steps further comprise display of said body-part or body-region scores in combined display with a body avatar or body image relative to which said body-part or body-region scores are displayed in respective proximity to corresponding parts or regions of the body avatar or body image.

[0037] Embodiments of the invention also include one or more non-transitory computer readable media having stored therein statements and instructions executable by one more processors to cause, when executed, performance of the any the steps recited in the preceding paragraphs.

[0038] Embodiments of the invention also include a computer implemented method executed by a computing system comprising one or more processors and one or more non- transitory computer readable media having stored therein statements and instructions executable by one more processors, of which at least one of said processors and at least one of said one or more non-transitory computer readable media are embodied in a camera-equipped mobile computing having a monoscopic camera whose field of view (FOV) is aimed or aimable in a manner encompassing a sports training mat comprising a plurality of machine-detectable markers on a topside of said sports training mat, said method comprising the steps recited in any of the preceding paragraphs.

[0039] Embodiments of the invention also include a sports training mat comprising a plurality of machine-detectable markers each occupying a respective discrete position on a topside of said sports training mat for detection thereof by a computer-vision algorithm for processing imagery, captured by a monoscopic camera of a camera-equipped mobile device, of user-performance of sports training exercises atop said mat, said machine detectable markers comprising at least a set of outer markers each residing proximate an outer perimeter of the mat for use as reference points to setup outer virtual boundaries of a training space overtop said mat, and a set of inner marks situated further inwardly from said outer perimeter of the mat for use as further reference points to setup internal virtual boundaries by which said training space is internally subdivided into a plurality of smaller zones.

[0040] In use, the mat is put into practice in combination with said camera-equipped mobile device.

[0041] Preferably, said camera-equipped mobile device comprises at least a subset of the one or more processors and non-transitory computer readable medium having stored therein the executable statements and instructions recited earlier in the preceding paragraphs.

[0042] Typically, said camera-equipped mobile device is a smartphone or tablet computer. The steps executed in the preceding method, and the steps executed by the one or more processors of the preceding system, may also comprise automated collection and analysis of biomechanical metrics, including at least one, in some instances a partial plurality, in any combination, and in other instances all of: bilateral movement data including arm and leg symmetry; postural metrics including shoulder balance, elbow swing, and hip rotation; stability measurements including shoulder and hip stability; sport-specific metrics including ball touches and head position; and movement cadence for left and right sides.

[0043] The steps executed in the preceding method, and the steps executed by the one or more processors of the preceding system, may also comprise automated analysis of kinesiology characteristics of each movement including at least one, in some instances a partial plurality, in any combination, and in other instances all of: movement group categorization; muscle group identification; movement speed classification; ability type definition; joint impact assessment; and energy expenditure calculation in METs.

[0044] The steps executed in the preceding method, and the steps executed by the one or more processors of the preceding system, may also comprise creation and storage of a player profile including at least one, in some instances a partial plurality, in any combination, and in other instances all of: age; gender; soccer experience level; dominant foot; club; and historical performance data.

[0045] The steps executed in the preceding method, and the steps executed by the one or more processors of the preceding system, may also comprise provision of adaptive difficulty scaling based on at least one, in some instances a partial plurality, in any combination, and in other instances all of: real-time biomechanical performance; kinesiology metrics; player profile data;

[0046] The steps executed in the preceding method, and the steps executed by the one or more processors of the preceding system, may also comprise implementation of a game system comprising at least one, in some instances a partial plurality, in any combination, and in other instances all of: skill-based scoring mechanisms derived from biomechanical metrics; progressive difficulty adjustment; achievement unlocks based on measured improvement; digital rewards tied to biomechanical performance; customizable avatars reflecting skill / level-up progression; multiplayer modes using synchronous and asynchronous options; and local, regional, national and / or worldwide digital tournaments with skill-based matchmaking.

[0047] The adaptive difficulty scaling may comprise: a. analyzing biomechanical metrics to determine current skill level; b. comparing performance against age and experience-appropriate benchmarks; and c. automatically adjusting at least one, in some instances a partial plurality, in any combination, and in other instances all of: i. training zone parameters; ii. success thresholds; iii. game mechanics difficulty; iv. reward systems; and d. maintaining optimal challenge levels through continuous assessment.

[0048] The benchmarks may also be influenced on the basis of gender.

[0049] The game system may comprise at least one, in some instances a partial plurality, in any combination, and in other instances all of: a. a serious game framework where: i. game progression is tied to measured skill improvement; ii. digital rewards require validated physical achievements; iii. social features use standardized performance metrics; and iv. competition mechanics reflect real skill differentials; b. multiple game modes including: i. single-player skill development; ii. multiplayer challenges; iii. asynchronous tournaments; and iv. skill-based battle modes; and c. achievement tracking comprising: i. biomechanical milestone tracking; ii. kinesiology-based progress measures; iii. skill-appropriate challenge completion; and iv. social comparison metrics.

[0050] The achievement tracking may also comprise habit forming metrics.

[0051] BRIEF DESCRIPTION OF THE DRAWINGS

[0052] For aiding the readers understanding of the forgoing summary of the invention and the detailed description of preferred embodiments that follows, appended hereto are a set of drawings, in which: Figure 1 is a block diagram of an inventive system for digitally supervised sports training using a training mat, a camera-equipped mobile device, and Al computer vision software for evaluating athletic movements of a sports training exercise in a geofenced training space mapped to a coordinate system using computer readable markers embodied on the mat;

[0053] Figure 2 is a top plan view of a preferred embodiment of the training mat of the inventive system of Figure 1 ;

[0054] Figure 3 is an annotated top plan view of the training mat schematically illustrating subdivision of the training space into smaller triangular zones whose boundaries, in the plane of the mat, run from marker to marker,

[0055] Figure 4 is a schematic illustration of a training session performed on the mat;

[0056] Figure 5 schematically illustrates a sequence of display screens displayed in the graphical user interface (GUI) of the software application of the mobile device during such training session;

[0057] Figure 6 schematically illustrates a results screen displayed in the GUI after completion of the training session, showing scored results of the computer executed evaluation of the user’s performance;

[0058] Figure 7 is a schematic flowchart illustration of a pose analysis process performed by the Al computer vision software of the system of Figure 1 on captured video frames of the training session;

[0059] Figure 8 is a schematic flowchart illustration of the computer executed evaluation of the user’s recorded user-performance against a virtual reference model composed of an equivalently captured and analysed demonstration of the training exercise by a professional athlete; and

[0060] Figure 9 shows timestamped frames of a stick model animation representative of either the recorded user performance of the training exercise, or the professional athlete’s demonstration of the training exercise, which animation is producible by, and demonstrates the meaning and usefulness of, the output data of the pose analysis process of Figure 7.

[0061] DETAILED DESCRIPTION OF PREFERRED EMBODIMENT(S)

[0062] In enabling disclosure of the inventive system and methodology summarized briefly above, with particular emphasis on the novel combination of the training mat and the computer vision motion tracking and evaluative analysis performed thereon, detailed description of at least one preferred embodiment of the invention is now given as follows, with particular reference to the figures briefly described above.

[0063] Figure 1 schematically shows the cooperatively related componentry of the inventive system 10. One such system component is an individual user’s training mat 12 on which such individual user 14 performs guided sports training exercises, typically in accompaniment by a sporting object 16 that is physically maneuvered by body movements of the user 14 during performance of any given one of the guided sports training exercises, and which user-manipulated sporting object 16 in the illustrated example is a soccer ball, but may vary in other implementation so the invention characterized by other types of sport. Another system component is a camera- equipped mobile device 18, typically embodied as a smartphone or tablet computer having at least one monoscopic camera, whose field of view is appropriately aimed, in use of the invention, to capture in its field of view (FOV) the laid out training mat 12 and a three-dimensional space residing directly thereover in which the user will ultimately perform the sports training exercise, which three dimensional space is therefore referred to herein as the training space. The mobile device 18, in the illustrated example, is one of a plurality of computing resources among which the Al computer vision software and data storage components necessary to perform any and all of the functionalities of the invention described and claimed herein are distributed, and which in the illustrated example include one or more cloud servers 20. Each computing resource comprises one or more computer processors, and one or more non-transitory computer readable media for storage of both data and executable statements and instructions of the software, for execution of the latter by said one or more computer processors.

[0064] Description is first made of the basic configuration of the training mat 12. The training mat 12 of the illustrated embodiment, illustrated in isolation in Figures 2 and 3, is intentionally simple in design, being a passive component having printed markers 22A — 22C thereon but lacking any electrically powered subcomponents. The key outer corner markers 22A situated nearest the outer perimeter 12A of the training mat 12 (shown with numerical indicia 1 , 3, 6, and 8 in the illustrated example) are used by the Al computer vision software to create a quadrilateral outer geofence, mapping the physical space overtop the mat into a digital coordinate system and defining the primary spatial boundaries 24A of an overall primary training space 24 above the training mat 12 for performance of the training session within that primary training space. Additional outer markers 22B (with numerical indicia 2, 4, 5, 7 in the illustrated example) and inner markers 22C (with alphabetic indicia A, B, C, D in the illustrated example) are strategically placed to enable precise triangulation within the space. They create smaller subdivided zones 26 of the primary training space to enable accurate position tracking within a defined space.

[0065] Before evaluated performance of a guided sports training exercise can take place, the Al computer vision software must first perform an initial geofence establishment, the first stage of which is to detect / setup a primary boundary. Here, the system, via Al computer vision analysis of captured digital imagery of the training mat 12, first identifies the four outer corner markers 22A, which act as the primary reference points. Reading these markers 22A, the Al computer vision algorithm creates a digital quadrilateral boundary 24A - essentially a virtual fence around the training area. This establishes the primary training space 24 and allows the system 10 to map the physical space into a digital coordinate system that the Al computer vision software can use for tracking and analysis of player movements during the training exercises. The next stage of the geofence establishment process is the creation of smaller triangulation zones. Here, the system, again via its Al computer vision software, creates smaller triangular zones using various marker combinations once the primary boundary has been established. These triangles (like 1 -A-2 or 2-3-B, named here according to the readable indicia of the three markers denoting the three vertices of each triangular zone) serve as precise spatial reference points. By breaking the larger space into smaller triangular zones with geofenced virtual zone boundaries 26A, the system can track position and movement, and evaluate user performance, more accurately than possible with just the outer boundary markers alone.

[0066] As elaborated upon below in greater detail, the Al computer vision software of the system uses these established triangular zones 26 to track player body position with enhanced precision. These zones enable accurate monitoring of movement patterns, ball control, and drill completion. The triangular nature of the zones provides multiple reference points for any given position, improving tracking accuracy. The triangulated zones also allow the system to perform a detailed analysis of player performance. This may include precise movement tracking, measurement of speed and agility, assessment of spatial awareness, and evaluation of technical skills. The multiple reference points provided by the triangulation system enable more accurate measurements than would be possible with simpler boundary tracking.

[0067] Training analysis is enhanced by increased accuracy through triangulation. The use of multiple reference points through triangulation can significantly improve the system's tracking capabilities. This enhanced accuracy enables better movement analysis, precise performance metrics, and reliable skill assessment. The triangulation approach compensates for the limitations of using a single mobile device camera (in contrast to prior art approaching using multiple cameras for depth perception), and also compensates for the limited processing capacity of old smartphones, and environmental challenges like lighting space constraints that can vary widely in users’ homes and multiple people being in front of the camera at the same time as the player. Using zone-based analysis, the system can actively monitor how the player moves between and within zones, tracking time spent in each area and analyzing movement patterns across zones. This spatial analysis can provide insights into a player's use of space, movement efficiency, and skill execution within specific areas.

[0068] The zone-based approach also enables better integration with game mechanics, introducing the ability to implement zone-based challenges, and introducing novel means of performance scoring. Regarding zone-based challenges, the defined zones serve as the foundation for creating specific training tasks and objectives. The system uses these zones to establish clear success criteria for different drills and to track achievement completion. This spatial structure enables the creation of varied and progressive training challenges. Regarding performance scoring, the triangulated zones enable accurate performance measurement and assessment. This precision allows for reliable progress tracking and fair competition metrics, as all players are measured against the same spatial references. The system can consistently evaluate performance across different sessions and between different players.

[0069] Having described the training mat 12 and the setup of the geofenced training space 24 its subdivision into smaller geofenced zones 26, Figure 4 schematically illustrates the performance of a guided sports training exercise by the user 14 and their sporting object 16 in the established training space 24 atop the training mat 12. The camera- equipped mobile device 18, continuing to have its camera FOV aimed at the now digitally-established and subdivided training space 24, is further employed here by the software to present the user 14 with instructional guidance on the sports training exercise to be performed, which guidance may typically be embodied, at least in part, by a video demonstration 28 of that sports training exercise, for example a previously recorded video of a professional athlete demonstrating that exercise on another training mat 12 with another mobile device 18.

[0070] Alternatively, video demonstration of the exercise may be an animated video demonstration, in which the bodily movements of the animated demonstrator may be animated using a virtual reference model derived from the body movement data captured from such prior recording of the professional athlete demonstrating that exercise. In the case of an animated video demonstration, the animated demonstrator may be an animated video of a user avatar that has been selected, created or customized by the user 14 in the software GUI. Whether animated or live action, the demonstration video may be accompanied by textual and / or auditory content as part of the instructional guidance, or the instructional guidance may be implemented in videoless fashion, instead being composed solely of such textual and / or auditory guidance.

[0071] The instructional guidance is conveyed to the user 14 through the graphical user interface (GUI) on the mobile device 18, as schematically illustrated in the top panel of the GUI display sequence of Figure 5, after which the training session is initiated, for example in response to user initiation thereof by user selection of a “start” command 30 in the GUI, also shown in the top panel of Figure 5. Video recording of the user’s performance of the training exercise in the training space is performed by the mobile device 18, and may be accompanied by concurrent playback of the video demonstration or other instructional guidance, as shown in the second panel of Figure 5 where the demonstration video 28 is visible during recording of the user performance, for example as a small subcomponent of the overall screen display during that recording process, a majority of which may be used for real-time display of the user video being recorded. The display screen during the recording stage may also include a running countdown timer 32 of an allotted time in which to complete the training exercise. As shown in the third and final panel of Figure 5, the GUI may feature a toggle 34 for user-selected hiding or viewing of the concurrent playback of the video demonstration 28 during the recording.

[0072] After completion of both the recording of the training session and the computer- automated evaluation of the performance by the Al computer vision software, as described in more detail below, user performance scores or other evaluative results from such evaluation are then presented to the user via the GUI of the mobile device 18, as denoted in the final step of the Figure 4 timeline and visually demonstrated in the GUI results screen display of Figure 6. The illustrated example of Figure 6 includes a plurality of evaluation scores, among which a subset are body-part or bodyregion scores 36 attributed to body movement evaluation of different parts or regions of the user’s body, displayed in the illustrated example in the context of a body avatar 38 or body image relative to which those body-part or body-region scores are displayed in respective proximity to corresponding parts or regions of the body avatar or body image. The illustrated example also provides an overall evaluation score 40 as an overarching metric of the user’s performance over the full spectrum of analyzed body movements. The illustrated example of Figure 6 is a single-user results screen showing the given user only the user’s own personal evaluation results, though the software may further implement a multi-user results screen that reports, in aggregate, a plurality of user-specific evaluation scores respectively assigned to a plurality of different users, based on equivalent evaluation of their respective performances of the sports training exercise.

[0073] Attention is now turned to Figure 7, which illustrates the workflow of the Al computer vision software when it comes to a pose-analysis process 100 performed on the video recording 42 of a training exercise (or drill), which process 100 may be substantially, if not entirely, the same whether it is being performed on video capture of a professional’s demonstrative performance of the training exercise as guidance for other users’ performance of that exercise, or video capture of the guided performance the training exercise by one of those other users (training users). First, at pose detection step 102, individual frames of the video recording 42 are analyzed with Al computer vision to locate human anatomical (skeletal) key points (bodily landmarks) of the videoed subject (professional or training user), of which earlier prototypes of the invention accounted for 32 such key points, a quantity that was subsequently increased 33 key points in more recently updated implementations. This step 102 derives an estimated pose of the videoed subject within that frame based on the coordinate points of the located key points. The present invention is not limited to any particular quantity of key points, which quantity is therefore denoted in the figure by a generic variable X. From the estimated pose, joint angle calculations are performed at step 104A for at least some of the located key points to derive joint angle values for those key points. This leverages the fact that the angle between two vectors originating from one key point (the subject joint) to two other key points (that are connected through that joint) denotes a joint angle of that one key point from which the vectors originate (e.g. elbow angle is derivable as the angular difference between a vector drawn from the elbow to the wrist and a vector drawn from the elbow to the shoulder). The present invention is not limited to any particular quantity of assessed joint angles, which is therefore denoted in the figure by a generic variable Y, though as a non-limiting example, the Applicant’s latest implementation of the software accounts for 24 joint angles (X = 33, Y = 24, where Y < X because not all of the anatomical key points used are joints).

[0074] At step 104B, additional useful information extracted from the pose detection step 102 is calculation of which of the located key points resides in which of the geofenced zones 26 of the training space 24, as is calculable using the coordinate locations of the located key points and the known virtual boundaries 26A of the geofenced zones 26. The summed results of steps 104A & 104B is that, for each analyzed frame of the recorded video 42, the software stores in memory a respective frame-specific datapoint of a time series of video frames, which datapoint contains: identification of located key points of that frame, identification of which geofenced zone 26 each located key point was found, and the angular values of any joint angles calculable (and thus calculated) from the located key points of that frame. So for example, in a video frame in which the subject’s left knee, right kneed, left hip, right hip, left shoulder, and right shoulder were found, of each these key points (in this case, each also being a joint for which a joint angle is calculable), each key point would be characterized in that frame’s datapoint by an identifier of the geofenced zone in which that key point was found, and a calculated value of the joint angle of that key point. Key points for which a joint angle is not calculable is instead denoted only by identification of the geofenced zone in which that key point was located. At step 106, the frame-specific datapoints are concatenated into a time-series sequence of datapoints denoting the changing body pose of the video subject over time with respective to the mat (relative to which the digital coordinate system used by the computer vision analysis was mapped in the geofence establishment procedure undertaken before the performance and video capture of the training exercise, as explained above). This dataset of concatenated frame-specific datapoints ultimately outputted from the pose-analysis process 100 of Figure 7 can thus be used as input to a comparative analysis of the body movements found in two different video recordings of the same training exercise to evaluate the performance of one performer of the training exercise against another.

[0075] Figure 8 schematically illustrates such comparative analysis, executed here for the purpose of evaluating the performance of a given training exercise by a training user 14 against a demonstrative performance of the training exercise by a professional athlete 114, whose derived dataset from the pose-analysis process 100 thus serves as a virtual reference model denoting an optimal or ideal performance of the training exercise against which to evaluate the performance by the training user 14. The professional user’s dataset may typically be stored at the server level 20 of the computing platform, given that it may serve as a shared standard against which the performances of many users can be evaluated. In instances where the pose-analysis process 100 of Figure 7 is performed locally on a user’s mobile device 18, the resultant dataset may be transferred to the server 20 via a WebSocket or an HTTP endpoint (e.g. depending on internet speed, and / or on particular mode of operation in the software is being used). That said, the present invention is not necessarily limited to server-side implementation of the comparative evaluation process 200, which may alternatively be executed locally on the user’s mobile device 18, for example based on acquisition of the professional’s pose-analysis dataset from the server 20 (e.g. in conjunction with retrieval / streaming of the demonstration video 28 therefrom).

[0076] In the comparative evaluation process 200 of Figure 8, the respective pose analysis datasets of the training user 14 and the professional athlete 114 are both used as input to a sequence matching algorithm 202 that exploits a time-series enabled machine learning system, for example using a chain of Long Short-Term Memory (LSTM) layers in a neural network, to determined a time warp needed to achieve optimal sequence matching between the datapoint sequences of the two datasets, based on which the degree of agreement / disagreement in joint angle and key point zone-occupation between the professional’s datapoints and the closest matching datapoints of the training user can be outputted as a constant sized vector that captures the differences between the training user’s and professional’s execution of the training exercise in relation to joint angulation and zone-occupation. This vector output can also be accompanied by outputting of a time warp score reflective of how well or poorly the training user’s body movement timing synched with the timing of the professional’s demonstrative performance of the exercise. The joint angle and zonal evaluation vector of the comparative evaluation process 200 is denoted in the figures as output 204A, and the time warp score as output 204B.

[0077] From the outputted vector 204A and time warp score 204B, various skill-specific scores can be calculated as metrics of different skill categories of the training user, such as technical execution, speed and agility, balance and power, and effort and endurance, from which an overall score is also derivable for output to the training user. Such scores may be normalized based on the training user’s demographics (e.g. age, gender, and skill level, or any subset thereof). As one non-limiting example of skill level categorization, the following skill level hierarchy may be employed: Recreational - Players in the recreational category are typically beginners or those who prefer a less-intense, non-competitive environment. Skill levels can vary widely, but the overall focus is on developing basic fundamentals, having fun, and learning the rules of the game. Players are not expected to have advanced technical skills; instead, the emphasis is on participation, learning through play, and building a foundation of coordination and ball control.

[0078] Developmental - At the developmental level, players possess a more intermediate skill set and demonstrate a greater interest in improving their technical abilities. Having moved beyond the basics, these players are focused on more structured learning, skill-focused drills, and grasping tactical concepts, such as positioning. While still focused on growth, players at this level are generally more athletic and have some previous league experience, making them moderately competitive.

[0079] Premier - Premier-level players are the most experienced and highly skilled in grassroots soccer. They possess advanced technical skills, exceptional speed, and a strong understanding of game strategy. These players are very serious about their development, demonstrate a high level of physical and mental ability, and are capable of making split-second decisions under pressure. They have mastered fundamental skills and are consistently executing passes, shots, and other plays with accuracy and precision.

[0080] The joint angles from the training user’s dataset, in raw form before standardization thereof, and for example in locally calculated fashion in real time on the mobile device 18, may be used to calculate biomechanic scores. These will vary by sport and may be application-dependent, but examples of which in the soccer context of the illustrated embodiment may include one or more of: head up, shoulder balance, elbow swing, hip rotation, knee up, and ball control.

[0081] Figure 9 schematically illustrates the usefulness of the described frame-specific datapoints extracted from the captured video frames as an effective representation of the user / professional pose in such image frames, not only for the described comparative evaluation of training users 14 against one or more professionals 114 (or even against one or more other training users of higher skill level), but also for other possible uses, such as reconstructed animation of a video captured performance of a training exercise, whether as a “stick animation” of the type shown in this figure, or to enable animation of a user avatar in a manner following the body movements digitally embodied in the captured datasets. In the illustrated example of Figure 9, the first three animation frames are representative of the datapoints of an early three frames near the start of a captured 20-second video of a 20-second training exercise, and the last three animation frames are representative of the datapoints of a late three frames of that captured video (in this case recorded at 30 frames per second (fps), thus ending with frame 600 at the expiration of the 20 second training exercise).

[0082] The animation frames schematically illustrate the anatomical / skeletal key points 50 (bodily landmarks) found in each corresponding frame of the captured and analyzed video, as well as illustration of the vectors 52 constructable between such key points, from which joint angles are derivable, the collective totality of which thus forms a skeletal approximation (stick figure) of the videoed subject. These illustrated examples also demonstrate the optional inclusion of an additional non-anatomical key point 50A, which denotes computer vision identification of the sporting object 16 (soccer ball, in this example). The pose detection process 100 may therefore also include locating of the sporting object key point 50A at pose detection step 102, and identification in step 104B of which of the geofenced zones 26 the sporting object key point 50A resides in.

[0083] Furthermore, the joint angle calculation step 104A may be a broader angle calculation step that not only calculates joint angles of the videoed subject’s pose, but also calculates angular measure of the sporting object’s location relative to one or more bodily references. Just as the angular relationship between two vectors drawn interconnectivity between three anatomical key points to represent body segments connected by a shared joint can be used to denote a joint angle of that joint, a vector drawn from an anatomical key point 50 to the sporting object key point 50A can be measured angularly of another vector drawn from that anatomical key point to another anatomical key point, which calculated angularity of those two vectors thus represents a positional or geometric relationship between the ball location and the subject’s body. The zone occupation of the sporting object 16 and one or more calculated angulations of the sporting object’s location relative to the subject’s body can thus be populated into the datapoint for each analyzed video frame in which the sporting object 16 is found, which sporting object positional data forms additional input to the comparative evaluation process 200 of Figure 8 to achieve more comprehensive evaluation that considers not only body movement, but also the athlete’s attempted reproduction of the intended ball / object movements of the training exercise.

[0084] Applicant’s technology is believed to outperform the competition, at least in part by its ability to provide personalized, near real-time feedback on a player's performance using just a smartphone camera and the proprietary training mat. Specifically, Applicant’s Al-powered computer vision technology can analyze a player's movements, ball control, and technique with high accuracy, all without the need for expensive motion capture equipment or wearable sensors. This is a significant leap beyond what Applicant’s known competitors offer.

[0085] Perceived advantages of at least some embodiments of Applicant’s innovative technology include:

[0086] 1. Precision: Applicant’s system can detect subtle movements and provide feedback on things like foot placement, ball contact, and body positioning with accuracy comparable to professional coaching.

[0087] 2. Accessibility: Unlike systems that require multiple cameras or special equipment, Applicant’s technology works with a standard smartphone, making it accessible to almost any family. 3. Real-time Processing: Applicant’s Al processes the near real-time video feed, offering immediate feedback rather than delayed analysis.

[0088] 4. Adaptive Learning: Applicant’s system continuously learns from the data it collects, allowing it to improve its analysis and recommendations over time.

[0089] 5. Comprehensive Analysis: Applicant’s system does not merely track basic metrics like speed or number of touches, but instead provides an in-depth analysis of technique, decision-making, and overall skill progression.

[0090] This technological edge allows Applicant to offer a level of personalized, professionalgrade training that was previously only available through in-person coaching but at a fraction of the cost and with the convenience of at-home use. Rather than merely tracking performance, actionable insights are outputted that can truly help young players improve their skills.

[0091] For computer vision implementation, at least some embodiments use deep learningbased computer vision techniques, specifically convolutional neural networks (CNNs) and pose estimation models. For temporal analysis, at least some embodiments employ recurrent neural networks (RNNs), particularly LSTMs, to analyze player movements over time. For multi-task learning, the model in at least some embodiments simultaneously performs player detection, pose estimation and skill assessment.

[0092] Applicant’s present working embodiment was designed for the sport of soccer, and used in combination with a purely conventional soccer ball with no integrated electronics for tracking ball movement, denoting embodiments in which all object tracking is executed by computer vision on the digital imagery captured by the digital camera of the mobile device (e.g. smartphone or tablet), which in some implementations may involve communication of this captured digital video imagery to a remote (e.g. cloud) server at which the computer vision algorithm processes the video data in real time, and communicates result data from the automated performance evaluation back to the mobile device for, typically visual, communication to the user, though as contemplated above, other embodiments may instead analyze the video locally on the mobile device.

[0093] In distinction over the aforementioned prior art, Applicant’s preferred embodiments uniquely use physical markers for precise, spatial mapping with dynamic geofencing for training zones, the use of a single mobile camera for tracking, and preferably local real-time processing of biomechanical data analysis.

[0094] Some embodiments may further implement a comprehensive gaming ecosystem anchored on the concept of serious games by merging digital and physical progress into a seamless experience. The game environment preferably includes storytelling and single-player mode, where players will receive positive feedback from the user interface of the mobile device software application to foster feelings of attention, approval, and praise for the child; multiplayer modes (both synchronous and asynchronous), where kids will be able to connect with friends and teammates to train, compete, and collaborate, building a community driven by physical challenges and digital tournaments. An achievement system based on level progression, which will unlock new challenges, badges, and digital collectibles to reflect skills development; a customizable avatar that reflects the player's personality and the feeling of appreciation; season passes, which will bring extra digital collectibles tied to training achievements (level-up); and a marketplace where players would be able to buy digital collectibles to further enhance their experience.

[0095] Some embodiments may combine training zones created by virtual geofencing, game mechanics that respond to real-world skill execution, reward systems tied to a biomechanical scoring system, and social features integrated with skill development. Collectively, there may be achieved a unique serious gaming platform where gaming elements directly reinforce proper technique, training goals are achieved through engaging gameplay, social features encourage consistent practice, and progress is measured both in-game achievements and real-world skills self-development.

[0096] Preferred embodiments of the training mat lack any sensors or other integrated electronic componentry in the mat itself, and instead have only the “passive” markings by which the computer-vision algorithm can identify the mat and use the markings to calibrate the system according to the occupied placement of the mat in the camera field of view of the mobile device in order to be able to setup the geofence bounded zones of the subdivided training space and track the user movements in the zone- subdivided 3D space from the digitally captured 2D imagery of the mat, user and ball, though addition of active subcomponentry to a mat that exploits its passive markings for equivalent purpose is also within the scope of the present invention.

[0097] Since various modifications can be made in the invention as herein above described, and many apparently widely different embodiments of same made, it is intended that all matter contained in the accompanying specification shall be interpreted as illustrative only and not in a limiting sense.

Claims

CLAIMS:

1. A system for digitally supervised sports training, said system comprising: a sports training mat comprising a plurality of machine-detectable markers on a topside of said sports training mat; a camera-equipped mobile device having a monoscopic camera whose field of view (FOV) is aimed or aimable in a manner encompassing said sports training mat and an overlying space thereabove; embodied at least partially by said camera-equipped mobile device, one or more processors and one or more non-transitory computer readable medium connected thereto and having stored therein statements and instructions executable by said one more processors to cause, when executed, performance of numerous steps comprising at least:(a) with the topside of said sports training mat situated within a field of view (FOV) of a singular monoscopic camera of said camera-equipped mobile device, capture digital imagery of the topside of said sports training mat using said singular monoscopic camera;(b) within said captured digital imagery of the topside of said sports training mat, detecting said machine-detectable markers using computervision;(c) using at least a subset of the detected markers as reference points, digitally defining outer virtual boundaries of a three-dimensional training space situated overtop of the mat, the topside of which denotes a bottom extremity of said training space;(d) communicating, to a user, instructional guidance regarding performance of a sports training exercise to be performed by said user;(e) digital capturing, by said singular monoscopic camera, of a userperformance of said sports training exercise; and(f) one of either:(i) automated tracking of body movements in captured digital imagery of said user-performance; or(ii) outputting of said captured digital imagery to another computing resource configured to perform said automated tracking of the body movements in said captured digital imagery of said user-performance.

2. The system of claim 1 wherein step (c) comprises subdividing the training space into a plurality of different zones.

3. The system of claim 2 wherein step (c) comprises defining the outer virtual boundaries of the training space using a first subset of the machine-detectable markers, and using a second subset thereof to subdivide the training space into the plurality of zones.

4. The system of 3 wherein said machine-detectable markers comprise outer markers situated proximate an outer perimeter of the sports training mat, and inner markers residing further inward from the outer perimeter of the sports training mat, and the first and second subset of markers respectively comprise the outer and inner markers.

5. A system for digitally supervised sports training, said system comprising: a sports training mat comprising a plurality of machine-detectable markers on a topside of said sports training mat; a camera-equipped mobile device having a monoscopic camera whose field of view (FOV) is aimed or aimable in a manner encompassing said sports training mat and an overlying space thereabove; embodied at least partially by said camera-equipped mobile device, one or more processors and one or more non-transitory computer readable medium connected thereto and having stored therein statements and instructions executable by said one more processors to cause, when executed, performance of numerous steps comprising at least:(a) with the topside of said sports training mat situated within a field of view (FOV) of a singular monoscopic camera of said camera-equipped mobile device, capture digital imagery of the topside of said sports training mat using said singular monoscopic camera;(b) within said captured digital imagery of the topside of said sports training mat, detecting said machine-detectable markers using computervision;(c) using at least a subset of the detected markers as reference points, digitally defining a plurality of zones collectively occupying a three-dimensional training space situated overtop of the mat, the topside of which denotes a bottom extremity of said training space;(d) communicating, to a user, instructional guidance regarding performance of a sports training exercise to be performed by said user;(e) digital capturing, by said singular monoscopic camera, of a userperformance of said sports training exercise; and(f) one of either:(i) automated tracking of body movements in captured digital imagery of said user-performance; or(ii) outputting of said captured digital imagery to another computing resource configured to perform said automated tracking of the body movements in said captured digital imagery of said userperformance.

6. The system of any one of claims 2 to 5 wherein step (f) comprises the tracking of said body movements, and said tracking is based, at least in part, by detected presence or movement within, among or between the plurality of zones.

7. The system of any one of claims 2 to 6 the zones are geofenced from one another, and geofenced virtual boundaries between said smaller zones are implemented for the tracking of said body movements in step (f).

8. The system of any one of claims 2 to 7 wherein step (f) comprises the tracking of said body movements, and said tracking of said body movements comprises: for each of at least some captured image frames of the user performance: detecting anatomical key points in the captured image frame and deriving therefrom a pose estimation of the user’s body;determining which of the detected anatomical key points reside in which of the zones; and storing a frame-specific datapoint of the user performance that comprises, for each detected anatomical key point, identification of which of the zones is occupied thereby; and concatenating a sequence of frame-specific datapoints as a dataset representative of said user performance.

9. The system of claim 8 wherein the steps further comprise performing a computer automated evaluation of the recorded user-performance against a virtual reference model composed of a comparative dataset of equivalent frame-specific datapoints derived from another captured performance of the sports training exercise by another person on an equivalent mat.

10. The system of claim 8 or 9 wherein said datapoint further comprises, for at least some of said anatomical key points, a respective joint angle calculated in accordance with said pose estimation.11 . The system of any one of claims 1 to 8 wherein the steps further comprise performing a computer automated evaluation of the recorded user-performance against a virtual reference model derived from another captured performance of the sports training exercise by another person.

12. The system of any one of claims 9 to 11 wherein the instructional guidance comprises a video recording of said another captured performance.

13. The system of any one of claims 9 to 11 wherein the instructional guidance comprises an animated reconstruction of said another captured performance, reconstructed from the virtual reference model.

14. The system of claim 13 wherein said animated reconstruction features an animated avatar of the user.

15. The system of any one of claims 9 to 14 wherein said computer automated evaluation comprises a time-series analysis of respective datasets from the userperformance and said another captured performance to evaluate the user-performance against said another captured performance based, at least in part, on time warping needed to achieve optimal sequence matching therebetween.

16. The system of any preceding claim wherein the sports training mat comprises a plurality of human discernible target indica on the topside thereof, and said instructional guidance comprises identification of one or more of said target indicia and identification of a user action to be taken relative thereto.

17. The system of claim 16 wherein at least a subset of said human discernible target indicia and at least a subset of the machine-readable markers of are of coincident location to on another on the topside of the mat.

18. The system of claim 17 wherein human discernible target indica are defined by, or are integral parts of, the machine-readable markers.

19. A computer implemented method executed by a computing system comprising one or more processors and one or more non-transitory computer readable media having stored therein statements and instructions executable by one more processors, of which at least one of said processors and at least one of said one or more non-transitory computer readable media are embodied in a camera-equipped mobile computing having a monoscopic camera whose field of view (FOV) is aimed or aimable in a manner encompassing a sports training mat comprising a plurality of machine-detectable markers on a topside of said sports training mat, said method comprising the steps recited in any preceding claim.

20. One or more non-transitory computer readable media having stored therein executable statements and instructions for execution by one or more processors to perform, when executed, the steps recited in any preceding claim.