Apparatus, head-mounted display, method
Patent Information
- Application Number
- PCT/EP2026/058518
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058518_01102026_PF_FP_ABST
Abstract
Description
[0001] Apparatus, head-mounted display, method
[0002] Field
[0003] The present disclosure relates to an apparatus for real-time coaching of a user, a headmounted display, and a method for real-time coaching of a user.
[0004] Background
[0005] Head-mounted displays (HMDs) are wearable devices equipped with visual output components that present digital content directly within the user’s field of view. These displays may incorporate additional technologies such as cameras, motion sensors, and augmented reality (AR) systems to capture or enhance visual information. HMDs can be used in various applications, including entertainment, training, and navigation.
[0006] Unmanned aerial vehicles (UAVs), commonly referred to as drones, are autonomous or remotely controlled aircraft equipped with cameras, sensors, and communication modules. Drones can capture visual and environmental data from an external perspective, providing aerial imaging and tracking capabilities. They are used in fields such as mapping, surveillance, and motion analysis.
[0007] Advancements in sensor technology, computer vision, and artificial intelligence have enabled improved tracking, imaging, and data processing in both wearable and external imaging systems. These technologies allow for enhanced visualization, analysis, and interaction across various industries, including leisure applications, industrial applications, or the like.
[0008] Summary
[0009] According to a first aspect, the disclosure provides an apparatus for real-time coaching of a user during a sporting activity. The apparatus includes processing circuitry configured to receive, in real-time during the sporting activity, first image data from an egocentric view of the user. The processing circuitry is further configured to determine, from the first image data, contextual conditions for the sporting activity. The processing circuitry is further configured to receive, in real-time during the sporting activity, second image data of the user from an exo-centric view. The processing circuitry is further configured to determine, from the second image data, one or more performance parameters of the user. The processing circuitry is furtherconfigured to determine, based on the contextual conditions and the one or more performance parameters of the user, a corrective coaching instruction for the user using a trained machine-learning model that is specific to the sporting activity. The processing circuitry is further configured to output the corrective coaching instruction to the user.
[0010] According to a second aspect, the disclosure provides a head-mounted display including at least one display area, at least one image sensor, and an apparatus for real-time coaching according to the first aspect.
[0011] According to a third aspect, the disclosure provides a method for real-time coaching of a user during a sporting activity. The method includes receiving, in real-time during the sporting activity, first image data from an egocentric view of the user. The method further includes determining, from the first image data, contextual conditions for the sporting activity. The method further includes receiving, in real-time during the sporting activity, second image data of the user from an exocentric view. The method further includes determining, from the second image data, one or more performance parameters of the user. The method further includes determining, based on the contextual conditions and the one or more performance parameters of the user, a corrective coaching instruction for the user using a trained machine-learning model that is specific to the sporting activity. The method further includes outputting the corrective coaching instruction to the user.
[0012] Further aspects are set forth in the dependent claims, the drawings, and the following description.
[0013] Brief description of the Figures
[0014] Some examples of apparatuses and / or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
[0015] Fig. 1 depicts an apparatus for real-time coaching according to the present disclosure;
[0016] Fig. 2 depicts an example of a head-mounted display according to the present disclosure;
[0017] Fig. 3 depicts a further example of a head-mounted display according to the present disclosure;Fig. 4 depicts an exemplary hardware configuration of a head-mounted display according to the present disclosure;
[0018] Fig. 5 depicts a flowchart of a method for real-time coaching according to the present disclosure;
[0019] Fig. 6 depicts an exemplary coaching scenario in which a system according to the present disclosure is applied in the context of skiing;
[0020] Fig. 7 depicts an exemplary coaching scenario in which a system according to the present disclosure is applied in the context of parkour;
[0021] Fig. 8 depicts an exemplary coaching scenario in which a system according to the present disclosure is applied in the context of racecars;
[0022] Fig. 9 depicts an exemplary coaching scenario in which a system according to the present disclosure is applied in the context of cycling; and
[0023] Fig. 10 depicts an exemplary coaching scenario in which a system according to the present disclosure is applied in the context of climbing.
[0024] Detailed Description
[0025] Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
[0026] Throughout the description of the figures same or similar reference numerals refer to same or similar elements and / or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and / or areas in the figures may also be exaggerated for clarification.
[0027] When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "atleast one of A and B" or "A and / or B" may be used. This applies equivalently to combinations of more than two elements.
[0028] If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and / or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and / or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or a group thereof.
[0029] Fig. 1 depicts an apparatus 100 for real-time coaching of a user during a sporting activity.
[0030] The apparatus may refer to any entity that includes one or more components arranged or interconnected to perform a particular function or set of functions as described herein. The apparatus 100 may include any physical device, system, structure, or the like, including standalone units and integrated arrangements. Hence, the apparatus 100 may include various configurations, and integrated systems that embody the concept disclosed herein.
[0031] The apparatus 100 includes processing circuitry 110. Processing circuitry may refer to one or more hardware components configured to execute operations on data, signals, instructions, or the like. For example, the processing circuitry 110 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a system-on-a-chip (SoC), a neuromorphic processor, a field programmable gate array (FPGA), or the like. The processing circuitry 110 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and / or non-volatile memory. For example, the apparatus 100 may include memory configured to store instructions, which when executed by the processing circuitry 110, cause the processing circuitry 110 to perform the steps and methods described herein.
[0032] As indicated above, the apparatus may be suitable for real-time coaching of a user during a sporting activityReal-time coaching may refer to the continuous or (near-)instantaneous provision of guidance, feedback, or instructions to the user based on data collected during the sporting activity. This term may encompass various forms of feedback, including auditory, visual, haptic, or other sensory outputs, and may be based on at least one of pre-programmed instructions, adaptive algorithms, real-time data analysis, or the like. Thereby, timely and relevant assistance to the user may be given.
[0033] A sporting activity (or sports activity) may refer to any physical exercise, movement, or performance of the user undertaken for fitness, recreation, competition, or training purposes. This term may include individual and team sports, indoor and outdoor activities, endurance and strength-based exercises, and any other form of physical exertion where real-time coaching may enhance performance, safety, or technique. It is not limited to formally recognized sports and encompasses both structured and unstructured physical activities.
[0034] Accordingly, a user may refer to an individual who interacts with, operates, or is an intended recipient of the real-time coaching provided by the apparatus. The term may encompass athletes, recreational participants, and any person engaging in a sporting activity, regardless of skill level, experience, or professional status. It may include individuals actively performing the sporting activity as well as those receiving guidance or training.
[0035] The processing circuitry 110 is configured to receive, in real-time during the sporting activity, first image data from an egocentric view of the user;
[0036] “Receiving” may refer to the act of acquiring, obtaining, or capturing data from a source, such as a sensor or camera, and making it available for processing, storage, or analysis within the apparatus. The reception may occur wirelessly or via wired communication, and it can involve direct input from an integrated (within the apparatus 100) or external (outside of the apparatus 100) imaging device.
[0037] An imaging device may refer to a hardware component or system configured to capture, record, or generate visual representations of a scene, object, or environment. This term may broadly encompass various types of image-capturing technologies, including but not limited to digital cameras, optical sensors, infrared cameras, depth cameras, stereo cameras, and other imaging systems capable of producing still images, video sequences, or other forms of visual data (image data).An imaging device may operate using different imaging techniques, such as visible light capture, thermal imaging, depth sensing, or multispectral imaging. It may be integrated into a larger system, such as a wearable device (a device worn by the user, such as a headmounted display / device, smart lenses, or the like), mobile equipment, or the like, and may communicate with other components for processing, analysis, or real-time feedback generation. The imaging device may function autonomously, under user control, or as part of an automated system designed to extract information from captured image data.
[0038] As indicated above, the image data may be received in real-time. Similar to the term “realtime coaching”, this may refer to the case that image data acquisition, transmission, and subsequent processing may occur with minimal latency, allowing the received information to be acted upon without significant delay. Real-time operation may ensure that the system can provide instantaneous or near-instantaneous responses, enabling immediate coaching, feedback, or analysis based on the captured image data.
[0039] The first image data may indicate an egocentric view of the user. In other words, the first image data may be captured from a perspective that aligns with what the user sees / perceives. This may be achieved using a camera mounted on the user's head, helmet, eyewear, or another body-worn or equipment-mounted position that provides a “first-person” viewpoint (i.e. , the egocentric view). The egocentric perspective may ensure that the captured data (roughly) represents the user's field of vision, allowing for the real-time feedback / coaching based on what the user is seeing or experiencing during the sporting activity.
[0040] The processing circuitry 110 is further configured to determine, from the first image data, contextual conditions for the sporting activity.
[0041] Determining may refer to identifying, analyzing, extracting, or inferring information through computational processing, algorithmic evaluation, or pattern recognition techniques. The determination process may involve various methods such as image processing, artificial intelligence (Al)-based analysis, object detection, motion tracking, or environmental assessment.
[0042] For this purpose, the first image data may be used as input, either directly, or a processed version thereof. From the first image data, contextual condition may be determined. Contextual conditions may refer to one or more of environmental, situational, or event-specific factors that may influence the sporting activity or the user during the sporting activity. These conditions may include, but are not limited to at least one of environmental factors (e.g., lighting conditions, weather, terrain, or playing surface characteristics), user-related factors (e.g.,such as posture, movement patterns, position, or interaction with equipment), and external elements (e.g., presence of other participants, obstacles, or objects relevant to the sporting activity).
[0043] The processing circuitry 110 is further configured to receive, in real-time during the sporting activity, second image data of the user from an exocentric view. In other words, similar to the first image data, second image data may be received and a repetitive description of terms described above is omitted. The difference to the first image data is that the second image data represent or indicate an exocentric view of the user (and of their surroundings / environ-ment). An exocentric view may be refer to the case that the image data may be captured from a perspective external to the user, rather than from the user's own field of view. This exocentric viewpoint may be provided by a stationary or mobile camera, a drone, a sensor array, or another imaging device positioned to observe the user from the outside. This external perspective may allow for an independent evaluation of the user's actions, form, or performance in relation to the surroundings, other participants, or relevant objects. For example, the user’s form of an exercise may better be evaluated based on the exocentric view instead of the egocentric view. However, their form may also depend on the context of the sporting activity which may be evaluated based on the egocentric view.
[0044] Accordingly, the processing circuitry 110 is further configured to determine, from the second image data, one or more performance parameters of the user.
[0045] A performance parameter may refer to a (at least one) measurable characteristic or attribute to quantify or describe the user's execution of the sporting activity. These parameters may include, but are not limited to biomechanical metrics (e.g., posture, stance, joint angles, range of motion, balance, or the like), kinematic data (e.g., velocity, acceleration, stride length, angular momentum, or the like), positional tracking (e.g., spatial location, movement trajectory, relative positioning in relation to at least one of objects, boundaries, other participants, or the like), technique / form evaluation (e.g., as symmetry, consistency, adherence to optimal movement patterns, or the like). Additionally or alternatively, the one or more performance parameters include a posture of the user and positioning of equipment used for the sporting activity. A performance parameter may relate to the user engaged in the sporting activity. Accordingly, a performance parameter may serve to assess at least one of the user's efficiency, accuracy, and effectiveness in performing the given sport or exercise.
[0046] The processing circuitry 110 is further configured to determine, based on the contextual conditions and the one or more performance parameters of the user, a corrective coachinginstruction for the user using a trained machine-learning model that is specific to the sporting activity.
[0047] Both the data generated from the first image data and from the second image data (i.e. , the contextual conditions and the performance parameter) may thus be used for generating the corrective coaching instruction. The corrective coaching instruction may refer to a feedback message, directive, or guidance generated to help the user improve their performance of the sporting activity in real-time. This instruction may be provided in various forms, including auditory cues, visual overlays, haptic feedback, or textual recommendations. The corrective coaching instructions may be adapted to adjust or optimize aspects of the user's technique, movement patterns, strategy, or the like.
[0048] In more detail, the corrective coaching instruction may refer to a specific, actionable piece of feedback generated to help a user improve their technique, execution, or decision-making during a sporting activity. The instruction may be derived from an analysis of contextual conditions and the user’s performance parameters. The corrective coaching instruction may address an identified deviation from an optimal or intended technique. It may help the user correct errors, inefficiencies, or suboptimal movement patterns that may affect performance, safety, or effectiveness. The corrective coaching instruction may be adapted to the user's current skill level, movement style, and specific demands of the specific sporting activity. The corrective coaching instruction may provide guidance, e.g., based on best practices, expert knowledge, or learned patterns of the machine-learning model (discussed below). It may facilitate learning and improvement by providing meaningful feedback that can be immediately applied or incorporated into training. It may reinforce positive behavior by not only correcting mistakes but also promoting proper technique. The feedback may be clear, concise, and actionable, ensuring that the user understands how to implement the correction. It may include step-by-step guidance or real-time cues (e.g., "keep your knees aligned with your toes during the squat"). The instruction may be adaptive, meaning it may adjust based on continuous user input / performance and environmental conditions. The nature of the corrective coaching instruction may depend on the sporting activity, the user's current performance, and specific aspects that should be corrected.
[0049] For example, a biomechanical correction may be carried out (such as form and posture adjustments). For example, when the user is running, the coaching instruction may be: "Shorten your stride to reduce impact force and improve efficiency." When the user is playing tennis, the coaching instruction may be: "Keep your wrist firm during the backhand stroke to increase shot stability." When the user is weightlifting, the coaching instruction may be: "Engage yourcore and keep your back straight while performing the deadlift." If the user is skiing, the coaching instruction may be: "Shift your weight forward to maintain control on icy slopes."
[0050] Additionally or alternatively, a kinematic adjustments (e.g., relating to movement and motion efficiency) may be carried out. If the user is cycling, a coaching instruction may be: "Increase your cadence to maintain optimal power output and reduce muscle fatigue." If the user is swimming, the coaching instruction may be: "Improve your arm pull symmetry to maintain a straight trajectory in the water." If the user is playing golf, the coaching instruction may be: "Slow down your backswing to enhance shot accuracy and control."
[0051] Additionally or alternatively, tactical or strategic guidance (e.g., relating to decision-making adjustments) may be given as a coaching instruction. If the user is playing soccer, the coaching instruction may be: "Position yourself closer to the opponent to improve defensive pressure." If the user is playing basketball, the coaching instruction may be: "Look for open teammates when facing a double-team to improve passing options." If the user is executing martial arts, the coaching instruction may be: "Adjust your stance width to maintain better balance when evading attacks."
[0052] Additionally or alternatively, sensory cues may be given to the user as the coaching instruction. For example, auditory feedback may be given based on a given rhythm that matches an intended running / jumping rhythm (or the like). A corresponding coaching instruction may then be: "Increase your step frequency to match the given rhythm for endurance running." On the other hand, haptic feedback may be given as a subtle vibration (e.g., in a smart wearable), e.g., to indicate that the user is tilting too far forward during a squat.
[0053] As indicated above, the corrective coaching instruction may be generated using a trained machine-learning model that is specific to the sporting activity.
[0054] A (trained) machine-learning (or machine-learned) model may refer to a data structure and / or set of rules representing a statistical model that the processing circuitry 110 may use at least one machine-learning output. It should be noted that the present disclosure may be carried out with one machine-learning or multiple machine-learning models. The data structure and / or set of rules may represent learned knowledge (e.g. based on training performed by a machine-learning algorithm as described below). In machine-learning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of training data.The machine-learning model may be trained based on a machine-learning algorithm. The term "machine-learning algorithm" may denote a set of instructions that are used to create, train or use a machine-learning model. For the machine-learning model to determine an output, the machine-learning model may be trained using training data, as commonly known. By training the machine-learning model with a large set of training data and associated training information, the machine-learning model may learn to determine the output. By training the machine-learning model using training information, the machine-learning model may learns a transformation between training input data and appropriate output data.
[0055] The machine-learning model may be trained using training input data. For example, the machine-learning model may be trained using a training method called "supervised learning". In supervised learning, the machine-learning model may be trained using a plurality of training samples, wherein each sample may include a plurality of input data values, and a plurality of desired output values, i.e., each training sample may be associated with a desired output value. By specifying both training samples and desired output values, the machine-learning model may learn which output value to provide based on an input sample that is similar to the samples provided during the training.
[0056] Apart from supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples may lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g. a classification algorithm or a similarity learning algorithm). Classification algorithms may be used as the desired outputs of the trained machine-learning model may be restricted to a limited set of values (categorical variables), i.e., the input is classified to one of the limited set of values. Similarity learning algorithms may be similar to classification algorithms but may be based on learning from examples using a similarity function that measures how similar two related objects are.
[0057] Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data may be supplied and an unsupervised learning algorithm may be used to find structure in the input data (e.g. by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data including a plurality of input values into subsets (clusters) so that input values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input values that are included in other clusters.Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the machine-learning model. In reinforcement learning, one or more software actors (called "software agents") may be trained to take actions in an environment. Based on the taken actions, a reward may be calculated. Reinforcement learning may be based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
[0058] Furthermore, additional techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and / or the machine-learning algorithm may include a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.
[0059] For example, the machine-learning model may be an Artificial Neural Network (ANN). An ANN may refer to a system that is inspired by biological neural networks, such as in retina or in brains. ANNs may include a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes. There may be three types of nodes: Input nodes that receive input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node may represent an artificial neuron. Each edge may transmit information from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs). The inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input. The weight of nodes and / or of edges may be adjusted in the learning process. In other words, the training of an ANN may include adjusting the weights of the nodes and / or edges of the ANN, i.e., to achieve a desired output for a given input.
[0060] Alternatively, the machine-learning model may include a different structure and, e.g., be a support vector machine, a random forest model or a gradient boosting model. Alternatively, the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
[0061] In examples, the machine-learning model may be a combination of any of above examples.In the context of the present disclosure, the machine-learning model may use the first image data (or the contextual conditions derived from the first image data) and the second image (or the one or more performance parameters derived from the second image data) data as input. For example, the contextual conditions may be derived by the same or a different machine-learning model. The same applies to the one or more performance parameters.
[0062] From the input data, the machine-learning model may extract the environmental conditions, such as at least one of environmental factors (e.g., surface conditions, weather, terrain type, lighting conditions, or the like), dynamic contexts (e.g., presence of obstacles, oppo-nent / player positions, spatial constraints, or the like), and real-time adjustments (e.g., changing conditions that require immediate adaptation (e.g., wind strength in outdoor sports)).
[0063] Furthermore, the one or more performance parameters may be extracted, such as at least one of biomechanical metrics (e.g., posture, stance, joint angles, movement symmetry, or the like), kinematic data (e.g., velocity, acceleration, stride length, angular momentum, or the like), positional tracking (e.g., the user’s relative position within the activity space, balance, body alignment, or the like).
[0064] The machine-learning model may further consider at least one of historical and expert-devel data, such as at least one of the user’s past performance data (e.g., previous movement patterns, improvement trends, common mistakes, or the like), expert reference data (e.g., professional athlete techniques, best-practice movement models, or the like) and training set data (e.g., pre-recorded datasets of correct and incorrect techniques across various skill levels).
[0065] Additionally or alternatively, the machine-learning model may consider additional sensor data, such as at least one of wearable sensor data (e.g., accelerometers, gyroscopes, pressure sensors, electromyography (EMG) signals, or the like), heart rate & fatigue metrics (e.g., physiological data affecting user performance and ability to correct movements, or the like)..
[0066] Based on the input data, the machine-learning model may generate the corrective coaching instruction (i.e., output data) including at least one of actionable feedback (e.g., specific movement corrections, real-time cues, and post-activity recommendations.
[0067] In some examples, the machine-learning model may adjust its feedback, for example based on the user’s progress.For training the machine-learning model, any of the above-mentioned model may be used, which may also depend on the specific sporting activity. As training data sources, a (labeled) set of data indicating people doing the specific sporting activity may be used. For example, data collected from professional athletes performing optimal techniques may be used and / or data from amateur and novice athletes with (labeled) movement errors may be used. It should be noted that the form used by a professional athlete is not necessarily the optimal form for the user at their training stage. The model may consider this accordingly, e.g., based on data indicating novices or intermediates with a similar body as the user doing an optimal movement.
[0068] Additionally or alternatively, annotated video and motion capture data, such as at least one of (high-speed) camera footage with biomechanical analysis, and motion capture recordings showing correct and incorrect movement patterns
[0069] If supervised learning is used, the model may be fed with the input data labeled with the correct and incorrect movement patterns. The model may learn to classify movements as optimal, suboptimal, or incorrect, e.g., based on expert annotations.
[0070] If unsupervised learning is used, the model may cluster and analyze movement patterns without predefined labels. It may detect recurring deviations from optimal technique that may not have been explicitly labeled. For example, common movement mistakes may be identified across multiple users by clustering similar errors.
[0071] If reinforcement learning is used, the model may continuously improve by testing corrective instructions and analyzing user responses. It may refine feedback by tracking whether the user successfully corrects their movements over time. For example, feedback intensity may be adjusted based on how well the user responds to previous corrections.
[0072] Additionally or alternatively, the model may apply at least one of validation and continuous learning. For example, the model may be validated against expert values. In such an example, the model’s corrective instructions may be tested and compared against expert human coaches' feedback. Fine-tuning may occur based on discrepancies between Al-generated corrections and human-coach (expert) recommendations. Continuous learning may be carried out based on user interaction. For example, the model may refine its coaching strategy as more data is collected from user sessions. Furthermore personalized coaching models may be developed for individual users based on historical improvement patterns.The processing circuitry 110 is then configured to output the corrective coaching instruction to the user. As indicated above, there may be various ways how the instruction is output, e.g., based on sound, text, speech, haptic feedback, or the like.
[0073] In some examples, the first image data are received from an image sensor of a head-mounted display worn by the user. The head-mounted display may be a wearable device such as augmented reality (AR) or virtual reality (VR) glasses, smart goggles, or a helmet-mounted display, incorporating one or more cameras or optical sensors that continuously or intermittently capture image data in real time.
[0074] In some examples, the second image data are received from an image sensor of an unmanned aerial vehicle (e.g., a drone) following the user. The drone may be equipped with one or more cameras or optical sensors that capture the exocentric view of the user. The drone may autonomously track the user’s location, e.g., using computer vision, GPS, motiontracking algorithms, or the like, ensuring that the second image data provide a consistent external viewpoint of the sporting activity.
[0075] By combining the first image data from the head-mounted display with the second image data from the drone, a comprehensive analysis of the user's performance and environmental context may be conducted. This integration may allow for real-time coaching, adaptive feedback, and enhanced motion analysis, enhancing accuracy and effectiveness of corrective coaching instructions.
[0076] In some examples, the processing circuitry is further configured to invoke a digital twin of an environment of the sporting activity and provide the contextual conditions and the one or more performance parameters to the digital twin for generating the corrective coaching instruction.
[0077] Invocation of a digital twin may refer to activating, initializing, or accessing a digital twin, i.e., a virtual, data-driven model that digitally replicates the real-world environment of the sporting activity. The digital twin may refer to a high-fidelity, dynamic, and continuously updated virtual representation of the physical setting in which the sporting activity is performed. The digital twin may be created using 3D spatial mapping, sensor fusion, computer vision, LiDAR scanning, GPS data, Al-based modeling, or the like, to reconstruct a realistic and interactive virtual environment. The digital twin may thus represent the environment of the sporting activity including all relevant external elements such as the playing field, training area, terrain, obstacles, slope, weather conditions, and any objects or participants that may influence the user’s performance. The digital twin may be deployed on cloud-based computing platforms, localprocessing units, or edge devices that receive and process real-time data from imaging devices and other sensors.
[0078] For updating the digital twin for further enhance the virtualization of the real environment, the contextual conditions may be provided to the digital twin to accurately reflect the sporting environment. For example, at least one of environmental factors (e.g., lighting, temperature, wind speed, humidity, surface conditions, or the like), spatial constraints (e.g., boundaries of the sporting activity, proximity to obstacles, positioning of other players or objects, or the like)and dynamic changes (e.g., moving objects, shifting weather conditions, variations in terrain, or the like) may be provided to the digital twin. In other words, in some examples, the processing circuitry 110 is further configured to adapt the environment in the digital twin based on the contextual conditions.
[0079] Furthermore, the one or more performance parameters may be provided to the digital twin to model the user’s performance within the digital twin. Thereby, performance analysis may be carried out int eh digital twin.
[0080] By integrating contextual conditions and performance parameters, the digital twin may create a real-time, data-enriched simulation of the user within the sporting environment, enabling advanced performance analysis and predictive modeling.
[0081] Accordingly, the corrective coaching instruction may be generated using the digital twin. Once the digital twin has been populated with the contextual and performance data, it may be used to simulate the user’s actions and compare them against optimal movement patterns or expert models, as discussed above. The machine-learning model discussed above may thus be applied to the simulated user in the digital twin.
[0082] In other words, in some examples, the processing circuitry 110 is further configured to simulate, based on the adapted environment in the digital twin and using the one or more performance parameters, the user doing the sporting activity in the digital twin. In such examples, the processing circuitry 110 may be further configured to apply, for the simulated user doing the sporting activity in the adapted environment in the digital twin, the machine-learning model to determine the corrective coaching instruction.
[0083] In some examples, the processing circuitry 110 is further configured to overlay, on the first image data, an indication of an improved performance determined based on the trainedperformance model. In such examples, the processing circuitry 110 may be further configured to cause output of the overlaid first image data for the user.
[0084] The overlaying may refer to displaying virtual elements directly onto the user's real-world perspective as captured by the image sensor of the HMD. The overlay may include using of augmented reality (AR) elements, such as graphical indicators, virtual objects, or animated guides, which are superimposed onto the user's field of view through the HMD display. If the HMD is a non-see-through HMD, a virtual reality (VR) environment may be generated based directly on the first image data or based on data of the digital twin in which the indication is overlayed. The overlay may be dynamically positioned in real-time to ensure accurate alignment with the real-world environment, allowing the user to perceive it as naturally integrated into their surroundings.
[0085] The indication of an improved performance may refer to a visual representation of how the user should carry out the sporting activity. This may take various forms, such as at least one of overlaying a representation of a virtual coach (e.g., a computer-generated avatar that performs the correct movement or follows the ideal trajectory, providing a real-time reference for the user), a motion path visualization (e.g., highlighted guidelines, arrows, or trajectory markers indicating where and how the user should move), a ghost overlay of correct technique (e.g., a semi-transparent version of the user’s body or equipment performing the corrected movement, allowing for direct comparison with the user’s current movement), and key Point highlights (e.g., markers or outlines emphasizing critical body positions, such as ideal knee bending angles or shoulder alignment).
[0086] In some examples, the contextual conditions are first contextual conditions. In such examples, the processing circuitry 110 may be further configured to determine, from the second image data, second contextual conditions of an environment of the sporting activity.
[0087] The second contextual conditions may provide additional or complementary environmental information that are not directly visible to the user, such as relative positioning within the sporting environment, distances to external objects, movement of other participants, broader spatial conditions affecting performance, obstacle that are not visible to the user, or the like.
[0088] By combining both first contextual conditions (from the user's perspective) and second contextual conditions (from an external viewpoint), the processing circuitry 110 may be able to generate a more comprehensive understanding of the sporting environment, enabling more accurate performance evaluation, real-time coaching, and adaptive feedback.For example, based on the second contextual conditions, if an obstacle (or an accident in front the user) or some other hazard is recognized, the processing circuitry 110 may be further configured to warn the user of the potential hazard.
[0089] In the following, examples of head-mounted displays (HMDs) are discussed. It should be understood that that the following description is not intended to be limiting to the elements and units described under reference of Figs. 2 to 4. Thus, according to the present disclosure, it may be sufficient when an HMD includes at least one display area, at least one image sensor, an apparatus for real-time coaching, as described under reference of Fig. 1 or any example relating to Fig. 1.
[0090] Fig. 2 is a diagram for explaining an external appearance example of an apparatus (or information processing system) 1001 according to the present disclosure. As illustrated in Fig. 2, the information processing system 1001 according to the present embodiment is configured as a head mounted display (HMD).
[0091] Fig. 2, depicts an HMD 1001 that includes an output mechanism unit 1011 and a mounting mechanism unit 1012. The mounting mechanism unit 1012 includes a mounting band 1013. In wearing the HMD 1001, a user wraps the mounting band 1013 around his / her head, so that the HMD 1001 is fixed to the user’s head. Note that the mounting band 1013 is not necessary wrapped around the user’s head as long as the HMD 1001 is fixed to the user’s head.
[0092] The output mechanism unit 1011 includes a housing 1014 having a shape that covers the user’s left and right eyes in a state in which the user wears the HMD 1001. The housing 1014 accommodates a display panel located so as to face the user’s eyes when the user wears the HMD 1001. The housing 1014 may further accommodate a lens located between the display panel (display unit 2005 (Fig. 4)) and the user’s eyes when the user wears the HMD 1001. The lens increases the user’s viewing angle. The display panel is divided into left and right regions on which stereo images corresponding to the parallax between the left and right eyes may be respectively displayed. The stereo images thus displayed may achieve stereoscopic view.
[0093] The HMD 1001 may further include speakers and earphones at positions corresponding to the user’s ears when the user wears the HMD 1001. In this example, the HMD 1001 includes a camera 1015 disposed on a front surface of the housing 1014. The camera 1015 capturesa moving image of a real space around the user in a field of view corresponding to the user’s line-of-sight.
[0094] The camera 1015, for example, includes a photodetection device such as an image sensor (e.g., a charge coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) or a ranging sensor, and an optical system such as an image forming lens. For example, in Fig. 2, the camera 1015 is a stereo camera that captures images of a space forward of the user from left and right viewpoints corresponding to the user’s left and right eyes. Note that the camera 1015 is not limited thereto. The camera 1015 may be a singlelens camera, or may be a multi-lens camera including three or more lenses. The camera 1015 may also be a combination of different types of sensors. In the application of hand tracking, the camera 1015 may be provided so as to capture an image of a space below the information processing system. In the application of eye tracking or face tracking, the camera 1015 may be provided so as to capture an image of the user’s eyes or an image of the user’s face.
[0095] The HMD 1001 also includes a sensor 2008 (Fig. 4). The sensor may include at least one of various sensors, such as an acceleration sensor, a gyro sensor, an angular velocity sensor, and a geomagnetic sensor, for deriving the motion, posture, position, and the like of the HMD 1001.
[0096] The HMD 1001 may be connected by wireless communication to another processing apparatus, or may establish a wired connection using a universal serial bus (USB) or the like.
[0097] In this case, the HMD 1001 may be configured to start an online application such as a game in which a plurality of users can participate via a network. In this case, the HMD 1001 performs predetermined processing on an image captured by the camera 1015. The HMD 1001 then generates a display image in the field of view of the camera 1015 and displays the display image.
[0098] Here, the content of the display image is not particularly limited, and various display images may be used depending on a function required of the system by the user, the content of a launched application, and the like.
[0099] For example, the HMD 1001 may perform certain processing on an image captured by the camera 1015, or may superimpose and draw a virtual object that interacts with an image of a real object. Alternatively, the HMD 1001 may draw a virtual world in a field of viewcorresponding to the user’s field of view on the basis of a captured image, a measurement value by a motion sensor in the sensor group of the HMD 1001, and the like.
[0100] Representative examples of these aspects include virtual reality (VR), augmented reality (AR), and mixed reality (MR). In addition, an image captured by the camera 1015 is used as a display image, so that a video see through (VST) form may be achieved, by which a real world can be seen through a screen on the HMD 1001.
[0101] Fig. 3 is a diagram for explaining an external appearance example of an information processing system 1101 according to the present disclosure. As illustrated in Fig. 3, the information processing system 1101 according to the present embodiment is configured as an eyewear-type HMD.
[0102] An HMD main body 1111 is used with the HMD main body 1111 mounted to the user’s head. The HMD main body 11 includes a front portion 1112, a right temple portion 1113 disposed on a right side of the front portion 1112, a left temple portion 1114 disposed on a left side of the front portion 1112, and a glass portion 1115 attached to a lower side of the front portion 1112. Note that although the glass portion 1115 illustrated in Fig. 3 is of an integrated type, the glass portion 1115 may be constituted of a pair of glasses separated from each other so as to cover the left and right eyes, respectively, or may be constituted of a single piece of glass so as to cover only one of the left and right eyes.
[0103] A display unit 1103 is of a see-through type, and is disposed on a surface of the glass portion 1115. The display unit 1103 performs AR display of a virtual object, in accordance with control by a processing circuit 2001. Note that the display unit 1103 may be of a non-see-through type. In this case, AR display is performed in such a manner that the display unit 1103 displays an image including an image currently captured by a camera 1104 and a virtual object superimposed on the currently captured image.
[0104] A camera 1104 includes, for example, a photodetection device such as an image sensor (e.g., a charge coupled device (CCD) sensor, a complemented metal oxide semiconductor (CMOS) sensor) ora ranging sensor, and an optical system such as an image forming lens. The camera 1104 is disposed outward on an outer surface of the front portion 1112. The camera 1104 captures an image of an object in a real space, and outputs image information obtained by the image capture, to the processing circuit 2001. In Fig. 3, for example, two cameras 1104 are disposed on the front portion 1112 with a predetermined spacing in a lateral direction. Note that The camera 1104 is not limited thereto. The camera 1104 may be a single-lenscamera, or may be a multi-lens camera including three or more lenses. The photodetection unit 1015 may also be a combination of different types of sensors. In the application of hand tracking, the camera 1104 may be provided so as to capture an image of a space below the information processing system. In the application of eye tracking or face tracking, the camera 1104 may be provided so as to capture an image of the user’s eyes or an image of the user’s face.
[0105] The eyewear-type HMD 1101 also includes a sensor 2008 (Fig. 4). The sensor may include at least one of various sensors, such as an acceleration sensor, a gyro sensor, an angular velocity sensor, and a geomagnetic sensor, for deriving the motion, posture, position, and the like of the HMD 1001.
[0106] Next, a hardware configuration example of the information processing system (the HMD 1001 or the eyewear-type HMD 1101) is described with reference to Fig. 4. As illustrated in Fig. 4, the hardware of the information processing system include the processing circuit 2001 , a memory 2002, a camera 2003, a display unit 2005, an input unit 2006, an output unit 2007, a sensor 2008, a communication interface (IF) 2009, an external network 2010, and a secondary storage device 2011 which are connected to each other via a bus 2012 and are capable of exchanging data and programs with each other.
[0107] The processing circuit 2001 operates on the basis of a program stored in the memory 2002 or the secondary storage device 2011, and controls the whole of the operation of the information processing system 1001 or 1101. The processing circuit is, for example, a processor that reads each program from the memory 2002 and executes each program, thereby achieving a function corresponding to each program thus read. The processor may include any one or more of, for example, a multicore processor, a controller, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or an equivalent individual logic circuit or integrated logic circuit. The processing circuit may be achieved using a plurality of separate chips.
[0108] The memory 2002 may include a memory of any format for storing data and executable software instructions, the memory being practicable using, for example, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electronically erasable programmable read-only memory (EEPROM), a semiconductor memory element such as a flash memory, a hard disk, an optical disk, or the like.The camera 2003 corresponds to the camera 1015 illustrated in Fig. 2 and the camera 1104 illustrated in Fig. 3, and includes a photodetection device such as an image sensor (e.g., a charge coupled device (CCD) sensor, a complemented metal oxide semiconductor (CMOS) sensor) or a ranging sensor, and an optical system such as an image forming lens.
[0109] The display unit 2005 is a display panel accommodated in a housing, and includes a liquid crystal display (LCD) or a display device made of an organic electroluminescence (EL) material or the like.
[0110] Although not illustrated in Figs. 2 and 3, the input unit 2006 includes an input device, such as a keyboard, a mouse, a touch panel, a microphone, or a controller, through which the user inputs an operation command. The input unit 2006 supplies various input signals to the control unit 2001.
[0111] The output unit 2007 includes an audio output device such as a speaker, a force sense presenting device, an odor presenting device, or the like. The output unit 2007 is controlled by the processing circuit 2001 to output a processing result in the form of a sound, a force sense, or an odor.
[0112] The sensor 2008 may include at least one of various sensors, such as an acceleration sensor, a gyro sensor, an angular velocity sensor, and a geomagnetic sensor, for deriving the motion, posture, position, and the like of each of the HMDs 1001 and 1101. The sensor 2008 may also include a biological sensor that senses human biological information and a pressure sensor that senses an input.
[0113] The communication interface 2009 is an interface via which each of the information processing systems 1001 and 1101 is connected to an external network 2010. The eyeweartype HMD 1101 communicates with a smartphone or an external device except a smartphone (e.g., a personal computer (PC), a server device on a network, or the like) in a wired or wireless manner. For example, the processing circuit 2001 receives data from another equipment via the communication interface 2009, and transmits data generated by the processing circuit 2001 to another equipment via the communication interface 2009.
[0114] An example of an information processing system to which the technology according to the present disclosure may be applied has been described above. The technology according to the present disclosure may be applied to any of the HMDs described under reference of Figs.
[0115] 2 to 4 without limiting the present disclosure in that regard.Fig. 5 depicts a flowchart of a method 200 for real-time coaching of a user during a sporting activity. The method 500 includes receiving, 210, in real-time during the sporting activity, first image data from an egocentric view of the user. The method 200 further includes determining, 220, from the first image data, contextual conditions for the sporting activity. The method 200 further includes receiving, 230, in real-time during the sporting activity, second image data of the user from an exocentric view. The method 200 further includes determining, 240, from the second image data, one or more performance parameters of the user. The method 200 further includes determining, 250, based on the contextual conditions and the one or more performance parameters of the user, a corrective coaching instruction for the user using a trained machine-learning model that is specific to the sporting activity. The method 200 further includes outputting, 260, the corrective coaching instruction to the user.
[0116] Fig. 6 depicts an exemplary coaching scenario according to the present disclosure in which a user’s 600 skiing is real-time coached. The user 600 wears an HMD 610 including a display showing a virtual coach skiing in front of the user 600. A camera of the HMD 610 is used to acquire first image data for determining contextual conditions, as discussed herein. Moreover, a drone 620 follows the user and acquires second image data for determining one or more performance parameters.
[0117] In other words, Fig. 6 illustrates a real-time coaching system for skiing that combines augmented reality, a head-mounted display, and an unmanned aerial vehicle to enhance the skier’s performance through visual overlays and external tracking. The skier, who is actively navigating a slope, is wearing the HMD 610 that captures the first image data from an egocentric perspective. The drone 620 that is positioned behind the skier (without limiting the present disclosure in that regard: position of the drone may depend on the circumstances), capturing second image data from an exocentric perspective. This external viewpoint allows for a broader analysis of the skier’s posture 630, technique, and positioning on the slope. The drone 620 continuously tracks the skier’s motion, providing valuable input to a trained performance model that processes the collected data.
[0118] By integrating both egocentric and exocentric image data, the system can analyze the skier’s performance parameters and contextual conditions of the environment. The skier receives real-time corrective coaching instructions, which are overlaid onto the first-person view in the head-mounted display. In this scenario, the augmented reality system ensures that the virtual coach or corrective guidance appears naturally aligned within the skier’s environment, enhancing the perception of real-time assistance.This setup allows for an advanced form of interactive coaching, where the skier can immediately adjust their technique based on both self-perceived movements and Al-driven analysis derived from external monitoring. The combination of augmented reality and drone-based tracking enables precise, adaptive feedback, improving the skier’s technique and overall efficiency on the slope.
[0119] In more general terms, in the context of skiing (without limiting the present disclosure in that regard), instead of conventional ski goggles, the skier may wear AR glasses, which superimpose vital information about the ski terrain. The glasses display details such as slope gradients, snow depth, and the location of icy patches or obstacles, gathered from the surrounding environment in real time.
[0120] A lightweight, autonomous drone may follow the skier and scan the environment using high-resolution cameras and LiDAR sensors. The drone constructs a 3D map of the surroundings, processing gradient changes, snow drifts, and hazard zones (e.g., to update the digital twin or instead of using a digital twin). This information is transmitted instantly to the AR glasses, allowing the skier to "see" and react to subtle terrain changes that may be difficult to notice unaided.
[0121] The system may leverage the computing power of Al enhanced processors (such as Sony IMX500 intelligent vision sensor module (without limiting the present disclosure in that regard) embedded within the drone and / or the AR glasses). Such a may enable real-time analysis of both standard and skier-specific postures, facilitating immediate feedback and corrections. By utilizing such a sensor’s computing capabilities, the drone may process visual data locally, eliminating latency and enhancing the efficiency of skiing practice. Skier movements may be captured and compared against pre-stored optimal skiing data, allowing the system to provide timely corrective suggestions through the AR glasses and earpiece. The virtual coach demonstrates the correct movements, enabling the skier to adjust their posture instantly.
[0122] The system may further include pose estimation algorithms that track the skier’s body movements, analyzing their posture and technique. This data may be compared to an optimal skiing form stored in the system, which has been pre-trained using professional skiing data. Whenever the skier’s posture deviates from the ideal, real-time feedback is delivered, e.g., via an earpiece, instructing the user to adjust their movements.The AR glasses may further display a virtual skiing instructor who demonstrates the ideal posture and movements for skiing. The virtual instructor’s position, movements, and pace may be dynamically adjusted according to the skier’s current skill level and the terrain conditions. This virtual coach may guide the skier in real-time, showing the correct motions to follow, making learning both visual and interactive.
[0123] To encourage progress and adherence to correct technique, the system may feature a scoring mechanism. Points may be awarded based on how close the skier’s posture matches the optimal skiing form and how well they follow the virtual instructor's movements. Scores may be presented on the AR display, allowing skiers to track their improvement over time.
[0124] For example, when a skier puts on the AR glasses (HMD), and the drone is activated to follow them as they begin skiing down a slope, the AR glasses may immediately show the skier a visual representation of the terrain ahead, highlighting a steep section and some icy areas on the side. As the skier descends, the an Al enhanced processor processes the skier’s movements in real-time, detecting that their knees are too rigid. The system delivers a voice command to the earpiece, prompting them to relax their knees and shift their weight. Simultaneously, the virtual instructor in the AR display executes the correct maneuver, giving the skier a visual cue to mimic. As the skier adjusts, the system calculates an improvement in their form and awards points based on their adherence to the suggested posture.
[0125] However, conventional ski training methods may be limited by the lack of real-time, on-the-go feedback. Traditional methods may rely on instructors, who may not always be available, and video analysis, which is post-factum and which may require significant downtime for review. Such approaches may result in slower learning and improvement rates. Moreover, skiers, particularly beginners, often struggle to accurately assess terrain conditions and correct their posture without expert guidance, increasing the risk of injury. These issues may also be relevant for other types of sports, such as exemplarily described in the following.
[0126] Fig. 7 illustrates a real-time coaching system for parkour, where a combination of a headmounted display (HMD), an augmented reality interface, and a drone-based external camera is used to analyze and improve the user’s movements. The scene captures a parkour athlete in motion, demonstrating a dynamic movement sequence that involves jumping, climbing, and scaling a vertical surface in an urban environment.
[0127] The head-mounted display worn by the athlete incorporates an image sensor that captures first-person (egocentric) image data, providing a real-time view of the environment from theathlete’s perspective. This HMD may also project augmented reality overlays into the user’s field of vision, such as trajectory markers or virtual coaching instructions, to assist in optimizing movement execution.
[0128] In addition to the egocentric view, a drone positioned nearby captures second image data from an exocentric perspective, offering an external viewpoint of the athlete’s body positioning, spatial trajectory, and environmental interactions. The drone follows the athlete’s motion and records real-time data on posture, movement efficiency, and potential areas for improvement. The combination of these two perspectives allows the system to extract both contextual conditions of the environment and performance parameters of the user, enabling an advanced level of motion analysis.
[0129] By integrating data from both perspectives, a machine-learning model processes the information to determine corrective coaching instructions tailored to the athlete’s parkour performance. These instructions may then be visually overlaid onto the first-person view in the HMD, guiding the athlete toward improved technique, more efficient movement paths, and better spatial awareness. In this particular example, the system could suggest an optimized jump angle, wall contact timing, or landing posture, helping the athlete refine their execution in real-time.
[0130] This intelligent coaching system may enhance skill development by providing adaptive, data-driven feedback, allowing athletes to adjust their movements based on both self-perceived execution and Al-driven motion analysis. By leveraging augmented reality, drone-based tracking, and machine-learning-assisted coaching, this system represents a sophisticated method for improving athletic performance in highly dynamic and complex sports such as parkour.
[0131] Fig. 8 depicts a real-time augmented coaching system for motorsport racing, integrating a head-mounted display (HMD), a virtual coach car 800, and a drone-based external camera to assist the driver in optimizing their racing performance. The perspective is from inside a Formula-style race car, with the driver’s hands gripping the steering wheel, which features multiple buttons and a digital dashboard for vehicle control and telemetry, but the present disclosure is not limited in that regard.
[0132] The head-mounted display worn by the driver captures first image data, providing a real-time visual feed of the racetrack from the driver’s perspective. Through augmented reality overlays, the virtual coach car 800 is projected onto the track ahead, guiding the driver toward theoptimal racing line, braking zones, and acceleration points. The virtual coach car serves as a dynamic reference, enabling the driver to follow an ideal trajectory to enhance speed and efficiency on the circuit.
[0133] Additionally, a drone positioned above the track captures second image data from an exo-centric perspective, providing an external view of the driver’s racing behavior and vehicle positioning. This drone continuously tracks the driver’s movements and transmits real-time performance data to a trained machine-learning model, which analyzes contextual conditions and performance parameters such as braking precision, cornering angles, and acceleration patterns.
[0134] By integrating both egocentric and exocentric image data, the system evaluates the driver’s execution against an Al-trained racing model, generating corrective coaching instructions to improve performance. These instructions are delivered via AR overlays in the HMD, ensuring that the driver receives immediate, data-driven feedback without diverting focus from the track.
[0135] This setup represents an advanced racing assistance system, combining real-time Al analysis, drone-based performance tracking, and augmented reality coaching to help drivers refine their technique, minimize lap times, and enhance overall racing efficiency.
[0136] Fig. 9 depicts a real-time augmented reality coaching system for mountain biking, integrating a head-mounted display (HMD), a virtual coach cyclist, and a drone-based tracking system to enhance the cyclist’s performance. The perspective is from the viewpoint of a cyclist who is mid-action, maneuvering on a rugged outdoor trail.
[0137] The head-mounted display worn by the cyclist captures first image data, providing a real-time visual representation of the terrain and surrounding environment. Through augmented reality overlays, a virtual coach cyclist 900 is displayed ahead on the trail, offering an optimized riding path, posture guidance, and ideal braking and acceleration points. This virtual cyclist acts as a real-time reference, allowing the user to mimic efficient movement techniques and improve their riding skills.
[0138] In addition to the egocentric view, a drone positioned behind and above the cyclist captures second image data from an exocentric perspective, providing an external viewpoint of the cyclist’s body position, bike control, and movement through the environment. The drone continuously tracks the cyclist’s performance and transmits real-time biomechanical andkinematic data to a machine-learning model, which processes factors such as balance, speed, body posture, and turn execution.
[0139] By integrating both egocentric and exocentric data, the system can analyze contextual conditions and performance parameters to generate corrective coaching instructions. These instructions are visually overlaid onto the head-mounted display, allowing the cyclist to immediately adjust their technique while maintaining focus on the trail. The system ensures that the virtual coach cyclist remains dynamically positioned, adapting to the cyclist’s real-world speed and trajectory.
[0140] This advanced augmented reality coaching system combines real-time Al-based motion analysis, drone-assisted performance tracking, and AR guidance to help cyclists refine their riding techniques, improve handling on rough terrain, and develop more efficient movement strategies for biking.
[0141] Fig. 10 illustrates a real-time augmented reality coaching system for rock climbing, integrating a head-mounted display (HMD), a virtual coach climber 1000, and a drone-based external camera to assist the climber in navigating and optimizing their ascent. The climber is scaling a steep rocky surface with a climbing route, and the system provides real-time guidance through AR overlays and external tracking.
[0142] The head-mounted display worn by the climber captures first-person (egocentric) image data, allowing for real-time visualization of the climbing route, surrounding terrain, and possible handholds and footholds. The HMD may project augmented reality overlays, including the virtual coach climber 1000, which appears within the climber’s field of view, guiding them toward optimal movement strategies, grip selection, and body positioning. Alternatively, optimal grip positions may be highlighted in color.
[0143] A drone positioned above the climber captures second image data from an exocentric perspective, providing a wider view of the climber’s progress, posture, and interaction with the environment. This drone-based external tracking system allows for an objective evaluation of the climber’s technique, efficiency, and positioning within the climbing route. The real-time data collected by the drone is transmitted to a machine-learning model, which analyzes contextual conditions (such as route difficulty and external obstacles) and performance parameters (such as grip stability, center of gravity, and movement efficiency).By integrating both egocentric and exocentric data, the system generates corrective coaching instructions that are overlaid onto the head-mounted display, providing instant feedback on climbing technique. The virtual coach climber, dynamically positioned within the climber’s field of view, acts as a real-time guide, demonstrating optimal movement sequences, weight distribution strategies, and recommended grip choices to enhance climbing performance.
[0144] This system represents an advanced Al-powered climbing assistant, combining real-time motion tracking, drone-assisted evaluation, and AR-enhanced coaching to improve technique, increase safety, and optimize route efficiency in challenging climbing environments.
[0145] The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.
[0146] The following examples relate to further embodiments of the present disclosure:
[0147] (1) An apparatus for real-time coaching of a user during a sporting activity. The apparatus includes processing circuitry configured to receive, in real-time during the sporting activity, first image data from an egocentric view of the user. The processing circuitry is further configured to determine, from the first image data, contextual conditions for the sporting activity. The processing circuitry is further configured to receive, in real-time during the sporting activity, second image data of the user from an exocentric view. The processing circuitry is further configured to determine, from the second image data, one or more performance parameters of the user. The processing circuitry is further configured to determine, based on the contextual conditions and the one or more performance parameters of the user, a corrective coaching instruction for the user using a trained machine-learning model that is specific to the sporting activity. The processing circuitry is further configured to output the corrective coaching instruction to the user.
[0148] (2) The apparatus of (1), wherein the first image data are received from an image sensor of a head-mounted display worn by the user, and wherein the second image data are received from an image sensor of an unmanned aerial vehicle following the user.
[0149] (3) The apparatus of (1) or (2), wherein the processing circuitry is further configured to invoke a digital twin of an environment of the sporting activity and provide the contextualconditions and the one or more performance parameters to the digital twin for generating the corrective coaching instruction.
[0150] (4) The apparatus of (3), wherein the processing circuitry is further configured to adapt the environment in the digital twin based on the contextual conditions.
[0151] (5) The apparatus of (4), wherein the processing circuitry is further configured to simulate, based on the adapted environment in the digital twin and using the one or more performance parameters, the user doing the sporting activity in the digital twin. In such examples, the processing circuitry may be further configured to apply, for the simulated user doing the sporting activity in the adapted environment in the digital twin, the machine-learning model to determine the corrective coaching instruction.
[0152] (6) The apparatus of any one of (1) to (5), wherein the processing circuitry is further configured to overlay, on the first image data, an indication of an improved performance determined based on the trained performance model. In such examples, the processing circuitry is further configured to cause output of the overlaid first image data for the user.
[0153] (7) The apparatus of any one of (1) to (6), wherein the one or more performance parameters include a posture of the user and positioning of equipment used for the sporting activity.
[0154] (8) The apparatus of any one of (1) to (7), wherein the contextual conditions are first contextual conditions. In such examples, the processing circuitry may be further configured to determine, from the second image data, second contextual conditions of an environment of the sporting activity.
[0155] (9) The apparatus of (7) or (8), wherein the processing circuitry is further configured to warn the user of a potential hazard based on the second contextual conditions.
[0156] (10) The apparatus of any one of (1) to (9), wherein the corrective coaching instruction is at least one of a visual instruction, a haptic instruction, and an auditive instruction.
[0157] (11) A head-mounted display including at least one display area, at least one image sensor, an apparatus for real-time coaching according to any one of (1) to (10).
[0158] (12) A method for real-time coaching of a user during a sporting activity. The method includes receiving, in real-time during the sporting activity, first image data from an egocentricview of the user. The method further includes determining, from the first image data, contextual conditions for the sporting activity. The method further includes receiving, in real-time during the sporting activity, second image data of the user from an exocentric view. The method further includes determining, from the second image data, one or more performance parameters of the user. The method further includes determining, based on the contextual conditions and the one or more performance parameters of the user, a corrective coaching instruction for the user using a trained machine-learning model that is specific to the sporting activity, outputting the corrective coaching instruction to the user.
[0159] (13) The method of (12), wherein the first image data are generated with an image sensor of a head-mounted display worn by the user, and wherein the second image data are generated with an image sensor of a unmanned aerial vehicle following the user.
[0160] (14) The method of (12) or (13), further including invoking a digital twin of an environment as a machine-learning model of the sporting activity and provide the first image data and the second image data to the digital twin for generating the corrective coaching instruction.
[0161] (15) The method of (14), further including adapting the environment in the digital twin based on the contextual conditions.
[0162] (16) The method of (15), further including simulating, based on the adapted environment in the digital twin and using the one or more performance parameters, the user doing the sporting activity in the digital twin. In such examples, the method further includes applying, using the simulated user doing the sporting activity in the adapted environment in the digital twin, the machine-learning model to determine the corrective coaching instruction.
[0163] (17) The method of any one of (12) to (16), further including overlaying, on the first image data, an indication of an optimal performance determined based on the trained performance model. In such examples, the method further includes displaying the overlaid first image data for the user.
[0164] (18) The method of any one of (12) to (17), wherein the one or more performance parameters include a posture of the user and equipment positioning.
[0165] (19) The method of any one of (12) to (18), wherein the contextual conditions are first contextual conditions. In such examples, the method further includes determining, from the second image data, second contextual conditions of an environment of the sporting activity.(20) The method of (19), further including warning the user of a potential hazard based on the second contextual conditions.
[0166] (21) The method of any one of (12) to (20), wherein the corrective coaching instruction is at least one of a visual instruction, a haptic instruction, and a auditive instruction.
[0167] (22) A non-transitory computer readable medium including instructions which, when carried out on processing circuitry, cause the processing circuitry to carry out the method of any one of (12) to (21).
[0168] (23) A computer program comprising instructions which, when the program is executed on a computer, causes the computer to carry out the method of any one of claims (12) to (21).
[0169] Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and / or contain machine-executable, processorexecutable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
[0170] It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and / or be broken up into several sub-steps, -functions, -processes or -operations.If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
[0171] The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.
Claims
ClaimsWhat is claimed is:
1. An apparatus for real-time coaching of a user during a sporting activity, the apparatus comprising processing circuitry configured to:receive, in real-time during the sporting activity, first image data from an egocentric view of the user;determine, from the first image data, contextual conditions for the sporting activity;receive, in real-time during the sporting activity, second image data of the user from an exo-centric view;determine, from the second image data, one or more performance parameters of the user;determine, based on the contextual conditions and the one or more performance parameters of the user, a corrective coaching instruction for the user using a trained machine-learning model that is specific to the sporting activity; andoutput the corrective coaching instruction to the user.
2. The apparatus of claim 1, wherein the first image data are received from an image sensor of a head-mounted display worn by the user, and wherein the second image data are received from an image sensor of an unmanned aerial vehicle following the user.
3. The apparatus of claim 1 , wherein the processing circuitry is further configured to:invoke a digital twin of an environment of the sporting activity and provide the contextual conditions and the one or more performance parameters to the digital twin for generating the corrective coaching instruction.
4. The apparatus of claim 3, wherein the processing circuitry is further configured to:adapt the environment in the digital twin based on the contextual conditions.
5. The apparatus of claim 4, wherein the processing circuitry is further configured to:simulate, based on the adapted environment in the digital twin and using the one or more performance parameters, the user doing the sporting activity in the digital twin;apply, for the simulated user doing the sporting activity in the adapted environment in the digital twin, the machine-learning model to determine the corrective coaching instruction.
6. The apparatus of claim 1 , wherein the processing circuitry is further configured to:overlay, on the first image data, an indication of an improved performance determined based on the trained performance model; andcause output of the overlaid first image data for the user.
7. The apparatus of claim 1 , wherein the one or more performance parameters include a posture of the user and positioning of equipment used for the sporting activity.
8. The apparatus of claim 1, wherein the contextual conditions are first contextual conditions, wherein the processing circuitry is further configured to:determine, from the second image data, second contextual conditions of an environment of the sporting activity.
9. The apparatus of claim 7, wherein the processing circuitry is further configured to:warn the user of a potential hazard based on the second contextual conditions.
10. The apparatus of claim 1, wherein the corrective coaching instruction is at least one of a visual instruction, a haptic instruction, and an auditive instruction.
11. A head-mounted display comprising:at least one display area;at least one image sensor; andan apparatus for real-time coaching according to any one of claims 1 to 10.
12. A method for real-time coaching of a user during a sporting activity, the method comprising:receiving, in real-time during the sporting activity, first image data from an egocentric view of the user;determining, from the first image data, contextual conditions for the sporting activity;receiving, in real-time during the sporting activity, second image data of the user from an exocentric view;determining, from the second image data, one or more performance parameters of the user;determining, based on the contextual conditions and the one or more performance parameters of the user, a corrective coaching instruction for the user using a trained machine-learning model that is specific to the sporting activity; andoutputting the corrective coaching instruction to the user.
13. The method of claim 12, wherein the first image data are generated with an image sensor of a head-mounted display worn by the user, and wherein the second image data are generated with an image sensor of a unmanned aerial vehicle following the user.
14. The method of claim 12, further comprising:invoking a digital twin of an environment as a machine-learning model of the sporting activity and provide the first image data and the second image data to the digital twin for generating the corrective coaching instruction.
15. The method of claim 14, further comprising:adapting the environment in the digital twin based on the contextual conditions.
16. The method of claim 15, further comprising:simulating, based on the adapted environment in the digital twin and using the one or more performance parameters, the user doing the sporting activity in the digital twin;applying, using the simulated user doing the sporting activity in the adapted environment in the digital twin, the machine-learning model to determine the corrective coaching instruction.
17. The method of claim 12, further comprising:overlaying, on the first image data, an indication of an optimal performance determined based on the trained performance model; anddisplaying the overlaid first image data for the user.
18. The method of claim 12, wherein the one or more performance parameters include a posture of the userand equipment positioning.
19. The method of claim 12, wherein the contextual conditions are first contextual conditions, and wherein the method further comprises:determining, from the second image data, second contextual conditions of an environment of the sporting activity.
20. The method of claim 19, further comprising:warning the user of a potential hazard based on the second contextual conditions.