Audio communications and analysis

A wearable device with beamforming microphones and AI processing enhances athlete communication and audio data analysis, addressing the need for device-free ears and improving broadcasting utilization.

WO2026003721A1PCT designated stage Publication Date: 2026-01-02VOKALO APS
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/056409
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-05
Filing Date
2025-06-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Athletes participating in sports activities desire to communicate with each other without wearing devices on their ears, and existing audio systems from participants are not effectively utilized in live broadcasting and post-production.

Method used

A wearable device with beamforming microphones captures audio data, processes it using AI models to remove undesired sounds, and establishes communication links with other devices, enabling real-time analysis and visualization of audio interactions.

Benefits of technology

Enables effective communication between athletes, facilitates real-time strategic coaching, and enhances audio data utilization in broadcasting and post-production, while keeping athletes' ears free during activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025056409_02012026_PF_FP_ABST
    Figure IB2025056409_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are system, method, and computer program product embodiments, and / or combinations and sub-combinations thereof, for audio communications and analysis. Embodiments comprise providing a device configured to be mounted in a wearable element. The wearable element is configured to be worn on a body of a user. The device comprises a plurality of microphones arranged in a beamforming configuration having at least a beamforming direction towards a mouth of the user when mounted in the wearable element. Audio data is captured via the plurality of microphones. The embodiments further comprise processing the audio data to remove undesired audio, establishing a communication link between the device and another device, and sending, via the communication link, the audio data from the device to the another device. The embodiments further comprise analyzing the audio data using an AI model.
Need to check novelty before this filing date? Find Prior Art

Description

AUDIO COMMUNICATIONS AND ANALYSISCROSS REFERENCE TO RELATED PATENT APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 663,331, entitled “AUDIO COMMUNICATIONS AND ANALYSIS,” filed on June 24, 2024 and the benefit of U.S. Provisional Patent Application No. 63 / 800,007, entitled “AUDIO COMMUNICATIONS AND ANALYSIS,” filed on May 5, 2025. The entire contents of the above referenced applications are incorporated by reference herein in their entireties.FIELD

[0002] The present disclosure is generally directed to capturing, sharing, and analyzing audio communications. In particular, the present disclosure relates to audio communications between participants in an activity.BACKGROUND

[0003] Participants in a sport activity (e.g., players, athletes, trainees, coaches) may wish to communicate with each other. A coach may desire to communicate to the players instructions and players may desire to communicate with other players on their team. However, athletes desire to have their ears free from devices when participating in the sport activity. In addition, audio included from sports production comes from commentators and surroundings, the audio from the participants (e.g., players, coaches and referees) is rarely available to be used in live broadcasting and post production.SUMMARY

[0004] Provided herein are system, apparatus, article of manufacture, method and / or computer program product embodiments, and / or combinations and sub-combinations thereof, for audio communications and analysis.

[0005] According to some embodiments, a method for audio communications is provided. In the method, a device is mounted in a wearable element. The wearableelement is configured to be worn on a body of a user and wherein the device comprises a plurality of microphones arranged in a beamforming configuration having a beamforming direction towards a mouth of the user when mounted in the wearable element. Audio data is captured via the plurality of microphones. At least one computer processor processes the audio data to remove undesired audio. A communication connection between the device and another device. The audio data is sent via the communication connection from the device to the other device.

[0006] According to some embodiments, an audio comprises a housing, a plurality of microphones, and processing circuitry. The housing is configured to be attached to a wearable element worn on a body of a user. The plurality of microphones are configured to capture audio data. The plurality of microphones are arranged in a beamforming configuration having a beamforming direction towards a mouth of the user when worn by the user. The processing circuitry is configured to capture audio data via the plurality of microphones, process, using at least one artificial intelligence (Al) model, the audio data to remove undesired audio, establish a communication link between the audio device and another audio device, and send via the communication link connection the audio data from the audio device to the another audio device.

[0007] According to some embodiments, a method comprises acquiring audio data from a plurality of audio devices, analyzing, by at least one computer processor using an artificial intelligence (Al) model, the audio data to generate communication data. The communication data comprises at least a number of interactions of a user of an audio device of the plurality of audio devices and / or one or more types of communications. The at least one computer processor generates a graphical representation of the communication data.

[0008] According to some embodiments, a method comprises acquiring a live audio stream from one or more audio devices associated with one or more participants in a sport activity, receiving a verbal command to tag an event in the live audio stream, synchronizing, using at least one computer processor, audio data from the live audio stream corresponding to the tagged event with stored participant data, and generating a visual representation based on the audio data and the tagged event, wherein the visual representation comprises video data corresponding to the audio data.

[0009] According to some embodiments, a method comprises acquiring audio data from one or more audio devices, acquiring video data for a period corresponding to the audiodata, synchronizing, using at least one computer processor, the audio data with the video data, identifying an event in the audio data or the video data based on a user input or using an artificial intelligence (Al) model, retrieving the audio data corresponding to the event using the synchronized audio data with the video data, and outputting the video data corresponding to the event and the retrieved audio data.

[0010] According to some embodiments, a method comprises acquiring audio data comprising a plurality of audio signals captured using a plurality of devices, identifying, using at least one computer processor, an audio signal having a highest decibel level among the plurality of audio signals, identifying an audio device from the plurality of devices corresponding to a source of the audio signal having the highest decibel level, determining a signal strength of a message from the audio device by identifying identical sound characteristics in respective audio signals of the plurality of devices, and determining a score corresponding to a likelihood of a message reception for each device based on a respective signal strength.

[0011] According to some embodiments, a method comprises acquiring audio data from a plurality of devices associated with a plurality of players, determining, using at least one computer processor, a communication metric for each player of the plurality of players based on the audio data, and generating a coaching instruction for one or more players of the plurality of players based on a respective communication metric.

[0012] According to some embodiments, a method comprises acquiring audio data from a device. The device is associated with a participant in a sport activity and is worn by the participant during the sport activity. The method also comprises detecting, using at least one computer processor, an impact event based on at least the audio data, determining one or more biometrics of the participant based on the audio data, and determining an intensity of the impact event based on the one or more biometrics.

[0013] Further features of the present disclosure, as well as the structure and operation of various embodiments, are described in detail below with reference to the accompanying drawings. It is noted that the present disclosure is not limited to the specific embodiments described herein. Such embodiments are presented herein for illustrative purposes only. Additional embodiments will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein.BRIEF DESCRIPTION OF DRAWINGS

[0014] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and to enable a person skilled in the relevant art(s) to make and use embodiments described herein.

[0015] FIG. l is a block diagram of an environment for a system for communication between users of a platform, in accordance with an embodiment of the present disclosure.

[0016] FIG. 2 is a block diagram of a device, in accordance with an embodiment of the present disclosure.

[0017] FIG. 3 is a block diagram of an artificial intelligence (Al) module, in accordance with an embodiment of the present disclosure.

[0018] FIG. 4 is a diagram that shows a processing flow for a system, in accordance with an embodiment of the present disclosure.

[0019] FIG. 5A is a schematic that shows a broadside beamforming configuration, in accordance with an embodiment of the present disclosure.

[0020] FIG. 5B is a schematic that shows an endfire beamforming configuration, in accordance with an embodiment of the present disclosure.

[0021] FIG. 5C is a schematic that shows a triangular configuration, in accordance with an embodiment of the present disclosure.

[0022] FIG. 6 is a schematic that shows a wearable element, in accordance with an embodiment of the present disclosure.

[0023] FIG. 7 is a flowchart for a method for establishing a communication connection between devices, in accordance with an embodiment of the present disclosure.

[0024] FIG. 8 is a flowchart for a method for generating a graphical representation of communication data, in accordance with an embodiment of the present disclosure.

[0025] FIG. 9 is a flowchart for a method for generating a visual representation of a tagged event, in accordance with an embodiment of the present disclosure.

[0026] FIG. 10 is a flowchart for a method for outputting video data, in accordance with an embodiment of the present disclosure.

[0027] FIG. 11 is a flowchart for a method for determining a likelihood of a message reception, in accordance with an embodiment of the present disclosure.

[0028] FIG. 12 is a flowchart for a method for generating a coaching instruction, in accordance with an embodiment of the present disclosure.

[0029] FIG. 13 is a flowchart for a method for determining an intensity of an impact during a sport activity, in accordance with an embodiment of the present disclosure.

[0030] FIG. 14A illustrates a user interface for displaying session data, in accordance with an embodiment of the present disclosure.

[0031] FIG. 14B illustrates a user interface for displaying communication data associated with a session, in accordance with an embodiment of the present disclosure.

[0032] FIG. 14C illustrates a user interface for displaying communication data associated with a session, in accordance with an embodiment of the present disclosure.

[0033] FIG. 14D illustrates a user interface for displaying communication data associated with a player, in accordance with an embodiment of the present disclosure.

[0034] FIG. 14E illustrates a user interface for displaying synchronized video data with audio data, in accordance with an embodiment of the present disclosure.

[0035] FIG. 15 shows a computer system for implementing various embodiments of this disclosure.

[0036] FIG. 16 is a flowchart for a method for tagging an event using a double tap, in accordance with an embodiment of the present disclosure.

[0037] The features of the present disclosure will become more apparent from the detailed description set forth below when takin in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements. Additionally, generally, the left-most digit(s) of a reference number identifies the drawings in which the reference number first appears. Unless otherwise indicated, the drawings provided throughout the disclosure should not be interpreted as to-scale drawings.DETAILED DESCRIPTION

[0038] Aspects of the present disclosure relate to a system, a device, and associated methodologies for audio communications. In particular, the present disclosure relates to capturing and analyzing audio data during a sport activity.

[0039] This specification discloses one or more embodiments that incorporate the features of the present disclosure. The disclosed embodiment(s) are provided as examples. The scope of the present disclosure is not limited to the disclosed embodiment(s). Claimed features are defined by the claims appended hereto.

[0040] The embodiment(s) described, and references in the specification to “one embodiment,” “an embodiment,” “an example embodiment,” etc., indicate that the embodiment(s) described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is understood that it is within the knowledge of one skilled in the art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0041] Spatially relative terms, such as “beneath,” “below,” “lower,” “above,” “on,” “upper” and the like, may be used herein for ease of description to describe one element or feature’s relationship to another element(s) or feature(s) as illustrated in the figures. The spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The apparatus may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein may likewise be interpreted accordingly.

[0042] The term “about,” “approximately,” or the like may be used herein to indicate a value of a quantity that may vary or be found to be within a range of values, based on a particular technology. Based on the particular technology, the terms may indicate a value of a given quantity that is within, for example, 1-20% of the value (e.g., ±1%, ±5% ±10%, ±15%, or ±20% of the value).

[0043] Embodiments of the disclosure may be implemented in hardware, firmware, software, or any combination thereof. Embodiments of the disclosure may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flash memory devices; electrical, optical, acoustical or other forms ofpropagated signals (e.g., carrier waves, infrared signals, digital signals, etc.), and others. Further, firmware, software, routines, and / or instructions may be described herein as performing certain actions. However, it should be appreciated that such descriptions are merely for convenience and that such actions in fact result from computing devices, processors, controllers, or other devices executing the firmware, software, routines, instructions, etc. In the context of computer storage media, the term “non-transitory” may be used herein to describe all forms of computer readable media, with the sole exception being a transitory, propagating signal.

[0044] As noted in the Background section above, athletes want to have their ears free from devices during play and training. Embodiments described herein include a device that may be worn on the body of an athlete. The body -worn device allows the athlete to have their ears free without compromising on audio quality. In addition, audio from the body-worn device may be used by live broadcasting feeds and / or in post programs. Furthermore, the more participants in the sport activity, the more audio should be listened to in order to find desired audio (e.g., great moments in the game such as interactions between players, reaction to scoring a goal). This is a time consuming task. The platform described herein analyzes the audio to identify interesting moments (e.g., desired audio) for broadcasting. The platform described herein may also filter out sensitive communication.

[0045] The system and associated platforms and methodologies described herein enhance communication between participants in an activity by providing models to analyze audio data. In addition, the system provides a user interface that incorporates visualizations, spatial player arrangements, and a color-coding scheme for the data. The system provides a multifaceted tool for enhancing communication and strategic decision-making in realtime or near real-time scenarios. In addition, the system and associated platforms and methodologies may analyze audio data from multiple audio sources (e.g., from devices associated with multiple participants) and may detect desired audio interactions.

[0046] The system and associated platforms and methodologies described herein may be used to analyze content, type, and tone of audio communications between participants of a sport activity (e.g., between players, between a player and a coach, or emotional statement by a player). The analysis may be used to provide real-time strategic coaching communications from the coach to one or more players, support the coach in making time-out decisions during the match, analyze the performance of the coach and non-players participants, and build a respective data based profile for each player. The analysis may be performed in real-time (e.g., the computer system reacts to events by performing tasks within a specific time interval such a threshold number of milliseconds) or offline.System Overview and Function

[0047] FIG. 1 is a block diagram of an environment 100 for a platform 102, in accordance with an embodiment of the present disclosure. Environment 100 may include platform 102, a device 104, a docking station 106, a network 108, an electronic device 110, a wearable element 112, a database 116, and a third party system 132.

[0048] Platform 102 may provide a cluster computing platform or a cloud computing platform to manage audio communications during an activity (e.g., sport activity). Platform 102 may acquire audio data from device 104 and / or docking station 106. Platform 102 may process the audio data. For example, platform 102 may perform or trigger a tagging process, a video synchronization process, generate one or more overlays based on the analysis of the audio data, or determine one or more biometrics of a participant based on the audio data. Platform 102 may operate on one or more servers and / or databases. The servers may be a variety of centralized or decentralized computing devices. For example, a server may be grid-computing resources, a virtualized computing resource, peer-to-peer distributed computing devices, a mobile device, a laptop computer, a desktop computer, or a combination thereof. The servers may be centralized in a single room, distributed across different rooms, distributed across different geographic locations, or embedded within network 108. In some aspects, platform 102 may be implemented using computer system 1500 described with reference to FIG. 15. In some aspects, as described further below, functions performed by platform 102 may be performed additionally or alternatively by docking station 106 or device 104. In some aspects, platform 102 may be used in a real time mode (e.g., during a sport activity) or in a non- real time mode (e.g., uploading and analyzing data associated with the sport activity after the sport activity has ended).

[0049] In some aspects, docking station 106 may hold a plurality of devices 104 and one or more electronic devices (e.g., a tablet). In some aspects, a number of devices that the docking station may hold may correspond to the number of players in a team for a specific sport activity (e.g., 11 players in a soccer game). In one example, docking station106 may have a capacity of 24 devices and one electronic device 110. In some aspects, a docking station may be associated with a single device 104. In some aspects, docking station 106 may include processing circuitry 126, a memory 128, and wireless charging circuitry 130.

[0050] In some aspects, docking station 106 may be configured as an edge device to facilitate device / player assignment pre-session and perform processing of the audio data. Docking station 106 may upload all data post-session to platform 102. Thus, the methodologies described herein may be used “offline” (e.g., without requiring a connection to platform 102) for session preparation and analysis. A session may refer to a match (or a game) or a training session. As described further below, during the session one or more devices 104 may be activated and worn by the players. Docking station 106 may be used to pair each device 104 with a player profile for each player. For example, using electronic device 110, a unique player identifier may be assigned to device 104. Device 104 may be also reassigned to a different player profile during the session or post session.

[0051] In some aspects, docking station 106 may be configured to provide power to device 104 and to electronic device 110 via charging circuitry 130 (e.g., wireless (transmitter coil) or wired (cable)) before the start of the session. In some aspects, docking station 106 may be used to simultaneously charge one or more devices 104 and electronic device 110.

[0052] In addition to providing power to device 104, docking station 106 may acquire and process audio data from device 104. For example, docking station 106 may acquire audio data from each active device in the session. In some aspects, docking station 106 may acquire audio data from device 104 via network 108. In addition to the audio data, docking station 106 may acquire session data from device 104. Session data may include data from sensors included in device 104. In order to process the audio data, docking station 106 may use processing circuitry 126 to run one or more artificial intelligence (Al) algorithms (e.g., machine learning models) or other algorithms (e.g., digital signal processing algorithms) to filter undesired audio (e.g., audio noise). Thereby ensuring that data analysis can be conducted locally even in the absence of a connection to platform 102.

[0053] Docking station 106 may upload the data (e.g., processed and unprocessed data) to platform 102 via network 108. In some aspects, docking station 106 may upload the datato platform 102 in real-time or near real-time. In some aspects, docking station 106 may upload the data at predefined intervals (e.g., at half time of the match) and / or when the session ends. The data acquired from device 104 may be stored in memory 128. After uploading the data to platform 102, the data may be purged from memory 128. In some aspects, a single cloud link (e.g., connection to platform 102) is used to upload all session data (e.g., audio data from all the devices) as opposed to maintaining multiple active cloud links from connected devices (e.g., up to 24 devices in a docking station). Using docking station 106 can ensure that session data is being trimmed and converted on a central device, reducing data package sizes significantly before transmission to platform 102 through wireless connections such as WiFi or cellular network.

[0054] A user may interact with platform 102 or docking station 106 using electronic device 110. In some aspects, electronic device 110 may be a tablet, a computing device, a laptop, or the like. Electronic device 110 may comprise a user interface such as a display. Using electronic device 110 and before the start of the session, the user may input the number of participants, may perform the participant device pairing, and may input any desired session settings as described further below. In addition, the user may indicate the start and the end of the session using electronic device 110.

[0055] After pairing the participant profile with device 104, device 104 may be worn by the participant to start the session (e.g., a game or a training session). In some aspects, the device is optimally positioned for audio input near the mouth. As discussed above, device 104 is configured to be positioned on a body of the participant using wearable element 112. In some aspects, wearable element 112 may be configured to hold device 104 near a mouth of the participant. For example, wearable element 112 may be a vest that comprises a slot or a pocket configured to hold device 104. A fabric material of wearable element 112 may be selected from a plurality of fabric materials based on the type of the activity or the preference of the user. In some aspects, device 104 may be positioned on a chest muscle of the participant to absorb external impacts. The placement of device 104 on or near a chest of a user (e.g., a player) provides the advantage of overcoming the challenges of being as close to the mouth as possible while being placed on a soft spot on the body to increase comfort and reduce external impact compared to being placed on a bone.

[0056] In some aspects, wearable element 112 may include a noise absorbing material. The placement of the noise absorbing material on wearable element 112 may be selectedsuch as to reduce noise from environmental conditions (e.g., wind) and friction noise due to fabrics rubbing against each other. The noise absorbing material may include fur or foam. Wearable element 112 is further described in relation to FIG. 6. In some aspects, the noise absorbing material may be attached to an outer surface of a housing of device 104.

[0057] As discussed above, device 104 may be used to capture audio data from each participant during the session. In some aspects, device 104 may be waterproof or sweatproof. For example, device 104 may include a waterproof or a water resistant micro- electro-mechanical system (MEMS) microphone. In some aspects, the device is small and is lightweight such as not to inconvenience the player. In some aspects, device 104 may include a rectangular housing (e.g., each linear dimension less than 100 mm) that is shock-proof such as device 104 may withstand accidental drop during play or training.

[0058] FIG. 2 is a block diagram of device 104, according to some aspects. Device 104 may comprise processing circuitry 202, a microphone 204, a speaker 206, a display 208, communication circuitry 210, localization circuitry 212, an image sensor 214, a motion sensor 216, charging circuitry 218, and audio input box 224.

[0059] In some aspects, microphone 204 is used to capture audio signals when a user (e.g., a player or a coach) communicates. Microphone 204 may be one or more microphones or one or more microphone arrays. In some aspects, microphone 204 may be a MEMS microphone. The one or more microphones may be positioned strategically in device 104 to achieve beamforming and remove mechanical noise. The one or more microphones may be arranged in a beamforming configuration having a beamforming direction toward the mouth of the participant. The beamforming configurations are further described in relation to FIGS. 5 A and 5B. In some aspects, the one or more microphones may be arranged in a configurations that supports beamforming in multiple directions.

[0060] In some aspects, display 208 may be an electronic paper display (e.g., E-Ink screen) that provides the advantage of having low energy consumption and easy visibility for the user. Display 208 may be used to display information to the user during setup or during the session. For example, device 104 may output via display 208 an indication of low battery level or a status of network connection. In addition, device 104 may output via display 208 identification data associated with the participant and the device. For example, device 104 may output via display 208 a player name, a profile identification, and / or a device identification.

[0061] Speaker 206 may be one or more speakers. In some aspects, speaker 206 is positioned in an upper portion of device 104 to be near the ear of the user. Speaker 206 may output audio data received from another player or the coach. A volume level of speaker 206 may be adjusted via a voice command received from the user. As described further below, the audio received from the coach may be from a participant in the sport activity (e.g., a human coach) or generated using one or more Al modules (Al coach).

[0062] Processing circuitry 202 may perform local analysis of the captured audio using one or more Al model and audio filtering. Processing circuitry 202 may include an Al module 220 and an audio filter module 222.

[0063] In some aspects, audio filter module 222 may implement one or more audio algorithms. The one or more audio algorithms may be used to filter undesired audio signals (e.g., undesired source). The one or more audio algorithms include, but are not limited to, band-pass filters, multiband compressors, noise reduction (e.g., audio noise) or filtering using digital signal processing algorithms, and beamforming. In some aspects, audio filter module 222 may process the audio data before uploading the audio data to platform 102 or outputting to docking station 106.

[0064] In some aspects, Al module 220 may perform audio filtering based on the audio characteristics of the participants. For example, Al module 220 may be trained to detect audio characteristics associated with the paired participant profile (e.g., deep learning model for voice recognition). Al module 220 may filter the noise based on the audio characteristics. Al module 220 may remove any audio that does not match the audio characteristic of the paired participant. In some aspects, Al module 220 may perform tasks similar to Al module 120.

[0065] In some aspects, communication circuitry 210 may enable device 104 to interact with other devices (e.g., platform 102) via network 108. For example, communication circuitry 210 may include a WiFi module to support network protocols such as a WiFi mesh.

[0066] In addition to capturing audio data, device 104 may capture video data using image sensor 214. In some aspects, image sensor 214 may include one or more cameras. For example, image sensor 214 may include a dual camera for stereo vision. Video data from image sensor 214 may be used for extended reality (xR) (e.g., augmented reality (AR), virtual reality (VR), and / or mixed reality (MR)). In some aspects, video data captured by image sensor 214 may be uploaded to platform 102. Platform 102 may usethe video data to synchronize the audio data with other data captured using other devices (e.g., audio data captured by devices associated with the other players in a team).

[0067] In addition, device 104 may capture motion data of the participant. In some aspects, motion sensor 216 may be used to determine a movement and an orientation of the player. In some aspects, motion sensor 216 may be an inertial measurement unit (IMU). The movement and the orientation of the user may be used by processing circuitry 202 when processing the audio data (e.g., to generate a visualization of the player movement and the player direction). Additionally, motion sensor 216 can be used to determine if device 104 is correctly positioned in wearable element 112 and provides auditory feedback if device 104 is positioned incorrectly. For example, based on data from motion sensor 216, processing circuitry 202 may determine whether device 104 is positioned correctly and generate the auditory feedback if the device is positioned incorrectly. In some aspects, the position of the device is determined. Then, the position may be compared with a device position configuration. In some aspects, an alert (e.g., the auditory feedback) may be output in response to determining that the position of the device fails to satisfy the device position configuration. The auditory feedback may include instructions to reposition the device or to rotate the device to the correct position. The auditory feedback can be generated based on a comparison between the data from motion sensor 216 and stored data indicative of the correct position of device 104.

[0068] In some aspects, localization circuitry 212 may determine a position of device 104. Localization circuitry 212 may include a global position system (GPS). Position data may be combined with the audio data and orientation data from motion sensor 216 to generate one or more communicative overlays as described further below. By acquiring position data from each device 104, platform 102 may determine the location of each device at a given coordinated universal time (UTC) time stamp. Platform 102 may combine the position data with the motion data (e.g., from each IMU) to determine the position and direction for all active devices.

[0069] Charging circuitry may be coupled to a battery of device 104. For wireless charging, charging circuitry 218 may comprise a receiver coil coupled to a battery of device 104. The receiver coil may detect a magnetic field from charging circuitry 130 of docking station 106. In some aspects, charging circuitry 218 of device 104 and charging circuity 130 of docking station 106 operates at the same resonant frequency such as multiple devices 104 may be charged at the same time.

[0070] Audio input box 224 may be used to connect device 104 to a head set (microphone and headset) of participants (e.g., players, coaches). Audio captured by the connected by the head set may be uploaded to platform 102 for analysis.

[0071] In some aspects, one or more modules described herein may be implemented on a gaming computer that may connect to platform 102 via network 108 or via docking station 106. A headset may be connected to the gaming computer to capture audio from an esport participant. The captured audio may be analyzed by platform 102 to generate various metrics and / or emotional states of the participant.

[0072] As discussed above, device 104 may capture audio data from the user. In some aspects, the audio data may include a voice command. The voice command may comprise a command to control a volume of speaker 206, request or initiate a call with another user or to an Al module (e.g., Al coach), or create a tag in the audio data. Using voice commands enables the participant to use the device hands-free and does not affect the user play. In addition, by using voice commands coaches can flexibly control communication lines to coaches and players. In some aspects, the voice command may be processed by Al module 220 using a speech recognition algorithm (e.g., a hidden Markov model, a deep neural network model).

[0073] In some aspects, by saying a certain “word + what to tag” players and coaches can tag events of the game while in play. The tags and associated audio data may be stored in database 116. Using platform 102, the players and the coaches may search for a verbal tag while reviewing the game or the training session. Platform 102 may also associate video data with the tag. In some aspects, if the tag include a name of a player, the player may receive the corresponding audio data. This approach streamlines the sports analysis workflow, allowing coaches and players to quickly tag, find, and access relevant video and audio content for review and improvement.

[0074] In some aspects, device 104 may automatically adjust a volume level by using information about movement and sound environment. Furthermore, device 104 may adjust the volume based on the distance between the devices. The distance between the devices may be received from platform 102. In some aspects, the volume level may be decreased when the devices are close together (e.g., during corner kicks, game stops).

[0075] Referring back to FIG. 1, platform 102 may include a live module 114, an audio module 118, Al module 120, a synchronization module 122, and a health module 124. Live module 114 may be implemented as a communication tool for users to interact witheach other during the session. For example, coaches may interact with athletes during a game. A user may interact with live module 114 using a user interface (e.g., electronic device 110). The user may select one or more users to communicate with. For example, a coach may select a player, a group of players, or an entire team to communicate with.

[0076] In some aspects, live module 114 may comprise a plurality of mode of operations. The plurality of modes may include a first mode “listen mode,” a second mode “speak mode,” and a third mode “two-way communication mode.” In the first mode “listen mode,” the user may listen to the selected users (e.g., players). However, the select users may not hear the user (e.g., coach). A one-way communication connection (e.g., audio communication) can be established between the device associated with the user and respective devices of the selected users. In the second mode “speak mode,” the user can communicate with the other users (e.g., other active devices in the session). However, the other users may not respond. For example, the coach may be able to speak with the players (e.g., selected players). However, the coach may not hear the players. A one-way communication connection can be established between the device associated with the user and all active devices in the session. In the third mode “dialogue mode,” two-way communication between the users is enabled. For example, the coach may speak to the players and may hear the players. A two-way communication connection can be established between a device associated with the coach and a respective device of each player of the team. Each device may receive or send audio data to another device.

[0077] In some aspects, platform 102 may translate a communication to a preferred language of the receiver. For example, platform 102 may translate the communication from a language spoken by the coach to the preferred language of the player. In some aspects, the preferred language may be stored in the user profile and loaded to device 104 at the start of the session. Once a communication is received from platform 102 or from another device 104, device 104 can translate the communication to the preferred language. For example, Al module 220 may comprise a convolutional neural network trained to translate the communication to the preferred language. In some aspects, the translation may be performed by platform 102. For example, Al module 120 may perform the translation prior to transmitting the audio to device 104. This can reduce the amount of communication that is not understood by players in international performance environments.

[0078] In some aspects, a user may be able to send a pre-recorded audio message to a user or a group of users during the session. A plurality of pre-recorded audio messages may be stored in a database 116. Then, the user may select the pre-recorded audio message for output on a specific device or all active devices in the session.

[0079] In some aspects, platform 102 may provide pre-recorded coaching instructions to the players. Using a user interface provided by platform 102, a coach may pre-record coaching instructions before the start of the session. The pre-recorded coaching instructions may be stored in database 116. During the session, the coach may select a pre-recorded coaching instruction and the desired audience (e.g., a player or a group of players). The pre-recorded coaching instruction may be output to the desired audience via speaker 206 of the respective device. The pre-recorded coaching instruction may be output using the preferred language of the player.

[0080] In some aspects, platform 102 may generate coaching instructions to the players using Al module 120. For example, platform 102 may generate the coaching instructions based on an emotional state of the user, a fatigue level, and other metrics including third party data associated with the player.

[0081] In some aspects, coaches can use device 104 as a headset connected via network 108 (e.g., Bluetooth, WiFi) to platform 102 and / or electronic device 110. This enables a scenario where the coaches have their “head free” but still have the flexibility to communicate with the relevant staff or players.

[0082] As discussed above, platform 102 may operate on one or more cloud based servers. This approach provides users with the full flexibility to access platform 102 from anywhere in the world. Platform 102 may structure, store, and analyze the data. The data may include audio data, video data, session data, third-party data (e.g., performance data of a player), or other data as would be appreciated by a person of ordinary skill in the art. In some aspects, audio data may be analyzed both quantitatively and qualitatively in realtime and offline.

[0083] In addition to analyzing the audio data, platform 102 may process the audio data to improve the audio quality in both post-session processing and in real-time or near realtime (e.g., live during the session). Audio module 118 may implement one or more audio algorithms to improve the audio quality. In some aspects, audio module 118 may include a digital signal processing (DSP) filter to perform noise cancelling to eliminate sounds such as wind, ambient sound, and noise from the audience or to filter out undesired audio.The one or more audio algorithms include, but are not limited to, band-pass filters, multiband compressors, noise reduction, and beamforming. The one or more audio algorithms may filter out all communications that is not from a target audio source. The target audio source may be an audio signal from device 104. In addition, the filtered audio data may be used for audio transcription. The filtered audio data may be converted to text using a speech recognition algorithm (e.g., a deep neural network). The text may be used by Al module 120 for audio analysis as described further below.

[0084] In some aspects, the one or more audio algorithms applied to the audio data may be selected based on the activity. Different models may be trained to suppress different types and sources of noise during the session (e.g., with audience or without audience, expected decibel level during the sporting event). For example, platform 102 may apply a first algorithm when the sport is soccer and apply a second different algorithm when the sport is tennis. In addition, different algorithms may be applied during play versus during practice.

[0085] As discussed above, sports communication varies significantly among different sports, clubs, and players. Thus, a flexible and adaptive approach is desired. Platform 102 uses Al techniques (e.g., ML techniques) to address the variability in communication. In some aspects, Al module 120 may comprise one or more models. Each model may be trained for a different sport or setting (e.g., play vs. training). In some aspects, the model may be selected based on the user input at a start of the session. For example, the user may input via electronic device 110 the type of sport (e.g., tennis vs soccer) and whether the session is play or training. The one or more models of Al module 120 are further described in relation with FIG. 3. The one or more models are trained for different sports, through clubs and down to individual athletes and coaches. Combined with the clear audio input, the technology described herein is less error prone and more precise.

[0086] In some aspects, Al module 120 may quantify and analyze spoken words and voice characteristics, ultimately providing valuable insights into sports communication dynamics, emotional states, mental health, team collaboration, team spirit, and team connectivity for decision making. This dual approach ensures a comprehensive understanding of sports communication. Furthermore, time series analysis may be applied to analyze communication patterns over time, test for correlations, and predict future events. The categorization enhances the interpretability of the analyzed data, allowing for targeted insights.

[0087] In some aspects, third party system 132 may include one or more servers and / or one or more databases. Platform 102 may connect to third party system 132 to retrieve or acquire data via network 108. The data may include video data, position data (from a GPS), and historical data associated with the participants. The data may be used by Al module 120 to generate coaching instructions. In some aspects, platform 102 may use the data to generate visual representation and communication heatmaps as further described below.

[0088] In some aspects, platform 102, docking station 106, and / or device 104 may monitor a health of device 104. For example, platform 102 may monitor a battery lifetime and a health of microphones and speakers associated with device 104. Platform 102 may output a warning to device 104 and / or to electronic device 110 that device 104 may need to be replaced or re-calibrated. Monitoring the health of device 104 may include monitoring the quality of audio captured by the microphones (e.g., feedback loop, speaker loop), monitoring the health of battery (e.g., number of charge discharge cycles, recharge level and charge time), and a change in the precision or accuracy in one or more of the Al models (e.g., decrease in the performance of speech recognition model 302).

[0089] Network 108 refers to a telecommunications network, such as a wired or wireless network. Network 108 can span and represent a variety of networks and network topologies. For example, network 108 can include wireless communication, wired communication, optical communication, ultrasonic communication, or a combination thereof. For example, satellite communication, cellular communication, Bluetooth, Infrared Data Association standard (IrDA), wireless fidelity (WiFi), and worldwide interoperability for microwave access (WiMAX) are examples of wireless communication that may be included in network 108. Cable, Ethernet, digital subscriber line (DSL), fiber optic lines, fiber to the home (FTTH), and plain old telephone service (POTS) are examples of wired communication that may be included in the network 108. Further, network 108 can traverse a number of topologies and distances. For example, network 108 can include a direct connection, personal area network (PAN), local area network (LAN), metropolitan area network (MAN), wide area network (WAN), or a combination thereof.

[0090] FIG. 3 is a block diagram of Al module 120, in accordance with an embodiment of the present disclosure. Al module 120 may comprise one or more of a speechrecognition model 302, an audio classification model 304, a multimodal generative model 306, and a text classification model 308.

[0091] In some aspects, speech recognition model 302 may perform speech recognition. Speech recognition model 302 may receive data 310 from device 104. Data 310 may include audio data. The audio data may comprise audio data captured by device 104 or by another device and uploaded to platform 102. Speech recognition model 302 and text classification model 308 may be trained in the context of a specific activity such as soccer. Training data may include unique words and expressions that are used in the specific activity. Text classification model 308 for soccer may be trained on a dataset that includes “calls” used in on-field communication such as “carry”, “settle”, “shoot”, or “shot”. Further, the training data may include short sentences, bilingual expressions, and names and nicknames of the players of the team. The output of speech recognition model 302 may be fed to text classification model 308. Text classification model 308 may generate text embeddings for data 310.

[0092] In addition to performing word analysis, platform 102 may evaluate voice characteristics using audio classification model 304. Audio classification model 304 may receive data 310 as an input. Audio classification model 304 may perform audio spectrogram analysis to determine the voice characteristics of data 310. Audio classification model 304 may implement one or more signal processing techniques to extract features from audio including, but not limited to spectrograms, and other audio representations. The voice characteristics may include spectral features such as pitch, tone, frequency or rhythm, volume, and speech rate. The voice characteristics may vary based on speaker’s emotional state. For example, a participant may speak more quickly and with a higher pitch when excited or anxious. Audio classification model 304 may determine an emotional state of the participant based on the voice characteristics or a change in the voice characteristics. Audio classification model 304 may include one or more machine learning algorithms to interpret the audio features and voice characteristics and output an emotional state. In some aspects, audio classification model 304 may receive text input from text classification model 308 in addition to speech recognition model 302. The integration may be achieved using feature concatenation or multimodal fusion. Audio features extracted from speech recognition model 302 or the one or more machine learning algorithms of audio classification model 304 may be combined with text embeddings generated by text classification model 308. Thus, audio classificationmodel 304 may not only analyze audio signal but also incorporates semantic information from textual inputs to enable a more nuanced understanding of the content being analyzed. In addition, audio classification model 304 may analyze background sounds and sound features. By adding this layer to the communication analysis, platform 102 provides a comprehensive understanding of context, emotions, and overall sound situations.

[0093] In some aspects, audio classification model 304 may be trained on a large datasets of audio recordings that are labeled with emotional states. The audio classification model 304 may recognize patterns in the voice characteristics that correlate with the specific emotions. The emotions may include happiness, sadness, anger, fear, or neutrality.

[0094] In some aspects, multimodal generative model 306 may receive data 310 and third party data 312 as inputs. Third party data 312 may include video data, biometric data, tactical data, event data, and positional and movement data. By performing time series analysis and correlation analysis, multimodal generative model 306 may identify complex patterns and relationships across multiple modalities. Thus, multimodal generative model 306 may make accurate predictions and provide actionable insights based on the integrated audio and third-party data.

[0095] To provide structured analysis, platform 102 may categorize communications in the audio data into various types and states. The categorization may be used to analyze communication dynamics, emotional states, mental health, team collaboration, team spirit, and team connectivity. Multimodal generative model 306 may extract one or more parameters from the audio data. The parameters may include, but are not limited to, positive and negative emotions, stress levels, calmness, emotional stimulation, supportiveness, confidence, task orientation, volume, speech rate, speech clarity, language mix, tactical communication, self-talk communication, and strategic communication. Multimodal generative model 306 may include a trained neural network classifier.

[0096] As discussed above, currently video analyses of sports sessions (e.g., training and matches) have limited audio functionality. The audio is generally captured from a microphone associated with the camera capturing the video data. Consequently, coaches and analysts are unable to zoom in on specific players' audio during critical moments. Platform 102 may output audio from one or multiple recorded players post-session (e.g., audio recorded by device 104). Furthermore, a user (e.g., a coach) can tag key events inthe game, including audio for selected players. Tags can be shared with both coaches and players. Platform 102 may receive a user input from the user to tag a specific event in the audio file. The user input may be a voice command. In some aspects, the tagging of the events may be performed in real time during a session of training or play (e.g., during a match). The event may include comer kicks, goals, or good / bad communication interactions. In some aspects, tags are automatically synchronized with data analysis on platform 102. This can create a unique combination of event data and data from Al models all displayed together with video and audio.

[0097] In some aspects, Al module 120 may automatically detect key audio moments in matches, while also ensuring the privacy and discretion of participants by intelligently filtering out sensitive audio content. For example, Al module 120 may automatically mute or unmute audio from a player or a coach depending on the communication. In some aspects, Al module 120 may automatically detect and filter contents or communications associated with tactics. In some aspects, custom triggers may be defined to identify the key audio moments. For example, Al module 120 may be trained to detect a communication from a player such as “shoot” as a key audio moment. This both enables a fully automated decision process or a hybrid model allowing for human oversight and decision-making.

[0098] In addition, Al module 120 may be trained to identify individual audio characteristics in order to filter out everything except audio inputs from the source as described further below.

[0099] In addition to analyzing the audio data, platform 102 may synchronize audio data with video data. Synchronization module 122 may perform synchronization between the audio data and the video data. In some aspects, synchronization module 122 may detect similar sound characteristics in the audio data associated with the video data (e.g., captured by the microphone of the camera) and the audio data captured by device 104. An Al model may be trained to detect the similar sound characteristics.

[0100] In some aspects, synchronization module 122 may use timestamps such as UTC timestamps associated with the audio files and video files to perform the synchronization. In some aspects, the Al model may detect events in the session (e.g., scoring a goal) that can be synced with the audio data (e.g., passing, claps, talks, whistle). The synchronized audio data and video data may be manually reviewed by a user. Changes made by the user may be used to retrain the Al model.

[0101] In addition to synchronizing the audio data with video data, synchronization module 122 may synchronize audio data and information from device 104 with external environmental parameters using UTC timestamps. The external environmental parameters may include but are not limited to rain, wind, and humidity. This provides enhanced precision in the communication analysis.

[0102] In some aspects, synchronization module 122 may synchronize audio data acquired from a plurality of devices 104. In some aspects, synchronization module 122 may synchronize time across the plurality of devices 104. As discussed, each device 104 of the plurality of devices may be worn by a participant (e.g., a player, a coach, a staff). It is often desired to have precise time alignment among the plurality of devices 104 to enable coherent aggregation, analysis, and playback of audio data captured from multiple sources (e.g., plurality of devices 104, external sources such as video cameras).

[0103] In some aspects, synchronization module 122 may employ adaptive or hybrid synchronization modes. Synchronization module 122 may identify the preferred time synchronization technique based on environmental conditions, device capabilities, and / or network availability. In some aspects, synchronization module 122 may apply one or more of the following techniques to synchronize the audio data: GPS-based synchronization techniques, pre-session synchronization, time beaconing, precision oscillator clock, post-session calibration. In some aspects, synchronization module 122 may employ a combination of two or more techniques.

[0104] In some aspects, synchronization module 122 may employ a GPS based synchronization technique to synchronize audio data among the plurality of devices. As discussed above, each device 104 may include localization circuitry 212. The localization circuitry 212 may include a GPS receiver module. The GPS receiver module may receive time signals from GPS satellites. The GPS signal provides an accurate universal time reference, enabling synchronization of internal clocks across the plurality of devices 104 to a common epoch. Synchronization module 122 may send a command to each device to enable GPS based synchronization.

[0105] In some aspects, synchronization module 122 may enable a pre-session synchronization technique using docking station 106. Prior to the start of a session, docking station 106 may transmit a reference timestamp to each device 104. In response to receiving the reference timestamp each device 104 may align an internal clock of device 104 with the reference timestamp. This approach provides the advantage ofreducing clock drift and ensuring consistency of timestamps across devices during the session.

[0106] In some aspects, environment 100 may include an edge device may periodically broadcast a time beacon to each device of the plurality of devices. The time beacon includes a current reference timestamp, allowing each device to correct for drift and maintain synchronized time over prolonged sessions. In some aspects, electronic device 110, docking station 106, or platform 102 may transmit the time beacon.

[0107] In some aspects, device 104 may include a high-precision crystal oscillator configured to minimize clock drift. The oscillator ensures that once synchronization is achieved (via GPS, beacon, or hub), the internal timekeeping remains accurate throughout the session.

[0108] In some aspects, platform 102 may acquire audio data from the plurality of devices 104. Synchronization module 122 may perform timestamp alignment. In addition, synchronization module may refine the timestamp alignment by comparing known events (e.g., whistles, keyword utterances, or other manual markers) across audio data received from the plurality of devices 104. For example, synchronization module 122 may detect the whistle in a first audio file or a first audio stream. Then, synchronization module 122 may detect the same whistle in other audio files or audio streams. In one example, synchronization module 122 may detect the whistle at a first timestamp. Then, synchronization module 122 may analyze the other audio files to detect the whistle within a predetermined period. The predetermined period may correspond to the timestamp + / - a threshold (e.g., timestamp + / - 1 second). In some aspects, synchronization module 122 may apply the calibration algorithm retrospectively to align all audio files acquired from the plurality of devices by applying timestamp offsets. Synchronization module 122 may determine the timestamp offsets based on the post-session processing. For example, after detecting a known event in two audio files (acquired from two devices), synchronization module 122 may identify a first timestamp corresponding to the event in a first audio file and identify a second timestamp corresponding to the event in a second audio file. Then, synchronization module 122 may determine the timestamp offset as a difference between the first timestamp and the second timestamp). Synchronization module 122 may apply the offset to the second audio file or the first audio file.

[0109] In some embodiments, a combination of two or more techniques may be employed. In some aspects, real-time synchronization (e.g., via GPS or beaconing) andpost-session correction is employed to achieve highly accurate temporal alignment of audio data across all devices.

[0110] In some aspects, communication data for each player or team may be visualized thereby enabling users to navigate between data points while listening to audio or observing data as an overlay in the video. For example, users can view all communication sequences for a specific player in a given category within a video timeline. This provides the advantage of allowing users to seamlessly jump between communication data tags and hear the actual content of the communication. Furthermore, external analytic platforms may be integrated with platform 102 to improve workflow for coaches. Platform 102 may add different layers on top of the video data, showing real-time data for a selected player, communication coverage, and communicative links / connectivity.[OHl] Platform 102 may generate player data. Platform 102 may use one or more Al methods (e.g., ML methods) to quantify and categorize in-game communication on an individual and team level and visualize it intuitively (e.g., using multimodal generative model 306). Furthermore, platform 102 may integrate data from an external platform with the player data generated based on the captured audio data. The data acquired from the external performance may include data associated with the player past performance in the sport. The integrated data may be visualized together with the communication data.

[0112] The Al analysis results may be output using a user interface. The analysis data are dynamically presented through interactive charts, graphs, and numerical representations, providing a granular insight into communication patterns. This real-time visualization empowers stakeholders, including coaches and team managers, to make prompt and informed decisions based on the identified insights.

[0113] Platform 102 may provide the user with a user interface. The user interface may comprise a configurable dashboard where a user can organize and customize data presentation according to specific preferences of the user. For example, coaches have the flexibility to choose metrics, such as communication types, archetypes, connectivity, communication quantity, player performance, mood, psychological status, team dynamics, and engagement, creating a tailored view that aligns with their strategic focus. Exemplary graphic interfaces are further described in relation to FIGS. 14A-14E.

[0114] Furthermore, to streamline coaches' decision-making processes and enhance navigability, platform 102 may comprise a sports-specific user interface allowing for arrangement of players in sport-specific formations to assess communication patterns indifferent team setups and strategic formation (e.g., but not limited to, addressing the back line of defenders in soccer / football). Thus, platform 102 may provide the user with a plurality of user interfaces. Based on the selection of the user, the specific user interface may be output. In some aspects, the user may select the activity type at the start of a session. Based on the selected activity type, platform 102 may identify and output a specific user interface.

[0115] In some aspects, platform 102 may use different attributes to represent and differentiate scores associated with various communication metrics. The different attributes may comprise a color-coding scheme to represent and differentiate the scores associated with the various communication metrics. The color-coded visual representation enhances the speed and efficiency with which coaches interpret and make decisions based on the communicative performance data.

[0116] In some aspects, the visual representation may include a communicative heatmap for one or more devices. Data from motion sensor 216 and microphone 204 may be used to determine a direction of the user corresponding to device 104. In addition, platform 102 may determine decibel levels from audio data captured by microphone 204. In some aspects, a distance of communication travel may be determined based on the decibel measurements and the direction of the user.

[0117] Platform 102 may generate various overlays based on the data. The overlays may be incorporated into the visual representation. Thus, a communicative heatmap for a certain device is generated and displayed. In some aspects, the communicative heatmap may use data from all active devices as reference points. This helps determine if a message is being heard or not, providing a clear visual representation of communication effectiveness within the team.

[0118] In some aspects, Al module 120 may identify a position and a direction of the user at a given timestamp. The position and direction data may be combined with audio data to generate information that comprises an overview of the network’s position, direction, and communication over time. The overlays may also include position data and orientation data.

[0119] In addition to providing communications between users during play and training, platform 102 may provide automatic coaching instructions. The automatic coaching instructions may be generated by Al module 120. Al module 120 may analyze the player’s and team’s communication in real-time and automatically provide customizedcoaching for individuals or teams based on various communication metrics. In some aspects, the automatic coaching instructions may be based on other data such as previous statistics associated with a participant, emotional state, and / or fatigue level.

[0120] In some aspects, platform 102 may generate the coaching instructions based on one or more triggers. The one or more triggers may be set by the user at the start of the session. Once the trigger is detected, platform 102 may generate the real-time coaching instructions. The triggers may be a voice command or predefined words detected in the audio data associated with the coach or one or more players.

[0121] In some aspects, platform 102 can automatically provide customized coaching for individuals or teams based on desired metrics generated from data acquired from an external system. For example, the data from past performances or data from external sensors may be acquired and used to generate the customized coaching for a player.

[0122] In some aspects, platform 102 may generate the coaching instructions based on data acquired from sensors. The sensors may include optical cameras. Platform 102 generates the coaching instructions based on what can be seen in videos and includes instructions on the team's playing style, tactics, technical abilities, physics, and mental state.

[0123] In some aspects, health module 124 may detect impacts and concussion based on at least the audio data. Health module 124 may combine audio data with data from motion sensor 216 to identify strong impacts and tally headers in soccer. Additionally, by analyzing the player's voice, breathing, and impact data from motion sensor 216, platform 102 can determine an intensity of the impact. For example, platform 102 may determine a likelihood or a risk of a concussion.

[0124] In some aspects, health module 124 may determine one or more biometrics of the participant in the session. Health module 124 may determine a hear rate of the user by analyzing the audio signal captured by device 104. In some aspects, device 104 may include an additional microphone positioned at a location that corresponds the location of the heart of the user. Health module 124 may use the audio signal from the additional microphone to determine the heart rate of the user.

[0125] In some aspects, health module 124 may analyze a player breath. Health module 124 may analyze a frequency, an intensity, and a respiratory pattern to determine a health and physical condition of the player. In addition, health module 124 may determine a fatigue level of the player based on the one or more biometrics such as the heart rate andthe respiratory pattern. Health module 124 may monitor the one or more biometrics to determine a mental state of the participant. For example, the heart rate may increase with stress, anxiety, or excitement. In addition, a lower heart rate variability (HRV) may be associated with higher stress levels. Changes in the respiratory pattern may also be used by health module 124 to determine the emotional state of the participant. For example, rapid, shallow breathing may indicate anxiety or stress while slower and deeper breathing might indicate a relaxed state. Al module 120 may generate instructions for the participant to control breathing in order to influence the emotional state of the participant.

[0126] FIG. 4 is a diagram that shows a processing flow 400 for platform 102, in accordance with an embodiment of the present disclosure. Processing flow 400 may be initiated by a user 402. In 404, user 402 may use electronic device 110 to prepare device 104. To prepare device 104, user 402 may input a session type. For example, user 402 may input whether the session corresponds to a match or to a training session. In addition, user 402 may indicate the type of sport such as soccer or American football. Based on the inputs from user 402, platform 102 may determine a total number of users (e.g., players and a coach). In some aspects, platform 102 may automatically detect the identity of the user using voice recognition techniques. Platform 102 may then associate device 104 to the user and may load to device 104 a profile associated with the identified user. Platform 102 may track the users by associating each user with a unique identifier. In 406, each device 104 is mounted on wearable element 112 and worn by each user. In 408, the user may manage the session by activating one or more features of platform 102. In 410, device 104 may process the audio data and upload to docking station 106. In 412, docking station 106 may apply one or more audio algorithms to filter the noise from the audio data. The filtered audio data may be uploaded to platform 102. In 414, one or more Al models may be applied (e.g., for voice recognition, for speech recognition, for detecting key moments). In 416, the output of the Al models may be uploaded to platform 102.

[0127] In some aspects, players and / or coaches request data from the game in real time or in near real-time by activating live features. For example, live audio algorithms 418 may be applied to audio data captured by one or more devices (e.g., active devices during the session). Output data from the live audio algorithms may be used for live communication 420 and / or as an input of the live AI / ML models 422. The output of the live AI / ML models may be input to the live dashboard 424. For example, data generated by the liveAI / ML models 422 may be added as overlays and presented on live dashboard 424. Live dashboard 424 may also output data from health module 124.

[0128] Platform 102 may use output data from the live audio algorithms 418 and live AI / ML models 422 for live broadcasting 426. For example, platform 102 may synchronize audio data with video data. In addition, platform 102 may identify key moments in the video data and identify the corresponding audio data for live broadcasting 426. In addition, or alternatively, platform 102 may identify key moments in audio data and filter out sensitive data for live broadcasting. Platform 102 may output the audio in the form of subtitle in one or more languages based on a user preference. The subtitles may be output during live broadcasting or during post-match analysis using electronic device 110.

[0129] FIG. 5A is a schematic that shows a beamforming configuration, in accordance with an embodiment of the present disclosure. Device 104 may comprise a plurality of microphones (e.g., microphone 502a, ..., microphone 502n). The plurality of microphones may be arranged in a beamforming array to enable beamforming towards the user (e.g., athlete). The beamforming array may boost a sound signal coming from a desired direction while attenuating other audio signals. For example, noises from sources other than the desired source (e.g., the athlete being recorded) may be attenuated. The noise sources may include spectators and other athletes.

[0130] In some aspects, the plurality of microphones may be arranged in a broadside configuration and arranged in device 104 such as the audio signal from the user is perpendicular to respective microphone axis of the plurality of microphones 502a, . . ., 502n. In some aspects, the plurality of microphones may be arranged in an endfire configuration as shown in FIG. 5B. The plurality of microphones may be positioned in device 104 such as the audio signal from the user is coming from the direction in front of the first microphone. The plurality of microphones may be positioned in a single array pointed towards the source of the audio signal (e.g., mouth of the user). In some aspects, other beamforming configurations may be implemented. For example, the plurality of microphones may be arranged in a circular or spherical array. In some aspects, the beamforming configurations supports beamforming in multiple directions (e.g., toward the mouth of the user and toward the ground (e.g., to capture the sound of a kick of the ball)). Signal processing algorithms may be implemented by processing circuitry 202 to determine a direction from which the audio signal has originated. In some aspects,processing circuitry 202 may detect and determine the origin of a distinct sound such as a whistle.

[0131] An additional microphone 504 may be used to capture structural and acoustical noise. The audio data captured by the additional microphone may be used to filter the structural and acoustical noise from the desired audio stream captured by plurality of microphones 502a, . . ., 502n. The filtering may be implemented using a transfer function including but not limited to a band-pass filter, a multiband compressor, beamforming, and noise reduction, including a cross-correlation filter between air-coupled and mechanically coupled microphones.

[0132] In some aspects, the plurality of microphones may be arranged in a triangular configuration in device 104 as shown in FIG. 5C. Additional microphone 504 may be used to capture noise (e.g., environmental noise). A first microphone 502a and a second microphone 502b may be used to capture audio from the participant. In some aspects, either first microphone 502a or second microphone 502b may be used. In some aspects, the plurality of microphones may be positioned in the same plane (e.g., at a same height from a surface of the housing of device 104). In some aspects, the plurality of microphones may be positioned at different heights from the surface of the housing. Other configurations may also be used. For example, a third microphone may be positioned at a center of the triangular configuration.

[0133] During the session, platform 102 may switch between one or more microphones from the plurality of microphones. In some aspects, athletes may communicate both loudly and quietly. Thus, the plurality of microphones are configured to support a wide dynamic range. In some cases, the audio input from a first microphone 502a (the one closest to the athlete) can be clipped. When this happens, platform 102 may use audio signal from microphone 502n (e.g., farther from the athlete) to minimize clipping in the audio input. In some aspects, the switching may be done by processing circuitry of device 104.

[0134] In some aspects, device 104 may comprise a plurality of microphone arrays. Device 104 may comprise a primary microphone array, a secondary microphone array, and an auxiliary microphone array. The primary microphone array may be used to capture audio data from the user. The secondary microphone array may be activated when device 104 detects clipping in the audio data. The auxiliary microphone may be used to capture structural and acoustical noise. In some aspects, the primary microphone array and thesecondary microphone array may be positioned such us to have a beamforming direction in a direction of the user. Secondary microphone array may be positioned farther from the mouth of the user. In some aspects, the primary microphone array and the secondary microphone array may have an endfire configuration. The auxiliary microphone may have a broadside configuration.

[0135] FIG. 6 is a schematic that shows a wearable element 112, in accordance with an embodiment of the present disclosure. FIG. 6 shows wearable element 112 as a vest. The vest may be worn under or over the jersey. However, wearable element 112 may be a cross body belt, a jersey, a belt pouch, or an armband. In some aspects, wearable element 112 comprises a pocket 602 configured to hold device 104 in a diagonal direction towards a mouth of the participant when worn. In some aspects, the pocket may be made from a noise absorbing material. In some aspects, speaker 206 is positioned in an upper section of device 104. As discussed previously herein, microphone 204 may have a beamforming direction 600 towards the face of the player to create best possible conditions for audio as shown in FIG. 6.Methods of Operation

[0136] FIG. 7 is a flowchart for a method 700 for establishing a communication connection (e.g., audio connection) between devices, according to an embodiment. Method 700 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 7, as will be understood by a person of ordinary skill in the art. Method 700 shall be described with reference to FIG. 1. However, method 700 is not limited to that example embodiment.

[0137] In 702, device 104 is provided to a participant in a session. Device 104 may be attached to a wearable element (e.g., wearable element 112). Device 104 may include a plurality of microphones arranged in a beamforming configuration. The beamforming configuration has a beamforming direction toward a mouth of the user when mounted in the wearable element and worn by the user.

[0138] In 704, device 104 may capture audio data via the plurality of microphones.

[0139] In 706, device 104 may process the audio data to remove undesired audio. The undesired audio may include audio noise. In some aspects, the undesired audio may include audio signal originating from sources other than device 104. In some aspects, processing the audio data may be performed by Al module 120. In some aspects, device 104 may identify audio characteristics in the audio data using an artificial intelligence model. The audio characteristics may be associated with the voice of the user. Device 104 may filter the audio that does not correspond to the audio characteristics of the user (e.g., audio wherein the source is not the user).

[0140] In some aspects, device 104 may determine a strength level of the captured audio data. Device 104 may switch between a first microphone of the plurality of microphones and a second microphone of the plurality of microphones based on the strength level of the audio data. For example, if the strength level exceeds a threshold, device 104 may switch to the second microphone. Device 104 may capture and use the audio data from the second microphone.

[0141] In 708, device 104 may establish a communication connection between device 104 and another device (e.g., another device 104, docking station 106, electronic device 110, platform 102).

[0142] In 710, device 104 may send via the communication connection the audio data and / or other data from the device to the another device. The another device may include another audio device (e.g., device 104, another computer system, a cloud service, docking station 106, electronic device 110, platform 102).

[0143] FIG. 8 is a flowchart for a method 800 for generating a graphical representation of communication data, according to an embodiment. Method 800 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 8, as will be understood by a person of ordinary skill in the art. Method 800 shall be described with reference to FIG. 1. However, method 800 is not limited to that example embodiment.

[0144] In 802, platform 102 may acquire audio data from a plurality of audio devices (e.g., device 104). The audio data may include a plurality of audio streams. In someaspects, the plurality of audio devices are associated with a plurality of players of a sport activity (e.g., soccer).

[0145] In 804, platform 102 may analyze the audio data to generate communication data using an Al model. In some aspects, platform 102 may detect a communication from a player of the plurality of players. In some aspects, platform 102 may determine a type of communication. For example, platform 102 may classify the communication as an orientation communication, a positive communication, a negative communication, a selftalk or a stimulation communication. The orientation communication may refer to a communication that includes guidance from a player to other players. The Al model may be trained to classify the orientation communication by detecting task specific command such as “move to the right” or “up” using a speech recognition algorithm. The self-talk communication may refer a communication by a participant addressed to the participant himself. For example, the self-talk communication may be a motivational phrase to boost confidence (e.g., “Keep moving”). In some aspects, a communication may be classified in more than one category. For example, a communication may be categorized as a self-talk communication and a positive communication (e.g., “Keep moving”) or a self-talk communication and a negative communication (e.g., “I am going to miss”).

[0146] In 806, platform 102 may generate a graphical representation of the communication data. In some aspects, the graphical representation may show a formation of the players during play or training.

[0147] FIG. 9 is a flowchart for a method 900 for generating a visual representation of a tagged event, according to an embodiment. Method 900 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 9, as will be understood by a person of ordinary skill in the art. Method 900 shall be described with reference to FIG. 1. However, method 900 is not limited to that example embodiment.

[0148] In 902, platform 102 may acquire a live audio stream from one or more audio devices associated with one or more participants in a sport activity. The live audio stream may be captured during a training event or a match.

[0149] In 904, platform 102 may receive a verbal command to tag an event in the live audio stream. In some aspects, platform 102 may detect the verbal command in audio data in the live audio stream. The tag may include a question for the coach, a tactical tag, or the like.

[0150] In 906, platform 102 may store the tagged event with corresponding participant data. Stored participant data may include a profile identification of the player, previous metrics (e.g., number of interactions in a previous session), or the like. Platform 102 may synchronize the audio data with video data that corresponds to the tagged event.

[0151] In 908, platform 102 may generate a visual representation of the tagged event. For example, platform 102 may display information associated with the tagged event based on an input from a participant such as the coach. The coach may review the one or more tagged events. The visual representation may comprise video data corresponding to the tagged events. In some aspects, platform 102 may display an indication of the tagged events in a video feed at corresponding time in the video data. In addition, platform 102 may output any third party data associated with the tagged event. The third party data may be received from third party system 132.

[0152] FIG. 10 is a flowchart for a method 1000 for outputting video data, according to an embodiment. Method 1000 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 10, as will be understood by a person of ordinary skill in the art. Method 1000 shall be described with reference to FIG. 1. However, method 1000 is not limited to that example embodiment.

[0153] In 1002, platform 102 may acquire audio data from one or more audio devices. Each audio device may be associated with a player.

[0154] In 1004, platform 102 may acquire video data for a period corresponding to the audio data. For example, platform 102 may obtain from an external system a video stream for the duration of the session.

[0155] In 1006, platform 102 may synchronize the audio data with the video data (e.g., using timestamps). In some aspects, platform 102 may synchronize the audio data with the video data using Al module 120 by identifying similar audio characteristics betweenaudio from audio devices (e.g., device 104) and audio from a video source. The audio characteristics may include audio associated with an event such as a whistle that is detected in both audio from audio devices and audio from the video source. Audio characteristics may include spoken language by a participant but also pitch, level, sound from an event (e.g., ball being kicked, head hitting a ball).

[0156] In 1008, platform 102 may identify an event in the audio data or the video data based on a user input or using an Al model. In some aspects, the event may include scoring a goal, a foul, a referee whistle, an injury of a player, or the like. In some aspects, the Al model may detect and identify the event based on sound characteristics associated with the event. In some aspects, platform 102 may identify the event in the video data.

[0157] In 1010, platform 102 may retrieve the audio data corresponding to the event using the synchronized audio data with the video data. In some aspects, platform 102 may identify audio data from one or more players that correspond to the event. In some aspects, a user may select one or more players and platform 102 may retrieve audio data corresponding to the one or more players (e.g., to review a player interaction prior to or after the event). In some aspects, platform 102 may retrieve video data corresponding to the event. In some aspects, platform 102 may identify video data that correspond to the event.

[0158] In 1012, platform 102 may output video data corresponding to the event and the retrieved audio data. In some aspects, platform 102 may output the audio data with the retrieved video data.

[0159] FIG. 11 is a flowchart for a method 1100 for determining a likelihood of a message reception, according to an embodiment. Method 1100 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 11, as will be understood by a person of ordinary skill in the art. Method 1100 shall be described with reference to FIG. 1. However, method 1100 is not limited to that example embodiment.

[0160] In 1102, platform 102 may acquire audio data comprising a plurality of audio signals or audio streams. The audio data may be acquired from a plurality of devices (e.g.,device 104) via network 108. In some aspects, the plurality of devices may correspond to all active devices during a session.

[0161] In 1104, platform 102 may identify the audio signal having the highest decibel level among the plurality of audio signals.

[0162] In 1106, platform 102 may identify a device from the plurality of devices from which the audio signal having the highest decibel level was acquired.

[0163] In 1108, platform 102 may determine a signal strength of a communication from the audio device by identifying identical sound characteristics in respective audio signals of the plurality of devices. For example, platform 102 may identify one or more devices that captured identical sound characteristics (compared to the communication) and may determine a decibel level for the communication in identified devices.

[0164] In 1110, platform 102 may determine a direction of a participant for each respective device based on data acquired from one or more sensors in the device (e.g., motion sensor 216).

[0165] In 1112, platform 102 may determine a score corresponding to a likelihood of a message reception for each device based on the respective signal level. In some aspects, platform 102 may generate a visual representation of a communicative heatmap based on at least the score.

[0166] In some aspects, platform 102 may determine position data of each device of the plurality of devices. The position data may be used to represent the position of each device in the visual representation. The position data may be determined based on a WIFI signal strength, global positioning system (GPS) data, or video data corresponding to the audio data.

[0167] FIG. 12 is a flowchart for a method 1200 for generating a coaching instruction, according to an embodiment. Method 1200 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 12, as will be understood by a person of ordinary skill in the art. Method 1200 shall be described with reference to FIG. 1. However, method 1200 is not limited to that example embodiment.

[0168] In 1202, platform 102 may acquire audio data from a plurality of devices associated with a plurality of players.

[0169] In 1204, platform 102 may determine a communication metric for each player of the plurality of players based on the audio data. The communication metric may include an average number of interactions by minute, a number of negative communications, or a metric comparing the average interactions in the current session with previous session.

[0170] In 1206, platform 102 may generate a coaching instruction for one or more players of the plurality of players based on the respective communication metric. As discussed previously, the coaching instruction may be generated by Al module 120. In some aspects, the coaching instruction is generated based on the communication metric or the video data associated with the audio data. The coaching instruction may indicate a change in the tactics based the metric. In some aspects, the coaching instruction may comprise outputting a soothing music to decrease a stress level of the player. In some aspects, the coaching instruction may be generated based on data acquired from a third party source (e.g., third party system 132). The data may include historical data associated with a team playing style, tactics, technical abilities, physics, and mental state.

[0171] In some aspects, the coaching instructions is provided to the device associated with each player in a language corresponding to the language stored in a respective profile of the one or more players. In some aspects, the language may be detected based on the communication from the player using a machine learning algorithm. The detected language is associated with the player and stored in a player profile as described previously herein.

[0172] FIG. 13 is a flowchart for a method 1300 for determining an intensity of an impact during a sport activity, according to an embodiment. Method 1300 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 13, as will be understood by a person of ordinary skill in the art. Method 1300 shall be described with reference to FIG. 1. However, method 1300 is not limited to that example embodiment.

[0173] In 1302, platform 102 may acquire audio data from a device. The device is associated with a participant in a sport activity and is worn by the participant during the sport activity.

[0174] In 1304, platform 102 may detect an impact event based on at least the audio data. For example, platform 102 may detect that the participant is involved in the impact event based on at least the audio data. In some aspects, Al module 120 may be trained to detecting the impact event by identifying sound characteristics associated with the event (e.g., sound characteristics of a fall, sound characteristics of a head collision).

[0175] In 1306, platform 102 may determine one or more biometrics of the player based on the audio data.

[0176] In 1308, platform 102 may determine an intensity of the impact event based on the one or more biometrics. In addition or alternatively, platform 102 may determine a fatigue level of the player based on the one or more biometrics.Example user interfaces

[0177] FIG. 14A illustrates a user interface for displaying session data, in accordance with an embodiment of the present disclosure. A user may access a landing page 1400 for data analysis. The user is provided with one or more graphical user interfaces (GUIs) to select or display previous session data. To access data associated with a previous session, the user is provided with a pane 1402 that includes a plurality of GUI objects associated with a list of previous sessions. Pane 1402 may also show the type of the session, the duration of the session, the corresponding coach, communication data and metrics associated with the session (e.g., average interaction length, average number of interaction, orientation communication percentage), an indication whether corresponding video data is available, and the detected language of the communications. Landing page 1400 may also display metrics associated with the previous sessions. For example, pane 1406 may show the number of past sessions that the user may access. Pane 1408 shows statistics associated with types of the previous sessions. For example, pane 1408 shows that 80% of the previous sessions are matches while 20% are training sessions. Landing page 1400 may also shows metrics associated with communication data. For example, pane 1410 shows an interaction type distribution (e.g., percentage of orientation communication, percentage of stimulation communication, percentage of positive evaluation, and percentage of negative evaluation). Pane 1412 may show an averagenumber of interactions per minute. Upon activating of GUI object 1404 associated with a previous session, the user is presented with a user interface associated with the previous session as shown in FIG. 14B.

[0178] FIG. 14B illustrates a user interface for displaying communication data associated with a session, in accordance with an embodiment of the present disclosure. A pane 1414 shows the communication data from the previous session. The communication data may include a total number of interactions, a length of the session, a distribution of the interaction types, a number of interactions per minute, and a graphical representation 1416 of a formation of the players.

[0179] In some aspects, graphical representation 1416 may show a distribution of the players on the field. For example, graphical representation 1416 may show a position and a name of the player along with an average interaction per minute during the session. The higher the number is the more vocal is the player. The graphical representation may be sport specific. The user may be provided with a “total” GUI object 1418, a “negative evaluation” GUI object 1420, an “orientation” GUI object 1422, a “positive evaluation” GUI object 1424, and a “stimulation” GUI object 1426. The user may activate one of the GUI objects to change the metric displayed in the graphical representation. For example, the user may activate “negative evaluation” GUI object 1420 to show the average number of “negative evaluation” for each player.

[0180] In addition to the visual representation of the players, pane 1414 may show the number of interactions with respect to time. Platform 102 may associate each interaction with a timestamp based on the timestamps received with the audio data. Thus, the user can track the development of interactions over time. Graph 1428 show the number of interactions as a function of time. The number of interactions may be associated with one player or the average interactions of all active players during the session as shown in FIG. 14C.

[0181] FIG. 14C illustrates a user interface for displaying communication data associated with a session, in accordance with an embodiment of the present disclosure. The user may be provided with a drop down list 1438 to select one or more players. For example, the user may select a player and the average interactions of the team in order to compare a specific player with the team. The user may also filter the interactions based on the type of interaction by selecting the respective GUI object.

[0182] FIG. 14D illustrates a user interface for displaying communication data associated with a player, in accordance with an embodiment of the present disclosure. The user may review communication data associated with a specific player. Upon selection of a player name or identification, the user may be provided with pane 1430. Pane 1430 may show the total number of sessions the user participated in. The user may filter the data by match or training session. In addition, pane 1430 may show metric associated with the interaction type (e.g., distribution of the interaction type). Pane 1430 may show a list of all the sessions including the date, the name of the session, the average number of interactions of the player in each session, and the percentage of orientation interactions. In addition, graph 1432 shows the number of interactions with respect of time for a particular session.

[0183] FIG. 14E illustrates a user interface for displaying synchronized video data with audio data, in accordance with an embodiment of the present disclosure. A video pane 1434 may display a video stream for a session and a graphical representation pane 1436 may show a representation of the position of the player with an indication of whether audio data is available. The indication changes as the video stream progress. For example, an audio symbol on a GUI object may be shown at a position of the player in the visual representation when audio is available. The user may listen to the audio by activating the respective GUI object. For example, a coach may review critical moments in the game while listening to the interactions of one or more players. The user may also tag the audio associated with the critical moment (e.g., midfielder communication).

[0184] As discussed above, by saying a certain “word+ what to tag” players and coaches can tag events of the game while in play. In some aspects, device 104 may activate a tag feature by detecting one or more movements initiated by the participant. By single tapping or double tapping on device 104, players and coaches can tag events. In some aspects, device 104 may detect a tap on device 104 or in the proximity to device 104. The tap may be detected by microphone 204 and / or motion sensor 216. As used herein, a tap may refer to a light strike, or other user movement that causes the motion sensor 216 to undergo acceleration. The tap may be on device 104 or in the proximity of device 104. For example, tapping on the wearable element where the device 104 is positioned (e.g., attached) may cause an acceleration spike in the output of motion sensor 216. In response to detecting a tap, device 104 may associate the tag with a timestamp corresponding to the tap. The timestamps corresponding to the tags may be stored in database 116.

[0185] In some aspects, motion sensor 216 may detect motion data (e.g., a movement and an orientation of the player). In some aspects, motion sensor 216 of device 104 may be an IMU. In some aspects, motion sensor 216 may be an accelerometer.

[0186] Device 104 may analyze data from motion sensor 216 to detect the tap. The data may include an output signal that represents an acceleration of the device as a function of time. Device 104 may identify the tap by comparing the acceleration (e.g., an absolute value of the acceleration) with a threshold. Strong positive or negative accelerations can be detected. Device 104 may detect a tap when the acceleration is greater than the threshold. Device 104 may activate the tag feature when accelerations that exceed the threshold are detected. However, some of these accelerations may represent false positives that do not correspond to taps by the player. For example, a player may unintentionally touch device 104. Accelerations may also be caused by physical contact between players.

[0187] In some aspects, device 104 may trigger the tagging action when the player double tap on device 104 or in proximity of device 104. Device 104 may determine if a double tap occurs by analyzing data from motion sensor 216. In some aspects, device 104 may detect a first tap (e.g., absolute value of the acceleration exceeds a threshold) and store timing information associated with the first tap. Device 104 may detect a second tap and store respective timing information. Then, device 104 may determine the time between the first tap and the second tap to determine whether a double tap has occurred. For example, if the time is within a predetermined range, then a double tap has occurred, and device 104 may activate the tag feature. Device 104 may associate a tag with the timestamp of the first tap or the second tap. In some aspects, a derivative of the acceleration may be used and compared with a threshold to detect the first tap and the second tap. In some aspects, the predetermined range may be user specific. For example, the predetermined range may be pre-calibrated for a player (e.g., multiple double tap are collected and analyzed to determine the predetermined range).

[0188] In some aspects, in response to detecting the double tap and activating the tag feature, device 104 may activate microphone 204. Microphone 204 may capture audio that comprise a verbal tag. Device 104 may associate the timestamp of the double tap with the verbal tag. For example, a player may double tap on device 104 and then say “defense”. Device 104 may detect the double tap using data from motion sensor 216. Then, device 104 may associate the verbal tag “defense” with the timestamp of the doubletap. Device 104 may add the verbal tag to categorized notes associated with device 104. For example, different categories may be predefined (e.g., “defense”, “feedback”, etc.). After the play session ends, the player or the trainer may review all timestamps associated with a category. Device 104 may group all timestamps associated with each category together. In some aspects, if device 104 does not detect a verbal tag after the double tap, device 104 may store the timestamp in an uncategorized category.

[0189] In some aspects, a different number of taps may trigger the tagging (e.g., three or more consecutive taps). Different actions may be triggered based on the number of detected taps or tap patterns. For example, three taps may be used to switch between the plurality of mode of operations (e.g., listen mode, speak mode, or two-way communication mode.) In some aspects, device 104 may detect a tap pattern (e.g., a double tap followed by a single tap, a single tap followed by a double tap) and trigger one or more actions associated with the tap pattern.

[0190] In some aspects, device 104 may transmit the output of motion sensor 216 to platform 102. Platform 102 may synchronize the output of motion sensor 216 with video content. Then, using timestamps of the tags, coaches and player may find and access relevant video and audio content.

[0191] In some aspects, an Al model may be used to detect a double tap. In some aspects, Al module 120 may be trained to detect a tap or a double tap on device 104. Al module 120 may receive as an input the acceleration or a derivative of the acceleration from motion sensor 216. Al module 120 may output a timestamp of each detected double tap. The timestamps may be associated with tags. In some aspects, Al model may be trained using data associated with each player or based for a specific sport activity.

[0192] Microphone 204 may capture audio associated with the tap. Device 104 may analyze the audio signal to detect audio characteristics associated with the tap or a double tap. In some aspects, device 104 may use outputs from motion sensor 216 and microphone 204 to detect whether a double tap has occurred. In some aspects, to minimize false positive, device 104 may trigger the tagging action when the double tap is detected by motion sensor 216 and by microphone 204.

[0193] FIG. 16 is a flowchart for a method 1600 for establishing a communication connection (e.g., audio connection) between devices, according to an embodiment. Method 1600 can be performed by processing logic that can comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g.,instructions executing on a processing device), or a combination thereof. It is to be appreciated that not all steps may be needed to perform the disclosure provided herein. Further, some of the steps may be performed simultaneously, or in a different order than shown in FIG. 16, as will be understood by a person of ordinary skill in the art. Method 1600 shall be described with reference to FIG. 1. However, method 1600 is not limited to that example embodiment.

[0194] As discussed above, device 104 is provided to a participant in a session. Device 104 may be attached to a wearable element (e.g., wearable element 112). Device 104 may include a plurality of microphones arranged in a beamforming configuration and a motion sensor (e.g., motion sensor 216).

[0195] In 1602, device 104 may detect a number of taps (e.g., a single tap, a double tap) using data from motion sensor 216 and / or microphone 204.

[0196] In 1604, device 104 may associate a tag with a timestamp corresponding to when the number of taps are detected. In some aspects, device 104 may further capture a verbal tag using microphone 204 and associate the verbal tag with the timestamp.Components of the System

[0197] FIG. 15 shows a computer system 1500, according to some embodiments. Various embodiments and components therein can be implemented, for example, using computer system 1500 or any other well-known computer systems. For example, the method steps of FIGS. 6-13 may be implemented via computer system 1500.

[0198] In some aspects, computer system 1500 may comprise one or more processors (also called central processing units, or CPUs), such as a processor 1504. Processor 1504 may be connected to a communication infrastructure or bus 1506.

[0199] In some aspects, one or more processors 1504 may each be a graphics processing unit (GPU). In an embodiment, a GPU is a processor that is a specialized electronic circuit designed to process mathematically intensive applications. The GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common to computer graphics applications, images, videos, etc.

[0200] In some aspects, computer system 1500 may further comprise user input / output device(s) 1503, such as monitors, keyboards, pointing devices, etc., that communicate with communication infrastructure 1506 through user input / output interface(s) 1502.Computer system 1500 may further comprise a main or primary memory 1508, such as random access memory (RAM). Main memory 1508 may comprise one or more levels of cache. Main memory 1508 has stored therein control logic (e.g., computer software) and / or data.

[0201] In some aspects, computer system 1500 may further comprise one or more secondary storage devices or memory 1510. Secondary memory 1510 may comprise, for example, a hard disk drive 1512 and / or a removable storage device or drive 1514. Removable storage drive 1514 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, tape backup device, and / or any other storage device / drive. Removable storage drive 1514 may interact with a removable storage unit 1518. Removable storage unit 1518 may comprise a computer usable or readable storage device having stored thereon computer software (control logic) and / or data. Removable storage unit 1518 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / any other computer data storage device. Removable storage drive 1514 reads from and / or writes to removable storage unit 1518 in a well- known manner.

[0202] In some aspects, secondary memory 1510 may comprise other means, instrumentalities or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 1500. Such means, instrumentalities or other approaches may comprise, for example, a removable storage unit 1522 and an interface 1520. Examples of the removable storage unit 1522 and the interface 1520 may comprise a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.

[0203] In some aspects, computer system 1500 may further comprise a communication or network interface 1524. Communication interface 1524 enables computer system 1500 to communicate and interact with any combination of remote devices, remote networks, remote entities, etc. (individually and collectively referenced by reference number 1528). For example, communication interface 1524 may allow computer system 1500 to communicate with remote devices 1528 over communications path 1526, which may be wired and / or wireless, and which may comprise any combination of LANs, WANs, theInternet, etc. Control logic and / or data may be transmitted to and from computer system 1500 via communications path 1526.

[0204] In some aspects, a non-transitory, tangible apparatus or article of manufacture comprising a non-transitory, tangible computer useable or readable medium having control logic (software) stored thereon is also referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 1500, main memory 1508, secondary memory 1510, and removable storage units 1518 and 1522, as well as tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 1500), causes such data processing devices to operate as described herein.

[0205] Based on the teachings contained in this disclosure, it will be apparent to those skilled in the relevant art(s) how to make and use embodiments of this disclosure using data processing devices, computer systems and / or computer architectures other than that shown in FIG. 15. In particular, embodiments may operate with software, hardware, and / or operating system implementations other than those described herein.

[0206] It is to be understood that the phraseology or terminology herein is for the purpose of description and not of limitation, such that the terminology or phraseology of the present disclosure is to be interpreted by those skilled in relevant art(s) in light of the teachings herein.

[0207] It is to be appreciated that the Detailed Description section, and not the Summary and Abstract sections, is intended to be used to interpret the claims. The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by the inventor(s), and thus, are not intended to limit the present disclosure and the appended claims in any way.

[0208] The present disclosure has been described above with the aid of functional building blocks illustrating the implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternate boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed.

[0209] While specific embodiments of the disclosure have been described above, it will be appreciated that embodiments of the present disclosure may be practiced otherwise than as described. The descriptions are intended to be illustrative, not limiting. Thus, itwill be apparent to one skilled in the art that modifications may be made to the disclosure as described without departing from the scope of the claims set out below.

[0210] The foregoing description of the specific embodiments will so fully reveal the general nature of the present disclosure that others can, by applying knowledge within the skill of the art, readily modify and / or adapt for various applications such specific embodiments, without undue experimentation, without departing from the general concept of the present disclosure. Therefore, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments, based on the teaching and guidance presented herein.

[0211] The breadth and scope of the protected subject matter should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

[0212] Exemplary embodiments are set out in the following numbered clauses.

[0213] Clause 1 : An audio device comprising: a housing configured to be attached to a wearable element worn on a body of a user; a plurality of microphones configured to capture audio data, wherein the plurality of microphones are arranged in a beamforming configuration having a beamforming direction towards a mouth of the user when worn by the user; and processing circuitry configured to: capture the audio data via the plurality of microphones; remove undesired audio from the audio data; establish a communication connection between the audio device and another device; and send, via the communication connection, the audio data from the audio device to the another device.

[0214] Clause 2: The audio device of clause 1, further comprising: a motion sensor configured to generate orientation data, and wherein the processing circuitry is further configured to: determine, based on the orientation data, a position of the audio device; and in response to determining that the position of the audio device fails to satisfy a device position configuration, generate an alert to the user.

[0215] Clause 3: The audio device of clause 1 or clause 2, further comprising:a noise absorbing material attached to an outer surface of the housing.

[0216] Clause 4: The audio device of any of clauses 1-3, wherein the plurality of microphones comprise a first microphone and a second microphone, and wherein the processing circuitry is further configured to: switch between the first microphone of the plurality of microphones and the second microphone of the plurality of microphones based on a strength level of the audio data; and capture the audio data via the second microphone of the plurality of microphones.

[0217] Clause 5: The audio device of any of clauses 1-4, wherein the processing circuitry is further configured to use at least one artificial intelligence (Al) model to remove the undesired audio from the audio data.

[0218] Clause 6: The audio device of any of clauses 1-5, wherein the processing circuity is further configured to: identify audio characteristics in the audio data using the at least one Al model, wherein the audio characteristics are associated with a voice of the user.

[0219] Clause 7: The audio device of any of clauses 1-6, wherein the plurality of microphones are arranged in a broadside beamforming configuration.

[0220] Clause 8: The audio device of any of clauses 1-7, wherein the plurality of microphones are arranged in an endfire beamforming configuration.

[0221] Clause 9: The audio device of any of clauses 1-8, wherein each of the plurality of microphones is a micro-electro-mechanical system (MEMS) microphones.

[0222] Clause 10: The audio device of any of clauses 1-9 , wherein the beamforming direction is a first beamforming direction and wherein the beamforming configuration supports beamforming in a plurality of beamforming directions including the first beamforming directions.

[0223] Clause 11 : The audio device of any of clauses 1-10, further comprising: an additional microphone configured to capture acoustical noise, and wherein the processing circuitry is further configured to: filter the audio data captured via the plurality of microphones based on at least the acoustical noise captured by the additional microphone.

[0224] Clause 12: A system comprising: an audio device comprising:a housing configured to be attached to a wearable element worn on a body of a user; a plurality of microphones configured to capture audio data, wherein the plurality of microphones are arranged in a beamforming configuration having at least a beamforming direction towards a mouth of the user when worn by the user; and processing circuitry configured to: capture the audio data via the plurality of microphones; remove undesired audio from the audio data; establish a communication connection between the audio device and another device; and send, via the communication connection, the audio data from the audio device to the another device; a docking station comprising: processing circuitry configured to acquire and process audio data from the audio device; and charging circuitry configured to provide power to the audio device.

[0225] Clause 13: The system of clause 12, wherein the processing circuitry of the audio device is further configured to: upload the audio data to a computer platform at predefined intervals.

[0226] Clause 14: A method for communication between devices, comprising: capturing audio data via a plurality of microphones of an audio device arranged in a beamforming configuration having at least a beamforming direction towards a mouth of a user when mounted in a wearable element; removing, by at least one computer processor, undesired audio from the audio data; establishing a communication link between the audio device and another device; and sending, via the communication link, the audio data from the audio device to the another device.

[0227] Clause 15: The method of clause 14, further comprising: providing a noise absorbing material within the wearable element or the audio device.

[0228] Clause 16: The method of clause 14 or clause 15, wherein the processing further comprises: identifying, using an artificial intelligence model, audio characteristics in the audio data, wherein the audio characteristics are associated with a voice of the user.

[0229] Clause 17: The method of any of clauses 14-16, further comprising: switching between a first microphone of the plurality of microphones and a second microphone of the plurality of microphones based on a strength level of the audio data; and capturing the audio data via the second microphone of the plurality of microphones.

[0230] Clause 18: The method of any of clauses 14-17, further comprising: determining, based on data from a motion sensor, a position of the audio device; and in response to determining that the position of the audio device fails to satisfy a device position configuration, generating an alert to the user.

[0231] Clause 19: The method of any of clauses 14-18, further comprising: capturing using an additional microphone acoustical noise, and filtering the audio data based on at least the acoustical noise.

[0232] Clause 20: A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: capturing audio data via a plurality of microphones of an audio device arranged in a beamforming configuration having at least a beamforming direction towards a mouth of a user when mounted in a wearable element; removing undesired audio from the audio data; establishing a communication link between the audio device and another device; and sending, via the communication link, the audio data from the audio device to the another device.

[0233] Clause 21 : The non-transitory computer-readable medium of clause 20, wherein the operations further comprise: identifying, using an artificial intelligence model, audio characteristics in the audio data, wherein the audio characteristics are associated with a voice of the user.

[0234] Clause 22: The non-transitory computer-readable medium of clause 20 or clause21, wherein the operations further comprise: switching between a first microphone of the plurality of microphones and a second microphone of the plurality of microphones based on a strength level of the audio data; and capturing the audio data via the second microphone of the plurality of microphones.

[0235] Clause 23 : A method for generating a graphical representation of communication data, comprising: acquiring audio data from a plurality of audio devices, wherein the plurality of audio devices are associated with a plurality of participants in a sport activity and each respective audio device is worn by each participant of the plurality of participants during the sport activity; generating communication data by analyzing the audio data using an artificial intelligence (Al) model; and generating the graphical representation of the communication data.

[0236] Clause 24: The method of clause 23, wherein the graphical representation comprises a visual representation of a player formation associated with the sport activity.

[0237] Clause 25: The method of clause 23 or clause 24, wherein generating the audio data comprises: detecting a communication from a participant of the plurality of participants; and determining a category associated with the communication, wherein the category is one of an orientation communication, a positive communication, a negative communication, a stimulation communication, or a self-talk communication; and wherein the communication data comprises at least a number of interactions for each category for each respective user of an audio device of the plurality of audio devices.

[0238] Clause 26: The method of clause 25, wherein the category is one of an orientation communication, a positive communication, a negative communication, a stimulation communication, or a self-talk communication.

[0239] Clause 27: The method of any of clauses 23-26, wherein the audio data comprises a live audio stream, and the method further comprises: detecting, in the live audio stream, a verbal command to tag an event; tagging a portion of the audio data based on the verbal command;synchronizing the tagged portion of the audio data with stored participant data; and generating the graphical representation based on the tagged audio data and the tagged event, wherein the graphical representation comprises video data corresponding to the audio data.

[0240] Clause 28: The method of any of clauses 23-27, further comprising: acquiring video data for a period corresponding to a duration of the audio data; synchronizing the audio data with the video data; identifying an event in the audio data or the video data based on a user input or using the artificial intelligence (Al) model; retrieving the audio data corresponding to the event using the synchronized audio data with the video data; and outputting the video data corresponding to the event and the retrieved audio data.

[0241] Clause 29: The method of clause 28, wherein: the plurality of audio devices are associated with a plurality of participants; the retrieved audio data comprises the audio data from one or more participants of the plurality of participants; and the one or more participants are identified based on the user input.

[0242] Clause 30: The method of any of clauses 23-29, wherein the audio data comprises a plurality of audio signals and the method further comprises: identifying an audio signal having a highest decibel level among the plurality of audio signals; identifying, from the plurality of audio devices, an audio device corresponding to a source of the audio signal having the highest decibel level; determining a signal strength of a message from the audio device by identifying identical sound characteristics in respective audio signals of the plurality of audio devices; and determining a score corresponding to a likelihood of a message reception for each audio device of the plurality of audio devices based on a respective signal strength.

[0243] Clause 31 : The method of clause 30, further comprising: generating a communicative heatmap based on at least the score.

[0244] Clause 32: The method of clause 31, further comprising:determining position data of each audio device of the plurality of audio devices, wherein the communicative heatmap further comprises the position data, and wherein the position data comprises an orientation of a participant associated with each audio device of the plurality of audio devices.

[0245] Clause 33: The method of clause 32, wherein the position data is determined based on a WiFi signal strength, global positioning system (GPS) data, or video data corresponding to the audio data.

[0246] Clause 34: The method of any of clauses 23-33, wherein the plurality of audio devices are associated with a plurality of participants of a sport activity, and the method further comprises: determining a communication metric for each participant of the plurality of participants based on the audio data; and generating a coaching instruction for one or more participants of the plurality of participants based on a respective communication metric.

[0247] Clause 35: The method of clause 34, wherein generating the coaching instruction comprises generating the coaching instruction based on video data corresponding to the audio data.

[0248] Clause 36: The method of clause 34, further comprising: transmitting the coaching instruction to the audio device associated with each participant of the one or more participants in a language corresponding to the language stored in a respective profile of the one or more participants.

[0249] Clause 37: The method of any of clauses 23-36, further comprising: synchronizing respective audio data of the audio data received from the plurality of audio devices.

[0250] Clause 38: A computing system, comprising: one or more memories; and at least one processor each coupled to at least one of the memories and configured to perform operations comprising: acquiring audio data from a plurality of audio devices, wherein the plurality of audio devices are associated with a plurality of participants in a sport activity and each respective audio device is worn by each participant of the plurality of participants during the sport activity;generating communication data by analyzing the audio data using an artificial intelligence (Al) model; and generating a graphical representation of the communication data.

[0251] Clause 39: The computing system of clause 38, wherein the graphical representation comprises a visual representation of a player formation associated with the sport activity.

[0252] Clause 40: The computing system of clause 38 or clause 39, wherein the operations further comprise: detecting a communication from a participant of the plurality of participants; and determining a category associated with the communication; and wherein the communication data comprises at least a number of interactions for each category for each respective user of an audio device of the plurality of audio devices.

[0253] Clause 41 : The computing system of clause 40, wherein the category is one of an orientation communication, a positive communication, a negative communication, a stimulation communication, or a self-talk communication.

[0254] Clause 42: The computer system of any of clauses 38-41, wherein the audio data includes a live audio stream, and the operations further comprise: detecting, in the live audio stream, a verbal command to tag an event; tagging a portion of the audio data based on the verbal command; synchronizing the tagged portion of the audio data with stored participant data; and generating the graphical representation based on the tagged audio data and the tagged event, wherein the graphical representation comprises video data corresponding to the audio data.

[0255] Clause 43 : A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: acquiring audio data from a plurality of audio devices, wherein the audio devices are associated with a plurality of participants in a sport activity and each respective audio device is worn by each participant of the plurality of participants during the sport activity; generating communication data by analyzing the audio data using an artificial intelligence (Al) model; and generating a graphical representation of the communication data.

[0256] Clause 44: The non-transitory computer-readable medium of clause 43, wherein the plurality of audio devices are associated with a plurality of participants of a sport activity; and wherein the graphical representation comprises a visual representation of a player formation associated with the sport activity.

[0257] Clause 45: The non-transitory computer-readable medium of clause 43 or clause 44, wherein the operations further comprise: detecting a communication from a participant of the plurality of participants; and determining a category associated with the communication, wherein the category is one of an orientation communication, a positive communication, a negative communication, a stimulation communication, or a self-talk communication; and wherein the communication data comprises at least a number of interactions for each category for each respective user of an audio device of the plurality of audio devices.

[0258] Clause 46: A method, comprising: acquiring audio data from an audio device, wherein the audio device is associated with a participant in a sport activity and is worn by the participant during the sport activity; determining, using at least one computer processor, one or more biometrics of the participant based on the audio data; and determining a status of the participant based on the one or more biometrics or the audio data.

[0259] Clause 47: The method of clause 46, wherein determining the status of the participant further comprises: detecting that the participant is involved in an impact event based on at least the audio data; and determining an intensity of the impact event based on the one or more biometrics.

[0260] Clause 48: The method of clause 46 or clause 47, wherein the status of the participant is indicative of a health status of the participant, and wherein determining the status of the participant further comprises determining the status of the participant based on the one or more biometrics.

[0261] Clause 49: The method of any of clauses 46-48, wherein the status comprises a fatigue level or a risk of concussion.

[0262] Clause 50: The method of any of clauses 46-49, further comprising:acquiring motion data from a motion sensor associated with the audio device; and wherein determining the status of the participant further comprises determining the status of the participant based on the one or more biometrics and the motion data.

[0263] Clause 51 : The method of any of clauses 46-50, further comprising: determining the one or more biometrics based on a content of the audio data.

[0264] Clause 52: The method of any of clauses 46-51, wherein the one or more biometrics comprise one or more of a heart rate and a respiratory pattern.

[0265] Clause 53: The method of any of clauses 46-52, further comprising: determining a mental state based on the one or more biometrics; generating a coaching instruction based on the mental state; and outputting the coaching instruction to the audio device.

[0266] Clause 54: A computing system, comprising: one or more memories; and at least one processor each coupled to at least one of the memories and configured to perform operations comprising: acquiring audio data from an audio device, wherein the audio device is associated with a participant in a sport activity and is worn by the participant during the sport activity; determining one or more biometrics of the participant based on the audio data; and determining a status of the participant based on the one or more biometrics or the audio data.

[0267] Clause 55: The computing system of clause 54, wherein the operations further comprise: detecting that the participant is involved in an impact event based on at least the audio data; and determining an intensity of the impact event based on the one or more biometrics.

[0268] Clause 56: The computing system of clause 54 or clause 55, wherein the status of the participant is indicative of a health status of the participant, and wherein the operations further comprise: determining the status based on the one or more biometrics.

[0269] Clause 57: The computing system of any of clauses 54-56, wherein the status comprises a fatigue level or a risk of concussion.

[0270] Clause 58: The computing system of any of clauses 54-57, wherein the operations further comprise: acquiring motion data from a motion sensor associated with the audio device; and determining the health status based on the one or more biometrics and the motion data.

[0271] Clause 59: The computing system of any of clauses 54-58, wherein the operations further comprise: determining the one or more biometrics based on a content of the audio data.

[0272] Clause 60: The computing system of any of clauses 54-59, wherein the one or more biometrics comprise one or more of a heart rate and a respiratory pattern.

[0273] Clause 61 : The computing system of any of clauses 54-60, wherein the operations further comprise: determining a mental state based on the one or more biometrics; generating a coaching instruction based on the mental state; and outputting the coaching instruction to the audio device.

[0274] Clause 62: A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: acquiring audio data from an audio device, wherein the audio device is associated with a participant in a sport activity and is worn by the participant during the sport activity; determining one or more biometrics of the participant based on the audio data; and determining a status of the participant based on the one or more biometrics or the audio data.

[0275] Clause 63 : The non-transitory computer-readable medium of clause 62, wherein the operations further comprise: detecting that the participant is involved in an impact event based on at least the audio data; and determining an intensity of the impact event based on the one or more biometrics.

[0276] Clause 64: The non-transitory computer-readable medium of clause 62 or clause 63, wherein the status of the participant is indicative of a health status of the participant, and wherein the operations further comprise: determining the health status based on the one or more biometrics.

[0277] Clause 65: The non-transitory computer-readable medium of any of clauses 62-64, wherein the status includes a fatigue level or a risk of concussion.

[0278] Clause 66: A method for triggering an action, comprising: acquiring data from an audio device, wherein the audio device is configured to be attached to a wearable element worn on a body of a user; determining, by at least one computer processor, a number of taps on the audio device by analyzing the data; and triggering an action based on the determined number of taps.

[0279] Clause 67: The method of clause 66, further comprising: associating a tag with a timestamp corresponding to when f the number of taps are detected.

[0280] Clause 68: The method of clause 66 or clause 67, wherein the action corresponds to a tagging action, and the method further comprises: activating a microphone of the audio device; capturing, using the microphone, a verbal tag; and associating the verbal tag with the timestamp of the determined number of taps.

[0281] Clause 69: The method of any of clauses 66-68, wherein the determined number of taps is equal to two.

[0282] Clause 70: The method of any of clauses 66-69, wherein the data corresponds to motion data detected by a motion sensor of the audio device.

[0283] Clause 71 : The method of clause 70, further comprising: detecting an acceleration spike in the motion data by comparing the motion data to a threshold.

[0284] Clause 72: The method of any of clauses 66-71, further comprising: detecting a first tap; detecting a second tap; determining a duration between the first tap and the second tap; and comparing the duration with a predetermined range to determine whether a double tap has occurred.

[0285] Clause 73: A computing system, comprising: one or more memories; and at least one processor each coupled to at least one of the memories and configured to perform operations comprising:acquiring data from an audio device, wherein the audio device is configured to be attached to a wearable element worn on a body of a user; determining, by at least one computer processor, a number of taps on the audio device by analyzing the data; and triggering an action based on the determined number of detected taps.

[0286] Clause 74: The computing system of clause 73, wherein the operations further comprise: associating a tag with a timestamp corresponding to when the number of taps are detected.

[0287] Clause 75: The computing system of clause 74, wherein the action corresponds to a tagging action, and the operations further comprise: activating a microphone of the audio device; capturing, using the microphone, a verbal tag; and associating the verbal tag with the timestamp of the determined number of taps.

[0288] Clause 76: The computing system of any of clauses 73-75, wherein the determined number of taps is equal to two.

[0289] Clause 77: The computing system of any of clauses 73-76, wherein the data corresponds to motion data detected by a motion sensor of the audio device.

[0290] Clause 78: The computing system of clause 77, wherein the operations further comprise: detecting an acceleration spike in the motion data by comparing the motion data to a threshold.

[0291] Clause 79: The computing system of any of clauses 73-78, wherein the operations further comprise: detecting a first tap; detecting a second tap; determining a duration between the first tap and the second tap; and comparing the duration with a predetermined range to determine whether a double tap has occurred.

[0292] Clause 80: A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:acquiring data from an audio device, wherein the audio device is configured to be attached to a wearable element worn on a body of a user; determining a number of taps on the audio device by analyzing the data; and triggering an action based on the determined number of taps.

[0293] Clause 81 : The non-transitory computer-readable medium of clause 80, wherein the operations further comprise: associating a tag with a timestamp corresponding to when the number of taps are detected.

[0294] Clause 82: The non-transitory computer-readable medium of clause 81, wherein the action corresponds to a tagging action, and the operations further comprise: activating a microphone of the audio device; capturing, using the microphone, a verbal tag; and associating the verbal tag with the timestamp of the determined number of taps.

[0295] Clause 83 : The non-transitory computer-readable medium of any of clauses 80-82, wherein the data corresponds to motion data detected by a motion sensor of the audio device.

[0296] Clause 84: The non-transitory computer-readable medium of clause 83, wherein the operations further comprise: detecting an acceleration spike in the motion data by comparing the motion data to a threshold.

[0297] Clause 85: The non-transitory computer-readable medium of any of clauses 80-84, wherein the operations further comprise: detecting a first tap; detecting a second tap; determining a duration between the first tap and the second tap; and comparing the duration with a predetermined range to determine whether a double tap has occurred.

Claims

WHAT IS CLAIMED IS:

1. An audio device comprising: a housing configured to be attached to a wearable element worn on a body of a user; a plurality of microphones configured to capture audio data, wherein the plurality of microphones are arranged in a beamforming configuration having at least a beamforming direction towards a mouth of the user when worn by the user; and processing circuitry configured to: capture the audio data via the plurality of microphones; remove undesired audio from the audio data; establish a communication connection between the audio device and another device; and send, via the communication connection, the audio data from the audio device to the another device.

2. The audio device of claim 1, further comprising: a motion sensor configured to generate orientation data, and wherein the processing circuitry is further configured to: determine, based on the orientation data, a position of the audio device; and in response to determining that the position of the audio device fails to satisfy a device position configuration, generate an alert to the user.

3. The audio device of claim 1, further comprising: a noise absorbing material attached to an outer surface of the housing.

4. The audio device of claim 1, wherein the plurality of microphones comprise a first microphone and a second microphone, and wherein the processing circuitry is further configured to:switch between the first microphone of the plurality of microphones and the second microphone of the plurality of microphones based on a strength level of the audio data; and capture the audio data via the second microphone of the plurality of microphones.

5. The audio device of claim 1, wherein the processing circuitry is further configured to use at least one artificial intelligence (Al) model to remove the undesired audio from the audio data.

6. The audio device of claim 1, wherein the processing circuity is further configured to: identify audio characteristics in the audio data using the at least one Al model, wherein the audio characteristics are associated with a voice of the user.

7. The audio device of claim 1, wherein the plurality of microphones are arranged in a broadside beamforming configuration.

8. The audio device of claim 1, wherein the plurality of microphones are arranged in an endfire beamforming configuration.

9. The audio device of claim 1, wherein each of the plurality of microphones is a micro- electro-mechanical system (MEMS) microphones.

10. The audio device of claim 1, wherein the beamforming direction is a first beamforming direction and wherein the beamforming configuration supports beamforming in a plurality of beamforming directions including the first beamforming direction.

11. The audio device of claim 1, further comprising: an additional microphone configured to capture acoustical noise, and wherein the processing circuitry is further configured to: filter the audio data captured via the plurality of microphones based on at least the acoustical noise captured by the additional microphone.

12. A system comprising: an audio device comprising: a housing configured to be attached to a wearable element worn on a body of a user; a plurality of microphones configured to capture audio data, wherein the plurality of microphones are arranged in a beamforming configuration having at least a beamforming direction towards a mouth of the user when worn by the user; and processing circuitry configured to: capture the audio data via the plurality of microphones; remove undesired audio from the audio data; establish a communication connection between the audio device and another device; and send, via the communication connection, the audio data from the audio device to the another device; and a docking station comprising: processing circuitry configured to acquire and process audio data from the audio device; and charging circuitry configured to provide power to the audio device.

13. The system of claim 12, wherein the processing circuitry of the audio device is further configured to: upload the audio data to a computer platform at predefined intervals.

14. A method for communication between devices, comprising: capturing audio data via a plurality of microphones of an audio device arranged in a beamforming configuration having at least a beamforming direction towards a mouth of a user when mounted in a wearable element; removing, by at least one computer processor, undesired audio from the audio data; establishing a communication link between the audio device and another device; andsending, via the communication link, the audio data from the audio device to the another device.

15. The method of claim 14, further comprising: providing a noise absorbing material within the wearable element or the audio device.

16. The method of claim 14, wherein the processing further comprises: identifying, using an artificial intelligence model, audio characteristics in the audio data, wherein the audio characteristics are associated with a voice of the user.

17. The method of claim 14, further comprising: switching between a first microphone of the plurality of microphones and a second microphone of the plurality of microphones based on a strength level of the audio data; and capturing the audio data via the second microphone of the plurality of microphones.

18. The method of claim 14, further comprising: determining, based on data from a motion sensor, a position of the audio device; and in response to determining that the position of the audio device fails to satisfy a device position configuration, generating an alert to the user.

19. The method of claim 14, further comprising: capturing using an additional microphone acoustical noise, and filtering the audio data based on at least the acoustical noise.

20. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:capturing audio data via a plurality of microphones of an audio device arranged in a beamforming configuration having at least a beamforming direction towards a mouth of a user when mounted in a wearable element; removing undesired audio from the audio data; establishing a communication link between the audio device and another device; and sending, via the communication link, the audio data from the audio device to the another device.

21. The non-transitory computer-readable medium of claim 20, wherein the operations further comprise: identifying, using an artificial intelligence model, audio characteristics in the audio data, wherein the audio characteristics are associated with a voice of the user.

22. The non-transitory computer-readable medium of claim 20, wherein the operations further comprise: switching between a first microphone of the plurality of microphones and a second microphone of the plurality of microphones based on a strength level of the audio data; and capturing the audio data via the second microphone of the plurality of microphones.

23. A method for generating a graphical representation of communication data, comprising: acquiring audio data from a plurality of audio devices, wherein the plurality of audio devices are associated with a plurality of participants in a sport activity and each respective audio device is worn by each participant of the plurality of participants during the sport activity; generating communication data by analyzing the audio data using an artificial intelligence (Al) model; and generating the graphical representation of the communication data.

24. The method of claim 23, wherein the graphical representation comprises a visual representation of a player formation associated with the sport activity.

25. The method of claim 23, wherein generating the audio data comprises: detecting a communication from a participant of the plurality of participants; and determining a category associated with the communication,; and wherein the communication data comprises at least a number of interactions for each category for each respective user of an audio device of the plurality of audio devices.

26. The method of claim 25, wherein the category is one of an orientation communication, a positive communication, a negative communication, a stimulation communication, or a self-talk communication.

27. The method of claim 23, wherein the audio data comprises a live audio stream, and the method further comprises: detecting, in the live audio stream, a verbal command to tag an event; tagging a portion of the audio data based on the verbal command; synchronizing the tagged portion of the audio data with stored participant data; and generating the graphical representation based on the tagged audio data and the tagged event, wherein the graphical representation comprises video data corresponding to the audio data.

28. The method of claim 23, further comprising: acquiring video data for a period corresponding to a duration of the audio data; synchronizing the audio data with the video data; identifying an event in the audio data or the video data based on a user input or using the artificial intelligence (Al) model; retrieving the audio data corresponding to the event using the synchronized audio data with the video data; and outputting the video data corresponding to the event and the retrieved audio data.

29. The method of claim 28, wherein: the plurality of audio devices are associated with a plurality of participants; the retrieved audio data comprises the audio data from one or more participants of the plurality of participants; and the one or more participants are identified based on the user input.

30. The method of claim 23, wherein the audio data comprises a plurality of audio signals and the method further comprises: identifying an audio signal having a highest decibel level among the plurality of audio signals; identifying, from the plurality of audio devices, an audio device corresponding to a source of the audio signal having the highest decibel level; determining a signal strength of a message from the audio device by identifying identical sound characteristics in respective audio signals of the plurality of audio devices; and determining a score corresponding to a likelihood of a message reception for each audio device of the plurality of audio devices based on a respective signal strength.

31. The method of claim 30, further comprising: generating a communicative heatmap based on at least the score.

32. The method of claim 31, further comprising: determining position data of each audio device of the plurality of audio devices, wherein the communicative heatmap further comprises the position data, and wherein the position data comprises an orientation of a participant associated with each audio device of the plurality of audio devices.

33. The method of claim 32, wherein the position data is determined based on a WiFi signal strength, global positioning system (GPS) data, or video data corresponding to the audio data.

34. The method of claim 23, wherein the plurality of audio devices are associated with a plurality of participants of a sport activity, and the method further comprises: determining a communication metric for each participant of the plurality of participants based on the audio data; and generating a coaching instruction for one or more participants of the plurality of participants based on a respective communication metric.

35. The method of claim 34, wherein generating the coaching instruction comprises generating the coaching instruction based on video data corresponding to the audio data.

36. The method of claim 34, further comprising: transmitting the coaching instruction to the audio device associated with each participant of the one or more participants in a language corresponding to the language stored in a respective profile of the one or more participants.

37. The method of claim 23, further comprising: synchronizing respective audio data of the audio data received from the plurality of audio devices.

38. A computing system, comprising: one or more memories; and at least one processor each coupled to at least one of the memories and configured to perform operations comprising: acquiring audio data from a plurality of audio devices, wherein the plurality of audio devices are associated with a plurality of participants in a sport activity and each respective audio device is worn by each participant of the plurality of participants during the sport activity; generating communication data by analyzing the audio data using an artificial intelligence (Al) model; and generating a graphical representation of the communication data.

39. The computing system of claim 38,wherein the graphical representation comprises a visual representation of a player formation associated with the sport activity.

40. The computing system of claim 38, wherein the operations further comprise: detecting a communication from a participant of the plurality of participants; and determining a category associated with the communication; and wherein the communication data comprises at least a number of interactions for each category for each respective user of an audio device of the plurality of audio devices.

41. The computing system of claim 40, wherein the category is one an orientation communication, a positive communication, a negative communication, a stimulation communication, or a self-talk communication.

42. The computing system of claim 38, wherein the audio data includes a live audio stream, and the operations further comprise: detecting, in the live audio stream, a verbal command to tag an event; tagging a portion of the audio data based on the verbal command; synchronizing the tagged portion of the audio data with stored participant data; and generating the graphical representation based on the tagged audio data and the tagged event, wherein the graphical representation comprises video data corresponding to the audio data.

43. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: acquiring audio data from a plurality of audio devices, wherein the audio devices are associated with a plurality of participants in a sport activity and each respective audio device is worn by each participant of the plurality of participants during the sport activity; generating communication data by analyzing the audio data using an artificial intelligence (Al) model; and generating a graphical representation of the communication data.

44. The non-transitory computer-readable medium of claim 43, wherein the plurality of audio devices are associated with a plurality of participants of a sport activity; and wherein the graphical representation comprises a visual representation of a player formation associated with the sport activity.

45. The non-transitory computer-readable medium of claim 43, wherein the operations further comprise: detecting a communication from a participant of the plurality of participants; and determining a category associated with the communication, wherein the category is one of an orientation communication, a positive communication, a negative communication, or a stimulation communication, or a self-talk communication; and wherein the communication data comprises at least a number of interactions for each category for each respective user of an audio device of the plurality of audio devices.

46. A method, comprising: acquiring audio data from an audio device, wherein the audio device is associated with a participant in a sport activity and is worn by the participant during the sport activity; determining, using at least one computer processor, one or more biometrics of the participant based on the audio data; and determining a status of the participant based on the one or more biometrics or the audio data.

47. The method of claim 46, wherein determining the status of the participant further comprises: detecting that the participant is involved in an impact event based on at least the audio data; and determining an intensity of the impact event based on the one or more biometrics.

48. The method of claim 46, wherein the status of the participant is indicative of a health status of the participant, and wherein determining the status of the participant further comprises determining the status of the participant based on the one or more biometrics.

49. The method of claim 46, wherein the status comprises a fatigue level or a risk of concussion.

50. The method of claim 46, further comprising: acquiring motion data from a motion sensor associated with the audio device; and wherein determining the status of the participant further comprises determining the status of the participant based on the one or more biometrics and the motion data.

51. The method of claim 46, further comprising: determining the one or more biometrics based on a content of the audio data.

52. The method of claim 46, wherein the one or more biometrics comprise one or more of a heart rate and a respiratory pattern.

53. The method of claim 46, further comprising: determining a mental state based on the one or more biometrics; generating a coaching instruction based on the mental state; and outputting the coaching instruction to the audio device.

54. A computing system, comprising: one or more memories; and at least one processor each coupled to at least one of the memories and configured to perform operations comprising: acquiring audio data from an audio device, wherein the audio device is associated with a participant in a sport activity and is worn by the participant during the sport activity; determining one or more biometrics of the participant based on the audio data; and determining a status of the participant based on the one or more biometrics or the audio data.

55. The computing system of claim 54, wherein the operations further comprise:detecting that the participant is involved in an impact event based on at least the audio data; and determining an intensity of the impact event based on the one or more biometrics.

56. The computing system of claim 54, wherein the status of the participant is indicative of a health status of the participant, and wherein the operations further comprise: determining the status based on the one or more biometrics.

57. The computing system of claim 54, wherein the status comprises a fatigue level or a risk of concussion.

58. The computing system of claim 54, wherein the operations further comprise: acquiring motion data from a motion sensor associated with the audio device; and determining the health status based on the one or more biometrics and the motion data.

59. The computing system of claim 54, wherein the operations further comprise: determining the one or more biometrics based on a content of the audio data.

60. The computing system of claim 54, wherein the one or more biometrics comprise one or more of a heart rate and a respiratory pattern.

61. The computing system of claim 54, wherein the operations further comprise: determining a mental state based on the one or more biometrics; generating a coaching instruction based on the mental state; and outputting the coaching instruction to the audio device.

62. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:acquiring audio data from an audio device, wherein the audio device is associated with a participant in a sport activity and is worn by the participant during the sport activity; determining one or more biometrics of the participant based on the audio data; and determining a status of the participant based on the one or more biometrics or the audio data.

63. The non-transitory computer-readable medium of claim 62, wherein the operations further comprise: detecting that the participant is involved in an impact event based on at least the audio data; and determining an intensity of the impact event based on the one or more biometrics.

64. The non-transitory computer-readable medium of claim 62, wherein the status of the participant is indicative of a health status of the participant, and wherein the operations further comprise: determining the health status based on the one or more biometrics.

65. The non-transitory computer-readable medium of claim 62, wherein the status includes a fatigue level or a risk of concussion.

66. A method for triggering an action, comprising: acquiring data from an audio device, wherein the audio device is configured to be attached to a wearable element worn on a body of a user; determining, by at least one computer processor, a number of taps on the audio device by analyzing the data; and triggering an action based on the determined number of taps.

67. The method of claim 66, further comprising: associating a tag with a timestamp corresponding to when f the number of taps are detected.

68. The method of claim 67, wherein the action corresponds to a tagging action, and the method further comprises: activating a microphone of the audio device; capturing, using the microphone, a verbal tag; and associating the verbal tag with the timestamp of the determined number of taps.

69. The method of claim 66, wherein the determined number of taps is equal to two.

70. The method of claim 66, wherein the data corresponds to motion data detected by a motion sensor of the audio device.

71. The method of claim 70, further comprising: detecting an acceleration spike in the motion data by comparing the motion data to a threshold.

72. The method of claim 66, further comprising: detecting a first tap; detecting a second tap; determining a duration between the first tap and the second tap; and comparing the duration with a predetermined range to determine whether a double tap has occurred.

73. A computing system, comprising: one or more memories; and at least one processor each coupled to at least one of the memories and configured to perform operations comprising: acquiring data from an audio device, wherein the audio device is configured to be attached to a wearable element worn on a body of a user; determining, by at least one computer processor, a number of taps on the audio device by analyzing the data; and triggering an action based on the determined number of detected taps.

74. The computing system of claim 73, wherein the operations further comprise: associating a tag with a timestamp corresponding to when the number of taps are detected.

75. The computing system of claim 74, wherein the action corresponds to a tagging action, and the operations further comprise: activating a microphone of the audio device; capturing, using the microphone, a verbal tag; and associating the verbal tag with the timestamp of the determined number of taps.

76. The computing system of claim 73, wherein the determined number of taps is equal to two.

77. The computing system of claim 73, wherein the data corresponds to motion data detected by a motion sensor of the audio device.

78. The computing system of claim 77, wherein the operations further comprise: detecting an acceleration spike in the motion data by comparing the motion data to a threshold.

79. The computing system of claim 73, wherein the operations further comprise: detecting a first tap; detecting a second tap; determining a duration between the first tap and the second tap; and comparing the duration with a predetermined range to determine whether a double tap has occurred.

80. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising: acquiring data from an audio device, wherein the audio device is configured to be attached to a wearable element worn on a body of a user;determining a number of taps on the audio device by analyzing the data; and triggering an action based on the determined number of taps.

81. The non-transitory computer-readable medium of claim 80, wherein the operations further comprise: associating a tag with a timestamp corresponding to when the number of taps are detected.

82. The non-transitory computer-readable medium of claim 81, wherein the action corresponds to a tagging action, and the operations further comprise: activating a microphone of the audio device; capturing, using the microphone, a verbal tag; and associating the verbal tag with the timestamp of the determined number of taps.

83. The non-transitory computer-readable medium of claim 81, wherein the data corresponds to motion data detected by a motion sensor of the audio device.

84. The non-transitory computer-readable medium of claim 83, wherein the operations further comprise: detecting an acceleration spike in the motion data by comparing the motion data to a threshold.

85. The non-transitory computer-readable medium of claim 80, wherein the operations further comprise: detecting a first tap; detecting a second tap; determining a duration between the first tap and the second tap; and comparing the duration with a predetermined range to determine whether a double tap has occurred.

Citation Information

Patent Citations

  • Mobile portable conversation microphone

    EP4117305A1

  • Wireless multi-user audio system

    US20140112495A1

  • Wearable smart device for hazard detection and warning based on image and audio data

    US20160210834A1

  • Microphone assembly

    US20210160613A1

  • Wearable device with directional audio

    US20220095049A1