Electronic keyboard with enhanced display

The electronic keyboard system addresses the lack of visual feedback and posture monitoring by integrating a wide display and analysis of user movements, providing real-time feedback to enhance learning and improve playing techniques.

WO2025240332A1PCT designated stage Publication Date: 2025-11-20POLARO INC

Patent Information

Application Number
PCT/US2025/028921
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-13
Filing Date
2025-05-12
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Existing electronic keyboards lack effective visual feedback and user posture monitoring features, making it difficult for users to learn and improve their playing techniques.

Method used

An electronic keyboard system with a display extending at least as wide as the keyboard, equipped with cameras and processors to analyze user movements and provide real-time feedback, including visual and haptic cues, to enhance learning and improve playing techniques.

Benefits of technology

The system provides intuitive visual feedback and real-time monitoring, improving learning accuracy and user engagement by aligning visual animations with physical keys, and adjusting difficulty based on user metrics, thus enhancing the learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025028921_20112025_PF_FP_ABST
    Figure US2025028921_20112025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and processors associated with an electronic keyboard are described. An example system includes an electronic keyboard with a threshold number of keys extending across a first width; and a display positioned above the electronic keyboard, the display extending at least as wide as the first width. The display provides real-time visual prompts aligned with each key, including the partially covered edge keys, and hosts an Al-driven user interface. A camera may capture a player's posture, hand position, gaze, and facial expressions. Transformer-based models may process inputs to generate adaptive feedback, emotion-aware coaching, and / or gamified scoring. Haptic key actuators, synchronized LEDs, and audio cues reinforce the feedback, while on-device encrypted inference and differential-privacy may filter and secure all biometric data.
Need to check novelty before this filing date? Find Prior Art

Description

ELECTRONIC KEYBOARD WITH ENHANCED DISPLAYCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Prov. Patent App. No. 63 / 646050 filed on May 13, 2024 and titled “ELECTRONIC KEYBOARD WITH ENHANCED DISPLAY,” the disclosure of which is hereby incorporated herein by reference in its entirety.BACKGROUNDTECHNICAL FIELD

[0002] The present disclosure relates to an electronic keyboard, such as a musical keyboard, with an enhanced display. In some embodiments, the disclosure relates to an interactive, Al-powered keyboard learning system with integrated visual feedback and user posture monitoring features.DESCRIPTION OF RELATED ART

[0003] Electronic keyboards, such as musical keyboards, typically have a musical range that extends between 61, 72, 76, 88, and so on, keys. An electronic keyboard may output audio to a speaker, such as one built into the electronic keyboard or one which is in wired or wireless communication with the electronic keyboard. Practitioners may play music based on sheet music, such as sheet music printed out or presented via a separate device (e.g., a tablet).SUMMARY

[0004] According to some embodiments, a system is described. The system includes an electronic keyboard with a threshold number of keys extending across a first width; and a display positioned above the electronic keyboard, the display extending at least as wide as the first width.

[0005] In some embodiments, the display may be configured to visually indicate note positions across the full range of keys, even if the physical screen slightly exceeds or falls short of the keyboard's exact width, such that each key, including partially covered edge keys, is associated with a corresponding visual guide.

[0006] According to some embodiments, a learning system is described. The learning system is configured to provide educational feedback to a player of a musical instrument. The learning system includes one or more cameras configured to capture features of the player playing the instrument; a processor configured to process the features and determine measurements for evaluation criteria corresponding to proper techniques for playing the instrument; and one or more audio / visual outputs configured to provide feedback to the player responsive to the measurements of the evaluation criteria.

[0007] In some embodiments, such evaluation criteria may include posture, reaction time, emotional state, and hand articulation, computed using transformer-based machine learning models trained to interpret user movements and expressions.

[0008] According to some embodiments, an electronic learning game system is described. The electronic learning system is configured to provide instruction in playing a piano. The electronic learning game system includes an electronic keyboard with keys extending across a first width; and a display positioned above the electronic keyboard, the display extending at least as wide as the first width; a processor configured to output to the display graphical information to prompt a player to actuate specific one or more keys of the electronic keyboard; sensors configured to determine the actuation of the one or more keys; the processor configured compare a presentation of the graphical information to the timing and selection of the one or more keys actually actuated by the player and calculate a score responsive to the comparison.

[0009] In some embodiments, the system further includes, or implements, gamified features such as awarding badges, tracking combo streaks, and adjusting difficulty based on real-time scoring and physiological metrics.

[0010] According to some embodiments, a method is described. The method includes capturing, via one or more cameras, images of a user interacting with an electronic keyboard having a display extending at least as wide as the keyboard; analyzing the images using one or more machine learning models to determine one or more of user posture, facial expression, eye gaze, and hand positioning; detecting actuation of one or more keys on the electronic keyboard during a musical exercise; determining a performance score based on timing, articulation, and accuracy of the detected key actuations relative to visual prompts displayed on the display; updating the display in real-time to reflect instructional feedback,in eluding visual cues, score updates, and corrective overlays; providing haptic, audio, or visual feedback to guide user correction or encourage continued performance; and storing user metrics and emotional response indicators to dynamically adjust lesson difficulty and generate longitudinal progress data.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] This disclosure is described herein with reference to drawings of certain embodiments, which are intended to illustrate, but not to limit, the present disclosure. It is to be understood that the accompanying drawings, which are incorporated in and constitute a part of this specification, are for the purpose of illustrating concepts disclosed herein and may not be to scale.

[0012] Figure l is a block diagram of an example electronic keyboard.

[0013] Figure 2 illustrates an example user interface associated with teaching or instructing a user to play the piano or keyboard.

[0014] Figure 3 is a flowchart of an example process for

[0015] Figures 4A-4J illustrate example views of an electronic keyboard according to the disclosed technologyDETAILED DESCRIPTION

[0016] This application describes an electronic keyboard which may be used, for example, for instruction regarding playing, or to otherwise help a user to learn to play, the keyboard, piano, and so on. As will be described, the electronic keyboard may include a keyboard portion and a display which may be positioned above, or otherwise proximate to, the keyboard portion. For example, the keyboard portion and display may form, in some embodiments, the electronic keyboard (e g., the keyboard portion and display may be mechanically connected). In some embodiments, the display may be attached to the keyboard portion via a pivot or hinge. In some embodiments, the display extend from the keyboard portion (e.g., at an angle, such as an angle away from the keyboard portion). The display may also serve as the primary interface for feedback and instruction, enabling integration of dynamic, interactive learning content.

[0017] In some embodiments, the display may be at least as wide as the keyboard portion. For example, the display may extend at least as far as the extremities of the keys ofthe keyboard. As an example, the keyboard portion may have 88 keys, and the display may extend across all 88 keys. In some embodiments, the display (e.g., the screen, such as the portion of the display not including the bezel) may be substantially a 1-to-l width with the keyboard portion (e.g., a left edge of a first key may correspond with a left edge of the screen and a right edge of a last key may correspond with a right edge of the screen). However, in some configurations, the physical display width may slightly exceed or fall short of the keyboard width; in such cases, the display is still configured to provide full-range note indication, including visual cues aligned with partially covered keys.

[0018] As will be described, the display may present visual animations reflecting keys the user is to press, with the key presses corresponding to a particular song or music. An example visual animation, such as illustrated in Figure 2, may inform one or more of a specific key to be pressed, a time at which the key is to be pressed, length of time associated with the key press, articulation to be provided to the key, one or more pedal presses to be applied, and so on. Example articulation may include, in some embodiments, keyboard expression information (e.g., velocity sensitivity, aftertouch or pressure sensitivity, displacement sensitivity, and so on). Since the display may be, in some embodiments, at least as wide as the keyboard, a visual animation for an individual key may be presented above, or otherwise substantially proximate to, the individual key. For example, above may reflect that, from a user’s viewpoint, the displayed key appears to correspond to a key of the keyboard (e.g., the display key may appear to be in line with the key, appear to extend from it, appear such that if the display were closed the displayed key would be substantially over the key, and so on). In this way, the user may quickly ascertain which visual animation is associated with which key of the electronic keyboard. This alignment facilitates intuitive association between visual content and physical key locations, improving learning accuracy.

[0019] The electronic keyboard may include one or more cameras which may be used to monitor the user’s movements. The electronic keyboard, such as via one or more processors included in the electronic keyboard optionally in combination with a cloud-based server, may analyze the user’s movements to inform the user’s progress regarding learning of the keyboard or a particular musical piece. For example, the cameras may be used to monitor key presses such as whether the user is using proper articulation, proper timing, and so on. The cameras may additionally be used to track example information or metrics (e.g., measurementsfor evaluation criteria, also referred to herein as feedback) including the user’s posture, hand movements (e.g., fingering, wrist movements), elbow position, facial expressions, reaction times, and so on. For example, the example information or metrics may reflect whether the user is properly playing a piece of music which is being presented as visual animations via the display. In some embodiments, facial expression data may be analyzed to detect frustration, confusion, engagement, happiness, excitement, or sadness, which can be used to adapt learning feedback and monitor emotional wellness over time.

[0020] In some embodiments, the one or more cameras may be positioned within an upper portion of the display (e.g., above a screen of the display). For example, a camera may be positioned substantially in the middle of the display in the upper portion. In this example, the camera may have a focal length wide enough to encompass the left-most key and the right-most keys of the keyboard portion. Additional cameras may be positioned on the sides or bottom of the display to provide alternative viewing angles for more accurate tracking of posture, gaze, and hand position.

[0021] To monitor the user’s movements, the electronic keyboard may execute one or more machine learning models. For example, the electronic keyboard may include one or more processors (e.g., central processing unit, graphics processing unit, neural processing unit, application specific integrated circuits, and so on) which compute forward passes through the one or more machine learning models based on images or video from the cameras. Example machine learning models may include neural networks, such as convolutional neural networks, vision transformers or other attention-based networks, and so on. As an example, a machine learning model may analyze images received at a particular, or adjustable, frequency (e.g., 30 Hz, 60 Hz, 120 Hz). In some cases, a learnable task-token transformer model may be employed to handle multiple tasks (e.g., expression detection, posture estimation, hand pose analysis) simultaneously using a shared feature representation.

[0022] With respect to monitoring hand movements, the machine learning model may track the user’s hand (e.g., individual fingers) in substantially real-time. As an example, a user’s finger may be represented as one or more joints connected via one or more segments or bones. For example, the machine learning model may be trained to associate the user’s fingers with an underlying skeletal model which may be used to track the user. As another example, the machine learning model may monitor position information associated with thefingers which is tracked over time based on the received images. For example, the machine learning model may be trained to segment individual fingers or otherwise identify location information in images (e.g., bounding boxes about individual fingers). These detections may further inform metrics such as fluidity of motion, consistency of technique, or corrective action needed for specific fingering errors.

[0023] Based on monitoring the user’s fingers, the electronic keyboard can determine, for example, metrics reflecting the user’s reaction times, fingering accuracy, articulation consistent with a score or direction, and so on. The electronic keyboard, such as via the one or more processors described herein, may access sheet music or other information defining or informing a song. The electronic keyboard may thus determine whether the user is playing correct keys based on analyzing received presses of the keys (e.g., based on received audio signals, or based on information indicating which keys are being pressed) in comparison to the music. As may be appreciated, the sheet music may indicate articulation, or the sheet music maybe associated with metadata indicating articulation or other aspects. The electronic keyboard may thus determine metrics, or other information, regarding the user’s articulation. These metrics may be combined into a composite score and displayed in real-time alongside visual coaching suggestions or progress graphs.

[0024] In some embodiments, all data or information described herein (e.g., metrics, summary information, and so on) may be processed locally (e.g., via a local computer or processor included in the electronic keyboard). In some embodiments, the electronic keyboard may be in communication with a server or cloud system via a network (e.g., the Internet). The server or cloud system may analyze images or video from the one or more cameras and determine metrics or summary information associated with the user’s playing. For example, the electronic keyboard may execute one or more machine learning models to determine particular information (e.g., proper fingering). In this example, the server may aggregate information across different playing sessions via analyzing the particular information and / or the videos of the different playing sessions. This aggregated information may reflect summary information associated with the user’s playing. The electronic keyboard may present this summary information, or the information may be accessible via a website or mobile application. To protect user privacy, sensitive data such as facial images or biometricpatterns may be encrypted and processed with differential privacy safeguards or on-device only, without cloud upload.

[0025] As described herein, the electronic keyboard may include a keyboard portion and a display. For example, the electronic keyboard may represent a combination of the keyboard and a display. In this example, the combination may be included in a same overall package or unit. As an example, the display may be connected to the keyboard portion via a case or mount of the electronic keyboard. The one or more processors described herein may be similarly included in the same overall package or unit. In some embodiments, the keyboard portion may be inserted into a case or mount which is connected, or otherwise in communication with, a display. For example, the user may select a preferred keyboard for insertion into the case or mount. In some embodiments, the case or mount may have the display portion built therein and may include a receiver (e.g., a drawer, sliding portion, sliding keyboard tray, and so on) to receive the keyboard. As an example, the case or mount may include a sliding mechanism, such as drawer slides, ball bearing slides, along with runners (e.g., drawer runners). Such integrated casing may also include speakers, cooling systems, or acoustic chambers designed to enhance sound resonance and learning immersion.

[0026] In some embodiments, the display may be separate from the keyboard portion. As an example, the user may use the user’s display (e.g., a TV, a computer display, and so on). For this example, the one or more processors may be in communication with the display and render a user interface for presentation. In some embodiments, an initial setup process may be performed to ensure that visual animations are presented proximate to corresponding keys of the keyboard portion. The system may employ computer vision to automatically calibrate note-to-key alignment, accounting for varied keyboard layouts and screen dimensions.

[0027] An example setup process, for example performed by the one or more processors, may include the display presenting a visual animation directed to a particular key (e g., middle C) and the user adjusting the position of the display and / or providing user input to adjust the visual animation. Another example setup process may include use of one or more cameras. For example, the cameras may be used to ascertain the display’s position relative to the keyboard portion. As another example, the cameras may be used to determine a width associated with the display. The width may be used, by the processors, to inform positioningof the visual animations relative to the keyboard portion. In some embodiments, the display may depict a graphical representation of the user’s keyboard portion. For example, the visual animations may be depicted as falling towards respective keys which form the keyboard portion. This graphical mapping may include real-time overlays, LED synchronization, or gamified elements showing accuracy streaks and reaction time prompts.

[0028] In some embodiments, the electronic keyboard, such as the one or more processors included therein, may receive information from the keyboard portion. For example, the information may include MIDI information. As another example, the information may include touch-based information from sensors included in the keyboard portion. For example, each key may have a sensor which informs how far the key is depressed. The sensor may additionally inform a pressure or force applied to the key. This information may be relayed to the one or more processor to be used to inform instruction of the user. Additionally, this information may be aggregated with, or otherwise used in conjunction with, the camera data. For example, the one or more processors may determine fingering, articulation, and so on, based on visual data reflecting the user’s movements along with the sensor data reflecting the actual pressing of the keys. Haptic feedback motors embedded in the keys may also activate based on timing, pressure accuracy, or to prompt corrections via tactile cues.

[0029] In some embodiments, the keyboard portion may have lights (e.g., LEDs) positioned under, or otherwise proximate to, each key. The lights may activate based on the visual animations. For example, a visual animation may reflect an animation which is moving to a lower portion of the display (e.g., towards the keyboard portion). In this example, as a visual animation associated with a key gets closer to the display, the key may increase in brightness. The key may also have a particular color selected based on the visual animation. Thus, the lights may be synchronized with the visual animations. These lighting cues may also be adaptive based on user skill level, providing visual hints or rhythm-based flashes in gamified learning modes.

[0030] Figure 1 is a block diagram of an example electronic keyboard 100. The keyboard 100 has a display 110 and a keyboard portion 120. One or more processors (e.g., processor 140) may be included in the electronic keyboard 100. The display 110 may be a touchscreen display which extends a similar width to, or is at least as wide as, the keyboard portion 120. In some embodiments, the touchscreen display may extend a substantially samewidth as the keyboard portion 120 such that individual keys presented on the touchscreen display are a same, or substantially similar, width as the keys of the keyboard portion 120. In some embodiments, the display 110 may be a touch screen display. In some embodiments, the display 110 may be a transparent display (e.g., a transparent OLED display) which is optionally touch sensitive. The keyboard portion 120 may include a threshold number of keys (e.g., 61, 66, 72, 76, 88, and so on). The display 110 may also serve as the primary output surface for real-time feedback, lesson content, emotion indicators, and fitness-related overlays such as heart rate zones, depending on user engagement.

[0031] In the illustrated example, the display 110 is presenting a user interface. The user interface may be rendered by one or more processors included in the electronic keyboard 100. The illustrated user interface may represent, as one example, a home screen which includes different application or software interfaces. For example, a video chat application is included in the user interface. In this example, the video chat may be between a piano or keyboard teacher and a user of the keyboard 100. Video of the user (e.g., from camera 130) may be routed via a network to the teacher. In some embodiments, the teacher may additionally receive at least a subset of the above-described tracked information (e.g. user’s posture, hand movements, elbow position, facial expressions, reaction times). Additionally, the teacher may receive metrics or scores associated with the tracked information. The metrics may be presented, in some embodiments, via a teacher dashboard on a user interface of a user device (e.g., laptop, tablet, smart phone, on the electronic keyboard, and so on). In some configurations, the teacher dashboard may visualize trends in user emotion or reaction speed to inform remote instruction strategy.

[0032] As described above, the one or more processors of the electronic keyboard 100 may analyze images and / or video from the camera 130). For example, visual features may be analyzed to determine measurements or metrics associated with evaluation criteria. Example evaluation criteria may include proper techniques for playing the instrument (e g., keyboard) such as proper fingering, articulation, and so on as described above. In some embodiments, the processors may identify skeletal movement of the user’s fingers. For these embodiments, the processors may access information identifying proper fingering. As an example, the proper fingering may reflect a professional-level playing of a piece of music. The processors may compare the determined fingering with the accessed information to determinemeasures or metrics indicative of fingering. The processors may additionally monitor certain measurements or metrics related to proper fingering, such as speed, estimate of pressure, whether the hands or fingers are smoothly moving or are jumpily moving, timing associated with fingering, and so on. These analyses may be synthesized into composite proficiency scores or feedback prompts, such as highlighting repetitive tension in the wrist or suggesting targeted slow practice on unstable passages.

[0033] With respect to the above, feedback may be provided which is responsiveness to the metrics or measurements. For example, the feedback may reflect graphical indicators related to the above-described measurements or metrics. As another example, warnings or textual descriptions may be presented informing the user to take certain actions. For example, a textual description may be presented which informs the user to move his / her hands in a more fluid motion. The feedback may additionally be output as audio information for the user. In certain embodiments, haptic actuators embedded in specific keys may deliver tactile cues (e.g., short pulses) to reinforce timing correction or hand repositioning recommendations. Feedback may also include overlays or other information, such as related to posture, emotional state, eye gaze, articulation, and so on as described herein.

[0034] The one or more processors may additionally analyze images and / or video to determine measurements associated with posture. For example, the user’s back posture may be analyzed based on estimating metrics associated with the user’s back. Example metrics may include a measure indicating an angle of the back with respect to the chair, information indicating positions of the user’s shoulders, and so on. A skeletal model may additionally be applied to the user, for example identifying orientation information associated with a skeletal model based on the user’s posture depicted in images and / or video (e.g., the model may be adjusted or adapted based on the user’s posture). The processors may optionally compare measured posture metrics with stored posture data indicating proper posture associated with playing the keyboard. With respect to the skeletal model, the processors may access a stored ideal template and present alignment overlays to the users.

[0035] The processors may output an evaluation associated with posture. As an example, the processors may cause audio and / or visual output. For this example, the output may include descriptions or summaries of the user’s posture. The output may additionally correct deviations from proper posture (e.g., alignment overlays, for example with respect to askeletal model). For example, as the user plays and starts to slouch the processors may output audio and / or visual requesting that the user sit upright. In some cases, real-time alerts may also include graphical overlays on the display showing ideal pose alignment side-by-side with the user's detected pose.

[0036] The one or more processors may additionally determine reaction time criteria. Reaction time may be based on how timely the user presses a key, moves his / her hand to an appropriate position, and so on. Similar to the above, the one or more processors may output results informing metrics or measurements associated with reaction time. The one or more processors may additionally output information to correct deviations from proper reaction time (e.g., the user interface may indicate that the user should anticipate adjustments in hand and / or finger positioning). The one or more processors may additionally determine hand position criteria, such as monitoring hand position and determining whether the user’s hand is properly positioned or situated over the keyboard. Similar to the above, the one or more processors may output results of the evaluation to correct deviations from a proper hand position. The one or more processors may determine eye focus criteria which may inform where the user is looking. For example, the processors may monitor eye position and / or orientation and determine a focal point (e.g., based on an intersection associated with a vector extending from each eye). The processors may output audio and / or visual output to correct deviations from a proper eye focus (e.g., the processors may request that the user focus on a portion of the user interface, such as the portion of Figure 2 illustrating the visual animation 210). In some embodiments, adjustments to visual tempo pacing or display brightness may be made based on prolonged reaction delays or gaze aversion, helping retain attention and reduce fatigue.

[0037] The output may, in some embodiments, include output presented via an application executing on a smartphone or user device of the user (e.g., a tablet, wearable device, AR or VR device). For example, the user may have a tablet positioned near them which outputs feedback associated with the user’s playing. The external device may display synchronized metrics or emotional state summaries derived from the artificial intelligence models, providing a second-screen experience to support learning analytics.

[0038] While Figure 1 illustrates the display 110 as extending across the width of the keyboard portion 120, in some embodiments the display 110 may be formed from two ormore displays. For example, the user may use two touchscreen tablets or a combination of tablets, laptops, or other displays which are positioned next to each other. With respect to tablets, the tablets may execute applications which cause output of the user interface. The tablets ay similarly detect their proximity and adjust the user interface to substantially seamlessly extend across the keyboard portion. In distributed display configurations, coordination of visual animations may be governed by a central processor module ensuring accurate spatial mapping of musical content.

[0039] Additionally, while Figure 1 illustrates the keyboard portion 120 as being an electronic keyboard, in some embodiments an acoustic piano or keyboard may be used. The camera 130 may be used to track the user’s movements as described above. A microphone may also be used to obtain audio that may be analyzed by one or more processors to determine notes and / or articulation information being played (e.g., via a machine learning model, fast Fourier transform, and so on). With respect to cameras, the visual information from the camera may be analyzed, for example via machine learning model(s), to determine whether the user is properly playing a piece of music. For example, and as illustrated in Figure 2, visual animations may be rendered on the display 110 which inform the user’s playing of particular keys. In embodiments in which an electronic keyboard is used, the keyboard 100 may use output from the keyboard (e.g., MIDI output) in combination with the visual data to determine whether the user is properly playing the piece based on the visual animations. In embodiments in which an acoustic piano or keyboard is used, the keyboard 100 may use the visual data to effectuate the determination. In such acoustic configurations, the system may estimate fingerkey contact and velocity using visual and audio correlation models without requiring direct electronic key data.

[0040] Figure 2 illustrates an example user interface 200 associated with teaching or instructing a user to play the piano or keyboard. The user interface 200 may be presented via a display, such as display 120 of Figure 1. A visual animation 210 is included in the user interface 200 which is positioned above a particular key. The visual animation 210, in some embodiments, may move down the user interface 200 at a speed representing how soon the user is to play the particular key (e.g., individual keys represent individual musical notes, with the visual animation reflecting the musical notes to play). For example, the visual animation 210 may be used to prompt the user to actuate specific keys of the keyboard. The visualanimation 210 may be adjusted to inform articulation, note length, and so on. The one or more processors described herein may generate the visual animation 210 based on a piece of music being played (e.., based on sheet music, based on stored or accessed information reflecting the music, the animation 210, and so on). The user interface 200, in the illustrated example, additionally includes a video chat 220, for example a chat with a teacher. Gamified overlays may also be displayed, including combo counters, accuracy scores, or health-based challenges such as maintaining heart rate targets during play.

[0041] The user interface 200 may present graphical information associated with timing and / or selection of the keys actually actuated by the user. The graphical information may optionally include a score responsive to the user’s proper timing and / or selection of keys. Scores may be presented as progress arcs, badge unlocks, or musical achievement ranks that adapt to skill level, providing both short- and long-term motivational cues.

[0042] Figure 3 is a flowchart of an example process 300 for interaction with an electronic keyboard. For convenience, the process 300 will be described as being performed by a system of one or more processors (e.g., the electronic keyboard 100. In the illustrated process, a user may be playing a piece of music (e.g., sheet music, a song, and so on). The system may be outputting visual animations associated with the piece of music and be monitoring responses by the user (e.g., key presses, emotional responses, posture, and so on as described herein).

[0043] At block 302, the system obtains images of the user. As described above, and as illustrated in Figures 4A-4J, the electronic keyboard may include one or more cameras configured to obtain images or video of the user.

[0044] At block 304, the system analyzes images based on one or more machine learning (ML) models (e.g., artificial intelligence models). As described herein, the ML models may include neural networks, transformer-based models, and so on. The ML models may be used, at least in part, to determine user posture, facial expression, eye gaze, and / or hand positioning information.

[0045] At block 306, the system detects actuation of keys. The user may interact with the electronic keyboard, such as via actuation of the keys. The system may receive information reflecting the actuation, for example via a MIDI signal, audio signal, communication signal, and so on.

[0046] At block 308, the system determines performance score(s). As described above, the system may determine timing, articulation, and / or accuracy of the detected key actuations relative to visual prompts displayed on the display. For example, the visual prompts may include the visual animation illustrated in Figure 2.

[0047] At block 310, the system updates display and provides feedback to guide user corrections. The feedback may include visual cues, score updates, and / or corrective overlays. For example, the display may provide corrections regarding articulation, timing, posture, and so on. As one example, the system may provide posture feedback that includes comparing the user's real-time skeletal configuration to a stored ideal template and presenting alignment overlays. The display may additionally update a score associated with accurate playing of the piece of music. The system may additionally provide haptic, audio, or visual feedback to guide user correction or encourage continued performance.

[0048] At block 312, the system stores user metrics and emotional response indicators. The system may user the stored information to dynamically adjust lesson difficulty and generate longitudinal progress data. For example, the system may simplify the piece of music (e.g., access a simplified version) or may make the piece of music more complex (e.g., include more of the piece of music, such as greater harmonic and / or rhythmic complexity). With respect to emotional response, the system may adjust lesson content in response to detected emotional states, including frustration, happiness, or confusion. The system may also modulate the brightness, tempo, or visual density of the display based on the user's gaze focus or reaction time latency.

[0049] In the process 300, the system may process biometric data (e.g., images, extracted information reflecting emotional state, posture, eye gaze, and so on as described herein, on-device using privacy-preserving encrypted computation.

[0050] Figures 4A-4J illustrate example views of an electronic keyboard according to the disclosed technology. These figures may also reflect various hardware integrations, such as embedded camera placements, actuator layouts, and LED lighting arrays coordinated with the display system.

[0051] Figure 4A illustrates a top view of the electronic keyboard.

[0052] Figure 4B illustrates a bottom view of the electronic keyboard.

[0053] Figure 4C illustrates a left view of the electronic keyboard.

[0054] Figure 4D illustrates a right view of the electronic keyboard.

[0055] Figure 4E illustrates a front view of the electronic keyboard.

[0056] Figure 4F illustrates a back view of the electronic keyboard.

[0057] Figure 4G illustrates a first isometric view of the electronic keyboard.

[0058] Figure 4H illustrates a second isometric view of the electronic keyboard.

[0059] Figure 41 illustrates a third isometric view of the electronic keyboard.

[0060] Figure 4J illustrates a fourth isometric view of the electronic keyboard.Example EmbodimentsThe above-described electronic keyboard may include one or more of the following.Example Hardware Components may include one or more of:• 4K screen enclosed in a casing integrated with the electronic keyboard• Touch panel embedded into the screen o Can be transparent touchscreen, such as OLED touchscreen• Wall mount for the instrument, enabling an electronic cable with a screen to be mounted on the wall.• Security cables to secure the keyboard on the wall.• Screen mount allowing for tilt adjustment, integrated with the touchscreen and keyboard.• Integrated case combined with the keyboard.• Screen-to-keyboard ratio substantially 1 :1. (Indication on the screen that matches note to note)• Wide-angle camera mounted at the top of the screen, covering substantially of the entire keyboard. o In some embodiments, two or more cameras mounted at the top and / or sides of the screen.• LED lights under the piano keys, available in red or multicolor, complementing the screen.• Built-in motherboard / computer within the touchscreen.• Array of a threshold number of speakers (e.g., 4, 6, and so on) incorporated into the shell.• A threshold number of additional speakers inside the piano keyboard (e.g., 2, 3, 4).• USB-C output and audio headphone ports on the side of the keyboard.• Cooling system within the screen.• Audio amplifier housed within the display case.• Power supply adapter contained within the display case.• Direct data connection from the piano keyboard to the computer integrated in the screen case.• Processor(s) (e.g., GPU, NPU) dedicated to running one or more local large language models (LLM) and / or machine learning computer vision algorithms in the display casing. LLMs may also be accessed via networked requests to outside systems executing the LLMs.• Example machine learning algorithms may include neural networks, such as convolutional neural networks, attention-based networks (e.g., vision transformers), and so on.• Array of microphones integrated into the screen and / or the keyboard.• Heavy-duty stand for the piano with display.• Display case which may be designed to enhance speaker sound as a resonator.• Mount and case for the piano keyboard, also designed as a resonator for speakers.• Fingerprint sensor integrated into the piano keyboard.• Haptic motors installed on each key and within the piano shell.• Permanent magnets positioned on each key and on the moving plank inside the piano case.• Electromagnets installed on each key and on the moving plank inside the piano case.• Adjustable weight settings for piano keys.• Fully weighted piano keys designed to complement the touch display.• Semi-weighted piano keys in combination with the touch display.• Non-weighted piano keys paired with the touch display.• Hammer action piano keys in combination with the touch display.• Horizontal hammer puller mechanism for piano keys.• Multicolored LEDs placed at the top of the piano keys (under the display).• LED lights mounted at the bottom of the display (as flashlights).• Foldable touchscreen display integrated with the piano keyboard (laptop style).• Foldable touchscreen display that opens 180 degrees like a book, with an adjustable angle.• Integration of all specified technology into both acoustic and digital grand pianos.• Integration of all specified technology into acoustic upright pianos and digital pianos designed to visually mimic an acoustic piano or harpsichord.• Red or multicolored LED lights positioned at the root of each piano key, synchronized with screen visuals.• Multiple wide and narrow-angle cameras mounted at the top, bottom, and sides of the screen to cover the entire keyboard, the player, and surrounding area.• Depth cameras and sensors mounted at the top and bottom front of the screen.• Multiple ultrasound proximity sensors embedded in the piano keyboard and display to measure distances of hands, people, and objects relative to different parts of the keyboard and screen.• Fingerprint sensor integrated into the piano keyboard.• Actuators mounted on each key.• Capacitive touch layer integrated into each piano key.• Secure processing modules using encrypted inference techniques (e.g., homomorphic encryption or trusted execution environments) to analyze visual and biometric data without exposing raw inputs.• On-device differential privacy filters to anonymize learning metrics and facial emotion outputs for user protection.Example Software Components may include one or more of:• Two-way video call functionality, enabling real-time transmission of MIDI and / or audio data with latency below a threshold for sessions involving two or more participants.• MIDI data transmission over video.• Music instrument audio transmission over video.• An operating system tailored for music education, creation, and Music Al.• Artificial intelligence (Al) and computer vision technology using camera to track posture, fingering, wrist movements, elbow position, facial expressions, reaction times, and so on.• Software algorithms to monitor and score / adjust posture, fingering, wrist movements, elbow position, facial expressions, reaction times, and so on.• Al and computer vision for tracking hand movements to teach music conducting.• Software algorithms for tracking and teaching music conducting via hand movements.• Al and computer vision to modify synthesizer sounds and other parameters (e.g., modulation, velocity, and so on) by hand movements (e.g., hand gestures).• Software algorithms to adjust synthesizer sounds and settings (such as modulation, velocity, and other parameters) through hand movements.• Al MIDI model to generate music, co-compose with users, and provide play-along capabilities.• Al teacher assistant to monitor visual, vocal, and sensor data from the keyboard to suggest optimal music learning, creating, and performing strategies.• Voice-controlled musical instrument selection through a Voice Assistant.• Capability to select and load different synthesizers from the internet, with control elements displayed on the screen.• Music teaching tools showing visual cues on the screen for playing notes.• Music teaching using red or multicolored LED lights on the keyboard to indicate notes.• Music teaching through display of musical sheets / notations on the screen.• Interactive music teaching by displaying notes on the screen with guidance to play by ear, followed by practice and repetition.• Full-screen display of conductor scores on the screen.• Software that adjusts audio output based on microphone input and the acoustic environment.• Software that detects vocal pitch for singing, offering display guidance and teaching through gamification.• Al LLM model, or network access to a system computing forward passes through an LLM, to mix multiple music tracks, both in real-time and non-real-time• Utilization of one or more microphones integrated into the touchscreen case to offer guidance and lessons for learning, playing, recording the following instruments:• Vocals and singing.• Utilization of one or multiple cameras integrated into the touchscreen to facilitate guidance and lessons on learning, playing, and recording various acoustic and digital instruments:• Includes strings, woodwind, brass, drums, percussion, and pianos.• Vocals and singing.• Emotional state detection module capable of identifying multiple states including frustration, confusion, engagement, happiness, excitement, and sadness, to adapt teaching content and monitor learner wellness.• Gamification engine providing real-time scoring, level progression, badge rewards, and physiological challenges (e.g., maintain BPM + note accuracy) for enhanced user motivation.• Real-time posture correction overlay with skeletal visual comparisons and attention- aware display adjustments.• Real-time FER (facial expression recognition) integrated with Al scoring models to adjust instruction dynamically.• Automatic difficulty scaling based on cumulative user metrics, including tempo responsiveness, emotional feedback, and posture compliance.• Multi-user session handling with individualized biometric monitoring and progress tracking per participant.• Integration with third-party wearables via open API for supplemental heart rate or movement data, if desired by user.Other Embodiments

[0061] All of the processes described herein may be embodied in, and fully automated, via software code modules executed by a computing system that includes one or more computers or processors. The code modules may be stored in any type of non-transitorycomputer-readable medium or other computer storage device. Some or all the methods may be embodied in specialized computer hardware.

[0062] Many other variations than those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein can be performed in a different sequence or can be added, merged, or left out altogether (for example, not all described acts or events are necessary for the practice of the algorithms). Moreover, in certain embodiments, acts or events can be performed concurrently, for example, through multi-threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines and / or computing systems that can function together.

[0063] The various illustrative logical blocks, modules, and engines described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can include electrical circuitry configured to process computer-executable instructions. In another embodiment, a processor includes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computational engine within an appliance, to name a few.

[0064] Conditional language such as, among others, “can,” “could,” “might” or “may,” unless specifically stated otherwise, are understood within the context as used in general to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or steps. Thus, such conditional language is not generally intended to imply that features, elements and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements and / or steps are included or are to be performed in any particular embodiment.

[0065] Disjunctive language such as the phrase “at least one of X, Y, or Z,” unless specifically stated otherwise, is understood with the context as used in general to present that an item, term, etc., may be either X, Y, or Z, or any combination thereof (for example, X, Y, and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y, or at least one of Z to each be present.

[0066] Any process descriptions, elements or blocks in the flow diagrams described herein and / or depicted in the attached figures should be understood as potentially representing modules, segments, or portions of code which include one or more executable instructions for implementing specific logical functions or elements in the process. Alternate implementations are included within the scope of the embodiments described herein in which elements or functions may be deleted, executed out of order from that shown, or discussed, including substantially concurrently or in reverse order, depending on the functionality involved as would be understood by those skilled in the art.

[0067] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items. Accordingly, phrases such as “a device configured to” are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, “a processor configured to carry out recitations A, B and C” can include a first processor configured to carry out recitation A working in conjunction with a second processor configured to carry out recitations B and C.

[0068] It should be emphasized that many variations and modifications may be made to the above-described embodiments, the elements of which are to be understood as beingamong other acceptable examples. All such modifications and variations are intended to be included herein within the scope of this disclosure.

Claims

WHAT IS CLAIMED IS:

1. A system comprising: an electronic keyboard with a threshold number of keys extending across a first width; and a display positioned above the electronic keyboard, the display extending at least as wide as the first width; wherein the display is configured to visually indicate note positions across the full range of keys including partially covered keys, and to display feedback.

2. The system of claim 1, wherein the keys represent musical notes.

3. The system of claim 1, wherein the system includes one or more processors configured to present a user interface via the display.

4. The system of claim 3, wherein the user interface presents visual animations that visually describe playing of a musical piece.

5. The system of claim 4, wherein an individual visual animation informs playing of an individual key of the electronic keyboard.

6. The system of claim 5, wherein the individual visual animation is presented above the individual key.

7. The system of claim 3, further comprising one or more cameras.

8. The system of claim 7, wherein the one or more processors track portions of the user based on image data from the one or more cameras.

9. The system of claim 8, wherein the processors track portions based on one or more machine learning models, and wherein the machine learning models include vision transformers configured for facial expression and posture analysis.

10. The system of claim 1, wherein the display is a touchscreen display configured to receive user input.

11. A learning system configured to provide educational feedback to a player of a musical instrument, the learning system comprising: one or more cameras configured to capture features of the player playing the instrument; a processor configured to process the features and determine measurements for evaluation criteria corresponding to proper techniques for playing the instrument; and one or more audio / visual outputs configured to provide feedback to the player responsive to the measurements of the evaluation criteria.

12. The leaning system of Claim 1 1 , wherein the instrument comprises a keyboard.

13. The learning system of Claim 12, wherein the keyboard outputs haptic feedback via one or more keys of the keyboard, and wherein the haptic feedback is configured to provide tactile responses during performance or correction events.

14. The learning system of Claim 12, wherein the keyboard extends across a first width, and wherein the audio / visual output includes a display extending at least as wide as the first width.

15. The learning system of Claim 11, wherein one of the measurements includes a posture of the player and the processor: processes the features from the one or more cameras to determine measurements of the posture; determines posture evaluation criteria responsive to a comparison of the measurements of the posture to posture data; and outputs results of the posture evaluation criteria to the audio / visual output to correct deviations from a proper posture for playing a piano.

16. The learning system of Claim 15, wherein the output results include overlaying a live skeletal outline on the display next to an ideal posture template.

17. The learning system of Claim 11, wherein one of the measurements includes a timing of the player and the processor: processes a reaction time between a prompt on the audio / visual output and actuation of one or more keys of the keyboard; determines reaction time evaluation criteria responsive to a comparison of the reaction time to reaction time data; and outputs results of the reaction time evaluation criteria to the audio / visual output to correct deviations from a proper reaction time for playing a piano.

18. The learning system of Claim 11, wherein one of the measurements includes a hand position of the player and the processor: processes the features from the one or more cameras to determine measurements of the hand position; determines hand position evaluation criteria responsive to a comparison of the measurements of the hand position to hand position data; and outputs results of the hand position evaluation criteria to the audio / visual output to correct deviations from a proper hand position for playing a piano.

19. The learning system of Claim 11, wherein one of the measurements includes an eye focus of the player and the processor: processes the features from the one or more cameras to determine measurements of the eye focus; determines eye focus evaluation criteria responsive to a comparison of the measurements of the eye focus to eye focus data; and outputs results of eye focus evaluation criteria to the audio / visual output to correct deviations from a proper eye focus for playing a piano.

20. The learning system of Claim 11 , wherein the audio / visual output comprises one or more smartphones.

21. The learning system of Claim 11, wherein the audio / visual output comprises one or more display panels.

22. The learning system of Claim 11, wherein the keyboard comprises a non-electronic piano.

23. An electronic learning game system configured to provide instruction in playing a piano, the learning game system comprising: an electronic keyboard with keys extending across a first width; and a display positioned above the electronic keyboard, the display extending at least as wide as the first width; a processor configured to output to the display graphical information to prompt a player to actuate specific one or more keys of the electronic keyboard; sensors configured to determine the actuation of the one or more keys; the processor configured compare a presentation of the graphical information to the timing and selection of the one or more keys actually actuated by the player and calculate a score responsive to the comparison.

24. The system of claim 23, wherein the processor is further configured to determine gamification information, and wherein to determine the gamification information the processor is configured to calculate scores, track performance streaks, and provide visual or audio rewards based on real-time performance metrics.

25. The system of claim 23, wherein the one or more processors are further configured to analyze facial expression data to determine emotional states including frustration, confusion, happiness, sadness, and engagement.

26. The system of claim 23, wherein data from the one or more cameras is processed in a secure inference module that preserves privacy using encrypted computations or differential privacy techniques.

27. A method of providing interactive music instruction using an Al-powered keyboard system, the method comprising: capturing, via one or more cameras, images of a user interacting with an electronic keyboard having a display extending at least as wide as the keyboard; analyzing the images using one or more machine learning models to determine one or more of: user posture, facial expression, eye gaze, and hand positioning; detecting actuation of one or more keys on the electronic keyboard during a musical exercise; determining a performance score based on timing, articulation, and accuracy of the detected key actuations relative to visual prompts displayed on the display; updating the display in real-time to reflect instructional feedback, including visual cues, score updates, and corrective overlays; providing haptic, audio, or visual feedback to guide user correction or encourage continued performance; and storing user metrics and emotional response indicators to dynamically adjust lesson difficulty and generate longitudinal progress data.

28. The method of claim 27, further comprising adjusting lesson content in response to detected emotional states, including frustration, happiness, or confusion.

29. The method of claim 27, further comprising modulating the brightness, tempo, or visual density of the display based on the user's gaze focus or reaction time latency.

30. The method of claim 27, wherein posture feedback includes comparing the user's real-time skeletal configuration to a stored ideal template and presenting alignment overlays.

31. The method of claim 27, wherein biometric data is processed on-device using privacy-preserving encrypted computation.

Citation Information

Patent Citations

  • Keyboard system with multiple cameras

    US20140251114A1

  • Smart piano system

    US20200193950A1

  • Method, device, system and apparatus for creating and / or selecting exercises for learning playing a music instrument

    US20220172640A1

Cited By

  • Emotional teaching aids

    TWI934879B