Computer-implemented system and method for adaptive music learning

GB2704268APending Publication Date: 2026-08-26MELSONIC LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
GB2025001328
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-01-30
Publication Date
2026-08-26

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method for generating and outputting instructional interactions based on analyzed audio data associated with a user device. The audio data is received 302 and analyzed by a comp
Need to check novelty before this filing date? Find Prior Art

Description

FIELD OF INVENTION

[0001] The disclosure relates to data processing and more particularly, to computer-implemented systems and methods for adaptive music learning. BACKGROUND

[0002] Conventional music learning systems and software face several limitations that hinder their effectiveness and adaptability. Most existing solutions require users to manually set the tempo or speed of the piece being played before analysis. This approach restricts flexibility, as the software expects notes at predetermined intervals, simplifying note detection but failing to accommodate variations in tempo or speed. Furthermore, these systems cannot dynamically detect and compute the tempo of performance in real-time, which limits their applicability in diverse learning environments.

[0003] US patent application US20210065664A1 filed by Anssi Klapuri et al. discloses a system that allows for determining a user's skill level in a particular aspect (or characteristic) of musical proficiency, such as rhythmic and instrument technique aspect, as opposed to estimating merely a vague overall skill level. In other words, by determining the user's skill characteristics objectively not only will it become possible to identify the particular skill in which the user needs to improve himself or herself but it is also possible to provide the user with virtual exercise content that is targeted at improving the particular skills.

[0004] Russian patent application RU2690863C1 filed by KOREN MOREL et al. discloses a system and method for teaching and studying music notation, ear training, and singing using a computerized device in accordance with embodiments.

[0005] Another significant limitation of existing systems is their inability to detect both individual notes and chords simultaneously. When multiple frequencies are played together, such as a chord strummed on a guitar, conventional solutions struggle to isolate and identify the individual notes within the chord. This deficiency restricts the effectiveness of these systems in providing comprehensive feedback for complex musical performances.

[0006] Additionally, current systems rely on instruments being pre-tuned to provide accurate feedback, creating a barrier for users who may not have perfectly tuned instruments. The existing music learning platforms require users to tune their instruments in advance, making accurate feedback impossible when an instrument is out of tune. Moreover, these systems cannot accommodate variations in tuning or scale, limiting their utility when the same musical piece is played at different frequencies or scales.

[0007] Furthermore, existing solutions fail to address the needs of learners of Indian classical instruments, such as the veena. There is a lack of teaching applications capable of analyzing performances on such instruments and providing accurate technical feedback. This gap in the market underscores the need for an advanced, adaptive music learning system and method capable of overcoming these challenges. SUMMARY

[0008] According to an embodiment of the disclosure, a computer-implemented method for adaptive music learning is described. The computer-implemented method includes receiving, by a computer, audio data associated with a user device of an entity. The computer-implemented method further includes analyzing, by the computer, the audio data to determine a set of individual notes and a set of chords. The computer-implemented method further includes detecting, by the computer, the onset of music in the audio data by analyzing differences in amplitude and audio frequency. The computer-implemented method further includes generating, by the computer, a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data; and outputting, by the computer, the first set of instructional interactions and the second set of instructional interactions.

[0009] According to one or more embodiments of the disclosure, a computer system for adaptive music learning is described. The computer system includes a processor set, a computer-readable storage media, and program instructions that are stored on the one or more computer-readable storage media. The program instructions are executable by the processor set to cause the processor set to receive audio data associated with a user device of an entity. The program instructions further cause the processor set to analyze the audio data to determine a set of individual notes and a set of chords. The program instructions further cause the processor set to detect the onset of music in the audio data by analyzing differences in amplitude and audio frequency. The program instructions further cause the processor set to generate a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data. The program instructions further cause the processor set to output the first set of instructional interactions and the second set of instructional interactions.

[0010] According to one or more embodiments of the disclosure, a computer program product for adaptive music learning is described. The computer program product includes a computer-readable storage media having program instructions stored on the computer-readable storage media to perform operations. The operations include receiving audio data associated with a user device of an entity. The operations further include analyzing the audio data to determine a set of individual notes and a set of chords. The operations further include detecting the onset of music in the audio data by analyzing differences in amplitude and audio frequency. The operations further include generating a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data. The operations further include outputting the first set of instructional interactions and the second set of instructional interactions.

[0011] Additional technical features and benefits are realized through the process of the disclosure. Embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed subject matter. For a better understanding, refer to the detailed description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The following description will provide details of preferred embodiments with reference to the following figures wherein:

[0013] FIG. 1 is a diagram that illustrates a computing environment for adaptive music learning, in accordance with an embodiment of the disclosure;

[0014] FIG. 2 is a diagram that illustrates an environment for adaptive music learning, in accordance with an embodiment of the disclosure; and

[0015] FIG. 3 illustrates a flowchart that illustrates an exemplary computer-implemented method for adaptive music learning, in accordance with an embodiment of the disclosure.

[0016] FIG. 4 illustrates a flowchart that illustrates a method for adjusting instrument tuning, in accordance with an embodiment of the disclosure.

[0017] FIG. 5 illustrates a flowchart that illustrates a method for analyzing samples, in accordance with an embodiment of the disclosure. DETAILED DESCRIPTION

[0018] The proposed Al-based adaptive music learning system and method accelerate the mastery of musical instruments through an engaging and interactive practice experience. The system leverages computing devices to ensure accessibility across commonly used platforms and uses advanced digital signal processing and AI algorithms to analyze musical performances at an individual note level. The system provides real-time gamified feedback, making practice sessions both enjoyable and motivating for learners. The present system generates detailed improvement suggestions in a gamified format and summarizes progress and areas of improvement, enabling instructors to effectively guide students.

[0019] One advantage of the proposed system is its ability to detect tempo and onset without requiring manual setup, enabling seamless analysis of various musical pieces. Additionally, the system identifies tuning issues and dynamically adjusts feedback to ensure accurate tonal guidance, making it a robust and adaptive tool for music education.

[0020] Various aspects of the disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of the machine logic included in computer program product (CPP) embodiments. With respect to any flowcharts, depending upon the technology involved, the operations can be performed in a different order than what is shown in a given flowchart. For example, again depending upon the technology involved, two operations shown in successive flowchart blocks may be performed in reverse order, as a single integrated operation, concurrently, or in a manner at least partially overlapping in time.

[0021] A computer program product embodiment (“CPP embodiment” or “CPP”) is a term used in the disclosure to describe any set of one, or more, storage media (also called “mediums”) collectively included in a set of one, or more, storage devices that collectively include machine readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device that can retain and store instructions for use by a computer processor. Without limitation, the computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these mediums include diskette, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), static random-access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device (such as punch cards or pits / lands formed in a major surface of a disc) or any suitable combination of the foregoing. A computer-readable storage medium, as that term is used in the disclosure, is not to be construed as storage in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through a fiber optic cable, electrical signals communicated through a wire, and / or other transmission media. As will be understood by those of skill in the art, data is typically moved at some occasional points in time during normal operations of a storage device, such as during access, de-fragmentation, or garbage collection, but this does not render the storage device as transitory because the data is not transitory while it is stored.

[0022] FIG. 1 is a diagram that illustrates a computing environment 100 for adaptive music learning, in accordance with an embodiment of the disclosure. With reference to FIG. 1, there is shown a computing environment 100 that contains an example of an environment for the execution of at least some of the computer code involved in performing the disclosed methods, such as an adaptive music learning module 120B. In addition to the adaptive music learning module 120B, computing environment 100 includes, for example, a computer 102, a wide area network (WAN) 104, an end-user device (EUD) 106, a remote server 108, a public cloud 110, and a private cloud 112. In this embodiment of the disclosure, the computer 102 includes a processor set 114 (including a processing circuitry 114A and a cache 114B), a communication fabric 116, a volatile memory 118, a persistent storage 120 (including an operating system 120A and the adaptive music learning module 120B, as identified above), a peripheral device set 122 (including a user interface (UI) device set 122A, a storage 122B, and a network module 124. The remote server 108 includes a remote database 108A. The public cloud 110 includes a gateway 110A, a cloud orchestration module HOB, a host physical machine set 110C, a virtual machine set HOD, and a container set 110E.

[0023] The computer 102 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of a computer or a mobile device now known or to be developed in the future that is capable of running a program, accessing a network or querying a database, such as a remote database 108 A. As is well understood in the art of computer technology, and depending upon the technology, the performance of a computer-implemented method may be distributed among multiple computers and / or between multiple locations. On the other hand, in this presentation of the computing environment 100, detailed discussion is focused on a single computer, specifically the computer 102, to keep the presentation as simple as possible. Computer 102 may be located in a cloud, even though it is not shown in a cloud in FIG. 1. On the other hand, computer 102 is not required to be in a cloud except to any extent as may be affirmatively indicated.

[0024] The processor set 114 includes one, or more, computer processors of any type now known or to be developed in the future. The processing circuitry 114A may be distributed over multiple packages, for example, multiple, coordinated integrated circuit chips. The processing circuitry 114A may implement multiple processor threads and / or multiple processor cores. The cache 114B may be memory that is located in the processor chip package(s) and is typically used for data or code that should be available for rapid access by the threads or cores running on the processor set 114. Cache memories are typically organized into multiple levels depending upon relative proximity to the processing circuitry 114A. Alternatively, some, or all, of the cache 114B for the processor set 114 may be located “off-chip.” In some computing environments, the processor set 114 may be designed for working with qubits and performing quantum computing.

[0025] Computer readable program instructions are typically loaded onto the computer 102 to cause a series of operations to be performed by the processor set 114 of the computers 102 and thereby effect a computer-implemented method, such that the instructions thus executed will instantiate the methods specified in flowcharts and / or narrative descriptions of computer-implemented methods included in this document (collectively referred to as “the disclosed methods”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 114B and the other storage media discussed below. The program instructions, and associated data, are accessed by the processor set 114 to control and direct the performance of the disclosed methods. In computing environment 100, at least some of the instructions for performing the disclosed methods may be stored in the dynamic modification of the adaptive music learning module 120B in persistent storage 120.

[0026] The communication fabric 116 is the signal conduction path that allows the various components of computer 102 to communicate with each other. Typically, this fabric is made of switches and electrically conductive paths, such as the switches and electrically conductive paths that make up buses, bridges, physical input / output ports, and the like. Other types of signal communication paths may be used, such as fiber optic communication paths and / or wireless communication paths.

[0027] The volatile memory 118 is any type of volatile memory now known or to be developed in the future. Examples include dynamic type random access memory (RAM) or static type RAM. Typically, the volatile memory 118 is characterized by a random access, but this is not required unless affirmatively indicated. In computer 102, the volatile memory 118 is located in a single package and is internal to computer 102, but alternatively or additionally, the volatile memory 118 may be distributed over multiple packages and / or located externally with respect to computer 102.

[0028] The persistent storage 120 is any form of non-volatile storage for computers that is now known or to be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is being supplied to computer 102 and / or directly to the persistent storage 120. The persistent storage 120 may be a read-only memory (ROM), but typically at least a portion of the persistent storage 120 allows the writing of data, deletion of data, and re-writing of data. Some familiar forms of the persistent storage 120 include magnetic disks and solid-state storage devices. The operating system 120A may take several forms, such as various known proprietary operating systems or open-source Portable Operating System Interfacetype operating systems that employ a kernel. The code included in the adaptive music learning module 120B typically includes at least some of the computer code involved in performing the disclosed methods.

[0029] The peripheral device set 122 includes the set of peripheral devices of computer 102. Data communication connections between the peripheral devices and the other components of computer 102 may be implemented in various ways, such as Bluetooth connections, Near-Field Communication (NFC) connections, connections made by cables (such as universal serial bus (USB) type cables), insertion-type connections (for example, secure digital (SD) card), connections made through local area communication networks and even connections made through wide area networks such as the internet. In various embodiments of the disclosure, the UI device set 122A may include components such as a display screen, speaker, microphone, wearable devices (such as goggles and smartwatches), keyboard, mouse, printer, touchpad, game controllers, and haptic devices. Storage 122B is external storage, such as an external hard drive, or insertable storage, such as an SD card. The storage 122B may be persistent and / or volatile. In some embodiments of the disclosure, storage 122B may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments of the disclosure where computer 102 is required to have a large amount of storage (for example, where computer 102 locally stores and manages a large database) then this storage may be provided by peripheral storage devices designed for storing very large amounts of data, such as a storage area network (SAN) that is shared by multiple, geographically distributed computers.

[0030] The network module 124 is the collection of computer software, hardware, and firmware that allows computer 102 to communicate with other computers through WAN 104. The network module 124 may include hardware, such as modems or Wi-Fi signal transceivers, software for packetizing and / or de-packetizing data for communication network transmission, and / or web browser software for communicating data over the internet. In some embodiments of the disclosure, network control functions, and network forwarding functions of the network module 124 are performed on the same physical hardware device. In various embodiments of the disclosure (for example, embodiments that utilize software-defined networking (SDN)), the control functions and the forwarding functions of the network module 124 are performed on physically separate devices, such that the control functions manage several different network hardware devices. Computer-readable program instructions for performing the disclosed methods can typically be downloaded to computer 102 from an external computer or external storage device through a network adapter card or network interface included in the network module 124.

[0031] The WAN 104 is any wide area network (for example, the internet) capable of communicating computer data over non-local distances by any technology for communicating computer data, now known or to be developed in the future. In some embodiments of the disclosure, the WAN 104 may be replaced and / or supplemented by local area networks (LANs) designed to communicate data between devices located in a local area, such as a Wi-Fi network. The WAN 104 and / or LANs typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.

[0032] The EUD 106 is any computer system that is used and controlled by an end user (for example, a customer of an enterprise that operates computer 102) and may take any of the forms discussed above in connection with computer 102. The EUD 106 typically receives helpful and useful data from the operations of computer 102. For example, in a hypothetical case where computer 102 is designed to provide a recommendation to an end user, this recommendation would typically be communicated from the network module 124 of computer 102 through WAN 104 to END 106. In this way, the END 106 can display, or otherwise present recommendations to an end user. In some embodiments of the disclosure, END 106 may be a client device, such as a thin client, heavy client, mainframe computer, desktop computer, and so on.

[0033] The remote server 108 is any computer system that serves at least some data and / or functionality to the computer 102. The remote server 108 may be controlled and used by the same entity that operates the computer 102. The remote server 108 represents the machine(s) that collect and store helpful and useful data for use by other computers, such as the computer 102. For example, in a hypothetical case where the computer 102 is designed and programmed to provide a recommendation based on historical data, then this historical data may be provided to the computer 102 from the remote database 108A of the remote server 108. In one of the implementations, the remote server 108 is configured to process the data (audio file) from the student's user device by analyzing the performance and comparing it to a reference piece, thereby generating feedback information for the student and / or teacher. In an embodiment, the remote server 108 operates on a public cloud, and may also be hosted on the computer 102 of the present invention.

[0034] The public cloud 110 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer capabilities, especially data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically leverages the sharing of resources to achieve coherence and economies of scale. The direct and active management of the computing resources of the public cloud 110 is performed by the computer hardware and / or software of the cloud orchestration module HOB. The computing resources provided by the public cloud 110 are typically implemented by virtual computing environments that run on various computers making up the computers of the host physical machine set HOC, which is the universe of physical computers in and / or available to the public cloud 110. Virtual computing environments (VCEs) typically take the form of virtual machines from the virtual machine set 110D and / or containers from the container set 110E. It is understood that these VCEs may be stored as images and may be transferred among and between the various physical machine hosts, either as images or after the instantiation of the VCE. The cloud orchestration module HOB manages the transfer and storage of images, deploys new instantiations of VCEs, and manages active instantiations of VCE deployments. Gateway 110A is the collection of computer software, hardware, and firmware that allows public cloud 110 to communicate through WAN 104.

[0035] Some further explanation of virtualized computing environments (VCEs) will now be provided. VCEs can be stored as “images”. A new active instance of the VCE can be instantiated from the image. Two familiar types of VCEs are virtual machines and containers. A container is a VCE that uses operating-system-level virtualization. This refers to an operating system feature in which the kernel allows the existence of multiple isolated user-space instances, called containers. These isolated user-space instances typically behave as real computers from the point of view of programs running in them. A computer program running on an ordinary operating system can utilize all resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and devices assigned to the container, a feature which is known as containerization.

[0036] The private cloud 112 is similar to public cloud 110, except that the computing resources are only available for use by a single enterprise. While the private cloud 112 is depicted as being in communication with the WAN 104, in various embodiments of the disclosure, a private cloud may be disconnected from the internet entirely and only accessible through a local / private network. A hybrid cloud is a composition of multiple clouds of different types (for example, private, community, or public cloud types), often respectively implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technology that enables orchestration, management, and / or data / application portability between the multiple constituent clouds. In this embodiment of the disclosure, the public cloud 110 and the private cloud 112 are both part of a larger hybrid cloud.

[0037] In operation, the present computer-implemented system and method leverage physics, digital signal processing (DSP), and artificial intelligence (AI) to analyze and evaluate musical performances. The system applies these technologies through a computer program to process musical samples provided by a student and accurately identify what the student has played. This identification is performed by comparing the student’s performance with a digital music score generated from a reference piece, enabling the system to determine the correctness of the student's playing and provide targeted feedback.

[0038] The application of physics includes the use of Fourier transforms, Fast Fourier Transforms (FFT), Constant-Q Transforms (CQT), harmonic analysis, and musical scales to dissect the musical samples into their constituent frequencies and harmonics. The DSP techniques utilized are based on the open-source library, librosa, which facilitates audio analysis and processing. Additionally, the present system employs a novel application of a random forest AI algorithm specifically adapted for analyzing digital signals representing music. This adaptation is particularly effective for Indian classical instruments such as the veena, as well as Western instruments like the guitar, and has the potential to extend to other instruments. The AI algorithm is trained to detect individual notes, chords, tempo, and other musical elements with a high degree of accuracy.

[0039] The creation of a digital music score from a reference piece involves existing computer algorithms; however, the present system’s unique combination of technologies and their implementation sequence result in a highly effective approach to accurately identifying musical elements in real-time. The identification process includes the precise isolation of notes and chords, correction for untuned instruments, and detection of musical onset and tempo, all without requiring manual input. This enables the system to dynamically adjust and provide accurate feedback even under varying conditions.

[0040] Through extensive iterative development, the inventors have identified a specific sequence of processing steps and unique variables that contribute to the system’s effectiveness. This combination enables the system to provide detailed, gamified feedback to students in realtime, significantly enhancing their learning experience. The system also delivers progress summaries and improvement suggestions for instructors to guide students more effectively.

[0041] Although invented for music education, the underlying technology has potential applications in other domains requiring regular rhythm analysis, such as detecting heartbeats, analyzing financial trading market trends, and studying astrophysical phenomena like pulsars. The versatility and precision of the invention make it a powerful tool for a range of signal-processing tasks, making it a valuable addition to both music learning and broader fields of study.

[0042] FIG. 2 is a diagram that illustrates an environment for adaptive music learning, in accordance with an embodiment of the disclosure. FIG. 2 is explained in conjunction with elements from FIG. 1. With reference to FIG. 2, there is shown a diagram of a network environment 200. The network environment 200 includes a computer system (hereinafter referred to as system 202), and a user device 204. The system 202 further includes an AI model 202A, and a digital signal processing model 202B. There is further shown a database 210. The network environment 200 further includes entity 208 associated with the user device 204. Examples of entity 208 include but are not limited to a music learner and an instructor. The network environment 200 further includes the WAN 104 of FIG. 1. In an embodiment of the disclosure, the user device 204 may be an exemplary embodiment of the EUD 106. Similarly, the computer system 202 may be an embodiment of the computer 102 in FIG. 1.

[0043] The system 202 may include suitable logic, circuitry, interfaces, and / or code that may be configured for adaptive music learning. System 202 is configured to receive audio data associated with a user device of an entity. System 202 is further configured to analyze the audio data to determine a set of individual notes and a set of chords. System 202 is configured to detect onset of music in the audio data by analyzing differences in amplitude and audio frequency. System 202 is configured to generate a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data. System 202 is configured to output the first set of instructional interactions and the second set of instructional interactions. System 202 is configured to apply a digital signal processing algorithm on the analyzed audio data to identify the set of individual notes. System 202 is configured to remove one or more harmonics to determine a primary frequency of each note. System 202 is configured to detect the set of chords by training an artificial intelligence (AI) model with a set of chord samples processed through the digital signal processing algorithm. System 202 is configured to apply the AI model to identify the set of chords within a set of audio segments. System 202 is configured to calibrate an untuned instrument by analyzing a set of open string frequencies and adjusting a set of parameters for analysis. System 202 is configured to assign a set of expected frequencies for each note based on a scale such as a standard scale or a Carnatic music scale. System 202 is configured to detect the tempo and onset of musical pieces without requiring user input. System 202 is configured to detect and correct a set of non-musical interruptions (such as ambient noise) by filtering signals using autocorrelation and amplitude rise analysis.

[0044] In an embodiment, the examples of the computer system 202 may include, but are not limited to, a server, a computing device, a virtual computing device, a mainframe machine, a computer workstation, a smartphone, a cellular phone, a mobile phone, a gaming device, or a consumer electronic (CE) device.

[0045] The user device 204 may include suitable logic, circuitry, interfaces, and / or code that may enable users to interact with the system 202. Examples of the user device 204 may include, but are not limited to, a computing device, a mainframe machine, a server, a computer workstation, a smartphone, a cellular phone, a mobile phone, a gaming device, a consumer electronic (CE) device, a head-mounted device, a Virtual Reality (VR) Headset, an Augmented Reality (AR) Device, a Mixed Reality (MR) Device, a Projection-based System, and / or any other device with computer vision display capabilities.

[0046] The display screen may include suitable logic, circuitry, and interfaces that may be configured to render an output generated by the system 202. In some embodiments of the disclosure, the display screen may be an external display device associated with the user device 204. The display screen may be a touch screen which may enable entity 208A to interact via the display screen. The touch screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. In accordance with an embodiment of the disclosure, the display screen may refer to a display screen of a head-mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electrochromic display, or a transparent display. In some embodiments of the disclosure, the display screen may be realized through several known technologies such as, but are not limited to, at least one of a liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology.

[0047] In an embodiment, the AI model 202A may correspond to a computer-based system or software that exhibits characteristics commonly associated with human intelligence. The AI model 202A may be designed to perform tasks that typically require human intelligence, such as problemsolving, learning, reasoning, perception, understanding natural language, and decision-making. AI systems can range from simple rule-based programs to sophisticated, self-learning systems.

[0048] The AI model 202A may be a sophisticated piece of software that leverages machine learning processes to understand music.

[0049] In an embodiment, the digital signal processing model 202B is configured to execute the digital signal processing algorithm on the analyzed audio data to accurately identify a set of individual notes within the musical performance. This process may involve breaking down the audio signal into its fundamental components, such as pitch, amplitude, and duration, to detect and isolate each note. The digital signal processing algorithm may employ techniques like spectral analysis, time-domain processing, and frequency-domain transformations to ensure precise identification of notes, even in complex musical compositions or scenarios with overlapping tones.

[0050] In an embodiment, the database 210 corresponds to an organized collection of data that may be stored and accessed electronically from a computer system (such as the system 202). The database 210 is configured to manage, store, retrieve, and update data efficiently. In an exemplary implementation, the structure of the database 210 typically involves tables, records, and fields that can be managed through various database management systems (DBMS). Examples of database 210 include but are not limited to, a relational database, a Non-Structured Query Language (SQL) database, a hierarchical database, a network database, a transactional database, a data warehouse, and a distributed database. In an embodiment, the database 210 is configured to store the audio data 204A associated with the user device 204 which may include the operating system data, software application data, and software application version data. Further, database 210 stores instructional interactions that may be used to train the AI model 202A.

[0051] In operation, the system 202 is configured to receive the audio data 204A associated with the user device 204. The audio data 204A includes details such as, but not limited to, musical performance data, musical instruction data, musical instruments data, one or more software applications associated with the user device 204, hardware specifications associated with user device 204, and one or more additional settings that influence the functionality of the one or more software applications on the user device 204. Further, the user device 204 is associated with an entity 208. For instance, the user device 204 corresponds to a specific hardware or a software platform that the entity (say a music learner) employs to access musical sessions or assessments.

[0052] According to an embodiment herein, the present system and method analyze and evaluate musical performances using a combination of audio signal processing, artificial intelligence (AI), and machine learning algorithms. The system processes input audio signals from users and compares them to reference samples to provide accurate and personalized feedback on the performance. It is designed to cater to both Western and Indian musical instruments, including but not limited to the guitar and veena.

[0053] The system utilizes two primary inputs: reference samples and test samples. Reference samples are pre-recorded audio files of a piece of music generated by a teacher, coach, or musician. These samples are processed to create a Digital Representation of the Reference Sample (DRRS), which serves as the standard for comparison. Test samples are audio files or real-time audio streams generated by the student, which are similarly analyzed to create a Digital Representation of the Test Sample (DRTS).

[0054] To process these inputs, the system follows a multi-step methodology. For reference samples, it identifies rhythmic patterns, extracts component frequencies, segments the sample based on rhythmic patterns, frequencies, and amplitude, and identifies dominant frequencies in each segment. This information is used to generate the DRRS, a cadence-independent digital representation. Test sample processing mirrors this approach but aligns with the DRRS to ensure accurate comparison. Both samples are then compared by analyzing their respective digital representations to detect similarities and differences in timing, notes, chords, and amplitude.

[0055] The system outputs the DRRS and DRTS as digital representations for comparison. Feedback is presented to users in multiple formats. Students receive visual feedback on the correctness of their notes and chords, audio playback of their performance, and improvement scores. Teachers are provided with detailed performance logs and progress reports for their students. Additionally, the app offers actionable suggestions to help users improve specific areas of their performance.

[0056] In an embodiment, the system dynamically detects tempo variations. Unlike existing solutions, the present system allows students to play at varying speeds during practice while maintaining accurate analysis. The system uses Short-Time Fourier Transforms (STFT), Constant-Q Transforms (CQT), and harmonic filtering to detect single notes and chords. For chord detection, a random forest AI model trained on chord samples further enhances accuracy. The system also features automatic tuning calibration, analyzing open string frequencies to adapt to untuned instruments. For veena, it adjusts expected note frequencies based on the first note played, aligning them with Carnatic scales.

[0057] Feedback is tailored to the user's skill level. At the basic level, it identifies correct notes without considering timing or speed. Intermediate feedback includes timing accuracy, while advanced feedback evaluates speed and precision against the reference. The system is also designed to tolerate errors, such as missed or additional notes, without cascading inaccuracies across subsequent segments. To further support the learning process, the app logs and analyzes historical performances, tracking progress over time and offering specific improvement suggestions.

[0058] Accordingly, one advantage of the present invention is that it supports a dynamic learning process where students can practice at their own pace without adhering to fixed speed or timing constraints. Reference samples can be generated by any player, offering flexibility and independence. The DRRS and DRTS representations enable accurate analysis irrespective of tempo variations, making the system highly adaptable. Additionally, the system is optimized for a wide range of instruments, offering a comprehensive tool for musicians of diverse backgrounds.

[0059] Beyond music education, the system's capabilities extend to other fields. In medical signal analysis, it can analyze rhythmic patterns like heartbeats or blood flow. In engineering, it can monitor rhythmic signals such as bearing wear in motors. These potential applications demonstrate the versatility of the invention, making it a valuable tool across multiple domains.

[0060] FIG. 3 illustrates a flowchart 300 that illustrates an exemplary computer-implemented method for adaptive music learning, in accordance with an embodiment of the disclosure. FIG. 3 is explained in conjunction with elements of FIG. 1, and FIG. 2. With reference to FIG. 3, there is shown a flowchart 300. The operations of the exemplary method may be executed by any computing system, for example, by the computer 102 of FIG. 1 or the system 202 of FIG. 2. The operations of the flowchart 300 may start at 302. At 302, the audio data associated with a user device of an entity is received. In this step, the system captures musical performance data from a user device, such as a smartphone or tablet, using an integrated user interface or external input. This input may include live audio recorded during practice sessions or pre-recorded audio files. The received audio data serves as the foundation for further analysis by the system.

[0061] At 304, a set of open string frequencies is analyzed and adjusted for calibration of an untuned instrument. The system detects the frequencies of open strings played by the user and compares them against expected values for a tuned instrument. If the strings are out of tune, the system recalibrates its parameters to account for the discrepancies, ensuring accurate feedback despite tuning issues.

[0062] At 306, a set of expected frequencies is assigned for each note based on a scale. The system assigns expected frequencies for each note based on the musical scale being used. This step allows the system to analyze the performance accurately across various scales, such as Western or Indian classical music.

[0063] At 308, signals are filtered to detect and correct a set of non-musical interruptions. The system filters out non-musical sounds, such as coughing or ambient noise, that could interfere with the analysis. By detecting and correcting these interruptions, the system ensures accurate feedback and a smooth learning experience for the user.

[0064] At 310, differences in amplitude and audio frequency are analyzed to detect onsets in the music in the audio data. The system detects when music starts by monitoring rapid changes in the amplitude and frequency of the audio signal. These changes indicate the beginning of a musical note or phrase. The system uses this capability to ignore non-musical sounds or pauses, ensuring precise analysis.

[0065] At 312, a digital signal processing algorithm is applied to the analyzed audio data to identify the set of individual notes. The system applies DSP techniques such as Short-Time Fourier Transform (STFT) and Constant-Q Transform (CQT) to isolate and analyze individual notes from the performance data. These methods identify the amplitude and frequency characteristics of each note in real-time.

[0066] At 314, none of the harmonics or one or more harmonics are removed to determine a primary frequency of each note. To ensure accurate identification, the system filters out harmonics (overtones) that are naturally generated alongside the fundamental frequency of a note. By isolating the primary frequency, the system ensures precise note detection.

[0067] At 316, the audio data is analyzed to determine a set of individual notes and a set of chords. The system processes the audio data using advanced digital signal processing (DSP) algorithms. These algorithms identify the fundamental frequencies within the audio data to differentiate between individual notes and chords. This analysis forms the basis for providing feedback and identifying performance accuracy. Examples of the DSP algorithms include but are not limited to Short-Time Fourier Transform (STFT) and Constant-Q Transform (CQT).

[0068] At 318, the AI model is applied to identify the set of chords within a set of audio segments. Using the trained AI model, the system analyzes audio segments from the user's performance to detect chords. This step enables the system to provide feedback on both individual notes and chords simultaneously. At 320, the sets of chords, notes, and timings generated by the analysis are compared with the corresponding set of chords, notes, and timings from a reference piece. This comparison involves evaluating the alignment of the musical elements, including the harmonic structure (chords), the individual musical pitches (notes), and their temporal arrangement (timings), to identify similarities, differences, or deviations.

[0069] At 322, tempo and onset of musical pieces are detected without requiring a user input. The system automatically identifies the tempo (speed) and onset of musical pieces by analyzing the timing and sequence of notes. This eliminates the need for manual input, enabling seamless analysis of performances at any tempo.

[0070] At 324, a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data are generated. The system generates two categories of feedback: 1. First set of instructional interactions includes real-time gamified feedback data designed to motivate and engage the learner. 2. Second set of instructional interactions includes detailed improvement suggestions and performance summary data for learners and instructors, highlighting areas for improvement and progress tracking.

[0071] At 326, the first set of instructional interactions and the second set of instructional interactions are outputted. The generated feedback is delivered to the user through the interface of the device. This output may include visual representations, textual feedback, or audio cues, allowing the learner to understand their performance and improve effectively.

[0072] In an embodiment, an artificial intelligence (AI) model is trained with a set of chord samples processed through the digital signal processing algorithm to detect the set of chords. The system trains the AI model using a large dataset of pre-recorded chord samples. Each sample undergoes DSP analysis to extract key characteristics, which are then used to build a robust AI model capable of recognizing various chords.

[0073] According to an embodiment herein, the method includes a step of summarizing progress and areas for improvement using graphical and textual reports accessible to both the learner and their instructor. The method includes a step of enabling seamless integration across various commonly used devices to ensure accessibility for all users. Examples of gamified feedback include real-time visual cues to indicate correct and incorrect notes; and interactive elements to motivate and engage the user during practice sessions. The method includes a step of generating tonal feedback dynamically adjusted to the detected tuning of the instrument; and ensuring feedback accuracy even when the instrument is tuned to different frequencies or scales.

[0074] FIG. 4 illustrates a flowchart depicting a method 400 for adjusting instrument tuning in accordance with an embodiment of the disclosure. The method 400 begins with receiving audio data associated with ‘open strings’ from a user device of an entity (step 402). The onset in the music in the audio data is then detected by analyzing variations in amplitude and frequency (step 404). The digital signal processing algorithm is applied to the analyzed audio data to identify individual notes (step 406). To determine the primary frequency of each note, zero, or one or more harmonics are removed (step 408).

[0075] Method 400 further includes a step 410 of calibrating an untuned instrument by analyzing the open string frequencies and adjusting parameters for further analysis. Expected frequencies for each note are assigned based on a predefined scale (step 412), and the calibration data is stored for future analyses of pieces played on the instrument with the same tuning (step 414).

[0076] FIG. 5 illustrates a flowchart depicting a method 500 for analyzing musical samples in accordance with an embodiment of the disclosure. The method 500 starts by receiving audio data from a user device (step 502). Non-musical interruptions in the audio data are detected and corrected by filtering the signals (step 504). The onsets in the music are detected by analyzing variations in amplitude and frequency (step 506). The audio data is then analyzed to identify individual notes and chords using previously stored calibration data for the instrument's tuning (step 508). A digital signal processing algorithm is applied to the analyzed audio data to identify individual notes (step 510), while zero, or one or more harmonics are removed to determine the primary frequency of each note (step 512). Chords are detected by training an AI model with processed chord samples (step 514), and the trained AI model is then applied to identify chords within the segmented audio data (step 516). Lastly, the tempo and onset of musical pieces are detected without requiring user input (step 518). This comprehensive method ensures accurate tuning, note identification, and analysis, enhancing the overall music learning and performance experience.

[0077] According to one or more embodiments of the disclosure, a computer program product for adaptive music learning is described. The computer program product includes a computer-readable storage media having program instructions stored on the computer-readable storage media to perform operations. The operations include receiving audio data associated with a user device of an entity. The operations further include analyzing the audio data to determine a set of individual notes and a set of chords. The operations further include detecting the onset of music in the audio data by analyzing differences in amplitude and audio frequency. The operations further include generating a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data. The operations further include outputting the first set of instructional interactions and the second set of instructional interactions.

[0078] Thus, the present system and method provide a significant advancement in music education technology. By combining cutting-edge signal processing and AI, it delivers an adaptive, precise, and personalized learning experience. The present system’s flexibility and comprehensive features set it apart from existing systems, offering transformative benefits for students, educators, and professionals alike.

[0079] The descriptions of the various embodiments of the disclosure have been presented for purposes of illustration but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

What is claimed is:

1. A computer-implemented method, comprising:receiving, by a computer, audio data associated with a user device of an entity;analyzing, by the computer, the audio data to determine a set of individual notes and a set of chords;detecting, by the computer, onset of music in the audio data by analyzing differences in amplitude and audio frequency;generating, by the computer, a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data; andoutputting, by the computer, the first set of instructional interactions and the second set of instructional interactions.

2. The computer-implemented method of claim 1, further comprising:applying, by the computer, a digital signal processing algorithm on the analyzed audio data to identify the set of individual notes; andremoving, by the computer, one or more harmonics to determine a primary frequency of each note.

3. The computer-implemented method of claim 1, further comprising:detecting, by the computer, the set of chords by training an artificial intelligence (AI) model with a set of chord samples processed through the digital signal processing algorithm; andapplying, by the computer, the AI model to identify the set of chords within a set of audio segments.

4. The computer-implemented method of claim 1, further comprising:calibrating, by the computer, an untuned instrument by analyzing a set of open string frequencies and adjusting a set of parameters for an analysis; andassigning, by the computer, a set of expected frequencies for each note based on a scale.

5. The computer-implemented method of claim 1, further comprising:detecting, by the computer, tempo, and onset of musical pieces without requiring a user input; anddetecting and correcting, by the computer, a set of non-musical interruptions by filtering signals.

6. A computer system, comprising:a processor set;one or more computer-readable storage media; andprogram instructions stored on the one or more computer-readable storage media, the program instructions executable by the processor set to cause the processor set to:receive audio data associated with a user device of an entity;analyze the audio data to determine a set of individual notes and a set of chords;detect onset of music in the audio data by analyzing differences in amplitude and audio frequency;generate a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data; andoutput the first set of instructional interactions and the second set of instructional interactions.

7. The computer system of claim 6, wherein the program instructions further cause the processor set to:apply a digital signal processing algorithm on the analyzed audio data to identify the set of individual notes; andremove one or more harmonics to determine a primary frequency of each note.

8. The computer system of claim 6, wherein the program instructions further cause the processor set to:detect the set of chords by training an artificial intelligence (AI) model with a set of chord samples processed through the digital signal processing algorithm; andapply the AI model to identify the set of chords within a set of audio segments.

9. The computer system of claim 6, wherein the program instructions further cause the processor set to:calibrate an untuned instrument by analyzing a set of open string frequencies and adjusting a set of parameters for an analysis; andassign a set of expected frequencies for each note based on a scale.

10. The computer system of claim 6, wherein the program instructions further cause the processor set to:detect tempo and onset of musical pieces without requiring a user input; anddetect and correct a set of non-musical interruptions by filtering signals.

11. A computer program product for adaptive music learning, the computer program product comprising:one or more computer-readable storage media; andprogram instructions stored on the one or more computer-readable storage media to perform operations comprising:receiving audio data associated with a user device of an entity;analyzing the audio data to determine a set of individual notes and a set of chords;detecting the onset of music in the audio data by analyzing differences in amplitude and audio frequency;generating a first set of instructional interactions and a second set of instructional interactions based on the analyzed audio data; andoutputting the first set of instructional interactions and the second set of instructional interactions.

Citation Information

Patent Citations

  • ViewWO2021/190660A1onEspacenetopensinnewtab

  • ViewUS20160253915A1onEspacenetopensinnewtab

  • ViewUS20110247479A1onEspacenetopensinnewtab

  • ViewUS20180122260A1onEspacenetopensinnewtab

  • ViewUS20120266738A1onEspacenetopensinnewtab