Intelligent musical instrument, terminal and music processing method
The interactive sheet music and AI model that combine smart musical instruments with terminals solve the problems of high learning barriers and insufficient real-time feedback in traditional musical instruments. Through real-time data collection and music theory optimization, it enables efficient practice and creation, breaking through the limitations of time and space.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-03-17
AI Technical Summary
Existing music education and performance methods have high learning barriers. Traditional instrument learning relies on music theory knowledge, and the practice process is limited by time and space, lacking real-time feedback and personalized guidance, making it difficult for users to create complete musical works.
By combining smart musical instruments with terminals, introducing interactive sheet music and AI models, real-time collection of performance actions, providing instant feedback and optimization suggestions, and using AI models for music theory optimization and creative expansion, a closed loop of 'performance—feedback—review—optimization—expansion' is formed.
It lowers the barrier to music learning, breaks through the limitations of time and space, enables users to practice and create efficiently in multiple scenarios, provides real-time feedback and music theory optimization, and helps users create complete musical works.
Smart Images

Figure CN121686984A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent audio processing technology, and in particular to intelligent musical instruments, terminals, and music processing methods. Background Technology
[0002] Current music education and performance methods present a high learning threshold. Firstly, learning traditional instruments often relies on a strong foundation in music theory, such as sight-reading ability, rhythmic understanding, and tonality comprehension. For beginners, the inability to read music, insufficient rhythmic sense, and difficulty finding a suitable tonality are major obstacles in the initial stages. Furthermore, traditional wind or keyboard instruments typically require complex fingerings and strong lung capacity, resulting in long learning periods and a tedious practice process, making it difficult for many users to persevere.
[0003] Secondly, music practice is often limited by time and space. Traditional instruments produce relatively loud noise during practice, which can easily disturb neighbors, requiring a private practice environment and ample practice time. In reality, most users find it difficult to balance work, study, and long hours of music training, further raising the barrier to entry for music learning.
[0004] Furthermore, although artificial intelligence technology has made some progress in music generation and education in recent years, such as automatically generating sheet music or accompaniment through algorithms, most of these achievements remain at the software level and lack interactive methods that integrate with the actual performance process. The sheet music generated or obtained by users lacks effective hardware interfaces and intuitive performance guidance, making it impossible to directly verify and experience the musical performance. At the same time, real-time guidance and personalized feedback during performance are also insufficient, making it difficult for users to obtain timely optimization suggestions regarding rhythm, timbre, or performance style.
[0005] Therefore, there is an urgent need for a method that combines simplified performance techniques, dynamic score guidance, and artificial intelligence assistance to lower the learning threshold, overcome time and space limitations, and provide users with real-time support for creation and performance. Summary of the Invention
[0006] The main purpose of this invention is to provide intelligent musical instruments and music processing methods, aiming to lower the entry barrier for music learning, so that even users who lack music theory knowledge can play notes through simplified operation and gradually build music cognition in the process. Furthermore, it breaks through the time and space limitations of the practice process, providing a more convenient music learning and performance method suitable for multiple scenarios; By effectively combining the score with the performance process, dynamic guidance of score information (notes, rhythm, dynamics, etc.) and real-time correspondence between performance actions can be achieved. With the help of artificial intelligence technology, learning and creation can be assisted, including music score recognition, humming to music score conversion, music theory optimization and expansion, so that users' inspiration and practice content can be transformed into complete musical works; Provide real-time feedback and optimization suggestions during the performance to compensate for the lack of teacher guidance and improve users' learning efficiency and performance experience.
[0007] To achieve the above objectives, the present invention provides a music processing method, the method comprising: The system receives music information and converts the music signal into an interactive score. The interactive score includes score information and a performance prompt signal associated with the score information. The performance prompt signal is used to prompt the performer to perform the performance action corresponding to the score information at the corresponding time. The interactive score provides performance output prompts to guide the performer to practice playing according to the prompts and record the performance content. The performance content is optimized based on a music AI model, resulting in an optimized performance.
[0008] In one embodiment of the present invention, receiving music information and converting the music signal into an interactive score includes: Receive music information and determine the format of the music information; Given that the music information is in the format of an audio file, audio decoding technology is used to analyze the audio features in the audio signal; Given that the music information is in the format of an image file, audio features in the sheet music image are obtained based on image recognition technology; Given that the music information is in the format of a sheet music file, identify the audio features in the sheet music file; Generate interactive musical scores based on audio features.
[0009] In one embodiment of the present invention, the audio features include notes, pitch, and timbre; The specific steps for generating interactive musical scores based on audio features are as follows: The prompting time of the prompt signal is determined based on the timestamp of the target note; The prompting content is determined based on the pitch and timbre of the target note; Establish a relationship between the prompt content and the timestamp of the target note.
[0010] In one embodiment of the present invention, the method further includes: Receives ambiguous instructions from the performer; Based on a music AI model, the fuzzy instructions input by the performer are converted into audio features, and the audio features are then converted into interactive musical scores. The method also includes converting the optimized performance content into interactive sheet music.
[0011] In one embodiment of the present invention, the optimization of the performance content based on the music AI model to obtain the optimized performance content includes: The audio signal is preprocessed to obtain a clean audio signal, and the preprocessed audio signal is input into an AI model for translation to convert the clean audio signal into music data that conforms to music grammar. The translation includes: The fundamental frequency of a sound signal is converted into pitch characteristics, and the time-domain duration of a sound signal is converted into note value. Note segmentation based on the continuity of audio data; Constructing musical phrase boundaries based on breakpoint detection; Melody contour extraction based on waveform analysis; Tonalities and chords are constructed based on a music grammar rule base.
[0012] In one embodiment of the present invention, the method further includes: Record performance data during the performance; After the performance, the performance process data is compared with the interactive score to generate review feedback information.
[0013] Furthermore, to achieve the above objectives, the present invention also discloses a terminal, comprising: Processor, memory, touchscreen, and communication module; The touch screen includes a display module and a touch module; The display module is used to display performance prompts and performance content; The touch module is used to perform performance operations using touch mode according to the performance prompt signal; The performance prompt signals are displayed on the display module at preset time intervals; The communication module is used to enable the terminal to establish a communication connection with the smart musical instrument; The memory and processor, and a music processing program stored on the memory and executable on the processor, the music processing program being configured to implement the steps of the music processing method as claimed in any one of claims 1 to 6.
[0014] In one embodiment of the present invention, the processor further includes: Music score recognition module; interactive music score generation module; The music score recognition module includes recognizing the format of music signals and issuing instructions corresponding to the format. The interactive music score generation module converts the music signal into an interactive music score according to the instructions. The music score recognition module and the interactive music score generation module are implemented by the processor executing the corresponding music processing program.
[0015] In one embodiment of the present invention, the memory further includes: Used to store performance content and interactive sheet music.
[0016] This invention also discloses an intelligent musical instrument, comprising: The main body of the instrument, used for playing practice; The control module, located in the main body of the instrument, is used to process the sound signals collected by the pickup module and control the prompting module to provide prompts based on the prompting signals fed back by the terminal. A sound pickup module, located on the main body of the instrument and connected to the control module, is used to pick up sound signals; An interactive module, located on the main body of the musical instrument and connected to the control module, is used to convert the sound signal and the content of the performance practice into a music signal and send it to the terminal based on the control of the control module, and is also used to receive prompt signals from the terminal. A prompting module is located on the main body of the instrument and connected to the control module, and is used to provide prompts based on the control of the control module in response to received prompting signals; The feedback module is located in the main body of the instrument and has a vibration motor inside that emits vibrations of corresponding frequency and amplitude in response to the performance content. A communication module, located in the main body of the instrument, is used to enable the smart instrument to establish a communication connection with the terminal. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of a music practice method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the process of creating interactive musical scores using an AI model, according to another embodiment of the present invention. Figure 3This is a flowchart illustrating a method for converting audio signals into interactive musical scores using an AI model, according to another embodiment of the present invention. Figure 4A This is a schematic diagram of a terminal performance interface and performance prompt signals according to another embodiment of the present invention; Figure 4B This is a schematic diagram illustrating the functional area division of the terminal performance interface according to another embodiment of the present invention. Figure 5 Here is a flowchart of a method for music theory optimization using fuzzy instructions combined with an AI model, according to another embodiment of the present invention; Figure 6 This is a schematic diagram of the system architecture of a terminal according to another embodiment of the present invention; Figure 7A This is a schematic diagram of the operation method of an intelligent musical instrument according to another embodiment of the present invention; Figure 7B This is a schematic diagram of the external structure of a smart musical instrument according to another embodiment of the present invention; Figure 8 This is a schematic diagram of the system architecture of an intelligent musical instrument according to another embodiment of the present invention; Figure 9A This is a schematic diagram of the first internal structure of an intelligent musical instrument according to another embodiment of the present invention; Figure 9B This is a schematic diagram of the second internal structure of a smart musical instrument according to another embodiment of the present invention; Figure 10 This is a functional schematic diagram of an intelligent musical instrument according to another embodiment of the present invention.
[0020] Explanation of icon numbers:
[0021] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Well-known modules, units, and their connections, links, communications, or operations are not shown or described in detail. Furthermore, the described features, architectures, or functions can be combined in any way in one or more embodiments. Those skilled in the art should understand that the various embodiments described below are only for illustrative purposes and not for limiting the scope of protection of the present invention. It is also readily understood that the modules, units, or processing methods in the various embodiments described herein and shown in the accompanying drawings can be combined and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] The definitions of various terms or methods used in the following embodiments are, except where logically impossible, generally defined as broad concepts that can be implemented under the premise of the content disclosed in the embodiments. Under this understanding, all specific subordinate limitations of the terms or methods should be considered as part of the invention and should not be narrowly interpreted or biased simply because the specification does not disclose such a specific limitation. Similarly, provided that it is logically feasible, the order of the steps in the method is flexible and varied, and all specific subordinate limitations in the broad concepts of various terms or methods fall within the scope of protection of this invention.
[0024] Because existing technologies lack intuitive prompts for users during performance, it is difficult to master rhythm and pitch, resulting in a high learning curve. Even with electronic sheet music displays, they can only be displayed statically and cannot interact with the performance in real time, so users cannot get immediate feedback when they make mistakes. The lack of data preservation and analysis during practice prevents users from reviewing and improving their weaknesses. Existing AI music generation tools are mostly focused on accompaniment or melody generation, failing to integrate with the user's actual performance process, resulting in a disconnect between practice and creation. When users create music, the lack of music theory optimization and expansion support often limits them to simple fragments, making it difficult to form a complete work.
[0025] The main solution in this application's embodiments is: By introducing interactive sheet music and AI models into smart musical instruments, a complete closed loop of "performance—feedback—review—optimization—expansion" is constructed. Examples include: The system converts audio signals or sheet music files into interactive scores with prompts, which can be displayed synchronously on smart instruments and terminals, providing users with intuitive performance guidance. During performance, the system captures the user's playing actions in real time and compares them with the interactive score, generating immediate feedback. After the performance, a review report is generated. An AI model is used to optimize the performance data using music theory, correcting deviations in pitch, rhythm, and mode. Based on the review results, a personalized practice plan is generated to guide the user in improving weak areas. In creative scenarios, the AI model can not only expand performance fragments into complete musical structures but also perform timbre rendering, virtual vocal singing, and automatic lyric writing, thereby lowering the barrier to music creation. Through the collaboration of smart instruments and terminals, users can practice efficiently and develop their practice results into complete musical works, achieving a unity of learning and creation.
[0026] Therefore, the present invention proposes a music processing method; it is understood that the intelligent musical instrument is equipped with a control device for storing and executing the following method. The control device can be implemented using a main controller, such as an MCU (Microcontroller Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or a SOC (System On Chip).
[0027] It is understood that the music theory terms used in this article include the following meanings: Pitch refers to the highness or lowness of a sound in the frequency dimension, and is usually determined by the fundamental frequency of the sound wave. In musical grammar, pitch corresponds to a specific musical note (such as C, D, E, etc.) or a piano key number.
[0028] The duration of a musical note refers to its length, which is represented in musical notation as a whole note, half note, quarter note, etc. The duration corresponds to the time domain length of the audio signal.
[0029] Beat refers to the basic unit of time and the cyclical structure of strong and weak beats in music. It is usually defined by tempo (BPM, beats per minute) and time signature (such as 4 / 4, 3 / 4).
[0030] Rhythm refers to the combination of different note values and pauses, forming the organizational rules of music in the time dimension. Rhythm is determined by the relative duration between notes and the position of accents.
[0031] A musical phrase is a melodic segment composed of several notes that has a complete meaning, similar to a sentence in language. The boundaries of a musical phrase are usually determined by pauses, breath points, or changes in the direction of the melody.
[0032] Melodic contour refers to the overall trend of pitch change over time, such as ascending, descending, or wavy. Melodic contour reflects the flow and emotional characteristics of music.
[0033] A mode refers to a scale system consisting of a set of pitches and its central tone (tonic). Common modes include major and minor. The mode determines the overall color of music.
[0034] Harmony refers to the vertical relationship between two or more notes. Through harmony, chords can be formed, expressing the layers and tension of music.
[0035] A chord is a sound composed of three or more notes stacked in thirds, such as a major triad, minor triad, or seventh chord. Chords are used in musical compositions for accompaniment and tonality support.
[0036] Musical grammar refers to the rules governing the organization of music in terms of pitch, rhythm, harmony, and phrasing, similar to the grammatical structure of a language. In this invention, musical grammar is used to guide the AI model in optimizing and expanding performance data.
[0037] It is understood that, in the embodiments of the present invention, the AI model includes, but is not limited to, the following models: Jukebox: A music generation model proposed by OpenAI that can directly generate complete music audio with vocals and accompaniment based on lyrics, style and artist features.
[0038] MuseNet: It can generate multi-instrument and multi-style music at the symbolic music level (such as MIDI), and supports cross-style mixing and melody arrangement.
[0039] DiffRhythm: A music generation framework based on the diffusion model, supporting the generation of complete songs with vocals and accompaniment.
[0040] Suno Music Generation Model (Suno AI): A commercial music generation platform that allows users to generate songs with vocals and accompaniment by inputting text prompts (such as style, mood, lyrics, etc.).
[0041] Udio music generation model (Udio): Supports music generation based on text prompts, and can provide song generation, segment extension and audio inpainting functions.
[0042] Artificial Intelligence Virtual Artist (AIVA): An early AI composition tool suitable for generating symphonies, film scores, and background music.
[0043] PerformanceNet: Transforms symbolic music (MIDI, piano reel) into audio output with performance expressiveness, enhancing the realism and detail of the generated music.
[0044] ACE-Step Music Generation Model (ACE-Step): Combining a diffusion model with a lightweight Transformer architecture, it supports lyric editing, track mixing, and melody control.
[0045] Reference Figure 1 In one embodiment of the present invention, the music processing method includes steps S100-S300, wherein: S100, receive music information and convert the music signal into an interactive score, the interactive score including score information and a performance prompt signal associated with the score information, the performance prompt signal being used to prompt the performer to perform the performance action corresponding to the score information at the corresponding time; S200: Based on the performance output prompt signal provided by the interactive score, prompt the performer to practice playing according to the prompt signal and record the performance content; The S300 optimizes the performance content based on a music AI model and obtains the optimized performance content.
[0046] Furthermore, music information is first received through a terminal device (including but not limited to smartphones, tablets, or computers with audio input interfaces). This music information can be an audio file (such as MP3 or WAV), a sheet music file (such as MIDI or MusicXML), or a scanned sheet music image file. After receiving the music information, the system calls the parsing module to process the music signal and generates an interactive sheet music based on the parsing results.
[0047] The interactive score contains score information and performance prompts associated with that score information. Specifically, the interactive score marks the performance time of each note on a timeline and prompts the performer in the display interface with cursors, symbols, or animation effects, enabling the performer to perform the corresponding performance action at the corresponding time.
[0048] Furthermore, during the performance, the terminal device provides real-time prompts to the performer based on the performance output prompts provided by the interactive score. These prompts can take the form of symbols, text, color changes displayed on the screen, or feedback from external instruments such as lights and vibrations. Simultaneously, the audio signals generated during the performance are captured by the sound pickup module and stored as part of the performance content.
[0049] Furthermore, after the performance, the performance content is input into a music AI model for optimization. The music AI model first extracts features from the performance content, such as identifying features like pitch, rhythm, and note values. Then, based on a trained music grammar rule base and a deep learning network, it corrects deviations in the performance, removes noise, and smooths unstable pitches or rhythms, thereby obtaining an optimized performance. Finally, the optimized performance result can be output as an audio file or sheet music for the performer to play back or review.
[0050] The optimized result can be converted back into interactive sheet music and fed back to the smart instrument. Users can then practice a second time based on the new interactive sheet music, thus forming a closed loop of "performance - feedback - optimization - repetition".
[0051] Through the above embodiments, the music practice method of the present invention can: Automatically convert audio or sheet music into interactive practice content to lower the learning threshold; It provides real-time prompts and recording during the performance to help users gradually develop a sense of rhythm and pitch. Artificial intelligence is used to optimize performance results based on music theory, making practice more efficient and composition more standardized.
[0052] It is understood that the interactive sheet music refers to sheet music data generated based on input audio signals or sheet music information that can be practiced by users, including but not limited to notes, rhythms, and key signatures of traditional sheet music, as well as performance prompts.
[0053] The prompts can be displayed through a touch screen, LEDs, vibration motors, or sound prompts to guide the performer to perform the performance at the corresponding time.
[0054] Furthermore, in this embodiment, the performance prompt signal refers to the prompt information output under the control of the interactive score, which is used to remind the user to perform specific performance actions.
[0055] The prompts include, but are not limited to, beat prompts (such as vibration or flashing lights), note prompts (corresponding key lights illuminate), and breathing direction prompts (blowing / inhaling direction indicator lights).
[0056] It is understood that, in this embodiment, performance feedback refers to the feedback information generated by the system based on the comparison results between the interactive score and the performance data during or after the user's performance.
[0057] Performance feedback includes, but is not limited to, information on pitch deviation, rhythm deviation, accuracy score, and performance suggestions.
[0058] As is understood, in this embodiment, debriefing practice refers to the process after a performance, where the system analyzes the recorded performance data, compares it with the interactive score, generates a debriefing report, and then generates a targeted practice plan based on the report. Debriefing practice forms a closed loop of "performance—feedback—optimization—re-practice".
[0059] It is understood that, in this embodiment, music theory optimization refers to the standardization and expansion of the performance results or user-input melody at the music theory level based on an artificial intelligence model.
[0060] Optionally, the optimization includes: Modal consistency optimization: Adjust notes that deviate from the key to notes that conform to the scale; Beat and rhythm optimization: Align irregular rhythms to standard beat points; Harmony and phrasing optimization: Add harmonic accompaniment to the melody or re-divide the phrasing; Creative expansion: Generate transitional sections, accompaniment sections, or complete musical forms based on existing melodies.
[0061] In this embodiment, the smart musical instrument includes a smart musical instrument and a terminal; A smart musical instrument refers to the hardware component of a smart musical instrument, which is used to receive interactive sheet music and perform performance operations.
[0062] It includes a sound output module (for outputting audio signals), a performance prompt module (for outputting prompt signals), a data recording module (for recording the performance process), and a communication module (for interacting with the terminal).
[0063] A terminal refers to the software that runs applications and the hardware platform on which they are hosted, such as a smartphone, tablet, or computer.
[0064] It is used to perform interactive score generation, AI optimization processing, performance data analysis, and user interface interaction.
[0065] The AI model includes, but is not limited to, artificial intelligence algorithm models used to perform functions such as music score recognition, voiceprint recognition, pitch detection, rhythm analysis, mode recognition, harmony generation, and melody expansion.
[0066] Furthermore, it can include convolutional neural networks, recurrent neural networks, generative adversarial networks, or large language models, etc.
[0067] Optional, refer to Figure 2 and Figure 3 Another embodiment of the present invention provides a format for analyzing music information, and a process for converting audio signals into interactive musical scores using an AI model, based on the above. Figure 1 The example shown; Includes the following steps: Step S110: Format Detection. In this step, the system performs format detection on the received input data to determine its type. The input data can be a musical score image, an audio signal, or a musical score file. Format detection allows different types of input data to be routed to their respective processing paths.
[0068] Step S121: Music score image input. When the input data is a music score image, the music score image is input into the system.
[0069] Step S131: Image recognition. Based on the image recognition algorithm, the input musical score image is recognized, and musical score elements such as notes, rhythm symbols, and key signatures are extracted to obtain the corresponding musical score data.
[0070] Step S122: Audio input. When the input data is an audio signal, the audio signal is input into the system.
[0071] Step S132: Preprocessing, preprocessing the input audio signal, including but not limited to noise reduction, framing, windowing and spectrum analysis, in order to extract features related to the music score later.
[0072] Step S142: AI processing. After preprocessing, an artificial intelligence algorithm is used to process the audio signal and extract musical score features, such as pitch, duration, and harmonic structure. This step converts continuous audio signals into symbolic feature data that can be used for score generation.
[0073] Step S123: Input sheet music format file. When the input data is a standard sheet music format file (such as MusicXML, MIDI, etc.), directly import the file and parse the sheet music element information contained therein.
[0074] Step S150: Score generation, which integrates the score elements obtained from image recognition, AI processing or score file parsing to generate a complete digital score.
[0075] Step S160: Output interactive score. Finally, the system outputs interactive score, which not only includes traditional musical notation information, but also combines performance prompts to provide interactive prompts and training assistance to the performer.
[0076] Furthermore, for audio input, the detailed steps based on steps S132, S142, and S150 are as follows: Step S1321: Digital acquisition of audio signals. The input analog audio signal is acquired and converted into a digital audio signal through an analog-to-digital converter module for subsequent feature extraction and processing.
[0077] Step S1322: Sound feature extraction. Based on the digital audio signal, acoustic feature parameters including spectrum, energy, and formants are extracted to reflect the timbre and dynamic characteristics of the audio signal.
[0078] Step S1323: Noise Reduction and Normalization. The acquired audio signal is denoised to remove background noise and irrelevant interference. The amplitude range of the audio signal is normalized to improve the stability of the signal and the accuracy of subsequent analysis.
[0079] Step S1421: Pitch and duration extraction. Pitch and duration information are extracted from the preprocessed audio signal to form basic data units corresponding to the notes.
[0080] Step S1422: Divide the natural tones into notes, and divide the extracted continuous pitch data into independent note units according to time, thereby obtaining discretized musical notation.
[0081] Step S1423: Pitch quantization, which maps the extracted pitch information to the standard scale to achieve pitch discretization and standardization, thereby ensuring that the generated notes meet the requirements of music theory.
[0082] Step S1424: Music theory structure and rhythm analysis. Based on the extracted note data, the musical phrase structure is analyzed in combination with music theory rules, and rhythmic features are extracted through waveform analysis.
[0083] Step S1425: Extract the melody outline through waveform analysis, and further extract the rhythmic pattern and rhythmic features of the melody through time domain and frequency domain analysis, providing a basis for subsequent score generation.
[0084] Step S1426: Phrase boundary detection. Based on the pauses, duration changes, and rhythmic features between notes, the boundaries of musical phrases are detected to achieve the division of the musical hierarchy.
[0085] Step S1427: Call the grammar rule library to perform music grammar analysis. Call the built-in music theory grammar rule library to perform grammar analysis on the segmented musical phrases to ensure that the generated score conforms to the music grammar rules.
[0086] Step S1501: Convert the physical quantities of the audio signal into musical notes.
[0087] Step S1502: Time value mapping, mapping the extracted time value information to the corresponding musical beats so that they can be combined into a complete musical score later.
[0088] Step S1503: Combining musical score elements, integrating multiple musical score elements such as notes, rhythms, and harmonies to form structured musical score data.
[0089] Step S1504: Format conversion and score generation. The structured score data is converted into a standardized, interactive digital score through a format conversion engine and output for user use.
[0090] Understandably, when a user inputs humming audio through the terminal's microphone, the system first preprocesses the audio signal, including: Background noise can be removed by using spectral subtraction or deep learning denoising models to obtain a clean speech signal; The continuous audio is divided into frames of fixed length and windowed to ensure temporal smoothness; Mel frequency cepstral coefficients (MFCC) and short-time Fourier transform (STFT) spectra are extracted as input features for the AI model.
[0091] The preprocessed audio features are input into a hybrid model based on convolutional neural networks and recurrent neural networks. The model performs the following steps: The fundamental frequency of the audio signal is extracted by a deep learning fundamental frequency detection module and quantized into pitch features, corresponding to the note heights in the musical score. The duration of a note is calculated using signal energy and autocorrelation function, and then quantized into its corresponding time value (e.g., quarter note, eighth note). Based on the abrupt change points of audio energy and the continuity of the fundamental frequency, the breakpoints between notes are identified to achieve the delineation of note boundaries; Detect long pauses or energy breaks between notes to divide the note sequence into multiple musical phrases; By using time series modeling, the trend of pitch change over time is obtained, and melodic lines are generated. It calls the built-in music grammar rule library, infers the mode based on the note set and interval relationship, and generates corresponding chord accompaniment information according to the pitch distribution.
[0092] After the above steps, the system obtains a set of music data that conforms to musical grammar, including: Note pitch sequence; Note value information; Phrase boundaries; Melody progression; Mode and chord information.
[0093] This music data can be stored as structured score files, including but not limited to MusicXML or MIDI formats.
[0094] The interactive music score creation in this embodiment includes the application creating interactive music scores based on generated music data. In this process: Transform note, time value, and mode information into visual musical notation; Add an interactive marker to each note. The interactive marker is used to drive the performance prompt module on the smart musical instrument to output a prompt signal. The prompt signal includes, but is not limited to, the key light illuminating, the vibration motor prompting the beat, and the breathing light indicating the direction of blowing and sucking. The generated interactive sheet music is sent to the smart musical instrument and simultaneously displayed on the terminal touch screen for users to practice playing.
[0095] In a preferred embodiment, the user can input two types of data sources via the terminal: one is an image of a printed or handwritten sheet music, and the other is a standard sheet music format file, such as a MIDI file or a MusicXML file. The terminal application processes the different inputs separately to generate interactive sheet music that can be executed by a smart musical instrument.
[0096] In a typical application scenario, the user first inputs an audio signal via a smart musical instrument or terminal. This signal could be a humming, a performance clip, or an external musical piece. Upon receiving the input, the system preprocesses the audio, performing noise reduction, frame segmentation, and feature extraction to obtain cleaner and more easily analyzable audio data. The preprocessed signal is then fed into an acoustic feature extraction module, which extracts the fundamental frequency, energy, and spectral characteristics of the audio, and converts these acoustic features into preliminary note data using a feature mapping model.
[0097] Next, the system enters the core AI processing stage. First, the fundamental frequency detection model extracts pitch and duration, obtaining the fundamental frequency information and duration of the notes. Then, the note segmentation module segments the natural sound stream based on the continuity of the fundamental frequency and energy changes, dividing the continuous sound signal into independent note units. Based on this, the note quantization module maps the segmented natural sounds to a standardized set of pitch and duration, thus forming a standardized note sequence. Next, the structure and rhythm analysis module analyzes the temporal relationships between notes using a rhythm recognition model, obtaining the beat pattern and rhythmic structure. The melody contour extraction module further extracts the melody direction based on the trend of pitch changes over time, while the phrase boundary detection module identifies complete phrase boundaries by analyzing features such as pauses and energy breaks. Finally, the music grammar analysis module combines a mode and harmony rule base to infer the tonality and harmony of the piece, thereby transforming the fragmented note data into a complete structure that conforms to music grammar.
[0098] In the output phase, the system further processes the optimized music data. First, it generates interactive sheet music using note and rhythm features, allowing the sheet music to be displayed as prompts on terminals or smart instruments to assist users in playing. Second, the system can also utilize a timbre rendering model to transform the same sheet music into different timbre representations, such as piano, guitar, or strings, and can select a virtual vocal synthesis module to convert the melody into a human voice. For scenarios requiring lyrics, the automatic lyric-filling module can generate lyrics by combining the rhythm and emotion of the melody and match the lyrics with the virtual vocals. Finally, all processing results are output as a musical work, allowing users not only to practice playing but also to obtain a complete musical piece optimized and expanded according to music theory.
[0099] In a preferred embodiment of the invention, the user first obtains an interactive musical score through a terminal. This score not only contains traditional note and rhythm information but also includes corresponding prompts. During practice, the intelligent musical instrument illuminates key lights, emits vibrations, or displays rhythm symbols on the touchscreen according to the instructions of the interactive score. The performer simply needs to follow these prompts to perform blowing, breathing, and key presses to complete a melody.
[0100] Optional, see reference Figure 4A and Figure 4B In one embodiment of the present invention, the terminal's touchscreen displays a music performance interactive interface, which provides interactive sheet music, rhythm prompts, and operation feedback information during the user's performance. The interface consists of four main display areas: a media content display area, a music content display area, a sheet music information display area, and an interactive sheet music display area.
[0101] The media content display area is primarily used to play music-related multimedia content, such as accompaniment videos, instructional videos, or user-recorded performance clips. This area can update in sync with the music rhythm and can also display demonstrative performance operations in instructional mode.
[0102] The music content display area is located at the top of the interface and is used to scroll through the lyrics and melody progress bar of the current piece. Users can simultaneously refer to the lyrics and rhythm progress for synchronized auditory and visual playing. In music without lyrics, this area can also be used to display text-based instruction or supplementary information, such as beat descriptions and dynamics indicators.
[0103] The music score display area shows simplified musical notation generated or converted by an AI model, using numerical symbols "1-7" to represent pitch, and adding symbols such as "~" and "⌒" above the symbols to indicate sustain or legato effects. This intuitive display in simplified musical notation allows users to quickly identify the order and duration of target notes.
[0104] The interactive sheet music display area is the core interactive area of the interface, used to display performance prompts generated by the interactive sheet music. These prompts are displayed in a falling trajectory format, with the prompt symbols moving downwards along the track. When a prompt symbol coincides with the judgment line at the bottom of the interface, the user is prompted to perform the corresponding performance operation. The interactive sheet music display area includes multiple parallel tracks, each corresponding to a touch area or physical button operation area, used to distinguish the input channels for different notes.
[0105] To make performance prompts more intuitive, the interface features fingering key indicator lines and bottom key indicator dots. The fingering indicator lines show the keys the user should press, corresponding from top to bottom to the three fingering keys on the smart instrument (fingering keys 1, 2, and 3). The bottom key indicator dots are located in the decision line area at the bottom of the track, indicating whether the user needs to perform an operation using the bottom keys, such as raising or lowering an octave. When the user presses bottom key 1, the overall pitch rises by one octave; when the user presses bottom key 2, the overall pitch falls by one octave, thus achieving wide range coverage.
[0106] In addition, the interface includes blow / inhale indicator arrows to prompt the user to perform "blow" or "inhale" airflow operations during performance. For example, an upward arrow indicates that inhalation is required, while a downward arrow indicates that blowing is required. Through airflow sensor detection, the system can determine the user's actual blowing and inhalation force and duration to further calculate the accuracy of note triggering.
[0107] The rhythm interval lines are used to indicate the time intervals between different rhythm beats, helping users master the playing rhythm pattern. The judgment line is the reference benchmark for the system to evaluate the timing of user operations. When the prompt symbol falls to the judgment line, the user should tap or long-press the touch area of the corresponding track to complete the playing of the corresponding note.
[0108] During the performance, the system performs real-time evaluation based on the synchronization between the prompt signal and the user's operation. When the user executes the operation precisely when the prompt signal reaches the judgment line, and the trigger time deviation is ≤20 milliseconds, the interface displays an "EXCELLENT" prompt; if the deviation is between 20 and 50 milliseconds, "GOOD" is displayed; if the timing is missed, "MISS" is displayed. Simultaneously, the system generates a performance timing evaluation result based on the accuracy of the user's operation, rhythmic stability, and airflow intensity.
[0109] In the background logic layer, the system collects the user's performance data in real time, including note trigger time, blowing and inhaling force, duration, and button operation status. After the performance, the system compares this data with the standard data of the interactive score one by one to analyze the differences in the user's performance. For example, it detects whether the notes are triggered at the correct time, whether the pitch is off, and whether the rhythm is stable.
[0110] The comparative analysis results will generate review feedback, which can take the form of charts or interactive prompts. The system can automatically mark phrases or notes with significant deviations in the performance and provide improvement suggestions, such as "the tempo is too fast" or "the pitch is too high." Users can select key passages for review based on the feedback, or adjust their performance style according to the system's suggestions.
[0111] Furthermore, in learning mode, the system can establish a correspondence between "target note - rhythm pattern - dynamic variation - performance action". By analyzing the user's playing habits and ambiguous commands, the system can dynamically adjust the tempo and brightness of the prompts. For example, when the AI model recognizes the ambiguous semantic command "more sad", the system will automatically slow down the tempo (below 60 BPM), slightly lower the pitch, and reduce the brightness, thereby achieving emotional performance guidance.
[0112] Optional, refer to Figure 5 In a preferred embodiment of the present invention, a method for optimizing a music AI model based on fuzzy commands is provided. This method uses fuzzy emotional descriptive words (including but not limited to "sad," "exhilarating," "relaxing," and "tense") input by the user to drive the music AI model to modify the music accordingly. This achieves the adjustment of musical effects without requiring the user to possess specific music theory knowledge or input explicit music theory parameters. Examples include: First, regarding the input and recognition of vague commands, in this embodiment, the user inputs vague emotional commands through the terminal interface. These commands are typically adjectives or phrases in natural language, such as "make the melody sadder," "make the music more exciting," or "want a gentler background." The system uses a natural language processing module to semantically analyze the input content, transforming the user's vague expression into semantic features that can be recognized by the AI model.
[0113] Secondly, the semantic feature mapping of the instructions maps the parsed fuzzy sentiment words to the latent dimensions of the music feature parameters. For example: “Sadness” corresponds to a drop in pitch, a slow rhythm, and an increase in the number of mutes; "Exhilarating" corresponds to increased volume, faster rhythm, and denser harmony; "Soothing" corresponds to a coherent melody, enhanced low frequencies, and weakened percussion. "Tense" corresponds to an increased use of dissonant chords and irregular rhythms.
[0114] It is important to emphasize that these mapping rules are not explicit music theory parameters, but rather adaptive adjustments of the weights within the trained music AI model and the multidimensional feature space.
[0115] Furthermore, regarding music data input and processing, in the specific implementation process, users can select different types of music input data, including: There are already music files (MIDI, audio files, etc.); User-improvised performance clips; The basic melody is automatically generated by the system.
[0116] The input music data first undergoes feature extraction, including information such as melodic lines, rhythmic patterns, timbre distribution, and harmonic structure. These features serve as the basic characteristics of the music to be modified and are then fed into the music AI model.
[0117] Furthermore, the AI model employs fuzzy optimization processing. After receiving the emotional features of fuzzy instructions, the music AI model performs multi-dimensional adjustments to the input music data: Melody adjustment: Different emotions can be expressed by choosing pitch, scale, and changing the melodic line. For example, the melody can be made more somber and expansive, or more uplifting and compact.
[0118] Rhythm and tempo adjustment: Change the note value, beat density, and overall tempo to achieve effects such as "soothing" or "exhilarating".
[0119] Harmony and tonality modification: Based on fuzzy instructions, select consonant or dissonant chords, or perform tonality modulation, thereby conveying different musical emotions.
[0120] Timbre and Dynamics Processing: By adjusting the timbre distribution, volume, and dynamic range of instruments, the overall atmosphere of the music is made to better match the vague emotional description.
[0121] Finally, the AI model generates modified versions of music, which are then output to the user's device in sheet music or audio format for listening and verification. If the user feels the effect is not as expected, they can input new, vague instructions (including but not limited to "a little sadder" or "not too rushed"), and the system will then perform another optimization iteration. Through this feedback loop, users can gradually obtain musical works that meet their emotional expectations.
[0122] It is understood that, through the above embodiments, the present invention achieves the following technical effects, including: Users do not need to have complex music theory knowledge to make music adjustments based on fuzzy language; The system can map fuzzy descriptions to specific musical features, enabling comprehensive modifications to melody, rhythm, harmony, timbre, and other aspects. It provides a more interactive and accessible method for music generation and optimization, enabling professional musicians to quickly experiment with emotional creations and allowing ordinary users to easily participate in music creation and modification.
[0123] Furthermore, in a preferred embodiment, a music generation model (Udio) is used to optimize the performance content, generating an optimized performance result that conforms to musical grammar and reflects the user's emotional intent. This includes the following steps: Audio preprocessing: During performance, the pickup module acquires the user's original performance audio signal. Since the original audio may contain environmental noise or background interference, the control module first performs preprocessing operations, including: Filter the audio signal to remove low-frequency or high-frequency noise; Energy normalization ensures that the amplitude of the audio signal remains within a uniform range under different performance intensities. By using silence detection technology to remove blank segments, a cleaner audio signal can be obtained.
[0124] The pre-processed clean audio signal is then passed as input to the Udio music generation model.
[0125] Audio translation and feature extraction: The clean audio signal is input to the front-end processing module of the Udio model. In this stage, speech-music feature translation is performed, including: Fundamental frequency analysis and pitch feature extraction: The fundamental frequency of the performance audio is detected by Fast Fourier Transform (FFT) and mapped to the corresponding pitch features; Time domain duration converted into note value: Through time envelope analysis, the duration of each note is extracted and converted into a note value that conforms to the rules of musical notation; Note segmentation: Using energy mutation points and spectral continuity detection methods, continuous audio data is segmented into individual note units; Phrase boundary construction: Based on the breakpoint detection algorithm, the pauses and structural turning points of the melody are identified to determine the natural boundaries of musical phrases.
[0126] Music grammar analysis and melody construction: After feature extraction, the Udio model calls its built-in music grammar rule library and generation engine to further process the extracted data. Melody contour extraction: By analyzing waveform and frequency changes, the overall direction of the melody is obtained, such as ascending, descending or cyclic structure; Tonality and chord construction: Combining the tonal distribution features and chord progression patterns learned from the training corpus, the melody is mapped to a tonal system that conforms to musical grammar, and appropriate chord accompaniment is matched for it.
[0127] Fuzzy emotion commands drive optimization. Users can input vague natural language commands via the terminal, such as "more sad," "more exciting," or "more soothing." The Udio model can translate these vague commands into specific music adjustment parameters, including: Under the "sad" instruction, lower the overall key, lengthen the note values, and weaken the rhythmic intensity; Under the "exhilarating" command, increase the tempo, enhance the harmonic density, and increase the proportion of high-frequency instruments; Under the "soothing" instruction, the dynamic range is reduced, dissonant intervals are decreased, and melodic coherence is improved.
[0128] Through this process, the model not only generates performance data that conforms to musical grammar, but also further optimizes it based on the user's vague emotional needs.
[0129] After the above processing, the Udio model outputs the optimized performance content. The output format may include: Digital sheet music: exported in MIDI or MusicXML format for easy further editing or archiving; Audio signal: Directly generates a complete audio file containing melody, harmony and accompaniment, which users can listen to in real time; Real-time feedback signal: The optimized melody is compared with the current performance, and a prompt signal is generated and returned to the performer through the instrument's prompt module or feedback module.
[0130] Optional, refer to Figure 6 One embodiment of the present invention provides a terminal, which includes a touch screen, a memory, a processor and a communication module. The components interact with each other through a data bus or an electrical signal path, thereby completing the various functional steps of the music processing method.
[0131] Furthermore, the touchscreen includes a display module and a touch module. The display module presents performance prompts and content graphically, such as displaying notes, beat points, or interactive markers on the screen in the form of a timeline or beat line. The touch module works in conjunction with the display module. When the performer performs touch operations based on the performance prompts, the touch module can sense the touch events in real time and transmit the touch signals to the processor for processing, thereby realizing the virtual performance function. The performance prompts are presented at preset time intervals, allowing the performer to accurately follow the prompts.
[0132] The processor, as the core control unit of the terminal, runs the music processing program stored in the memory, scheduling and coordinating the work of various modules. In one specific embodiment, the processor executes the functions of the score recognition module and the interactive score generation module by calling the music processing program. The memory stores data related to music processing, including performance content generated by the performer during practice, interactive score files, and the music processing program itself. In one specific embodiment, the memory can also save performance logs and review data for subsequent AI model optimization or review training. In this way, the terminal can not only support real-time performance and display but also provide access to and comparison of historical data, ensuring the integrity and traceability of practice.
[0133] The communication module is used to establish a communication connection between the terminal and the smart musical instrument, supporting wired or wireless data interaction.
[0134] Optionally, refer to Figure 7A , Figure 7B and Figure 8 Another embodiment of the present invention provides an intelligent musical instrument, the main body of which integrates a communication module, a prompting module, a control module, a sound pickup module, an interaction module, and a feedback module. The various modules work together through electrical connection and data interaction to realize complete music interaction and practice functions.
[0135] For example, in practical applications, the main body of the musical instrument serves as the core carrier for performance, not only supporting the user's playing operations but also incorporating various sensing and interactive components. A sound pickup module is located inside the instrument's main body to collect sound signals generated by the user during performance, such as airflow sounds, key presses, or string vibrations. These sound signals are transmitted to the control module via circuitry. The control module analyzes and processes the signals, extracting characteristic parameters such as pitch and note value, and then combines this information with prompts returned from the terminal for further analysis.
[0136] Upon receiving a prompt signal from the terminal, the control module triggers the prompt module accordingly. The prompt module provides intuitive cues to the performer through lights, symbols, or sound signals, enabling the performer to execute the corresponding performance action at the correct time. Simultaneously, the prompt module maintains two-way interaction with the control module, ensuring that the prompt output is consistent with the performance rhythm.
[0137] The interaction module, acting as a bridge component in the system, is directly connected to the control module and undertakes the function of converting bidirectional information flows. On one hand, the interaction module converts the sound signals collected by the sound pickup module and processed by the control module, along with relevant data from performance practice, into standardized music signals and sends them to the terminal for interactive score comparison, AI optimization, or data storage. On the other hand, the interaction module is also responsible for receiving prompt signals from the terminal and transmitting them to the control module to further trigger the prompt and feedback modules.
[0138] The feedback module, located within the instrument's main body and electrically connected to the interaction module, integrates a vibration motor. During performance, the feedback module emits physical vibrations of corresponding frequency and amplitude based on the performance content and feedback from the terminal. The performer can directly feel this feedback through touch with their fingers or mouth, thus gaining a more immersive and real-time performance experience. This vibrational feedback can simulate the rhythm and dynamics of notes, helping performers practice even in the absence of external sound.
[0139] The communication module ensures the interconnection between the smart musical instrument and the terminal, supporting wired interfaces or wireless methods (such as Bluetooth or Wi-Fi) to achieve high-speed and stable data interaction. Through the communication module, the smart musical instrument can upload the collected performance data to the terminal in real time, while simultaneously receiving control or prompt signals from the terminal, ensuring the closed-loop operation of the entire system.
[0140] In a preferred embodiment, the intelligent musical instrument includes, but is not limited to, an interactive electronic tube instrument based on a combination of breath-driven and key fingering. Its core feature is that through multi-dimensional input of blowing / inhaling actions, fingering key combinations, mouthpiece biting, and thumb pitch-changing keys, the corresponding note signal triggering and pitch-changing control are realized, thereby achieving the seven notes of a natural scale with minimal operations.
[0141] During performance, when the performer blows into the mouthpiece, the system identifies the airflow direction as "positive breath" and enters the blowing register; when the performer inhales, the system identifies it as "reverse breath" and enters the inhalation register. For the blowing direction, if the performer does not press any finger keys, the output note is "do"; when only the first finger key is pressed, the output is "mi"; when only the second finger key is pressed, the output is "sol"; when only the third finger key is pressed, the output is "ti". For the inhalation direction, if no finger key is pressed, the output is "re"; pressing the first finger key outputs "fa"; pressing the second finger key outputs "la"; pressing the third finger key outputs "high do". Thus, blowing corresponds to the odd-numbered notes of the natural scale (1, 3, 5, 7), and inhalation corresponds to the even-numbered notes (2, 4, 6, high 1), forming a complete octave range.
[0142] To achieve a wider range, the instrument features two octave-switching buttons at the bottom. Pressing the bottom left button raises the overall pitch by one octave, while pressing the bottom right button lowers it by one octave. This allows users to freely switch between the low, middle, and high registers without altering the fingering, achieving a range spanning three octaves. For example, blowing air with all three fingering buttons in their normal positions produces a middle "do"; pressing the bottom left button produces a high "do"; and pressing the bottom right button produces a low "do." Similarly, inhaling air allows for rapid switching between different octave registers using button combinations.
[0143] This fingering system fully integrates a three-dimensional input mechanism of breath direction, finger position, and octave control in its performance logic. It eliminates the complex operations of "multi-key combinations" and "mouthpiece pitch changes" found in traditional solutions, enabling performers to achieve full-scale performance with a single airflow and single-key action. Through an internal audio signal mapping algorithm, the system maps each breath and key combination to a unique pitch output signal, ensuring accurate note recognition and low response latency.
[0144] Optional, refer to Figure 9A , Figure 9B and Figure 10 In a preferred embodiment, the smart musical instrument adopts a modular mechanical structure design, which is composed of a shell assembly, interactive components, control circuits and feedback modules, ensuring both portability and a good interactive experience and performance feedback capability.
[0145] The smart musical instrument's casing consists of an upper shell, a middle shell, and a bottom shell, which enclose and support the internal functional components. The upper shell features finger placement keys and a breathing indicator light, providing users with intuitive operation and prompts. The middle shell houses the mainboard, subboard, battery, and airflow sensor, among other core components. The bottom shell, when assembled with the middle shell, forms a complete sealed cavity, ensuring mechanical strength and safety.
[0146] In terms of the interaction module, the instrument is equipped with multiple input units, including a key interaction unit, a blow / inhale interaction unit, a press interaction unit, and a fingertip interaction unit. The key interaction unit is mainly used to simulate the fingering operations of traditional musical instruments; the blow / inhale interaction unit is located at the front of the instrument and works with an airflow sensor to detect the user's breathing intensity and direction; the press interaction unit is used to generate additional control signals by applying pressure with the thumb; and the fingertip interaction unit provides auxiliary operation by wearing a fingertip on the little finger. These interaction components are connected to the processing module via a circuit board, enabling real-time acquisition of user movements and conversion into performance data.
[0147] The control unit consists of a processor, a memory, and a control module. The processor is responsible for executing preset music processing programs and recognizing and parsing the acquired motion signals; the memory is used to store performance data, interactive scores, and practice plans; the control module is responsible for signal distribution and logical judgment, and transmits the processed results to the prompting module and the feedback module.
[0148] The prompt module is located on the upper surface of the housing and is used to display performance prompts, such as indicating the user's key presses, rhythms, and breathing operations through lights or symbols; the feedback module includes a linear vibration motor, which is used to generate tactile feedback during performance, so that the user can get immediate confirmation of operation.
[0149] In addition, the instrument is equipped with a pickup module and a communication module. The pickup module can collect the performer's voice signal or external accompaniment signal and provide it to the processor for further music synthesis or comparison; the communication module supports wired interface and wireless transmission, ensuring that the instrument can interact with terminal devices to realize the distribution of interactive scores, the uploading of performance data, and the invocation of AI optimization.
[0150] In a preferred embodiment, when the smart musical instrument of the present invention executes interactive musical scores, the performance prompt signal is not only output on the smart musical instrument, but also synchronously displayed on the display module of the terminal.
[0151] Furthermore, after the interactive sheet music is sent to the smart instrument, the performance prompt module will illuminate indicator lights, emit vibrations, or emit other forms of prompts at corresponding times to guide the performer in performing the music. At the same time, the application on the terminal will synchronously display corresponding prompt information on the display module, such as highlighting the notes to be played on the screen interface, displaying flashing beat markers, or indicating the breathing direction of the performance on the score.
[0152] In this way, users can operate the instrument using both physical cues from the instrument itself and intuitive graphical guidance from the display module, achieving a dual-cue effect. This synchronization mechanism significantly enhances the learning experience, making it particularly suitable for beginners to practice rhythm and note following in the early stages.
[0153] In another embodiment, the smart musical instrument of the present invention further supports the transmission of audio signals between the smart musical instrument and the terminal.
[0154] During performance, users generate playing commands through button presses or blowing / drawing operations on the smart instrument, and the sound module synthesizes corresponding audio signals based on these commands. In addition to playing the audio signal locally on the smart instrument, the audio signal is also transmitted to the terminal in real time via signal transmission channels (including but not limited to Bluetooth, Wi-Fi, or USB interface).
[0155] After receiving the audio signal, the terminal's main control chip can process it, such as enhancing sound effects, equalizing volume, compensating for delay, or mixing it with the accompaniment track. The processed audio signal will then be transmitted to the terminal's playback device, including the built-in speaker or external headphones, to output high-fidelity, optimized performance sound to the user.
[0156] In one scenario, when a user uses external headphones, they can practice music without making noise or disturbing others.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0158] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A music processing method characterized by, The method comprises: receiving music information and converting the music signal into an interactive score, the interactive score comprising score information and a performance prompt signal associated with the score information, the performance prompt signal being used to prompt the performer to perform a performance action corresponding to the score information at a corresponding time; providing a performance output prompt signal according to the performance prompt signal provided by the interactive score to prompt the performer to perform a performance practice according to the prompt signal and record the performance content; optimizing the performance content based on a music AI model and obtaining the optimized performance content.
2. The music processing method of claim 1, wherein, The receiving music information and converting the music signal into an interactive score comprises: receiving music information and determining the format of the music information; in the case where it is determined that the format of the music information is an audio file, using an audio decoding technology to analyze the audio features in the audio signal; in the case where it is determined that the format of the music information is a picture file, obtaining the audio features in the score picture based on an image recognition technology; in the case where it is determined that the format of the music information is a score file, identifying the audio features in the score file; generating an interactive score based on the audio features.
3. The music processing method of claim 2, wherein, The audio features at least include notes, pitches and timbres; The generating an interactive score based on the audio features specifically comprises: determining a prompt time of the prompt signal based on the timestamp of the target note; determining the prompt content of the prompt signal based on the pitch and timbre of the target note; establishing an association between the prompt content and the timestamp of the target note.
4. The music processing method of claim 1, wherein, The method further comprises: receiving a fuzzy instruction input by the performer; The fuzzy instruction includes voice input, text input or natural language input; performing semantic recognition and sentiment analysis on the fuzzy instruction based on a music AI model to extract audio features corresponding to the fuzzy instruction; The audio features at least include pitch changes, rhythm speeds, volume strengths and rhythm patterns; establishing a mapping relationship between the semantic features of the fuzzy instruction and the audio features, and converting the fuzzy instruction into an interactive score based on the mapping relationship to guide the performer to perform a performance operation; The prompt signal at least includes one of light flashing of an intelligent musical instrument, vibration motor vibration, or note track, beat symbol on a terminal screen, for prompting the performer to perform a performance action corresponding to the score information at a corresponding time; The method further comprises converting the optimized performance content into an interactive score.
5. The music processing method of claim 1, wherein, The optimizing the performance content based on a music AI model and obtaining the optimized performance content comprises: preprocessing the audio signal to obtain a pure audio signal, and inputting the preprocessed audio signal into an AI model for translation to convert the pure audio signal into music data conforming to a music grammar; The translation comprises: converting the fundamental frequency of the sound signal into a pitch feature and converting the time domain duration of the sound signal into a note time value; performing note segmentation based on the continuity of the audio data; constructing a phrase boundary based on breakpoint detection; extracting a melody contour based on waveform analysis; constructing tonality and chords based on a music grammar rule library.
6. The music processing method of claim 1, wherein, The method further comprises: recording performance process data in the performance process; After the performance, the performance process data is compared with the interactive score to generate review feedback information.
7. A terminal, characterized by comprising: The application comprises: a processor, a memory, a touch screen and a communication module; the touch screen comprises a display module and a touch module; the display module is used to display performance prompt signals and performance content; the touch module is used to perform performance operations in a touch manner according to the performance prompt signals; the performance prompt signals are presented on the display module at set time intervals; the communication module is used to establish a communication connection between the terminal and the intelligent musical instrument; the memory and the processor, and a music processing program stored in the memory and executable on the processor, wherein the music processing program is configured to implement the steps of the music processing method according to any one of claims 1 to 6.
8. The terminal according to claim 7, characterized by The processor further comprises: a score recognition module and an interactive score generation module; the score recognition module comprises a format recognition module for recognizing the format of a music signal and a command issuing module for issuing a command corresponding to the format according to the format; the interactive score generation module converts the music signal into an interactive score according to the command; the score recognition module and the interactive score generation module are implemented by the processor executing a corresponding music processing program.
9. The terminal according to claim 7, characterized by The memory further comprises: for storing performance content and interactive scores.
10. A smart musical instrument, characterized by, The application comprises: a musical instrument body for performance practice; a control module arranged on the musical instrument body and used to process sound signals collected by a sound pickup module and control a prompt module to give prompts based on prompt signals fed back by a terminal; a sound pickup module arranged on the musical instrument body and connected to the control module and used to pick up the sound signals; an interaction module arranged on the musical instrument body and connected to the control module and used to convert the sound signals and performance practice content into music signals and send them to a terminal based on the control of the control module, and further used to receive prompt signals fed back by the terminal; a prompt module arranged on the musical instrument body and connected to the control module and used to give prompts to the received prompt signals based on the control of the control module; a feedback module arranged on the musical instrument body and connected to the interaction module and used to give responses with corresponding frequencies and amplitudes following the performance content; a communication module arranged on the musical instrument body and used to establish a communication connection between the intelligent musical instrument and the terminal.