Synchronous dynamic display system and method for numbered musical notation, stave and lyrics in karaoke scene

By designing a system that dynamically displays simplified musical notation, standard musical notation, and lyrics simultaneously, the problem of insufficient guidance in dual-musical-notation integration in karaoke systems has been solved. This enables personalized assistance for users with different musical backgrounds, improving the karaoke experience and singing effect.

CN121884752APending Publication Date: 2026-04-17ZHONGRUI SHENGSHI WUHAN EDUCATION TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGRUI SHENGSHI WUHAN EDUCATION TECHNOLOGY DEVELOPMENT CO LTD
Filing Date
2026-01-21
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing karaoke systems lack dynamic guidance that integrates simplified musical notation, standard musical notation, and lyrics, failing to meet the needs of users with different musical backgrounds. Furthermore, existing dynamic notation technology is not adapted to karaoke scenarios and cannot achieve real-time synchronization between dual notation and vocal performance.

Method used

Design a system for synchronously and dynamically displaying musical notation, standard musical notation, and lyrics in a karaoke setting. The system includes a cloud server module, a user terminal module, a data parsing module, a synchronization control module, a dynamic display module, and an audio acquisition and analysis module. It enables synchronous display of musical notation, standard musical notation, and lyrics, as well as providing singing assistance. It achieves precise synchronization and personalized feedback through MIDI timing algorithms and AI pitch analysis algorithms.

Benefits of technology

It achieves precise synchronization of simplified musical notation, standard musical notation, and lyrics, adapts to the usage habits of users with different musical backgrounds, improves users' karaoke experience and singing level, provides personalized pitch and rhythm feedback, and covers the needs of users at all levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884752A_ABST
    Figure CN121884752A_ABST
Patent Text Reader

Abstract

The invention discloses a synchronous dynamic display system and method for numbered musical notation, stave and lyrics in a karaoke scene, and belongs to the technical field of music entertainment and intelligent interaction. The system comprises a cloud server module, a user terminal module, a data analysis module, a synchronous control module, a dynamic display module and an audio acquisition and analysis module, the cloud server module stores the standard audio accompaniment, the bispectrum original data, the lyric text data and the audio track time sequence data after correlation calibration; the data analysis module analyzes the resources into structured time sequence data; the synchronous control module establishes a precise synchronous relationship among bispectrums, lyrics, accompaniment and voice; the dynamic display module realizes three-layer visual synchronous display; and the audio acquisition and analysis module monitors singing deviation and synchronously feeds back the singing deviation. According to the method, the blank of bispectrum fusion guidance of the existing karaoke system is filled, different music basic user requirements are covered, the singing experience and level are improved, and the method can be widely applied to scenes such as online karaoke and offline KTV.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of music entertainment and intelligent interactive technology, specifically to a system and method for synchronously and dynamically displaying musical notation, staff notation, and lyrics suitable for karaoke scenarios. Background Technology

[0002] With the development of the digital entertainment industry, karaoke has become an important part of people's daily leisure and entertainment. User demand for online karaoke platforms (such as WeSing and Changba) and offline KTV venues continues to grow. Most mainstream karaoke systems now feature dynamic lyrics display, and some systems have also added dynamic pitch lines to help users master rhythm and pitch, enhancing the singing experience.

[0003] However, existing karaoke systems still have significant shortcomings: On the one hand, the auxiliary display methods of existing systems are mostly concentrated on lyrics and pitch lines, or only support the display of simplified musical notation, lacking dynamic guidance that integrates simplified musical notation, standard musical notation, and lyrics. This fails to meet the needs of users with weak musical foundations for intuitive simplified musical notation, and is also difficult to adapt to the usage habits of users with professional musical foundations who are accustomed to standard musical notation, resulting in insufficient singing guidance for users of different levels. On the other hand, existing dynamic notation technologies (such as BBQE dynamic simplified musical notation and Zhibeizhe AI ​​intelligent music notation) are mainly applied to instrumental performance and music teaching scenarios. Their functional design is not adapted to the core needs of karaoke scenarios. They cannot achieve real-time synchronization of dual notation with karaoke accompaniment and vocal performance, nor can they be deeply integrated with lyrics for display, making it difficult to directly apply them to karaoke systems.

[0004] Furthermore, while existing dynamic feedback technologies in music (such as the real-time performance feedback system of Magic Music Score) can track and provide feedback on performance status, these technologies focus on instrumental performance scenarios and are not optimized for the characteristics of human vocal singing. Therefore, they cannot meet the personalized assistance needs of karaoke users for pitch and rhythm. Thus, there is an urgent need for a display system and method that can adapt to karaoke scenarios, achieve precise synchronization of simplified notation, standard musical notation, and lyrics, and provide personalized assistance based on the characteristics of human vocal singing. This would fill the gap in existing technology, cover the needs of users with different musical backgrounds, and improve the user's karaoke experience and singing effect. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a system and method for synchronously and dynamically displaying sheet music, staff notation, and lyrics in a karaoke setting. This system achieves precise synchronization between sheet music, staff notation, lyrics, accompaniment, and vocal performance. It also adapts to the usage habits of users with different musical backgrounds, provides intuitive melody and rhythm guidance, and offers personalized feedback based on the characteristics of vocal performance, thereby comprehensively improving the user's karaoke experience and singing level.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a system for synchronous dynamic display of simplified musical notation, staff notation, and lyrics in a karaoke setting, comprising a cloud server module, a user terminal module, a data parsing module, a synchronization control module, a dynamic display module, and an audio acquisition and analysis module. The modules work together to achieve synchronous display of simplified musical notation, staff notation, and lyrics, as well as singing assistance. The cloud server module stores a song resource library, which includes standard audio accompaniment, lyrics text data, original simplified musical notation data, original staff notation data, and track timing data for each song. The original simplified musical notation data includes information such as notes, time signatures, key signatures, and durations. The original staff notation data includes information such as the staff chart, note positions, clefs, sharps, flats, and durations. The track timing data is used to link simplified musical notation notes, staff notation notes, lyrics, and standard audio accompaniment on a unified timeline. Simultaneously, the cloud server module also receives requests from the user terminal module to enable the transmission and updating of song resources. User terminal module: As a user interaction carrier, it is used to receive user operation commands (such as song selection, playback control, display parameter adjustment, musical notation switching, etc.) and transmit the commands to the cloud server module and synchronization control module; at the same time, it is used to display dynamic simplified musical notation, dynamic staff notation, lyrics and singing auxiliary information, and to collect the user's vocal audio. The data parsing module receives song resources transmitted from the cloud server module and processes the raw data of simplified musical notation, staff notation, lyrics, and track timing. It converts the raw simplified musical notation data into dynamically renderable structured simplified musical notation note data and the raw staff notation data into dynamically renderable structured staff notation note data. Both types of data include information such as note position, pitch, duration, and beat correspondence. The module also breaks down the lyrics into phrases and sentences, establishing a mapping between lyrics and the timeline. Synchronization Control Module: As the core control unit of the system, it is used to establish the temporal synchronization relationship between dynamic simplified musical notation, dynamic staff notation, lyrics, standard audio accompaniment, and user vocal performance audio; based on the audio track timing data, it acquires the playback progress of the standard audio accompaniment in real time, synchronously driving the flow of simplified musical notation notes, the flow of staff notation notes, and the scrolling display of lyrics; at the same time, it receives vocal performance analysis data transmitted by the audio acquisition and analysis module, dynamically adjusts the display rhythm of the dual notation and lyrics, and achieves real-time synchronization between vocal performance, dual notation, and lyrics; Dynamic Display Module: Receives drive signals from the synchronization control module to achieve synchronized visual display of simplified musical notation, staff notation, and lyrics. It employs a three-layer rendering method, setting up a staff notation layer, a simplified musical notation layer, and a lyrics display layer within the same display interface. The dual-notation display layer highlights the currently sung note in real time and uses the same dynamic cursor to indicate the singing progress. It supports personalized settings such as dual-notation scaling, custom colors, and individual hiding / showing of musical notation. The lyrics display layer scrolls word by word or line by line, precisely synchronized with the progress of the dual-notation display layer. Simultaneously, dynamic pitch lines can be overlaid to visually compare the user's real-time singing pitch with the standard pitch of the dual-notation. Audio Acquisition and Analysis Module: This module is used to acquire the user's vocal audio in real time through the audio acquisition device (such as a microphone) of the user terminal module, analyze and process the audio signal, and extract key data such as singing pitch and rhythm. The extracted data is compared with the standard pitch and rhythm in the original data of the simplified musical notation and the original data of the staff notation to generate analysis results such as pitch deviation and rhythm deviation. The analysis results are then transmitted to the synchronization control module and displayed to the user in a visual manner through the dynamic display module.

[0007] Furthermore, in the song resource library of the cloud server module, the original data of the simplified musical notation, the original data of the staff notation, the lyrics text data, and the timing data of the audio tracks for each song are generated and correlated and calibrated by professional notation software to ensure the timing consistency of the four types of data. The original data of the simplified musical notation supports MusicXML and JPW formats, and the original data of the staff notation supports multiple mainstream music notation formats such as MusicXML and MIDI formats. The data parsing module is compatible with the parsing and processing of multiple formats, realizing the synchronous parsing and conversion of dual-spectrum data.

[0008] Furthermore, the synchronization control module employs a MIDI timing-based synchronization algorithm. Through precise timestamp matching, it controls the synchronization error between dynamic musical notation, dynamic staff notation, lyrics, and standard audio accompaniment to within 0.03 milliseconds, ensuring that the displayed rhythm is completely consistent with the accompaniment rhythm. Simultaneously, it supports a dynamic rhythm adjustment function. When the audio acquisition and analysis module detects that the user's singing rhythm deviation exceeds a preset threshold, the synchronization control module can adaptively adjust the display progress of the dual scores and lyrics to guide the user back to the correct rhythm.

[0009] Furthermore, the dynamic display module supports switching between multiple display modes, including a three-element fusion mode of "simplified musical notation + staff notation + lyrics", a "simplified musical notation + lyrics" mode, a "staff notation + lyrics" mode, a "simplified musical notation only" mode, a "staff notation only" mode, and a "lyrics only" mode. Users can choose according to their own musical background. The dual-music display layers support independent customization of note size, color, and music score background. The lyrics display layer supports adjustments to font, color, and scrolling speed to adapt to the visual habits and usage needs of different users.

[0010] Furthermore, the audio acquisition and analysis module employs an AI pitch analysis algorithm, which, combined with the pitch fluctuation characteristics of human vocal performance, performs noise reduction processing on the acquired vocal audio signal before pitch extraction, thereby improving the accuracy of pitch analysis. The analysis results are synchronously fed back to the corresponding notes in the bispectral notation through a dynamic display module using color markings. For example, correctly sung notes are displayed in green, notes that are too low are displayed in blue, notes that are too high are displayed in red, and rhythm lag / lead is visually displayed through dynamic cursor offset, ensuring consistency in bispectral feedback.

[0011] A method for synchronously and dynamically displaying simplified musical notation, standard musical notation, and lyrics in a karaoke setting, based on the aforementioned system, includes the following steps: S1: The user selects the target karaoke song through the user terminal module, sets the display mode (such as "simplified notation + staff notation + lyrics"), and sends a song resource acquisition request to the cloud server module; S2: After receiving the request, the cloud server module extracts the standard audio accompaniment, lyrics text data, simplified musical notation data, staff notation data, and audio track timing data of the target song from the song resource library and transmits them to the user terminal module. S3: The data parsing module parses and processes the received resources to generate structured lyrics timing data, simplified musical notation timing data, and staff notation timing data, ensuring that the timelines of the three types of data correspond accurately. S4: The user starts karaoke playback through the user terminal module. The synchronization control module starts standard audio accompaniment playback based on the audio track timing data, and establishes a unified synchronization association between the accompaniment playback progress and the lyrics timing data, the simplified musical notation timing data, and the staff notation timing data. S5: Driven by the synchronization control module, the dynamic display module displays dynamic simplified musical notation, dynamic staff notation, and lyrics in real time according to the display mode set by the user. The same dynamic cursor links the highlighted notes of the two scores and the scrolling of lyrics to indicate the singing progress. At the same time, the audio acquisition and analysis module starts human voice acquisition, extracts the pitch and rhythm data of the user's singing in real time, and compares it with the standard data of the two scores to generate analysis results. S6: The synchronization control module dynamically adjusts the display progress of the musical score and lyrics based on the analysis results of the audio acquisition and analysis module to ensure consistency with the user's singing rhythm; the dynamic display module synchronously feeds back the pitch and rhythm deviations to the musical score and lyrics interface in the form of color markings and cursor offsets. S7: After the karaoke session, the system can generate a singing analysis report. Based on the pitch and rhythm analysis data throughout the session, combined with the note deviations corresponding to the two scores, it provides users with targeted suggestions for singing improvement. Beneficial effects

[0012] Compared with the prior art, the present invention has the following beneficial effects: The present invention achieves precise synchronization of simplified musical notation and staff notation with lyrics, karaoke accompaniment, and vocal performance, filling the gap in the lack of dual-notation fusion guidance in existing karaoke systems. It not only meets the needs of users with weak music foundation for the intuitiveness of simplified musical notation, but also adapts to the usage habits of professional users with staff notation, covering the needs of users at all levels, effectively reducing the difficulty of singing and improving the accuracy of singing. This invention addresses the core needs of karaoke scenarios by optimizing the dual-score fusion display logic. Through designs such as three-layer rendering, free switching between multiple display modes, and independent personalized settings for dual scores, it differs from existing dual-score technologies that are only applicable to musical instrument performances, significantly improving the user's interactive experience and scenario adaptability. This invention combines the characteristics of human vocal performance, optimizes audio acquisition and analysis algorithms, and achieves synchronous visual feedback on pitch and rhythm deviations in the bispectral and lyrics. With the dynamic rhythm adaptive adjustment function, it can guide users to correct singing deviations in a timely manner and help users of different skill levels improve their singing level in a targeted manner. The system architecture of this invention has strong compatibility and can be seamlessly integrated into existing online karaoke platforms, offline KTV song selection systems, and home karaoke devices. The song resource library can be expanded with bispectral data based on existing sheet music resources without the need to reconstruct the existing system, making it easy to achieve industrial application and promotion. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the module structure of the system in an embodiment of the present invention; the relationships between the modules in the diagram are as follows: the cloud server module is communicatively connected to the user terminal module and the data parsing module, respectively, for transmitting song resources; the user terminal module is connected to the synchronization control module, the audio acquisition and analysis module, and the dynamic display module, respectively, to realize command interaction, audio acquisition, and information display; the output end of the data parsing module is connected to the synchronization control module, for transmitting structured time-series data; the synchronization control module communicates bidirectionally with the dynamic display module and the audio acquisition and analysis module, respectively, to realize synchronization drive and data feedback.

[0014] Figure 2 This is a schematic diagram of the process steps of the method in an embodiment of the present invention.

[0015] Figure 3 This is a schematic diagram of the interface for synchronous and dynamic display of musical notation and lyrics in an embodiment of the present invention; (Note: The attached diagram is for illustrative purposes only, and the actual implementation can be designed according to specific requirements) Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0017] like Figure 1As shown, this embodiment provides a system for synchronously and dynamically displaying musical notation, staff notation, and lyrics in a karaoke scenario, including a cloud server module, a user terminal module, a data parsing module, a synchronization control module, a dynamic display module, and an audio acquisition and analysis module. The cloud server module stores a massive amount of song resources. Each song is associated with standard audio accompaniment, lyrics text data, original music score data in MusicXML format, original staff notation data in MusicXML format, and track timing data. The timing of these four types of data is calibrated using professional notation software. The user terminal module uses a smartphone and supports operations such as song selection, playback control, notation switching, and parameter adjustment. It also features microphone audio acquisition and screen display functions. The data parsing module uses a JavaScript parsing engine to quickly parse the dual-music score data and lyrics text data, generating structured timing data. The synchronization control module uses a MIDI timing-based synchronization algorithm to ensure that the synchronization error is controlled within 0.03 milliseconds. The dynamic display module uses Canvas technology to achieve three-layer rendering, supporting the synchronized display of dual scores and lyrics, as well as personalized settings. The audio acquisition and analysis module integrates an AI pitch analysis algorithm, which can achieve noise reduction and accurate analysis of human voice audio, and synchronously feed it back to the dual scores.

[0018] like Figure 2 As shown in the figure, this embodiment provides a method for synchronously and dynamically displaying simplified musical notation, staff notation, and lyrics based on the above system. The specific steps are as follows: S1: Users search for and select the target song "Jasmine Flower" through the karaoke APP (user terminal module) on their smartphones, select the "simplified notation + staff notation + lyrics" display mode through the "notation switching" button on the interface, and click the "karaoke" button to send a resource acquisition request to the cloud server module; S2: After receiving the request, the cloud server module extracts the standard audio accompaniment, lyrics text data, MusicXML format simplified score raw data, MusicXML format staff score raw data and track timing data of "Jasmine Flower" from the song resource library, and transmits them to the user terminal module through the network. S3: The data parsing module parses the received resources, converts the original simplified musical notation data into structured simplified musical notation note timing data containing note positions, pitches, and durations, converts the original staff notation data into structured staff notation note timing data containing staff tables, note positions, and clef information, splits the lyrics text data into several phrases such as "What a beautiful jasmine flower," and establishes a precise mapping relationship between the three types of data and the same time axis; S4: When the user clicks "Start Singing", the synchronization control module starts playing the standard audio accompaniment. Based on the audio track timing data, the accompaniment playback progress is associated with the simplified musical notation timing data, the staff notation timing data, and the lyrics timing data to determine the initial synchronization timestamp. S5: Under the drive of the synchronization control module, the dynamic display module displays a three-layer interface on the mobile phone screen (such as...). Figure 3 As shown in the image, the upper layer is the staff notation display layer 1, the middle layer is the numbered musical notation display layer 2, and the lower layer is the lyrics display layer 3. The red dynamic cursor 4 moves from left to right with the accompaniment progress, synchronously highlighting the corresponding notes 5 (displayed in green) in the double score. The lyrics display layer 3 synchronously scrolls to display "What a beautiful jasmine flower", precisely matching the progress of the double score. At the same time, the audio acquisition and analysis module acquires the user's vocal audio through the mobile phone microphone, performs noise reduction processing, extracts pitch data, and compares it with the standard pitch of the double score. When the system detects that the pitch of a note sung by the user is too low, the dynamic display module marks the corresponding notes in both scores in blue, and the dynamic pitch line 6 displays the deviation between the current pitch and the standard pitch. The synchronization control module fine-tunes the scrolling speed of the scores and lyrics according to the user's singing rhythm to ensure that the user's singing and the display progress are consistent. If the user wants to view the staff score separately, they can quickly switch to the "staff score + lyrics" mode through the score switching button 8. S7: After the karaoke session ends, the system generates a performance report based on the full analysis data. Combining the discrepancy between the two musical scores, it points out problems such as "the pitch of the high notes is too low" and "the tempo of the chorus is a bit too fast," and provides improvement suggestions such as "strengthening high note practice" and "following the discrepancy cursor to control the rhythm."

[0019] In this embodiment, users can independently customize the size and color of the notes in the simplified musical notation and the standard musical notation (e.g., set the standard notes in the simplified musical notation to black, the standard notes in the standard musical notation to dark gray, and the highlighted notes in both musical notation to green) through the settings interface of the karaoke app. They can also adjust the font and scrolling speed of the lyrics and switch the display mode as needed, effectively improving the singing experience and accuracy for users with different musical backgrounds.

Claims

1. A system for synchronously and dynamically displaying musical notation, staff notation, and lyrics in a karaoke setting, characterized in that, The system includes a cloud server module, a user terminal module, a data parsing module, a synchronization control module, a dynamic display module, and an audio acquisition and analysis module. The cloud server module stores a song resource library containing standard audio accompaniment, lyrics text data, original numbered musical notation data, original staff notation data, and track timing data. It also receives requests from the user terminal module and transmits song resources. The user terminal module receives user operation commands, displays relevant information, and acquires vocal audio. The data parsing module parses the song resources, generating structured lyrics timing data, numbered musical notation timing data, and staff notation timing data. The synchronization control module establishes the timing synchronization relationship between dynamic numbered musical notation, dynamic staff notation, lyrics, standard audio accompaniment, and vocal audio, driving the collaborative work of each module. The dynamic display module provides layered, visual, and synchronized display of numbered musical notation, staff notation, and lyrics, along with deviation feedback. The audio acquisition and analysis module acquires and analyzes vocal audio, generating pitch and rhythm deviation data and synchronously feeding it back to the dual-spectrum display.

2. The system according to claim 1, characterized in that, The original data of simplified musical notation, staff notation, lyrics, and track timing in the song resource library have been calibrated by professional notation software. The original data of simplified musical notation supports MusicXML and JPW formats, and the original data of staff notation supports MusicXML and MIDI formats. The data parsing module is compatible with multiple formats.

3. The system according to claim 1, characterized in that, The synchronization control module adopts a MIDI timing-based synchronization algorithm, with synchronization error controlled within 0.03 milliseconds, and supports dynamic rhythm adjustment function, which can adaptively match the user's singing rhythm.

4. The system according to claim 1, characterized in that, The dynamic display module supports switching between multiple display modes such as "simplified musical notation + staff notation + lyrics", "simplified musical notation + lyrics", and "staff notation + lyrics". The dual-music display layer supports independent personalized settings, and the lyrics display layer supports adjustments to font, color, and scrolling speed.

5. The system according to claim 1, characterized in that, The audio acquisition and analysis module uses an AI pitch analysis algorithm to first reduce noise in the human voice audio signal and then extract pitch data. The deviation results are visualized and fed back synchronously on the bispectral and lyrics through color marking and cursor offset.

6. A method for synchronously and dynamically displaying musical notation, staff notation, and lyrics in a karaoke scenario, implemented based on the system described in any one of claims 1-5, characterized in that... Includes the following steps: S1: The user selects the target song through the user terminal module, sets the display mode, and sends a resource acquisition request to the cloud server module; S2: The cloud server module transmits the standard audio accompaniment, dual-spectrum raw data, lyrics text data, and audio track timing data of the target song to the user terminal module; S3: The data parsing module parses the resources and generates structured timing data; S4: The synchronization control module starts playing the accompaniment and establishes a unified synchronous association for multiple data. S5: The dynamic display module synchronously displays the dual scores and lyrics, and the audio acquisition and analysis module collects and analyzes the vocal data. S6: The synchronization control module adaptively adjusts the display progress, and the dynamic display module synchronously provides feedback on deviation information. S7: A singing analysis report is generated after the karaoke session ends.