Audio Annotation on Musical Scores via Voice Input
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies do not efficiently allow instructions or comments to be recorded in association with display content representing a temporal transition of a music performance, particularly on portable devices like tablet terminals, making it difficult to mark positions on musical scores and input text while playing an instrument.
Innovation Solution
A method and apparatus that display content representing a temporal transition of a music performance, allowing users to record audio comments by designating time positions on the display, storing the recording time and position association, and enabling easy reproduction of audio comments at desired times, with the option to move and reposition marks on the display content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a musical score is displayed on a tablet terminal and the user attempts to input text or marks by traditional methods, then the musical score can be displayed, but the text input and mark designation require considerable time and labor, and it is difficult to execute while holding a musical instrument
Solution Approach 1:
The patent replaces traditional mechanical text input methods (keyboard typing, menu selection) with voice-based audio input. The user can speak instructions naturally while holding the instrument, and the system converts speech to text and associates it with the displayed musical score position, eliminating the need for manual text input operations.
Solution Approach 2:
The patent introduces audio recording and recognition technology as an intermediary between the user's intent and the musical score annotation. Instead of directly inputting text, the user speaks, and the audio system mediates by converting speech to text and automatically associating it with the current score position, streamlining the annotation process.
2Productivity
If traditional mark designation methods are used on displayed musical score, then marks can be added to symbols, but the process requires selecting mark types and manual positioning, which is time-consuming and labor-intensive
Solution Approach 1:
The system automatically performs mark designation based on audio input and current display position. When the user speaks an instruction, the system automatically creates the appropriate mark at the relevant position on the musical score without requiring the user to manually select mark types or position them, making the system serve itself in the annotation process.
Solution Approach 2:
The system prepares the annotation environment in advance by displaying the musical score with interactive positions ready for annotation. The audio input system is pre-configured to capture and process instructions, and the mark association logic is pre-established, so that when the user speaks, the entire annotation process executes automatically without requiring step-by-step user configuration.
3Loss of information
If audio recording is performed without association with display content time positions, then audio can be recorded, but the audio cannot be precisely linked to specific moments in the musical performance representation
Solution Approach 1:
The system continuously monitors the display position and timing information from the musical score representation, and uses this feedback to automatically associate audio recordings with the correct time positions. The audio recording function receives real-time feedback about the current score position and synchronizes the recording timestamp accordingly, ensuring precise linkage without complex manual configuration.
Data Source
AI summary
In an ensemble performance, performance lesson or the like, a musical score (display content) is displayed and input audio is recorded. The input audio includes notes, comments or the like uttered by a human player or an instructor. During recording of the input audio is received a user input designating a desired time position on the musical score displayed on the display device. A recording time of the input audio based on a time point at which the user input has been received and the time position on the musical score designated by the user input are stored into a storage device in association with each other. An icon is displayed in association with the time position on the musical score designated by the user input. Once the icon is selected, voice based on the recording time recorded in association with the icon is reproduced to sound the comments, etc.


