Information processing method, program, and information processing device
The information processing method converts musical notation into acoustic signals, addressing the complexity and space issues of braille scores, enabling accessible and efficient auditory understanding of musical content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- YAMAHA CORP
- Filing Date
- 2022-02-10
- Publication Date
- 2026-05-26
AI Technical Summary
Braille musical scores require specialized knowledge and take up three times the paper surface area, making them difficult for beginners and small children to understand and time-consuming to read.
An information processing method that generates acoustic signals representing sounds related to musical notation using a computer system, allowing users to understand musical scores through sound.
Enables understanding of musical scores without the need for specialized knowledge and reduces reading time by presenting content audibly, making it accessible to beginners and children.
Smart Images

Figure 0007865019000001 
Figure 0007865019000002 
Figure 0007865019000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for assisting in understanding the content of musical scores.
Background Art
[0002] Conventionally, braille musical scores have been used to enable visually impaired persons to understand the content of musical scores. For example, Patent Document 1 below discloses a musical score automatic braille translation system that translates musical score data into braille using a computer and prints out the translated musical score data on a braille typewriter to generate a braille musical score.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Understanding braille musical scores requires specialized knowledge, which poses a problem that it is difficult for beginners and small children to understand. Also, when comparing a normal musical score with a braille musical score, the braille musical score requires about three times the paper surface area to represent the same content score, which poses a problem that it takes time to read. Considering the above circumstances, one aspect of the present disclosure aims to present the content of a musical score by sound.
Means for Solving the Problems
[0005] To solve the above problems, an information processing method according to one aspect of the present disclosure is realized by a computer system, and based on musical score data representing a musical score including one or more performance symbols, an acoustic signal representing a sound related to the performance symbol is generated.
[0006] A program according to one aspect of this disclosure causes a computer system to function as a generation unit that generates an acoustic signal representing a sound related to a musical notation based on musical score data representing a musical score that includes one or more musical notations.
[0007] An information processing device according to one aspect of the present disclosure includes a generation unit that generates an acoustic signal representing a sound related to a musical score, based on musical score data representing a musical score including one or more musical scores. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram illustrating the configuration of the information processing device 10 according to the first embodiment. [Figure 2] This diagram illustrates the structure of the data stored in the storage device 12. [Figure 3] This is a diagram showing the types of musical symbols. [Figure 4] This is a block diagram illustrating the functional configuration of the control device 11. [Figure 5] This diagram illustrates the instruction reception screen provided by the instruction reception unit 30. [Figure 6] This diagram illustrates the instruction reception screen provided by the instruction reception unit 30. [Figure 7] This diagram illustrates the instruction reception screen provided by the instruction reception unit 30. [Figure 8] This diagram illustrates the instruction reception screen provided by the instruction reception unit 30. [Figure 9] This diagram illustrates the instruction reception screen provided by the instruction reception unit 30. [Figure 10] This diagram schematically illustrates the processing of the text generation unit 32. [Figure 11] This is an example diagram of a musical score. [Figure 12] This diagram schematically shows the timing of when text-to-speech is read aloud. [Figure 13] This diagram schematically shows the timing of when text-to-speech is read aloud. [Figure 14] This diagram illustrates the display screen during musical score reading. [Figure 15] This flowchart illustrates the specific steps involved in the process by which the control device 11 executes the music reading application. [Figure 16] This diagram illustrates the instruction reception screen provided by the instruction reception unit 30. [Figure 17] This diagram illustrates the display screen while the table of contents display function is running. [Figure 18] This is a block diagram illustrating the functional configuration of the control device 11A in the third embodiment. [Modes for carrying out the invention]
[0009] A: First Embodiment Figure 1 is a block diagram illustrating the configuration of an information processing device 10 according to the first embodiment. The information processing device 10 is a computer system comprising a control device 11, a storage device 12, a sound collection device 13, a sound emission device 14, an operating device 15, and a display device 16. The information processing device 10 can be implemented as an information terminal such as a smartphone, tablet terminal, or personal computer. In this embodiment, the information processing device 10 is assumed to be a smartphone. The information processing device 10 can be implemented as a single device, or as a group of devices configured separately from each other (for example, a client-server system).
[0010] The control device 11 consists of one or more processors that control each element of the information processing device 10. For example, the control device 11 consists of one or more types of processors such as a CPU (Central Processing Unit), SPU (Sound Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit).
[0011] The storage device 12 is one or more memories that store the program PG (see FIG. 2) executed by the control device 11 and various data used by the control device 11. The storage device 12 is composed of, for example, a known recording medium such as a magnetic recording medium or a semiconductor recording medium, or a combination of multiple types of recording media. Note that a portable recording medium detachable from the information processing device 10 or a recording medium (e.g., cloud storage) to which the control device 11 can perform writing or reading via a communication network may be used as the storage device 12.
[0012] The sound collection device 13 detects ambient sound (air vibrations) and outputs it as an acoustic signal. The sound collection device 13 is, for example, a microphone. Note that a sound collection device 13 separate from the information processing device 10 may be connected to the information processing device 10 by wire or wirelessly.
[0013] The sound playback device 14 reproduces the sound represented by the acoustic signal. The sound playback device 14 is, for example, a speaker or headphones. Note that a D / A converter that converts the acoustic signal from digital to analog and an amplifier that amplifies the acoustic signal are omitted from the illustration for convenience. Also, a sound playback device 14 separate from the information processing device 10 may be connected to the information processing device 10 by wire or wirelessly.
[0014] The operation device 15 is an input device that receives instructions from the user. The operation device 15 is, for example, an operator operated by the user or a touch panel T that detects contact by the user. In this embodiment, it is assumed that the touch panel T is used as the operation device 15. In this case, the touch panel T also serves as the operation device 15 and the display device 16 described later. Note that an operation device 15 (e.g., a mouse or a keyboard) separate from the information processing device 10 may be connected to the information processing device 10 by wire or wirelessly.
[0015] The display device 16 displays images under the control of the control device 11. For example, various display panels such as liquid crystal display panels or organic EL (electroluminescence) panels can be used as the display device 16. Alternatively, a separate display device 16 may be connected to the information processing device 10 by wire or wireless connection.
[0016] Figure 2 is a diagram illustrating the configuration of data stored in the storage device 12. Figure 3 is a diagram showing the types of musical symbols. As shown in Figure 2, the storage device 12 stores a program PG executed by the control device 11, audio data VD, musical score data SD, and symbol text data TD. In this embodiment, program PG is a program for executing a musical score reading application. The musical score reading application is an application that generates acoustic signals indicating sounds related to various information written in the musical score corresponding to the musical score data SD, and reproduces said acoustic signals. More specifically, the musical score reading application reads aloud the text corresponding to the musical symbols, enabling the user to understand the contents of the musical score using their hearing. Hereinafter, in this embodiment, reading a musical score corresponds to generating acoustic signals indicating sounds related to the musical symbols contained in the musical score, and reproducing said acoustic signals.
[0017] The audio data VD is data for generating synthesized speech that reads musical scores. The audio data VD is a speech synthesis library containing multiple phoneme fragments. Each phoneme fragment is a single phoneme (e.g., a vowel or consonant), which is the smallest unit of linguistic meaning, or a phoneme chain formed by linking multiple phonemes. In this embodiment, the audio data VD includes male voice data representing a male voice and female voice data representing a female voice. In addition, in this embodiment, the audio data VD includes Japanese voice data that pronounces Japanese and English voice data that pronounces English. That is, the audio data VD includes at least four combinations of two genders and two languages. By using multiple types of audio data VD, it is possible to change the type of speech that reads aloud, for example, depending on the part of the musical score or the type of symbol in the score, thereby improving convenience.
[0018] The musical score data SD is data representing the musical score of a piece of music. The musical score data SD is distributed, for example, by being distributed via a network from a distribution device such as a web server (not shown), or by a recording medium containing the musical score data SD being sold in a store. The musical score data SD is general data that can be obtained regardless of whether the user is hearing-impaired or not. In this embodiment, the musical score data SD describes the content of the musical score of a piece of music using a specific data description language. Specifically, the musical score data SD is a file for musical score representation (for example, a MusicXML format file) in which elements of the musical score, such as musical symbols, are expressed as logical information.
[0019] The musical score data SD is distributed to the information processing device 10 via a communication network such as the internet from, for example, a distribution device (typically a web server), and then stored in the storage device 12. Multiple musical score data SDs may be stored in the storage device 12. Generally, one musical score data SD is created for each musical piece.
[0020] A musical score is a representation of a piece of music using musical symbols, including performance markings. As shown in Figure 3, musical symbols include note symbols, clefs, time signatures, key signatures, and performance markings. Note symbols include notes, rests, and accidentals added to notes in the score. Clefs are written at the left end of the staff and specify the relationship between the position on the staff and the pitch of the note. Time signatures specify the number of beats in a measure and the type of note that constitutes one beat. Key signatures are a set of accidentals used to specify the key of the piece.
[0021] Performance markings are added to musical scores to indicate nuances that cannot be expressed by notes and rests alone. Performance markings include tempo markings such as adagio and andante, expression markings such as affettuoso and agitato, dynamic markings such as fortissimo and pianissimo, articulation markings such as tenuto and staccato (hereinafter referred to as "articulation markings"), repeat markings such as da capo and segno, ornamental markings such as trill and turn, abbreviation markings such as ottava alta and ottava bassa, and technique markings indicating instrument-specific playing techniques such as pedal and pizzicato. In this embodiment, finger numbers, which specify the fingers to use when playing the notes written in the score, are also included in the performance markings.
[0022] As shown in Figure 2, the musical score data SD contains attribute information B for each musical symbol that makes up the target musical piece. Attribute information B is information that defines the musical attributes of each musical symbol and includes a beat identifier B1 and a symbol identifier B2. Beat identifier B1 is information that specifies the temporal position of the musical symbol in the target musical piece. If the musical symbol is a note symbol or performance mark, the number of beats from the beginning of the target musical piece to the musical symbol (for example, the beat number counted with an eighth note as one beat) is preferably used as beat identifier B1. Symbol identifier B2 is information for identifying the type of musical symbol. For example, if the musical symbol is a note symbol, symbol identifier B2 includes the note name (note number) and note value. The note name represents the pitch of the note, and the note value represents the duration of the note on the musical score. If the musical symbol is something other than a note symbol, a string indicating the name of the musical symbol is preferably used as symbol identifier B2.
[0023] Furthermore, the musical score data SD includes tempo information TP, which specifies the tempo of the music indicated by the score. The tempo information TP includes, for example, the number of beats per minute (unit time) and the type of note that constitutes one beat.
[0024] Furthermore, the sheet music data SD includes sheet music image data MD. The sheet music image data MD is data representing the image of the sheet music of the target song (hereinafter referred to as "sheet music image"). Specifically, an image file (for example, a PDF file) that represents the sheet music image as a planar image in raster or vector format is suitable as the sheet music image data MD.
[0025] Symbol text data TD is data containing text corresponding to musical symbols written in a musical score. Symbol text data TD includes a symbol identifier C1, a name text C2, and a semantic text C3. Symbol identifier C1 is information for identifying the type of musical symbol and is in the same format as symbol identifier B2 in attribute information B. The symbol identifier C1 corresponding to a note symbol may consist only of the note name.
[0026] Name text C2 is text indicating the name of the symbol identified by symbol identifier C1. Symbol identifier C1 and name text C2 may be the same string. If the musical symbol is a note symbol, name text C2 is "Do", "Re", etc. If the musical symbol is a performance mark, name text C2 is "Crescendo", "Forte", etc. In this embodiment, the sheet music reading application is capable of reading sheet music in multiple languages. For example, the languages that can be selected during reading are Japanese or English. Therefore, name text C2 includes Japanese text indicating the name of the musical symbol and English text indicating the name of the musical symbol.
[0027] Semantic text C3 is text that indicates the meaning of the symbol identified by symbol identifier C1. For example, if the name of the musical symbol is "adagio," then semantic text C3 is "slowly." If the musical symbol is a note symbol, semantic text C3 may not be provided. Semantic text C3 also includes Japanese text that indicates the meaning of the musical symbol and English text that indicates the meaning of the musical symbol.
[0028] Furthermore, a non-verbal notification sound may be used as the name text C2 corresponding to a rest. In this case, when generating the acoustic signal, an acoustic signal representing a non-verbal notification sound is generated as the sound related to the rest. A non-verbal notification sound is a sound that does not have linguistic meaning, and includes, for example, a metronome sound, a click sound, a beep sound, etc. By using a non-verbal notification sound as the sound related to a rest, the user can immediately understand that the sound corresponds to a rest when the notification sound is played. In this embodiment, a click sound is used as the name text C2 corresponding to a rest. The click sound corresponding to a quarter rest is "click", and the click sound corresponding to an eighth note is "click". In addition, the click sounds corresponding to rests are not limited to "click" and "click", but may be, for example, "hmm" for a quarter rest and "ooh" for an eighth rest.
[0029] Additionally, as the name text C2 corresponding to the rest, you may use words that are used to indicate rests when keeping time, such as "un" for a quarter rest and "u" for an eighth rest.
[0030] Figure 4 is a block diagram illustrating the functional configuration of the control device 11. The control device 11 generates acoustic signals representing sounds related to musical symbols that make up a musical score (hereinafter referred to as "symbol sounds") based on musical score data SD representing a musical score. The musical score represented by the musical score data SD contains one or more performance symbols. Therefore, it can also be said that the control device 11 generates acoustic signals representing sounds related to performance symbols based on musical score data SD representing a musical score containing one or more performance symbols. Sounds related to performance symbols are, for example, sounds that indicate the name of the performance symbol, or sounds that indicate words corresponding to the meaning of the performance symbol. In this embodiment, the name of the performance symbol corresponds to the performance symbol name text C2, and words corresponding to the meaning of the performance symbol correspond to the performance symbol meaning text C3. In addition, in this embodiment, the musical score includes note symbols in addition to performance symbols. Therefore, the control device 11 generates acoustic signals representing sounds related to performance symbols and sounds related to note symbols. Sounds related to note symbols are, for example, sounds that indicate the name of the note indicated by the note symbol.
[0031] The acoustic signal is a signal that causes the sound emission device 14 to reproduce symbolic sounds. The control device 11 executes a program PG stored in the storage device 12 to generate and reproduce the acoustic signal, thereby realizing multiple functions (instruction receiving unit 30, text generation unit 32, speech synthesis unit 34, performance analysis unit 38, output control unit 40).
[0032] The instruction receiving unit 30 receives instructions from the user for the operating device 15. The instruction receiving unit 30 displays a screen for receiving user instructions on the touch panel T, for example. The user inputs instructions by touching the screen displayed on the touch panel T.
[0033] Figures 5 to 9 illustrate the instruction reception screens provided by the instruction reception unit 30. When the sheet music reading application is launched, the instruction reception unit 30 displays the instruction reception screen SC1 for selecting the song to be read aloud on the touch panel T, as shown in Figure 5, for example. The reception screen SC1 displays, for example, NA1 to NA5, which indicate the data names of the sheet music data SD stored in the storage device 12. The user specifies the sheet music data SD to be read aloud by touching the display NA1 to NA5 corresponding to the desired sheet music data SD. In the example in Figure 5, the sheet music data "yyy.xml" corresponding to display NA2 is specified. When the user touches the OK button BT in this state, the selection of the sheet music data "yyy.xml" is confirmed. In subsequent reception screens, the user's selection instruction is confirmed by touching the OK button BT. Note that instead of the displays NA1 to NA5 indicating the data names, the title of the song corresponding to the sheet music data SD may be displayed.
[0034] When the musical score data SD is specified, the instruction receiving unit 30 displays the instruction receiving screen SC2 for selecting the musical staff to be read aloud from the musical score data SD on the touch panel T, as shown in Figure 6, for example. The instruction receiving screen SC2 displays options NB1 and NB2 for specifying the musical staff to be read aloud from the grand staff. Option NB1 specifies the reading aloud of the right-hand musical staff located in the upper staff. Option NB2 specifies the reading aloud of the left-hand musical staff located in the lower staff. The user specifies the musical staff to be read aloud by checking at least one of the checkboxes CK for option NB1 or NB2.
[0035] When a musical staff to be read aloud is specified, the instruction receiving unit 30 displays a receiving screen SC3 for selecting the type of symbol to be read aloud on the touch panel T, as shown in Figure 7, for example. The receiving screen SC3 displays options NC1 to NC11 for specifying the type of symbol to be read aloud. Option NC1 specifies the reading of musical note symbols. Option NC2 specifies the reading of performance markings. Of the musical symbols shown in Figure 3, clefs, time signatures, and key signatures rarely change within a single musical score, so in this embodiment, they are not included as continuously read aloud targets. On the other hand, the user may also be able to specify whether or not to include clefs, time signatures, and key signatures as read aloud targets, similar to musical note symbols and performance markings.
[0036] Furthermore, for performance markings, it is possible to specify in more detail the types of markings to be read aloud. Option NC3 specifies the reading of tempo markings. Option NC4 specifies the reading of expression markings. Option NC5 specifies the reading of dynamic markings. Option NC6 specifies the reading of articulation markings. Option NC7 specifies the reading of repeat signs. Option NC8 specifies the reading of ornamental markings. Option NC9 specifies the reading of omission marks. Option NC10 specifies the reading of technique markings. Option NC11 specifies the reading of finger numbers.
[0037] In other words, a musical score represented by the score data SD contains multiple performance markings, and each of these performance markings belongs to one of several classifications. These classifications correspond to the types of performance markings shown in Figure 3. The instruction receiving unit 30 accepts the selection of at least one of the multiple classifications of performance markings. The control device 11 generates an acoustic signal for the performance markings belonging to one or more of the selected classifications from among the multiple performance markings.
[0038] When an item to be read aloud is specified, the instruction receiving unit 30 displays the instruction receiving screen SC4 for selecting settings when reading aloud the musical score data SD on the touch panel T, as shown in Figure 8, for example. In the upper area E1 of the receiving screen SC4, options ND1 and ND2 are displayed for specifying the information to be output when reading the musical score aloud. Option ND1 specifies that only the musical score should be read aloud. That is, option ND1 specifies that only the audio should be output. Option ND2 specifies that in addition to reading the musical score, the musical score image should be displayed. That is, option ND2 specifies that both audio and an image should be output. Options ND1 and ND2 can be selected using radio buttons. The user specifies the information to be output when reading the musical score aloud by touching the radio button corresponding to option ND1 or the radio button corresponding to option ND2.
[0039] Furthermore, in the lower area E2 of the reception screen SC4, options NE1 to NE4 are displayed for specifying the tempo of reading the musical score. Option NE1 specifies that the score will be read at the tempo specified in the musical score. If option NE1 is selected, it may not be possible to read all the symbols specified in reception screen SC3, depending on the relationship between the reading tempo and the number of syllables to be read. In this case, the number of symbols to be read will be reduced accordingly. Option NE2 specifies that all the symbols specified in reception screen SC3 will be read, regardless of the tempo specified in the musical score. Option NE3 specifies that the score will be read at a tempo synchronized with the user's performance of the musical score. Option NE4 allows the user to specify any tempo. In the example shown in the diagram, the user specifies the reading tempo by specifying the number of beats per minute.
[0040] In addition, in relation to specifying the reading tempo, it may be possible to set, for example, the speech time per syllable during reading (in other words, the number of syllables read per unit of time). For example, if option NE1 is selected, the shorter the speech time per syllable, the more symbols specified on the reception screen SC3 can be read aloud. Also, if option NE2 is selected, the shorter the speech time per syllable, the faster the reading can be completed.
[0041] When the OK button BT on the reception screen SC4 is touched, the instruction reception unit 30 displays a reception screen SC5 for further setting selection instructions on the touch panel T, as shown in Figure 9, for example. In the upper area E3 of the reception screen SC5, options NF1 to NF2 for specifying the language to be used when reading the musical score aloud are displayed. Option NF1 specifies reading in Japanese. Reading in Japanese means, for example, using "Do, Re, Mi, Fa, So, La, Si" as note names, and using Japanese pronunciation when pronouncing the names of musical notation. Option NF2 specifies reading in English. Reading in English means, for example, using "C, D, E, F, G, A, B" as note names, and using English pronunciation when pronouncing the names of musical notation. Note that languages other than Japanese and English may be specified on the reception screen SC5. In this case, the symbol text data TD includes the name text C2 and semantic text C3 of the language in question.
[0042] Additionally, in the reception screen SC5, area E4 in the middle displays options NG1 to NG2 for specifying what to read when the musical notation is read aloud. Option NG1 specifies that the name of the musical notation will be read aloud. If option NG1 is selected, the name text C2 from the symbol text data TD will be read aloud when the musical notation is read aloud. Option NG2 specifies that the words or phrases indicating the meaning of the musical notation will be read aloud. If option NG2 is selected, the meaning text C3 from the symbol text data TD will be read aloud when the musical notation is read aloud.
[0043] Furthermore, in the lower area E5 of the reception screen SC5, options NH1 to NH2 for specifying the type of voice to be read aloud are displayed. For example, if both the right-hand staff and the left-hand staff are specified as the staff to be read aloud on the reception screen SC2, text indicating multiple musical note symbols may be read aloud simultaneously. In this embodiment, in order to make it easier for the user to identify the text being read aloud, the voice for reading the right-hand staff and the voice for reading the left-hand staff can be specified as different types. That is, the instruction reception unit 30 can individually set the type of voice for each of the multiple parts of the musical piece. In this embodiment, the voice types can be specified as male voice and female voice. Option NH1 specifies that the voice for reading the right-hand staff should be either male or female. Option NH2 specifies that the voice for reading the left-hand staff should be either male or female.
[0044] Furthermore, it is possible to specify different types of voices for reading out musical notation, for example, and for reading out performance markings. Also, if four or more types of voices can be specified, it is possible to specify that the musical notation on the staff for the right hand, the musical notation on the staff for the left hand, the performance markings on the staff for the right hand, and the performance markings on the staff for the left hand be read out by different voices.
[0045] Furthermore, if the sound emission device 14 is a stereo speaker, it may be configured to output the sound of reading the staff for the right hand from the right speaker, and the sound of reading the staff for the left hand from the left speaker. Also, if the sound emission device 14 is a stereo speaker, it may be specified that the speaker outputting the sound of reading the musical notation and the speaker outputting the sound of reading the performance markings are different.
[0046] In addition, for example, when reading out a chord, the user may be able to choose whether to read out each note that makes up the chord individually or to read out the chord name that corresponds to the chord. In this case, for example, the name text C2 of the symbol text data TD may store the text indicating the note names of each note that make up the chord, and the semantic text C3 may store the text indicating the chord name of the chord.
[0047] Once these settings are complete and the OK button BT in Figure 9 is pressed, the control device 11 starts the process of generating an audio signal. The touch panel T also displays a button, for example, to instruct the start of reading the musical score (hereinafter referred to as the "start playback button"). The user presses the start playback button at an appropriate time to begin reading the musical score.
[0048] The text generation unit 32 shown in Figure 4 generates text that represents the content of the musical score. Figure 10 is a schematic diagram showing the processing of the text generation unit 32. The text generation unit 32 reads the musical score data SD specified on the reception screen SC1 shown in Figure 5 (S1). The text generation unit 32 classifies the musical score data SD into right-hand data, which represents the staff for the right hand, and left-hand data, which represents the staff for the left hand (S2). Of the right-hand data and left-hand data, the data corresponding to the staff to be read aloud, as specified on the reception screen SC2 shown in Figure 6, is the subject of subsequent processing.
[0049] The musical staff data to be read aloud includes attribute information B for all types of musical symbols (S3). The text generation unit 32 extracts the attribute information B of the symbols to be read aloud, as specified on the reception screen SC3 shown in Figure 7, from the musical staff data to be read aloud, and arranges them in chronological order based on the beat identifier B1 (S4).
[0050] The text generation unit 32 compares the symbol identifier B2 of the extracted attribute information B with the symbol identifier C1 of the symbol text data TD, and reads either the name text C2 or the semantic text C3 corresponding to the symbol identifier C1 (S5). Whether to read the name text C2 or the semantic text C3 is determined by which of the options NG1 or NG2 is selected on the reception screen SC5 shown in Figure 9. Furthermore, whether to read Japanese text or English text is determined by which of the options NF1 or NF2 is selected on the reception screen SC5. In the diagram, these selections are referred to as "specification of content to be read aloud". The read texts are arranged in the same order (chronological order) as the attribute information B. Through the above processing, text indicating the content of the musical score (hereinafter referred to as "read-aloud text") is generated (S6).
[0051] Furthermore, the text generation unit 32 adds timing labels to the spoken text (S7). Timing labels are information that identifies the timing of when the spoken text is read. Here, even if the spoken text has the same content, the reading speed will differ depending on the tempo specified for reading the musical score on the reception screen SC4 shown in Figure 8. Therefore, the text generation unit 32 adds timing labels to the spoken text according to the tempo setting for reading the musical score.
[0052] Figure 11 is an example of a musical score. Figures 12 and 13 schematically show the timing of the text being read aloud. The musical score G shown in Figure 11 shows the first two measures of the musical score data (yyy.xml) specified as the text to be read aloud on the reception screen SC1 shown in Figure 5. The musical score G includes both the score for the right hand and the score for the left hand. Based on the tempo information TP, the musical score G is specified to have 120 beats per minute, with one quarter note representing one beat.
[0053] For example, in the reception screen SC4 shown in Figure 8, if reading at the tempo of the music (option NE1) is selected, a timing label is added to ensure that reading occurs at the timings shown in Figure 12. Figure 12 shows the right-hand reading sound indicating the reading of the right-hand score, and the left-hand reading sound indicating the reading of the left-hand score. A time axis is shown between the right-hand reading sound and the left-hand reading sound. One division on the time axis (t1) is based on the shortest note in the score G, which is an eighth note. Based on the tempo of the score G described above, one division on the time axis (t1) is 0.25 seconds.
[0054] This explains the right-hand reading function. When reading begins, first, for 0.25 seconds (period P1), "mezzo piano," "staccato," and "Mi" are read aloud. The order in which "mezzo piano," "staccato," and "Mi" are read aloud is arbitrary. Next, for 0.25 seconds (period P2), "staccato" and "Fa" are read aloud. Comparing period P1 and period P2, period P1 has more syllables read per unit time, so the reading speed for period P1 needs to be faster than for period P2. Next, for 0.5 seconds (period P3), "So" is read aloud. Comparing period P2 and period P3, period P3 has fewer syllables read per unit time, so the reading speed for period P3 needs to be slower than for period P2. Next, for 0.5 seconds (period P4), "Mi" is read aloud. Comparing period P3 and period P4, the number of syllables read per unit time is the same for both, so the reading speeds for period P3 and period P4 are approximately the same. Next, during a 0.5-second period (period P5), a click sound ("click") indicating a quarter rest is read aloud.
[0055] In other words, if the reading at the tempo of the music is specified, the control device 11 generates an acoustic signal such that when the target point in time, which is progressing at a speed corresponding to the tempo specified by the tempo information TP, reaches the point corresponding to the performance marking, the sound related to that performance marking is emitted.
[0056] Furthermore, if, for example, during period P1, the number of syllables to be read per unit of time is large and may be difficult for the user to understand, the text generation unit 32 may reduce the amount of text to be read aloud. For example, if the speech time for one syllable during reading is predetermined, it is possible to determine whether or not to read the text aloud at the time the text is generated. The text generation unit 32 determines whether or not the text can be read aloud within the given time based on the generated words, the tempo of the music, and the speech time for one syllable.
[0057] If the text generation unit 32 determines that it cannot read the text within the allotted time, it may remove the text corresponding to the musical notation from the text to be read and only read the musical note notation. Alternatively, if the text generation unit 32 determines that it cannot read the text within the allotted time, it may overlap the sounds of each notation. For example, during period P1, it may overlap the sounds of "mezzo piano," "staccato," and "mi."
[0058] Furthermore, if the text generation unit 32 determines that reading cannot be performed within the time limit during a period in which multiple musical notations are to be read aloud (for example, period P1 in Figure 12), it may choose to read aloud some of the multiple musical notations and exclude the remaining ones. Taking period P1 as an example, the text generation unit 32 may choose to read aloud either "mezzo piano" or "staccato" and not the other. In other words, when "mezzo piano" and "staccato" are pronounced together, the control device 11 selects either "mezzo piano" or "staccato" and includes the text corresponding to the selected musical notation in the text to be read aloud. "Mezzo piano" is an example of a sound related to the first musical notation, and "staccato" is an example of a sound related to the second musical notation.
[0059] The instruction receiving unit 30 may also allow the user to set the priority of reading aloud for each classification of musical notation. In this case, the text generation unit 32 will delete the text of musical notation belonging to the lowest priority classification from the text to be read aloud.
[0060] Alternatively, for example, non-verbal sounds corresponding to each musical notation may be predetermined, and if it is determined that the reading cannot be done within the allotted time, the non-verbal sounds may be included in the text to be read aloud instead of the text corresponding to the musical notation.
[0061] This section describes the left-hand reading function. The left-hand score shows a triad consisting of three pitch classes. When reading begins, the chord in the first measure is read aloud. The reading of the chord in the first measure continues for 2 seconds. In this embodiment, the chord is not pronounced as "Do-Mi-So," but rather "Do," "Mi," and "So" are read aloud as separate notes. If the readings of "Do," "Mi," and "So" begin simultaneously, the user may not be able to distinguish between the notes. Therefore, as shown in Figure 12, the start timing of the readings of "Do," "Mi," and "So" may be slightly staggered. "Slightly" means, for example, a time shorter than the time corresponding to the shortest note in the score. In the case of score G, a time shorter than 0.25 seconds, which corresponds to an eighth note, corresponds to "slightly."
[0062] In this way, by reading the musical score aloud in time with the tempo of the music, users can grasp the rhythm of the music along with the pitch of the notes indicated in the score.
[0063] Furthermore, for example, if "Read all items" (option NE2) is selected on the reception screen SC4 shown in Figure 8, timing labels are added to ensure that the items are read aloud at the timings shown in Figure 13. When reading all items, the number of syllables read per unit time is kept constant. Figure 13 shows only the reading sounds for the right hand, omitting the reading sounds for the left hand. In Figure 13, the symbol Mti (where i is an integer from 1 to 9) indicates the metronome sound. The user can understand the divisions of beats by the metronome sound Mti.
[0064] When the reading begins, the phrase "Measure 1" is read first, indicating that the first measure is about to be read. Next, the metronome tone Mt1 is sounded, followed by the readings of "mezzo piano," "staccato," "Mi," "staccato," and "Fa." Then, the metronome tone Mt2 is sounded, followed by the reading of "So." Similarly, between subsequent metronome tones Mti, the text indicating the pitch of the note symbol is read aloud. Note that the last note of the second measure is a half note and has a length of two beats. In this case, after the metronome tone Mt7, the reading of "Mi" is extended and continued as "Mi," and the metronome tone Mt8 is sounded again, repeating "Mi." After the reading of "Mi" has finished, the metronome tone Mt9 is sounded.
[0065] For the left-hand reading sounds, it is preferable to synchronize them with the right-hand reading sounds. For example, after the metronome sound Mt1 is pronounced, the reading of "Do," "Mi," and "So" continues until just before the phrase "Measure 2" is read aloud. The timing of the start of the reading of "Do," "Mi," and "So" may be slightly staggered as described above. After that, after the metronome sound Mt5 is pronounced, the reading of "C," "Re," and "So" continues until just before the metronome sound Mt9 is pronounced.
[0066] Thus, the control device 11 may generate an acoustic signal indicating the sound produced for a musical notation, regardless of the tempo of the music. This allows the user to understand all of the specified types of notations and to fully grasp the contents of the musical score.
[0067] Furthermore, in the reception screen SC4 shown in Figure 8, if the user specifies reading aloud at a tempo synchronized with their performance of the musical score (option NE3), the timing of the reading aloud cannot be predicted in advance, so the text generation unit 32 does not need to add timing labels. Also, if the user specifies an arbitrary tempo (option NE4), the tempo of the musical piece in the explanation for reading aloud at the tempo of the musical piece (option NE1) described above should be replaced with the tempo specified by the user, and the same processing should be performed.
[0068] In the reception screen SC4 shown in Figure 8, option NE4 allowed the user to specify an arbitrary tempo by specifying the number of beats per minute. However, this is not the only option; for example, the user may specify the progress of the reading using an operator. The operator may be, for example, an operation button displayed on the touch panel T. Also, if, for example, the information processing device 10 is connected to a musical instrument played by the user, the operator may be a component of the instrument. For example, if the instrument is a piano, the pedal can be used as the operator. In this case, for example, if the user presses the damper pedal once, the reading may progress in units of one note symbol or one measure, and if the user presses the soft pedal once, the reading may be reversed in units of one note symbol or one measure, and so on.
[0069] The speech synthesis unit 34 shown in Figure 4 generates an acoustic signal using the text to be read aloud generated by the text generation unit 32 and the audio data VD. The speech synthesis unit 34 is an example of a generation unit. The speech synthesis unit 34 sequentially selects the phonetic elements corresponding to the text to be read aloud from among the multiple phonetic elements contained in the audio data VD, adjusts the pitch of each phonetic element, and then connects them to generate an acoustic signal. The pitch of sounds related to musical notes in the text to be read aloud may match the pitch of the musical note, or it may be a predetermined pitch. The acoustic signal generated by the speech synthesis unit 34 is supplied to the sound emission device 14, and sounds representing musical notation are reproduced from the sound emission device 14.
[0070] The performance analysis unit 38 analyzes the user's performance of an instrument. For example, the performance analysis unit 38 analyzes the position in the musical piece where the user is playing the instrument (performance position). For example, the performance analysis unit 38 captures the sound of the instrument's performance using the sound acquisition device 13 and analyzes the pitch and duration of the sound. The performance analysis unit 38 compares the pitch of the analyzed sound with the pitch of the notes on the musical score data SD and sequentially analyzes the performance position in the musical piece at each of multiple points in time on the time axis.
[0071] Furthermore, if the instrument is an electronic instrument, for example, the performance analysis unit 38 may obtain performance information from the electronic instrument indicating the operating state of the instrument. The operating state, for example, if the electronic instrument is an electronic piano, includes the identifier of the pressed key and the pressure applied. In this case, the performance analysis unit 38 uses the performance information to map the performance position at each point in time onto the musical score.
[0072] In the first embodiment, the performance analysis unit 38 only needs to operate when the user specifies reading the score aloud at a tempo synchronized with the user's performance of the musical score (option NE3) on the reception screen SC4 shown in Figure 8.
[0073] The output control unit 40 controls the output of sound based on an acoustic signal and the output of a musical score image based on musical score image data MD. For example, in the reception screen SC4 shown in Figure 8, if only the reading of the musical score is specified (option ND1), the output control unit 40 causes the sound output device 14 to play the sound represented by the acoustic signal generated by the speech synthesis unit 34. Also, in the reception screen SC4, if both the reading of the musical score and the display of a musical score image are specified (option ND2), the output control unit 40 causes the sound output device 14 to play the sound represented by the acoustic signal generated by the speech synthesis unit 34, and also displays the musical score image data MD on the display device 16.
[0074] Figure 14 illustrates the display screen during musical score reading. For example, after the user has finished inputting selection instructions to the reception screen SC5 shown in Figure 9, when the user touches the start playback button displayed on the touch panel T, the sound of the musical score being read, which is the sound represented by the acoustic signal, is output from the sound emission device 14, and the display on the touch panel T, which is the display device 16, switches to the display screen SC6 shown in Figure 14. The display screen SC6 shows a message 601 indicating that the musical score reading is playing, a musical score image 602, a pause button 604, a fast forward button 606, a rewind button 608, a listen-back button 610, and an end button 612.
[0075] The musical score image 602 is an image displaying the musical score image data MD contained in the musical score data SD to be read aloud. A bar 603 indicating the reading position is superimposed on the musical score image 602. The output control unit 40 scrolls the musical score image 602 based on the timing labels attached to the text to be read aloud. At this time, the output control unit 40 adjusts the scrolling speed of the musical score image 602 so that the musical symbols being read aloud and the bar 603 are superimposed. Alternatively, instead of displaying the reading position with the bar 603, the musical symbols to be read aloud may be highlighted.
[0076] In Figure 14, a musical score using a five-line staff is shown as an example of musical score image 602, but it is not limited to this; for example, a piano roll may be displayed as musical score image 602. Also, in the reception screen SC4 shown in Figure 8, if only musical score reading is specified (option ND1), musical score image 602 will not be displayed.
[0077] The pause button 604, fast-forward button 606, rewind button 608, listen again button 610, and end button 612 accept operations related to the reading of the musical score. When the pause button 604 is pressed, the output control unit 40 pauses the reading of the musical score. When the fast-forward button 606 is pressed, the output control unit 40 fast-forwards the reading of the musical score. For example, if the fast-forward button 606 is touched once, the output control unit 40 changes the reading position to the beginning of the measure following the measure containing the current reading position. When the rewind button 608 is pressed, the output control unit 40 rewinds the reading of the musical score. For example, if the rewind button 608 is touched once, the output control unit 40 changes the reading position to the beginning of the measure containing the current reading position. When the listen again button 610 is pressed, the output control unit 40 restarts reading the musical score from the beginning. In other words, the output control unit 40 changes the reading position to the beginning of the first measure of the musical score being read aloud. When the end button 612 is pressed, the output control unit 40 stops reading the musical score being read aloud.
[0078] Furthermore, the user may be allowed to specify the starting position for reading the musical score. For example, while the output control unit 40 is waiting for an instruction to start reading, it displays a start playback button and a musical score image on the touch panel T. Referring to Figure 14, the start playback button is displayed in place of message 601 on the display screen SC6. The user scrolls the musical score image 602 so that the bar 603 aligns with the position on the musical score image 602 where they want to start reading. When the start playback button is touched in this state, reading begins from the position on the musical score image 602 where the bar 603 aligns. Alternatively, the user may specify the starting position for reading the musical score by, for example, specifying the measure number where they want to start reading.
[0079] Furthermore, when the user selects to read aloud at a tempo synchronized with their performance of the musical score (option NE3) on the reception screen SC4 shown in Figure 8, the output control unit 40 adjusts the timing of the read-aloud sound output based on the performance position analyzed by the performance analysis unit 38. The output control unit 40 adjusts the output speed of the read-aloud sound so that, for example, a position predetermined beats ahead of the performance position on the musical score is read aloud. The predetermined beat may be specified by the user.
[0080] As another example, the output control unit 40 may, for example, when the performance position is the Nth measure (where N is an integer greater than or equal to 1), read out the musical symbols contained in the N+1 measure just before the performance of the Nth measure ends. Just before the performance of the Nth measure ends means, for example, after the last note of the Nth measure has been played. This is modeled after, for example, a choir conductor previewing the lyrics to be sung next for the choir members.
[0081] Figure 15 is a flowchart illustrating the specific steps involved in the process by which the control device 11 executes the music reading application. For example, the music reading application is launched in response to instructions from the user to the operating device 15.
[0082] When the sheet music reading application is launched, the control device 11 (instruction receiving unit 30) displays the reception screens SC1 to SC5 shown in Figures 5 to 9 to receive various specifications from the user regarding the reading of the sheet music (S100). These specifications include, for example, specifying the sheet music data SD to be read aloud, specifying the type of symbols to be read aloud, and specifying the language to be read aloud.
[0083] The control device 11 (text generation unit 32) generates spoken text based on the various specifications received in S100, the musical score data SD, and the symbol text data TD (S102). The control device 11 (speech synthesis unit 34) uses the spoken text and the speech data VD to generate an audio signal that reads the spoken text by speech synthesis (S104). The control device 11 (output control unit 40) waits until the user instructs it to read the musical score (S106: NO). When the user instructs the control device 11 (output control unit 40) to read the musical score (S106: YES), it plays the audio signal from the sound emission device 14 (S108) and terminates the process according to this flowchart.
[0084] Note that the generation of the text to be read aloud (S102) and the generation of the sound signal (S104) may be performed after the user gives the instruction to read the musical score aloud (S106: YES).
[0085] As described above, in the first embodiment, an acoustic signal representing the sound related to a performance mark is generated based on a musical score data SD containing one or more performance marks. Therefore, the performance marks included in the musical score can be grasped by hearing, making it easier for people with visual impairments, beginners who are not accustomed to reading music, or young children to understand the musical score.
[0086] Furthermore, in the first embodiment, the sound related to the performance marking is either a sound indicating the name of the performance marking or a sound indicating a phrase corresponding to the meaning of the performance marking. When the sound related to the performance marking is a sound indicating the name of the performance marking, the notation on the musical score can be accurately grasped. Also, when the sound related to the performance marking is a sound indicating a phrase corresponding to the meaning of the performance marking, even if the user has little knowledge of performance marking and cannot understand its meaning from the name of the performance marking alone, they can still grasp the content indicated by the musical score.
[0087] Furthermore, in the first embodiment, the sound corresponding to the musical notation is produced at a timing corresponding to the tempo of the music. This makes it easier for the user to understand the position of the musical notation within the music, thereby improving convenience.
[0088] Furthermore, in the first embodiment, when the sound related to the first performance marking and the sound related to the second performance marking are played simultaneously, either the first performance marking or the second performance marking is selected to generate the acoustic signal. This prevents the sound related to the first performance marking and the sound related to the second performance marking from being played simultaneously, thereby improving the audibility of the sounds related to the performance markings.
[0089] Furthermore, in the first embodiment, sounds corresponding to musical notation are produced regardless of the tempo of the music. This prevents overlapping sounds related to musical notation, thereby improving the audibility of the sounds related to musical notation.
[0090] Furthermore, in the first embodiment, an acoustic signal is generated for musical notation belonging to a selected classification from among multiple classifications. Therefore, the sound related to the musical notation required by the user can be selectively produced, thereby improving convenience.
[0091] Furthermore, in the first embodiment, in addition to sounds related to performance markings, an acoustic signal representing sounds related to musical note symbols is generated. Therefore, musical note symbols included in the score can be perceived aurally, making it even easier to understand the score.
[0092] Furthermore, in the first embodiment, a non-verbal notification sound is emitted as a sound related to a rest. This allows the user to immediately recognize that the sound emitted corresponds to a rest.
[0093] B: Second Embodiment A second embodiment will now be described. For elements whose function is the same as in the first embodiment in each of the embodiments described below, the same reference numerals as in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.
[0094] In the first embodiment, the information processing device 10 read aloud a musical score data SD specified by the user from among multiple musical score data SDs stored in the device. However, even if a list of data names for musical score data SDs is displayed, such as the reception screen SC1 shown in Figure 5, the user may not be able to identify the musical score data SD corresponding to the desired song. In the second embodiment, portions of multiple musical score data SDs are read aloud sequentially, enabling the user to identify the musical score data SD corresponding to the desired song. The function of sequentially reading portions of multiple musical score data SDs is hereinafter referred to as the "table of contents display function".
[0095] Figure 16 is an example of the instruction reception screen by the instruction reception unit 30. In the second embodiment, when the sheet music reading application is started, the instruction reception unit 30 displays a menu selection instruction reception screen SC7 on the touch panel T, for example, as shown in Figure 16. The reception screen SC7 displays options NI1 and NI2. Option NI1 specifies the reading of the sheet music data SD selected by the user, as shown in the first embodiment. When option NI1 is touched, the instruction reception unit 30 displays the reception screen SC1 shown in Figure 5 and receives the user's specification of the sheet music data SD to be read aloud.
[0096] Option NI2 specifies the execution of the table of contents display function. In option NI2, the table of contents display function is referred to as "melody table of contents". When option NI2 is selected, the text generation unit 32 generates text for reading aloud a portion of the musical score represented by each of the multiple musical score data SDs stored in the storage device 12. A portion of the musical score may include, for example, performance markings and note symbols.
[0097] A portion of a musical score is, for example, part or all of a specific structural section (hereinafter referred to as "structural section") among several sections that divide a musical piece according to its musical meaning. Structural sections include, for example, sections such as the intro, A section, B section, chorus, and outro. Specifically, the text generation unit 32 generates text for reading aloud the "chorus" structural section of a musical piece for each of the multiple musical score data SDs. Alternatively, the text generation unit 32 generates text for reading aloud the "intro" structural section (a predetermined number of measures at the beginning of the score) of a musical piece for each of the multiple musical score data SDs.
[0098] Furthermore, the score data SD may include information indicating the correspondence between positional information (e.g., measure number) within the score and structural sections. Also, when the table of contents display function is executed, the instruction receiving unit 30 may receive various instructions as in the first embodiment (see Figures 6 to 9). The speech synthesis unit 34 generates an acoustic signal using the text to be read aloud generated from each score data SD and the audio data VD. The output control unit 40 causes the sound output device 14 to reproduce the sound based on the acoustic signal.
[0099] Figure 17 illustrates the display screen during the execution of the table of contents display function. Display screen SC8 shows displays NA1 to NA5, which indicate the data names of the musical score data SD stored in the storage device 12. Displays NA1 to NA5 are arranged vertically, and parts of the musical score are read aloud in order, starting with the musical score data "xxx.xml" shown by display NA1. When the reading of the musical score data "xxx.xml" is finished, the reading of the musical score data "yyy.xml" shown by display NA2 begins. The display corresponding to the musical score data being read aloud (display NA2 in Figure 17) may be displayed with a different background color from the other displays.
[0100] Furthermore, in the display screen SC8 of Figure 17, the user may be allowed to select which of the multiple musical score data SDs stored in the storage device 12 are to be read aloud in the table of contents display function. In addition, the user may be allowed to specify the order in which the musical score data SDs are read aloud in the table of contents display function.
[0101] Furthermore, the display screen SC8 displays the same pause button 604, fast forward button 606, rewind button 608, listen again button 610, and end button 612 as the display screen SC6 shown in Figure 14. Even while the table of contents display function is running, the user can use the pause button 604, fast forward button 606, rewind button 608, listen again button 610, and end button 612 to perform operations related to the reading of the musical score.
[0102] In other words, in the second embodiment, the musical score data SD is the first musical score data, and the storage device 12 also stores second musical score data that is different from the first musical score data. The control device 11 generates a first sound signal that indicates the sound related to the performance markings and the sound related to the note markings contained in a part of the first musical score corresponding to the first musical score data, and generates a second sound signal that indicates the sound related to the performance markings and the sound related to the note markings contained in a part of the second musical score corresponding to the second musical score data. The control device 11 also sequentially plays the first sound signal and the second sound signal to the sound output device 14. For example, the first musical score data is the musical score data "xxx.xml", and the second musical score data is the musical score data "yyy.xml".
[0103] According to the second embodiment, the control device 11 selects a portion from each of the multiple musical score data and sequentially plays the sounds related to the performance markings and note symbols contained in the selected portion. This allows the user to easily understand which musical score each of the multiple musical score data SDs corresponds to, and to quickly select the desired musical score data SD from among the multiple musical score data SDs.
[0104] C: Third Embodiment A third embodiment will now be described. For elements whose function is the same as in the first embodiment in each of the embodiments described below, the same reference numerals as in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.
[0105] In the first embodiment, the information processing device 10 read aloud the musical score data SD. In the third embodiment, in addition to reading aloud the musical score data SD, the information processing device 10 assists the user in making their performance sound as close as possible to the notes indicated in the musical score.
[0106] Figure 18 is a block diagram illustrating the functional configuration of the control device 11A in the third embodiment. The control device 11A includes a performance evaluation unit 42 in addition to the configuration of the control device 11 (see Figure 4) according to the first embodiment. The performance evaluation unit 42 evaluates the user's performance of the instrument based on the analysis results from the performance analysis unit 38. In the first embodiment, the performance analysis unit 38 analyzed the position of the instrument played by the user. In the third embodiment, in addition to analyzing the position of the instrument, the performance analysis unit 38 also analyzes the loudness of the sound produced by the instrument.
[0107] The performance evaluation unit 42 evaluates whether the user's performance conforms to the musical symbols in the score. More specifically, the performance evaluation unit 42 detects the difference between the sound produced by the user playing the piece and the sound indicated by the musical symbols in the score representing the piece, and determines whether the difference falls outside a predetermined acceptable range.
[0108] For example, the evaluation of whether a performance follows the musical notation is performed by detecting the difference between the pitch of the sound played by the user and the pitch of the note on the musical score, and the difference between the duration of the sound played and the note value on the musical score. The performance evaluation unit 42 evaluates that the smaller the above difference, the more the performance follows the musical notation on the score, i.e., the higher the skill of the performance.
[0109] Users can set an acceptable range for differences based on their own playing skill level, for example. Generally, it is thought that the more skilled a user is, the smaller the acceptable range of differences will be. If there are any parts where the difference falls outside the acceptable range, the text generation unit 32 generates text that points out those parts. Specifically, it generates text that reads out the correct pitch and note value written in the score, as well as the pitch and note value in the user's performance, for example, "Right hand, measure 2, 'F, D, E' was written as 'F, E, D'." Such text is called "support text."
[0110] Furthermore, whether or not the performance follows the performance markings is determined for each performance marking. The performance evaluation unit 42 evaluates the performance by detecting the difference between the volume of the sound played by the user and the volume of the sound played in accordance with the dynamic markings, for example, when the performance marking is a dynamic marking. The performance evaluation unit 42 also evaluates the performance by detecting the difference between the duration of the sound played by the user and the duration of the sound played in accordance with the articulation markings, for example, when the performance marking is an articulation marking. The performance evaluation unit 42 evaluates the performance as being in accordance with the performance markings in the score, i.e., the higher the skill level of the performer, the smaller the difference.
[0111] The user sets an acceptable range for differences, for example, based on their own playing skill level. If there are any parts where the difference falls outside the acceptable range, the text generation unit 32 generates support text that points out those parts. Specifically, it reads aloud performance markings written in the score, such as "Right hand, measure 1, 'staccato, E, staccato, F,' the staccato is not strong enough," and generates support text indicating that the user's performance did not reflect the performance markings.
[0112] The speech synthesis unit 34 generates an acoustic signal using the support text and the audio data VD. The text indicating the pitch of a musical note symbol within the support text may be read aloud at a pitch corresponding to that note. The output control unit 40 causes the sound emission device 14 to reproduce the sound based on the acoustic signal.
[0113] The playback of the audio signal may be performed after the user has finished playing the piece, or it may be performed during the performance. If performed during the performance, the output control unit 40 may immediately play the support text if, for example, a difference occurs that falls outside the acceptable range. In this case, after playing the support text, the output control unit 40 may play an audio prompt to play the section where the difference outside the acceptable range occurred (the section pointed out by the support text) again. When the user plays the section pointed out by the support text, the performance evaluation unit 42 evaluates whether the performance conforms to the musical notation in the score and repeats the above process. This encourages the user to repeatedly practice sections they find difficult (sections where it is difficult to play in accordance with the musical notation in the score), and the user can efficiently learn to play the piece represented in the score.
[0114] Alternatively, for example, the user's performance could be recorded, and the portion of the recording corresponding to the section pointed out in the support text could be played back along with a reading of the support text.
[0115] Furthermore, when generating support text, performance markings may be read aloud at all times, regardless of whether there are differences. In this case, for example, if there is a large difference between the performance markings and the performance, the volume of the read-aloud voice may be increased (the larger the difference, the louder the read-aloud voice), so that the user can understand whether the performance is in accordance with the performance markings.
[0116] In other words, in the third embodiment, the control device 11 acquires the sound of the music played by the user, and detects the difference between the sound indicated by the musical symbols included in the musical score representing the music and the sound of the performance. If the difference falls outside a predetermined tolerance range, the control device 11 generates an acoustic signal indicating the sound related to the musical symbols included in the part of the musical score corresponding to the location where the difference occurred.
[0117] According to the third embodiment, the user can grasp the difference between their own performance and what is indicated in the musical score, and efficiently learn to perform the piece indicated in the score. Specifically, the control device 11 indicates the location of the difference in the musical score to the user by reading aloud the musical notation at the location where the difference occurred. This allows the user to intuitively grasp the location in the musical score where the difference occurred, compared to, for example, if only the location in the score (such as the measure number) was mechanically read aloud. In addition, the third embodiment verbalizes the content of the user's performance. For example, the control device 11 reads aloud the pitch and note value in the performance made by the user. This allows the user to objectively grasp the content of their own errors.
[0118] D: Variant The following are examples of specific modifications that may be added to each of the embodiments exemplified above. Multiple embodiments selected from the following examples may be merged as appropriate, provided they do not contradict each other.
[0119] (1) In each of the above-described embodiments, the speech synthesis unit 34 performed segment-type speech synthesis, but the method of speech synthesis is not limited to the above examples. For example, statistical model-type speech synthesis using a deep neural network or a statistical model such as an HMM (Hidden Markov Model) may be used.
[0120] (2) In each of the above forms, the musical score data SD was used to read aloud the sounds related to the symbols contained in the musical score. The use of the musical score data SD is not limited to this, and the musical score data SD may also be used to play back performance sounds. Specifically, for example, the information processing device 10 may play back the performance sounds of the left-hand staff of the musical score represented by the musical score data SD, and at the same time read aloud the musical symbols in the right-hand staff. The user practices playing with their right hand while listening to the sounds of the musical symbols being read aloud. By playing back the performance sounds of the left-hand staff, the user can efficiently learn the timing of playing with their right hand and the harmony of the piece.
[0121] (3) The information processing device 10 may be implemented by a server device that communicates with an information device such as a smartphone or tablet terminal. For example, the information processing device 10 receives a specification of musical score data SD from the information device and generates an acoustic signal by speech synthesis processing using the specified musical score data SD. The information processing device 10 transmits the acoustic signal generated by the speech synthesis processing to the information device. The acoustic signal is then played back by the information device.
[0122] (4) The functions of the information processing device 10 (instruction receiving unit 30, text generation unit 32, speech synthesis unit 34, performance analysis unit 38, output control unit 40, performance evaluation unit 42) are realized through the cooperation of one or more processors constituting the control device 11 and the program PG stored in the storage device 12, as described above.
[0123] (5) In each of the above-described configurations, the user visually inspects the items displayed on the touch panel T and performs touch operations on the touch panel T when making various settings or giving instructions in the sheet music reading application. However, this is not the only configuration; for example, information (such as selection items in the settings) may be presented to the user by voice reading. In addition, input from the user to the information processing device 10 may be performed by voice input. In particular, when the sheet music reading application is used by a visually impaired person, the voice-based configuration is effective.
[0124] The above program can be provided in a form stored on a computer-readable recording medium and installed on the computer. The recording medium is, for example, a non-transitory recording medium, such as an optical recording medium (optical disc) like a CD-ROM, but it also includes any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium except for transient propagation signals, and volatile recording media are not excluded. Furthermore, in a configuration in which a distribution device distributes the program via a communication network, the recording medium that stores the program in the distribution device corresponds to the aforementioned non-transitory recording medium.
[0125] E: Addendum From the forms exemplified above, the following configuration can be understood, for example.
[0126] An information processing device according to one aspect of this disclosure (Aspect 1) is implemented by a computer system and generates an acoustic signal representing a sound related to a musical notation based on musical score data representing a musical score that includes one or more musical notations.Therefore, the musical notations included in the musical score can be grasped by hearing, making it easier for people with visual impairments, beginners who are not accustomed to reading music, and young children to grasp the musical score.
[0127] In the specific example of Embodiment 1 (Embodiment 2), the sound related to the performance marking is either a sound indicating the name of the performance marking or a sound indicating a phrase corresponding to the meaning of the performance marking. In the above embodiments, when the sound related to the performance marking is a sound indicating the name of the performance marking, the description on the musical score can be accurately grasped. Furthermore, when the sound related to the performance marking is a sound indicating a phrase corresponding to the meaning of the performance marking, even if the user has little knowledge of performance markings and cannot understand the meaning of the performance marking by name alone, they can still grasp the content indicated by the musical score.
[0128] In a specific example of Embodiment 1 or Embodiment 2 (Embodiment 3), the musical score data includes tempo information specifying the tempo of the musical piece indicated by the score, and in generating the sound signal, the sound signal is generated such that when the target time point, which is the speed corresponding to the tempo specified by the tempo information, reaches the time point corresponding to the musical notation, a sound related to the musical notation is emitted. In the above embodiments, a sound corresponding to the musical notation is emitted at a timing corresponding to the tempo of the musical piece. Therefore, the user can more easily grasp the position of the musical notation within the musical piece, thereby improving convenience.
[0129] In any one specific example (Aspect 4) of aspects 1 to 3, the one or more performance symbols include a first performance symbol and a second performance symbol, and in the generation of the sound signal, when the sound related to the first performance symbol and the sound related to the second performance symbol are sounded together, either the first performance symbol or the second performance symbol is selected, and the sound information representing the sound related to the selected performance symbol is generated. In the above aspects, when the sound related to the first performance symbol and the sound related to the second performance symbol are sounded together, either the first performance symbol or the second performance symbol is selected, and the sound information is generated. Therefore, the sound related to the first performance symbol and the sound related to the second performance symbol are not sounded together, and the intelligibility of the sound related to the performance symbol can be improved.
[0130] In a specific example of Embodiment 1 (Embodiment 5), the musical score data includes tempo information specifying the tempo of the musical piece indicated by the score, and in the generation of the sound signal, an sound signal is generated that indicates the sound corresponding to the performance markings, regardless of the tempo of the musical piece. In the above embodiment, the sound corresponding to the performance markings is emitted regardless of the tempo of the musical piece. Therefore, the sounds corresponding to the performance markings are not emitted repeatedly, and the audibility of the sounds corresponding to the performance markings can be improved.
[0131] In any one specific example (Aspect 6) of Embodiments 1 to 5, the one or more performance symbols are multiple performance symbols, each of the multiple performance symbols belongs to one of the multiple classifications, and the selection of at least one of the multiple classifications is accepted. In the generation of the sound signal, the sound signal is generated for the performance symbols belonging to the one or more classifications selected from the multiple performance symbols. In the above embodiment, the sound signal is generated for the performance symbols belonging to the selected classification. Therefore, the sound related to the performance symbols that the user needs can be selectively produced, improving convenience.
[0132] In any one specific example (7th embodiment) of embodiments 1 to 6, the musical score includes musical notation in addition to the performance markings, and generating the acoustic signal includes generating the acoustic signal that represents the sound related to the performance markings and the sound related to the musical notation. In the above embodiments, an acoustic signal is generated that represents the sound related to the musical notation in addition to the sound related to the performance markings. Therefore, the musical notation included in the musical score can be grasped by hearing, making it even easier to grasp the musical score.
[0133] In a specific example of Embodiment 7 (Embodiment 8), in the generation of the acoustic signal, an acoustic signal representing a non-verbal notification sound is generated as the sound related to the rest. In the above embodiments, a non-verbal notification sound is used as the sound related to the rest. Therefore, the user can immediately recognize that the sound related to the rest corresponds to the sound when it is played.
[0134] In a specific example of Embodiment 7 or Embodiment 8 (Embodiment 9), the musical score data is first musical score data, and the generation of the sound signal includes the generation of a first sound signal indicating sounds related to the performance markings and sounds related to the note symbols contained in a portion of the first musical score corresponding to the first musical score data, and the generation of a second sound signal indicating sounds related to the performance markings and sounds related to the note symbols contained in a portion of the second musical score data corresponding to a second musical score data different from the first musical score data, and further includes sequentially playing the first sound signal and the second sound signal on a sound-emitting device. In the above embodiments, a portion is selected from each of the multiple musical score data, and the sounds related to the performance markings and note symbols contained in the selected portion are played sequentially. Therefore, the user can easily understand which musical score each of the multiple musical score data corresponds to, and can quickly select the desired musical score data from among the multiple musical score data.
[0135] A program according to one aspect of the present disclosure (Aspect 10) causes a computer system to function as a generation unit that generates an acoustic signal representing a sound related to a musical score, based on musical score data representing a musical score that includes one or more musical notations.
[0136] An information processing device according to one aspect of the present disclosure (Aspect 11) includes a generation unit that generates an acoustic signal representing a sound related to a musical score, based on musical score data representing a musical score that includes one or more musical scores.
[0137] Here, braille musical scores require approximately three times the amount of paper compared to regular musical scores to represent the same content, and reading them takes longer. Therefore, for example, if a user forgets the title of a song and wants to find the desired score based on its content, they have to spend time reading multiple scores, which presents a challenge.
[0138] An information processing device according to one aspect of the present disclosure is implemented by a computer system and generates a first sound signal indicating a sound related to a musical symbol contained in a portion of a first musical score corresponding to first musical score data containing one or more musical symbols, generates a second sound signal indicating a sound related to a musical symbol contained in a portion of a second musical score corresponding to second musical score data different from the first musical score data, which contains one or more musical symbols, and sequentially reproduces the first sound signal and the second sound signal to a sound-emitting device.
[0139] Furthermore, instrument instruction aims to enable students to correctly interpret musical symbols written in sheet music and to accurately play the notes indicated by those symbols. However, it is sometimes impossible for a performer to determine whether they have correctly interpreted the musical symbols in the sheet music or whether they are accurately playing the notes indicated by those symbols. For this reason, performers generally seek instruction from an instructor. However, it is not practical to have an instructor constantly by their side while practicing, and opportunities to receive feedback on their performance are limited.
[0140] An information processing device according to one aspect of the present disclosure is implemented by a computer system and acquires performance sounds, which are sounds produced when a user plays a musical piece; detects the difference between the performance sounds and the sounds indicated by musical symbols included in a musical score representing the musical piece; and if the difference falls outside a predetermined tolerance range, generates an acoustic signal indicating the sound related to the musical symbols included in the part of the musical score corresponding to the location where the difference occurred. [Explanation of Symbols]
[0141] 10... Information processing device, 11, 11A... Control device, 12... Memory device, 13... Sound collection device, 14... Sound emission device, 15... Operation device, 16... Display device, 30... Instruction reception unit, 32... Text generation unit, 34... Speech synthesis unit, 38... Performance analysis unit, 40... Output control unit, 42... Performance evaluation unit, PG... Program, SD... Score data, T... Touch panel, TD... Symbol text data, VD... Audio data.
Claims
1. Based on musical score data representing a musical score containing one or more performance markings, an acoustic signal representing a sound related to the performance marking is generated. The sounds related to the performance markings are sounds that indicate the name of the performance markings, sounds that indicate words or phrases corresponding to the meaning of the performance markings, or non-verbal sounds that correspond to the performance markings. In generating the aforementioned acoustic signals, the system generates an acoustic signal that represents a sound related to the performance markings, which is sounded at the tempo of the musical piece indicated by the musical score, or an acoustic signal that represents a sound related to the performance markings, which is sounded at a tempo synchronized with the user's performance of the musical score. Information processing methods implemented by computer systems.
2. The aforementioned musical score data includes tempo information that specifies the tempo of the musical piece indicated by the musical score, Generating an acoustic signal that indicates a sound related to a performance mark, which is to be played at the tempo of the musical piece indicated by the musical score, includes generating the acoustic signal such that the sound related to the performance mark is played when the target time point, which is progressing through the musical piece at a speed corresponding to the tempo specified by the tempo information, reaches the time point corresponding to the performance mark. The information processing method according to claim 1.
3. To generate an acoustic signal indicating a sound related to the musical notation that is produced at a tempo synchronized with the user's performance of the musical score, The acoustic signal is generated such that a sound is produced corresponding to a musical notation located a predetermined beat ahead of the performance position on the musical score, or The method includes generating the sound signal such that, when the performance position is in the Nth measure (where N is an integer of 1 or more), after the last note of the Nth measure is played, a sound related to the performance markings in the N+1th measure is produced. The information processing method according to claim 1 or 2.
4. The one or more performance markings mentioned above include a first performance marking and a second performance marking, In generating the aforementioned acoustic signal, when the sound related to the first performance mark and the sound related to the second performance mark are played simultaneously, either the first performance mark or the second performance mark is selected, and the acoustic signal representing the sound related to the selected performance mark is generated. The information processing method according to any one of claims 1 to 3.
5. The aforementioned one or more performance symbols are multiple performance symbols, Each of the aforementioned performance symbols belongs to one of several classifications, The system accepts the selection of at least one of the aforementioned multiple classifications. In generating the aforementioned sound signals, the sound signals are generated for the performance symbols belonging to the one or more classifications selected from the plurality of performance symbols. The information processing method according to any one of claims 1 to 4.
6. The aforementioned musical score includes, in addition to the performance markings, note symbols, Generating the aforementioned acoustic signal includes generating the acoustic signal that represents the sound related to the performance marking and the sound related to the musical note marking. The information processing method according to any one of claims 1 to 5.
7. In generating the aforementioned acoustic signals, an acoustic signal representing a non-verbal notification sound is generated as a sound related to a rest. The information processing method according to claim 6.
8. The aforementioned musical score data is the first musical score data, The generation of the aforementioned acoustic signal is Generation of a first sound signal that indicates the sound related to the performance markings and the sound related to the note markings included in a portion of the first musical score corresponding to the first musical score data, This includes generating a second sound signal that represents the sound related to the performance markings included in a portion of the second score corresponding to a second score data different from the first score data, and the sound related to the note markings, The present invention further includes sequentially reproducing the first acoustic signal and the second acoustic signal in a sound-emitting device. The information processing method according to claim 6 or 7.
9. Based on musical score data representing a musical score containing one or more performance markings, an acoustic signal representing a sound related to the performance marking is generated. The sounds related to the performance markings are sounds that indicate the name of the performance markings, sounds that indicate words or phrases corresponding to the meaning of the performance markings, or non-verbal sounds that correspond to the performance markings. In generating the aforementioned sound signals, a generation unit generates sound signals that represent sounds related to the performance markings, which are pronounced at the tempo of the musical piece indicated by the musical score, or sound signals that represent sounds related to the performance markings, which are pronounced at a tempo synchronized with the user's performance of the musical score. A program that makes a computer system function.
10. Based on musical score data representing a musical score containing one or more performance markings, an acoustic signal representing a sound related to the performance marking is generated. The sounds related to the performance markings are sounds that indicate the name of the performance markings, sounds that indicate words or phrases corresponding to the meaning of the performance markings, or non-verbal sounds that correspond to the performance markings. In generating the aforementioned sound signals, a generation unit generates sound signals that represent sounds related to the performance markings, which are pronounced at the tempo of the musical piece indicated by the musical score, or sound signals that represent sounds related to the performance markings, which are pronounced at a tempo synchronized with the user's performance of the musical score. An information processing device equipped with the following features.