Electronic musical instrument, method and program

The electronic musical instrument addresses the issue of excessive lyrics progression by using a method that considers time differences and chord discrimination time to control lyrics progression, achieving appropriate control during performances.

JP7673786B2Active Publication Date: 2025-05-09CASIO COMPUTER CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023214342
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-05-09
Estimated Expiration
2039-12-23

AI Technical Summary

Technical Problem

Existing electronic musical instruments struggle to properly control the progression of lyrics when multiple notes are played simultaneously, leading to excessive lyrics progression.

Method used

The electronic instrument employs a method where lyrics progression is controlled by determining the time difference between pitch specifications and using chord discrimination time to decide whether to maintain or progress the lyrics index.

Benefits of technology

This approach allows for appropriate control of lyrics progression during performances, ensuring that lyrics do not progress excessively even when multiple keys are pressed simultaneously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007673786000001
    Figure 0007673786000001
  • Figure 0007673786000002
    Figure 0007673786000002
  • Figure 0007673786000003
    Figure 0007673786000003
Patent Text Reader

Abstract

To appropriately control the advance of lyrics according to musical performance.SOLUTION: An electronic musical instrument according to an aspect of the present disclosure comprises a plurality of performance operators that are associated with pieces of pitch data different from each other, and a processor. The processor determines whether a chord is designated in response to a user operation to the plurality of performance operators, when determining that the chord is designated, instructs the utterance of a singing voice according to first lyrics at every pitch designated in response to the user operation, and when determining that the chord is not designated, instructs the utterance of the singing voice according to the first lyrics at one pitch designated in response to the user operation, and instructs the utterance of a singing voice according to second lyrics subsequent to the first lyrics at one remaining pitch designated in response to the user operation.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to electronic musical instruments, methods and programs. [Background technology]

[0002] In recent years, the use of synthetic voices has been expanding. In this context, if there were an electronic musical instrument that could not only play automatically, but also progress lyrics according to the key presses of the user (player) and output synthetic voice corresponding to the lyrics, it would be desirable because it would enable more flexible expression of synthetic voices.

[0003] For example, Patent Document 1 discloses a technique for progressing lyrics in synchronization with a performance based on a user's operation using a keyboard or the like. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 4735544 Summary of the Invention [Problem to be solved by the invention]

[0005] However, when multiple sounds can be generated simultaneously using a keyboard or the like, if the lyrics are advanced simply each time a key is pressed, the lyrics may advance too far when multiple keys are pressed simultaneously.

[0006] Therefore, one of the objects of the present disclosure is to provide an electronic musical instrument, method, and program that can appropriately control the progression of lyrics during performance. [Means for solving the problem]

[0007] An electronic musical instrument according to one aspect of the present disclosure includes: Among the pitches designated by the performance, if the time difference between the timing at which the pitch was designated in the past and the timing at which the pitch was designated most recently is within the chord discrimination time, Latest pitch According to the specification of The lyrics are determined to be maintained by not advancing the lyrics index. and when a chord of the same lyrics is sounded in both cases, and the lowest note being sounded does not change due to the latest pitch designated during the sounding of the chord and outside the chord discrimination time, the lyrics index is advanced to determine that lyrics are progressing. Control it as follows. Effect of the Invention

[0008] According to one aspect of the present disclosure, the progression of lyrics played can be appropriately controlled. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram showing an example of the external appearance of an electronic musical instrument 10 according to an embodiment. [Diagram 2] FIG. 2 is a diagram showing an example of a hardware configuration of a control system 200 of the electronic musical instrument 10 according to an embodiment. [Diagram 3] FIG. 3 is a diagram illustrating an example of the configuration of the voice learning unit 301 according to an embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the waveform data output unit 302 according to an embodiment. [Diagram 5] FIG. 5 is a diagram illustrating another example of the waveform data output unit 302 according to an embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of a flowchart of a lyrics progression control method according to an embodiment. [Figure 7] FIG. 7 is a diagram illustrating an example of a flowchart of a lyric progression determination process based on chord voicing. [Figure 8] FIG. 8 is a diagram showing an example of the lyric progression controlled using the lyric progression determination process. [Figure 9] FIG. 9 is a diagram illustrating an example of a flowchart of the synchronization process. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Singing two or more notes in a part that is originally composed of one syllable per note (syllable style) is also called melismatic singing. Melisma singing may be interpreted as fake singing, fist singing, etc.

[0011] The inventors came up with the lyric progression control method disclosed herein by noting that a characteristic of melisma is that the pitch can be freely changed while maintaining the previous vowel, when performing melismatic singing on an electronic musical instrument equipped with a singing voice synthesis sound source.

[0012] According to one aspect of the present disclosure, it is possible to control so that lyrics do not progress during a melisma. Also, even when multiple keys are pressed simultaneously, it is possible to appropriately control whether lyrics progress or not.

[0013] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the following description, the same parts are denoted by the same reference numerals. Since the same parts have the same names, functions, etc., detailed description will not be repeated.

[0014] In addition, in the present disclosure, "lyric progression", "lyric position progression", "singing position progression", etc. may be read as interchangeable terms. In addition, in the present disclosure, "lyrics are not advanced", "lyric progression is not controlled", "lyrics are held", "lyrics are suspended", etc. may be read as interchangeable terms.

[0015] (Electronic Musical Instruments) 1 is a diagram showing an example of the appearance of an electronic musical instrument 10 according to an embodiment. The electronic musical instrument 10 may include a switch (button) panel 140b, a keyboard 140k, a pedal 140p, a display 150d, and a speaker 150s.

[0016] The electronic musical instrument 10 is a device for receiving input from a user via controls such as a keyboard and switches, and for controlling performance, lyric progression, etc. The electronic musical instrument 10 may be a device having a function of generating sounds according to performance information such as MIDI (Musical Instrument Digital Interface) data. The device may be an electronic musical instrument (such as an electronic piano or synthesizer), or may be an analog musical instrument equipped with sensors and the like and configured to have the functions of the controls described above.

[0017] The switch panel 140b may include switches for operating the volume specification, settings such as sound source and tone, song (accompaniment) selection, song playback start / stop, song playback settings (tempo, etc.), and the like.

[0018] The keyboard 140k may have multiple keys as performance operators. The pedal 140p may be a sustain pedal that sustains the sound of a pressed key while the pedal is pressed, or a pedal for operating an effector that processes the tone, volume, etc.

[0019] In the present disclosure, the terms "sustain pedal," "pedal," "foot switch," "controller," "switch," "button," "touch panel," etc. may be interchangeable. In the present disclosure, "pressing a pedal" may be interchangeable with "operating a controller."

[0020] The keys may be called performance operators, pitch operators, timbre operators, direct operators, first operators, etc. The pedals may be called non-performance operators, non-pitch operators, non-timbre operators, indirect operators, second operators, etc.

[0021] The display 150d may display lyrics, musical scores, various setting information, etc. The speaker 150s may be used to emit sounds generated by playing.

[0022] The electronic musical instrument 10 may be capable of generating and converting at least one of MIDI messages (events) and Open Sound Control (OSC) messages.

[0023] The electronic musical instrument 10 may be referred to as a controller 10, a lyric progression controller 10, or the like.

[0024] The electronic musical instrument 10 may communicate with a network (such as the Internet) via at least one of wired and wireless communication (e.g., Long Term Evolution (LTE), 5th generation mobile communication system New Radio (5G NR), Wi-Fi (registered trademark), etc.).

[0025] The electronic musical instrument 10 may hold in advance vocal data (which may be called lyric text data, lyric information, etc.) relating to lyrics whose progression is to be controlled, or may transmit and / or receive the vocal data via a network. The vocal data may be text written in a musical score description language (e.g., MusicXML), may be written in a MIDI data storage format (e.g., Standard MIDI File (SMF) format), or may be text provided in a normal text file.

[0026] In addition, the electronic musical instrument 10 may obtain the content sung by the user in real time via a microphone or the like provided in the electronic musical instrument 10, and apply voice recognition processing to the content to obtain text data as singing voice data.

[0027] FIG. 2 is a diagram showing an example of a hardware configuration of a control system 200 of the electronic musical instrument 10 according to an embodiment.

[0028] A central processing unit (CPU) 201, a ROM (read only memory) 202, a RAM (random access memory) 203, a waveform data output section 211, a key scanner 206 to which the switch (button) panel 140b, keyboard 140k, and pedal 140p of Figure 1 are connected, and an LCD controller 208 to which an LCD (Liquid Crystal Display) as an example of the display 150d of Figure 1 is connected are each connected to a system bus 209.

[0029] A timer 210 for controlling the sequence of automatic performance may be connected to the CPU 201. The CPU 201 may be called a processor, and may include an interface with peripheral circuits, a control circuit, an arithmetic circuit, a register, and the like.

[0030] The functions of each device may be realized by loading specific software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations and control communication via the communication device 1004, reading and / or writing of data in the memory 1002 and storage 1003, etc.

[0031] 1 by executing a control program stored in ROM 202 while using RAM 203 as a working memory. In addition to the control program and various fixed data, ROM 202 may also store vocal data, accompaniment data, and song data including these.

[0032] The CPU 201 is equipped with a timer 210 used in this embodiment, which counts the progress of an automatic performance in the electronic musical instrument 10, for example.

[0033] The waveform data output unit 211 may include a sound source LSI (large scale integrated circuit) 204, a voice synthesis LSI 205, etc. The sound source LSI 204 and the voice synthesis LSI 205 may be integrated into one LSI.

[0034] The singing voice waveform data 217 and song waveform data 218 output from the waveform data output unit 211 are converted into an analog singing voice output signal and an analog musical sound output signal by D / A converters 212 and 213, respectively. The analog musical sound output signal and the analog singing voice output signal may be mixed in a mixer 214, and the mixed signal may be amplified by an amplifier 215 and then output from the speaker 150s or an output terminal.

[0035] A key scanner (scanner) 206 constantly scans the key-on / key-off state of the keyboard 140k in FIG. 1, the switch operation state of the switch panel 140b, the pedal operation state of the pedal 140p, and the like, and issues an interrupt to the CPU 201 to notify it of state changes.

[0036] The LCD controller 208 is an IC (integrated circuit) that controls the display state of an LCD, which is an example of the display 150d.

[0037] It should be noted that the above system configuration is merely an example and is not limited to this. For example, the number of each circuit included is not limited to this. Electronic musical instrument 10 may have a configuration that does not include some circuits (mechanisms), or may have a configuration in which the function of one circuit is realized by multiple circuits. It may also have a configuration in which the functions of multiple circuits are realized by one circuit.

[0038] Furthermore, electronic musical instrument 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc., and some or all of the functional blocks may be realized by the hardware. For example, CPU 201 may be implemented by at least one of these pieces of hardware.

[0039] <Generating Acoustic Models> Fig. 3 is a diagram showing an example of the configuration of a voice training unit 301 according to an embodiment. The voice training unit 301 may be implemented as a function executed by a server computer 300 that is externally present and separate from the electronic musical instrument 10 of Fig. 1. The voice training unit 301 may be built into the electronic musical instrument 10 as a function executed by the CPU 201, the voice synthesis LSI 205, or the like.

[0040] The voice training unit 301 and the waveform data output unit 302 described below that realize the voice synthesis in this disclosure may be implemented based on, for example, a statistical voice synthesis technique based on deep learning.

[0041] The speech training unit 301 may include a training text analysis unit 303 , a training acoustic feature extraction unit 304 , and a model training unit 305 .

[0042] In the voice learning unit 301, for example, recordings of a singer singing a number of songs in an appropriate genre are used as the learning singing voice data 312. Also, as the learning singing voice data 311, lyrics texts of each song are prepared.

[0043] The training text analysis unit 303 receives training singing data 311 including lyrics text and analyzes the data. As a result, the training text analysis unit 303 estimates and outputs a training language feature sequence 313 which is a discrete numerical sequence expressing phonemes, pitches, etc. corresponding to the training singing data 311.

[0044] The training acoustic feature extraction unit 304 inputs and analyzes training singing voice data 312 collected via a microphone or the like by a singer singing lyrics text corresponding to the training singing voice data 311 in accordance with the input of the training singing voice data 311. As a result, the training acoustic feature extraction unit 304 extracts and outputs a training acoustic feature sequence 314 expressing voice features corresponding to the training singing voice data 312.

[0045] In the present disclosure, the training acoustic feature sequence 314 and the acoustic feature sequence corresponding to the acoustic feature sequence 317 described later include acoustic feature data (which may be called formant information, spectral information, etc.) that models the human vocal tract, and vocal cord sound source data (which may be called sound source information) that models the human vocal cords. For example, Mel-cepstrum, Line Spectral Pairs (LSP), etc. can be used as the spectral information. For example, a fundamental frequency (F0) and a power value that indicate the pitch frequency of human voice can be used as the sound source information.

[0046] The model learning unit 305 estimates, by machine learning, an acoustic model that maximizes the probability that the training acoustic feature sequence 314 is generated from the training language feature sequence 313. That is, the relationship between the language feature sequence, which is text, and the acoustic feature sequence, which is speech, is represented by a statistical model called an acoustic model. The model learning unit 305 outputs model parameters that represent the acoustic model calculated as a result of the machine learning as the learning result 315. Therefore, the acoustic model corresponds to a trained model.

[0047] As the acoustic model represented by the learning result 315 (model parameters), a hidden Markov model (HMM) may be used.

[0048] The HMM acoustic model may learn how the characteristic parameters of the vocal voice, such as the vibration of the vocal cords and the vocal tract characteristics, change over time when a singer sings lyrics according to a certain melody. More specifically, the HMM acoustic model may be a model of the spectrum, fundamental frequency, and their time structures calculated from the training vocal data, on a phoneme-by-phoneme basis.

[0049] First, the processing of the speech training unit 301 in Fig. 3 in which the HMM acoustic model is adopted will be described. A model training unit 305 in the speech training unit 301 may input the training language feature sequence 313 output by the training text analysis unit 303 and the training acoustic feature sequence 314 output by the training acoustic feature extraction unit 304, to train an HMM acoustic model that maximizes the likelihood.

[0050] The spectral parameters of singing voices can be modeled using a continuous HMM. On the other hand, the logarithmic fundamental frequency (F0) is a variable-dimensional time-series signal that takes continuous values ​​in voiced sections and has no value in unvoiced sections, so it cannot be directly modeled using a normal continuous or discrete HMM. Therefore, we use MSD-HMM (Multi-Space probability Distribution HMM), an HMM based on a probability distribution in a multi-space corresponding to variable dimensions, and simultaneously model the mel-cepstrum as a multidimensional Gaussian distribution for the spectral parameters, and the logarithmic fundamental frequency (F0) as a Gaussian distribution in one-dimensional space for voiced sounds and zero-dimensional space for unvoiced sounds.

[0051] It is also known that the characteristics of the phonemes that make up a singing voice vary under the influence of various factors, even if the acoustic characteristics of the phonemes are the same. For example, the spectrum and logarithmic fundamental frequency (F0) of a phoneme, which is a basic phonetic unit, differ depending on the singing style, tempo, or the lyrics and pitch of the preceding and following songs. Such factors that affect acoustic features are called contexts.

[0052] In the statistical speech synthesis process of an embodiment, an HMM acoustic model that takes context into consideration (context-dependent model) may be adopted to accurately model the acoustic features of speech. Specifically, the training text analysis unit 303 may output a training language feature sequence 313 that takes into consideration not only the phonemes and pitch of each frame, but also the immediately preceding and following phonemes, the current position, the immediately preceding and following vibratos, accents, etc. Furthermore, context clustering based on a decision tree may be used to improve the efficiency of the combination of contexts.

[0053] For example, the model training unit 305 may generate, as the training result 315, a state duration decision tree for determining state duration from a training language feature sequence 313 corresponding to the context of a large number of phonemes related to state duration extracted by the training text analysis unit 303 from the training singing data 311.

[0054] In addition, the model training unit 305 may generate, as the training result 315, a Mel-Cepstral parameter decision tree for determining Mel-Cepstral parameters from the training acoustic feature sequence 314 corresponding to a large number of phonemes related to the Mel-Cepstral parameters extracted by the training acoustic feature extraction unit 304 from the training singing voice data 312.

[0055] Furthermore, the model learning unit 305 may generate, as the learning result 315, a logarithmic fundamental frequency decision tree for determining the logarithmic fundamental frequency (F0) from the training acoustic feature sequence 314 corresponding to a large number of phonemes related to the logarithmic fundamental frequency (F0) extracted by the training acoustic feature extraction unit 304 from the training singing voice data 312. Note that the voiced and unvoiced sections of the logarithmic fundamental frequency (F0) may be modeled as one-dimensional and zero-dimensional Gaussian distributions, respectively, by the MSD-HMM corresponding to the variable dimension, and a logarithmic fundamental frequency decision tree may be generated.

[0056] Instead of or together with the HMM-based acoustic model, an acoustic model based on a deep neural network (DNN) may be adopted. In this case, the model learning unit 305 may generate model parameters representing a nonlinear conversion function of each neuron in the DNN from the language feature to the acoustic feature as the learning result 315. The DNN makes it possible to express the relationship between the language feature sequence and the acoustic feature sequence using a complex nonlinear conversion function that is difficult to express using a decision tree.

[0057] Furthermore, the acoustic models of the present disclosure are not limited to these, and any voice synthesis method may be adopted as long as it is a technology that uses statistical voice synthesis processing, such as an acoustic model that combines HMM and DNN.

[0058] The learning results 315 (model parameters) may be stored in the ROM 202 of the control system of the electronic musical instrument 10 of FIG. 2 when the electronic musical instrument 10 of FIG. 1 is shipped from the factory, as shown in FIG. 3, and may be loaded from the ROM 202 of FIG. 2 to a singing voice control unit 306 (described later) in the waveform data output unit 211 when the electronic musical instrument 10 of FIG. 1 is powered on.

[0059] The learning results 315 may be downloaded to the singing voice control unit 306 in the waveform data output unit 211 from an external source such as the Internet via the network interface 219 by the performer operating the switch panel 140b of the electronic musical instrument 10, as shown in FIG. 3, for example.

[0060] <Speech synthesis based on acoustic models> FIG. 4 is a diagram illustrating an example of the waveform data output unit 302 according to an embodiment.

[0061] The waveform data output unit 302 includes a processing unit (which may be called a text processing unit, a pre-processing unit, etc.) 306, a singing voice control unit (which may be called an acoustic model unit) 307, a sound source 308, a singing voice synthesis unit (which may be called a vocalization model unit) 309, etc.

[0062] The waveform data output unit 302 inputs singing voice data 215 including lyrics and pitch information instructed by the CPU 201 via the key scanner 206 in Fig. 2 based on key depressions of the keyboard 140k in Fig. 1, and synthesizes and outputs singing voice waveform data 217 corresponding to the lyrics and pitch. In other words, the waveform data output unit 302 executes a statistical voice synthesis process that synthesizes singing voice waveform data 217 corresponding to the singing voice data 215 including lyrics text by predicting it using a statistical model called an acoustic model set in the singing voice control unit 306.

[0063] When song data is being played back, the waveform data output section 302 also outputs song waveform data 218 that corresponds to the corresponding song playback position.

[0064] The processing unit 307 inputs singing voice data 215 including information on the phonemes and pitches of the lyrics specified by the CPU 201 in Fig. 2 as a result of a performer's performance in sync with the automatic performance, for example, and analyzes the data. The singing voice data 215 may include, for example, data on the nth note (which may be called the nth note) (for example, pitch and note duration data), singing voice data of the nth note, etc.

[0065] For example, the processing unit 307 may determine the presence or absence of lyric progression based on note on / off data, pedal on / off data, etc. acquired from the operation of the keyboard 140k and the pedal 140p, in accordance with a lyric progression control method described later, and acquire the singing voice data 215 corresponding to the lyrics to be output. The processing unit 307 may then analyze the pitch data designated by the key depression and the acquired singing voice data 215 to generate a linguistic feature sequence 316 expressing phonemes, parts of speech, words, etc. corresponding to the acquired singing voice data 215, and output the linguistic feature sequence 316 to the singing voice control unit 306.

[0066] The singing data may be information including at least one of the following: lyrics (characters), syllable type (start syllable, middle syllable, end syllable, etc.), lyrics index, corresponding pitch (correct pitch), and corresponding pronunciation period (e.g., pronunciation start timing, pronunciation end timing, pronunciation duration) (correct pronunciation period).

[0067] For example, in the example of FIG. 4, the singing voice data 215 may include information on singing voice data of the nth lyric corresponding to the nth note (n=1, 2, 3, 4, ...) and the specified timing at which the nth note should be played (nth singing voice playback position).

[0068] The singing voice data 215 may include information (such as data in a specific audio file format, MIDI data, etc.) for playing an accompaniment (song data) corresponding to the lyrics. When the singing voice data is represented in SMF format, the singing voice data 215 may include a track chunk in which data related to the singing voice is stored, and a track chunk in which data related to the accompaniment is stored. The singing voice data 215 may be read from the ROM 202 to the RAM 203. The singing voice data 215 is stored in a memory (for example, the ROM 202, the RAM 203) before it is played.

[0069] The electronic musical instrument 10 may control the progress of the automatic accompaniment based on events indicated by the singing voice data 215 (for example, meta events (timing information) indicating the timing and pitch of lyrics, MIDI events indicating note-on or note-off, or meta events indicating beats).

[0070] The singing control unit 306 estimates an acoustic feature sequence 317 corresponding to the language feature sequence 316 input from the processing unit 307 and the acoustic model set as the learning result 315, and outputs formant information 318 corresponding to the estimated acoustic feature sequence 317 to the singing synthesis unit 309.

[0071] For example, when an HMM acoustic model is adopted, the singing control unit 306 connects HMMs by referring to a decision tree for each context obtained by the language feature sequence 316, and predicts the acoustic feature sequence 317 (formant information 318 and vocal cord sound source data 319) that maximizes the output probability from each of the connected HMMs.

[0072] When a DNN acoustic model is employed, the singing voice control unit 306 may output an acoustic feature sequence 317 on a frame-by-frame basis in response to a phoneme sequence of a linguistic feature sequence 316 input on a frame-by-frame basis.

[0073] In FIG. 4, a processing unit 307 obtains musical instrument sound data (pitch information) corresponding to the pitch of the pressed key from a memory (which may be the ROM 202 or the RAM 203 ), and outputs the data to a sound source 308 .

[0074] The sound source 308 generates a sound source signal (which may be called instrument sound waveform data) of instrument sound data (pitch information) corresponding to the sound to be generated (note-on) based on the note-on / off data input from the processing unit 307, and outputs it to the singing voice synthesis unit 309. The sound source 308 may execute control processes such as envelope control of the sound to be generated.

[0075] The singing voice synthesis unit 309 forms a digital filter that models the vocal tract based on a series of formant information 318 sequentially input from the singing voice control unit 306. The singing voice synthesis unit 309 also uses the sound source signal input from the sound source 309 as an excitation source signal, applies the digital filter, and generates and outputs singing voice waveform data 217 as a digital signal. In this case, the singing voice synthesis unit 309 may be called a synthesis filter unit.

[0076] The singing voice synthesis unit 309 may be capable of employing various voice synthesis methods, including the cepstrum voice synthesis method and the LSP voice synthesis method.

[0077] In the example of FIG. 4, the singing voice waveform data 217 that is output uses the instrument sound as the sound source signal, and therefore loses some fidelity compared to the singer's singing voice, but the singing voice retains both the atmosphere of the instrument sound and the vocal quality of the singer, allowing for effective singing voice waveform data 217 to be output.

[0078] The sound source 309 may operate to process the instrument sound waveform data and output the output of other channels as song waveform data 218. This makes it possible to produce accompaniment sounds using normal instrument sounds, or to produce instrument sounds for a melody line and simultaneously vocalize the melody.

[0079] 5 is a diagram showing another example of the waveform data output unit 302 according to an embodiment. Contents that overlap with those in FIG. 4 will not be described repeatedly.

[0080] 5 estimates an acoustic feature sequence 317 based on the acoustic model as described above. Then, the singing control unit 306 outputs formant information 318 corresponding to the estimated acoustic feature sequence 317 and vocal cord sound source data (pitch information) 319 corresponding to the estimated acoustic feature sequence 317 to the singing synthesis unit 309. The singing control unit 306 may estimate an estimate of the acoustic feature sequence 317 that maximizes the probability that the acoustic feature sequence 317 is generated.

[0081] The singing synthesis unit 309 may generate data (which may be called, for example, singing waveform data of the nth lyric corresponding to the nth note) for generating a signal by applying a digital filter that models the vocal tract based on the series of formant information 318 to, for example, a pulse train (in the case of voiced phonemes) that is periodically repeated at the fundamental frequency (F0) and power value contained in the vocal cord sound source data 319 input from the singing control unit 306, or white noise (in the case of unvoiced phonemes) having the power value contained in the vocal cord sound source data 319, or a signal obtained by mixing these, and outputting the data to the sound source 308.

[0082] The sound source 308 generates and outputs singing voice waveform data 217 of a digital signal from the singing voice waveform data of the nth lyric corresponding to the note to be pronounced (note-on) based on the note-on / off data input from the processing unit 307.

[0083] In the example of Figure 5, the output vocal waveform data 217 is a signal that is completely modeled by the vocal control unit 306, as it is a sound generated by the sound source 308 based on the vocal cord sound source data 319 as a sound source signal, and therefore it is possible to output vocal waveform data 217 that is very faithful to the singer's singing voice and has a natural singing voice.

[0084] In this way, the voice synthesis disclosed herein differs from existing vocoders (a technique in which spoken words are input through a microphone and replaced with instrument sounds for synthesis), in that it can output synthetic voice by operating the keyboard without the user (performer) having to sing (in other words, without the user having to input a voice signal to be pronounced in real time into electronic musical instrument 10).

[0085] As described above, by adopting the statistical voice synthesis processing technology as the voice synthesis method, it is possible to realize a memory capacity that is significantly smaller than that of the conventional voice unit synthesis method. For example, in an electronic musical instrument using the voice unit synthesis method, a memory with a storage capacity of several hundred megabytes is required for voice unit data, but in this embodiment, a memory with a storage capacity of only a few megabytes is sufficient to store the model parameters of the learning result 315. This makes it possible to realize a lower-priced electronic musical instrument, and makes it possible for a wider range of users to use a high-quality singing voice performance system.

[0086] Furthermore, in the conventional segment data method, since the segment data needs to be adjusted manually, it takes a huge amount of time (years) and effort to create data for singing performance, but in the creation of model parameters of the learning result 315 for the HMM acoustic model or DNN acoustic model according to this embodiment, almost no data adjustment is required, so the creation time and effort is reduced to a fraction of that. This also makes it possible to realize a lower-priced electronic musical instrument.

[0087] In addition, general users can use the learning function built into the server computer 300 and voice synthesis LSI 205 available as a cloud service to learn their own voice, the voice of a family member, or the voice of a celebrity, and use that as a model voice to play a singing voice on an electronic musical instrument. In this case, it is also possible to realize a singing voice performance that is much more natural and high quality than before, on a lower-cost electronic musical instrument.

[0088] (Lyric progression control method) A lyric progression control method according to an embodiment of the present disclosure will be described below. Each lyric progression control method may be used by the processing unit 307 of the electronic musical instrument 10 described above, or the like.

[0089] The subject of operation (electronic musical instrument 10) in each of the following flowcharts may be interpreted as either the CPU 201 or the waveform data output unit 211 (or the tone generator LSI 204 and voice synthesis LSI 205 therein), or a combination of these. For example, the CPU 201 may execute a control processing program loaded from the ROM 202 to the RAM 203 to perform each operation.

[0090] Note that an initialization process may be performed at the start of the flow shown below. The initialization process may include interrupt processing, lyric progression, deriving TickTime, which is the reference time for automatic accompaniment, tempo setting, song selection, song loading, instrument sound selection, and other button-related processing.

[0091] The CPU 201 can detect operations of the switch panel 140b, the keyboard 140k, the pedals 140p, etc., based on an interrupt from the key scanner 206 at an appropriate timing, and execute the corresponding processing.

[0092] In the following, an example of controlling the progression of lyrics is shown, but the target of the progression control is not limited to this. For example, based on the present disclosure, the progression of any character string, sentence (e.g., news script), etc. may be controlled instead of lyrics. In other words, the lyrics of the present disclosure may be read as characters, character strings, etc.

[0093] 6 is a diagram showing an example of a flowchart of a lyrics progression control method according to an embodiment of the present invention. Note that, although the generation of synthetic voice in this example is based on FIG. 4, it may be based on FIG.

[0094] First, the electronic musical instrument 10 assigns 0 to a lyric index (also represented as "n") indicating the current position in the lyrics, and to a note number (also represented as "SKO") indicating the highest note of the key being pressed (step S101). Note that if the lyrics are to be started from the middle (for example, from the previously stored position), a value other than 0 may be assigned to n.

[0095] The lyrics index may be a variable indicating which syllable (or character) from the beginning of the lyrics corresponds to when the entire lyrics are considered as a character string. For example, the lyrics index n may indicate the singing voice data at the nth playback position of the singing voice data 215 shown in Fig. 4, Fig. 5, etc. In the present disclosure, the lyrics corresponding to one lyrics position (lyric index) may correspond to one or more characters constituting one syllable. The syllables included in the singing voice data may include various syllables such as only vowels, only consonants, or a consonant + vowel.

[0096] Step S101 may be performed when a performance starts (for example, when playback of song data starts), when vocal data is read, or the like.

[0097] The electronic musical instrument 10 may play back song data (musical accompaniment) corresponding to the lyrics in response to, for example, a user's operation (step S102). The user can perform key depression operations in time with the musical accompaniment to advance the progression of the lyrics and to perform the performance.

[0098] The electronic musical instrument 10 determines whether the playback of the song data that was started in step S102 has ended (step S103). If the playback has ended (step S103-Yes), the electronic musical instrument 10 may end the process of this flowchart and return to a standby state.

[0099] In this case, the electronic musical instrument 10 may read the vocal data designated by the user's operation in step S102 as a progression control target, and may determine in step S103 whether or not all of the vocal data has progressed.

[0100] If the playback of the song data has not ended (step S103-No), the electronic musical instrument 10 determines whether or not a new key has been pressed (a note-on event has occurred) (step S111). If a new key has been pressed (step S111-Yes), the electronic musical instrument 10 performs a lyric progression determination process (a process for determining whether or not to progress the lyrics) (step S112). An example of this process will be described later. Then, the electronic musical instrument 10 determines whether or not there is lyric progression (whether or not it has been determined that the lyrics should be progressed) as a result of the lyric progression determination process (step S113).

[0101] If it is determined that the lyrics should be progressed (step S113-Yes), the electronic musical instrument 10 increments the lyrics index n (step S114). This increment is basically an increment of 1 (substituting n+1 for n), but a value greater than 1 may be added depending on the result of the lyrics progression determination process in step S112, etc.

[0102] After incrementing the lyrics index, the electronic musical instrument 10 acquires acoustic feature data (formant information) of the n-th singing voice data from the singing voice control unit 306 (step S115).

[0103] On the other hand, if it is not determined that the lyrics should be advanced (step S113-No), the electronic musical instrument 10 does not change the lyrics index (maintains the value of the lyrics index). In this case, step S115 is unnecessary, and therefore the process can be simplified.

[0104] After step S115 or S113-No, the electronic musical instrument 10 instructs the sound source 309 to generate an instrument sound (generate instrument sound waveform data) at a pitch corresponding to the key pressed (step S116).Then, the electronic musical instrument 10 instructs the singing voice synthesis unit 309 to add formant information of the n-th singing voice data to the instrument sound waveform data output from the sound source 308 (step S117).

[0105] For a note that is already being pronounced, the electronic musical instrument 10 may continue to output the same note (or a vowel of the same note) without progressing the lyrics, or may output a note based on the lyrics that have progressed. When pronouncing a note that corresponds to the same lyric index value as a note that is already being pronounced, the electronic musical instrument 10 may output the vowel of the lyrics. For example, when the lyric "Sle" is already being pronounced and the same lyric is to be newly pronounced, the electronic musical instrument 10 may newly pronounce the sound "e."

[0106] If no new key has been pressed (step S111-No), the electronic musical instrument 10 determines whether or not a new key has been released (a note-off event has occurred) (step S121). If a new key has been released (step S121-Yes), the electronic musical instrument 10 performs a muting process on the corresponding vocal data (step S122). The electronic musical instrument 10 also updates the currently sounded note management table (step S123).

[0107] Here, the note management table may manage the note number of the key that is being sounded (that is being pressed) and the time when the key pressing started. In step S123, the electronic musical instrument 10 may delete information related to the muted note from the note management table.

[0108] The electronic musical instrument 10 also assigns the note number of the highest note being sounded to SKO (step S124).

[0109] Next, the electronic musical instrument 10 judges whether or not all the keys are off (step S125). If all the keys are off (step S125-Yes), the electronic musical instrument 10 performs a synchronization process between the lyrics and the song (accompaniment) (step S126). The synchronization process will be described later.

[0110] After steps S117, S125-No and S126, the process returns to step S103.

[0111] The electronic musical instrument 10 of the present disclosure may be capable of simultaneously generating multiple sounds by using synthesized voices with different voice tones for each sound. For example, when the user presses four notes, the electronic musical instrument 10 may synthesize and output voices corresponding to soprano, alto, tenor, and bass voice tones in order from the highest note.

[0112] <Lyric progression determination process> The lyrics progression determination process in step S102 will be described in detail below.

[0113] 7 is a diagram showing an example of a flowchart of a lyric progression determination process based on chord voicing. In other words, this process corresponds to a process of determining lyric progression based on which pitch (which may be interpreted as "what pitch" or "which part") of a chord has changed due to key pressing.

[0114] The electronic musical instrument 10 updates the note management table during sound production (step S112-1). Here, information regarding the note of the newly pressed key is added to the note management table. The key press time corresponding to the new key press in step S111 may be called the current key press time, the latest key press time, etc.

[0115] The electronic musical instrument 10 judges whether the newly pressed note is higher than the SKO (step S112-2). If the newly pressed note is higher than the SKO (step S112-2-Yes), the electronic musical instrument 10 assigns the note number of the newly pressed note to the SKO and updates the SKO (step S112-3). The electronic musical instrument 10 then judges that lyrics are progressing (step S112-11). This is because it is considered that the highest note (soprano part) usually corresponds to the melody.

[0116] If the newly pressed note is not higher than SKO (step S112-2-No), the electronic musical instrument 10 judges whether or not the difference between the latest key press time and the previous key press time is within the chord discrimination time (step S112-4). Step S112-4 can be rephrased as, for example, a step of judging whether or not the difference between the key press time of the newly pressed note and the key press time of the previous key press (or the key press time i times ago (i is an integer)) is within the chord discrimination time. The previous key press time preferably corresponds to a key that has been continuously pressed even in the latest key press time.

[0117] Here, the chord discrimination time is a time (period) for determining that a plurality of sounds produced within that time are simultaneous chords, and determining that a plurality of sounds produced outside that time are independent sounds (e.g., sounds of a melody line) or arpeggios. The chord discrimination time may be expressed in milliseconds or microseconds, for example.

[0118] The chord discrimination time may be obtained from a user input or may be derived based on the tempo of the song. The chord discrimination time may be referred to as a predetermined set time, a set time, or the like.

[0119] If the difference between the most recent key press time and the previous key press time is within the chord discrimination time (step S112-4-Yes), the electronic musical instrument 10 determines that the pressed keys are a simultaneous chord (a chord has been specified), and determines to maintain the lyrics (do not progress the lyrics) (step S112-12).

[0120] If there is no previous key pressing time within the chord discrimination time (step S112-4-No), the electronic musical instrument 10 judges whether the number of currently pressed keys is equal to or exceeds a predetermined number and whether the newly pressed key corresponds to a specific note among all the pressed keys (step S112-5). If step S112-4-No, the electronic musical instrument 10 may determine that the chord designation has been released or that no chord has been designated.

[0121] The number of currently pressed keys may be determined from the number of notes present in the note management table. The predetermined number may be, for example, four notes (assuming four voices: soprano, alto, tenor, and bass) or eight notes. The specific note may be the lowest note (corresponding to the bass part) among all the pressed notes, or the i-th (i is an integer) highest or lowest note. These predetermined numbers, specific notes, etc. may be set by a user operation, or may be specified in advance.

[0122] If the result of step S112-5 is Yes, the electronic musical instrument 10 determines that the lyrics are to be maintained (step S112-12).If the result of step S112-5 is No, the electronic musical instrument 10 determines that the lyrics are to be progressed (step S112-11).

[0123] According to the process of step S112-4, when multiple keys are pressed with the intention of creating a chord, it is undesirable for the lyrics to advance by the number of keys, so that the lyrics can be advanced by only one step.

[0124] According to the lyric progression determination process shown in FIG. 7, for example, it is possible to progress the lyrics if there are multiple sounds (melody) with a large time difference between their pronunciations, rather than multiple sounds (so-called simultaneous chords (harmony)) with a small time difference between their pronunciations.

[0125] For example, when the highest key changes with the key pressing of a chord (step S112-2-Yes), the lyrics can be advanced according to the highest key pressed. Also, if the top note of the chord that will be responsible for the melody is maintained, the lyrics can be controlled not to advance. This is expected to be effective when reproducing a polyphonic chorus.

[0126] Also, when the lowest key is changed (step S112-5-Yes), the lyrics are not advanced according to the lowest key. With this configuration, even if the pitch of only the lowest note of the chord changes, which may correspond to the bass part of a four-part chorus, the lyrics are not advanced as long as the chord of the higher parts is maintained.

[0127] In addition, when the pressed key other than the lowest note is changed (step S112-5-No), the lyrics can be controlled to advance according to the pressed key. With this configuration, the lyrics can be appropriately advanced when the part that can take charge of the melody in the four-part chorus is played independently instead of as a chord.

[0128] Incidentally, the question "whether or not the newly pressed note is higher than the SKO" in step S112-2 may be interpreted as "whether or not the newly pressed note corresponds to the melody part".

[0129] Note that the question in step S112-5 “whether the current number of pressed keys is equal to or greater than a predetermined number and whether the newly pressed note corresponds to a specific note among all the notes currently being pressed” may be interpreted as “whether the newly pressed note does not correspond to a melody part (or corresponds to a harmony part)”.

[0130] Information on which notes correspond to the melody (or harmony) parts for each range of lyrics may be given in advance. For example, the information may indicate that the melody part of the lyrics corresponding to lyrics indexes = 0 to 10 is the highest note among the notes that can be pressed, and the melody part of the lyrics corresponding to lyrics indexes = 11 to 20 is the lowest note among the notes that can be pressed, etc.

[0131] The information may include at least one of the following: information indicating which highest note corresponds to the melody (or harmony) part, information indicating which range (e.g., hiA to hiG#) corresponds to the melody (or harmony) part, etc.

[0132] Based on the above information, the electronic musical instrument 10 may recognize, for example, the highest note (soprano part) in the A melody as the melody, and the third highest note (tenor part) in the chorus as the melody, and use this information for lyric control.

[0133] Fig. 8 is a diagram showing an example of lyric progression controlled using the lyric progression determination process. In this example, a case where the user presses the keys according to the musical score shown in the figure will be described. For example, the treble clef may be pressed by the user's right hand, and the bass clef may be pressed by the user's left hand. Also, "Sle", "e", "ping", "heav", "en" and "ly" correspond to lyric indexes 1-6, respectively.

[0134] It is assumed that the chord discrimination time is shorter than an eighth note (for example, the length of a 32nd note), and that the predetermined number in step S102-17 above is 4, and that the specific note is the lowest note.

[0135] First, at timing t1, four keys are pressed. The electronic musical instrument 10 performs the lyric progression determination process of Fig. 7, and because step S112-2 is Yes, determines in step S112-11 that the lyrics should be progressed. Then, in step S114, the electronic musical instrument 10 increments the lyrics index by 1, and generates and outputs the lyrics "Sle" using four synthetic voices.

[0136] Next, at timing t2, while continuing to press the right hand key, the user moves the left hand to the "D" key. This D is the lowest note that the electronic musical instrument 10 should sound at t2. The electronic musical instrument 10 performs the lyric progression determination process of FIG. 7, and as step S112-5 is Yes, determines in step S112-12 not to progress the lyrics. The electronic musical instrument 10 then generates and outputs the D using the vowel (e) of "Sle" that is already being sounded, while maintaining the lyric index. The electronic musical instrument 10 continues to sound the other three voices.

[0137] Similarly, at t3, the electronic musical instrument 10 outputs the lyric "e" using sounds corresponding to the four keys, and at t4, the lyrics are maintained and only the lowest note is updated. Also, at t5, the electronic musical instrument 10 outputs the lyric "ping" using sounds corresponding to the four keys, and at t6, the lyrics are maintained and only the lowest note is updated.

[0138] In the example section t1-t6 in Figure 8, the lyrics of the upper triad were assigned one syllable per note, and the lyrics progressed with each keystroke. On the other hand, the bass part was assigned one syllable (melisma) per two notes, and because it was judged to be the lowest note of the four voices, there were some parts where the lyrics did not progress with each keystroke.

[0139] <Synchronization process> The synchronization process may be a process for matching the position of lyrics with the playback position of the current song data (accompaniment). This process allows the position of lyrics to be moved appropriately when the lyrics position is exceeded due to too many keys being pressed, or when the lyrics position does not advance as expected due to not enough keys being pressed.

[0140] FIG. 9 is a diagram illustrating an example of a flowchart of the synchronization process.

[0141] The electronic musical instrument 10 acquires the playback position of the song data (step S126-1).The electronic musical instrument 10 then determines whether the playback position matches the (n+1)th vocal playback position (step S126-2).

[0142] The (n+1)th vocal reproduction position may indicate a desired timing at which the (n+1)th note is reproduced, which is derived in consideration of the total note duration of vocal data up to the nth note.

[0143] If the song data playback position and the (n+1)th vocal playback position match (step S126-2-Yes), the synchronization process may be terminated. If not (step S126-2-No), the electronic musical instrument 10 may obtain the Xth vocal playback position that is closest to the song data playback position (step S126-3), assign X-1 to n (step S126-4), and terminate the synchronization process.

[0144] If no accompaniment is being played, the synchronization process may be omitted. Also, if the appropriate timing for sounding lyrics is derived based on the singing voice data, even if no accompaniment is being played, the electronic musical instrument 10 may perform a process for adjusting the position of the lyrics to the position when they would have been sounded appropriately, depending on the time elapsed from the start of performance to the present, the number of key presses, etc.

[0145] According to the embodiment described above, lyrics can be smoothly progressed even when multiple keys are pressed simultaneously.

[0146] (Modification) 4, 5, etc. may be switched on / off based on the user's operation of the switch panel 140b. When it is off, the waveform data output unit 211 may be controlled to generate and output a sound source signal of musical instrument sound data of a pitch corresponding to the pressed key.

[0147] Some steps may be omitted from the flowcharts in Fig. 6 and the like. When a determination process is omitted, the determination may be interpreted as proceeding along a route of always "Yes" or always "No" in the flowchart.

[0148] It is sufficient for electronic musical instrument 10 to be able to control at least the position of lyrics, and it is not necessary for electronic musical instrument 10 to generate or output sounds corresponding to lyrics. For example, electronic musical instrument 10 may transmit sound waveform data generated based on key depressions to an external device (such as server computer 300), and the external device may generate / output synthetic voice based on the sound waveform data.

[0149] The electronic musical instrument 10 may perform control to display lyrics on the display 150d. For example, lyrics near the current lyrics position (lyrics index) may be displayed, or lyrics corresponding to sounds being pronounced or sounds that have been pronounced may be displayed in color so that the current lyrics position can be identified.

[0150] The electronic musical instrument 10 may transmit at least one of singing voice data, information on the current lyric position, etc. to the external device. The external device may control display of lyrics on its own display based on the received singing voice data, information on the current lyric position, etc.

[0151] In the above example, the electronic musical instrument 10 is a keyboard-like instrument, but is not limited to this. The electronic musical instrument 10 may be any instrument that has a configuration that allows the user to specify the timing of sound generation, such as an electric violin, an electric guitar, a drum, or a trumpet.

[0152] For this reason, a "key" in this disclosure may be interpreted as a string, a valve, another performance operator for specifying pitch, any performance operator, etc. A "key press" in this disclosure may be interpreted as striking a key, picking, playing, operating an operator, etc. A "key release" in this disclosure may be interpreted as stopping a string, stopping playing, stopping (not operating) an operator, etc.

[0153] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. The means for realizing each functional block is not particularly limited. That is, each functional block may be realized by one physically combined device, or may be realized by two or more physically separated devices connected by wire or wirelessly.

[0154] In addition, terms explained in this disclosure and / or terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings.

[0155] The information, parameters, etc. described in this disclosure may be expressed using absolute values, may be expressed using relative values ​​from a predetermined value, or may be expressed using other corresponding information. Furthermore, the names used for parameters, etc. in this disclosure are not limiting in any respect.

[0156] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0157] Information, signals, etc. may be input and output via multiple network nodes. Input and output information, signals, etc. may be stored in a specific location (e.g., memory) or may be managed using a table. Input and output information, signals, etc. may be overwritten, updated, or added. Output information, signals, etc. may be deleted. Input information, signals, etc. may be transmitted to another device.

[0158] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0159] Additionally, software, instructions, information, etc. may be transmitted or received over a transmission medium. For example, if the software is transmitted from a website, server, or other remote source using wired and / or wireless technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave, etc.), then these wired and / or wireless technologies are included within the definition of transmission media.

[0160] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched according to implementation. In addition, the processing procedures, sequences, flow charts, etc. of each aspect / embodiment described in this disclosure may be reordered unless inconsistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0161] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0162] Any reference to an element using a designation such as "first," "second," etc., used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must precede the second element in some way.

[0163] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Further, when used in this disclosure, the term "or" is not intended to be an exclusive or.

[0164] In this disclosure, where articles have been added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0165] Regarding the above embodiment, the following supplementary notes are disclosed. (Appendix 1) A plurality of performance operators (e.g., keys) each associated with a different pitch data (e.g., note number); A processor (e.g., a CPU 201), the processor comprising: Determine whether or not a chord has been designated in response to user operations on the plurality of performance operators (for example, determine whether or not a plurality of pitch data corresponding to each of the pitches designated in response to the user operations (in other words, data on the note numbers included in the note-on data) have been acquired within a chord discrimination time (which may be within a few milliseconds or may be almost simultaneous)); When it is determined that the chord has been designated, instructing pronunciation of a singing voice corresponding to the first lyrics (e.g., "Sle" in FIG. 8) at each pitch designated in response to a user operation, When it is determined that the chord is not specified, instruct the pronunciation of a singing voice corresponding to the first lyrics at one pitch specified in response to a user operation, and instruct the pronunciation of a singing voice corresponding to a second lyric following the first lyrics (e.g., "e" in FIG. 8) at the remaining pitch specified in response to a user operation. Electronic musical instrument.

[0166] (Appendix 2) The processor, determining whether or not pitch data acquired in response to a user operation after it is determined that the chord has been designated and before it is determined that the designation of the chord has been released is the lowest pitch among the designated pitches; When it is determined that the note is the lowest note, the pronunciation of the singing voice corresponding to the second lyrics is not instructed, but the pronunciation of the singing voice corresponding to the first lyrics is instructed at the pitch of the determined lowest note (in other words, when it is the lowest note, the lyrics do not progress); When the note is not determined to be the lowest note, the pronunciation of the singing voice corresponding to the first lyrics is not instructed, but the pronunciation of the singing voice corresponding to the second lyrics is instructed (in other words, when the note is not the lowest note, the lyrics are progressed). 2. An electronic musical instrument as defined in appended claim 1.

[0167] (Appendix 3) The processor, Instructs playback of accompaniment data (song data), Determine whether or not all pitch designations have been released in response to a user operation (in other words, whether or not all keys are off); when it is determined that all pitch designations have been released, a first playback position in singing voice text data including the first lyrics data and the second lyrics data, which is a first playback position of lyrics to be sung in response to a next user operation, is changed to a second playback position corresponding to a playback position in the accompaniment data (in other words, a synchronization process is performed); 3. An electronic musical instrument as defined in claim 1 or 2.

[0168] (Appendix 4) The processor, acquiring a plurality of musical instrument sound data (e.g., data on musical instrument sounds such as brass sounds) corresponding to the acquired plurality of pitch data; when it is determined that the chord has been designated, by providing formant information corresponding to the first lyrics to each of the plurality of musical instrument sound data corresponding to each pitch designated in response to a user operation (in other words, by performing voice synthesis), the pronunciation of a singing voice corresponding to the first lyrics at each pitch designated in response to a user operation is instructed without the user having to sing; when it is determined that the chord is not specified, formant information corresponding to the first lyrics and formant information corresponding to the second lyrics are added to each of the plurality of musical instrument sound data corresponding to each pitch specified in response to a user operation, thereby instructing the pronunciation of the singing voice corresponding to the first lyrics and the second lyrics without the user having to sing. 4. An electronic musical instrument according to any one of claims 1 to 3.

[0169] (Appendix 5) The processor, When it is determined that the chord is specified, inputting the data of the first lyrics into a trained model, thereby assigning formant information output by the trained model to each of the plurality of musical instrument sound data; When it is determined that the chord is not specified, the method inputs the first lyrics data into a trained model to provide formant information output by the trained model, and inputs the second lyrics data into the trained model to provide formant information output by the trained model to each of the plurality of musical instrument sound data. 5. An electronic musical instrument as described in appended claim 4.

[0170] (Appendix 6) The trained model is generated by machine learning using singing voice data of a certain singer as training data, and outputs formant information indicating acoustic features of the singing voice of the certain singer in response to input of lyrics data. 6. An electronic musical instrument as described in appended claim 5.

[0171] (Appendix 7) Electronic musical instrument computers, determining whether or not a chord has been designated in response to user operations on a plurality of performance operators; when it is determined that the chord has been designated, instructing pronunciation of a singing voice corresponding to the first lyrics at each pitch designated in response to a user operation; when it is determined that the chord is not specified, instructing pronunciation of a singing voice corresponding to the first lyrics at one pitch specified in response to a user operation, and instructing pronunciation of a singing voice corresponding to second lyrics subsequent to the first lyrics at the remaining pitch specified in response to a user operation; method.

[0172] (Appendix 8) Electronic musical instrument computers, determining whether or not a chord has been designated in response to user operations on a plurality of performance operators; when it is determined that the chord has been designated, instructing pronunciation of a singing voice corresponding to the first lyrics at each pitch designated in response to a user operation; when it is determined that the chord is not specified, instructing pronunciation of a singing voice corresponding to the first lyrics at one pitch specified in response to a user operation, and instructing pronunciation of a singing voice corresponding to second lyrics subsequent to the first lyrics at the remaining pitch specified in response to a user operation; program.

[0173] Although the invention according to the present disclosure has been described in detail above, it is clear to those skilled in the art that the invention according to the present disclosure is not limited to the embodiments described in the present disclosure. The invention according to the present disclosure can be implemented as modified and altered forms without departing from the spirit and scope of the invention defined based on the description of the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the invention according to the present disclosure.

Claims

1. When the time difference between the timing of the most recent pitch specification and the timing of the most recent pitch specification is within the chord discrimination time, the lyrics index is not advanced in accordance with the most recent pitch specification, thereby determining that the lyrics have been maintained, and the chords of the same lyrics are played in both cases. When the lowest note being sounded does not change due to the latest pitch designated during the sounding of the chord and outside the chord discrimination time, it is determined that lyrics are progressing by advancing the lyrics index. An electronic musical instrument that can be controlled in this way.

2. When the lowest note being pronounced changes due to the latest pitch specified during the pronunciation of the chord and outside the chord discrimination time, it is determined that the lyrics are maintained by not advancing the lyrics index.

2. The electronic musical instrument according to claim 1.

3. When the highest note being pronounced changes due to the latest pitch, the lyrics index is advanced to determine that lyrics are progressing.

2. The electronic musical instrument according to claim 1.

4. If the latest pitch is higher than the highest pitch being produced, it is determined that lyrics are progressing by advancing the lyrics index.

4. The electronic musical instrument according to claim 1, wherein the electronic musical instrument is controlled as follows.

5. Electronic musical instrument computers, If the time difference between the timing of the previous pitch specification and the timing of the latest pitch specification is within the chord discrimination time, the lyrics index is not advanced in accordance with the latest pitch specification, thereby determining that the lyrics are to be maintained, and the chords of the same lyrics are sounded in both cases; When the lowest note being sounded does not change due to the latest pitch designated during the sounding of the chord and outside the chord discrimination time, it is determined that lyrics are progressing by advancing the lyrics index. method.

6. Electronic musical instrument computers, If the time difference between the timing of the previous pitch specification and the timing of the latest pitch specification is within the chord discrimination time, the lyrics index is not advanced in accordance with the latest pitch specification, thereby determining that the lyrics are to be maintained, and the chords of the same lyrics are sounded in both cases; When the lowest note being sounded does not change due to the latest pitch designated during the sounding of the chord and outside the chord discrimination time, it is determined that lyrics are progressing by advancing the lyrics index. program.

Citation Information

Patent Citations

  • Electronic musical instrument

    JP1992349497A

  • Electronic musical instrument

    JP2003099056A

  • Device and program for synthesizing singing voice

    JP2008170592A

  • Controller, synthetic singing sound creation device and program

    JP2016206496A

  • Pronunciation control device, pronunciation control method, and program

    JP2017194594A