Electronic device, electronic musical instrument, method and program

The electronic musical instrument synchronizes vocal and accompaniment data through operator detection, addressing the issue of inappropriate lyric advancement during simultaneous key presses, thereby improving performance control and expression.

JP7740315B2Active Publication Date: 2025-09-17CASIO COMPUTER CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023187620
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-09-17
Estimated Expiration
2039-12-23

AI Technical Summary

Technical Problem

Existing electronic musical instruments face issues with lyrics advancing too far when multiple keys are pressed simultaneously, leading to inappropriate progression control during performances.

Method used

An electronic musical instrument that utilizes a control unit to synchronize vocal data playback with accompaniment data by detecting operations on performance operators, allowing for appropriate control of lyric progression, especially during melismatic singing.

Benefits of technology

Enables precise control over lyric progression, ensuring synchronized and appropriate advancement of lyrics even when multiple keys are pressed simultaneously, enhancing musical expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740315000001
    Figure 0007740315000001
  • Figure 0007740315000002
    Figure 0007740315000002
  • Figure 0007740315000003
    Figure 0007740315000003
Patent Text Reader

Abstract

To provide an electronic musical instrument, a method and a program for appropriately controlling lyric progression for performance.SOLUTION: A control system 200 of an electronic musical instrument is equipped with a plurality of first performance operators each of which is associated with different pitch data, and a pedal. When a first user operation to a first performance operator is detected while no operation to the pedal is detected, and a second user operation is detected after the first user operation, as well as an instruction is issued to pronounce with a singing voice according to a first lyric corresponding to a first note in response to the first user operation, an instruction is issued to pronounce with a singing voice according to a second lyric corresponding to a second note following the first note in response to the second user operation. When the first user operation to the first performance operator is detected while the operation to the pedal is detected, and the second user operation is detected after the first user operation, as well as an instruction is issued to pronounce with a singing voice according to the first lyric in response to the first user operation, no instruction is issued to the pronounce with a singing voice according to the second lyric in response to the second user operation.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an electronic device, an electronic musical instrument, a method, and a program. [Background technology]

[0002] In recent years, the use of synthetic speech has been expanding. In this context, it would be desirable to have an electronic musical instrument that not only performs automatically, but also can progress lyrics in response to key presses by the user (player) and output synthetic speech corresponding to the lyrics, as this would enable more flexible expression of synthetic speech.

[0003] For example, Patent Document 1 discloses a technique for progressing lyrics in synchronization with a performance based on a user's operation using a keyboard or the like. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 4735544 Summary of the Invention [Problem to be solved by the invention]

[0005] However, when multiple notes can be simultaneously produced using a keyboard or the like, if the lyrics simply progress each time a key is pressed, the lyrics will advance too far when multiple keys are pressed simultaneously.

[0006] Therefore, one of the objects of the present disclosure is to provide an electronic musical instrument, method, and program that can appropriately control the progression of lyrics during performance. [Means for solving the problem]

[0007] An electronic device according to one aspect of the present disclosure includes: During playback of accompaniment data, after detecting an operation on a second performance operator that stops the progression of vocal data that is being progressed in response to detection of an operation on a first performance operator, Operation of the first performance control and the operation of the second performance operator is not detected. detection In the case of The vocal playback position of the vocal data At the playback position of the accompaniment dataThe control unit executes synchronization processing to match the signals. [Effects of the Invention]

[0008] According to one aspect of the present disclosure, the progression of lyrics played can be appropriately controlled. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram showing an example of the appearance of an electronic musical instrument 10 according to an embodiment. [Figure 2] FIG. 2 is a diagram showing an example of the hardware configuration of a control system 200 of the electronic musical instrument 10 according to an embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of the voice learning unit 301 according to an embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the waveform data output unit 302 according to an embodiment. [Figure 5] FIG. 5 is a diagram illustrating another example of the waveform data output unit 302 according to an embodiment. [Figure 6] FIG. 6 is a diagram showing an example of a flowchart of a lyrics progression control method according to an embodiment. [Figure 7] FIG. 7 is a diagram showing an example of a flowchart of the sound generation process of the n-th singing voice data. [Figure 8] FIG. 8 is a diagram showing an example of lyric progression controlled using the lyric progression determination process. [Figure 9] FIG. 9 is a diagram illustrating an example of a flowchart of the synchronization process. DETAILED DESCRIPTION OF THE INVENTION

[0010] The use of two or more notes in a part originally composed of one syllable per note (syllable style) is also called melismatic singing. Melisma singing can also be interpreted as falsetto, fist-squeezing, etc.

[0011] The inventors came up with the lyric progression control method of the present disclosure by noting that a characteristic of melisma is that the pitch can be freely changed while maintaining the previous vowel when performing melismatic singing on an electronic musical instrument equipped with a singing voice synthesis sound source.

[0012] According to one aspect of the present disclosure, it is possible to control so that lyrics do not progress during a melisma. Furthermore, even when multiple keys are pressed simultaneously, it is possible to appropriately control whether lyrics progress.

[0013] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the following description, identical parts are designated by the same reference numerals. Since identical parts have the same names, functions, etc., detailed description thereof will not be repeated.

[0014] In the present disclosure, terms such as "lyric progression," "lyric position progression," and "singing position progression" may be interchangeable. In the present disclosure, terms such as "do not progress lyrics," "do not control lyric progression," "hold lyrics," and "suspend lyrics" may be interchangeable.

[0015] (electronic musical instrument) 1 is a diagram showing an example of the appearance of an electronic musical instrument 10 according to an embodiment. The electronic musical instrument 10 may include a switch (button) panel 140b, a keyboard 140k, pedals 140p, a display 150d, and speakers 150s.

[0016] The electronic musical instrument 10 is a device that receives input from a user via controls such as a keyboard and switches, and controls the performance, lyric progression, etc. The electronic musical instrument 10 may be a device that has the function of generating sounds according to performance information such as MIDI (Musical Instrument Digital Interface) data. The device may be an electronic musical instrument (such as an electronic piano or synthesizer), or may be an analog musical instrument equipped with sensors and configured to have the functions of the above-mentioned controls.

[0017] The switch panel 140b may include switches for operating the volume setting, sound source, tone setting, song (accompaniment) selection, song playback start / stop, song playback settings (tempo, etc.), etc.

[0018] The keyboard 140k may have multiple keys as performance controls. The pedal 140p may be a sustain pedal that sustains the sound of the pressed key while the pedal is depressed, or a pedal for operating an effector that processes the tone, volume, etc.

[0019] In the present disclosure, the terms sustain pedal, pedal, foot switch, controller (operator), switch, button, touch panel, etc. may be interchangeable. In the present disclosure, depressing a pedal may be interchangeable with operating a controller.

[0020] The keys may be called performance operators, pitch operators, timbre operators, direct operators, first operators, etc. The pedals may be called non-performance operators, non-pitch operators, non-timbre operators, indirect operators, second operators, etc.

[0021] The display 150d may display lyrics, musical scores, various setting information, etc. The speaker 150s may be used to emit sounds generated by playing.

[0022] The electronic musical instrument 10 may be capable of generating and converting at least one of MIDI messages (events) and Open Sound Control (OSC) messages.

[0023] The electronic musical instrument 10 may be referred to as a control device 10, a lyric progression control device 10, or the like.

[0024] The electronic musical instrument 10 may communicate with a network (such as the Internet) via at least one of wired and wireless communication (e.g., Long Term Evolution (LTE), 5th generation mobile communication system New Radio (5G NR), Wi-Fi (registered trademark), etc.).

[0025] The electronic musical instrument 10 may store in advance vocal data (which may also be called lyric text data, lyric information, etc.) relating to lyrics for which progression control is to be performed, or may transmit and / or receive the vocal data via a network. The vocal data may be text written in a musical notation description language (e.g., MusicXML), or may be written in a MIDI data storage format (e.g., Standard MIDI File (SMF) format), or may be text provided in an ordinary text file.

[0026] In addition, the electronic musical instrument 10 may acquire the content of what the user sings in real time via a microphone or the like provided in the electronic musical instrument 10, and apply voice recognition processing to this to obtain text data as singing voice data.

[0027] FIG. 2 is a diagram showing an example of the hardware configuration of a control system 200 of the electronic musical instrument 10 according to an embodiment.

[0028] A central processing unit (CPU) 201, a ROM (read only memory) 202, a RAM (random access memory) 203, a waveform data output unit 211, a key scanner 206 to which the switch (button) panel 140b, keyboard 140k, and pedal 140p of Figure 1 are connected, and an LCD controller 208 to which an LCD (Liquid Crystal Display) as an example of the display 150d of Figure 1 is connected, are each connected to a system bus 209.

[0029] A timer 210 for controlling the sequence of automatic performance may be connected to the CPU 201. The CPU 201 may also be called a processor, and may include an interface with peripheral circuits, a control circuit, an arithmetic circuit, a register, and the like.

[0030] The functions of each device may be realized by loading specified software (programs) onto hardware such as processor 1001, memory 1002, etc., so that processor 1001 performs calculations and controls communication via communication device 1004, reading and / or writing of data in memory 1002 and storage 1003, etc.

[0031] The CPU 201 executes a control program stored in the ROM 202 while using the RAM 203 as a work memory, thereby carrying out the control operations of the electronic musical instrument 10 shown in Fig. 1. In addition to the control program and various fixed data, the ROM 202 may also store vocal data, accompaniment data, and song data including these.

[0032] The CPU 201 is equipped with a timer 210 used in this embodiment, which counts the progress of an automatic performance in the electronic musical instrument 10, for example.

[0033] The waveform data output unit 211 may include a sound source LSI (large scale integrated circuit) 204, a voice synthesis LSI 205, etc. The sound source LSI 204 and the voice synthesis LSI 205 may be integrated into one LSI.

[0034] The singing voice waveform data 217 and song waveform data 218 output from the waveform data output unit 211 are converted into an analog singing voice output signal and an analog musical sound output signal by D / A converters 212 and 213, respectively. The analog musical sound output signal and the analog singing voice output signal may be mixed by a mixer 214, and the mixed signal may be amplified by an amplifier 215 and then output from the speaker 150s or an output terminal.

[0035] A key scanner (scanner) 206 constantly scans the key-on / key-off state of the keyboard 140k in FIG. 1, the switch operation state of the switch panel 140b, the pedal operation state of the pedal 140p, and the like, and notifies the CPU 201 of the state change by issuing an interrupt.

[0036] The LCD controller 208 is an integrated circuit (IC) that controls the display state of an LCD, which is an example of the display 150d.

[0037] Note that this system configuration is an example and is not limited to this. For example, the number of each circuit included is not limited to this. Electronic musical instrument 10 may have a configuration that does not include some circuits (mechanisms), or may have a configuration in which the function of one circuit is realized by multiple circuits, or may have a configuration in which the functions of multiple circuits are realized by one circuit.

[0038] Furthermore, electronic musical instrument 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by such hardware. For example, CPU 201 may be implemented by at least one of these pieces of hardware.

[0039] <Generating Acoustic Models> Fig. 3 is a diagram showing an example of the configuration of a voice training unit 301 according to an embodiment. The voice training unit 301 may be implemented as a function executed by a server computer 300 that exists externally, separate from the electronic musical instrument 10 shown in Fig. 1. The voice training unit 301 may also be built into the electronic musical instrument 10 as a function executed by the CPU 201, the voice synthesis LSI 205, or the like.

[0040] The voice training unit 301 and the waveform data output unit 302 described below that realize the voice synthesis in the present disclosure may be implemented based on, for example, a statistical voice synthesis technique based on deep learning.

[0041] The speech training unit 301 may include a training text analysis unit 303 , a training acoustic feature extraction unit 304 , and a model training unit 305 .

[0042] In the voice learning unit 301, for example, recordings of a singer singing a plurality of songs in an appropriate genre are used as the learning singing voice data 312. Also, the learning singing voice data 311 is prepared as the lyric text of each song.

[0043] The training text analysis unit 303 receives training singing data 311 including lyric text and analyzes the data. As a result, the training text analysis unit 303 estimates and outputs a training language feature sequence 313, which is a discrete numerical sequence expressing phonemes, pitches, etc. corresponding to the training singing data 311.

[0044] The training acoustic feature extraction unit 304 receives and analyzes training singing voice data 312 collected via a microphone or the like by a singer singing lyrics text corresponding to the training singing voice data 311 in accordance with the input of the training singing voice data 311. As a result, the training acoustic feature extraction unit 304 extracts and outputs a training acoustic feature sequence 314 representing features of voice corresponding to the training singing voice data 312.

[0045] In the present disclosure, the training acoustic feature sequence 314 and the acoustic feature sequence corresponding to the acoustic feature sequence 317 described later include acoustic feature data (which may be referred to as formant information, spectral information, etc.) that models the human vocal tract, and vocal cord sound source data (which may be referred to as sound source information) that models the human vocal cords. Examples of the spectral information that can be used include Mel-cepstrum and Line Spectral Pairs (LSP). Examples of the sound source information that can be used include a fundamental frequency (F0) and a power value that indicate the pitch frequency of human speech.

[0046] The model training unit 305 estimates, by machine learning, an acoustic model that maximizes the probability that a training acoustic feature sequence 314 will be generated from the training language feature sequence 313. That is, the relationship between the language feature sequence, which is text, and the acoustic feature sequence, which is speech, is represented by a statistical model called an acoustic model. The model training unit 305 outputs model parameters that represent the acoustic model calculated as a result of machine learning as the training result 315. Therefore, this acoustic model corresponds to a trained model.

[0047] An HMM (Hidden Markov Model) may be used as the acoustic model represented by the learning result 315 (model parameters).

[0048] The HMM acoustic model may learn how the vocal tract vibration and vocal cord characteristic parameters change over time when a singer sings lyrics to a melody. More specifically, the HMM acoustic model may be a phoneme-by-phoneme model of the spectrum, fundamental frequency, and their time structure obtained from training vocal data.

[0049] First, we will explain the processing of speech training unit 301 in Fig. 3, which employs an HMM acoustic model. Model training unit 305 in speech training unit 301 may input training language feature sequence 313 output by training text analysis unit 303 and training acoustic feature sequence 314 output by training acoustic feature extraction unit 304, to train an HMM acoustic model that maximizes the likelihood.

[0050] The spectral parameters of singing voices can be modeled using a continuous HMM. However, the logarithmic fundamental frequency (F0) is a variable-dimensional time-series signal that takes continuous values ​​in voiced sections and has no value in unvoiced sections, so it cannot be directly modeled using conventional continuous or discrete HMMs. Therefore, we use a multi-space probability distribution (MSD-HMM), an HMM based on a probability distribution in multispace that corresponds to variable dimensions, to simultaneously model the mel-cepstrum as a multidimensional Gaussian distribution for the spectral parameters, and the logarithmic fundamental frequency (F0) as a Gaussian distribution in one-dimensional space for voiced sounds and zero-dimensional space for unvoiced sounds.

[0051] Furthermore, it is known that the characteristics of the phonemes that make up singing voices vary due to the influence of various factors, even if the acoustic characteristics of the phonemes are the same. For example, the spectrum and logarithmic fundamental frequency (F0) of a phoneme, which is a basic phonetic unit, differ depending on the singing style, tempo, surrounding lyrics, pitch, etc. Factors that affect such acoustic features are called contexts.

[0052] In one embodiment of the statistical speech synthesis process, an HMM acoustic model (context-dependent model) that takes context into account may be employed to accurately model the acoustic features of speech. Specifically, the training text analysis unit 303 may output a training linguistic feature sequence 313 that takes into account not only the phonemes and pitch of each frame, but also the immediately preceding and succeeding phonemes, the current position, the immediately preceding and succeeding vibrato, and accent. Furthermore, context clustering based on a decision tree may be used to efficiently combine contexts.

[0053] For example, the model training unit 305 may generate, as the training result 315, a state duration decision tree for determining state duration from the training language feature sequence 313 corresponding to the context of a large number of phonemes related to state duration extracted by the training text analysis unit 303 from the training singing data 311.

[0054] Furthermore, the model training unit 305 may generate, as the training result 315, a mel-cepstral parameter decision tree for determining mel-cepstral parameters from the training acoustic feature sequence 314 corresponding to a large number of phonemes related to the mel-cepstral parameters extracted from the training singing voice data 312 by the training acoustic feature extraction unit 304.

[0055] Furthermore, the model training unit 305 may generate, as the training result 315, a logarithmic fundamental frequency decision tree for determining the logarithmic fundamental frequency (F0) from the training acoustic feature sequence 314 corresponding to a large number of phonemes related to the logarithmic fundamental frequency (F0) extracted by the training acoustic feature extraction unit 304 from the training singing voice data 312. Note that the voiced and unvoiced sections of the logarithmic fundamental frequency (F0) may be modeled as one-dimensional and zero-dimensional Gaussian distributions, respectively, by an MSD-HMM compatible with variable dimensions, and a logarithmic fundamental frequency decision tree may be generated.

[0056] Note that an acoustic model based on a deep neural network (DNN) may be adopted instead of or in addition to an HMM-based acoustic model. In this case, the model training unit 305 may generate model parameters representing a nonlinear transformation function of each neuron in the DNN from linguistic features to acoustic features as the training result 315. The DNN makes it possible to represent the relationship between a linguistic feature sequence and an acoustic feature sequence using a complex nonlinear transformation function that is difficult to represent using a decision tree.

[0057] Furthermore, the acoustic models of the present disclosure are not limited to these, and any speech synthesis method may be adopted as long as it is a technology that uses statistical speech synthesis processing, such as an acoustic model that combines HMM and DNN.

[0058] The learning results 315 (model parameters) may be stored in the ROM 202 of the control system of the electronic musical instrument 10 of FIG. 2 when the electronic musical instrument 10 of FIG. 1 is shipped from the factory, as shown in FIG. 3, and may be loaded from the ROM 202 of FIG. 2 to the singing voice control unit 306 (described later) in the waveform data output unit 211 when the electronic musical instrument 10 is powered on.

[0059] For example, as shown in FIG. 3, the learning result 315 may be downloaded from an external source, such as the Internet, to the singing voice control unit 306 in the waveform data output unit 211 via the network interface 219 by the performer operating the switch panel 140b of the electronic musical instrument 10.

[0060] <Speech synthesis based on acoustic models> FIG. 4 is a diagram illustrating an example of the waveform data output unit 302 according to an embodiment.

[0061] The waveform data output unit 302 includes a processing unit (which may be called a text processing unit, a preprocessing unit, etc.) 306, a singing voice control unit (which may be called an acoustic model unit) 307, a sound source 308, a singing voice synthesis unit (which may be called a vocalization model unit) 309, etc.

[0062] The waveform data output unit 302 receives singing voice data 215 including lyrics and pitch information instructed by the CPU 201 via the key scanner 206 of Fig. 2 based on key depressions on the keyboard 140k of Fig. 1, synthesizes and outputs singing voice waveform data 217 corresponding to the lyrics and pitch. In other words, the waveform data output unit 302 executes statistical voice synthesis processing to synthesize singing voice waveform data 217 corresponding to the singing voice data 215 including lyrics text by predicting it using a statistical model called an acoustic model set in the singing voice control unit 306.

[0063] Furthermore, when song data is played back, the waveform data output section 302 outputs song waveform data 218 corresponding to the corresponding song playback position.

[0064] The processing unit 307 receives singing voice data 215 including information on the phonemes, pitch, etc. of the lyrics specified by the CPU 201 in Fig. 2 as a result of a performer's performance in sync with the automatic performance, for example, and analyzes the data. The singing voice data 215 may include, for example, data on the nth note (which may also be called the nth note) (for example, pitch and note duration data), singing voice data of the nth note, etc.

[0065] For example, the processing unit 307 may determine whether or not lyrics are progressing based on note-on / off data, pedal-on / off data, etc. acquired from the operation of the keyboard 140k and pedals 140p, in accordance with a lyrics progression control method described later, and acquire singing voice data 215 corresponding to the lyrics to be output. Then, the processing unit 307 may analyze the pitch data specified by the key depression and the acquired singing voice data 215 to generate a linguistic feature sequence 316 representing phonemes, parts of speech, words, etc., corresponding to the acquired singing voice data 215, and output the linguistic feature sequence 316 to the singing voice control unit 306.

[0066] The singing voice data may be information including at least one of the lyrics (characters), syllable type (start syllable, middle syllable, end syllable, etc.), lyrics index, corresponding pitch (correct pitch), and corresponding pronunciation period (e.g., pronunciation start timing, pronunciation end timing, pronunciation duration) (correct pronunciation period).

[0067] For example, in the example of Figure 4, the singing data 215 may include information on singing data of the nth lyric corresponding to the nth note (n = 1, 2, 3, 4, ...) and the specified timing at which the nth note should be played (nth singing playback position).

[0068] The singing voice data 215 may include information (such as data in a specific audio file format or MIDI data) for playing the accompaniment (song data) corresponding to the lyrics. When the singing voice data is represented in SMF format, the singing voice data 215 may include a track chunk in which data related to the singing voice is stored and a track chunk in which data related to the accompaniment is stored. The singing voice data 215 may be read from the ROM 202 to the RAM 203. The singing voice data 215 is stored in memory (for example, the ROM 202 or the RAM 203) before performance.

[0069] The electronic musical instrument 10 may control the progress of the automatic accompaniment based on events indicated by the singing data 215 (for example, meta events (timing information) that indicate the timing and pitch of lyrics, MIDI events that indicate note-on or note-off, or meta events that indicate the beat).

[0070] The singing control unit 306 estimates a corresponding acoustic feature sequence 317 based on the language feature sequence 316 input from the processing unit 307 and the acoustic model set as the learning result 315, and outputs formant information 318 corresponding to the estimated acoustic feature sequence 317 to the singing synthesis unit 309.

[0071] For example, when an HMM acoustic model is adopted, the singing control unit 306 connects HMMs by referring to a decision tree for each context obtained from the language feature sequence 316, and predicts the acoustic feature sequence 317 (formant information 318 and vocal cord sound source data 319) that maximizes the output probability from each of the connected HMMs.

[0072] When a DNN acoustic model is employed, the singing voice control unit 306 may output an acoustic feature sequence 317 on a frame-by-frame basis in response to a phoneme sequence of a linguistic feature sequence 316 input on a frame-by-frame basis.

[0073] In FIG. 4, the processing unit 307 acquires instrument sound data (pitch information) corresponding to the pitch of the pressed key from a memory (which may be the ROM 202 or the RAM 203 ), and outputs it to the sound source 308 .

[0074] The sound source 308 generates a sound source signal (which may also be called instrument sound waveform data) of instrument sound data (pitch information) corresponding to the sound to be generated (note-on) based on the note-on / off data input from the processing unit 307, and outputs the signal to the singing synthesis unit 309. The sound source 308 may also perform control processing such as envelope control of the generated sound.

[0075] The singing synthesis unit 309 forms a digital filter that models the vocal tract based on the series of formant information 318 sequentially input from the singing control unit 306. The singing synthesis unit 309 also uses the sound source signal input from the sound source 309 as an excitation source signal, applies the digital filter, and generates and outputs singing voice waveform data 217 as a digital signal. In this case, the singing synthesis unit 309 may be called a synthesis filter unit.

[0076] The singing voice synthesis unit 309 may be capable of employing various voice synthesis methods, including the cepstrum voice synthesis method and the LSP voice synthesis method.

[0077] In the example of Figure 4, the singing voice waveform data 217 that is output uses the sound of an instrument as the sound source signal, so it is somewhat less faithful than the singer's singing voice, but it is a singing voice that retains both the atmosphere of the instrument sound and the vocal quality of the singer, so that effective singing voice waveform data 217 can be output.

[0078] The sound source 309 may process the instrument sound waveform data and also output the output of other channels as song waveform data 218. This allows for operations such as generating accompaniment sounds using normal instrument sounds, or generating instrument sounds for a melody line and simultaneously vocalizing the melody.

[0079] 5 is a diagram showing another example of the waveform data output unit 302 according to an embodiment. Contents that overlap with those in FIG. 4 will not be described again.

[0080] 5 estimates an acoustic feature sequence 317 based on the acoustic model, as described above. Then, the singing control unit 306 outputs formant information 318 corresponding to the estimated acoustic feature sequence 317 and vocal cord sound source data (pitch information) 319 corresponding to the estimated acoustic feature sequence 317 to the singing synthesis unit 309. The singing control unit 306 may estimate the acoustic feature sequence 317 so as to maximize the probability that the acoustic feature sequence 317 is generated.

[0081] The singing synthesis unit 309 may generate data (which may be called, for example, singing waveform data of the nth lyric corresponding to the nth note) for generating a signal by applying a digital filter that models the vocal tract based on the series of formant information 318 to, for example, a pulse train (in the case of voiced phonemes) that is periodically repeated at the fundamental frequency (F0) and power value included in the vocal cord sound source data 319 input from the singing control unit 306, or white noise (in the case of unvoiced phonemes) having the power value included in the vocal cord sound source data 319, or a signal obtained by mixing these, and outputting the data to the sound source 308.

[0082] Based on the note-on / off data input from the processing unit 307, the sound source 308 generates and outputs singing waveform data 217 of a digital signal from the singing waveform data of the nth lyric corresponding to the note to be pronounced (note-on).

[0083] In the example of Figure 5, the output singing waveform data 217 is a signal that is completely modeled by the singing control unit 306, as it uses the sound generated by the sound source 308 based on the vocal cord sound source data 319 as the sound source signal, and can output singing waveform data 217 that is very faithful to the singer's singing voice and has a natural singing voice.

[0084] In this way, unlike existing vocoders (a method of synthesizing speech by inputting spoken words via a microphone and replacing them with instrument sounds), the voice synthesis disclosed herein can output synthesized speech by operating the keyboard without the user (performer) having to sing (in other words, without the user having to input voice signals that are pronounced in real time into the electronic musical instrument 10).

[0085] As explained above, by adopting statistical speech synthesis processing technology as the speech synthesis method, it is possible to realize a significantly smaller memory capacity than conventional speech segment synthesis methods. For example, electronic musical instruments using speech segment synthesis require hundreds of megabytes of memory for storing speech segment data, but in this embodiment, only a few megabytes of memory is required to store the model parameters of the learning result 315. This makes it possible to realize electronic musical instruments at a lower price, and makes it possible for a wider range of users to use high-quality singing voice performance systems.

[0086] Furthermore, in conventional fragment data methods, manual adjustment of fragment data is required, which requires an enormous amount of time (years) and effort to create data for vocal performance, but in the creation of model parameters for the training result 315 for an HMM acoustic model or DNN acoustic model according to this embodiment, almost no data adjustment is required, so the creation time and effort is reduced to a fraction of that. This also makes it possible to realize a more affordable electronic musical instrument.

[0087] Furthermore, general users can use the learning functions built into the server computer 300, speech synthesis LSI 205, etc., available as a cloud service, to train their own voice, the voice of a family member, or the voice of a celebrity, etc., and use this as a model voice to play a singing voice on an electronic musical instrument. In this case, too, it becomes possible to realize singing voice performances with much more natural and high-quality sound than ever before, on a lower-cost electronic musical instrument.

[0088] (Lyric progression control method) Lyric progression control methods according to embodiments of the present disclosure are described below. Each lyric progression control method may be used by the processing unit 307 of the electronic musical instrument 10 described above, for example.

[0089] The subject of the operations in each of the following flowcharts (electronic musical instrument 10) may be interpreted as either the CPU 201, the waveform data output unit 211 (or the tone generator LSI 204 and voice synthesis LSI 205 therein), or a combination of these. For example, each operation may be performed by the CPU 201 executing a control processing program loaded from the ROM 202 to the RAM 203.

[0090] Note that an initialization process may be performed before starting the flow described below. This initialization process may include interrupt processing, lyric progression, deriving TickTime, which serves as the reference time for automatic accompaniment, tempo setting, song selection, song loading, instrument sound selection, and other button-related processing.

[0091] The CPU 201 can detect operations of the switch panel 140b, keyboard 140k, pedals 140p, etc. based on an interrupt from the key scanner 206 at an appropriate timing, and can perform corresponding processing.

[0092] Note that, although an example of controlling the progression of lyrics is shown below, the target of progression control is not limited to this. For example, based on the present disclosure, instead of lyrics, the progression of any character string, sentence (e.g., a news script), etc. may be controlled. In other words, the lyrics of the present disclosure may be interchangeably read as characters, character strings, etc.

[0093] 6 is a diagram showing an example of a flowchart of a lyrics progression control method according to an embodiment. Note that, although the generation of synthetic speech in this example is based on FIG. 5, it may also be based on FIG.

[0094] First, the electronic musical instrument 10 assigns 0 to a lyrics index (also referred to as "n") indicating the current position in the lyrics (step S101). Note that if lyrics are to be started from the middle (for example, starting from the previously stored position), a value other than 0 may be assigned to n.

[0095] The lyric index may be a variable indicating which syllable (or character) from the beginning of the lyrics when the entire lyrics are considered as a character string. For example, the lyric index n may indicate the singing data at the nth playback position in the singing data 215 shown in Figures 4 and 5. In the present disclosure, the lyrics corresponding to one lyric position (lyric index) may correspond to one or more characters constituting one syllable. The syllables included in the singing data may include various syllables, such as vowels only, consonants only, or a combination of a consonant and a vowel.

[0096] Step S101 may be performed when a performance starts (for example, when playback of song data starts), when singing voice data is read, or the like.

[0097] The electronic musical instrument 10 may play song data (accompaniment) corresponding to the lyrics in response to, for example, a user's operation (step S102). The user can perform key operations in time with the accompaniment to advance the lyrics and perform the performance.

[0098] The electronic musical instrument 10 determines whether the playback of the song data that began in step S102 has finished (step S103). If the playback has finished (step S103-Yes), the electronic musical instrument 10 may end the processing of this flowchart and return to a standby state.

[0099] In this case, the electronic musical instrument 10 may read the vocal data designated by the user's operation as the object of progression control in step S102, and determine in step S103 whether or not all of the vocal data has progressed.

[0100] If playback of the song data has not ended (step S103—No), the electronic musical instrument 10 determines whether the pedal is on (whether the pedal is being depressed) (step S111). If the pedal is on (step S111—Yes), the electronic musical instrument 10 determines whether a new key has been pressed (a note-on event has occurred) (step S112). If a new key has been pressed (step S112—Yes), the electronic musical instrument 10 increments the lyric index n (step S113). This increment is basically by 1 (n is assigned n+1), but a value greater than 1 may be added.

[0101] After incrementing the lyric index, the electronic musical instrument 10 performs sound generation processing for the nth vocal data (step S114). An example of this processing will be described later. Then, the electronic musical instrument 10 decrements n by the same value as the increment in step S113 (in FIG. 6, n is assigned n-1) (step S115). In other words, if the pedal is on, n is maintained before and after the key is pressed, and the lyrics do not progress.

[0102] Next, the electronic musical instrument 10 determines whether a new key has been released (a note-off event has occurred) (step S116). If a new key has been released (step S116-Yes), the electronic musical instrument 10 performs a muting process on the corresponding vocal data (step S117).

[0103] Next, the electronic musical instrument 10 determines whether the pedal is off and all the keys are off (step S118). If the pedal is off and all the keys are off (step S118-Yes), the electronic musical instrument 10 performs synchronization processing between the lyrics and the song (accompaniment) (step S119). The synchronization processing will be described later.

[0104] On the other hand, if the pedal is off (step S111-No), the electronic musical instrument 10 determines whether a new key has been pressed (a note-on event has occurred) (step S122). If a new key has been pressed (step S122-Yes), the electronic musical instrument 10 increments the lyric index n (step S123). This increment is basically by 1 (substituting n+1 for n), but a value greater than 1 may be added.

[0105] After incrementing the lyrics index, the electronic musical instrument 10 performs sound generation processing for the n-th singing voice data (step S124). This processing may be the same as the processing in step S114.

[0106] In other words, when the pedal is off, n is incremented before and after the key is pressed, so the lyrics progress.

[0107] Next, the electronic musical instrument 10 determines whether a new key has been released (a note-off event has occurred) (step S126). If a new key has been released (step S126-Yes), the electronic musical instrument 10 performs muting processing on the corresponding vocal data (step S127).

[0108] After steps S119, S126-No and S127, the process returns to step S103.

[0109] Note that S113 and S115 may be omitted. This allows the pronunciation process to be performed without progressing the lyrics. If S113 and S115 are present, the singing voice data pronounced by S114 will be the (n+1)th data, but if S113 and S115 are not present, the singing voice data pronounced by S114 will be the nth data.

[0110] The determination in S111 may be reversed, that is, may be interpreted as whether the pedal is off or not (Yes if the pedal is off).

[0111] For the sound that is already being played, the electronic musical instrument 10 may continue to output the same sound (or the vowel of the same sound) without advancing the lyrics, or may output the sound based on the advanced lyrics. Also, when the electronic musical instrument 10 pronounces a sound corresponding to the value of the same lyric index as the sound that is already being played, it may be output so as to pronounce the vowel of the lyric. For example, when the lyric "Sle" is already being pronounced and the same lyric is newly pronounced, the electronic musical instrument 10 may newly pronounce the sound "e".

[0112] In addition, when the electronic musical instrument 10 of the present disclosure simultaneously pronounces a plurality of sounds, each sound may be pronounced using a synthesized voice of a different timbre. For example, when the user presses four keys, the electronic musical instrument 10 may perform voice synthesis and output so as to correspond to the voices of soprano, alto, tenor, and bass timbres in order from the highest sound.

[0113] <Pronunciation processing of the n-th singing voice data> The pronunciation processing of the n-th singing voice data in step S114 will be described in detail below.

[0114] FIG. 7 is a diagram showing an example of a flowchart of the pronunciation processing of the n-th singing voice data.

[0115] The processing unit 307 of the electronic musical instrument 10 inputs the pitch data specified by key pressing and the n-th singing voice data to the singing voice control unit 306 (step S114-1).

[0116] Then, the singing voice control unit 306 of the electronic musical instrument 10 estimates the acoustic feature quantity series 317 based on the input, and outputs the corresponding formant information 318 and vocal tract sound source data (pitch information) 319 to the singing voice synthesis unit 309. Also, the singing voice synthesis unit 309 generates what may be called the n-th singing voice waveform data (the singing voice waveform data of the n-th lyric corresponding to the n-th note) based on the input formant information 318 and vocal tract sound source data (pitch information), and outputs it to the sound source 308. Then, the sound source 308 acquires the n-th singing voice waveform data from the singing voice synthesis unit 309 (step S114-2).

[0117] The electronic musical instrument 10 performs sound generation processing on the acquired n-th singing voice waveform data using the sound source 308 (step S114-3).

[0118] FIG. 8 is a diagram showing an example of lyric progression controlled using the lyric progression determination process. In this example, a case where the user presses the keys according to the musical score shown in the figure is explained. For example, the treble clef may be pressed with the user's right hand, and the bass clef may be pressed with the user's left hand. Also, lyric indexes 1-6 correspond to "Sle," "e," "ping," "heav," "en," and "ly," respectively.

[0119] Also, assume that the user turns on the pedal at t1 and turns it off at t2. Similarly, assume that the user turns on the pedal at t3 and turns it off before t5. Similarly, assume that the user turns on the pedal at t5 and turns it off before the scheduled start of the next measure.

[0120] First, at timing t1, four keys are pressed. The electronic musical instrument 10 performs the determination process of Fig. 6, and since steps S111 and S112 are Yes, in step S113 the lyric index is incremented by 1, and the lyric "Sle" is generated and output using four-part synthesized sounds. In step S115, the lyric index is reset to its original value.

[0121] Next, at timing t2, the user moves his left hand to the "D" key while continuing to press the right hand key. The electronic musical instrument 10 performs the determination process of FIG. 6, and since step S111 is determined to be No, in step S123 the lyric index is incremented by 1, and the note "D" is generated using the lyric "Sle" and output. The electronic musical instrument 10 continues to produce the other three notes.

[0122] Similarly, at t3, the electronic musical instrument 10 outputs the lyric "e" using the notes corresponding to the four keys, and at t4, it updates only the notes that are newly pressed for the lyric "e." Also, at t5, the electronic musical instrument 10 outputs the lyric "ping" using the notes corresponding to the four keys, and at t6, it updates only the notes that are newly pressed for the lyric "ping."

[0123] In the example section t1-t6 in Figure 8, the lyrics of the upper triad were assigned one melody per note, and the lyrics progressed with each key press. On the other hand, the bass clef part was assigned one melody per two notes, and there were some parts where the lyrics did not progress with each key press due to pedal operation.

[0124] <Synchronization processing> The synchronization process may be a process of aligning the lyrics position with the playback position of the current song data (accompaniment). This process allows the lyrics position to be moved appropriately when the lyrics position is exceeded due to too many keys being pressed or when the lyrics position is slower than expected due to not enough keys being pressed.

[0125] FIG. 9 is a diagram illustrating an example of a flowchart of the synchronization process.

[0126] The electronic musical instrument 10 acquires the playback position of the song data (step S119-1), and then determines whether the playback position matches the (n+1)th vocal playback position (step S119-2).

[0127] The (n+1)th vocal reproduction position may indicate a desired timing at which the (n+1)th note is reproduced, which is derived in consideration of the total note length of the vocal data up to the nth note.

[0128] If the song data playback position and the (n+1)th vocal playback position match (step S119-2-Yes), the synchronization process may be terminated. If not (step S119-2-No), the electronic musical instrument 10 may obtain the Xth vocal playback position closest to the song data playback position (step S119-3), assign X-1 to n (step S119-4), and terminate the synchronization process.

[0129] Note that if no accompaniment is being played, the synchronization process may be omitted. Also, if the appropriate timing for sounding lyrics is derived based on the singing voice data, even if no accompaniment is being played, the electronic musical instrument 10 may perform a process for adjusting the position of the lyrics to the position that would be obtained if they were properly sounded, depending on the elapsed time from the start of performance to the present, the number of key presses, etc.

[0130] According to the embodiment described above, lyrics can be smoothly progressed even when multiple keys are pressed simultaneously.

[0131] (Variation) 4, 5, etc. may be switched on / off based on the user's operation of the switch panel 140b. When it is off, the waveform data output unit 211 may be controlled to generate and output a sound source signal of musical instrument sound data of a pitch corresponding to the key pressed.

[0132] Some steps may be omitted from the flowcharts such as Figure 6. When a determination process is omitted, the determination may be interpreted as proceeding to a route of always "Yes" or always "No" in the flowchart.

[0133] The electronic musical instrument 10 is only required to be able to control the position of lyrics, and is not necessarily required to generate or output sounds corresponding to the lyrics. For example, the electronic musical instrument 10 may transmit sound waveform data generated based on key depressions to an external device (such as the server computer 300), and the external device may generate / output synthetic voices based on the sound waveform data.

[0134] The electronic musical instrument 10 may control the display 150d to display lyrics. For example, lyrics near the current lyrics position (lyric index) may be displayed, or lyrics corresponding to sounds being produced or sounds that have been produced may be displayed in color so that the current lyrics position can be identified.

[0135] The electronic musical instrument 10 may transmit at least one of singing voice data, information on the current lyric position, etc. to an external device. The external device may control the display of lyrics on its own display based on the received singing voice data, information on the current lyric position, etc.

[0136] In the above example, the electronic musical instrument 10 is a keyboard instrument such as a keyboard, but is not limited to this. The electronic musical instrument 10 may be any device that allows the user to specify the timing of sound generation, such as an electric violin, electric guitar, drums, or trumpet.

[0137] Therefore, the term "key" in this disclosure may be interpreted as a string, a valve, another performance operator for specifying pitch, any performance operator, etc. The term "key pressing" in this disclosure may be interpreted as striking a key, picking, playing, operating an operator, etc. The term "key release" in this disclosure may be interpreted as stopping a string, stopping playing, stopping (not operating) an operator, etc.

[0138] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the means for realizing each functional block is not particularly limited. That is, each functional block may be realized by a single physically coupled device, or may be realized by two or more physically separate devices connected by wire or wirelessly.

[0139] In addition, terms explained in this disclosure and / or terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings.

[0140] The information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or other corresponding information. Furthermore, the names used for parameters, etc. in this disclosure are not limiting in any way.

[0141] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0142] Information, signals, etc. may be input / output via multiple network nodes. The input / output information, signals, etc. may be stored in a specific location (e.g., memory) or may be managed using a table. The input / output information, signals, etc. may be overwritten, updated, or added. Output information, signals, etc. may be deleted. Input information, signals, etc. may be transmitted to another device.

[0143] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0144] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.

[0145] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched depending on the implementation. Furthermore, the order of the processing procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless inconsistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the specific order presented.

[0146] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0147] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0148] When used in this disclosure, the terms "include," "including," and variations thereof are intended to be inclusive, similar to the term "comprising." Furthermore, when used in this disclosure, the term "or" is not intended to be an exclusive or.

[0149] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0150] The following notes are provided regarding the above embodiments. (Appendix 1) a plurality of first performance operators (e.g., keys) each associated with a different pitch data (e.g., note number); a second performance control (e.g., a pedal); a processor, the processor comprising: When a first user operation on the first performance operator is detected while no operation on the second performance operator is detected, and a second user operation on the first performance operator after the first user operation is detected, instruct the vocals to pronounce a first lyric (e.g., "Sle" in FIG. 8) corresponding to a first note (e.g., a note with a pitch of A5 in FIG. 8) in accordance with the first user operation, and instruct the vocals to pronounce a second lyric (e.g., "e" in FIG. 8) corresponding to a second note following the first note (e.g., a note with a pitch of E5 in FIG. 8) in accordance with the second user operation (in other words, progress the lyrics in accordance with the second user operation); When a first user operation on the first performance operator is detected while an operation on the second performance operator is being detected, and a second user operation on the first performance operator after the first user operation is detected, control is performed so that the pronunciation of a singing voice corresponding to the first lyrics is instructed in response to the first user operation, and the pronunciation of a singing voice corresponding to the second lyrics is not instructed in response to the second user operation (in other words, the lyrics are not advanced in response to the second user operation). Electronic musical instrument.

[0151] (Appendix 2) The processor: When the first user operation and the second user operation are detected while an operation on the second performance operator is being detected, instruct the pronunciation of the singing voice corresponding to the first lyrics in accordance with the second user operation (for example, instruct the pronunciation to extend the vowels of the singing voice of the first lyrics as they are, after changing the pitch in accordance with the second user operation); 1. An electronic musical instrument as described in Appendix 1.

[0152] (Appendix 3) The processor: when the first user operation and the second user operation are detected in a state where no operation on the second performance operator is detected, instruct, in response to the first user operation, to produce a singing voice corresponding to the first lyrics at a first pitch designated by the first user operation, and instruct, in response to the second user operation, to produce a singing voice corresponding to the second lyrics at a second pitch designated by the second user operation; when the first user operation and the second user operation are detected in a state in which an operation on the second performance operator is detected, instruct, in response to the first user operation, to produce a singing voice corresponding to the first lyrics at a first pitch designated by the first user operation, and instruct, in response to the second user operation, to produce a singing voice corresponding to the first lyrics at a second pitch designated by the second user operation; 1. An electronic musical instrument as defined in Appendix 1 or 2.

[0153] (Appendix 4) The processor: Instruct the playback of accompaniment data (e.g., song data), In response to detection of a change from a state in which a user operation on the first performance operator is detected to a state in which it is not detected, a determination is made as to whether an operation on the second performance operator is detected and whether a user operation on any of the first performance operators is detected (in other words, a determination is made as to whether the pedal is off and all keys are off); when no operation on the second performance operator is detected and no user operation on any of the first performance operators is detected, a first playback position in singing voice text data including the first lyrics data and the second lyrics data, which is to be sung in response to a next user operation, is changed to a second playback position corresponding to a playback position in the accompaniment data (in other words, performing a synchronization process); 4. An electronic musical instrument according to any one of appendices 1 to 3.

[0154] (Appendix 5) The processor: When the first user operation and the second user operation are detected while no operation on the second performance operator is detected, inputting the first lyric data into a trained model in response to the first user operation to instruct pronunciation according to the singing voice data output by the trained model, and inputting the second lyric data into the trained model in response to the second user operation to instruct pronunciation according to the singing voice data output by the trained model; When the first user operation and the second user operation are detected while an operation on the second performance operator is detected, the first lyric data is input into a trained model in response to the first user operation, thereby instructing pronunciation according to the singing voice data output by the trained model, and the first lyric data is input into the trained model in response to the second user operation, thereby instructing pronunciation according to the singing voice data output by the trained model. 5. An electronic musical instrument according to any one of appendices 1 to 4.

[0155] (Appendix 6) The trained model is generated by machine learning using singing voice data of a certain singer as training data, and outputs singing voice data inferred from the singing voice of the certain singer in response to input lyric data. 1. An electronic musical instrument as described in Appendix 5.

[0156] (Appendix 7) Electronic musical instrument computers, when a first user operation on a first performance operator is detected while no operation on a second performance operator is detected, and a second user operation on the first performance operator after the first user operation is detected, instruct the device to produce a singing voice according to first lyrics corresponding to a first note in response to the first user operation, and instruct the device to produce a singing voice according to second lyrics corresponding to a second note following the first note in response to the second user operation; When a first user operation on the first performance operator is detected while an operation on the second performance operator is being detected, and a second user operation on the first performance operator after the first user operation is detected, control is performed to instruct the pronunciation of a singing voice corresponding to the first lyrics in response to the first user operation, and not to instruct the pronunciation of a singing voice corresponding to the second lyrics in response to the second user operation. method.

[0157] (Appendix 8) Electronic musical instrument computers, a step of instructing, in response to the first user operation, to produce a vocal sound according to first lyrics corresponding to a first note, and in response to the second user operation, to produce a vocal sound according to second lyrics corresponding to a second note following the first note, when a first user operation on a first performance operator is detected while no operation on a second performance operator is detected, and a second user operation on the first performance operator after the first user operation is detected; and when a first user operation on the first performance operator is detected while an operation on the second performance operator is being detected, and a second user operation on the first performance operator after the first user operation is detected, execute a procedure of instructing the pronunciation of a singing voice corresponding to the first lyrics in response to the first user operation, and controlling the pronunciation of a singing voice corresponding to the second lyrics not to be instructed in response to the second user operation. program.

[0158] Although the invention according to the present disclosure has been described in detail above, it is clear to those skilled in the art that the invention according to the present disclosure is not limited to the embodiments described in the present disclosure. The invention according to the present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the invention as defined by the description of the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not impose any limiting meaning on the invention according to the present disclosure.

Claims

1. during playback of accompaniment data, an operation to a second performance operator is detected to stop the progression of vocal data that is being advanced in response to the detection of an operation to a first performance operator, and then, when neither the operation to the first performance operator nor the operation to the second performance operator is detected, a synchronization process is executed to align the vocal playback position of the vocal data with the playback position of the accompaniment data. An electronic device that includes a control unit.

2. the first performance operator includes a keyboard; The electronic device according to claim 1 .

3. The control unit controlling the singing voice reproduction position of the singing voice data not to advance in response to a user operation on the first performance operator while an operation on the second performance operator is being detected; 3. The electronic device according to claim 1 or 2.

4. The voice at the singing voice playback position of the singing voice data is generated based on acoustic feature data output from a trained model by inputting singing voice data corresponding to lyrics into the trained model. The electronic device according to claim 1 .

5. The trained model is generated by machine learning using singing voice data of a certain singer as training data, and outputs singing voice data inferred from the singing voice of the certain singer in response to input of lyrics data.

5. The electronic device according to claim 4.

6. The electronic device according to claim 1; the first performance operator; Electronic musical instruments, including

7. A method in which at least one processor of an electronic device performs the process of any one of claims 1 to 5.

8. A program causing at least one processor of an electronic device to execute the process according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Electronic musical instrument

    JP1992349497A

  • Electronic music instrument, control method of electronic music instrument, and program

    JP2019219569A

  • Electronic musical instrument, method, and program

    JP2021099462A

  • Apparatus and program for vocal synthesis

    JP4735544B2