Acoustic output system, acoustic output device, information processing device, and speech sound generation method

The acoustic output system allows electronic musical instruments to generate speech by synthesizing spectral parameters with acoustic data, enabling voice production through user interactions.

JP2026056710APending Publication Date: 2026-04-02CASIO COMPUTER CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing electronic musical instruments cannot produce a voice based on acoustic data output from a speaker.

Method used

An acoustic output system comprising an acoustic output device and an information processing device, where the acoustic output device generates acoustic data in response to user operation and outputs it to the information processing device, which synthesizes spectral parameters with the acoustic data to generate speech data, which is then output to the acoustic output device to produce speech.

Benefits of technology

The system enables an electronic musical instrument to produce a singing voice based on acoustic data, allowing users to enjoy making the instrument produce speech through playing operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026056710000001_ABST
    Figure 2026056710000001_ABST
Patent Text Reader

Abstract

Based on the audio data output from the audio output device in response to user input, the audio output device is made to emit sound. [Solution] The sound output system comprises an electronic musical instrument 2 as a sound output device and a terminal device 3. The electronic musical instrument 2 generates sound data in response to user operation and outputs the generated sound data to the terminal device 3. The terminal device 3 generates speech data by combining spectral parameters with the sound data output from the electronic musical instrument 2 and outputs the generated speech data to the electronic musical instrument 2. The electronic musical instrument 2 produces sound based on the speech data output from the terminal device 3.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an acoustic output system, an acoustic output device, an information processing device, and a voice pronunciation method.

Background Art

[0002] There is known an electronic musical instrument that generates and outputs singing voice output data based on musical sound output data (excitation source signal) for a voice source generated and output by a sound source LSI based on a user's performance operation.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the technique described in Patent Document 1, it is not possible to cause a voice to be pronounced on the electronic musical instrument based on waveform data of musical sound (referred to as acoustic data) output from a speaker of the electronic musical instrument.

[0005] The present invention has been made in view of the above problems, and an object thereof is to cause a voice to be pronounced on an acoustic output device based on acoustic data output from the acoustic output device.

Means for Solving the Problems

[0006] To solve the above problems, the present invention provides an acoustic output system comprising an acoustic output device and an information processing device, wherein the acoustic output device generates acoustic data in response to user operation and outputs the generated acoustic data to the information processing device, the information processing device generates speech data by synthesizing spectral parameters with the acoustic data output from the acoustic output device and outputs the generated speech data to the acoustic output device, and the acoustic output device produces speech based on the speech data output from the information processing device. [Effects of the Invention]

[0007] According to the present invention, an acoustic output device can be made to produce sound based on acoustic data output from that device. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows an example of the overall configuration of the electronic musical instrument system of the present invention. [Figure 2] Figure 1 is a block diagram showing the functional configuration of an electronic musical instrument. [Figure 3] Figure 1 is a diagram illustrating how the sound changes when the control buttons of an electronic musical instrument are operated. [Figure 4] Figure 1 is a block diagram showing the functional configuration of the terminal device. [Figure 5] Figure 4 is a flowchart showing the flow of external waveform input processing performed by the CPU. [Figure 6] Figure 5 is a flowchart showing the flow of the Note On / Off event generation process. [Figure 7] This graph plots the envelope data generated by the Note On / Off event generation process over time. [Figure 8] Figure 5 is a flowchart showing the flow of the acoustic data processing process. [Figure 9] Figure 4 is a flowchart showing the flow of the singing voice generation process executed by the CPU. [Figure 10]This diagram schematically illustrates the process from operating the electronic musical instrument to producing vocal sounds using the electronic instrument in this embodiment. [Modes for carrying out the invention]

[0009] The embodiments for carrying out the present invention will be described below with reference to the drawings. However, the embodiments described below are subject to various technically preferred limitations for carrying out the present invention. Therefore, the technical scope of the present invention is not limited to the embodiments and illustrated examples below.

[0010] As shown in Figure 1, the electronic musical instrument system 1 (sound output system) according to this embodiment is configured by connecting an electronic musical instrument 2 (sound output device) and a terminal device 3 (information processing device) via a communication interface I (or communication network N).

[0011] The electronic instrument 2 is equipped with a performance control 206 and generates acoustic data (which may be described as an excitation source) in response to the user's operation of the performance control 206, and produces (outputs) musical tones based on the generated acoustic data. Furthermore, when the electronic instrument 2 is connected to the terminal device 3 via a communication interface I (or communication network N) and the user operates the performance control 206, the electronic instrument 2 generates acoustic data in response to the operation of the performance control 206 and outputs the generated acoustic data to the terminal device 3. When the terminal device 3 outputs voice data (which may be described as singing voice data) in response to the output of the acoustic data, the electronic instrument 2 acquires the voice data and produces a singing voice (speech) based on the acquired voice data. It should be noted that the audio data in this proposal is not MIDI data. In other words, the audio data in this proposal does not include data that represents "commands" for playing sound, such as MIDI data. The audio data in this proposal is audio data. That is, it is waveform data like the kind obtained when acquiring external sound from a microphone.

[0012] In this embodiment, as shown in FIG. 1, the electronic musical instrument 2 is a cat-shaped sound output device. However, the proposed sound output device includes electronic musical instruments, electronic toys, electronic stringed instruments, electronic wind instruments, electronic percussion instruments, etc.

[0013] FIG. 2 is a block diagram showing the functional configuration of the control system of the electronic musical instrument 2 in FIG. 1. As shown in FIG. 2, the electronic musical instrument 2 includes a CPU (Central Processing Unit) 201 connected to a timer 210, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, a sound source unit 204, a performance operator 206, a mouth opening / closing unit 207, and a communication unit 208, which are respectively connected to a bus 209. Further, a D / A converter 211 is connected to the sound source unit 204. The acoustic data, which is the waveform data of the musical sound output from the sound source unit 204, is converted into an analog signal by the D / A converter 211, amplified by an amplifier 213, and then output as a musical sound such as an instrument sound from a speaker 214. Also, the voice data (singing voice waveform data) from the terminal device 3 acquired by the communication unit 208 is converted into an analog signal by the D / A converter 211, amplified by the amplifier 213, and then output as a singing voice from the speaker 214.

[0014] The CPU 201 as the control unit is a processor that executes the control operation of the electronic musical instrument 2 in FIG. 1 by executing the program stored in the ROM 202 while using the RAM 203 as a work memory. The CPU 201 may be composed of a plurality of CPUs. In this case, the plurality of CPUs may be involved in common processing, or alternatively, the plurality of CPUs may independently execute different processes in parallel.

[0015] For example, when the performance operator 206 is operated, the CPU 201 causes the sound source unit 204 to generate acoustic data according to the operation of the performance operator 206, and outputs the musical sound based on the generated acoustic data through the D / A converter 211, the amplifier 213, and the speaker 214.

[0016] Also, when the performance operator 206 is operated while being connected to the terminal device 3 via the communication unit 208, the CPU 201 causes the sound source unit 204 to generate acoustic data according to the operation of the performance operator 206, and outputs the generated acoustic data to the terminal device 3 via the communication unit 208. When the voice data generated based on the acoustic data in the terminal device 3 is acquired by the communication unit 208, the CPU 201 causes a voice to be pronounced based on the acquired voice data. That is, the CPU 301 causes the voice based on the acquired voice data to be output via the D / A converter 211, the amplifier 213, and the speaker 214.

[0017] The ROM 202 stores programs and various fixed data, etc. The RAM 203 is a volatile semiconductor memory and forms a work area for temporarily storing various data and programs.

[0018] The sound source unit 204 has a waveform ROM in which acoustic data for generating musical sounds is stored. Here, the musical sound is a musical sound having a timbre emitted by the electronic musical instrument 2 according to the operation of the performance operator 206. The sound source unit 204 reads acoustic data from, for example, a waveform ROM (not shown) based on the pitch information and volume information (velocity value) corresponding to the operation of the performance operator 206 according to the control instruction from the CPU 201, and outputs it to the D / A converter 211. The sound source unit 204 is not limited to the PCM (Pulse Code Modulation) sound source method, and may use other sound source methods such as the FM (Frequency Modulation) sound source method.

[0019] The performance operator 206 is an operator for the user to control the pitch and volume (velocity value). As shown in FIG. 3, the performance operator 206 has a performance operator 206a for controlling the pitch and a performance operator 206b for controlling the volume. In the present embodiment, the right hand of the cat of the electronic musical instrument 2 serves as the performance operator 206a for controlling the pitch, and the left hand serves as the performance operator 206b for controlling the volume.

[0020] For example, when a user touches the right hand of the cat on the electronic instrument 2 and changes the height of the right hand, the performance control 206a outputs a signal to the CPU 201 detecting the height of the right hand. The CPU 201 outputs pitch information corresponding to the detection signal from the performance control 206a to the sound source unit 204. For example, as shown in Figure 3, when it is detected that the right hand is set to the lowest position, it outputs pitch information for the lowest note that the electronic instrument 2 can output (e.g., C), and as the position of the right hand rises, the pitch increases, and when it is detected that the right hand is set to the highest position, it outputs pitch information for the highest note that the electronic instrument 2 can output (e.g., G).

[0021] For example, when a user touches the left hand of the cat on the electronic instrument 2 and changes the height of the left hand, the performance control 206b outputs a detection signal for the height of the left hand to the CPU 201. The CPU 201 outputs volume information corresponding to the detection signal from the performance control 206b to the sound source unit 204. For example, as shown in Figure 3, when the left hand is set to the lowest position, the volume information for the smallest sound that the electronic instrument 2 can output is output, and as the position of the left hand rises, the volume increases, and when the left hand is set to the highest position, the volume information for the largest sound that the electronic instrument 2 can output is output.

[0022] The mouth opening / closing unit 207 has a mechanism that opens and closes the mouth of the cat on the electronic musical instrument 2 based on control from the CPU 201.

[0023] The communication unit 208 transmits and receives data with external devices such as terminal devices 3 connected via a communication network N such as the Internet or a communication interface I such as a USB (Universal Serial Bus) cable.

[0024] The terminal device 3 acquires the acoustic data output from the electronic instrument 2 using the communication unit 307, synthesizes spectral parameters (which may also be expressed as spectral envelope or acoustic features) with the acquired acoustic data to generate audio data, and outputs the generated audio data to the electronic instrument 2.

[0025] As shown in Figure 4, the terminal device 3 is a computer comprising a CPU 301, ROM 302, RAM 303, storage unit 304, operation unit 305, display unit 306, communication unit 307, etc., and each unit is connected by a bus 308. Examples of terminal devices 3 include tablet PCs (Personal Computers), notebook PCs, and smartphones.

[0026] The CPU 301, acting as the control unit, is a processor that controls the operation of each part of the terminal device 3 by reading and executing various programs, including the singing voice generation application 302a, stored in the ROM 302, while using the RAM 303 as work memory. The CPU 301 may be composed of multiple CPUs. In this case, multiple CPUs may be involved in common processing, or multiple CPUs may independently execute different processing in parallel.

[0027] ROM302 is a non-temporary recording medium readable by the CPU301 as a computer, and stores various data, including the singing voice generation application 302a and the trained model 302b. The singing voice generation application 302a is an application program for the CPU301 to execute the singing voice generation function described later. The trained model 302b is generated by machine learning on multiple datasets consisting of musical score data (lyrics data (text information of lyrics) and pitch data (including information on note length)) of multiple songs, and audio data of each song sung by a certain singer. When the trained model 302b receives lyric data and pitch data of any song (even a phrase), it infers a set of singing voice parameters (called singing voice information) to produce a singing voice equivalent to that of the singer who generated the trained model 302b when singing the input song. The pitch data input to the trained model 302b may be tailored to the song, or it may be a fixed value. If it is a fixed value, for example, it is preferable to use C3 as the reference tone, or E4 for a female voice.

[0028] RAM303 is a volatile semiconductor memory that forms a work area for temporarily storing various data and programs. In this embodiment, for example, a voice generation buffer 303a used by the voice generation application 302a is formed in RAM303.

[0029] The memory unit 304 is composed of a non-volatile semiconductor memory or an HDD (Hard Disk Drive), etc., and stores various data. The memory unit 304 may also store the singing voice generation application 302a and the trained model 302b.

[0030] The control unit 305 consists of push-button switches and a touch panel attached to the display unit 306. The control unit 305 detects user operations on the push-button switches and touch operations on the screen, and outputs operation signals to the CPU 301.

[0031] The display unit 306 consists of an LCD (Liquid Crystal Display), an EL (Electro Luminescence) display, etc., and displays various information according to the display information instructed by the control unit 11.

[0032] The communication unit 307 transmits and receives data with external devices such as electronic musical instruments 2 connected via a communication network N such as the Internet or a communication interface I such as a USB (Universal Serial Bus) cable.

[0033] Next, the operation of the electronic musical instrument system 1 will be described. In the terminal device 3, when the operation unit 305 instructs the generation of singing parameters and the lyrics data and pitch data of any song (or phrase, hereinafter the same) to be played by the electronic musical instrument 2 are input via the communication unit 307, etc., the CPU 301 causes the trained model 302b to generate singing information. That is, the CPU 301 inputs the input lyrics data and pitch data to the trained model 302b, causes the trained model 302b to infer a group of singing parameters, and stores the singing information, which is the inferred group of singing parameters, in the storage unit 304. Note that the lyrics data and pitch data may be stored in the storage unit 304 in advance. In addition, accompaniment data (sound waveform data of the accompaniment) corresponding to the lyrics data and pitch data may be stored in the storage unit 304 in association with the lyrics data and pitch data.

[0034] Here, we will explain the vocal information. Each section of a song, divided into predetermined time units in the time direction, is called a frame, and the trained model 302b generates vocal parameters on a frame-by-frame basis. That is, the vocal information for a song generated by the trained model 302b consists of multiple vocal parameters (a time-series group of vocal parameters) on a frame-by-frame basis. The vocal parameters on a frame-by-frame basis include spectral parameters (the frequency spectral envelope of the voice being produced) and fundamental frequency F0 parameters (the base pitch frequency of the voice being produced, which can also be described as the excitation source).

[0035] Furthermore, when the terminal device 3 is instructed to generate audio data by operating the operation unit 305, the CPU 301 starts the singing voice generation application 302a and executes the following processes.

[0036] First, the CPU 301 initializes the buffer (vocal generation buffer 303a), various variables (previous average amplitude value, i), flags (Note On Flag), arrays, parameters, etc., used by the vocal generation application 302a. The CPU 301 also prompts the user to select a song from the songs whose vocal information is stored in the memory unit 304 that the electronic instrument 2 should play, and reads the vocal information of the selected song into the RAM 303. The CPU 301 also displays a message on the display unit 306, for example, "Please connect electronic instrument 2," to notify the user to connect electronic instrument 2. Furthermore, while the vocal generation application 302a is running, the CPU 301 executes the external waveform input processing shown in Figure 5 and the vocal generation processing shown in Figure 9 at predetermined intervals (at predetermined time intervals).

[0037] The user connects the electronic instrument 2 and the terminal device 3 via communication and plays the electronic instrument 2. Specifically, as shown in Figure 3, the user plays the cat-shaped electronic instrument 2 by raising and lowering the hands, which are the performance controls 206a and 206b. The CPU 201 of the electronic instrument 2 generates sound data via the sound source unit 204 in response to the operation of the performance controls 206a and 206b, and transmits the sound data to the terminal device 3 via the communication unit 208.

[0038] The CPU 301 of the terminal device 3 performs the external waveform input processing shown in Figure 5 at predetermined intervals, thereby generating Note On or Note Off events that trigger the generation of voice data in the singing voice generation process described later, based on the acoustic data output from the electronic instrument 2, and optimizing the acoustic data as excitation source waveform data for singing voices.

[0039] In the external waveform input processing, first, the CPU 301 acquires the acoustic data output from the electronic instrument 2 and acquired by the communication unit 307 up to the size of the singing voice generation buffer 303a and stores it in the singing voice generation buffer 303a (step S1).

[0040] Next, the CPU 301 executes the Note On / Off event generation process (step S2). As shown in Figure 6, in the Note On / Off event generation process, first, the CPU 301 obtains the maximum amplitude value (absolute value) of the acoustic data in the singing voice generation buffer 303a (step S201).

[0041] Next, the CPU 301 calculates the current average amplitude value (step S202). The current average amplitude value can be calculated using the following (Equation 1). Note that the previous average amplitude value is a variable set in step S203, which will be described later, and its initial value is 0. Envelope data is generated using the following (Equation 1). Current average amplitude = (Previous average amplitude + Current maximum amplitude) ÷ 2 ... (Equation 1)

[0042] In this embodiment, envelope data is generated by a moving average of the maximum amplitude values, but the method of generating envelope data is not limited to this. For example, envelope data may be generated by performing a Fourier transform (FFT) on the input acoustic data, obtaining the signal, performing a Hilbert transform (multiplying with a 90-degree phase shift), and then performing an inverse Fourier transform. Furthermore, in order to optimize the Note On / Note Off timing, the method of generating envelope data may be changed when generating Note On events and when generating Note Off events.

[0043] Next, the CPU 301 sets the current average amplitude value to the previous average amplitude value (variable) (step S203).

[0044] Next, CPU301 determines whether the Note On Flag is set to OFF (step S204). The Note On Flag is set to On when a Note On event is generated and to OFF when a Note Off event is generated. The initial value is set to OFF.

[0045] If the CPU 301 determines that the Note On Flag is set to OFF (step S204; YES), it determines whether the current average amplitude value, which is the envelope data, is greater than a preset Note On threshold (first threshold) (step S205). If the CPU 301 determines that the current average amplitude value is less than or equal to the preset Note On threshold (step S205; NO), it proceeds to step S3 in Figure 5.

[0046] If the CPU 301 determines that the current average amplitude value is greater than a preset Note On threshold (step S205; YES), it generates a Note On event and outputs it to the singing voice generation process (step S206). Then, the CPU 301 sets the Note On Flag to On (step S207) and proceeds to step S3 in Figure 5. The Note On event is an event (note on data) that indicates that a performance operation (Note On) has been performed on the electronic instrument 2.

[0047] On the other hand, if in step S204 it is determined that the Note On Flag is not set to OFF (i.e., it is set to ON) (step S204; NO), the CPU 301 determines whether the current average amplitude value is smaller than the preset Note Off threshold (second threshold) (step S208). Here, Note On threshold > Note Off threshold.

[0048] If the CPU 301 determines that the current average amplitude value is equal to or greater than the preset Note Off threshold (step S208; NO), it proceeds to step S3 in Figure 5.

[0049] If the CPU 301 determines that the current average amplitude value is smaller than the preset Note Off threshold (step S208; YES), it generates a Note Off event and outputs it to the singing voice generation process (step S209). Then, the CPU 301 sets the Note On Flag to Off (step S210) and proceeds to step S3 in Figure 5. The Note On event is an event (note off data) that indicates that the performance operation has been canceled (Note Off) on the electronic instrument 2.

[0050] Figure 7 is a graph plotting the envelope data (current average amplitude value) generated by the Note On / Off event generation process over time. A Note On event is generated at time T1, as shown in Figure 7, and a Note Off event is generated at time T2.

[0051] Returning to Figure 5, in step S3, the CPU 301 performs sound data processing (step S3). As shown in Figure 8, in the sound data processing, the CPU 301 first determines whether the electronic instrument 2 is in Note On or Release mode (step S301). The CPU 301 determines that the instrument is in Note On mode if a Note On event has been generated but a Note Off event has not yet been generated (i.e., the Note On Flag is set to On). The CPU 301 also determines that the instrument is in Release mode if a Note Off event has been generated but the value of the sound data has not yet become 0.

[0052] If the CPU 301 determines that the electronic instrument 2 is neither Note On nor Released (step S301; NO), it sets the values ​​of all acoustic data stored in the vocal generation buffer 303a to 0 (step S302) and proceeds to step S4 in Figure 5.

[0053] If the CPU 301 determines that the electronic instrument 2 is in Note On or Release mode (step S301; YES), it sets the variable i to 0 (step S303) and acquires a noise waveform (noise data) (step S304). As the noise waveform, a predetermined length of PCM (Pulse Code Modulation) waveform can be used, such as an actual voiceless noise waveform, white noise, or pink noise waveform. The noise waveform data is stored in advance in, for example, the ROM 302 or the storage unit 304.

[0054] Next, the CPU 301 determines whether or not the electronic instrument 2 is in the release stage (step S305). If it determines that the electronic instrument 2 is not in the release stage (step S305; NO), the CPU 301 proceeds to step S310.

[0055] If the CPU determines that electronic instrument 2 is in the release phase (step S305; YES), the CPU 301 subtracts the release coefficient (step S306). The release coefficient is a coefficient used to attenuate the noise waveform during the release phase. For example, the CPU 301 sets the initial value of the release coefficient to 100 / 100 and subtracts 2 / 100 each time.

[0056] Next, the CPU 301 determines whether the release coefficient is < 0 (step S307). If it determines that the release coefficient is < 0 (step S307; YES), the CPU 301 sets the release coefficient to 0 (step S308) and proceeds to step S309. If it determines that the release coefficient is not < 0 (step S307; NO), the CPU 301 proceeds to step S309.

[0057] In step S309, the CPU 301 uses the value obtained by multiplying the noise waveform value by the release coefficient as the noise waveform (step S309), and proceeds to step S310.

[0058] In step S310, the CPU 301 multiplies the value of waveform[i] (amplitude value) by an amplification coefficient and adds a noise coefficient to obtain the value of waveform[i] (step S310). Waveform[i] is the i-th acoustic data from the beginning among the acoustic data in the singing voice generation buffer 303a. The amplification coefficient is a coefficient for amplifying the value of waveform[i]. That is, the amplification coefficient > 1. Since the pronunciation of consonants contains noise components, adding a noise coefficient to waveform[i] makes it possible to approximate the pronunciation of a singing voice. The amplification coefficient may be predetermined, or it may be changed based on the maximum amplitude value so as constant a distortion as possible can be obtained.

[0059] Next, the CPU 301 determines whether waveform[i] > clip level (step S311). The clip level is a predetermined value that is the upper limit of the amplitude value of waveform[i]. If it is determined that waveform[i] > clip level (step S311; YES), the CPU 301 replaces the value of waveform[i] with the clip level (i.e., the upper limit) (step S312) and proceeds to step S313. In other words, if waveform[i] > clip level, the value of waveform[i] is clipped up to the clip level (i.e., the upper limit). The processing in steps S311 to S312 can increase the harmonics in the acoustic data, so the acoustic data can be made closer to waveform data that has the characteristics of vocal cords. If it is determined that waveform[i] > clip level (the value of waveform[i] is less than or equal to the clip level) (step S311; YES), the CPU 301 proceeds to the processing in step S313.

[0060] In the example above, the processing method described was amplification and clipping of the acoustic data to approximate the waveform data representing the characteristics of the vocal cords, but the processing method is not limited to this. For example, one could perform a Fourier transform (FFT) on both the acoustic data and the pre-prepared vocal cord waveform data, then approximate the values ​​of the Fourier-transformed acoustic data with those of the Fourier-transformed speech waveform data, and finally perform an inverse Fourier transform.

[0061] In step S313, the CPU 301 determines whether the value obtained by incrementing i is less than the number of data points in the singing voice generation buffer 303a (step S313). If it determines that the value obtained by incrementing i is less than the number of data points in the singing voice generation buffer 303a (step S313; YES), the CPU 301 increments i (step S314) and returns to step S304. The CPU 301 repeatedly executes the processes in steps S304 to S314 until it determines that the value obtained by incrementing i is greater than or equal to the number of data points in the buffer. Through the processes in steps S304 to S314, the CPU 301 can optimize the acoustic data as an excitation source for the singing voice.

[0062] In step S313, if the CPU 301 determines that the value obtained by incrementing i is greater than or equal to the number of data items in the buffer (step S313; NO), the CPU 301 proceeds to step S4 in Figure 5.

[0063] In step S4 of Figure 5, the CPU 301 outputs the processed acoustic data as excitation source waveform data to the singing voice synthesis process (step S4), and terminates the external waveform input process.

[0064] Furthermore, the CPU 301 of terminal device 3 executes the singing voice generation process shown in Figure 9 at predetermined intervals. The singing voice generation process starts the singing voice synthesis process when a Note On event is generated in the external waveform input process, and stops the singing voice synthesis process when a Note Off event is generated.

[0065] In the singing voice generation process, first, the CPU 301 determines whether or not a Note On event has been generated (step S11).

[0066] If the CPU 301 determines that a Note On event has been generated (step S11; YES), it starts the singing voice synthesis process (step S12) and then terminates the singing voice generation process. The singing voice synthesis process generates audio data by combining the spectral parameters of the singing voice information stored in the memory unit 304 with the acoustic data (excitation source waveform data) output from the external waveform input process, and then outputs (transmits) the generated audio data to the electronic instrument 2 via the communication unit 307. When the CPU 301 outputs (transmits) the audio data to the electronic instrument 2 via the communication unit 307 during the singing voice synthesis process, it may also output the corresponding accompaniment data to the electronic instrument 2 at the same time.

[0067] If it is determined that no Note On event has been generated (Step S11; NO), the CPU 301 determines whether or not a Note Off event has been generated (Step S13). If it is determined that no Note Off event has been generated (Step S13; NO), the CPU 301 terminates the singing voice generation process. If it is determined that a Note Off event has been generated (Step S13; YES), the CPU 301 stops the singing voice synthesis process (Step S14) and terminates the singing voice generation process.

[0068] In the electronic musical instrument system 1 described above, as shown in Figure 10, when the user raises and lowers the hand portion of the cat-shaped electronic musical instrument 2 to perform a performance operation, acoustic data corresponding to the performance operation is output to the terminal device 3. The terminal device 3 synthesizes the acoustic data from the electronic musical instrument 2 with spectral parameters generated by the trained model 302b to generate sound data, which is output to the electronic musical instrument 2. The electronic musical instrument 2 outputs a singing voice from the speaker 214 based on the sound data received from the terminal device 3. At this time, the CPU 201 of the electronic musical instrument 2 opens and closes the mouth opening / closing part 207 in accordance with the output of the singing voice. In other words, when the user performs a performance operation on the electronic musical instrument 2, the cat's mouth on the electronic musical instrument 2 moves in response to the performance operation and produces a singing voice (sings). Therefore, the user can enjoy producing a singing voice from the electronic musical instrument 2 by performing a performance operation on it.

[0069] As described above, the electronic musical instrument system 1 comprises an electronic musical instrument 2 and a terminal device 3. The electronic musical instrument 2 generates acoustic data in response to user operation and outputs the generated acoustic data to the terminal device 3. The terminal device 3 synthesizes spectral parameters with the acoustic data output from the electronic musical instrument 2 to generate speech data and outputs the generated speech data to the electronic musical instrument 2. The electronic musical instrument 2 produces singing voices based on the speech data output from the terminal device 3.

[0070] Therefore, the electronic musical instrument system 1 can cause the electronic musical instrument 2 to produce a singing voice based on the acoustic data output from the electronic musical instrument 2. In other words, it can cause the electronic musical instrument 2, which does not have a function to generate audio data based on the user's playing operations, to produce a singing voice based on the playing operations. The user can enjoy making the electronic musical instrument 2 produce a singing voice by performing the playing operations.

[0071] Furthermore, the CPU 301 of the terminal device 3 generates note-on data indicating that a performance operation has been performed on the electronic instrument 2, based on the acoustic data output from the electronic instrument 2, and generates audio data based on the generated note-on data. For example, the CPU 301 generates envelope data from the acoustic data, and when the generated envelope data reaches a first threshold, it generates note-on data, and generates audio data based on the generated note-on data. Therefore, the timing of note-on (performance operation timing) can be detected from the waveform acoustic data, and the electronic instrument 2 can be made to emit a singing voice at the timing corresponding to the performance operation.

[0072] Furthermore, the CPU 301 of terminal device 3 generates note-off data indicating that the performance operation has been canceled on the electronic instrument 2 when the envelope data generated from the acoustic data falls below a second threshold which is smaller than the first threshold, and stops generating audio data based on the generated note-off data. Therefore, the timing of the note-off (the timing when the performance operation is canceled) can be detected from the acoustic data, which is a waveform, and the electronic instrument 2 can be made to mute the singing voice at the timing corresponding to the performance operation.

[0073] The descriptions in the above embodiments are merely preferred examples of the acoustic output system, acoustic output device, information processing device, and voice generation method according to the present invention, and are not limited thereto.

[0074] For example, the output level of the singing voice based on the generated audio data may be controlled by calculating a velocity value based on the acoustic data and multiplying the calculated velocity value by the audio data as a coefficient. For example, the CPU 301 calculates the difference between the current average amplitude value and the previous average amplitude value of the acoustic data, and calculates a velocity value based on the calculated difference value. Then, it multiplies the audio data by a coefficient based on the calculated velocity value and outputs it to the speaker 214. The velocity value can be calculated, for example, by the following (Equation 2). Velocity value = Maximum velocity value × Difference value ÷ Maximum difference value ... (Equation 2) Alternatively, a conversion table that pre-defines the relationship between the difference value and the velocity value may be stored in the storage unit 304, and the velocity value may be derived from the difference value based on the conversion table.

[0075] Furthermore, while the above embodiment discloses examples using semiconductor memory such as ROM or a hard disk as a computer-readable medium for the program, the invention is not limited to this example. Other computer-readable mediums that can be used include SSDs and portable recording media such as CD-ROMs. Carrier waves can also be used as a medium for providing program data via a communication line.

[0076] Although embodiments of the present invention have been described above, the technical scope of the present invention is not limited to the embodiments described above, but is determined based on the claims. Furthermore, equivalent scopes of the present invention that have been modified from the claims but are not related to the essence of the present invention are also included in the technical scope of the present invention. [Explanation of Symbols]

[0077] 1 Electronic musical instrument system, 2 Electronic musical instruments, 201 CPU, 206 Performance control unit, 208 Communication unit, 214 Speaker, 3 Terminal device, 301 CPU, 307 Communication unit

Claims

1. It comprises an audio output device and an information processing device, The aforementioned acoustic output device is The system generates acoustic data in response to user operations and outputs the generated acoustic data to the information processing device. The aforementioned information processing device is Audio data is generated by combining spectral parameters with the acoustic data output from the aforementioned acoustic output device. The generated audio data is output to the audio output device. The aforementioned acoustic output device is Based on the audio data output from the information processing device, the device produces sound. Audio output system.

2. The system generates acoustic data in response to user operations, and the communication unit outputs the generated acoustic data to the information processing device. In the aforementioned information processing device, the communication unit acquires the audio data generated based on the acoustic data. The system will produce sound based on the acquired audio data. An acoustic output device equipped with a control unit.

3. The audio data output from the audio output device is acquired by the communication unit. The acquired acoustic data is combined with spectral parameters to generate audio data. The generated audio data is output to the audio output device by the communication unit. An information processing device equipped with a control unit.

4. The control unit, Based on the aforementioned acoustic data, note-on data indicating that a performance operation has been performed on the acoustic output device is generated, and based on the generated note-on data, the audio data is generated. The information processing apparatus according to claim 3.

5. The control unit, Envelope data is generated from the aforementioned acoustic data, and when the generated envelope data reaches a first threshold, the note-on data is generated. The information processing apparatus according to claim 4.

6. The control unit, When the envelope data falls below a second threshold, which is smaller than the first threshold, the sound output device generates note-off data indicating that the playback operation has been canceled, and based on the generated note-off data, the generation of the sound data is stopped. The information processing apparatus according to claim 5.

7. A method for producing sound in an acoustic output system comprising an acoustic output device and an information processing device, The aforementioned acoustic output device is The system generates acoustic data in response to user operations and outputs the generated acoustic data to the information processing device. The aforementioned information processing device is Audio data is generated by combining spectral parameters with the acoustic data output from the aforementioned acoustic output device. The generated audio data is output to the audio output device. The aforementioned acoustic output device is Based on the audio data output from the information processing device, the device produces sound. Audio pronunciation method.

Citation Information

Patent Citations

  • Electronic musical instrument, electronic musical instrument control method, and program

    JP6835182B2