Information processing device, electronic musical instrument, information processing system, information processing method, and storage medium

By generating sound data suitable for the vocal range, the problem of incomplete vocal range coverage in the prior art is solved, and natural singing synthesis is achieved.

CN115116414BActive Publication Date: 2025-09-23CASIO COMPUTER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210270338.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-24
Filing Date
2022-03-18
Publication Date
2025-09-23
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

When the existing technology uses an electronic keyboard to synthesize singing, the range of sound is not fully covered, resulting in unnatural vocal switching.

Method used

By detecting the designated pitch, the first and second sound models are used to generate third data corresponding to the pitch, thereby synthesizing a singing voice that fits the vocal range.

Benefits of technology

It achieves complete coverage of the vocal range, eliminates the unnatural feeling of vocal switching, and generates singing that is more in line with the musical range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116414B_ABST
    Figure CN115116414B_ABST
Patent Text Reader

Abstract

An information processing device, an electronic musical instrument, an information processing system, an information processing method, and a storage medium are provided. A processor detects a specified pitch (S201). The processor reads first sound data (223) of a first sound model (221) and second sound data (224) of a second sound model (222) corresponding to the detected pitch from a sound model (220), such as a database system, and generates deformation data based on them (S202). The processor outputs a sound based on the generated deformation data (S203). The sound is output in an optimal range that matches the key range of the music.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority based on Japanese Patent Application No. 2021-45284 (filing date: March 18, 2021), Japanese Patent Application No. 2021-117857 (filing date: July 16, 2021), and Japanese Patent Application No. 2021-190167 (filing date: November 24, 2021). Technical Field

[0002] The present invention relates to an information processing device, an electronic musical instrument, an information processing system, an information processing method and a storage medium. Background Art

[0003] The following prior art is known: based on the stored lyric data, corresponding parameters and co-articulation parameters are read from a phoneme database, the corresponding sound is synthesized and output through the resonance peak synthesis sound source unit, and clear consonants are emitted through the PCM sound source, thereby synthesizing high-quality singing sounds corresponding to the lyric data (for example, refer to Japanese Patent Gazette No. 3233036).

[0004] The human vocal range is typically around two octaves. Therefore, if the aforementioned prior art technology is applied to a 61-key electronic keyboard, assigning a single vocalist to all the keys results in a range that cannot be fully covered by a single voice. On the other hand, even if multiple voices are used to cover the range, there will be an unnatural and uncomfortable feeling when the vocalists switch places. Summary of the Invention

[0005] Therefore, an object of the present invention is to enable generation of audio data suitable for a sound range.

[0006] An information processing device according to an example of the technical solution detects a designated pitch and generates third data corresponding to the designated pitch based on first data of a first sound model and second data of a second sound model different from the first sound model.

[0007] According to the present invention, it is possible to generate audio data suitable for a sound range. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 This is an operation explanation diagram of the first embodiment.

[0009] Figure 2 This is a flowchart showing an outline of the operation of the first embodiment.

[0010] Figure 3 It is a diagram showing an example of the appearance of an electronic keyboard instrument according to the second embodiment.

[0011] Figure 4This is a block diagram showing a hardware configuration example of a control system of an electronic keyboard instrument according to a second embodiment.

[0012] Figure 5 This is a block diagram showing a configuration example of a speech synthesis LSI according to the second embodiment.

[0013] Figures 6A to 6C This is a diagram for explaining the operation of the formant interpolation processing unit according to the second embodiment.

[0014] Figure 7 This is a flowchart showing an example of the main processing of singing voice synthesis executed by the CPU in the second embodiment.

[0015] Figure 8 This is a flowchart showing an example of speech synthesis processing executed by the speech synthesis unit of the speech synthesis LSI in the second, third, and fourth embodiments.

[0016] Figure 9 This is a flowchart showing a detailed example of the singing voice optimization process executed by the formant interpolation processing unit of the speech synthesis LSI 405 in the second, third, and fourth embodiments.

[0017] Figure 10 This is a diagram showing a connection configuration of a third embodiment in which a sound synthesizing unit and an electronic keyboard instrument operate independently.

[0018] Figure 11 This is a diagram showing an example of the hardware configuration of the sound synthesis section in the third embodiment in which the sound synthesis section and the electronic keyboard instrument operate independently.

[0019] Figure 12 This is a flowchart showing an example of the main processing of singing voice synthesis in the third and fourth embodiments.

[0020] Figure 13 This is a diagram showing a connection configuration of a fourth embodiment in which a part of the sound synthesizing unit and the electronic keyboard instrument operate independently.

[0021] Figure 14 This is a diagram showing an example of the hardware configuration of a sound synthesis unit in a fourth embodiment in which a part of the sound synthesis unit operates independently from the electronic keyboard instrument.

[0022] Figure 15 This is a block diagram showing a partial configuration example of a speech synthesis LSI and a speech synthesis unit according to the fourth embodiment. DETAILED DESCRIPTION

[0023] Hereinafter, embodiments for implementing the present invention will be described in detail with reference to the accompanying drawings. First, a first embodiment will be described.

[0024] The range of a person's singing voice as an example of sound is usually about two octaves. On the other hand, for example, when realizing a singing function as an information processing device, the range may be specified to exceed the range of a person's singing voice and reach about five octaves, for example.

[0025] Therefore, in the first embodiment, for example, Figure 1 As shown, for the two-octave range 1 on the bass side, the first singing model obtained by modeling a lower pitch, such as a male singing voice, is assigned, and for the two-octave range 2 on the treble side, the second singing model obtained by modeling a higher pitch, such as a female singing voice, is assigned.

[0026] Furthermore, in the first embodiment, for example, Figure 1 As shown, to the non-overlapping range 3 of about two octaves in the middle between ranges 1 and 2, a male-female mid-range singing voice transformed from the first range singing voice of range 1 and the second range singing voice of range 2 is assigned.

[0027] Figure 2 This is a flowchart showing an example of a sound generation process (for example, a singing voice generation process) executed by at least one processor (hereinafter referred to as “processor”) of the information processing device according to the first embodiment.

[0028] First, the processor detects a designated pitch (step S201). When the information processing device is implemented as an electronic musical instrument, for example, the electronic musical instrument includes a performance operating member 210. The processor detects the designated pitch based on the pitch designation data 211 detected by the performance operating member 210.

[0029] Here, the information processing device includes, for example, an acoustic model 220 serving as a database system. The processor reads first acoustic data (first data) 223 of the first acoustic model 221 and second acoustic data (second data) 224 of the second acoustic model 222 from the acoustic model 220, for example, the database system. The processor then generates deformed data (third data) based on the first acoustic data 223 and the second acoustic data 224 (step S202). More specifically, if the acoustic model 220 is a human singing voice model, the processor generates the deformed data based on an interpolation operation between the formant frequency of the first singing voice data corresponding to the first acoustic data 223 and the formant frequency of the second singing voice data corresponding to the second acoustic data 224.

[0030] Here, for example, the first sound model 221 stored as the sound model 220 may include a learned model that has learned the first sound (e.g., the singing voice of the first singer), and similarly, the second sound model 222 stored as the sound model 220 may include a learned model that has learned the second sound (e.g., the singing voice of the second singer).

[0031] The processor outputs the sound based on the deformation data generated in step S202 (step S203).

[0032] Here, for example, there may be a non-overlapping range between the first range corresponding to the first acoustic model 221 and the second range corresponding to the second acoustic model 222, and the pitch detected in step S201 may be included in this non-overlapping range. Furthermore, the deformed data generated in step S202 can be generated even when there is no acoustic model corresponding to the range of the specified song. However, even if there is a range that overlaps with the first and second ranges, the present invention can be applied to generate deformed data based on the acoustic data of multiple acoustic models.

[0033] In the sound generation process of the first embodiment described above, if the music belongs to, for example, Figure 1 For the bass side of range 1, the processor infers the resonance peak frequency based on the first singing voice model corresponding to the first sound model 221, for example, male-like singing voice, which is pre-assigned to the range 1, and outputs the singing voice of the first range corresponding thereto.

[0034] Furthermore, if the music belongs to e.g. Figure 1 For the high-pitched range 2, the processor infers the resonance peak frequency based on the second singing voice model corresponding to the second sound model 222 of female-like singing voice pre-assigned to the range 2, and outputs the singing voice of the corresponding second range.

[0035] On the other hand, if the music belongs to e.g. Figure 1 The processor passes the range 3 between the ranges 1 and 2. Figure 2 The processing of step S202 generates deformed data and outputs sound based on, for example, the first singing voice data corresponding to the first sound data 223 of the first singing voice model of the first sound model 221 corresponding to the male-like singing voice, and the second singing voice data corresponding to the second sound data 224 of the second singing voice model of the second sound model 222 corresponding to the female-like singing voice.

[0036] As a result of the above processing, it is possible to output, for example, singing voice in an optimal range that better matches the key range of the music.

[0037] Next, the second embodiment will be described. Figure 2The object is a singing voice model that models a human singing voice, and the model corresponding to the sound model 220 is used as the target. Figure 3 This figure shows an example of the appearance of an electronic keyboard instrument 300 according to a second embodiment. The electronic keyboard instrument 300 includes a keyboard 301 composed of a plurality of keys serving as operating elements; a first switch panel 302 for controlling various settings, such as volume setting, tempo setting for automatic lyrics playback, and the start of automatic lyrics playback; and a second switch panel 303 for selecting songs and instrument tones. Furthermore, each key of the keyboard 301 is equipped with an LED (Light Emitting Diode) 304. This LED 304 illuminates at full brightness when the key containing it is the next key to be designated for automatic lyrics playback, and at half its full brightness when the key is the next key to be designated for automatic lyrics playback. Furthermore, although not specifically shown, the electronic keyboard instrument 300 is equipped with speakers on the inside, side, or back of the instrument to play musical sounds and vocals.

[0038] Figure 4 This shows the second embodiment. Figure 3 FIG. 4 is a diagram showing an example of the hardware configuration of a control system 400 of an electronic keyboard instrument 300. Figure 4 In the control system 400, a CPU (Central Processing Unit) 401, a ROM (Read Only Memory) 402, a RAM (Random Access Memory) 403, a sound source LSI (Large Scale Integrated Circuit) 404, a sound synthesis LSI 405, Figure 3 The keyboard 301, the key scanner 406 connected to the first switch panel 302 and the second switch panel 303, and Figure 3 The CPU 401 is connected to an LED controller 407 connected to the LEDs 304 provided on each key of the keyboard 301, and a network interface 408 for exchanging MIDI data with an external network. Furthermore, the CPU 401 is connected to a timer 410 for controlling the sequence of automatic reproduction of vocal data. Furthermore, musical sound output data 418 and vocal sound output data 417 output from the sound source LSI 404 and the sound synthesis LSI 405, respectively, are converted by D / A converters 411 and 412 into analog musical sound output signals and analog vocal sound output signals, respectively. The analog musical sound output signals and analog vocal sound output signals are mixed by a mixer 413, and after being amplified by an amplifier 414, the mixed signal is output from a speaker or output terminal (not specifically shown).

[0039] The CPU 401 executes the control program stored in the ROM 402 while using the RAM 403 as a working memory, thereby executing Figure 3The ROM 402 stores, in addition to the control program and various control data, performance instruction data including lyrics data to be described later.

[0040] The CPU 401 is provided with a timer 410 for timing the progress of automatic reproduction of the performance instruction data in the electronic keyboard instrument 300 , for example.

[0041] The sound source LSI 404 reads musical sound waveform data from, for example, a waveform ROM (not shown) according to a sound generation control instruction from the CPU 401, and outputs the data to the D / A converter 411. The sound source LSI 404 is capable of generating a maximum of 256 sounds simultaneously.

[0042] When the speech synthesis LSI 405 receives lyrics information (text data of lyrics) and pitch information about pitch as singing voice data 415 from the CPU 401 , it synthesizes corresponding singing voice data (singing voice output data 417 ) and outputs it to the D / A converter 412 .

[0043] The key scanner 406 continuously scans Figure 3 The key-press / key-release status of the keyboard 301 and the switch operation status of the first switch panel 302 and the second switch panel 303 are interrupted to the CPU 401 to transmit the status changes.

[0044] LED controller 407 is a Figure 3 An IC (integrated circuit) that controls the display status of each LED 304 included in each key on the keyboard 301.

[0045] Figure 5 This is a block diagram showing a configuration example of the speech synthesis unit 500 in the second embodiment. Figure 4 A function performed by the sound synthesis LSI 405.

[0046] The sound synthesis unit 500 is input from Figure 4 The processor of the voice synthesis unit 500 synthesizes and outputs singing voice output data 417 based on singing voice data 415 including lyrics information, pitch information, and range information, as instructed by the CPU 401. At this time, the processor of the voice synthesis unit 500 performs the following vocalization processing: Based on the target sound source information 512 output from the acoustic model unit 501 and the target spectrum information 513 output from the acoustic model unit 501 via the formant interpolation processing unit 506, the processor performs the following vocalization processing: For the acoustic model set in the acoustic model unit 501, the processor outputs singing voice output data 417 that infers the singer's singing voice, corresponding to the singing voice data 415 including lyrics information, pitch information, and range information input from the CPU 401. The voice synthesis unit 500 is implemented, for example, based on the technology described in Japanese Patent No. 6610714.

[0047] The details of the basic operation of the speech synthesis unit 500 are disclosed in the aforementioned patent document. The operation of the speech synthesis unit 500 including the operation unique to the second embodiment will be described below.

[0048] The speech synthesis unit 500 includes a text analysis unit 502, an acoustic model unit 501, an utterance model unit 503, and a formant interpolation processing unit 506. The formant interpolation processing unit 506 is a unit related to functions unique to the second embodiment.

[0049] In the second embodiment, the sound synthesis unit 500 performs the following statistical sound synthesis processing: by making predictions using a statistical model such as the acoustic model set in the acoustic model unit 501, singing sound output data 417 corresponding to singing data 415 including lyrics, pitch and range as text of lyrics is synthesized.

[0050] The text parsing unit 502 inputs Figure 4 The CPU 401 specifies singing voice data 415, which includes information related to lyrics, pitch, and range. The data is analyzed. As a result, the text analysis unit 502 generates a language feature sequence 507 representing phonemes, parts of speech, and words corresponding to the lyrics in the singing voice data 415, as well as pitch information 508 corresponding to the pitches in the singing voice data 415, and provides these to the acoustic model unit 501.

[0051] Furthermore, the text analysis unit 502 generates range information 509 corresponding to the range in the singing voice data 415 and provides it to the formant interpolation processing unit 506. If the range indicated by the range information 509 falls within the range of the first range, which is the currently set range, the formant interpolation processing unit 506 requests spectrum information 510 of the first range (hereinafter referred to as "first range spectrum information 510") from the acoustic model unit 501.

[0052] The first sound range spectrum information 510 may be expressed as first spectrum information, first spectrum data, first sound data, or first data.

[0053] On the other hand, if the range represented by the range information 509 does not fall within the range of the first range, which is the current range, but falls within the range of another new range, the formant interpolation processing unit 506 replaces the new range with the first range and requests the first range spectrum information 510 from the acoustic model unit 501.

[0054] On the other hand, when the range represented by the range information 509 does not enter the range of any range including the first range but enters the range of the range between the above-mentioned first range and another second range, the formant interpolation processing unit 506 requests both the first range spectral information 510 and the second range spectral information 511 (hereinafter referred to as "second range spectral information 511") from the acoustic model unit 501.

[0055] The second sound range spectrum information 511 may be expressed as second spectrum information, second spectrum data, second sound data, or second data.

[0056] The acoustic model unit 501 receives the aforementioned language feature value sequence 507 and pitch information 508 from the text analysis unit 502 , and receives a request specifying the aforementioned range from the formant interpolation processing unit 506 .

[0057] As a result, the acoustic model unit 501 uses, for example, the acoustic model set as the learning result through machine learning to infer the first sound range spectrum, or the first sound range spectrum / second sound range spectrum corresponding to the phoneme that maximizes the generation probability, and provides it to the resonance peak interpolation processing unit 506 as the first sound range spectrum information 510 or the first sound range spectrum information 510 / second sound range spectrum information 511, respectively.

[0058] Furthermore, the acoustic model unit 501 estimates the sound source corresponding to the phoneme that maximizes the generation probability using the acoustic model, and provides it as target sound source information 512 to the sound source generation unit 504 in the utterance model unit 503 .

[0059] The resonance peak interpolation processing unit 506 provides the first sound range spectrum information 510 or the spectrum information obtained by interpolating the first sound range spectrum information 510 and the second sound range spectrum information 511 (hereinafter referred to as "interpolated spectrum information") as the target spectrum information 513 to the synthesis filter unit 505 in the vocal model unit 503.

[0060] The target spectrum information 513 may also be expressed as deformed data or third data.

[0061] The vocalization model unit 503 generates singing voice output data 417 corresponding to the singing voice data 415 by inputting the target sound source information 512 output from the acoustic model unit 501 and the target spectrum information 513 output from the formant interpolation processing unit 506. Figure 4 The D / A converter 412 outputs the signal through the mixer 413 and the amplifier 414 and emits the signal from a speaker (not shown).

[0062] The acoustic feature quantities output by the acoustic model unit 501 include spectral information modeling the human vocal tract and sound source information modeling the human vocal cords. Parameters for the spectral information can include, for example, line spectral pairs (LSPs), line spectral frequencies (LSFs), or improved versions of these, such as Mel-LSPs (hereinafter referred to as "LSPs"), which can efficiently model multiple formant frequencies characteristic of the human vocal tract. Therefore, the first-range spectral information 510 or second-range spectral information 511 output from the acoustic model unit 501, or the target spectral information 513 output from the formant interpolation processing unit 506, can serve as frequency parameters based on, for example, these LSPs.

[0063] As other examples of the parameters of the spectrum information, cepstrum or mel cepstrum may be used.

[0064] As sound source information, the fundamental frequency (F0) representing the pitch frequency of a human voice and its power value (in the case of voiced phonemes) or the power value of white noise (in the case of unvoiced phonemes) can be used. Therefore, the target sound source information 512 output from the acoustic model unit 501 can be used as parameters of the F0 and power value described above.

[0065] The vocalization model unit 503 includes a sound source generation unit 504 and a synthesis filter unit 505. The sound source generation unit 504 is the part that models the human vocal cords. By sequentially inputting a sequence of target sound source information 512 input from the acoustic model unit 501, the sound source generation unit 504 generates sound source input data 514 consisting of, for example, a pulse train that periodically repeats the fundamental frequency (F0) and power value contained in the target sound source information 512 (for voiced phonemes), white noise having the power value contained in the target sound source information 512 (for unvoiced phonemes), or a signal composed of a mixture of these.

[0066] The synthesis filter unit 505 is a part that models the human vocal tract. It forms an LSP digital filter that models the vocal tract based on the LSP frequency parameters contained in the target spectrum information 513 that is sequentially input from the acoustic model unit 501 via the formant interpolation processing unit 506. The digital filter is excited by the sound source input data 514 input from the sound source generation unit 504 as an excitation source signal, and the synthesis filter unit 505 outputs filter output data 515 of a digital signal. The filter output data 515 is Figure 4After the D / A converter 412 converts the sound into an analog singing voice output signal, it is mixed with the analog musical sound output signal output from the sound source LSI 404 via the D / A converter 411 in the mixer 413. After the mixed signal is amplified by the amplifier 414, it is output from a speaker or output terminal not specifically shown in the figure.

[0067] The sampling frequency of the vocal sound output data 417 is, for example, 16 kHz. Furthermore, when LSF parameters obtained through LSP analysis are used as parameters for the first sound range spectrum information 510, the second sound range spectrum information 511, and the target spectrum information 513, the update frame period is, for example, 5 milliseconds, the analysis window length is, for example, 25 milliseconds, the window function is, for example, a Blackman window, and the number of analyses is, for example, 10.

[0068] Based on Figure 3 、 Figure 4 and Figure 5 The overall operation of the second embodiment of the structure of the present invention is briefly described. First, the CPU 401 guides the performance of the music by the performer based on the performance guidance data including at least lyrics information, pitch information and timing information. Specifically, in Figure 4 In the embodiment of the present invention, the CPU 401 sequentially reads out a series of performance instruction data groups for automatic reproduction, which include at least lyrics information, pitch information, and timing information, stored in the ROM 402 as a memory, and automatically reproduces the lyrics information and pitch information contained in the performance instruction data group at a timing corresponding to the timing information contained in the performance instruction data group. The timing can be based on, for example, a timing synchronized with a set performance beat. Figure 4 The interrupt processing of the timer 410 is controlled.

[0069] At this time, the CPU 401 instructs the user to perform key operations in synchronization with the automatic reproduction and perform performance learning (performance practice) by indicating the keys on the keyboard 301 corresponding to the pitch information of the automatic reproduction. More specifically, in the processing of the performance guidance, the CPU 401 synchronizes with the timing of the automatic reproduction, for example, as Figure 3 As indicated by the key with two LEDs 304 lighting up, the LED 304 of the key (operating member) corresponding to the next automatically reproduced pitch information is lit with a stronger brightness, for example, the maximum brightness, and the LED 304 of the key corresponding to the next automatically reproduced pitch information is lit with a weaker brightness, for example, half the brightness of the maximum brightness.

[0070] Next, CPU401 obtains the player's Figure 3 The information related to the performance operation of pressing or releasing the keys on the keyboard 301 is the performance information.

[0071] Next, when the key-pressing timing (operation timing) and the key-pressing pitch (operation pitch) of the key on the keyboard 301 being learned correctly correspond to the timing information and pitch information of the automatic reproduction, the CPU 401 transmits the automatically reproduced lyrics information and pitch information as singing voice data 415 to the player at the key-pressing timing. Figure 5 As a result, as described above, the sound source input data 514 output by the sound source generating unit 504, to which the target sound source information 512 output by the acoustic model unit 501 is set, excites the digital filter of the synthesis filter unit 505 formed based on the target spectrum information 513 output by the acoustic model unit 501 via the formant interpolation processing unit 506, thereby outputting the filter output data 515. The filter output data 515 is used as Figure 4 The singing voice output data 417 is output.

[0072] The singing data 415 may be information including at least one of lyrics (text data), syllable type (start syllable, middle syllable, end syllable, etc.), lyrics index, corresponding pitch (correct pitch) and corresponding pronunciation period (for example, pronunciation start timing, pronunciation end timing, pronunciation length (duration)) (correct pronunciation period).

[0073] For example, as in Figure 5 As illustrated in , the singing voice data 415 may include singing voice data of n-th lyrics corresponding to n-th (n=1, 2, 3, 4, ...) notes, and information on a predetermined timing (n-th singing voice reproduction position) at which the n-th note should be reproduced.

[0074] The vocal data 415 may include information (data in a specific sound file format, MIDI data, etc.) for playing the accompaniment (song data) corresponding to the lyrics. If the vocal data is represented in the SMF format, the vocal data 415 may include a track chunk storing data related to the vocals and a track chunk storing data related to the accompaniment. The vocal data 415 may be read from the ROM 402 into the RAM 403. The vocal data 415 is stored in a memory (e.g., the ROM 402 or RAM 403) before being played.

[0075] In addition, the electronic keyboard instrument 300 can control the progress of automatic accompaniment based on events represented by the singing data 415 (for example, meta events (timing information) indicating the vocalization timing and pitch of lyrics, MIDI events indicating note on or note off, or meta events indicating beats, etc.).

[0076] Here, in the acoustic model unit 501, an acoustic model of singing voice is set as a result of learning by machine learning, for example. As described in the first embodiment, the singing range of a person is generally about two octaves. On the other hand, Figure 3 The keyboard 301 shown has, for example, 61 keys covering 5 octaves.

[0077] Therefore, in the second embodiment, an acoustic model that is the result of learning a low-pitched voice, such as a male singing voice, through machine learning is assigned to the key range 1 of the two octaves on the bass side of the 61-key keyboard 301, and an acoustic model that is the result of learning a high-pitched voice, such as a female singing voice, through machine learning is assigned to the key range 2 of the two octaves on the treble side.

[0078] Furthermore, in the first embodiment, the middle male and female singing voices transformed from the first range singing voice of key range 1 and the second range singing voice of key range 2 are assigned to the key range 3 of the central two octaves of the 61-key keyboard 301 .

[0079] Here, for example, in the vocal data 415 pre-loaded from the ROM 402 to the RAM 403, a meta-event as the beginning may be maintained indicating that the overall average of the music containing the vocal data 415 belongs to Figure 1 Which key field data among the key fields 1, 2, and 3 illustrated in . Figure 5 The text analysis unit 502 may receive key range data as part of the singing voice data 415 from the CPU 201 at the start of singing voice synthesis. The text analysis unit 502 may provide the range information 509 corresponding to the key range data to the formant interpolation processing unit 506 at the start of singing voice synthesis.

[0080] At the start of singing synthesis, the formant interpolation processing unit 506 determines whether the range indicated by the range information 509 belongs to Figure 1 The formant interpolation processing unit 506 determines that the range indicated by the sound source information 319 belongs to the range of the keyboard range 1, 2, or 3. Figure 1 In the case of either the exemplified key range 1 or key range 2, the key range 1 or 2 is set as the first range, and then a request is made to the acoustic model unit 501 to access the acoustic model of the first range.

[0081] As a result, after the start of singing synthesis, the acoustic model unit 501 uses the acoustic model of the first range requested from the formant interpolation processing unit 506 to infer the first range spectrum corresponding to the phoneme that maximizes the generation probability for the language feature sequence 507 and pitch information 508 received from the text analysis unit 502, and provides it to the formant interpolation processing unit 506 as the first range spectrum information 510.

[0082] Through the above control actions, if the music as a whole belongs to Figure 3 For key range 1 on the bass side of the keyboard 301, the acoustic model unit 501 estimates the spectrum based on the acoustic model of, for example, male singing voice, which is pre-assigned to key range 1, and outputs the corresponding first sound range spectrum information 510. Furthermore, the formant interpolation processing unit 506 provides the first sound range spectrum information 510 output from the acoustic model unit 501 as is, as target spectrum information 513, to the synthesis filter unit 505 in the vocalization model unit 503.

[0083] Furthermore, if the piece as a whole belongs to e.g. Figure 1 For key range 2 on the high-pitched side, the acoustic model unit 501 estimates the spectrum based on the acoustic model of, for example, a female singing voice that is pre-assigned to key range 2, and outputs the corresponding first sound range spectrum information 510. Furthermore, the formant interpolation processing unit 506 provides the first sound range spectrum information 510 output from the acoustic model unit 501 as is, as target spectrum information 513, to the synthesis filter unit 505 in the vocalization model unit 503.

[0084] On the other hand, if the music as a whole belongs to e.g. Figure 1 If the key range 3 is in the middle, the resonance peak interpolation processing unit 506 sets the key range 1 and key range 2 on both sides of the key range 3 as the first range and the second range respectively, and then requests the acoustic model unit 501 to access the acoustic models of both the first range and the second range.

[0085] The acoustic model unit 501 outputs two sets of spectrum information: first-range spectrum information 510 corresponding to the spectrum inferred from the acoustic model of male-like singing voices, and second-range spectrum information 511 corresponding to the spectrum inferred from the acoustic model of female-like singing voices, which are pre-assigned to keys 1 and 2 on either side of key range 3. Furthermore, the formant interpolation processing unit 506 calculates interpolated spectrum information by interpolating the first-range spectrum information 510 and the second-range spectrum information 511, and provides this interpolated spectrum information as target spectrum information 513, which is a deformation of the interpolated spectrum information, to the synthesis filter unit 505 within the vocalization model unit 503.

[0086] The target spectrum information 513 may be expressed as deformed data (third sound data), third spectrum information, or the like.

[0087] As a result of the above processing, the filter output data 515 synthesized by the target spectrum information 513 can be output from the synthesis filter unit 505 as the singing sound output data 417, wherein the target spectrum information 513 is based on information of an acoustic model that is the result of machine learning of the singing voice in the optimal key range that is well matched to the key range of the entire music.

[0088] Figures 6A to 6C: is an operation explanation diagram of the formant interpolation processing unit 506. Figures 6A to 6C In each of the graphs shown, the horizontal axis represents frequency [Hz] and the vertical axis represents power [dB].

[0089] Figure 6A 601 is schematically shown Figure 1 The graph of the vocal tract spectrum characteristics of a certain voiced phoneme, such as a male, in key range 1 is shown in FIG. The vocal tract spectrum characteristics 601 of key range 1 can be formed by an LSP digital filter formed based on the LSP parameters L1[i] (1≦i≦N, N is the number of LSP analyses) calculated by LSP analysis. In addition, in FIG. 6 , for the sake of simplicity of explanation, the number of LSP analyses is shown as N=6, but in reality, for example, N=10. In the vocal tract spectrum characteristics 601, F1[1] is the first formant frequency of key range 1, and F1[2] is the second formant frequency of key range 1. The formant frequency is the frequency that forms a peak in the vocal tract spectrum characteristics, which determines the difference in voiced phonemes such as "a", "i", "u", "e", and "o" emitted by the human vocal tract, and also determines the difference in voice quality between males and females. In fact, there are higher-order formant frequencies, but for the sake of simplicity of explanation, the higher-order formant frequencies above the third order are omitted here. The frequency intervals between the LSP parameters L1[i] can well model the spectral characteristics of the human vocal tract. In particular, the sharpness of the peak of the resonance peak frequency (the width of the frequency interval between the top and the bottom of the peak) and the intensity (power) can be represented by the frequency intervals between adjacent LSP parameters L1[i].

[0090] If the piece as a whole belongs to e.g. Figure 1 For key range 1 on the bass side, the acoustic model unit 501 estimates a spectrum based on the acoustic model of, for example, a male singing voice that is pre-assigned to key range 1, and outputs LSP parameters L1[i] (1≦i≦N) corresponding to the spectrum as first sound range spectrum information 510. Furthermore, the formant interpolation processing unit 506 supplies the LSP parameters of the first sound range spectrum information 510 output from the acoustic model unit 501 as they are to the synthesis filter unit 505 in the vocalization model unit 503 as LSP parameters of the target spectrum information 513.

[0091] Figure 6B The 602 is for Figure 6A The same voiced phonemes are schematically represented Figure 1 Graph showing the vocal tract spectrum characteristics of, for example, a female voice in key range 2. Vocal tract spectrum characteristics 602 for key range 2 can be realized using an LSP digital filter formed using LSP parameters L2[i] (1≦i≦N, where N is the number of LSP analyses) calculated based on LSP analysis. In vocal tract spectrum characteristics 602, F2[1] is the first formant frequency of key range 2, and F2[2] is the second formant frequency of key range 2. Figure 6B each element in Figure 6A is the same as the case of

[0092] If the whole piece of music belongs to, for example, Figure 1 the key range 2 on the bass side, the acoustic model unit 501 infers the spectrum according to the acoustic model of, for example, a female-like voice pre-assigned to this key range 2, and outputs the LSP parameters L2[i] (1 ≤ i ≤ N) corresponding to this spectrum as the first pitch range spectrum information 510. And, the formant interpolation processing unit 506 provides the LSP parameters of the first pitch range spectrum information 510 output from the acoustic model unit 501 as the LSP parameters of the target spectrum information 513 to the synthesis filter unit 505 in the sound generation model unit 503 as they are.

[0093] Comparing Figure 6A with Figure 6B it can be seen that Figure 1 the difference between the male-like voice in the key range 1 and the female-like voice in the key range 2 of Figure 5 is significantly shown as the difference in the pitch frequency in the target sound source information 512 of Figure 5 (the female is about twice that of the male). In addition, regarding the formant frequency, it is known that the first formant frequency F2[1] and the second formant frequency F2[2] of the female-like voice in the key range 2 are respectively higher frequencies than the first formant frequency F1[1] and the second formant frequency F1[2] of the male-like voice in the key range 1 (refer to the following literature).

[0094] [粕谷等,“年龄、性别带来的日语5个元音的音调频率和共振峰频率的变化”(日语原文:年齢,性別による日本語5母音のピッチ周波数とホルマント周波数の変化),声学学会杂志(日语原文:音響学会誌)24,,6(1968)]

[0095] In addition, for the sake of easy understanding of the explanation, the Figure 6A vocal tract spectrum characteristics 601 corresponding to the same voiced phoneme and Figure 6B the vocal tract spectrum characteristics 602 of Figure 6B slightly exaggerate and depict the difference in the formant frequency.

[0096] Figure 6C The 603 of Figure 6A 、 Figure 6B schematically represents Figure 1The graph of the vocal tract spectrum characteristics of the key range 3, for example, a voice between a male and a female, is shown in FIG. The first formant frequency F3[1] in the vocal tract spectrum characteristics 603 of the key range 3 has a frequency intermediate between the first formant frequency F1[1] of the male-like voice of the key range 1 and the first formant frequency F2[1] of the female-like voice of the key range 2. Similarly, the second formant frequency F3[2] in the vocal tract spectrum characteristics 603 of the key range 3 has a frequency intermediate between the second formant frequency F1[2] of the male-like voice of the key range 1 and the second formant frequency F2[2] of the female-like voice of the key range 2.

[0097] That is, it can be seen that the vocal tract spectrum characteristics 603 of the male-female singing voice in key range 3 can be calculated through frequency domain interpolation processing based on the vocal tract spectrum characteristics 601 of the male-like voice in key range 1 and the vocal tract spectrum characteristics 602 of the female-like voice in key range 2.

[0098] Specifically, it is known that the above-mentioned LSP parameters have a frequency dimension, and thus have excellent interpolation characteristics in the frequency domain. Therefore, in the second embodiment, when the entire music belongs to, for example, Figure 1 In the case of the middle key range 3, as described above, the acoustic model unit 501 outputs two spectral information items, namely, the first range spectral information 510 corresponding to the spectrum inferred based on the acoustic model of male-like singing voice and the second range spectral information 511 corresponding to the spectrum inferred based on the acoustic model of female-like singing voice, which are pre-assigned to the key ranges 1 and 2 on both sides of the key range 3.

[0099] Furthermore, the formant interpolation processing unit 506 calculates the LSP parameters L3[i] of the key range 3 as the interpolated spectrum information by performing the interpolation operation represented by the following equation (1) between the LSP parameters L1[i] of the first-range spectrum information 510 and the LSP parameters L2[i] of the second-range spectrum information 511. Here, N is the number of LSP analyses.

[0100] L3[i]=(L1[i]+L2[i]) / 2(1≦i≦N)…(1)

[0101] Figure 5 The formant interpolation processing unit 506 uses the LSP parameter L3[i] (1≦i≦N) calculated by the operation of the above formula (1) as Figure 5 The target spectrum information 513 is provided to the synthesis filter unit 505 in the utterance model unit 503.

[0102] As a result of the above processing, the synthesis filter unit 505 can output filter output data 515 synthesized using target spectrum information 513 having optimal vocal tract spectrum characteristics that well match the key range of the entire music as singing voice output data 417.

[0103] The following pairs have Figures 3 to 5 The detailed operation of the second embodiment of the structure is described. Figure 7 This is a flowchart showing an example of the main processing of the singing voice synthesis in the second embodiment. Figure 4 The CPU 401 loads the singing voice synthesis program stored in the ROM 402 into the RAM 403 and executes the processing.

[0104] First, the CPU 401 substitutes an initial value "1" for the lyrics index variable n as a variable on the RAM 403 indicating the current position of the lyrics, and initializes the first range variable as a variable on the RAM 403 indicating the current range, for example, Figure 1 In addition, when starting the lyrics from the middle (for example, starting from the last storage position), a value other than "0" can be substituted into the lyrics index variable n.

[0105] The lyrics index variable n can also be a variable that indicates the syllable (or character) corresponding to the first syllable (or character) when the lyrics are viewed as a string. For example, the lyrics index variable n can be represented by Figure 5 The singing voice data for the nth reproduction position of the singing voice data 415 represented in . Furthermore, in the present invention, the lyrics corresponding to the position of one lyric (the value of the lyric index variable n) may correspond to one or more characters constituting one syllable. The syllables included in the singing voice data may include vowels only, consonants only, consonants plus vowels, and other syllables.

[0106] Next, before the start of the singing voice synthesis, the CPU 401 sends a signal to the voice synthesis LSI 405 indicating that the overall average of the music to be reproduced is Figure 1 The key range data of which key range among the key ranges 1, 2, and 3 illustrated in the example is read from RAM 403, and the key range data is included in the singing voice data 415 of the designated range, and the singing voice data 415 is sent to the designated range. Figure 4 The sound synthesis LSI 405 sends it (step S702).

[0107] Then, CPU401 repeatedly performs a series of processing from steps S703 to S710 while incrementing the value of the lyrics index variable n by +1 each time in step S707, until it is determined in step S710 that the reproduction of the singing data has ended (there is no longer any singing data corresponding to the new value of the lyrics index variable n), thereby performing singing synthesis processing.

[0108] In a series of repeated processing from step S703 to step S710, the CPU 401 first determines Figure 4 The key scanner 406 will Figure 3The keyboard 301 is scanned to determine whether there is a new key (step S703).

[0109] If the determination in step S703 is YES, the CPU 401 reads the singing voice data of the n-th lyric indicated by the value of the lyric index variable n on the RAM 403 from the RAM 403 (step S704 ).

[0110] Next, the CPU 401 transmits the singing voice data 415 indicating the progress of the singing voice including the singing voice data read out in step S704 to the speech synthesis LSI 405 (step S705 ).

[0111] Furthermore, the CPU 401 specifies the pitch corresponding to the key pressed by the player on the keyboard 301 detected by the key scanner 406, and specifies the pitch of the player's key. Figure 3 The pronunciation instruction of the musical instrument sound pre-specified on the switch panel 303 is sent to the sound source LSI 404 as the pronunciation control data 416 (step S706).

[0112] As a result, the sound source LSI 404 generates musical sound output data 418 corresponding to the sound production control data 416. This musical sound output data 418 is converted into an analog musical sound output signal by the D / A converter 411. This analog musical sound output signal is mixed in the mixer 413 with the analog singing sound output signal output from the sound synthesis LSI 405 via the D / A converter 412. The mixed signal is amplified by the amplifier 414 and then output from a speaker or output terminal (not shown).

[0113] In addition, the process of step S706 may not be performed. In this case, no musical sound corresponding to the key operation performed by the performer is produced, and the key operation is used only for the progress of the singing voice synthesis.

[0114] Next, the CPU 401 increments the value of the lyrics index variable n by +1 (step S707 ).

[0115] After the processing of step S707 or after the determination of step S703 is "No", the CPU 401 determines Figure 4 The key scanner 406 will Figure 3 The keyboard 301 is scanned to determine whether there is a new key (step S708).

[0116] If the determination in step S708 is "YES," the CPU 401 instructs the voice synthesis LSI 405 to mute the singing voice corresponding to the pitch of the key release detected by the key scanner 406, and instructs the sound source LSI 404 to mute the musical sound corresponding to the pitch (step S709). Consequently, the voice synthesis LSI 405 and the sound source LSI 404 execute the corresponding mute operations.

[0117] After processing step S709 or if the determination in step S708 is "No", CPU 401 determines whether there is no singing data corresponding to the value of the lyrics index variable n incremented in step S707 on RAM 403 and the reproduction of the singing data has ended (step S710).

[0118] If the determination in step S710 is "NO", the CPU 401 returns to the process in step S703 to proceed with the singing voice synthesis process.

[0119] If the determination in step S710 is "Yes", the CPU 401 ends the Figure 7 The process of singing voice synthesis is illustrated in the flowchart.

[0120] Figure 8 In the second embodiment, Figure 4 The flowchart shows an example of speech synthesis processing performed by a processor (not specifically shown) of speech synthesis LSI 405. This processing is performed by the processor executing a speech synthesis processing program stored in a memory (not specifically shown) within speech synthesis LSI 405. Alternatively, this processing may be a hybrid of hardware and software, such as a DSP (digital signal processor) or an FPGA (field programmable gate array).

[0121] The processor of the sound synthesis LSI 405 realizes, for example, by executing the sound synthesis processing program. Figure 5 The following description of each process is actually performed by the above processor, but in order to make the description easier to understand, it is assumed that Figure 5 The processing performed by each part is described.

[0122] first, Figure 5 The text analysis unit 502 is repeatedly determining whether Figure 4 The CPU 401 is in a standby state for processing the received singing voice data 415 (the determination process of step S801 is repeated as “No”).

[0123] When the singing voice data 415 is received from the CPU 401 and the determination in step S801 is "Yes", the text analysis unit 502 determines whether the range is specified by the received singing voice data 415 (see Figure 7 (Step S702) (Step S802).

[0124] If the determination result in step S802 is "Yes", the range information 509 is passed from the text analysis unit 502 to the formant interpolation processing unit 506. The subsequent operations are those of the formant interpolation processing unit 506.

[0125] The formant interpolation processing unit 506 performs the singing voice optimization process (the above is step S803). Figure 9 After the vocal optimization process at step S803, the text analysis unit 502 returns to the standby process of the vocal data 415 at step S801.

[0126] After receiving the singing voice data 415 again and the determination of step S801 becomes "yes", if the determination of step S802 in the text analysis unit 502 becomes "no", the received singing voice data 415 indicates the progress of the lyrics (see Figure 7 The text analysis unit 502 analyzes the lyrics and pitches contained in the singing data 415. As a result, the text analysis unit 502 generates a language feature sequence 507 representing the phonemes, parts of speech, and words corresponding to the lyrics in the singing data 415, and pitch information 508 corresponding to the pitches in the singing data 415, and provides these to the acoustic model unit 501.

[0127] On the other hand, through the singing optimization processing of step S803 performed before the start of singing synthesis, a request for obtaining the first sound range spectrum information 510 or a request for obtaining the first sound range spectrum information 510 and the second sound range spectrum information 511 is sent from the formant interpolation processing unit 506 to the acoustic model unit 501.

[0128] Based on the above information, the formant interpolation processing unit 506 obtains the singing voice optimization processing in step S803 described later from the acoustic model unit 501. Figure 9 Each LSP parameter of the first sound range spectrum information 510 requested from the acoustic model unit 501 in step S903 or S908 is stored in the RAM 403 (step S804).

[0129] Next, the formant interpolation processing unit 506 determines whether the interpolation flag stored in the RAM 403 has been set to "1" by the singing voice optimization processing in step S803 described later, that is, whether the interpolation processing is being executed (step S805).

[0130] If the judgment of step S805 is "No" (interpolation processing is not performed), the resonance peak interpolation processing unit 506 will set the arrangement variables of the target spectrum information 513 on RAM403 as is, using the LSP parameters of the first range spectrum information 510 obtained from the acoustic model unit 501 and stored in RAM403 in step S804 (step S806).

[0131] If the determination in step S805 is "yes" (interpolation processing is executed), the formant interpolation processing unit 506 obtains the formant interpolation processing unit 506 from the acoustic model unit 501 in the singing voice optimization processing in step S803 described later. Figure 9 The LSP parameters of the second sound range spectrum information 511 requested from the acoustic model unit 501 in step S908 are stored in the RAM 403 (step S807).

[0132] Then, the formant interpolation processing unit 506 performs formant interpolation processing (step S808). Specifically, the formant interpolation processing unit 506 performs the interpolation processing operation of the above formula (1) between the LSP parameters L1[i] of the first sound range spectrum information 510 stored in the RAM 403 in step S804 and the LSP parameters L2[i] of the second sound range spectrum information 511 stored in the RAM 403 in step S807, thereby calculating the LSP parameters L3[i] of the interpolated spectrum information and storing them in the RAM 403.

[0133] After step S808 , the formant interpolation processing unit 506 sets each LSP parameter L3 [i] of the interpolation spectrum information stored in the RAM 403 in step S808 to the array variable of the target spectrum information 513 on the RAM 403 (step S809 ).

[0134] After step S806 or S809, the formant interpolation processing unit 506 provides the target sound source information 512 output from the acoustic model unit 501 to the sound source generation unit 504 of the vocalization model unit 503. Furthermore, the formant interpolation processing unit 506 sets the LSP parameters of the target spectrum information 513 stored in RAM 403 in step S806 or S809 to the LSP digital filter of the synthesis filter unit 505 within the vocalization model unit 503 (step S810). The CPU 401 then returns to the standby processing of the singing voice data 415 in step S801, which is being executed by the text analysis unit 502.

[0135] As a result of the above processing, the vocal model unit 503 excites the LSP digital filter of the synthesis filter unit 505 provided with the above-mentioned target spectrum information 513 through the sound source input data 514 output from the sound source generation unit 504 provided with the above-mentioned target sound source information 512, thereby outputting the filter output data 515 as the singing sound output data 417.

[0136] Figure 9 Yes Figure 8 Flowchart of a detailed example of the singing voice optimization process in step S803. Figure 5 The formant interpolation processing unit 506 performs the above.

[0137] First, the formant interpolation processing unit 506 obtains information on the musical range (key range) set in the musical range information 509 transmitted from the text analysis unit 502 (step S901 ).

[0138] Next, the formant interpolation processing unit 506 determines whether the overall range of the music set in the singing data 415 obtained in step S901 (refer to the description of step S702) is within the range of the first range set as the current range in the first range variable stored in RAM403 (step S902).

[0139] In addition, in the first range variable, for example, the initial setting is Figure 1 Key field 1 (refer to Figure 7 Step S701).

[0140] If the determination in step S902 is “Yes”, the formant interpolation processing unit 506 requests the acoustic model unit 501 for spectrum information corresponding to the first range set in the first range variable (step S903 ).

[0141] Then, the formant interpolation processing unit 506 sets the value "0" indicating that the interpolation process is not to be performed to the interpolation flag variable on the RAM 403 (step S904). Figure 8 If the reference is made in step S805, the determination in step S805 becomes "No" and the interpolation process is not performed. Then, the formant interpolation processing unit 506 ends the process in step S806. Figure 9 As shown in the flowchart Figure 8 The singing optimization process of step S803 is performed.

[0142] If the range of the entire musical piece set in the vocal data 415 obtained in step S901 is not within the range of the first range, the determination in step S902 is "No", then the formant interpolation processing unit 506 determines whether there is a new range other than the first range that includes the range of the entire musical piece (for example, Figure 1 Key field 2) (step S905).

[0143] If the determination in step S905 is “Yes”, the formant interpolation processing unit 506 replaces the value of the first range variable indicating the current range on the RAM 403 with the value indicating the new range (step S906 ).

[0144] Then, the formant interpolation processing unit 506 requests the spectrum information corresponding to the first range set in the first range variable from the acoustic model unit 501 (step S903), and sets the value "0" to the interpolation flag variable on the RAM 403 (step S904). Then, the formant interpolation processing unit 506 ends the process of Figure 9 As shown in the flowchart Figure 8 The singing optimization process of step S803 is performed.

[0145] When the range of the entire music piece set in the singing data 415 obtained in step S901 is not within the range of the first range (the judgment of step S902 is "No") and there is no new range outside the first range (the judgment of step S905 is also "No"), the formant interpolation processing unit 506 determines whether the range of the entire music piece is between the current range represented by the first range variable and the other second range (step S907).

[0146] If the determination in step S907 is “Yes”, the formant interpolation processing unit 506 requests the acoustic model unit 501 to provide both spectral information corresponding to the first range set in the first range variable and spectral information corresponding to the second range determined in step S907 (step S908 ).

[0147] Then, the formant interpolation processing unit 506 sets the value "1" indicating the execution of the interpolation process to the interpolation flag variable on the RAM 403 (step S909). Figure 8 If the formant interpolation processing unit 506 is referenced in step S805, the determination in step S805 is "yes", and the interpolation processing is performed in step S808. Figure 9 As shown in the flowchart Figure 8 The singing optimization processing of step S803 is performed.

[0148] If the judgment in step S907 is "No", the formant interpolation processing unit 506 cannot determine the range. In this case, the formant interpolation processing unit 506 maintains the current range, requests the spectrum information corresponding to the first range set in the first range variable from the acoustic model unit 501 (step S903), and sets the value "0" to the interpolation flag variable on the RAM 403 (step S904). Then, the formant interpolation processing unit 506 ends the process in Figure 9 As shown in the flowchart of Figure 8 The singing optimization processing of step S803 is performed.

[0149] In the second embodiment described above, before the start of singing voice synthesis, the singing voice data 415 of the designated range is sent to the Figure 4 The sound synthesis LSI 405 sends the sound to the sound synthesis LSI 405. In the sound synthesis LSI 405, before the start of the singing voice synthesis, the formant interpolation processing unit 506 performs singing voice optimization processing based on the singing voice data 415 specifying the above-mentioned range received via the text analysis unit 502, thereby controlling the range requested to the acoustic model unit 501. In contrast, the formant interpolation processing unit 506 of the sound synthesis LSI 405 can also control the range of the singing voice based on the pitch included in the singing voice data 415 for each singing voice. Through this processing, for example, when the range of the music to be synthesized by the singing voice spans, for example, Figure 1Even in a case where the key range 1, 2, and 3 is large, an appropriate acoustic model can be selected based on the singing voice data 415 at the time of utterance and uttered using the utterance model unit 503.

[0150] In the second embodiment described above, the formant interpolation processing unit 506 performs the Figure 9 In the singing voice optimization process illustrated in the flowchart of FIG, a discrimination process ( Figure 9 Steps S902, S905 or S907, etc.). In contrast, it is also possible to Figure 4 The ROM 402 etc. are prepared in advance for each range (for example Figure 1 Tables are set for key ranges 1, 2, and 3 (for example, key ranges 1, 2, and 3) to indicate whether a single key range 1 (when the range is key range 1) is sufficient, a single key range 2 (when the range is key range 2) is sufficient, or whether interpolation processing between key ranges 1 and 2 is required (when the range is key range 3). Furthermore, the formant interpolation processing unit 506 can also perform vocal optimization processing by referring to this table. With this embodiment, even if the key range settings or interpolation settings become complex, it is possible to consistently appropriately select a range and determine whether interpolation processing is required by referring to the table of settings such as whether interpolation is required.

[0151] Furthermore, in the second embodiment described above, in the utterance model unit 503, the sound source input data 514 for exciting the synthesis filter unit 505 is Figure 5 The sound source generating unit 504 generates the target sound source information 512 from the acoustic model unit 501. In contrast, the sound source input data 514 may be generated not by the sound source generating unit 504 but by Figure 4 The sound source LSI 404 generates a part of the musical sound output data 418 for the sound source using a specific sound channel. With this structure, as the singing sound output data 417, a singing sound that retains the characteristics of the specific musical sound generated by the sound source LSI 404 can be generated interestingly.

[0152] In the second embodiment described above, the acoustic model set in the acoustic model unit 501 is learned through machine learning using learning sheet data including learning lyrics information, learning pitch information, and learning range information, and a singer's learning vocal data. However, in addition to the acoustic model obtained through machine learning, an acoustic model using a conventional phoneme database can also be used.

[0153] The second embodiment described above is an information processing device of the present invention. Figure 4 and Figure 5The embodiment shown is that the sound synthesis LSI 405 and the sound synthesis unit 500 as one of its functions are built into the control system 400 of the electronic keyboard instrument 300. Alternatively, the sound synthesis LSI and the sound synthesis unit as one of its functions (hereinafter collectively referred to as the "sound synthesis unit") and the electronic instrument may be separate devices. Figure 10 and Figure 11 These are diagrams showing a connection form between a sound synthesizing unit and an electronic keyboard instrument, and a hardware configuration example of the sound synthesizing unit according to a third embodiment in which the sound synthesizing unit and the electronic keyboard instrument operate independently.

[0154] like Figure 10 As shown, in the third embodiment, the second embodiment Figure 4 The sound synthesis LSI 405 shown and its function as a Figure 5 The sound synthesis unit 500 shown is installed in, for example, a tablet terminal or smartphone (hereinafter referred to as "tablet terminal, etc.") 1001 as dedicated hardware or software (application), and the electronic musical instrument can be configured as, for example, an electronic keyboard instrument 1002 without a sound synthesis function.

[0155] Figure 11 Yes means having Figure 10 FIG is a diagram showing an example of a hardware configuration of a tablet terminal 1001 according to a third embodiment of the present invention. Figure 11 In the example, CPU 1101, ROM 1102, RAM 1103, sound synthesis LSI 1106, D / A converter 1107 and amplifier 1108 have the same Figure 4 The output of the amplifier 1108 is connected to a speaker or earphone terminal (not shown in the figure) built into the tablet terminal 1001. The touch panel display 1104 provides the same Figure 3 The switch panels 302 and 303 have the same function.

[0156] In having Figure 10 and Figure 11In the third embodiment of the configuration example, the tablet terminal 1001 and the electronic keyboard instrument 1002 communicate wirelessly based on a standard called MIDI over Bluetooth Low Energy (hereinafter referred to as "BLE-MIDI"). BLE-MIDI is a standard for wireless communication between musical instruments that enables communication using the standard specification MIDI (Musical Instrument Digital Interface) for communication between musical instruments over the wireless standard Bluetooth Low Energy (registered trademark). The electronic keyboard instrument 1002 can communicate with the BLE-MIDI communication interface 1105 ( Figure 11 In this state, key information or key release information including pitch information specified by the electronic keyboard instrument 1002 is notified in real time to the singing voice synthesis application executed on the tablet terminal 1001 via BLE-MIDI.

[0157] In addition, a MIDI communication interface connected to the electronic keyboard instrument 1002 via a wired MIDI cable may be used instead of the BLE-MIDI communication interface 1105 .

[0158] In the third embodiment, Figure 10 The electronic keyboard instrument 1002 does not have a built-in sound synthesis LSI, while the tablet terminal 1001 has a built-in sound synthesis LSI 1106 ( Figure 11 ). And, in Figure 11 In the embodiment, the CPU 1101 of the tablet terminal 1001, for example, processes a singing synthesis application by executing the same Figure 7 The flowchart shown is the same Figure 12 The main processing illustrated in the flowchart is executed with Figure 7 The control process of the singing voice synthesis is the same as that described in the flowchart of FIG. Figure 12 In the illustrated flowchart, Figure 7 The steps with the same step numbers as those in the illustrated flowchart are executed Figure 7 The same treatment is applied to the case. Figure 12 In the illustrated flowchart, Figure 7 The flowchart shown in the example omits the Figure 4 Part of the processing of steps S706 and S709 of the sound source LSI 404.

[0159] Furthermore, the CPU 1101 monitors whether key-up information and key-down information are received from the electronic keyboard instrument 1002 via the BLE-MIDI communication interface 1105 .

[0160] If the CPU 1101 receives a keystroke from the electronic keyboard instrument 1002, it executes the same Figure 7 That is, in the case where the determination in step S1201 is "yes", CPU 1101 reads out the singing voice data of the nth lyric represented by the value of the lyric index variable n on RAM 1103 from RAM 403 ( Figure 12 Step S704).

[0161] Next, CPU 1101 indicates that Figure 12 The singing voice data 415 (refer to the singing voice data 415 of the singing voice progress) read in step S704 Figure 5 ) is built into tablet terminals, etc. 1001 Figure 11 The sound synthesis LSI1106 sends ( Figure 12 Step S705).

[0162] On the other hand, if the CPU 1101 receives the key release information from the electronic keyboard instrument 1002, it executes the same Figure 7 The same processing is performed as part of step S709. Figure 12 In the case where the determination in step S1202 is "yes", the CPU 1101 sends a signal to the built-in Figure 11 The voice synthesis LSI 1106 instructs the muting of the singing voice corresponding to the pitch of the key release included in the key release information ( Figure 12 Step S1203).

[0163] Through the above Figure 12 The control processing of steps S705 and S1203 is repeated, and the tablet terminal 1001 has a built-in Figure 11 The sound synthesis LSI 1106 performs the same operation as described above in the second embodiment. Figure 5 The sound synthesis unit 500 is also composed of Figure 8 、 Figure 9 As a result, for example, in the sound synthesis LSI 1106, singing sound output data equivalent to the singing sound output data 417 of the second embodiment is generated. This singing sound output data is output from the built-in speaker of the tablet terminal 1001, or is sent from the tablet terminal 1001 to the electronic keyboard instrument 1002 and output from the built-in speaker of the electronic keyboard instrument 1002, thereby enabling the sound to be produced in synchronization with the performance operation of the electronic keyboard instrument 1002.

[0164] Next, a fourth embodiment will be described. Figure 13This is a diagram showing a connection configuration of a fourth embodiment in which a portion of a sound synthesizing unit and an electronic keyboard instrument operate independently. Figure 14 13 is a diagram showing an example of the hardware configuration of a tablet terminal 1301 or the like corresponding to the speech synthesis unit of the fourth embodiment. Figure 15 This is a block diagram showing a configuration example of a portion of a speech synthesis LSI and a speech synthesis unit according to the fourth embodiment.

[0165] In the above Figure 5 In the second embodiment of the module structure, the sound synthesis unit 500 is a unit including Figure 4 On the other hand, in the third embodiment described above, Figure 5 The sound synthesis unit 500 is used as Figure 10 1001 built-in tablet terminals Figure 11 In the third embodiment, the tablet terminal 1001 has a built-in voice synthesis LSI 1106. Figure 11 The sound synthesis LSI 1106 has the same function as that included in the second embodiment. Figure 4 The control system 400 has the same functions as the sound synthesis LSI 405 built into the electronic keyboard instrument.

[0166] In the fourth embodiment, the electronic keyboard instrument 1302 and the tablet terminal 1301 are connected by, for example, a USB cable 1303. In this case, the control system of the electronic keyboard instrument 1302 has Figure 4 The control system 400 of the electronic keyboard instrument 300 of the second embodiment has a similar modular structure and incorporates a sound synthesis LSI 405. On the other hand, in the fourth embodiment, unlike the third embodiment, the tablet terminal 1301 may be a conventional terminal computer, rather than a built-in sound synthesis LSI. Figure 14 This shows the fourth embodiment. Figure 13 FIG is a diagram showing an example of a hardware configuration of a tablet terminal 1301. Figure 14 In the embodiment, the CPU 1401, the ROM 1402, the RAM 1403 and the touch panel display 1404 have the same Figure 11 The CPU 1101, ROM 1102, RAM 1103 and touch panel display 1104 have the same functions. USB (Universal Serial Bus) communication interface 1405 is as shown in FIG. Figure 13As shown, the USB cable 1303 connecting the tablet terminal 1301 and the electronic keyboard 1302 drives the transmission and reception of signals between the electronic keyboard 1302. Although not shown in the figure, a similar USB communication interface is also installed on the electronic keyboard 1302 side.

[0167] Furthermore, if the data capacity allows, a wireless communication interface such as Bluetooth (a registered trademark of Bluetooth SIG, Inc., U.S.) or Wi-Fi (a registered trademark of Wi-Fi Alliance, U.S.) may be used instead of the wired USB communication interface.

[0168] In the fourth embodiment Figure 15 In, with Figure 5 The modules with the same number as the modules in the block diagram have the same Figure 5 The same function as in the case of the fourth embodiment. Figure 15 The utterance model unit 503 (speech synthesis filter unit) is separated from the speech synthesis unit 1501 and built into the same structure as in the second embodiment. Figure 4 In the sound synthesis LSI 405 within the control system 400.

[0169] On the other hand, the fourth embodiment Figure 15 The acoustic model unit 501, the text analysis unit 502 and the formant interpolation processing unit 506 in the speech synthesis unit 1501 are different from the above-mentioned second embodiment. Figure 5 The functional units of the text analysis unit 502 and the formant interpolation processing unit 506 in the speech synthesis unit 500 are the same.

[0170] Specifically, these processes are for tablet terminals etc. 1301 Figure 14 The CPU 1401 executes the processing of the speech synthesis program read from the ROM 1402 to the RAM 1403. The CPU 1401 executes the speech synthesis program to perform the same operation as that in the third embodiment. Figure 12 In the fourth embodiment, the CPU 1401 executes the operations performed by the processors in the sound synthesis LSI 405 in the second embodiment. Figure 8 The flow chart of the sound synthesis process is shown as an example, and Figure 8 Details of step S803 Figure 9 The flowchart illustrates the singing voice optimization processing.

[0171] However, CPU1401, in Figure 12 In step S705, the Figure 12The singing voice data 415 indicating the progress of the singing voice read out in step S704 (see Figure 5 ), not to the sound synthesis LSI, but to the Figure 8 The sound synthesis processing illustrated in the flowchart is passed.

[0172] And, as Figure 15 As shown, CPU1401, in Figure 8 In step S810 of the sound synthesis process shown in the flowchart of Figure 8 The target spectrum information 513 generated in step S806 or S809 is generated together with the target sound source information 512 output from the acoustic model unit 501. Figure 14 The USB communication interface 1405 is connected via Figure 13 The USB cable 1303 is connected to the sound synthesis LSI 405 in the electronic keyboard instrument 1302 (see Figure 4 ) and the voice model part 503 of the action is sent.

[0173] As a result, the sound synthesis LSI 405 ( Figure 4 ) generates singing voice output data 417. The singing voice output data 417 is the same as that of the second embodiment. Figure 4 The analog singing voice output signal is converted into an analog singing voice output signal by a D / A converter 412. The analog singing voice output signal is mixed with the analog musical sound output signal in a mixer 413, and the mixed signal is amplified by an amplifier 414 and output from a speaker or output terminal (not shown).

[0174] As described above, in the fourth embodiment, the function of the sound synthesis LSI 405 of the electronic keyboard instrument 1302 and the singing voice synthesis function of the tablet terminal 1301 can be combined to produce sounds synchronized with the performance operation of the electronic keyboard instrument 1302.

[0175] Alternatively, the acoustic model unit 501 including the learned model unit may be built into an information processing device such as a tablet terminal 1301 or a server, and the generation unit for generating the third sound data, such as the formant interpolation processing unit 506, may be built into the electronic keyboard 1302. In this case, the first sound range spectrum information 510 and the second sound range spectrum information 511 are transmitted from the information processing device to the electronic keyboard 1302.

[0176] While the disclosed embodiments and their advantages have been described in detail above, a person skilled in the art can make various changes, additions, and omissions without departing from the scope of the present invention clearly described in the claims.

[0177] In addition, the present invention is not limited to the above-mentioned embodiments and can be variously modified in the implementation stage without departing from the scope of its purpose. In addition, the functions performed in the above-mentioned embodiments can be implemented in combination as much as possible. The above-mentioned embodiments include various levels, and various inventions can be extracted by appropriately combining the multiple components disclosed. For example, even if some components are deleted from all the components shown in the embodiment, as long as the effect can be obtained, the structure with the deleted components can be extracted as an invention.

Claims

1. An information processing device, characterized in that Detect the specified pitch; generating third data corresponding to the specified pitch based on first data output by the first sound model and second data output by a second sound model different from the first sound model; The first sound model corresponds to the first sound range; The second sound model corresponds to a second sound range different from the first sound range; There is a non-overlapping range between the first range and the second range; The above-specified pitches are included in the above-mentioned non-overlapping pitch ranges.

2. The information processing device according to claim 1, wherein The first voice model includes a learned model that has learned the singing voice of the first singer; The second voice model includes a learned model obtained by learning the singing voice of a second singer different from the first singer.

3. The information processing device according to claim 1 or 2, wherein: The third data is generated based on an interpolation operation between the formant frequency corresponding to the first data and the formant frequency corresponding to the second data.

4. The information processing device according to claim 1 or 2, wherein: If the acoustic model does not correspond to the musical range of the designated music, the third data is generated.

5. An electronic musical instrument, characterized in that have: The information processing device according to any one of claims 1 to 4; and Play operation element, used to specify the pitch.

6. An electronic musical instrument, characterized in that Having a performance operating member for specifying pitch; outputting pitch data corresponding to the specified pitch to an information processing device; acquiring data from the information processing device in response to the output, the data including first data corresponding to a first acoustic model obtained by learning the singing voice of a first singer and second data corresponding to a second acoustic model obtained by learning the singing voice of a second singer; synthesizing the sound based on the acquired data; The first sound model corresponds to the first sound range; The second sound model corresponds to a second sound range different from the first sound range; There is a non-overlapping range between the first range and the second range; The specified pitch is included in the non-overlapping range.

7. The electronic musical instrument according to claim 6, wherein The data further includes third data generated by the information processing device based on the first data output by the first sound model and the second data output by the second sound model; The above-mentioned sound is synthesized based on the above-mentioned third data.

8. The electronic musical instrument according to claim 6, wherein The data includes the first data output by the first sound model and the second data output by the second sound model; generating third data based on the first data and the second data; The above-mentioned sound is synthesized based on the generated third data.

9. An information processing system, characterized in that: have: The electronic musical instrument according to any one of claims 6 to 8; and The information processing device transmits data corresponding to the first sound model and the second sound model to the electronic musical instrument in response to acquisition of the pitch data transmitted from the electronic musical instrument.

10. An information processing method, characterized in that: The information processing device generates third data corresponding to the designated pitch based on the first data of the first sound model and the second data of the second sound model; The first sound model corresponds to the first sound range; The second sound model corresponds to a second sound range different from the first sound range; There is a non-overlapping range between the first range and the second range; The above-specified pitches are included in the above-mentioned non-overlapping pitch ranges.

11. The information processing method according to claim 10, wherein: The first voice model includes a learned model that has learned the singing voice of the first singer; The second voice model includes a learned model obtained by learning the singing voice of a second singer different from the first singer.

12. The information processing method according to claim 10 or 11, wherein: The information processing device generates the third data based on an interpolation operation between a formant frequency corresponding to the first data and a formant frequency corresponding to the second data.

13. The information processing method according to claim 10 or 11, wherein: If the sound model does not correspond to the musical range of the designated music, the third data is generated.

14. A storage medium, characterized in that A program is stored that enables the information processing device to implement the following functions: Detecting that a specified pitch is contained in a non-overlapping range; The third data corresponding to the above-specified pitch is generated based on the first data of the first sound model corresponding to the first sound range and the second data of the second sound model corresponding to the second sound range different from the above-mentioned first sound range, and there is the above-mentioned non-overlapping sound range between the above-mentioned first sound range and the above-mentioned second sound range.

Citation Information

Patent Citations

  • Game machine

    JP2021045284A

  • Information processing device, terminal device, information processing method and program

    JP2021117857A

  • Heater box and mobile dryness system

    JP2021190167A

  • Device and method for synthesizing musical sound

    JP1998240264A

  • Information processing method and information processing device

    JP2020076843A