Speech processing method, speech processing system, and program

The voice processing method addresses monotonous timbre in conventional voice synthesis by adding acoustic components to phoneme intervals, resulting in audio signals with diverse and natural timbres.

JP2026079175APending Publication Date: 2026-05-15YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
YAMAHA CORP
Filing Date
2024-10-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Conventional voice synthesis technologies often produce monotonous auditory impressions due to a lack of variety in timbre.

Method used

A voice processing method that generates voice signals by adding acoustic components to specific phoneme intervals, using a computer system to acquire and mix phoneme type and interval data, thereby creating audio signals with diverse timbres.

Benefits of technology

The method produces audio signals with varied and natural auditory impressions by integrating instrument sounds with voice signals, enhancing the timbre and coherence of synthesized voices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026079175000001_ABST
    Figure 2026079175000001_ABST
Patent Text Reader

Abstract

It generates audio signals with a variety of timbres. [Solution] The audio processing method is implemented by a computer system that acquires audio material data X specifying a phoneme type W and a phoneme interval for each of a plurality of phonemes arranged in time series, and generates an audio signal Z corresponding to the audio material data X. In generating the audio signal Z, for the first phoneme among the plurality of phonemes, an audio signal Z is generated in which a first acoustic component Y is added to the phoneme interval of the first phoneme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technology for synthesizing sounds such as voices.

Background Art

[0002] Voice synthesis technologies that input text information and generate voices have been proposed conventionally. For example, Patent Document 1 discloses a technology for generating smooth and natural voices.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional voice synthesis technologies focus on achieving a uniform goal of generating auditorily natural voices. Therefore, there is a problem that the auditory impression of the synthesized voice tends to be monotonous. In view of the above circumstances, one aspect of the present disclosure aims to generate voice signals with various timbres.

Means for Solving the Problems

[0005] In order to solve the above problems, a voice processing method according to one aspect of the present disclosure is a voice processing method realized by a computer system that acquires voice material data specifying a phoneme type and a phoneme section for each of a plurality of phonemes arranged in time series, and generates a voice signal according to the voice material data. In the generation of the voice signal, for a first phoneme among the plurality of phonemes, a voice signal in which a first acoustic component is added to the phoneme section of the first phoneme is generated.

[0006] To solve the above problems, an audio processing system according to one aspect of the present disclosure comprises a data acquisition unit that acquires audio material data specifying a phoneme type and a phoneme interval for each of a plurality of phonemes arranged in time series, and a signal generation unit that generates an audio signal corresponding to the audio material data, wherein the signal generation unit generates the audio signal in which a first acoustic component is added to the phoneme interval of a first phoneme among the plurality of phonemes.

[0007] To solve the above problems, a program according to one aspect of the present disclosure is a program that causes a computer system to function as a data acquisition unit that acquires speech material data specifying a phoneme type and a phoneme interval for each of a plurality of phonemes arranged in time series, and a signal generation unit that generates a speech signal corresponding to the speech material data, wherein the signal generation unit generates the speech signal in which a first acoustic component is added to the phoneme interval of a first phoneme among the plurality of phonemes. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram illustrating the configuration of the voice processing system in the first embodiment. [Figure 2] This is a schematic diagram of audio material data. [Figure 3] This is a schematic diagram of the acoustic component group. [Figure 4] This is a schematic diagram of the generation of an audio signal. [Figure 5] This block diagram illustrates the function of a control device that generates audio material data. [Figure 6] This block diagram illustrates the function of a control device that generates an audio signal. [Figure 7] This is a schematic diagram of the mixing of audio material data and acoustic components in the first embodiment. [Figure 8] This is a flowchart of the audio signal generation process in the first embodiment. [Figure 9] This is a schematic diagram of the mixing of audio material data and acoustic components in the second embodiment. [Figure 10] It is a schematic diagram of the mixing of voice material data and acoustic components in the third embodiment. [Figure 11] It is a schematic diagram of the mixing of voice material data and acoustic components in the fourth embodiment. [Figure 12] It is a flowchart of the voice signal generation process in the fourth embodiment. [Figure 13] It is a schematic diagram of the mixing of voice material data and acoustic components in the fifth embodiment. [Figure 14] It is a schematic diagram of the generation of a voice signal in the sixth execution embodiment. [Figure 15] It is a schematic diagram of the generation of a voice signal in the seventh embodiment. [Figure 16] It is a block diagram illustrating the functional configuration of the control device in the eighth embodiment. [Figure 17] It is a flowchart of the voice signal generation process in the eighth embodiment. [Figure 18] It is a block diagram illustrating the functional configuration of the control device in the ninth embodiment. [Figure 19] It is an explanatory diagram of the training process in the ninth embodiment. [Figure 20] It is a flowchart of the training process in the ninth embodiment. [Figure 21] It is a schematic diagram of the replacement of voice material data and acoustic components in Modification 1.

Embodiments for Carrying Out the Invention

[0009] A: First Embodiment FIG. 1 is a block diagram illustrating the configuration of the voice processing system 100 in the first embodiment. The voice processing system 100 generates a voice signal Z representing the waveform of a specific voice. The voice represented by the voice signal Z is, for example, the voice pronounced by a virtual speaker, or the singing voice pronounced by a virtual singer singing a piece of music, etc.

[0010] The voice processing system 100 includes an operation device 10, a control device 11, a storage device 12, and a sound output device 13. The voice processing system 100 is realized by, for example, a portable information device such as a smartphone or a tablet terminal, or a portable or stationary information device such as a personal computer. Note that the voice processing system 100 can be realized not only as a single device but also as a plurality of devices separately configured from each other.

[0011] The operation device 10 is an input device that receives an instruction from a user. The operation device 10 is, for example, an operator that the user operates or a touch panel that detects contact by the user. Note that an operation device 10 (for example, a mouse or a keyboard) separate from the voice processing system 100 may be connected to the voice processing system 100 by wire or wirelessly.

[0012] The control device 11 is composed of one or more processors that control each element of the voice processing system 100. For example, the control device 11 is composed of one or more types of processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0013] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is composed of a known recording medium such as a magnetic recording medium or a semiconductor recording medium, for example. The storage device 12 may be composed of a combination of a plurality of types of recording media. Further, a portable recording medium detachable from the voice processing system 100 or a recording medium (for example, cloud storage) on which the control device 11 can perform writing or reading via a communication network may be used as the storage device 12.

[0014] The sound-emitting device 13 reproduces sound waves under the control of the control device 11. The sound-emitting device 13 is, for example, an output device such as a speaker or headphones. Specifically, the sound-emitting device 13 reproduces the sound represented by the audio signal Z generated by the control device 11. The sound-emitting device 13, separate from the audio processing system 100, may be connected to the audio processing system 100 by wire or wireless connection.

[0015] Here, we will explain the generation of the audio signal Z. The control device 11 generates audio material data X from the synthesis data M stored in the storage device 12, and generates the audio signal Z by mixing the audio material data X with the acoustic component Y.

[0016] The synthesis data M is time-series data that specifies the content of the singing sounds produced when a virtual singer sings a song. For each of the multiple notes that make up the singing sound, the synthesis data M specifies the pitch, duration of sound production, and syllables. A syllable contains at least one phoneme. The synthesis data M is time-series data arranged in a sequence of note data that instructs the singing sound by specifying the syllable and pitch, and time data that specifies the time of reading each note data. The note data specifies, for example, the pitch and intensity of a syllable. The time data specifies, for example, the timing of reading consecutive note data. The synthesis data M is, for example, data in a format compliant with the MIDI (Musical Instrument Digital Interface) standard. The synthesis data M is stored in the storage device 12.

[0017] Audio material data X is generated from synthesis data M and stored in storage device 12. Figure 2 is a schematic diagram of audio material data X. Audio material data X includes material signal V and phoneme data A.

[0018] The source signal V is a signal that represents the waveform of the singing sound. In other words, the source signal V consists of a time series of samples that represent the sound pressure of the singing sound.

[0019] Phoneme data A specifies the pronunciation interval (hereinafter referred to as "phoneme interval") and the phoneme type W of each phoneme in the source signal V. Each phoneme interval is specified by the time when the pronunciation of each phoneme begins (i.e., the start point) and the time when the pronunciation ends (i.e., the end point). The phoneme type W of each phoneme corresponds to the syllable specified for each note by the synthesis data M. The phoneme type W is assumed to be the type of phoneme that makes up the syllable (i.e., the phonetic symbol). The phoneme types are vowels, which are harmonic components containing a fundamental component and multiple overtone components, and consonants, which are non-harmonic components. In other words, the speech source data X specifies the phoneme interval and phoneme type W for each of the multiple phonemes arranged in time series.

[0020] Acoustic component Y is a signal that represents the waveform of an instrument sound. Specifically, an instrument sound is a percussive sound with a sharp attack immediately after sound production. For example, the sounds of instruments such as bass drums, cymbals, castanets, and maracas are used as acoustic component Y. In other words, acoustic component Y is a nonharmonic component, just like consonants.

[0021] Each acoustic component Y is associated with a phoneme type W whose instrument sound is audibly similar to that of the acoustic component Y. The acoustic component Y is stored in the memory device 12. Specifically, the acoustic component Y is stored in the memory device 12 as an acoustic component group C. Figure 3 is a schematic diagram of the acoustic component group C. The acoustic component group C contains multiple acoustic components Yn (where n is a natural number). Each of the multiple acoustic components Yn represents the waveform of an instrument sound produced by different types of instruments. As mentioned above, each of the multiple acoustic components Yn is associated with a phoneme type W whose instrument sound is audibly similar to that of the acoustic component Y. That is, the acoustic component group C contains multiple acoustic components Yn corresponding to multiple phoneme types Wn. For example, among the multiple phoneme types Wn, phoneme type W1 [s] corresponds to acoustic component Y1 among the multiple acoustic components Yn.

[0022] Figure 4 is a schematic diagram of the generation of the audio signal Z. The control device 11 generates the audio signal Z by mixing the source signal V and the acoustic component Y in the audio source data X. For example, the source signal V and the acoustic component Y are added together.

[0023] The acoustic component Y is mixed into each phoneme interval specified by the phoneme data A in the source signal V. Specifically, the acoustic component Y is added to the source signal V so that the starting point of the phoneme specified by the phoneme data A coincides with the starting point of the acoustic component Y. In other words, the starting point of the phoneme specified by the phoneme data A and the starting point of the acoustic component Y1 corresponding to the phoneme type W1 of that phoneme coincide on the time axis. Therefore, compared to a form in which the starting point of the phoneme specified by the phoneme data A and the starting point of the acoustic component Y mixed into that phoneme interval do not coincide on the time axis, a more audibly natural speech signal Z can be generated.

[0024] As described above, the control device 11 generates an audio signal Z by mixing the acoustic component Y into the phoneme section of the source signal V for the phoneme specified by the phoneme data A. Therefore, compared to a configuration in which an audio signal Z is generated in which phonemes exist individually as specified by the audio source data X, it is possible to generate an audio signal Z with a variety of timbres.

[0025] Furthermore, the acoustic component Y corresponds to the phoneme type W of the phoneme represented by the material signal V being mixed. Specifically, the acoustic component Y that is audibly similar to the phoneme type W of the phoneme specified by the phoneme data A is added to the material signal V corresponding to that phoneme interval. For example, if the phoneme type W1 of the phoneme specified by the phoneme data A is [s], the acoustic component Y1 corresponding to phoneme type W1 is added to the material signal V corresponding to that phoneme interval. Therefore, compared to a configuration in which the acoustic component Y is mixed regardless of the phoneme type W of the phonemes being mixed, it is possible to generate audio signals Z with diverse timbres while maintaining the audible impression represented by the phoneme type W of the phonemes being mixed.

[0026] The audio processing system 100 generates audio material data X from synthesis data M and generates audio signal Z from audio material data X by executing a program stored in the storage device 12 using the control device 11.

[0027] Figure 5 is a block diagram illustrating the function of the control device 11 in generating audio material data X. The control device 11 implements multiple functions (control data generation unit 21, material signal generation unit 22) for generating audio material data X from synthesis data M by executing a program stored in the storage device 12. The functions of the control device 11 may be implemented by a collection of multiple devices (i.e., a system), or some or all of the functions of the control device 11 may be implemented by a dedicated electronic circuit (e.g., a signal processing circuit).

[0028] The control data generation unit 21 generates control data D from the synthesis data M stored in the storage device 12. The control data D includes pitch data B and phoneme data A. Pitch data B represents the temporal change (pitch curve) of the pitch of the singing sound represented by the synthesis data M. As described above, phoneme data A specifies each phoneme interval in the source signal V and the phoneme type W of each phoneme.

[0029] The control data generation unit 21 generates control data D by performing predetermined calculations on the synthesis data M. For example, the control data generation unit 21 generates control data D using a generative model composed of a deep neural network (DNN) or the like. The generative model is a statistical estimation model that learns the relationship between the synthesis data M and the control data D through machine learning.

[0030] The material signal generation unit 22 generates a material signal V from the control data D. As mentioned above, the material signal V is a signal that represents the waveform of the singing sound. The material signal generation unit 22 generates the material signal V by performing predetermined calculations on the control data D. For example, the material signal generation unit 22 generates the material signal V using a generation model composed of a deep neural network or the like. The generation model is a statistical estimation model that has learned the relationship between the control data D and the material signal V through machine learning.

[0031] As described above, the control device 11 (control data generation unit 21 and material signal generation unit 22) generates speech material data X, which includes phoneme data A and material signal V, from synthesis data M.

[0032] Figure 6 is a block diagram illustrating the function of the control device 11 in generating an audio signal Z. The control device 11 implements multiple functions (data acquisition unit 23 and audio signal generation unit 24) for generating an audio signal Z from audio material data X by executing a program stored in the storage device 12. The functions of the control device 11 may be implemented by a collection of multiple devices (i.e., a system), or some or all of the functions of the control device 11 may be implemented by a dedicated electronic circuit (e.g., a signal processing circuit).

[0033] The data acquisition unit 23 acquires audio material data X from the storage device 12. The data acquisition unit 23 includes a first acquisition unit 231 and a second acquisition unit 232.

[0034] The first acquisition unit 231 sequentially reads out the source signal V contained in the audio source data X unit by unit period. Each unit period is a period (frame) with a sufficiently short duration compared to each phoneme interval. Specifically, the first acquisition unit 231 reads out each sample constituting the source signal V in chronological order for each unit period.

[0035] The second acquisition unit 232 acquires the acoustic component Y corresponding to the phoneme type W of the phoneme represented by the source signal V from the acoustic component group C. The second acquisition unit 232 sequentially acquires the acoustic component Y corresponding to the phoneme type W of each phoneme interval corresponding to a consonant phoneme among the multiple phoneme intervals specified by the phoneme data A for the source signal V from the acoustic component group C. Specifically, the second acquisition unit 232 acquires the acoustic component Y corresponding to the phoneme type W of each phoneme whenever the start point of a consonant phoneme interval specified by the phoneme data A for the source signal V arrives.

[0036] The audio signal generation unit 24 generates an audio signal Z by performing a mixing process 240 that mixes the source signal V and the acoustic component Y. Specifically, the audio signal generation unit 24 generates an audio signal Z by adding the acoustic component Y corresponding to the phoneme type W of the consonant represented by the source signal V to the phoneme section of that consonant in the source signal V. Note that the audio signal generation unit 24 is just one example of a "signal generation unit".

[0037] Here, we will explain the mixing of audio source data X and acoustic components Y in the audio signal Z.

[0038] Figure 7 is a schematic diagram of the mixing of audio material data X and acoustic component Y. Figure 7 illustrates the material signal V and phoneme data A represented by the audio material data X, and the acoustic component Y acquired by the second acquisition unit 232. The synthesis data M used to generate the material signal V is also shown. Specifically, the section of the material signal V corresponding to the note specified by the synthesis data M (hereinafter referred to as "note of interest Mi") is shown.

[0039] The syllable designated for the note Mi is "su". The syllable "su" consists of the consonant [s] and the vowel [u]. Therefore, the section of the source signal V corresponding to the note Mi includes the phoneme section corresponding to the consonant [s] and the phoneme section corresponding to the vowel [u]. The source signal generation unit 22 generates the source signal V such that the starting point t2 of the vowel [u] coincides with the starting point of the note Mi on the time axis. That is, the phoneme section of the consonant [s] starts at a point earlier than the starting point of the note Mi.

[0040] Acoustic component Y is mixed only into the consonant phoneme intervals of the source signal V. Acoustic component Y does not overlap with the vowel phoneme intervals of the source signal V. As mentioned above, the starting point of the consonant phoneme interval specified by phoneme data A and the starting point of acoustic component Y corresponding to the phoneme type W of that consonant are the same time. Specifically, the starting point t1 of the consonant [s] and the starting point of acoustic component Y corresponding to the consonant [s] coincide on the time axis. On the other hand, Figure 7 illustrates the case where the time length of acoustic component Y is shorter than the time length of the phoneme interval of the consonant [s]. Therefore, the ending point of acoustic component Y is located before the starting point t2 of the vowel [u]. Thus, compared to the form in which acoustic component Y is mixed into both the consonant [s] phoneme interval and the vowel [u] phoneme interval of the source signal V, it is possible to generate audio signals Z with diverse timbres in which the consonant [s] is emphasized while maintaining the auditory impression of the vowel [u].

[0041] Figure 8 is a flowchart illustrating the specific steps of the process by which the voice processing system 100 of the first embodiment generates a voice signal Z (hereinafter referred to as the "voice signal generation process"). For example, the voice signal generation process is started in response to instructions from the user to the operating device 10.

[0042] The control device 11 (data acquisition unit 23) acquires audio material data X from the storage device 12 (Sb1). The audio material data X includes phoneme data A and material signals V. Specifically, the control device 11 (first acquisition unit 231) sequentially reads out the material signals V included in the audio material data X for each unit period.

[0043] The control device 11 determines whether the start of a phoneme interval specified by phoneme data A in the speech material data X has arrived (Sb2). If it determines that the start of a phoneme interval has arrived (Sb2:Yes), it determines whether the phoneme at which the start of the phoneme interval has arrived is a consonant (Sb3). If it determines that the phoneme at which the start of the phoneme interval has arrived is a consonant (Sb3:Yes), the control device 11 (second acquisition unit 232) acquires the acoustic component Y corresponding to the phoneme type W of the consonant from the acoustic component group C of the storage device 12 (Sb4).

[0044] The control device 11 (speech signal generation unit 24) generates a speech signal Z from the speech material data X and acoustic component Y (Sb5). Specifically, it generates the speech signal Z by performing a mixing process 240 which mixes (for example, adds) the acoustic component Y corresponding to the phoneme type W of the consonant specified by the phoneme data A with the phoneme section of the material signal V for that consonant. On the other hand, if it is determined that the start point of the phoneme section has not yet arrived (Sb2: No), or if it is determined that the phoneme for which the start point has arrived is a vowel (Sb3: No), the material signal V read from the storage device 12 in step Sb1 becomes the speech signal Z.

[0045] The control device 11 outputs the audio signal Z to the sound emission device 13 (Sb6). As a result, sound is reproduced in which the acoustic component Y is mixed with the consonants.

[0046] The control device 11 determines whether the termination condition is met (Sb7). The termination condition is, for example, that termination is instructed by the user through an operation on the operating device 10, or that processing has been completed for all of the audio material data X representing the entire synthesis data M. If it is determined that the termination condition is not met (Sb7: No), the audio signal generation process moves to step Sb1 and repeats the processing from the acquisition of the audio material data X onward (Sb1 to Sb7) for the next unit period. If it is determined that the termination condition is met (Sb7: Yes), the audio signal generation process terminates.

[0047] B: Second Embodiment A second embodiment of this disclosure will now be described. For elements whose function is the same as in the first embodiment in each of the embodiments described below, the same reference numerals as in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.

[0048] Figure 9 is a schematic diagram of the mixing of audio material data X and acoustic component Y in the second embodiment.

[0049] The control device 11 (speech signal generation unit 24) mixes the source signal V with the acoustic component Y in the generation of the speech signal Z (Sb5) in the speech signal generation process, such that the endpoint of the phoneme interval of the consonant specified by the phoneme data A coincides with the endpoint of the acoustic component Y. Specifically, the control device 11 (speech signal generation unit 24) adjusts the position of the acoustic component Y on the time axis with respect to the source signal V so that the endpoint of the acoustic component Y coincides with the endpoint of the phoneme interval of the consonant. Therefore, the endpoint t2 of the consonant [s] and the endpoint of the acoustic component Y corresponding to the consonant [s] coincide on the time axis. On the other hand, in Figure 9, it is assumed that the time length of the acoustic component Y is shorter than the time length of the phoneme interval of the consonant [s]. Therefore, the starting point of the acoustic component Y is located after the starting point t1 of the consonant [s]. Therefore, compared to a form where the endpoint t2 of the consonant [s] and the endpoint of the acoustic component Y do not coincide on the time axis, it is possible to generate an audio signal Z that gives the impression of being more cohesive audibly.

[0050] C: Third Embodiment The mixing of audio material data X and acoustic component Y in the third embodiment will be described. Figure 10 is a schematic diagram of the mixing of audio material data X and acoustic component Y in the third embodiment.

[0051] The control device 11 (speech signal generation unit 24), in the generation of the speech signal Z (Sb5) in the speech signal generation process, stretches and compresses the acoustic component Y to match the time length of the phoneme interval of the phoneme specified by the phoneme data A, and mixes the stretched and compressed acoustic component Y with the source signal V. The time length of the phoneme interval of the consonant in the source signal V is variable. Therefore, the control device 11 (speech signal generation unit 24) identifies the time length of the phoneme interval of the consonant from the phoneme data A and stretches and compresses the time length of the acoustic component Y to match that time length. Figure 10 illustrates the case where the time length of the acoustic component Y is shorter than the time length of the phoneme interval of the consonant [s]. Therefore, the control device 11 (speech signal generation unit 24) stretches the acoustic component Y so that the endpoint of acoustic component Y1 and the endpoint t2 of the consonant [s] coincide on the time axis. Therefore, in the reproduction of the speech signal Z, there is the advantage that the duration for which the consonant [s] and the acoustic component Y are pronounced are aligned.

[0052] D: Fourth Embodiment The generation of the audio signal Z in the fourth embodiment will now be described. In the first embodiment, the starting point of the phoneme interval of a consonant and the starting point of the acoustic component Y corresponding to the phoneme type W of the consonant coincided on the time axis. However, in the fourth embodiment, the configuration in which the starting point of the phoneme interval of a consonant and the starting point of the acoustic component Y corresponding to the phoneme type W of the consonant coincide is not essential.

[0053] Figure 11 is a schematic diagram of the mixing of audio material data X and acoustic component Y in the fourth embodiment. In generating the audio signal Z (Sb5), the control device 11 (audio signal generation unit 24) mixes acoustic component Y1 corresponding to the phoneme type W1 of the consonant into the phoneme section of the material signal V, similar to the first embodiment, and also mixes acoustic component Y2 corresponding to the phoneme type W2 of the vowel into the phoneme section of the material signal V. The configuration and method for mixing acoustic component Y2 into the phoneme section of the vowel is the same as the configuration and method for mixing acoustic component Y1 into the phoneme section of the consonant in the first embodiment. Therefore, the control device 11 (audio signal generation unit 24) mixes the material signal V and acoustic component Y2 such that the starting point t2 of the phoneme section of the vowel [u] and the starting point of acoustic component Y2 coincide on the time axis. As explained above, acoustic component Y1 is mixed into the phoneme interval of the consonant [s] in the source signal V, and acoustic component Y2 is mixed into the phoneme interval of the vowel [u] in the source signal V. Therefore, compared to a configuration in which acoustic component Y is not mixed into the phoneme interval of the vowel [u] in the source signal V, it is possible to generate an audio signal Z with a wider variety of timbre combinations.

[0054] Figure 12 is a flowchart illustrating the specific steps of the speech signal generation process in the fourth embodiment. In the speech processing method of the fourth embodiment, step Sb3, which determines whether the phoneme at which the start of the phoneme interval arrives is a consonant, is omitted from the processing similar to that of the first embodiment. Specifically, when the control device 11 determines that the start of the phoneme interval has arrived (Sb2: Yes), the control device 11 (second acquisition unit 232) acquires the acoustic component Y corresponding to the phoneme type W of the phoneme from the acoustic component group C of the storage device 12 (Sb4). The procedure for the speech signal generation process is the same as in the first embodiment, except that step Sb3 is omitted.

[0055] E: Fifth Embodiment The generation of the audio signal Z in the fifth embodiment will now be described. Figure 13 is a schematic diagram of the mixing of audio material data X and acoustic component Y in the fifth embodiment.

[0056] The control device 11 (speech signal generation unit 24) mixes acoustic component Y across the consonant phoneme interval and the vowel phoneme interval in the source signal V. That is, acoustic component Y is mixed so as to straddle the boundary between the consonant phoneme interval and the vowel phoneme interval. Specifically, the starting point of acoustic component Y is located within the consonant phoneme interval specified by phoneme data A, and the ending point of acoustic component Y is located within the vowel phoneme interval following that consonant. Specifically, the control device 11 (speech signal generation unit 24) adjusts the position of acoustic component Y on the time axis relative to the source signal V so that the starting point of acoustic component Y is located within the consonant phoneme interval and the ending point of acoustic component Y is located within the vowel phoneme interval. Therefore, for example, the ending point of acoustic component Y is located after the starting point t2 of the vowel [u] and before the ending point t3 of the vowel [u]. On the other hand, Figure 13 illustrates the case where the duration of acoustic component Y is shorter than the duration of the phoneme interval of the consonant [s]. Therefore, the starting point of acoustic component Y is located after the starting point t1 of the consonant [s] and before the ending point t2 of the consonant [s]. Consequently, compared to a configuration in which acoustic component Y is mixed only in the phoneme interval of the consonant [s] in the source signal V, it is possible to synthesize a more diverse timbre of audio signal Z in which acoustic component Y is also mixed in the phoneme interval of the vowel [u] in the source signal V.

[0057] F: Sixth Embodiment The generation of the audio signal Z in the sixth embodiment will now be described. The control device 11 (audio signal generation unit 24) further performs the first adjustment 241 and the second adjustment 242 in the generation of the audio signal Z (Sb5) in the audio signal generation process.

[0058] Figure 14 is a schematic diagram of the generation of the audio signal Z in the sixth embodiment. The control device 11 (audio signal generation unit 24) performs a first adjustment 241 to adjust the volume of the source signal V by multiplying the source signal V by an adjustment value (gain) α. Specifically, the control device 11 (audio signal generation unit 24) generates a source signal α·V after the first adjustment 241, in which the volume has been adjusted according to the adjustment value α.

[0059] The control device 11 (sound signal generation unit 24) performs a second adjustment 242, which adjusts the volume of the sound component Y by multiplying the sound component Y by an adjustment value β. Specifically, the control device 11 (sound signal generation unit 24) generates the sound component β·Y after the second adjustment 242, in which the volume has been adjusted according to the adjustment value β.

[0060] The control device 11 sets adjustment values ​​α and β in response to instructions from the user via the operating device 10, for example. That is, the volume balance between the source signal V and the acoustic component Y in the audio signal Z changes in response to instructions from the user.

[0061] As described above, the control device 11 (sound signal generation unit 24) generates the sound signal Z by performing a mixing process 240 that mixes the material signal α·V after the first adjustment 241 and the acoustic component β·Y after the second adjustment 242. Therefore, according to the sixth embodiment, compared to a configuration in which the first adjustment 241 and the second adjustment 242 are not performed, it is possible to adjust the balance of volume between the consonant phonemes and the acoustic component Y corresponding to those consonants in the sound signal Z.

[0062] G: Seventh Embodiment The generation of the audio signal Z in the seventh embodiment will now be described. The audio signal generation unit 24 of the seventh embodiment crossfades the source signal V and the acoustic component Y. Similar to the sixth embodiment, the audio signal generation unit 24 of the seventh embodiment performs a first adjustment 241 on the source signal V and a second adjustment 242 on the acoustic component Y.

[0063] Figure 15 is a schematic diagram of the generation of the audio signal Z in the seventh embodiment. The control device 11 (audio signal generation unit 24) crossfades the source signal V and the acoustic component Y by performing a mixing process 240 which mixes the source signal α·V after the first adjustment 241 and the acoustic component β·Y after the second adjustment 242. Specifically, the control device 11 (audio signal generation unit 24) crossfades the source signal V and the acoustic component Y in the consonant phoneme interval by adjusting the adjustment values ​​α and β.

[0064] The control device 11 (sound signal generation unit 24) increases the adjustment value α from 0 to 1 over time, for example, in the phoneme interval of a consonant. Specifically, the adjustment value α is 0 at the starting point t1 of the phoneme interval of the consonant [s], increases over time, and becomes 1 at the ending point t2 of the phoneme interval of the consonant [s]. Therefore, the volume of the source signal α·V after the first adjustment 241 is minimum at the starting point t1 of the phoneme interval of the consonant [s], increases over time, and becomes maximum at the ending point t2 of the phoneme interval of the consonant [s].

[0065] The control device 11 (speech signal generation unit 24) maintains the adjustment value α at 1 in the vowel phoneme interval of the source signal V. Specifically, the adjustment value α is maintained at 1 from the start point t2 to the end point t3 of the vowel [u] phoneme interval. Therefore, the volume of the source signal α·V after the first adjustment 241, from the start point t2 to the end point t3 of the vowel [u] phoneme interval, is constant within the vowel [u] phoneme interval.

[0066] Furthermore, the control device 11 (sound signal generation unit 24) decreases the adjustment value β from 1 to 0 over time, for example, in the consonant phoneme interval of the source signal V. Specifically, the adjustment value β is 1 at the starting point t1 of the consonant [s] phoneme interval, decreases over time, and becomes 0 at the ending point t2 of the consonant [s] phoneme interval. Therefore, the volume of the acoustic component β·Y after the second adjustment 242 is at its maximum at the starting point t1 of the consonant [s] phoneme interval and decreases over time. In Figure 15, it is assumed that the time length of the acoustic component Y is shorter than the time length of the consonant [s] phoneme interval. Therefore, the volume of the acoustic component β·Y after the second adjustment 242 is at its minimum at the ending point of the acoustic component Y.

[0067] As described above, the control device 11 (sound signal generation unit 24) performs a mixing process 240 that mixes the material signal α·V after the first adjustment 241 and the acoustic component β·Y after the second adjustment 242, thereby generating a sound signal Z that gradually switches from the acoustic component Y to the material signal V. Therefore, compared to a configuration in which the volume of the material signal V and the volume of the acoustic component Y change abruptly in the consonant phoneme section, it is possible to generate a sound signal Z in which the material signal V and the acoustic component Y in the consonant phoneme section are smoothly linked in a way that gives a naturally auditory impression.

[0068] H: Eighth Embodiment The generation of the audio signal Z in the eighth embodiment will now be described. Figure 16 is a block diagram illustrating the functional configuration of the control device 11 in the eighth embodiment. The control device 11 functions as an audio recognition unit 25 in addition to the elements similar to those in the first embodiment (data acquisition unit 23, audio signal generation unit 24) by executing a program stored in the storage device 12.

[0069] The speech recognition unit 25 performs speech recognition on the speech signal Z. The speech recognition unit 25 can optionally employ known technologies. For example, an acoustic model such as an HMM (Hidden Markov Model), a language model representing linguistic constraints, and a word dictionary containing a large number of registered words may be used. The speech recognition unit 25 may also perform speech recognition on speech recognition using an estimation model such as a deep neural network.

[0070] The speech recognition unit 25 estimates a sequence of recognized phonemes by performing speech recognition on the speech signal Z. The sequence of recognized phonemes is the result of speech recognition on the speech signal Z and is a time series of phonemes that represent the content of the speech signal Z.

[0071] The control device 11 (second acquisition unit 232) compares the recognized phoneme sequence with the time series of phonemes specified by phoneme data A (hereinafter referred to as the "target phoneme sequence"). Specifically, the control device 11 (second acquisition unit 232) compares the recognized phoneme sequence estimated by speech recognition of the speech signal Z with the target phoneme sequence and determines whether they match or not. If the recognized phoneme sequence and the target phoneme sequence match, the control device 11 outputs the speech signal Z to the sound emission device 13.

[0072] If the recognized phoneme sequence and the target phoneme sequence do not match, the control device 11 selects an acoustic component Y and generates a speech signal Z. The control device 11 (second acquisition unit 232) acquires another acoustic component Y corresponding to the phoneme type W of the consonant specified in the phoneme data A. Specifically, the control device 11 (second acquisition unit 232) acquires an acoustic component Y2 from the storage device 12 that is different from the acoustic component Y1 that was mixed into the speech-recognized speech signal Z. The control device 11 (speech signal generation unit 24) generates the speech signal Z by mixing the acoustic component Y2 with the source signal V. Once the speech signal Z is generated, the control device 11 (speech recognition unit 25) performs speech recognition on the speech signal Z again. As described above, the control device 11 iteratively changes the acoustic component Y mixed into the source signal V until the recognized phoneme sequence and the target phoneme sequence match. Therefore, compared to a configuration in which speech recognition is not performed on the speech signal Z, the content of the speech signal Z is more easily accurately grasped by the listener.

[0073] Figure 17 is a flowchart illustrating the specific steps of the speech signal generation process in the eighth embodiment. In the speech processing method of the eighth embodiment, estimation of the recognized phoneme sequence (Sb8) and determination of whether the target phoneme sequence and the recognized phoneme sequence match (Sb9) are added to the process similar to that of the first embodiment. Specifically, the control device 11 (speech recognition unit 25) estimates the recognized phoneme sequence as a result of speech recognition of the speech signal Z (Sb8). Then, the control device 11 (second acquisition unit 232) determines whether the recognized phoneme sequence and the target phoneme sequence match (Sb9).

[0074] If the recognized phoneme sequence and the target phoneme sequence do not match (Sb9: No), the speech signal generation process proceeds to step Sb4. That is, the process from the acquisition of acoustic component Y onward (Sb4 to Sb9) is repeated until the recognized phoneme sequence and the target phoneme sequence match (Sb9: Yes). However, in step Sb4, the control device 11 (second acquisition unit 232) acquires an acoustic component Y that is different from the acoustic component Y that was mixed with the speech-recognized speech signal Z.

[0075] If the recognized phoneme sequence matches the target phoneme sequence (Sb9: Yes), the control device 11 outputs the audio signal Z to the sound emission device 13 (Sb6). The procedure for generating the audio signal is the same as in the first embodiment, except that steps Sb8 and Sb9 are added.

[0076] I: Ninth Embodiment The generation of the audio signal Z in the ninth embodiment will now be described. Figure 18 is a block diagram illustrating the functional configuration of the control device 11 in the ninth embodiment.

[0077] The control device 11 (data acquisition unit 23) acquires audio material data X from the storage device 12. The control device 11 (audio signal generation unit 24) generates an audio signal Z from the audio material data X. In the ninth embodiment, a trained generation model G is used to generate the audio signal Z. The generation model G is a statistical model that has learned the relationship between audio material data X and audio signal Z through prior machine learning. The audio signal Z generated using the generation model G is an audio signal Z generated by adding acoustic component Y to the phoneme interval of consonants specified by phoneme data A in the material signal V, as exemplified in other embodiments.

[0078] The generative model G is implemented by a combination of a program that causes the control device 11 to perform an operation to generate an audio signal Z from audio material data X, and multiple variables (weights and biases) applied to the operation. The multiple variables are set by machine learning (especially deep learning) using multiple training data T and stored in the storage device 12. The control device 11 (audio signal generation unit 24) generates the audio signal Z by processing the audio material data X with the trained generative model G.

[0079] For example, a deep neural network (DNN) can be used as the generative model G. However, the configuration of the generative model G is arbitrary and not limited to the examples given above.

[0080] This section describes the machine learning process for generative model G. Figure 19 is an explanatory diagram of the process (hereinafter referred to as "training process") for establishing generative model G using machine learning. The control device 11, by executing the program stored in the storage device 12, functions not only as the elements exemplified in Figure 18 (data acquisition unit 23 and audio signal generation unit 24), but also as the training processing unit 50 shown in Figure 19. The training processing unit 50 establishes generative model G through machine learning using multiple training data T.

[0081] Multiple training data sets T are stored in the memory device 12. Each of the multiple training data sets T consists of a combination of training audio material data Xt and training audio signals Zt. Each training data set T is training data in which the training audio material data Xt and training audio signals Zt are correlated with each other. The training processing unit 50 establishes a generative model G by machine learning using the multiple training data sets T.

[0082] Figure 20 is a flowchart of the training processing unit 50. For example, the training process is started in response to instructions from the user to the operating device 10.

[0083] When the training process begins, the control device 11 (training processing unit 50) selects one of the multiple training data T stored in the memory device 12 (hereinafter referred to as "selected training data T") (Sc1). The control device 11 (training processing unit 50) iteratively updates multiple variables of the initial or provisional generative model G (hereinafter referred to as "provisional generative model Gp") using the selected training data T (Sc2~Sc4).

[0084] The control device 11 generates an audio signal Z by processing the audio material data X of the selected training data T using a provisional generative model Gp (Sc2). The control device 11 calculates a loss function that represents the error between the audio signal Z generated by the provisional generative model Gp and the audio signal Z of the selected training data T (Sc3). The control device 11 updates several variables of the provisional generative model Gp so that the loss function is reduced (ideally minimized) (Sc4). For example, backpropagation is used to update each variable according to the loss function.

[0085] The control device 11 determines whether a predetermined termination condition has been met (Sc5). The termination condition is, for example, that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not met (Sc5: No), the control device 11 selects the unselected training data T stored in the memory device 12 as the new selected training data T (Sc1). That is, the process of updating multiple variables of the provisional generative model Gp (Sc2~Sc4) is repeated until the termination condition is met (Sc5: Yes). If the termination condition is met (Sc5: Yes), the control device 11 terminates the training process. The provisional generative model Gp at the time the termination condition is met is finalized as the trained generative model G.

[0086] As can be understood from the above explanation, the generative model G learns the latent relationship between the audio source data Xt and the audio signal Zt in multiple training data sets T. Therefore, the trained generative model G outputs a statistically valid audio signal Z for unknown audio source data X under these relationships.

[0087] As described above, in the ninth embodiment, a trained generative model G is used to generate the speech signal Z. Therefore, under the latent relationship between the speech material data X and the speech signal Z in the multiple training data T used for machine learning, a variety of speech signals Z that represent a reasonable speech for an unknown speech material data X can be generated.

[0088] As can be understood from the above explanation, the process of generating an audio signal Z to which acoustic component Y is added includes both the process of generating an audio signal Z by mixing acoustic component Y with a source signal V (i.e., the first to eighth embodiments) and the process of generating an audio signal Z to which acoustic component Y is added using a generation model G (i.e., the ninth embodiment).

[0089] J: Variant The following are examples of specific modifications that may be added to each of the embodiments exemplified above. Two or more embodiments may be arbitrarily selected from the following examples and merged as appropriate, provided they do not contradict each other.

[0090] (1) In the first to eighth embodiments, a configuration was shown in which the mixing of the acoustic component Y and the material signal V is additive. However, the mixing of the acoustic component Y and the material signal V is not limited to the above examples. Any known synthesis technique can be arbitrarily employed for mixing the acoustic component Y and the material signal V. For example, multiplication of the acoustic component Y and the material signal V is also included in "mixing".

[0091] (2) In the first to eighth embodiments, the acoustic component Y is mixed into the phoneme interval of the consonant specified in the phoneme data A of the material signal V. However, the phoneme interval of the consonant specified in the phoneme data A of the material signal V may be replaced with the acoustic component Y. Figure 21 is a schematic diagram of the replacement of the audio material data X and the acoustic component Y in Modification 1.

[0092] In Figure 21, it is assumed that the duration of acoustic component Y matches the phoneme interval of the consonant [s]. Therefore, the control device 11 (speech signal generation unit 24) generates a speech signal Z in which only acoustic component Y is pronounced from the start point t1 to the end point t2 of the consonant [s] within the phoneme interval of the consonant specified in the phoneme data A of the source signal V. In other words, a speech signal Z is generated in which the consonant [s] is replaced with acoustic component Y. Thus, compared to a configuration in which the consonant [s] and acoustic component Y are pronounced in parallel, it is possible to generate a speech signal Z with a perceptually natural timbre that is easily perceived as speech produced from a single sound source.

[0093] As can be understood from the above explanation, the process of "generating an audio signal in which a first acoustic component is added to the phoneme interval of a first phoneme" in claim 1 includes not only the process of generating an audio signal Z by mixing the acoustic component Y with the source signal V, but also the process of generating an audio signal Z by replacing a part of the source signal V with the acoustic component Y.

[0094] (3) In each embodiment, a configuration in which acoustic component Y is added to the consonant phoneme intervals in the source signal V is shown. However, acoustic component Y may be added only to the vowel phoneme intervals in the source signal V, or acoustic component Y may be added to both the consonant phoneme intervals and the vowel phoneme intervals in the source signal V.

[0095] (4) In each embodiment, the acoustic component Y is a percussion sound, but the acoustic component Y is not limited to the examples given above. Examples of acoustic component Y include instrument sounds, ambient sounds, sound effects, and speech. The type of acoustic component Y may be either a recorded sound or a synthesized sound. The frequency components of the acoustic component Y may be either harmonic or non-harmonic components. For example, a configuration in which the acoustic component Y consists only of instrument sounds, or a configuration in which the acoustic component Y consists only of non-harmonic components, cannot be said to generate a variety of audio signals Z.

[0096] In a configuration where the acoustic component Y is a harmonic component, the control device 11 (speech signal generation unit 24) may generate a speech signal Z in which an acoustic component Y exhibiting the same pitch as the vowel following the consonant is added to the phoneme interval of the consonant in the source signal V. Therefore, compared to a configuration where the pitch of the acoustic component Y is unrelated to the pitch of the vowel, it is possible to generate a speech signal Z that has a more audibly natural and unified sound in terms of pitch.

[0097] Alternatively, in a configuration where the acoustic component Y is a harmonic instrument sound, the control device 11 (speech signal generation unit 24) may generate a speech signal Z by adding an acoustic component Y, which indicates a pitch corresponding to the pitch of a vowel, to the phoneme interval of the vowel in the source signal V. Therefore, compared to a configuration where the pitch of acoustic component Y is unrelated to the pitch of the vowel, it is possible to generate a speech signal Z that gives a more diverse auditory impression. Note that "corresponding to pitch" means that the pitch of the vowel and the pitch of acoustic component Y are the same, or that the pitch of the vowel and the pitch of acoustic component Y are in a consonant relationship.

[0098] (5) In each embodiment, the synthesis data M may be not only the singing sound of a virtual singer, but also the utterance of a virtual speaker, or the singing sound or utterance of a real person.

[0099] (6) In each embodiment, the sound component Y may include a watermark. A watermark is an audio signal embedded in which encrypted text information, etc., is made imperceptible to humans. The configuration in which the sound component Y includes a watermark has the advantage that the audio signal Z can be separated into the sound component Y and the source signal V by known separation and extraction processes. The sound watermark is used to prevent the unauthorized use of the audio signal Z. For example, if the watermark included in the sound component Y is configured to indicate the purpose of use, it is possible to determine, for example, whether the purpose of use of the audio signal Z is the speaker's voice or the singer's voice. Also, if the watermark included in the sound component Y is configured to indicate the copyright holder, for example, the copyright holder of the audio signal Z can be identified from the sound component Y.

[0100] (7) In the first embodiment, • Configuration 1: A configuration that generates an audio signal Z in which an acoustic component Y is added to the phoneme interval of the first phoneme among multiple phonemes, • Configuration 2: A configuration that generates an audio signal Z in which, for at least one first phoneme among multiple phonemes, an acoustic component corresponding to the phoneme type of the first phoneme is added from among multiple acoustic components of different timbres. This was illustrated as an example. Configuration 1 and Configuration 2 can exist independently of each other. Therefore, one of Configuration 1 and Configuration 2 is not essential to the other.

[0101] For example, for configuration 1, configuration (configuration 2) is not essential, in which the acoustic component Y added to the source signal V corresponds to the type of phoneme in the phoneme section of the source signal V to which the acoustic component Y is added. Also, for configuration 2, configuration (configuration 1) is not essential, in which the audio signal Z is generated by adding the acoustic component Y to the phoneme section of the phoneme specified by the phoneme data A.

[0102] (8) The notation "nth" (where n is a natural number) in this application is used solely as a formal and convenient label to distinguish each element in notation and has no substantive meaning whatsoever. Therefore, there is no room for restrictive interpretation of the position or manufacturing order of each element based on the notation "nth".

[0103] K:Additional notes From the forms exemplified above, the following configuration can be understood, for example.

[0104] A speech processing method according to one aspect of the present disclosure (Aspect 1) is a speech processing method implemented by a computer system that acquires speech material data specifying the phoneme type and phoneme interval for each of a plurality of phonemes arranged in time series, and generates a speech signal corresponding to the speech material data, wherein in the generation of the speech signal, for the first phoneme among the plurality of phonemes, a first acoustic component is added to the phoneme interval of the first phoneme to generate the speech signal. In the above aspect, a speech signal is generated in which the first acoustic component is added to the phoneme interval of the first phoneme. Therefore, compared to a configuration in which a speech signal is generated in which the first phoneme exists alone as instructed by the speech material data, it is possible to generate a speech signal with a variety of timbres. Note that "addition" means that the first acoustic component is included in the phoneme interval of the first phoneme in the speech signal. In other words, it is sufficient that the first acoustic component is included in the phoneme interval of the first phoneme in the speech signal as a result, and the processing in the process of generating the speech signal is irrelevant. The first acoustic component may be added to the phoneme interval of the first phoneme in the audio signal, or the phoneme interval may be replaced by the first acoustic component. "Phoneme type" refers to the type of phoneme (i.e., phonetic symbol) that constitutes a syllable. "Phoneme interval" refers to the pronunciation interval of a phoneme. Specifically, the pronunciation interval of a phoneme is specified by the start time (i.e., the beginning point) and end time (i.e., the ending point) of the pronunciation of that phoneme.

[0105] In a specific example of Embodiment 1 (Embodiment 2), in the generation of the speech signal, the starting point of the first phoneme and the starting point of the first acoustic component are at the same time. In this embodiment, the time when the pronunciation of the first phoneme begins and the time when the pronunciation of the first acoustic component begins coincide. Therefore, compared to a form in which the pronunciation of the first phoneme and the pronunciation of the first acoustic component begin at different times, it is possible to generate a speech signal that sounds more natural to the ear.

[0106] In a specific example of Embodiment 1 or Embodiment 2 (Embodiment 3), in the generation of the audio signal, the time length of the first acoustic component is stretched or compressed to match the time length of the first phoneme. In the above embodiments, the first phoneme and the first acoustic component are pronounced from the same point in time and end at the same point in time. Therefore, there is an advantage in that the pronunciation periods of the first phoneme and the first acoustic component are synchronized.

[0107] In any specific example of Embodiments 1 to 3 (Embodiment 4), in the audio signal, the volume of the first phoneme changes over time during a transition period that includes at least a portion of the phoneme interval of the first phoneme, and the volume of the first acoustic component changes in the opposite direction to the change in the volume of the first phoneme during the transition period. In the above embodiments, the audio signal is generated by the crossfading of the volume of the first phoneme and the volume of the first acoustic component within the transition period. Therefore, compared to a configuration in which the volume of the first phoneme and the volume of the first acoustic component change abruptly, it is possible to generate an audio signal in which the first phoneme and the first acoustic component are smoothly linked with a perceptually natural impression. Note that "crossfading" is a process in which the first phoneme and the first acoustic component are mixed while increasing the volume of the first phoneme over time and decreasing the volume of the first acoustic component over time. However, the process may also involve gradually decreasing the volume of the first phoneme and gradually increasing the volume of the first acoustic component while mixing the first phoneme and the first acoustic component.

[0108] In any specific example (5) of Embodiments 1 to 4, in generating the audio signal, a first audio signal is generated in which the first acoustic component is added to the phoneme interval of the first phoneme among the plurality of phonemes. If the phoneme sequence representing the result of speech recognition of the first audio signal does not match the phoneme sequence represented by the audio material data, a second audio signal is generated in which an acoustic component different from the first acoustic component is added. In the above embodiment, if the spoken content of the reproduced sound of the first audio signal is difficult to hear, a second audio signal is newly generated in which an acoustic component different from the first audio signal is added. Therefore, the spoken content of the reproduced sound of the second audio signal is easier for the listener to accurately grasp compared to the reproduced sound of the first audio signal. However, in generating the second audio signal, it is also possible to add only the newly selected acoustic component to the audio material data, or to add the newly selected acoustic component to the first audio signal.

[0109] In any specific example of Embodiments 1 to 5 (Embodiment 6), the first acoustic component is a musical instrument sound. In the above embodiments, since the first acoustic component is a musical instrument sound, it is possible to generate diverse and characteristic audio signals that are a mixture of human voice and musical instrument sound, compared to a configuration in which the first acoustic component is a human voice.

[0110] In any specific example of Embodiments 1 to 6 (Embodiment 7), the first acoustic component is an acoustic component corresponding to the phoneme type of the first phoneme, among a plurality of acoustic components with different timbres. In the above embodiments, the timbre of the added first acoustic component is selected according to the phoneme type of the first phoneme. Therefore, compared to a configuration in which the first acoustic component is added regardless of the phoneme type of the first phoneme, it is possible to generate audio signals with diverse timbres while maintaining the auditory impression represented by the phoneme type of the first phoneme.

[0111] In any specific example (8) of Embodiments 1 to 7, the plurality of phonemes include a first phoneme which is a consonant and a second phoneme which is a vowel which follows the first phoneme to form a syllable, and the pitch of the first sound component corresponds to the pitch of the second phoneme. In the above embodiments, the first sound component is added to the first phoneme, but the pitch of the first sound component is selected to correspond to the pitch of the second phoneme. Therefore, compared to a configuration in which the pitch of the first sound component is unrelated to the pitch of the second phoneme, it is possible to generate an audio signal that sounds natural and unified in terms of pitch. Note that "corresponding" means a musically harmonious relationship. Specifically, this means that the pitch of the second phoneme and the pitch of the first sound component are the same, or that the pitch of the second phoneme and the pitch of the first sound component are in a consonant relationship.

[0112] In any specific example (9) of Embodiments 1 to 8, the end point of the first phoneme and the end point of the first acoustic component are at the same time during the generation of the audio signal. In the above embodiments, the end point of the pronunciation of the first phoneme and the end point of the pronunciation of the first acoustic component coincide. Therefore, compared to a form in which the pronunciation of the first phoneme and the pronunciation of the first acoustic component end at different times, it is possible to generate an audio signal in which the first phoneme and the first acoustic component give the impression of being audibly unified.

[0113] In any specific example of Embodiments 1 to 9 (Embodiment 10), the plurality of phonemes include a first phoneme which is a consonant and a second phoneme which is a vowel that follows the first phoneme to form a syllable, and in the generation of the speech signal, the first acoustic component is not added to the second phoneme. In the above embodiments, the first acoustic component is added only to the first phoneme which is a consonant among the phonomes composed of a consonant and a vowel. Therefore, compared to a configuration in which the first acoustic component is added to both the first and second phonemes, it is possible to generate speech signals with diverse timbres in which the first phoneme which is a consonant is emphasized while maintaining the auditory impression of the second phoneme which is a vowel.

[0114] In a specific example of Embodiment 10 (Embodiment 11), a second acoustic component having a different timbre from the first acoustic component is added to the second phoneme. In the above embodiments, the first acoustic component is added to the first phoneme, and the second acoustic component is added to the second phoneme. Therefore, compared to a configuration in which no acoustic component is added to the second phoneme, or a configuration in which the first acoustic component is added to both the first and second phonemes, it is possible to generate audio signals with a wider variety of timbre combinations.

[0115] In any specific example of Embodiments 1 to 8 (Embodiment 12), the plurality of phonemes include a first phoneme which is a consonant and a second phoneme which is a vowel that follows the first phoneme to form a syllable, and in the generation of the speech signal, the endpoint of the first acoustic component is within the phoneme interval of the second phoneme. In the above embodiments, the duration of sound production of the first acoustic component overlaps not only with the duration of sound production of the first phoneme but also with the duration of sound production of the second phoneme. Therefore, compared to a configuration in which the range to which the first acoustic component is added is limited to the duration of sound production of the first phoneme, it is possible to synthesize speech signals with a wider variety of timbres, as the first acoustic component also overlaps with the second phoneme.

[0116] In any specific example of Embodiments 1 to 7 and Embodiment 9 (Embodiment 13), the first phoneme is a vowel, and in the speech signal, the pitch of the first phoneme corresponds to the pitch of the first sound component. In the embodiments described above, the pitch of the first sound component is selected for the pitch of the first phoneme. Therefore, compared to a configuration in which the pitch of the first sound component does not correspond to the pitch of the first phoneme, it is possible to generate speech signals that give a more diverse auditory impression. Note that "corresponding" means a musically harmonious relationship. Specifically, this means that the pitch of the second phoneme and the pitch of the first sound component are the same, or that the pitch of the second phoneme and the pitch of the first sound component are in a consonant relationship.

[0117] In any specific example of Embodiments 1 to 13 (Embodiment 14), the audio signal is generated by mixing the acoustic signal of the first phoneme with the first acoustic component. In the above embodiments, since the acoustic signal of the first phoneme and the acoustic signal of the first acoustic component are mixed, the process of generating the audio signal can be simplified compared to a configuration in which the audio signal is generated by a generation model. Note that "mixing signals" means adding or multiplying the acoustic signal of the first phoneme and the acoustic signal of the first acoustic component.

[0118] In a specific example of Embodiment 14 (Embodiment 15), in the generation of the audio signal, a first adjustment is performed to adjust the volume of the first phoneme, and a second adjustment is performed to adjust the volume of the first acoustic component. In the mixing, the acoustic signal representing the first phoneme after the first adjustment is performed and the acoustic signal representing the first acoustic component after the second adjustment is performed are mixed. In the above embodiment, the relative volume of the first phoneme and the volume of the first acoustic component are adjusted. Therefore, compared to a configuration in which the first and second adjustments are not performed, it is possible to balance the volume of the first phoneme and the first acoustic component in the audio signal.

[0119] In any specific example of Embodiments 1 to 13 (Embodiment 16), in the generation of the audio signal, an audio signal is generated in which the first phoneme is replaced by the first acoustic component. In the above embodiments, an audio signal is generated in which the first phoneme is replaced by the first acoustic component. Therefore, compared to a configuration in which the first phoneme and the first acoustic component are pronounced in parallel, it is possible to generate an audio signal with a perceptually natural timbre that is easily perceived as sound produced from a single sound source.

[0120] In any specific example of Embodiments 1 to 13 (Embodiment 17), the generation of the audio signal is performed by processing the audio material data with a machine learning-trained generative model. In the embodiments described above, the audio signal is generated by processing the audio material data with a machine learning-trained generative model. Therefore, under the latent relationship between the audio material data and the audio signal in the multiple training datasets used for machine learning, a variety of audio signals that represent a valid voice for unknown audio material data can be generated.

[0121] A speech processing system according to one aspect of the present disclosure (Aspect 18) is a speech processing method implemented by a computer system that acquires speech material data specifying phoneme intervals of a plurality of phonemes and generates a speech signal corresponding to the speech material data, wherein in generating the speech signal, for at least one first phoneme among the plurality of phonemes, a speech signal is generated in which an acoustic component corresponding to the phoneme type of the first phoneme is added from among a plurality of acoustic components of different timbres. In the above aspect, the timbre of the added first acoustic component is selected according to the phoneme type of the first phoneme. Therefore, compared to a configuration in which the first acoustic component is added regardless of the phoneme type of the first phoneme, it is possible to generate speech signals of diverse timbres while maintaining the auditory impression represented by the first phoneme.

[0122] A speech processing system according to one aspect of the present disclosure (Aspect 19) comprises a data acquisition unit that acquires speech material data specifying a phoneme type and a phoneme interval for each of a plurality of phonemes arranged in time series, and a signal generation unit that generates a speech signal corresponding to the speech material data, wherein the signal generation unit generates a speech signal in which a first acoustic component is added to the phoneme interval of a first phoneme among the plurality of phonemes. In the above aspect, a speech signal is generated in which a first acoustic component is added to the phoneme interval of a first phoneme. Therefore, compared to a configuration in which a speech signal is generated in which the first phoneme exists alone as instructed by the speech material data, it is possible to generate a speech signal with a variety of timbres.

[0123] A speech processing program according to one aspect of the present disclosure (Aspect 20) is a program that causes a computer system to function as a data acquisition unit that acquires speech material data specifying the phoneme type and phoneme interval for each of a plurality of phonemes arranged in time series, and a signal generation unit that generates a speech signal corresponding to the speech material data, wherein the signal generation unit generates a speech signal in which a first acoustic component is added to the phoneme interval of a first phoneme among the plurality of phonemes. In the above aspect, a speech signal is generated in which a first acoustic component is added to the phoneme interval of a first phoneme. Therefore, compared to a configuration in which a speech signal is generated in which the first phoneme exists alone as instructed by the speech material data, it is possible to generate a speech signal with a variety of timbres. [Explanation of Symbols]

[0124] 10...Operating device, 11...Control device, 12...Storage device, 13...Sound emission device, 21...Control data generation unit, 22...Material signal generation unit, 23...Data acquisition unit, 24...Speech signal generation unit, 25...Speech recognition unit, 50...Training processing unit, 100...Speech processing system, 231...First acquisition unit, 232...Second acquisition unit, 240...Mixing processing, 241...First adjustment, 242...Second adjustment, A...Phoneme data, B...Pitch data T, C...Acoustic component group, D...Control data, G...Generative model, Gp...Provisional generative model, M...Synthesis data, Mi...Note of interest, T...Training data, t1...Starting point of consonant [s], t2...Ending point of consonant [s] and starting point of vowel [u], t3...Ending point of vowel [u], V...Source signal, W...Phoneme type, X...Speech source data, Y...Acoustic component, Z...Speech signal, α...Adjustment value for the first adjustment, β...Adjustment value for the second adjustment.

Claims

1. For each of the multiple phonemes arranged in chronological order, we obtain audio material data that specifies the phoneme type and phoneme interval. Generate an audio signal corresponding to the aforementioned audio material data. A speech processing method implemented by a computer system, In generating the aforementioned audio signal, For the first phoneme among the plurality of phonemes, the audio signal is generated in which the first acoustic component is added to the phoneme interval of the first phoneme. Audio processing methods.

2. The plurality of phonemes include a first phoneme which is a consonant and a second phoneme which is a vowel which follows the first phoneme to form a syllable. In generating the aforementioned audio signal, The first acoustic component is not added to the second phoneme. The audio processing method according to claim 1.

3. The second phoneme is given a second acoustic component that has a different timbre from the first acoustic component. The audio processing method according to claim 2.

4. In generating the aforementioned audio signal, For the first phoneme among the plurality of phonemes, a first sound signal is generated in which the first acoustic component is added to the phoneme interval of the first phoneme. If the sequence of phonemes representing the result of speech recognition of the first audio signal does not match the sequence of phonemes represented by the audio material data, a second audio signal is generated in which a different acoustic component from the first acoustic component is added. The audio processing method according to claim 1.

5. In the aforementioned audio signal, The volume of the first phoneme changes over time during a transition period that includes at least a portion of the phoneme interval of the first phoneme. The volume of the first acoustic component changes in the opposite direction to the change in volume of the first phoneme during the transition period. The audio processing method according to claim 1.

6. In generating the aforementioned audio signal, The acoustic signal of the first phoneme and the first acoustic component are mixed. The audio processing method according to claim 1.

7. In generating the aforementioned audio signal, Perform a first adjustment to adjust the volume of the first phoneme, A second adjustment is performed to adjust the volume of the first sound component. In the aforementioned mixing, The acoustic signal representing the first phoneme after the first adjustment is performed is mixed with the acoustic signal representing the first acoustic component after the second adjustment is performed. The audio processing method according to claim 6.

8. In generating the aforementioned audio signal, The first phoneme is replaced by the first acoustic component to generate an audio signal. The audio processing method according to claim 1.

9. In generating the aforementioned audio signal, The audio signal is generated by processing the audio material data using a machine learning-based generative model. The audio processing method according to claim 1.

10. The first acoustic component is an acoustic component among a plurality of acoustic components of different timbres that corresponds to the phoneme type of the first phoneme. The audio processing method according to claim 1.

11. A data acquisition unit that acquires audio material data specifying the phoneme type and phoneme interval for each of multiple phonemes arranged in chronological order, A signal generation unit that generates an audio signal corresponding to the aforementioned audio material data. A voice processing system comprising, The signal generation unit, For the first phoneme among the plurality of phonemes, the audio signal is generated in which the first acoustic component is added to the phoneme interval of the first phoneme. Voice processing system.

12. A data acquisition unit that acquires audio material data specifying the phoneme type and phoneme interval for each of multiple phonemes arranged in chronological order, and Signal generation unit that generates an audio signal corresponding to the aforementioned audio material data It is a program that makes a computer system function as follows: The signal generation unit, For the first phoneme among the plurality of phonemes, the audio signal is generated in which the first acoustic component is added to the phoneme interval of the first phoneme. program.