Audio signal conversion device and program
The audio signal conversion device enhances realism by segmenting speech into phonemes and moras, adjusting acoustic characteristics for each phoneme and speaker, ensuring high-quality sound propagation from a recording to a target position.
Patent Information
- Application Number
- JP2021187109
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2041-11-17
AI Technical Summary
Existing audio systems fail to adequately simulate the radiation characteristics of human voices, particularly for each phoneme or speaker, leading to a lack of realism in content spaces.
An audio signal conversion device that segments speech into phonemes and moras, compensates for the first acoustic characteristics using acoustic data specific to each phoneme, and adds second acoustic characteristics based on target positions, adjusting for speaker attributes to enhance realism.
The device improves the sense of realism by ensuring consistent and high-quality sound propagation from a recording position to a target position, accounting for individual speaker characteristics.
Smart Images

Figure 0007805137000006 
Figure 0007805137000007 
Figure 0007805137000008
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio signal conversion device and program, for example, a technique for adjusting a human voice emitted from an object depending on the user's position. [Background technology]
[0002] In recent years, the practical application of audio systems such as object-based audio and AR / VR (Augmented Reality / Virtual Reality) audio, which play content that combines audio signals and audio metadata, has been progressing (Non-Patent Documents 1-5). Object-based audio and AR / VR are characterized by receiving audio signals and associated audio metadata, rendering the audio signals according to the user's position in the content space, and playing back sound based on the rendered audio signals. To further enhance the sense of realism of these contents, it has been proposed to add acoustic characteristics (radiation characteristics) for each radiation direction to the sound waves emitted from the object that serves as the sound source and then play them back. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Rec. ITU-R BS.2076-1 “Audio Definition Model” (2017) [Non-patent document 2] Rec. ITU-R BS.2125-0 “A serial representation of the Audio Definition Model” (2019) [Non-patent document 3] ISO / IEC 23008-3:2019 “Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio, Second edition” (2019) [Non-patent document 4] ETSI TS 103 190-2 V1.2.1 “AC-4 Part 2” (2018) [Non-patent document 5] ATSC Standard: A / 342 Part 3(2017) Summary of the Invention [Problem to be solved by the invention]
[0004] However, the radiation characteristics of a human voice (sometimes referred to as "human voice" in this application) can differ for each phoneme or speaker. Therefore, while radiation characteristics given using a simple model such as a hard sphere can ensure real-time performance, they cannot adequately simulate the radiation characteristics of an actual human voice.
[0005] The present invention has been made to solve the above-mentioned problems, and one object of the present invention is to provide an audio signal conversion device and program that can improve the sense of realism of human voices in a content space. [Means for solving the problem]
[0006] [1] One aspect of the present invention is an acoustic characteristic generating unit that identifies, for each mora that segments a speech pronunciation, acoustic characteristic data that indicates the sound radiation characteristics of each phoneme that constitutes the mora, and determines, from the acoustic characteristic data, a first acoustic characteristic that corresponds to a recording position based on a sound source position and a second acoustic characteristic that corresponds to a target position based on the sound source position; and an acoustic characteristic conversion filtering unit that compensates for the first acoustic characteristic of a first phoneme-based speech signal that corresponds to the phoneme, and adds the second acoustic characteristic to generate a second phoneme-based speech signal, The acoustic feature generation unit classifies the acoustic feature data for each speaker attribute, and determines the first acoustic feature and the second acoustic feature using the acoustic feature data corresponding to the input speaker attribute.It is an audio signal conversion device. According to the configuration of [1], the first acoustic characteristics corresponding to the recording position of the first phoneme-specific speech signal for each phoneme constituting each mora are converted into the second acoustic characteristics corresponding to the target position. Since the acoustic characteristics are adjusted so that output speech having the same quality as the sound propagating from the person to the target position for each mora is obtained, it is possible to improve the sense of realism without impairing the acoustic characteristics of the speech in units of mora. Furthermore, since the first acoustic characteristic and the second acoustic characteristic are determined using acoustic characteristic data corresponding to the input speaker attributes, an output voice having a sound quality according to the speaker can be obtained.
[0008] [2] One aspect of the present invention is the above-mentioned audio signal conversion device, comprising: an audio signal division unit that divides an input audio signal into first phoneme-specific audio signals of phoneme-specific periods based on a period during which each phoneme is pronounced; and an acoustic characteristic conversion filter generation unit that determines conversion filter coefficients by dividing the second acoustic characteristic in the frequency domain for the phoneme-specific periods by the first acoustic characteristic, wherein the acoustic characteristic conversion filtering unit may convert the frequency characteristic of the first phoneme-specific audio signal in the phoneme-specific periods into first conversion coefficients that indicate the frequency characteristic of the first phoneme-specific audio signal, multiply the first conversion coefficients by the conversion filter coefficients to calculate second conversion coefficients, convert the second conversion coefficients into a second periodic function in the time domain, and extract one period of the signal from the second periodic function as the second phoneme-specific audio signal. [2] According to this configuration, the first acoustic characteristic corresponding to the sound collection position is compensated for in each phoneme period for the first phoneme-specific sound signal, and the second acoustic characteristic corresponding to the target position is added, thereby obtaining sound with stable sound quality for each phoneme.
[0009] [3] One aspect of the present invention is the above-mentioned audio signal conversion device, which includes an audio signal connection unit, wherein the audio signal division unit divides a period including the pronounced period and a period of a predetermined length adjacent to the pronounced period into a first phoneme-specific audio signal as the phoneme-specific period, and the audio signal connection unit may mix the second phoneme-specific audio signal of the current phoneme with the second phoneme-specific audio signal of the previous phoneme in an overlapping period in which the second phoneme-specific audio signal of the current phoneme overlaps with the second phoneme-specific audio signal of the previous phoneme so that the ratio of the second phoneme-specific audio signal of the current phoneme increases over time. [3] According to the configuration of (1), during an overlap period in which the second phoneme-specific audio signals overlap between adjacent phonemes, the ratio of the second phoneme-specific audio signal of the current phoneme becomes relatively larger as time passes compared to the second phoneme-specific audio signal of the previous phoneme. This eliminates or mitigates abrupt changes in acoustic characteristics between phonemes, thereby reducing the sense of incongruity between phonemes.
[0010] [4] One aspect of the present invention is the above-mentioned speech signal conversion device, which may include a morpheme analysis unit that divides a character string into morphemes, and a phoneme division unit that determines a phoneme string consisting of phonemes that indicate the pronunciation of the morphemes, and a mora string consisting of moras that segment the pronunciation. [4] According to the configuration of (1), in addition to the phoneme sequence that forms the pronunciation of the input character string, the mora sequence that segments the pronunciation is obtained. Therefore, the acoustic characteristics according to the recording position can be converted into the acoustic characteristics according to the target position, taking into account the transition of the acoustic characteristics between the phonemes in each mora.
[0011] [5] One aspect of the present invention is to provide a computer that executes the following: [4] The present invention may also be a program for causing the audio signal conversion device to function as any one of the audio signal conversion devices described above. [5] According to the configuration, the first acoustic characteristic corresponding to the recording position of the first phoneme-specific audio signal for each phoneme constituting a mora is converted into the second acoustic characteristic corresponding to the target position. Therefore, an output audio having the same quality as the sound propagating from a person to the target position for each mora can be obtained. This can improve the sense of realism of the speech for each mora. Furthermore, since the first acoustic characteristic and the second acoustic characteristic are determined using acoustic characteristic data corresponding to the input speaker attributes, an output voice having a sound quality according to the speaker can be obtained. [Effects of the Invention]
[0012] According to the present invention, it is possible to improve the sense of realism of human voices in a content space. [Brief explanation of the drawings]
[0013] [Figure 1]1 is an explanatory diagram for explaining an overview of a speech processing system according to an embodiment of the present invention; [Figure 2] 1 is a schematic block diagram illustrating an example of the functional configuration of an audio signal conversion device according to an embodiment of the present invention. [Figure 3] 10 is a table showing an example of input data for audio signal conversion processing according to the present embodiment. [Figure 4] FIG. 2 is a schematic block diagram illustrating an example of the configuration of an acoustic characteristic generation unit according to the present embodiment. [Figure 5] FIG. 10 is a diagram showing a first example of a correspondence relationship between moras and acoustic characteristic data. [Figure 6] FIG. 10 is a diagram showing a second example of the correspondence between moras and acoustic characteristic data. [Figure 7] FIG. 10 is a diagram showing a third example of the correspondence between moras and acoustic characteristic data. [Figure 8] FIG. 10 is a diagram showing a fourth example of the correspondence between moras and acoustic characteristic data. [Figure 9] FIG. 10 is a diagram showing a fifth example of the correspondence between moras and acoustic characteristic data. [Figure 10] 10 is a flowchart illustrating an example of an audio signal conversion process according to the present embodiment. [Figure 11] FIG. 10 is a schematic block diagram illustrating an example of a functional configuration of an audio signal conversion device according to a modified example of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] (overview) Hereinafter, an embodiment of the present invention will be described with reference to the drawings. First, an outline of the present embodiment will be described. Fig. 1 is an explanatory diagram for explaining the outline of a speech processing system 1 according to the present embodiment. The audio processing system 1 includes an audio signal conversion device 10 and a microphone 20. The audio signal conversion device 10 compensates for a first acoustic characteristic and adds a second acoustic characteristic to an audio signal (recorded audio signal) input from the microphone 20. The audio signal conversion device 10 outputs a sound reproduced according to the audio signal (converted audio signal) including components of the second acoustic characteristic to a speaker for presenting to a user Usr. The first acoustic characteristic corresponds to the propagation characteristic of audio from an object Obj serving as a sound source installed at a predetermined position (sometimes referred to as the "sound source position" in this application) to the microphone 20 installed at a recording position (recording coordinates (r1, φ1, θ1)). r1, φ1, and θ1 respectively represent the distance from the origin of the recording position, the azimuth angle, and the elevation angle. The second acoustic characteristic corresponds to the propagation characteristic of audio from the object Obj installed at the sound source position to a target position (target coordinates (r2, φ2, θ2)). r2, φ2, and θ2 respectively indicate the distance from the origin of the target position, the azimuth angle, and the elevation angle. The target position corresponds to the position where the user Usr aims to listen to the sound. The sound radiation characteristics (directional characteristics) that make up the sound propagation characteristics can differ depending on the phonemes contained in each mora.
[0015] Note that a mora refers to a segmental unit of speech having a certain duration. A mora is also called a beat. One mora generally corresponds to a sound represented by one Japanese kana character (including long vowels, glottal stops, and double consonants). However, one yo-on (written as one discarded kana and the kana character immediately preceding it) also corresponds to one mora. Therefore, a mora does not uniquely correspond to a syllable. In this embodiment, a mora may refer to one phoneme belonging to one mora or a phoneme string consisting of multiple phonemes. Furthermore, the following explanation mainly uses Japanese as an example.
[0016] The voice signal conversion device 10 receives text data representing a character string in parallel with a voice signal. The voice signal conversion device 10 performs morphological analysis on the character string to segment it into individual words, and then divides the segmented words into a mora string and a phoneme string representing the pronunciation of the segmented words. The voice signal conversion device 10 segments the recorded voice signal into first phoneme-specific voice signals for each phoneme corresponding to each mora, compensates for first acoustic characteristics of the first phoneme-specific voice signals, and adds second acoustic characteristics to generate second phoneme-specific voice signals. The voice signal conversion device 10 outputs a converted voice signal obtained by connecting the generated second phoneme-specific voice signals between phonemes. In other words, the voice signal conversion device 10 reflects different human voice radiation characteristics for each phoneme constituting each mora in the recorded voice signal according to the positional relationship between the object Obj and the user Usr, thereby simulating a converted voice signal collected at a target position. Reproducing the playback sound heard by a user actually located at the recording position improves the sense of realism.
[0017] The audio signal conversion device 10 is useful for allowing individual users to experience content by assuming that they are located at any position in the content space. The audio signal conversion device 10 can be applied to, for example, the playback and production of AR / VR content. Note that the following description mainly focuses on an example in which the sound source position is located at the origin O of a three-dimensional coordinate system. The origin O may be the center of the head of a person, which is the object Obj, or the position of the lips.
[0018] Next, an example of the functional configuration of the audio signal conversion device 10 according to this embodiment will be described. Fig. 2 is a schematic block diagram showing an example of the functional configuration of the audio signal conversion device 10 according to this embodiment. An input audio signal and text data representing a character string are input to the audio signal conversion device 10 in association with each other, and an output audio signal is output. The input audio signal and the output audio signal correspond, for example, to the recorded audio signal and converted audio signal shown in FIG. 1. The input audio signal represents a voice spoken by a person at a recording location. The input audio signal may be an audio signal recorded in real time on the spot, or may be an audio data file in a predetermined format (e.g., WAV format) input from another device. The character string is a sequence consisting of one or more characters representing the spoken content of the voice related to the input audio signal. The text data may be a data file independent of the input audio signal, or may be metadata accompanying the input audio signal. The input audio signal and the text file may each have synchronization information (e.g., time code) for each predetermined section (segment, e.g., one sentence) and may be associated with each other for each section.
[0019] The speech signal conversion device 10 includes a morphological analysis unit 112, a phoneme division unit 114, a speech signal division unit 116, an acoustic characteristic generation unit 122, an acoustic characteristic conversion filter generation unit 124, an acoustic characteristic conversion filtering unit 126, and a speech signal connection unit 128.
[0020] The morphological analysis unit 112 performs morphological analysis on the character strings shown in the text data input thereto using a known morphological analysis method, dividing the string into one or more morphemes to obtain a morpheme string. A morpheme is the smallest unit of expression that has meaning in a language. In Japanese, a word corresponds to a morpheme. Words are broadly divided into independent words and adjunct words. An independent word is a morpheme that appears alone in a sentence. An adjunct word is a morpheme that is not used alone but appears attached to other morphemes. Depending on the language, a morpheme may be a unit smaller than a word. The morphological analysis unit 112 outputs the obtained morpheme string to the phoneme division unit 114.
[0021] The phoneme division unit 114 determines a phoneme string including one or more phonemes indicating the pronunciation of each morpheme constituting the morpheme string input from the morpheme analysis unit 112, and a mora string including one or more moras obtained by segmenting the pronunciation. The phoneme division unit 114 pre-stores morpheme data (morpheme lexicon) indicating the phoneme string indicating the pronunciation of each morpheme and the mora string obtained by segmenting the pronunciation. The phoneme division unit 114 can identify the phoneme string and mora string for each morpheme by referring to the morpheme data. In Japanese, one mora is generally associated with one or two phonemes. The phoneme division unit 114 adds information about the corresponding mora to each phoneme constituting the phoneme string. The phoneme division unit 114 outputs phoneme string information indicating the determined phoneme string to the speech signal division unit 116 and the acoustic characteristics generation unit 122. The phoneme sequence information may be represented as data in which phoneme labels indicating individual phonemes are arranged in the order of the phonemes.
[0022] The phoneme division unit 114 determines, for each phoneme constituting a predetermined phoneme sequence in the input speech signal input thereto, a period during which the phoneme is pronounced as a phoneme-specific period. The phoneme division unit 114, for example, calculates acoustic features (e.g., Mel-cepstrum) for each frame of a predetermined length (e.g., 10 to 50 ms) from the input speech signal, and calculates likelihood for each phoneme candidate calculated using a known acoustic model. The acoustic model may be a mathematical model that indicates the correspondence between acoustic features and phonemes, such as a Hidden Markov Model (HMM). The phoneme division unit 114 can determine, as the phoneme-specific period for the determined phoneme, the period during which the likelihood is highest for the determined phoneme. For each phoneme, the phoneme division unit 114 determines, as a division time, the time at which the phoneme is separated from the next phoneme to be pronounced, and outputs the determined division time to the speech signal division unit 116. The determined division time is the end of the phoneme interval of the phoneme of interest and the start of the phoneme interval of the next phoneme to be uttered.
[0023] The audio signal division unit 116 divides the input audio signal input thereto into first phoneme-specific audio signals for each phoneme by dividing the input audio signal into phoneme-specific periods of the phonemes indicated in the phoneme sequence input thereto from the phoneme division unit 114, using the division times input thereto from the phoneme division unit 114. The audio signal division unit 116 may label the corresponding first phoneme-specific audio signals with phoneme labels indicating the phonemes. The audio signal division unit 116 outputs the divided first phoneme-specific audio signals to the acoustic characteristic conversion filtering unit 126. The phoneme labels may be described as metadata attached to the first phoneme-specific audio signals, or may be described as file names of data files in which the first phoneme-specific audio signals are stored.
[0024] The acoustic characteristic generation unit 122 stores acoustic characteristic data in advance for each phoneme that makes up a mora. Even for phonemes that have the same sound value, acoustic characteristic data corresponding to phonemes that belong to different moras is set separately. For example, acoustic characteristic data for the phoneme / a / that makes up the mora "ka" is stored separately from acoustic characteristic data for the phoneme / a / that makes up the mora "sa." The acoustic characteristic data is data that indicates the radiation characteristics of sound from a sound source. The radiation characteristics include directional characteristics for each frequency, that is, information indicating the intensity for each radiation direction from the sound source.
[0025] The acoustic feature generation unit 122 identifies acoustic feature data corresponding to the phonemes that make up each mora, for each mora that makes up the mora string that is input from the phoneme division unit 114. Using the identified acoustic feature data, the acoustic feature generation unit 122 determines, for each frequency, the acoustic feature that corresponds to the recording coordinates that are input to the unit itself, as the first acoustic feature. The acoustic characteristic generation unit 122 determines the distance attenuation amount according to the distance from the sound source to the recording position using a predetermined distance attenuation model. As the distance attenuation model, for example, the inverse square law of distance can be used. The acoustic characteristic generation unit 122 can calculate the first acoustic characteristic by multiplying the distance attenuation amount determined for each frequency by the intensity in the recording direction.
[0026] The acoustic characteristic generation unit 122 can determine, as the second acoustic characteristic, an acoustic characteristic corresponding to the target coordinates input to the unit for each frequency, using a method similar to that for the first acoustic characteristic. That is, the acoustic characteristic generation unit 122 determines the strength in the target direction for each frequency using the identified acoustic characteristic data, and determines the distance attenuation amount for each frequency using a predetermined distance attenuation model. The acoustic characteristic generation unit 122 can calculate the second acoustic characteristic by multiplying the strength in the target direction by the distance attenuation amount. The first acoustic characteristic and the second acoustic characteristic can each be given as a complex number for each frequency. The acoustic characteristic generation unit 122 outputs the first acoustic characteristic and the second acoustic characteristic determined for each phoneme to the acoustic characteristic conversion filter generation unit 124. An example configuration of the acoustic characteristic generation unit 122 will be described later.
[0027] The acoustic characteristic conversion filter generation unit 124 determines, for each phoneme, a conversion filter coefficient for converting the first acoustic characteristic corresponding to the recording position input from the acoustic characteristic generation unit 122 into a second acoustic characteristic corresponding to the target position. For example, the acoustic characteristic conversion filter generation unit 124 calculates a quotient obtained by dividing the second acoustic characteristic by the first acoustic characteristic for each frequency as a conversion filter coefficient in the frequency domain. The acoustic characteristic conversion filter generation unit 124 outputs the calculated conversion filter coefficient to the acoustic characteristic conversion filtering unit 126.
[0028] The acoustic characteristic conversion filtering unit 126 performs filtering on the first phoneme-specific audio signal input from the audio signal division unit 116 using the conversion filter coefficients for that phoneme input from the audio signal division unit 116 to generate a second phoneme-specific audio signal. In the filtering process, the acoustic characteristic conversion filtering unit 126, for example, performs a Fourier transform on the first phoneme-specific audio signal to calculate first conversion coefficients in the frequency domain, and multiplies the first conversion coefficients calculated for each frequency by the conversion filter coefficients to calculate second conversion coefficients. The acoustic characteristic conversion filtering unit 126 performs an inverse Fourier transform on the calculated second conversion coefficients to obtain a second phoneme-specific audio signal in the time domain. Through the filtering process, the first phoneme-specific audio signal is compensated with the first acoustic characteristic, and the second acoustic characteristic is added, thereby converting components of the first acoustic characteristic into components of the second acoustic characteristic. The acoustic characteristic conversion filtering unit 126 outputs the generated second phoneme-specific audio signal to the audio signal connection unit 128. An example of an algorithm for the acoustic characteristic conversion process will be described later.
[0029] The acoustic characteristic conversion filter generation unit 124 may perform an inverse Fourier transform on the conversion filter coefficients in the frequency domain to calculate conversion filter coefficients in the time domain and output them to the acoustic characteristic conversion filtering unit 126. In this case, the acoustic characteristic conversion filtering unit 126 performs a convolution operation as a filtering process using the conversion filter coefficients input from the acoustic characteristic conversion filter generation unit 124.
[0030] The audio signal connection unit 128 connects the second phoneme-based audio signals for each phoneme input from the acoustic characteristic conversion filtering unit 126 between phonemes and outputs them as an output audio signal. The audio signal connection unit 128, for example, temporarily stores (buffers) the input second phoneme-based audio signals and starts outputting the second phoneme-based audio signal for the current phoneme immediately after outputting the second phoneme-based audio signal for the previous phoneme, which is the immediately preceding phoneme, in the order in which the second phoneme-based audio signals are input.
[0031] Next, an example of input data for the speech signal conversion processing according to this embodiment will be described. FIG. 3 is a table showing an example of input data for the speech signal conversion processing according to this embodiment. In this example, it is assumed that an input speech signal representing speech conveying "It's cold this morning" and text data representing the speech content as a character string are input. The character string representing the speech content is divided into a morpheme string "It's cold this morning" by the morpheme analysis unit 112. The morpheme string is decomposed into a mora string "Kesawasamui" as the entire speech content and a phoneme string " / k / / e / / s / / a / / w / / a / / s / / a / / m / / u / / i / " by the phoneme division unit 114. The mora "ke" is associated with the phoneme string / k / / e / . Each of the other moras is associated with two or one phoneme. The recording coordinates and target coordinates are determined for each phoneme by the acoustic feature generation unit 122. In the example shown in FIG. 3, the recording coordinates and target coordinates are given as polar coordinates with the origin at the center of the head of the person, which is the object Obj. The recording position where the input voice signal is recorded is fixed at a distance of 1 m from the front of the speaker. Here, it is assumed that the voice is recorded in a stable environment such as a studio. On the other hand, the target coordinates change for each mora. It is assumed that the target position where the user is located moves freely within the virtual space. The acoustic characteristics generation unit 122 acquires the recording coordinates and the target coordinates using phonemes or moras as processing units, and the voice conversion process is performed for each processing unit, thereby mitigating fluctuations in sound quality within the processing unit and achieving stabilization. This improves the sound quality of the pronunciation for each mora.
[0032] Next, the algorithm for the audio signal conversion process will be described. The audio signal conversion device 10 according to this embodiment receives an input audio signal s [r1] (t) is input. [r1] is a three-dimensional vector [r1, φ1, θ1] indicating the recording coordinates. [...] indicates the vector.... t indicates time. The origin O of the coordinate system is taken to be, for example, the position of the center of a person's lips. r1, φ1, and θ1 indicate the distance from the origin, the azimuth angle, and the elevation angle, respectively. The audio signal division unit 116 divides the input audio signal s [r1](t) is divided into the first phoneme speech signal for each phoneme. m The first phoneme-specific speech signal for pm,[r1] (t)(t m ≦t≦t m+1 ) is expressed as t m is the phoneme p m Therefore, the time t m From time t m+1 The period until phoneme p m This corresponds to the phoneme period. The acoustic characteristic conversion filtering unit 126 converts the first phoneme-based speech signal s pm,[r1] Phoneme p in (t) m A periodic function s with a period of each phoneme period pm,[r1] '(t) is Fourier transformed to obtain the first transformation coefficient S pm,[r1] (ω) is calculated. The first conversion coefficient S pm,[r1] (ω) is expressed by equation (1).
[0033]
number
[0034] The acoustic characteristic conversion filtering unit 126 calculates the first conversion coefficient S in the frequency domain as shown in equation (2). pm,[r1] (ω) has a second acoustic characteristic D(p m ,r2,φ2,θ2,ω) to obtain the first acoustic characteristic D(p m ,r1,φ1,θ1,ω) to obtain the second conversion coefficient S pm,[r2] (ω) can be calculated. m ,r2,φ2,θ2,ω) is a function of the target coordinate [r2] (=[r2, φ2, θ2]).
[0035]
number
[0036] Here, the first acoustic characteristic D(p m ,r1,φ1,θ1,ω) is the direction-dependent component D(p m,φ1,θ1,ω) and the distance-dependent component R(r1), which can be expressed as the product of these. Similarly, the second acoustic characteristic D(p m ,r2,φ2,θ2,ω) is the direction-dependent component D(p m ,φ2,θ2,ω) and the distance-dependent component R(r2). In this case, equation (2) can be rewritten as equation (3).
[0037]
number
[0038] It is practically acceptable to assume that the distance-dependent components R(r1) and R(r2) have distance attenuation characteristics based on the inverse square law of distance. In this case, equation (3) can be approximated as equation (4).
[0039]
number
[0040] However, as the target position approaches the person, the distance r2 becomes smaller, so the second conversion coefficient S pm,[r2] Therefore, a predetermined upper limit is set in the acoustic characteristic conversion filtering unit 126, and the second conversion coefficient S pm,[r2] For example, if the acquired distance r2 is less than a predetermined lower limit (for example, 10 to 50 cm), the acoustic characteristic conversion filtering unit 126 sets the distance r2 to the lower limit and adjusts the second conversion coefficient S pm,[r2] (ω) may be calculated. This results in the second conversion coefficient S pm,[r2] Overflow of (ω) is avoided.
[0041] The acoustic characteristic conversion filtering unit 126 calculates the second conversion coefficient S as shown in equation (5). pm,[r2] (ω) is inverse Fourier transformed to obtain the phoneme p of the second phoneme-specific speech signal. m A periodic function s with a period of each phoneme period pm,[r2] '(t) can be calculated.
[0042]
number
[0043] The acoustic characteristic conversion filtering unit 126 uses a periodic function s pm,[r2] '(t) for one period is used as the second phoneme speech signal s pm,[r2] It can be extracted as (t). The speech signal connection unit 128 connects the second phoneme-specific speech signals s obtained for each phoneme. pm,[r2] (t) are arranged in the order of phonemes on the time axis to generate the output speech signal s [r2] (t) where the output audio signal s [r2] (t) is [s p1,[r2] (t) s p2,[r2] (t) …].
[0044] Next, an example of the configuration of the acoustic characteristic generation unit 122 according to this embodiment will be described. FIG. 4 is a schematic block diagram showing an example of the configuration of the acoustic characteristic generation unit 122 according to this embodiment. The acoustic characteristic generation unit includes an acoustic characteristic data storage unit 122a and an acoustic characteristic calculation unit 122b. In the example of FIG. 4, the acoustic characteristic generation unit 122 receives input of speaker attributes in addition to sound collection coordinates and target coordinates. As will be described later, the input speaker attributes are used to determine first acoustic characteristics corresponding to the sound collection coordinates and second acoustic characteristics corresponding to the target coordinates. The speaker attributes are expressed using element information such as gender, age, and body type. Information that affects the sound radiation characteristics of the head and its surroundings can be used as the element information.
[0045] The acoustic characteristic data storage unit 122a stores acoustic characteristic data (acoustic characteristic data by mora / phoneme and radiation direction) classified by samples (samples) of predetermined speaker attributes. A plurality of predetermined types of element information or combinations thereof are used as speaker attributes for classification. Speaker attributes AC in FIG. 4 indicate classifications of individual speaker attributes. For gender, which is element information of speaker attributes, male and female can be used as samples or their elements. For age, for example, a representative value for each pre-classified age group can be used as a sample or its elements. For the correspondence between age groups and representative values, for example, a 15-year-old group can be used as a representative value for 13 to 17 years old (juveniles), a 24-year-old group can be used as a representative value for 18 to 30 years old (young adults), and a 45-year-old group can be used as a representative value for 31 to 59 years old (prime adults). For body type, for example, a representative value for each pre-classified height range can be used as a sample or its elements. As a correspondence relationship between height ranges and representative values, for example, a set of 155cm for a representative value of 150cm to 160cm, 165cm for a representative value of 160cm to 170cm, and 175cm for a representative value of 170cm to 180cm can be used as samples or elements thereof.
[0046] The acoustic characteristic data corresponding to each speaker attribute sample is, for example, data indicating the directional characteristics of sounds uttered for each frequency of each phoneme belonging to a mora. The directional characteristics are expressed as relative complex intensities for a plurality of predetermined discrete radiation directions (discrete radiation directions). The relative complex intensities correspond to the ratios of the complex intensities in each discrete radiation direction to the complex intensity in a predetermined radiation direction (e.g., φ=0°, θ=0°) as a reference intensity. The complex intensities are derived from the audio signals collected when the phonemes are uttered for a typical sample. For example, SOFA (Spatially Oriented Format for Acoustics) can be used as the data format of the acoustic characteristic data.
[0047] The acoustic characteristic data storage unit 122a stores sound collection coordinates, target coordinates, and speaker attributes. The acoustic characteristic data storage unit 122a evaluates the input speaker attribute samples and identifies acoustic characteristic data corresponding to the evaluated samples. The identified acoustic characteristic data is read and output to the acoustic characteristic calculation unit 122b. The sound collection coordinates, target coordinates, and speaker attributes can be set in response to operation signals input from an operation input unit (not shown) at different times. The operation input unit can be an input device that receives a user operation and outputs an operation signal corresponding to the received operation, such as a mouse, a touch sensor, a keyboard, a dial, or a knob. If speaker attributes are not input, the acoustic characteristic data storage unit 122a may store acoustic characteristic data corresponding to predetermined speaker attributes as default values, or may store acoustic characteristic data last used before startup.
[0048] Each time sound collection coordinates are input, the acoustic characteristic calculation unit 122b updates the sound collection coordinates used to calculate the first acoustic characteristic to the newly input sound collection coordinates. Each time target coordinates are input, the acoustic characteristic calculation unit 122b updates the target coordinates used to calculate the second acoustic characteristic to the newly input target coordinates. If the sound collection coordinates and the target coordinates are equal, the acoustic characteristic calculation unit 122b does not convert the acoustic characteristics of the input audio signal. In this case, the acoustic characteristic calculation unit 122b outputs a gain, which is a constant real value independent of frequency, as a conversion filter coefficient to the acoustic characteristic conversion filtering unit 126. The acoustic characteristic conversion filtering unit 126 simply multiplies the first phoneme-specific audio signal by the constant gain to generate a second phoneme-specific audio signal, and the acoustic characteristics of the first phoneme-specific audio signal are not substantially converted.
[0049] If the acoustic property data contains a relative complex intensity (sometimes simply referred to as "intensity" in this application) corresponding to the same radiation direction as the recording direction, which is the direction indicated by the recording coordinates, that intensity is determined as the intensity in the recording direction. The recording direction is represented by the azimuth and elevation angles of the recording coordinates. If the acoustic property data does not contain an intensity corresponding to the same radiation direction as the recording direction indicated by the recording coordinates, the intensity in the recording direction can be determined by interpolating the intensities corresponding to multiple radiation directions (three or more in three-dimensional space) surrounding the recording direction. Any of the interpolation methods can be used, such as linear interpolation, spline interpolation, and interpolation using spherical harmonic expansion.
[0050] The acoustic characteristic calculation unit 122b determines the distance attenuation amount according to the distance from the sound source to the recording position using a predetermined distance attenuation model (for example, the inverse square law of distance). The distance to the recording position is expressed as the distance in recording coordinates. The acoustic characteristic calculation unit 122b multiplies the distance attenuation amount determined for each frequency by the intensity in the recording direction to calculate the first acoustic characteristic.
[0051] The acoustic characteristic calculation unit 122b can calculate the second acoustic characteristic for the target coordinates using the same method as for the first acoustic characteristic for the recording coordinates. That is, the acoustic characteristic calculation unit 122b calculates, for each frequency, the distance attenuation based on the intensity in the target direction and the distance to the target position, and multiplies the calculated distance attenuation by the intensity in the target direction to calculate the second acoustic characteristic. The acoustic characteristic calculation unit 122b outputs the calculated first acoustic characteristic and second acoustic characteristic to the acoustic characteristic conversion filter generation unit .
[0052] 4, the acoustic characteristic data storage unit 122a is exemplified as a case where speaker attributes are input, but input of speaker attributes may be omitted. In that case, the acoustic characteristic data does not need to be classified according to speaker attributes.
[0053] Next, examples of the types of correspondence between moras and acoustic characteristic data in Japanese will be given. 5 shows a first example of the correspondence between moras and acoustic characteristic data. In the example shown, mora I consists of one phoneme p. Phoneme p is a vowel. Phoneme p is associated with acoustic characteristic data Ip. Figure 6 shows a second example of the correspondence between mora and acoustic characteristic data. I consists of two phonemes, q and s, which are a consonant and a vowel, respectively. The phonemes q and s are associated with acoustic characteristic data IIq and IIs, respectively.
[0054] When deriving the directional characteristics for the first phoneme q belonging to mora II consisting of two phonemes, the speech signal of an initial portion within a predetermined period from the beginning of phoneme q may be adopted, and the speech signal of the remaining subsequent portion may be discarded.When deriving the directional characteristics for the subsequent phoneme s belonging to mora II, the speech signal of an initial portion within a predetermined period from the beginning of phoneme s may be discarded, and the speech signal of the remaining subsequent portion may be adopted.Since the portion in which a transient change in acoustic features related to phoneme transitions in mora II appears appears, it is possible to derive significant phoneme-specific radiation characteristics using the stable portion of the acoustic features.
[0055] The moras illustrated in Figures 5 and 6 occur independently during speech. In contrast, the mora III illustrated in Figure 7 is a long sound. A long sound is formed by a continuation of the phoneme p that forms the vowel of the immediately preceding mora I or the phoneme s that forms the vowel of the mora II. Therefore, for the phoneme p or s of mora III, the acoustic feature data storage unit 122a may use the same acoustic feature data as the acoustic feature data IIIp1 or IIIs1 corresponding to the phoneme p of the immediately preceding mora I or the vowel s of the mora II.
[0056] Mora IV shown in FIG. 8 is a geminate consonant. A geminate consonant is a sound corresponding to the kana character "っ." A geminate consonant follows a vowel and is preceded by a phoneme q that is the same as the phoneme q that forms the immediately following consonant. Therefore, the acoustic feature data storage unit 122a may use, for the phoneme q of mora IV, acoustic feature data that is the same as the acoustic feature data IVq2 that corresponds to the phoneme q that forms the immediately following consonant of mora II.
[0057] The phoneme t of mora V illustrated in FIG. 9 is a nasal sound. A nasal sound corresponds to the kana character "n." A nasal sound is a phoneme t that follows the immediately preceding vowel phoneme p1 or s1 and forms a predetermined consonant. The phoneme t is one of / n / , / m / , and / ng / depending on the immediately following phoneme q2 that forms the consonant of mora II. Therefore, the acoustic characteristic data storage unit 122a may store information indicating the correspondence between the phoneme q2 and the phoneme t, and acoustic characteristic data corresponding to one of / n / , / m / , and / ng / may be used for the phoneme t.
[0058] Next, an example of the audio signal conversion process according to this embodiment will be described below with reference to a flowchart of FIG. (Step S102) The voice signal conversion device 10 acquires the input voice signal and text data of a character string indicating the speech content in association with each other. (Step S104) The morphological analysis unit 112 performs morphological analysis on the character string indicated in the text data, and breaks it down into a string of morphemes. (Step S106) The phoneme division unit 114 refers to the morpheme data and determines, for each morpheme, a phoneme string and a mora string that indicate the pronunciation of the morpheme.
[0059] (Step S108) The phoneme division unit 114 obtains phoneme labels indicating the phonemes in the order of the phonemes constituting the phoneme string, and outputs the phoneme labels to the speech signal division unit 116 and the acoustic feature generation unit 122. (Step S110) The phoneme division unit 114 uses a predetermined mathematical model to determine division times that divide the period in which each phoneme is pronounced, for each phoneme that makes up the phoneme sequence determined for the input speech signal. The speech signal division unit 116 divides the input speech signal into first phoneme-based speech signals using the determined division times.
[0060] (Step S112) The acoustic feature generation unit 122 acquires a speaker attribute, determines a sample of the speaker attribute corresponding to the acquired speaker attribute, and identifies acoustic feature data corresponding to the determined sample. (Step S114) The acoustic characteristic generating unit 122 acquires recording coordinates. (Step S116) The acoustic characteristic generation unit 122 acquires the target coordinates. (Step S118) The acoustic characteristic generation unit 122 determines whether the recording coordinates are identical to the target coordinates. If the acoustic characteristic generation unit 122 determines that the recording coordinates are identical to the target coordinates (step S118 YES), it sets a conversion filter exhibiting a constant gain in the acoustic characteristic conversion filtering unit 126, and essentially stops the acoustic characteristic conversion process. Thereafter, the process proceeds to step S136. If the acoustic characteristic generation unit 122 determines that the recording coordinates are not identical to the target coordinates (step S118 NO), the process proceeds to step S120.
[0061] (Step S120) The acoustic characteristic generation unit 122 determines whether or not the identified acoustic characteristic data stores an acoustic characteristic corresponding to a radiation direction equal to the recording direction forming the recording coordinates. If the acoustic characteristic generation unit 122 determines that the acoustic characteristic is stored (YES in step S120), the acoustic characteristic generation unit 122 proceeds to processing in step S122. If the acoustic characteristic generation unit 122 determines that the acoustic characteristic is not stored (NO in step S120), the acoustic characteristic generation unit 122 proceeds to processing in step S124. (Step S122) The acoustic characteristic generation unit 122 selects an acoustic characteristic corresponding to the recording direction from the identified acoustic characteristic data. The acoustic characteristic generation unit 122 calculates a distance attenuation amount according to the distance to the recording coordinates using a predetermined distance attenuation model. The acoustic characteristic generation unit 122 calculates a first acoustic characteristic by multiplying the selected acoustic characteristic by the calculated distance attenuation amount. (Step S124) The acoustic characteristic generation unit 122 interpolates acoustic characteristics for each of a plurality of radiation directions sandwiching the recording direction from the identified acoustic characteristic data, and calculates acoustic characteristics for the recording direction. The acoustic characteristic generation unit 122 calculates a distance attenuation amount according to the distance to the recording coordinates using a predetermined distance attenuation model. The acoustic characteristic generation unit 122 calculates a first acoustic characteristic by multiplying the selected acoustic characteristic by the calculated distance attenuation amount.
[0062] (Step S126) The acoustic characteristic generation unit 122 determines whether or not the identified acoustic characteristic data stores an acoustic characteristic corresponding to an emission direction equal to the target direction forming the target coordinates. If the acoustic characteristic generation unit 122 determines that the acoustic characteristic is stored (YES in step S126), the acoustic characteristic generation unit 122 proceeds to processing in step S128. If the acoustic characteristic generation unit 122 determines that the acoustic characteristic is not stored (NO in step S126), the acoustic characteristic generation unit 122 proceeds to processing in step S130. (Step S128) The acoustic characteristic generation unit 122 selects an acoustic characteristic corresponding to the target direction from the identified acoustic characteristic data. The acoustic characteristic generation unit 122 calculates a distance attenuation amount according to the distance to the target coordinates using a predetermined distance attenuation model. The acoustic characteristic generation unit 122 calculates a second acoustic characteristic by multiplying the selected acoustic characteristic by the calculated distance attenuation amount. (Step S130) The acoustic characteristic generation unit 122 interpolates acoustic characteristics for each of a plurality of radiation directions sandwiching the target direction from the identified acoustic characteristic data, and calculates acoustic characteristics for the target direction. The acoustic characteristic generation unit 122 calculates a distance attenuation amount according to the distance to the target coordinates using a predetermined distance attenuation model. The acoustic characteristic generation unit 122 calculates a second acoustic characteristic by multiplying the selected acoustic characteristic by the calculated distance attenuation amount.
[0063] (Step S132) The acoustic characteristic conversion filter generation unit 124 determines conversion filter coefficients that compensate for the first acoustic characteristic and add the second acoustic characteristic. (Step S134) The acoustic characteristic conversion filtering unit 126 filters the first phoneme-based audio signal using the determined conversion filter coefficients to generate a second phoneme-based audio signal. (Step S136) The voice signal connection unit 128 buffers the second phoneme-specific voice signals generated for each phoneme in that order. (Step S138) The audio signal connection unit 128 connects the buffered second phoneme-based audio signals in the order of the phonemes. (Step S140) The audio signal connection unit 128 outputs the connected second phoneme-specific audio signals.
[0064] (Variation) Next, a modified example of this embodiment will be described. The audio signal conversion device 10 may include at least an acoustic characteristic generation unit 122, an acoustic characteristic conversion filter generation unit 124, an acoustic characteristic conversion filtering unit 126, and an audio signal connection unit 128. In the speech signal conversion device 10 according to the modified example shown in FIG. 11, the morphological analysis unit 112, the phoneme division unit 114, and the speech signal division unit 116 shown in FIG. 2 are omitted. The first phoneme-specific speech signals are input to the acoustic characteristic conversion filtering unit 126 in the order of the phonemes that appear in a series of utterances. Information about the phonemes and the moras to which the phonemes belong is input to the acoustic characteristic generation unit 122 in the order of the phonemes. The functional configurations of the acoustic characteristic generation unit 122, the acoustic characteristic conversion filter generation unit 124, the acoustic characteristic conversion filtering unit 126, and the speech signal connection unit 128 will be described with reference to the description of the speech signal conversion device 10 illustrated in FIG. 2.
[0065] The speech signal conversion device 10 illustrated in FIG. 11 may also include a speech recognition unit (not shown). The speech recognition unit performs speech recognition processing on an input speech signal input thereto, analyzes a mora sequence and a phoneme sequence that constitute the pronunciation of the speech represented by the input speech signal, and decomposes the input speech signal into a first phoneme-specific speech signal for each phoneme that constitutes the phoneme sequence. The speech recognition unit calculates acoustic features for each frame of the input speech signal and can determine the uttered phonemes and their phoneme-specific periods based on the acoustic features calculated using a known acoustic model (e.g., HMM). The speech recognition unit can identify morphemes that constitute the utterance content based on the obtained phoneme sequence using a known language model (e.g., n-gram) and determine the mora sequence using the above-mentioned method. The speech recognition unit outputs the obtained first phoneme-specific speech signal to the acoustic feature conversion filtering unit 126. The speech recognition unit associates the mora sequence with the obtained phoneme sequence and outputs the resulting signal to the acoustic feature generation unit 122.
[0066] The acoustic characteristic conversion filtering unit 126 may calculate the first transform coefficient by performing a short-time Fourier transform on the second phoneme-specific audio signal for each frame of a predetermined time length (for example, 10 to 50 ms) that is sufficiently shorter than the phoneme period. In this case, the transform filter coefficient by which the first transform coefficient is multiplied, and therefore the frequency resolution of the first acoustic characteristic and the second acoustic characteristic, may be determined in advance based on the time length corresponding to one frame.
[0067] The audio signal division unit 116 may define, as a phoneme-specific period, a period including a period in which each phoneme is spoken and a margin period of a predetermined time length (e.g., 5 to 50 ms) immediately before and / or after the period. The audio signal division unit 116 may divide, as a first phoneme-specific audio signal, a portion of the input audio signal included in the defined phoneme-specific period. Since the second phoneme-specific audio signal is obtained by filtering the time length of the first phoneme-specific audio signal as described above, an overlapping period occurs in which the second phoneme-specific audio signal corresponding to the current phoneme at that time overlaps with the second phoneme-specific audio signal corresponding to the previous phoneme immediately before the current phoneme. The audio signal connection unit 128 may mix (crossfade) the second phoneme-specific audio signal of the current phoneme and the second phoneme-specific audio signal of the previous phoneme during the overlapping period so that the intensity ratio of the second phoneme-specific audio signal of the current phoneme becomes relatively large over time. This reduces the time change in acoustic features due to phoneme switching, thereby eliminating or reducing the sense of incongruity caused by a significant change in acoustic features. Although the above description mainly uses Japanese speech as an example, the present invention is not limited to this. This embodiment can also be applied to languages in which pronunciation can be segmented into moras, or speech in which moras can be conceived as speech units. Examples of such languages include Finnish, Hawaiian, Italian, and Spanish. Additionally, the microphone 20 does not necessarily have to be a separate body from the audio signal conversion device 10, but may be configured integrally with the audio signal conversion device 10.
[0068] As described above, the audio signal conversion device 10 (FIGS. 2 and 11) according to this embodiment is equipped with an acoustic characteristic generation unit 122 (FIGS. 2, 4 and 11) that identifies, for each mora that segments the pronunciation of audio, acoustic characteristic data that indicates the audio radiation characteristics of each phoneme that constitutes the mora, and determines, from the identified acoustic characteristic data, a first acoustic characteristic that corresponds to a recording position based on the sound source position and a second acoustic characteristic that corresponds to a target position based on the sound source position, and an acoustic characteristic conversion filtering unit 126 (FIGS. 2 and 11) that compensates for the first acoustic characteristic of the first phoneme-specific audio signal that corresponds to the phoneme, and adds the second acoustic characteristic to generate a second phoneme-specific audio signal. With this configuration, the first acoustic characteristic corresponding to the recording position of the first phoneme-specific speech signal is converted into the second acoustic characteristic corresponding to the target position for each phoneme constituting a mora. Since the acoustic characteristics are adjusted so that output speech having the same quality as the sound propagating from the person to the target position for each mora is obtained, it is possible to improve the sense of realism without impairing the acoustic characteristics of speech in units of mora.
[0069] Furthermore, the acoustic feature generation unit 122 (FIG. 4) may classify acoustic feature data for each speaker attribute, and determine the first acoustic feature and the second acoustic feature using the acoustic feature data corresponding to the input speaker attribute. With this configuration, the first acoustic characteristic and the second acoustic characteristic are determined using the acoustic characteristic data corresponding to the input speaker attribute, so that an output voice having a sound quality according to the speaker can be obtained.
[0070] The audio signal conversion device 10 (FIG. 2) may further include an audio signal division unit 116 (FIG. 2) that divides an input audio signal into first phoneme-specific audio signals for phoneme-specific periods based on the periods during which each phoneme is pronounced, and an acoustic characteristic conversion filter generation unit 124 (FIG. 2) that determines conversion filter coefficients by dividing second acoustic characteristics in the frequency domain for the phoneme-specific periods by the first acoustic characteristics. The acoustic characteristic conversion filtering unit 126 (FIG. 2) may convert the frequency characteristics of the first phoneme-specific audio signal for the phoneme-specific periods into first conversion coefficients that indicate the frequency characteristics of the first phoneme-specific audio signal, multiply the first conversion coefficients by the conversion filter coefficients to calculate second conversion coefficients, convert the second conversion coefficients into a second periodic function in the time domain, and extract one period of the signal from the second periodic function as the second phoneme-specific audio signal. With this configuration, the first acoustic characteristic corresponding to the sound collection position is compensated for in each phoneme period to the first phoneme-specific sound signal, and the second acoustic characteristic corresponding to the target position is added, thereby obtaining sound with stable sound quality for each phoneme.
[0071] The speech signal conversion device 10 (FIG. 2) may also include a speech signal connection unit 128 (FIG. 2). The speech signal division unit 116 (FIG. 2) may divide a period including a period in which each phoneme is pronounced and a period of a predetermined time length adjacent to that period into first phoneme-specific speech signals as phoneme-specific periods. The speech signal connection unit 128 may mix the second phoneme-specific speech signal of the current phoneme and the second phoneme-specific speech signal of the previous phoneme in an overlapping period in which the second phoneme-specific speech signal of the current phoneme and the second phoneme-specific speech signal of the previous phoneme overlap so that a ratio of the second phoneme-specific speech signal of the current phoneme increases over time. With this configuration, during the overlap period when the second phoneme-specific audio signals overlap between adjacent phonemes, the ratio of the second phoneme-specific audio signal of the current phoneme becomes relatively larger compared to the second phoneme-specific audio signal of the previous phoneme over time. This eliminates or mitigates sudden changes in acoustic characteristics between phonemes, thereby reducing the sense of incongruity between phonemes.
[0072] The speech signal conversion device 10 (FIG. 2) may also include a morpheme analysis unit 112 (FIG. 2) that divides a character string into morphemes, and a phoneme division unit 114 (FIG. 2) that determines a phoneme string consisting of phonemes that indicate the pronunciation of the morpheme and a mora string consisting of moras that segment the pronunciation. This configuration not only provides a phoneme sequence that forms the pronunciation of the input character string, but also a mora sequence that segments the pronunciation. Therefore, it is possible to convert the acoustic characteristics corresponding to the recording position into the acoustic characteristics corresponding to the target position, taking into account the transition of acoustic characteristics between phonemes in each mora.
[0073] Note that parts of the speech signal conversion device 10 described above, such as the morphological analysis unit 112, the phoneme division unit 114, the speech signal division unit 116, the acoustic characteristics generation unit 122, the acoustic characteristics conversion filter generation unit 124, the acoustic characteristics conversion filtering unit 126, and the speech signal connection unit 128, or any combination thereof, may be implemented by a computer. In this case, a program for implementing each control function may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed. Note that the term "computer system" used here refers to a computer system built into the speech signal conversion device 10, including an operating system (OS) and hardware such as peripheral devices. The term "computer-readable recording medium" refers to portable media such as a flexible disk, a magneto-optical disk, a read-only memory (ROM), and a CD-ROM, as well as storage devices such as a hard disk built into a computer system. Furthermore, the term "computer-readable recording medium" may include a medium that dynamically stores a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, or a medium that stores a program for a certain period of time, such as a volatile memory within a computer system that serves as a server or client in such a case. The program may also be one that realizes part of the above-mentioned functions, or one that can realize the above-mentioned functions in combination with a program already stored in the computer system.
[0074] Furthermore, part or all of the voice signal conversion device 10 in the above-described embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the voice signal conversion device 10 may be individually implemented as a processor, or part or all of the blocks may be integrated into a processor. The integrated circuit implementation method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit implementation technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used.
[0075] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like are possible within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0076] 1...Speech processing system, 10...Speech signal conversion device, 112...Morphological analysis unit, 114...Phoneme division unit, 116...Speech signal division unit, 122...Acoustic characteristic generation unit, 124...Acoustic characteristic conversion filter generation unit, 126...Acoustic characteristic conversion filtering unit, 128...Speech signal connection unit, 20...Microphone
Claims
1. For each mora that segments the pronunciation of the speech, acoustic characteristic data indicating the sound emission characteristics of each phoneme that constitutes the mora are identified; an acoustic characteristic generating unit that determines, from the acoustic characteristic data, a first acoustic characteristic corresponding to a recording position based on a sound source position and a second acoustic characteristic corresponding to a target position based on the sound source position; an acoustic characteristic conversion filtering unit that compensates for the first acoustic characteristic of a first phoneme-based speech signal corresponding to the phoneme and adds the second acoustic characteristic to generate a second phoneme-based speech signal, The acoustic characteristic generation unit The acoustic characteristic data is classified for each speaker attribute, The first acoustic characteristic and the second acoustic characteristic are determined using the acoustic characteristic data corresponding to the input speaker attribute. Audio signal conversion device.
2. a voice signal dividing unit that divides an input voice signal into first phoneme-specific voice signals of phoneme-specific periods based on a period during which each phoneme is pronounced; an acoustic characteristic conversion filter generation unit that determines a conversion filter coefficient by dividing the second acoustic characteristic in the frequency domain in the phoneme-specific period by the first acoustic characteristic; The acoustic characteristic conversion filtering unit converting the first phoneme-specific speech signal into a first conversion coefficient indicating a frequency characteristic of the first phoneme-specific speech signal in the phoneme-specific period; multiplying the first transform coefficients by the transform filter coefficients to calculate second transform coefficients; The second transform coefficients are transformed into a second periodic function in the time domain, and one period of the signal is extracted from the second periodic function as the second phoneme-specific speech signal.
2. The audio signal conversion device according to claim 1.
3. an audio signal connection; The audio signal dividing unit Dividing a period including the pronounced period and a period of a predetermined time length adjacent to the pronounced period into a first phoneme-based speech signal as the phoneme-based period; The audio signal connection unit During an overlapping period in which the second phoneme-specific voice signal of the current phoneme and the second phoneme-specific voice signal of the previous phoneme overlap, the second phoneme-specific voice signal of the current phoneme and the second phoneme-specific voice signal of the previous phoneme are mixed so that a ratio of the second phoneme-specific voice signal of the current phoneme increases with the passage of time.
3. The audio signal conversion device according to claim 2.
4. a morpheme analysis unit that divides a character string into morphemes; a phoneme division unit that determines a phoneme string consisting of phonemes that indicate the pronunciation of the morpheme and a mora string consisting of moras that segment the pronunciation. The audio signal conversion device according to any one of claims 1 to 3.
5. A program for causing a computer to function as the audio signal conversion device according to any one of claims 1 to 4.
Citation Information
Patent Citations
IEC23008-3
ITRBS.2076-1
ITRBS.2125-0
Bone conduction microphone output signal reproduction device
JP1996079868A
Voice processor and processing method
JP2009042552A