Audio processing device
Patent Information
- Application Number
- JP2025031191
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
AI Technical Summary
【0006】 本発明によれば、音声を明瞭化することができる。
Smart Images

Figure 2026144090000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech processing apparatus that processes speech uttered by a speaker. [Background Art]
[0002] Conventionally, apparatuses that process audio signals to obtain more preferable speech have been known. For example, in the apparatus described in Patent Document 1, a component in a specific band is extracted from an original audio signal, the extracted signal component is multiplied to generate a high-frequency component, and the high-frequency component is added to the original signal. [Prior Art Documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Unexamined Patent Publication No. Hei 5-266582 [Summary of the Invention] [Problems to be Solved by the Invention]
[0004] However, as in the apparatus described in Patent Document 1 above, it is difficult to clarify speech only by adding high-frequency components to the original speech. [Means for Solving the Problems]
[0005] A speech processing apparatus according to one aspect of the present invention includes: an audio signal acquisition unit that acquires an audio signal; and an overtone conversion unit that identifies a fundamental tone included in the audio signal acquired by the audio signal acquisition unit, and converts non-integer order overtones of the fundamental tone into integer order overtones of the fundamental tone. [Effects of the Invention]
[0006] According to the present invention, speech can be clarified. [Brief Description of the Drawings]
[0007] [Figure 1] Block diagram showing a configuration of main parts of a speech processing apparatus according to an embodiment of the present invention. [Figure 2A] A diagram showing an example of a sound waveform that includes the fundamental tone and integer harmonics. [Figure 2B] A diagram showing an example of a speech waveform that includes the fundamental tone and non-integer harmonics. [Figure 3] A diagram showing an example of the waveform of a composite wave of the fundamental frequency and integer harmonics. [Figure 4] A flowchart showing an example of the harmonic conversion process performed by the controller in Figure 1. [Figure 5] A flowchart showing an example of the background music generation process executed by the controller in Figure 1. [Modes for carrying out the invention]
[0008] Embodiments of the present invention will be described below with reference to Figures 1 to 5. The voice processing device according to the embodiment of the present invention processes the voices spoken by a speaker at seminars, lectures, etc. It is known that in person-to-person communication, auditory information has the second greatest influence after visual information and has a greater influence than linguistic information (Mehrabian's Law). Therefore, in this embodiment, the voice processing device is configured as follows to process and clarify the voices spoken by the speaker in real time, making it easier for participants and listeners to understand the content.
[0009] Figure 1 is a block diagram showing the main components of an audio processing device (hereinafter referred to as "device") 100 according to an embodiment of the present invention. As shown in Figure 1, the device 100 comprises an audio input unit 10 such as a microphone, a controller 20 that processes audio signals input via the audio input unit 10, and an audio output unit 30 such as a speaker that outputs the audio signals processed by the controller 20.
[0010] The controller 20 is comprised of a computer including a processor such as a CPU, memory such as RAM and ROM, and other peripheral circuits. The controller 20 is connected to an audio input unit 10 and an audio output unit 30. The controller 20 has an audio signal acquisition unit 21, a Fourier transform unit 22, a harmonic conversion unit 23, an inverse Fourier transform unit 24, an audio signal output unit 25, a speech recognition unit 26, and a BGM generation unit 27, and functions as the audio signal acquisition unit 21, Fourier transform unit 22, harmonic conversion unit 23, inverse Fourier transform unit 24, audio signal output unit 25, speech recognition unit 26, and BGM generation unit 27.
[0011] The audio signal acquisition unit 21 acquires the audio signal of the speaker's voice, which is input via the audio input unit 10. The Fourier transform unit 22 decomposes the audio signal acquired by the audio signal acquisition unit 21 into frequency (frequency domain) components using a Fourier transform. The frequency resolution (frequency bandwidth) of the Fourier transform is set within a range that ensures real-time audio processing.
[0012] Figures 2A and 2B show examples of speech waveforms. Figure 2A shows an example of a speech waveform that includes the fundamental tone, which is the main component of speech, and integer harmonics with wavelengths that are integer multiples of the fundamental tone. Figure 2B shows an example of a speech waveform that includes the fundamental tone and non-integer harmonics with wavelengths that are not integer multiples of the fundamental tone.
[0013] As shown in Figure 2A, the waveform of a sound containing the fundamental frequency and integer harmonics has a clear period, is nearly symmetrical in the vertical direction, and corresponds to a highly intelligible sound. Such a sound signal can be decomposed by Fourier transform into the frequency of the fundamental frequency (e.g., 440 Hz) and the frequencies of the integer harmonics (e.g., the frequency of the second harmonic, 880 Hz, and the frequency of the third harmonic, 1320 Hz).
[0014] On the other hand, as shown in FIG. 2B, the waveform of speech including a fundamental frequency and non-integer harmonics has an unclear period, no vertical symmetry is observed, and corresponds to speech with low clarity. Such an audio speech signal can be decomposed by Fourier transform into the fundamental frequency (e.g., 490 Hz) and frequencies of non-integer harmonics (e.g., 590 Hz which is the frequency of the 1.2nd harmonic, and 1660 Hz which is the frequency of the 3.4th harmonic).
[0015] The harmonic transforming unit 23 identifies the lowest frequency among the frequencies (frequency domain) decomposed by the Fourier transform unit 22 as the fundamental frequency, thereby specifying the fundamental tone included in the speech signal acquired by the speech signal acquiring unit 21. In addition, for frequencies of sounds other than the fundamental frequency among the frequencies (frequency domain) decomposed by the Fourier transform unit 22, it determines whether the frequency is an integer multiple of the fundamental frequency, thereby determining whether the frequency is an integer harmonic frequency or a non-integer harmonic frequency.
[0016] The harmonic transforming unit 23 attenuates the speech signal at the frequency (frequency domain) identified as a non-integer harmonic frequency to an extremely low level (signal intensity) using a notch filter or the like, thereby removing the non-integer harmonic speech signal from the speech signal acquired by the speech signal acquiring unit 21.
[0017] Further, the harmonic transforming unit 23 rounds the order of the original non-integer harmonics included in the speech signal acquired by the speech signal acquiring unit 21 to specify the closest integer order, and generates an integer harmonic speech signal having the same level (signal intensity) as the original non-integer harmonics. In this case, for example, the 1.2nd harmonic is converted into the 1st harmonic (i.e., the fundamental tone), and the 3.4th harmonic is converted into the 3rd harmonic.
[0018] Alternatively, the order of non-integer harmonics included in the speech signal acquired by the speech signal acquiring unit 21 may be rounded up, and an integer harmonic speech signal having the smallest integer order greater than the order of the non-integer harmonics before conversion may be generated. In this case, for example, the 1.2nd harmonic is converted into the 2nd harmonic, and the 3.4th harmonic is converted into the 4th harmonic.
[0019] It may be preset in advance whether the converted integer harmonics are even harmonics or odd harmonics. For example, when converting a non-integer harmonic to an integer harmonic by rounding up its order, a 1.2nd harmonic is converted to a 2nd harmonic if even harmonics are set, and is converted to a 3rd harmonic if odd harmonics are set.
[0020] The converted integer harmonics may be a combination of even harmonics and odd harmonics, and the ratio of even harmonics to odd harmonics may be preset in advance to achieve a desired sound quality. For example, when rounding up the order of a non-integer harmonic and converting it to an integer harmonic containing even harmonics and odd harmonics at a ratio of 2:3, a 1.2nd harmonic with a level (signal intensity) of 100 is converted to an integer harmonic including a 2nd harmonic with a level of 40 and a 3rd harmonic with a level of 60.
[0021] The harmonic conversion unit 23 converts non-integer harmonics into integer harmonics by removing the audio signal of non-integer harmonics from the audio signal acquired by the audio signal acquisition unit 21 and generating the audio signal of integer harmonics.
[0022] The inverse Fourier transform unit 24 re-synthesizes, through inverse Fourier transform, the audio signal acquired by the audio signal acquisition unit 21 from which the audio signal of non-integer harmonics has been removed by the harmonic conversion unit 23, and the audio signal of integer harmonics generated by the harmonic conversion unit 23. In addition to the fundamental tone included in the original audio signal acquired by the audio signal acquisition unit 21 and the audio signal of integer harmonics generated by the harmonic conversion unit 23, the re-synthesized audio signal may also include the integer harmonics that were included in the original audio signal. The re-synthesized audio signal is output via the audio signal output unit 25.
[0023] Figure 3 shows an example of the waveform of a composite wave of the fundamental tone and integer harmonics, specifically the waveform of a composite wave of the fundamental tone and the second harmonic. As shown in Figure 3, when the fundamental tone and integer harmonics are combined, a composite wave is generated that corresponds to a highly intelligible sound with the same period as the fundamental tone and a vertically symmetrical shape. Similarly, even when the harmonics combined with the fundamental tone are integer harmonics of the third harmonic or higher, or when integer harmonics of multiple frequencies (frequency domains) are combined, a composite wave is generated that corresponds to a highly intelligible sound with the same period as the fundamental tone and a vertically symmetrical shape.
[0024] Figure 4 is a flowchart showing an example of the harmonic conversion process performed by the controller 20 in Figure 1. As shown in Figure 4, first, in step S1, an audio signal input via the audio input unit 10 is acquired. Next, in step S2, the audio signal acquired in step S1 is decomposed into frequencies (frequency domain) by Fourier transform to identify the fundamental frequency. Next, in step S3, it is determined whether or not non-integer harmonics are included in the frequencies (frequency domain) decomposed in step S2. If affirmative in step S3, the process proceeds to step S4; if negative in step S3, the process proceeds to step S5. In step S4, non-integer harmonics are converted to integer harmonics by removing the non-integer harmonics identified in step S3 and generating an audio signal of integer harmonics with an equivalent level (signal intensity). Next, in step S5, the audio signal decomposed into frequencies (frequency domain) in step S2 and converted in step S4 as necessary is recombined by inverse Fourier transform. Next, in step S6, the audio signal recombined in step S5 is output via the audio signal output unit 25.
[0025] The sound a person produces varies depending on the shape of their resonant cavities (the trachea, pharyngeal cavity, oral cavity, and nasal cavity located above the vocal cords). As a result, some speakers may naturally have poor voice clarity. While voice training can correct the shape of the resonant cavities during vocalization, it can be a significant burden for some individuals. Device 100 clarifies the speech by converting non-integer harmonics contained in the speaker's voice to integer harmonics in real time and outputting them, making the content easier for attendees and listeners to understand. Furthermore, because the conversion to integer harmonics is performed without changing the level (signal intensity) of the non-integer harmonics contained in the original speech signal, the balance between the fundamental tone and harmonics (integer and non-integer harmonics) is maintained, suppressing any unnatural feeling caused by the conversion.
[0026] In Figure 1, the speech recognition unit 26 performs speech recognition on the speech signal acquired by the speech signal acquisition unit 21 at predetermined intervals. The BGM generation unit 27 determines a chord progression based on the recognition results from the speech recognition unit 26 and generates background music (BGM) for the determined chord progression. The BGM generated by the BGM generation unit 27 is also output via the speech signal output unit 25.
[0027] For example, speech recognition can determine whether the speaker's current topic is cheerful or serious. If it's a cheerful topic, a cheerful chord progression is determined; if it's a serious topic, a serious chord progression is determined. Topics can shift even within a single-theme seminar or lecture. By determining chord progressions and generating and playing background music according to topic changes, it's possible to enhance the concentration of participants and audience members, making the content of seminars and lectures easier to understand.
[0028] Figure 5 is a flowchart showing an example of the BGM generation process performed by the controller 20 in Figure 1. As shown in Figure 5, first, in step S10, an audio signal input via the audio input unit 10 is acquired. Next, in step S11, speech recognition is used to recognize the content of speech over a predetermined period and determine the type of topic. Next, in step S12, a chord progression in accordance with the topic determined in step S11 is determined and BGM is generated. Next, in step 13, the BGM generated in step S12 is output via the audio signal output unit 25.
[0029] According to embodiments of the present invention, the following effects can be achieved. (1) The device 100 includes an audio signal acquisition unit 21 that acquires an audio signal, and a harmonic conversion unit 23 that identifies the fundamental tone contained in the audio signal acquired by the audio signal acquisition unit 21 and converts non-integer harmonics of the fundamental tone into integer harmonics of the fundamental tone (steps S1 to S4 in Figures 1 and 4). In this way, the audio can be made clearer by converting non-integer harmonics contained in the audio into integer harmonics. Furthermore, by maintaining a balance between the fundamental tone and harmonics (integer harmonics and non-integer harmonics), the unnatural feeling caused by converting non-integer harmonics into integer harmonics can be suppressed.
[0030] (2) The harmonic conversion unit 23 converts non-integer harmonics into integer harmonics whose order is the integer closest to the order of the non-integer harmonics. In this case, the unnaturalness caused by converting non-integer harmonics to integer harmonics can be further suppressed.
[0031] (3) The harmonic conversion unit 23 converts non-integer harmonics into integer harmonics whose order is greater than the order of the non-integer harmonics and is the smallest integer. In this case as well, the unnaturalness caused by converting non-integer harmonics to integer harmonics can be further suppressed. Furthermore, even when non-integer harmonics between the fundamental tone and the second harmonic are included, the balance between the fundamental tone and the harmonics can be maintained.
[0032] (4) Integer harmonics are either even or odd harmonics of the fundamental tone, as set in advance. In other words, it is possible to pre-set whether to convert non-integer harmonics to even or odd harmonics, which not only clarifies the sound but also allows for a preferred sound quality.
[0033] (5) The device 100 further includes a speech recognition unit 26 that recognizes the speech signal acquired by the speech signal acquisition unit 21, and a BGM generation unit 27 that determines a chord progression based on the recognition result by the speech recognition unit 26 and generates background music for the determined chord progression (steps S10 to S13 in Figures 1 and 5). This makes it possible to play background music that is in line with the content of the speaker's speech, thereby increasing the audience's concentration on the content of the speaker's speech.
[0034] In the above embodiment, as shown in Figure 1, an example was described in which a single controller 20 functions as an audio signal acquisition unit 21, a harmonic conversion unit 23, a speech recognition unit 26, a background music generation unit 27, etc. However, the audio signal acquisition unit, harmonic conversion unit, speech recognition unit, and background music generation unit are not limited to these. For example, a controller 20a such as an audio interface equipped with a DSP (Digital Signal Processor) that functions as an audio signal acquisition unit 21a, a harmonic conversion unit 23, etc., and a controller 20b such as a personal computer that functions as an audio signal acquisition unit 21b, a speech recognition unit 26, a background music generation unit 27, etc., may be provided separately. The above description is merely an example, and the present invention is not limited by the embodiments and modifications described above, as long as the features of the present invention are not impaired. It is also possible to arbitrarily combine one or more of the above embodiments and modifications, and to combine modifications with each other. [Explanation of Symbols]
[0035] 10 Audio input unit, 20 Controller, 21 Audio signal acquisition unit, 22 Fourier transform unit, 23 Harmonic conversion unit, 24 Inverse Fourier transform unit, 25 Audio signal output unit, 26 Voice recognition unit, 27 BGM generation unit, 30 Audio output unit, 100 Audio processing device (device)
Claims
1. Audio signal acquisition unit that acquires audio signals, A sound processing device comprising: an audio signal acquisition unit that identifies the fundamental tone contained in the audio signal acquired by the audio signal acquisition unit, and a harmonic conversion unit that converts non-integer harmonics of the fundamental tone into integer harmonics of the fundamental tone.
2. In the voice processing device according to claim 1, The harmonic conversion unit is characterized by converting the non-integer harmonic into an integer harmonic whose order is the integer closest to the order of the non-integer harmonic.
3. In the voice processing device according to claim 1, The harmonic conversion unit is characterized by converting the non-integer harmonic into an integer harmonic whose order is greater than the order of the non-integer harmonic and is the smallest integer.
4. In the audio processing device according to any one of claims 1 to 3, The speech processing device is characterized in that the integer harmonics are either even harmonics or odd harmonics of the fundamental tone, as set in advance.
5. In the audio processing device according to any one of claims 1 to 3, A speech recognition unit that performs speech recognition on the speech signal acquired by the speech signal acquisition unit, A speech processing device further comprising: a BGM generation unit that determines a chord progression based on the recognition result from the speech recognition unit and generates background music for the determined chord progression.
Citation Information
Patent Citations
sound equipment
JP1993266582A