Audio compensation program, device and method using harmonic and background sounds
By generating harmonic and background sound signals to mask missing audio portions, the method addresses the discomfort caused by high-power white noise, offering a seamless and improved subjective quality in audio communication.
Patent Information
- Application Number
- JP2022148871
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-09-20
AI Technical Summary
Existing audio communication technologies struggle to effectively compensate for long periods of packet loss in digital audio signals, leading to discomfort and degraded subjective quality due to the insertion of high-power white noise, which fails to provide a seamless listening experience.
The use of harmonic and background sound signals generated based on preceding audio signal components or preset conditions to create a compensation signal that masks the missing portion, with the harmonic sound signal mimicking speech vowels and the background sound signal enhancing the continuous listening effect.
The proposed method provides a less unpleasant audio compensation that improves subjective quality by creating a seamless listening experience, reducing discomfort and enhancing the perceived continuity of speech communication.
Smart Images

Figure 0007725436000001 
Figure 0007725436000002
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for compensating for missing portions of an audio signal. [Background technology]
[0002] Currently, digital audio communication using audio packets is widely used. In digital audio communication, packet loss, channel capacity limitations, and other factors can cause temporal gaps in the audio signal. Audio signals with such gaps can cause discomfort to listeners and significantly degrade the subjective quality (user perceived quality). Furthermore, the occurrence of gaps in audio signals is a major problem in multimedia communication and streaming playback over the Internet.
[0003] For example, in voice transmissions where real-time performance is not a major requirement, the missing portion can be compensated for by repeatedly transmitting the same voice data.In contrast, in voice transmissions where real-time performance is highly required, such as voice communications, packet loss concealment (PLC) has been proposed as a method of compensating for missing portions.
[0004] PLC is a loss-concealing method standardized in Non-Patent Document 1 (ITU-T Recommendation G.711, Appendix I), and specifically, waveform replacement and predictive replacement methods are proposed as examples of this PLC. However, this standardized PLC can only handle packet losses with a duration of 60 ms (milliseconds) or less. In fact, even with this PLC, loss portions exceeding 60 ms become silent. This is because it is no longer possible to physically restore the audio waveform in an audio signal that has experienced a significant amount of packet loss. However, currently, packet losses exceeding 60 ms occur frequently, and this standardized PLC is unable to accurately address such packet loss.
[0005] Therefore, loss compensation methods that can cope with the occurrence of long periods of packet loss are currently being studied. For example, as disclosed in Non-Patent Document 2, the part of the human brain that controls hearing in particular has various mechanisms that enable smooth speech communication even under poor conditions. One of these is the continuum effect (phoneme repair phenomenon, phonological repair phenomenon), which is explained in detail in Non-Patent Documents 3 and 4.
[0006] The continuum effect is a phenomenon in which missing parts of an audio signal are filled with an acoustic signal unrelated to the audio, causing sounds that should be interrupted to be perceived as being smoothly connected, and is considered to be one of the auditory illusions. In order to elucidate the mechanism by which this auditory illusion occurs and to apply this illusion to audio signal processing, active research is currently being conducted to generate an effective continuum effect, as disclosed in, for example, Non-Patent Document 5. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] ITU-T Recommendation G.711, Appendix I, “A high quality low-complexity algorithm for packet loss concealment with G.711”, 1999 [Non-patent document 2] RM Warren, “Auditory Perception: A New Analysis and Synthesis”, Cambridge University Press, Cambridge, 1999. [Non-patent document 3] GA Miller and JCR Licklider, “The intelligibility of interrupted speech”, Journal of the Acoustical Society of America, vol.22, pp.167-173, 1950 [Non-patent document 4] RM Warren, “Perceptual restoration of missing speech sounds”, Science. 167, pp.392-393, 1970 [Non-patent document 5] M. Kashino, “Phonemic Restoration: The brain creates missing speech sounds”, Acoustical Science and Technology, 27(6), pp.318-321, 2006 [Non-patent document 6] BCJ Moore, “An Introduction to the Psychology of Hearing”, 5th Edition, Emerald Group Publishing Ltd, 2003 Summary of the Invention [Problem to be solved by the invention]
[0008] However, as pointed out in Non-Patent Document 5, in order to effectively produce the above-mentioned continuous hearing effect, it has been considered necessary that (Condition 1) an acoustic signal with sufficiently large power is inserted into the missing part to produce a masking effect, and (Condition 2) an acoustic signal is inserted into the missing part without any gaps, so that no break in sound is perceived at the start of the missing part.
[0009] Therefore, the only method proposed to fill in the missing parts of a speech signal using the continuum effect has been to insert high-power broadband noise (white noise) seamlessly into the missing parts. However, this conventional method still leaves the listener with a sense of discomfort despite the speech completion process, making it difficult to improve the subjective quality of speech communication, for example.
[0010] Therefore, an object of the present invention is to provide a sound compensation program, device, and method that can perform sound compensation that is less unpleasant. [Means for solving the problem]
[0011] According to the present invention, there is provided an audio compensation program for compensating for a missing portion of an audio signal, comprising: a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal; Current Complement a compensation signal inserting means for inserting a compensation signal into the missing portion; to make the computer function 、 The background sound generating means generates the background sound signal based on the sound signal acquired as the ambient sound when the ratio between the amplitude of the signal portion of the voice section in the sound signal acquired as the ambient sound and the amplitude of the signal portion other than the voice section is equal to or less than a predetermined value. Characterized by A voice compensation program is provided. According to the present invention, there is also provided an audio compensation program for compensating for a missing portion of an audio signal, the program comprising: a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal by combining the harmonic sound signal and the background sound signal or by connecting the harmonic sound signal and the background sound signal; a compensation signal inserting means for inserting the compensation signal into the missing portion; to make the computer function, When the time length of the missing portion is equal to or longer than a predetermined length, the compensation signal generating means generates the compensation signal by combining the harmonic sound signal and the background sound signal using a synthesis ratio that is a monotonically increasing function of the ratio between the amplitude of the signal portion of the voice section in the signal portion before the missing portion and the amplitude of the signal portion other than the voice section. A voice compensation program is provided.
[0012] As one embodiment of the audio compensation program according to the present invention, it is also preferable that the harmonic sound generating means determines the fundamental frequency of the signal portion before the missing portion, or obtains the fundamental frequency as the preset frequency condition, and generates the harmonic sound signal having a fundamental frequency component related to the fundamental frequency and a harmonic frequency component related to a harmonic frequency relative to the fundamental frequency.
[0013] It is also preferable that the harmonic sound generating means generates the harmonic sound signal whose amplitude initially takes a value determined based on the amplitude of the signal portion preceding the missing portion and then decreases over time.Furthermore, it is also preferable that the compensation signal includes the harmonic sound signal at least at the beginning of the signal, and that the compensation signal inserting means inserts the compensation signal by connecting it to the signal portion immediately preceding the missing portion.
[0015] Also , back In an embodiment relating to a background sound signal, the background sound generating means When the ratio of the amplitude of the signal portion of the voice section in the sound signal acquired as the surrounding sound to the amplitude of the signal portion other than the voice section is equal to or less than a predetermined value, In the acoustic signal For signal parts that are not speech segments, Alternatively, it is also preferable to generate the background sound signal by performing a weighting process and / or a filtering process having frequency dependency corresponding to the amplitude spectrum of the signal portion of the voice section on the signal portion before the missing portion that is not a voice section.
[0016] Furthermore, it is also preferable that the background sound generation means generates the background sound signal based on the acoustic signal acquired as the ambient sound when the ratio between the amplitude of the signal portion of the voice section in the acoustic signal acquired as the ambient sound and the amplitude of the signal portion that is not the voice section is equal to or less than a predetermined value.
[0017] Furthermore, in an embodiment relating to this background sound signal, it is also preferable that the compensation signal generating means generates the compensation signal by combining the harmonic sound signal with the background sound signal, or by connecting the harmonic sound signal followed by the background sound signal.
[0018] Furthermore, in the case of generating the compensation signal by synthesis, when the time length of the missing portion is equal to or longer than a predetermined value, the compensation signal generating means preferably generates the compensation signal by synthesizing the harmonic sound signal and the background sound signal using a synthesis ratio that is a monotonically increasing function of the ratio between the amplitude of the signal portion of the voice section in the signal portion before the missing portion and the amplitude of the signal portion that is not the voice section.
[0019] According to the present invention, there is also provided an audio compensation program for compensating for a missing portion of an audio signal, the program comprising: a harmonic sound generating means for generating a harmonic sound signal having a specific frequency component based on a signal portion preceding the missing portion in the audio signal or based on a pre-prepared acoustic signal; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal when a ratio of the amplitude of a signal portion in a voice section in the signal portion before the missing portion to the amplitude of a signal portion other than the voice section is lower than a predetermined threshold value; Current Complement a compensation signal inserting means for inserting a compensation signal into the missing portion; A voice compensation program is provided that causes a computer to function as a
[0020] Book According to the present invention, there is further provided an audio compensation device for compensating for a missing portion of an audio signal, comprising: a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal; Current Complement a compensation signal inserting means for inserting a compensation signal into the missing portion; With death, The background sound generating means generates the background sound signal based on the sound signal acquired as the ambient sound when the ratio between the amplitude of the signal portion of the voice section in the sound signal acquired as the ambient sound and the amplitude of the signal portion other than the voice section is equal to or less than a predetermined value. Characterized by An audio compensation device is provided.
[0021] According to the present invention, there is further provided an audio compensation system for compensating for a missing portion of an audio signal, comprising: a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal; Current Complement a compensation signal inserting means for inserting a compensation signal into the missing portion; With death, The background sound generating means generates the background sound signal based on the sound signal acquired as the ambient sound when the ratio between the amplitude of the signal portion of the voice section in the sound signal acquired as the ambient sound and the amplitude of the signal portion other than the voice section is equal to or less than a predetermined value. Characterized byA voice compensation system is provided.
[0022] According to the present invention, there is further provided an audio compensation method for compensating for a missing portion of an audio signal, the method comprising the steps of: generating a harmonic sound signal, which is an acoustic signal having specific frequency components, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; A step of generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; generating a compensation signal from the harmonic sound signal and the background sound signal; Current Complement inserting a compensation signal into the missing portion; and In the step of generating a background sound signal, when a ratio between an amplitude of a signal portion of a voice section in the sound signal acquired as the ambient sound and an amplitude of a signal portion other than the voice section is equal to or less than a predetermined value, the background sound signal is generated based on the sound signal acquired as the ambient sound. child A computer-implemented method for audio compensation is provided, comprising: [Effects of the Invention]
[0023] According to the audio compensation program, device and method of the present invention, audio compensation can be performed in a less unpleasant manner. [Brief explanation of the drawings]
[0024] [Figure 1] 1 is a functional block diagram showing a functional configuration of an embodiment of a voice compensation device according to the present invention; [Figure 2] 1 is a spectrogram (voiceprint) for explaining an embodiment of a voice compensation method according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0025] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0026] [Audio compensation device] 1 is a functional block diagram showing the functional configuration of an embodiment of a voice compensation device according to the present invention, in which functional components not related to voice compensation processing are omitted.
[0027] 1, which is an embodiment of the audio compensation device according to the present invention, is a device that compensates for missing portions of an audio signal received by, for example, a communication interface unit 101, and outputs the audio signal that has undergone audio compensation processing from, for example, a speaker 105. In this embodiment, the audio compensation processing is performed substantially in real time. As a result, for example, suitable digital audio communication can be performed with other communication terminals (other smartphones 1 in FIG. 1), and multimedia communication, streaming playback, and the like can be performed suitably.
[0028] Generally, digital audio communication using audio packets often results in time-defective portions due to packet loss, channel capacity limitations, and other factors. Such audio signal defects can also occur in multimedia communication over the Internet, streaming playback, and other similar applications. The smartphone 1 is capable of compensating for such audio signal defects.
[0029] Specifically, in order to perform such voice compensation processing, the smartphone 1: (A) a harmonic sound generating unit 112 that generates a "harmonic sound signal" that is an acoustic signal having specific frequency components based on a signal portion before the missing portion in the (received) audio signal or based on a preset frequency condition; (B) a compensation signal inserting unit 111c that inserts a "compensation signal" including the generated "harmonic sound signal" into the missing part; It is characterized by having:
[0030] Here, the "harmonic signal" in (A) above is an acoustic signal having specific frequency components, as described above. In contrast, the envelope of the amplitude spectrum in the voice section of a speech signal usually forms multiple peaks. These peaks (the corresponding frequency bands) are called formants, and are named in ascending order of frequency as the first formant, second formant, etc. These formants are important parts that correspond to vowels in speech, and vowels in speech are largely determined by the positions of the first and second formants on the frequency axis.
[0031] Therefore, by generating a "harmonic signal" having frequency components roughly corresponding to at least the first and second formants of a certain audio signal and connecting the generated "harmonic signal" to the end of this audio signal, it is possible to give the listener the perception (illusion) that the audio related to this audio signal is continuing to occur without interruption even in the missing section.
[0032] Furthermore, such a "harmonic sound signal" can be a signal with a smaller intensity (amplitude) than, for example, the high-power white noise that has traditionally been used to compensate for speech defects. Furthermore, as described above, the "harmonic sound signal" is not noise, but rather a signal that produces a perceived sound closer to a vowel in speech. The smartphone 1 inserts a "compensation signal" containing such a suitable "harmonic sound signal" into the missing portion, thereby enabling speech compensation that is less unpleasant and (in the case of voice communication) has improved subjective quality.
[0033] Incidentally, the smartphone 1 further includes the following as an embodiment to be described in detail later: (C) a background sound generation unit 113 that generates a "background sound signal" that is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band based on an audio signal that has been prepared in advance or that has been acquired as ambient sound, or based on a signal portion preceding a missing portion in a (received) audio signal; (D) a compensation signal generation unit 114 that generates a "compensation signal" from the generated "harmonic sound signal" and "background sound signal"; By using this "background sound signal" as the "compensation signal", it is possible to achieve a continuous listening effect, and audio loss due to longer packet losses can also be suitably repaired.
[0034] Furthermore, in the present invention, the harmonic sound generation unit 112 (A), the background sound generation unit 113 (C), and the compensation signal generation unit 114 (D) may be included in separate devices, and the generated compensation signals may be transmitted to an audio compensation device (e.g., a smartphone) equipped with the compensation signal insertion unit 111c (B). Furthermore, for example, the harmonic sounds and the background sounds may be generated by separate devices, or may be generated and prepared in advance. In any case, these functional components as a whole constitute the audio compensation system according to the present invention.
[0035] [Device configuration, audio compensation program and method] The functional configuration of a smartphone 1 as an embodiment of a voice compensation device according to the present invention will be described in more detail below. Similarly, in the functional block diagram of Fig. 1, the smartphone 1 includes a communication interface 101, a voice data buffer 102, an acoustic signal storage unit 103, a compensation signal storage unit 104, a speaker (SP) 105, a microphone (MC) 106, and a processor / memory (a processing system with a memory function). The processor / memory stores a voice compensation program according to the present invention, has a computer function, and executes the voice compensation program to perform voice compensation processing.
[0036] Here, the audio compensation device according to the present invention is not limited to a smartphone, but may also be a device dedicated to audio compensation processing equipped with the audio compensation program according to the present invention, a general-purpose cloud server, a non-cloud server, a personal computer (PC), a notebook or tablet computer, a wearable device, or even other communication terminal capable of receiving audio signals (including audio signals of various content).
[0037] The processor memory also includes, as functional components, an audio signal processing unit 111 including a codec unit 111a, a loss detection unit 111b, and a compensation signal insertion unit 111c; a harmonic sound generation unit 112 including a fundamental wave / harmonic generation unit 112a; a background sound generation unit 113 including a speech section spectrum generation unit 113a and a weighting processing unit 113b; a compensation signal generation unit 114 including a synthesis ratio determination unit 114a; and an input / output control unit 121. These functional components can be considered to be functions realized by executing an audio compensation program according to the present invention stored in the processor memory. The processing flow shown by arrows connecting the functional components of the smartphone 1 in the functional block diagram of FIG. 1 can also be understood as one embodiment of an audio compensation method according to the present invention.
[0038] In the embodiment described below, the smartphone 1 performs compensation processing for missing portions of an audio signal in audio communication. However, the smartphone 1 can also perform audio compensation processing in communication related to audio signals other than audio communication (for example, multimedia communication, streaming playback, etc.) in the same manner as described below.
[0039] <Audio signal processing means> 1, the audio signal processing unit 111 is a functional component that processes audio signals related to audio communication and enables digital audio communication (transmission and reception of audio packets) with other communication terminals via the communication interface unit 101 and the communication equipment 2. First, the codec unit 111a (a) Decode the received (encoded) voice packets to generate a voice signal (digital voice waveform data), and (b) Generates a voice packet by encoding the voice signal (digital voice waveform data) acquired from the microphone 106 through the input / output control unit 121. The processing is carried out.
[0040] Here, the audio signal generated in (a) above is output from the speaker 105 via the input / output control unit 121, and the audio packet generated in (b) above is transmitted from the communication interface unit 101. In this embodiment, the audio signal (digital audio waveform data) described above can be data in various formats, such as waveform data on the time axis, amplitude spectrum, power spectrum, spectrogram, etc.
[0041] Next, the loss detection unit 111b detects loss portions in the generated audio signal. Specifically, for each frame of the audio signal (digital audio waveform data), the time variation of amplitude (signal strength), amplitude spectrum, and power spectrum are obtained, and a time period in which the amplitude or power is below a predetermined lower threshold is identified and determined as a loss portion. Alternatively, the loss portion may be determined by detecting a code 'sil' that indicates a silent period and is added to the decoded audio signal. The compensation signal insertion unit 111c will be described in detail later.
[0042] <Harmonic sound generation means> Also in the functional block diagram of Figure 1, the harmonic sound generation unit 112 generates a harmonic sound signal based on (a) the signal portion preceding the detected missing portion in the received audio signal, which in this embodiment is the frame immediately preceding it within a predetermined short period of time, or (b) based on preset frequency conditions.
[0043] Specifically, in this embodiment, first, the fundamental wave / harmonic wave generating unit 112a of the harmonic sound generating unit 112 performs the following: (a) determining the fundamental frequency of the signal portion immediately preceding the defect, for example by frequency analysis; or (b) Obtain a fundamental frequency as a preset frequency condition (for example, 130 Hz, which is expected to have a masking effect on speech), A fundamental wave having a fundamental frequency component related to this fundamental frequency and harmonics (e.g., 2nd to 50th harmonics) having harmonic frequency components related to harmonic frequencies of this fundamental frequency are generated. Here, the fundamental frequency of the audio signal is usually 500 Hz or less, for example, a frequency between 100 Hz and 200 Hz.
[0044] Next, the harmonic sound generation unit 112 combines the generated fundamental wave and these harmonics at a predetermined amplitude ratio (for example, an amplitude ratio in which the higher the harmonic, the smaller the amplitude relative to the amplitude of the fundamental wave) to generate a harmonic sound signal (hereinafter also referred to as a harmonic complex sound signal) having the fundamental wave frequency component and the harmonic frequency component.
[0045] Generally, a speech signal has a harmonic structure in which acoustic energy is concentrated in a specific frequency band and roughly corresponds to a vowel. A harmonic complex sound signal generated based on the fundamental frequency of such a speech signal (or a fundamental frequency set in consideration of the masking effect on speech) can have a masking effect on at least the first and second formants of the speech signal. Furthermore, by seamlessly connecting such a harmonic complex sound signal to the speech signal portion immediately preceding the missing portion, it is possible to give the listener the perception (illusion) that the speech corresponding to that speech signal portion is occurring continuously without interruption. Furthermore, such a harmonic complex sound signal is a signal that produces a perception of sound similar to a speech vowel, thereby enabling speech compensation with less discomfort (unnaturalness) and improved subjective quality in speech communication.
[0046] If the duration of the missing portion exceeds a predetermined time (e.g., 60 ms), the harmonic complex sound signal does not have to be generated for the entire missing portion. In other words, it is preferable to generate only the first predetermined time (e.g., 60 ms) of the missing portion to create a continuous hearing effect.
[0047] Furthermore, in this embodiment, the amplitude of this harmonic complex sound signal initially takes a value determined based on the amplitude of the signal portion immediately before the loss (e.g., the same value as that amplitude), and is then adjusted to a signal that decreases over time, for example, a signal whose amplitude attenuates by 20% over 200 ms. A harmonic complex sound signal with such an amplitude can be a signal with a lower intensity (amplitude) than, for example, high-power white noise, which has traditionally been used to compensate for speech losses. Furthermore, such a harmonic sound signal is a signal that is closer to a speech signal (whose sound pressure level is not constant over time) and, as described above, produces a perceived sound similar to a vowel in speech. As a result, speech compensation can be performed with less discomfort (noisiness) and improved subjective quality in speech communication.
[0048] It is also preferable that the harmonic complex sound signal generated as described above is buffered or stored in the acoustic signal storage unit 103. In particular, it is also preferable that (b) the harmonic complex sound signal generated based on preset frequency conditions is stored in advance in the acoustic signal storage unit 103, and is read out as needed and used for the voice compensation processing.
[0049] <Background sound generation means> Similarly, in the functional block diagram of FIG. 1, the background sound generation unit 113 includes: (A) For example, in the audio signal storage unit 103, an audio signal prepared in advance, (a) an acoustic signal captured as ambient sound (at the receiving end), for example, by a microphone 106; or (c) The portion of the received voice signal (including the surrounding sounds of the transmitting side) before the missing portion Based on the above, a background sound signal is generated, which is an acoustic signal having an amplitude spectrum that peaks in a predetermined low frequency band.
[0050] Specifically, in this embodiment, the background sound generation unit 113 performs one or both of a weighting process and a filtering process having frequency dependency corresponding to the amplitude spectrum of the signal portion of the voice section on the signal portion that is not a voice section in the acoustic signal (a) or (b) (or in the signal portion (c)) to generate a background sound signal. Note that the amplitude spectrum of the voice section is generated by the voice section spectrum generation unit 113a, and the weighting process and filtering process are performed by the weighting processing unit 113b.
[0051] For example, the background sound generation unit 113 may generate an amplitude spectrum for a (certain) long period of time in the signal portion of the voice section, and perform filtering on the signal portion that is not in the voice section using a band-pass filter having a gain that matches the frequency dependency of this amplitude intensity, thereby generating a background sound signal. Here, the pass band of the band-pass filter may be, for example, a frequency band from 50 Hz to 500 Hz that can mask the first and second formants of a normal voice signal.
[0052] Alternatively, the non-speech signal portion may be subjected to a normal band-pass filtering process, and each frequency component in the processed signal portion may be weighted according to the frequency dependence of the amplitude intensity. Alternatively, such weighting may be performed directly on the non-speech signal portion without band-pass filtering. Still further, it is also possible to perform only normal band-pass filtering without performing such weighting.
[0053] In any case, in the background sound generation process of this embodiment, it is also preferable to use signals corresponding to background sounds present in a real environment, such as environmental sounds (environmental noise), as the acoustic signals (signal portions) of (A) to (C) above. Generally, environmental noise (such as typical air conditioner noise) is a sound in which acoustic energy is concentrated in the low frequency band. In this regard, it is known that a sound of a certain frequency exerts a stronger masking effect on high frequency sounds than on sounds of lower frequencies (see Non-Patent Document 6). Therefore, by using such environmental noise concentrated in the low frequency band as a background sound signal, it is possible to perform a more natural and less unpleasant audio compensation process (loss repair process). From this, it can be understood that it is also highly preferable to use the receiving-side acoustic signal (environmental sound) of (B) above as the acoustic signal for generating background sounds.
[0054] As a further preferred modification, it is also preferable to select the signal that serves as the basis for generating the background sound signal from among the above (A) to (C) depending on the content of the acoustic signal acquired as the ambient sound (on the receiving side) in (B) above. More specifically, for example, if the ratio between the amplitude of the signal portion of the speech section in the acoustic signal in (B) above and the amplitude of the signal portion other than the speech section (hereinafter abbreviated as the audio background amplitude ratio) is less than 1 (a negative value in dB), i.e., if the receiving side is in a high-noise environment, the background sound signal may be generated using the acoustic signal in (B) above, taking into consideration the reduction in loudness due to noise after loss repair. On the other hand, if this audio background amplitude ratio is 1 or more (0 dB or more in dB), the background sound signal may be generated using the acoustic signal prepared in advance in (A) above. Incidentally, in this specific example, the signal portion (including the surrounding sounds on the transmitting side) before the missing portion (c) above is set not to be used (selected) because the amplitude of non-voice sections is often suppressed by a noise suppressor, and the continuous listening effect on the receiving side cannot be expected to be as great as with the acoustic signal (b) above.
[0055] Furthermore, when the ratio (audio background amplitude ratio) between the amplitude of the signal portion of the speech section in the sound signal acquired as the ambient sound in (B) above (or (C) above) and the amplitude of the signal portion other than the speech section is lower than a predetermined threshold (when the environmental noise is relatively loud), the background sound generation unit 113 may generate a background sound signal based on the sound signal acquired as the ambient sound. In this case, the compensation signal is generated from this generated background sound signal and a harmonic sound signal. On the other hand, in cases other than those described above, it is also possible to set the background sound signal not to be generated. In this case, the compensation signal is generated from only the harmonic sound signal (harmonic complex sound signal).
[0056] Furthermore, as an alternative to the above-described embodiment, the compensation signal inserted into the missing portion may not include a harmonic sound signal (harmonic complex sound signal), but may be the background sound signal generated as described above. In this case, too, it is possible to perform audio compensation that is less unpleasant than the conventional insertion of white noise.
[0057] <Compensation signal generation means> Also in the functional block diagram of FIG. 1 , the compensation signal generation unit 114 generates a compensation signal from the generated harmonic sound signal (harmonic complex sound signal) and the generated background sound signal. Alternatively, as described above, the compensation signal generation unit 114 may use the generated harmonic sound signal (harmonic complex sound signal) as the compensation signal as is, or may use the generated background sound signal as is. Here, it is also preferable to set the duration of the compensation signal to be the same as the duration of the missing portion of the audio signal. However, if the compensation signal is the harmonic sound signal (harmonic complex sound signal) as is and the duration of this missing portion exceeds a predetermined time (e.g., 60 ms), the compensation signal may be a signal having a duration of the predetermined time (60 ms) sufficient to produce the continuous hearing effect.
[0058] In this embodiment, the compensation signal generator 114 specifically: (a) combining the generated harmonic sound signal with the generated background sound signal; or (b) Connect the generated harmonic sound signal to the generated background sound signal, Here, when combining the two signals as in (a) above, they may simply be combined at a predetermined combining ratio (the ratio of the amplitudes of the two signals used in combining), for example, a combining ratio less than 1.
[0059] As a further preferred modification, when the time length of the missing portion of (c) above is equal to or longer than a predetermined length (for example, equal to or longer than 60 ms), the compensation signal generating unit 114 preferably generates the compensation signal by synthesizing the harmonic sound signal and the background sound signal using a synthesis ratio that is a monotonically increasing function of the ratio (audio background amplitude ratio) between the amplitude of the signal portion of the voice section in the signal portion before the missing portion and the amplitude of the signal portion that is not the voice section.
[0060] In this case, for example, first, the amplitude of the generated harmonic complex sound signal is adjusted to the same value as the amplitude of the signal in the speech section in the signal portion before the missing portion. This makes it possible to prevent the perception of a break or large fluctuation in sound near the start of the missing portion, and to achieve the effect of continuous listening. Next, the amplitude of the generated background sound signal is adjusted to 0 dB based on the amplitude of the adjusted harmonic complex sound signal. After these adjustments, the synthesis ratio of the harmonic complex sound signal and the background sound signal is set to (a) If the duration of the missing part is 60 ms or less, it is set to 1 (0 dB). (b) If the duration of the missing portion exceeds 60 ms, the threshold may be set to 0.7 (-3 dB) if the audio background amplitude ratio in the signal portion preceding the missing portion is 2 (6 dB) or less, or 1 (0 dB) if the audio background amplitude ratio exceeds 2 (6 dB).
[0061] In this way, by adjusting the synthesis ratio and generating a compensation signal, it is possible to perform audio compensation that causes less discomfort (a sense of incongruity or loudness) as a result of the adjustment being performed appropriately.
[0062] As yet another modification, the compensation signal generation unit 114 may be configured to generate a compensation signal from the harmonic sound signal and the background sound signal when the ratio (audio background amplitude ratio) between the amplitude of the signal portion of the speech section in the signal portion before the missing portion in (c) above and the amplitude of the signal portion that is not the speech section exceeds a predetermined threshold (when environmental noise is relatively small), and otherwise generate a compensation signal from only the background sound signal. Alternatively, it may be configured to generate a compensation signal from only the harmonic sound signal when this audio background amplitude ratio exceeds a predetermined threshold, and otherwise generate a compensation signal from the harmonic sound signal and the background sound signal.
[0063] <Means for inserting compensation signal> 1, the compensation signal inserting unit 111c of the audio signal processing unit 111 inserts the generated compensation signal into the detected missing portion. Here, in this embodiment, the compensation signal is a signal that includes a harmonic sound signal at least at the beginning of the signal, either (a) as a synthesized signal with a background sound signal, or (b) in the form of a harmonic sound signal followed by a background sound signal.
[0064] In this embodiment, the compensation signal inserting unit 111c inserts such a compensation signal into the signal portion immediately preceding the missing portion of the received audio signal, so as to seamlessly connect the signal, thereby giving the listener the perception (illusion) that the audio related to the received audio signal continues to occur without interruption even in the missing section.
[0065] In this embodiment, the audio signal that has been subjected to this suitable audio compensation processing is output to the speaker 105 in real time (via the input / output control unit 121). In reality, the time length of an audio packet is, for example, 20 ms, but the processing time for the compensation signal generation and insertion processing described above can be kept to less than 20 ms. As a result, such real-time audio communication can be realized.
[0066] [Example] Fig. 2 shows a spectrogram (voiceprint) for explaining an embodiment of the voice compensation method according to the present invention. The spectrogram is a graph that shows the amplitude intensity of the voice signal by its color (although it is shown as light and dark in Fig. 2).
[0067] Figure 2(A) shows a received speech signal with two missing portions on the time axis. To make the missing portions clearer compared to the original speech signal that was originally intended to be received, the two missing portions (missing time intervals) are shaded in black. Figure 2(B) shows the result of a conventional white noise insertion process in which high-amplitude white noise is inserted into these missing portions. A speech signal that has undergone this type of conventional speech completion processing is perceived as unpleasant (unnatural, noisy) by the listener.
[0068] Next, Figure 2(C) shows the result of harmonic sound insertion processing in Example 1, in which a compensation signal consisting of a harmonic complex sound signal is inserted into these missing parts. Here, the inserted harmonic complex sound signal has a fundamental frequency of 130 Hz and is adjusted so that its amplitude intensity is the same as that of the signal part immediately preceding the missing part. It can be seen that the result of this harmonic sound insertion processing largely maintains the structure of the signal part immediately preceding the missing part, particularly the frequency structure of the vowels, even in the missing part, without any interruption. In fact, when listening to the speech signal resulting from this harmonic sound insertion processing, it was confirmed that the speech related to this speech signal had a reduced unpleasantness (sense of strangeness and annoyance).
[0069] Finally, Fig. 2(D) shows the result of harmonic and background synthetic sound insertion processing in which a compensation signal consisting of harmonic and background synthetic sound signals is inserted into these missing parts as Example 2. Here, the inserted harmonic and background synthetic sound signals are (a) a harmonic sound signal having a fundamental frequency of 130 Hz and an amplitude intensity adjusted to be the same as the amplitude intensity of the signal portion immediately before the missing portion; (b) A background sound signal generated by applying a band-pass filter with a passband of 50 to 500 Hz to the prepared acoustic signal. are generated by combining them at a predetermined combining ratio.
[0070] The results of this harmonic and background synthetic sound insertion process show that the low frequency bands corresponding to at least the first and second formants in the signal portion immediately preceding the missing portion are masked in the missing portion, and the frequency structure at higher frequencies in that signal portion is largely maintained in the missing portion.In fact, when the speech signal resulting from this harmonic and background synthetic sound insertion process was listened to, it was confirmed that the speech related to this speech signal also had a reduced unpleasantness (sense of incongruity and noisiness).
[0071] As explained in detail above, according to the present invention, the missing portion of the speech signal is compensated for using the compensation signal composed of the harmonic sound signal and / or background sound signal as explained above, thereby making it possible to perform speech compensation that is less unpleasant. Furthermore, if the present invention is applied to speech communications, it can also contribute to improving the subjective quality of speech communications.
[0072] Furthermore, by using the less unpleasant voice compensation signal generated by the present invention, it is possible to provide various high-quality online learning opportunities to children and students not only in urban areas but also in rural areas. In other words, the present invention can contribute to Goal 4 of the Sustainable Development Goals (SDGs) led by the United Nations, "Ensure inclusive and equitable quality education and promote lifelong learning opportunities for all."
[0073] Furthermore, for example, it is possible to provide high-quality online business meetings and vocational training opportunities to adults in rural areas as well as urban areas by using the less unpleasant voice compensation signals generated by this invention. In other words, this invention can contribute to the achievement of Goal 8 of the United Nations' SDGs, "Promote inclusive and sustainable economic growth, employment and decent work for all."
[0074] With respect to the various embodiments of the present invention described above, various changes, modifications, and omissions within the scope of the technical spirit and perspective of the present invention may be easily made by those skilled in the art. The above description is merely an example and is not intended to be limiting in any way. The present invention is limited only by the claims and their equivalents. [Explanation of symbols]
[0075] 1. Smartphone (voice compensation device) 101 Communication Interface 102 Audio data buffer 103 Acoustic signal storage section 104 Compensation signal storage section 105 Speaker (SP) 106 Microphone (MC) 111 Audio signal processing section 111a codec section 111b Defect detection unit 111c Compensation signal insertion section 112 Harmonic sound generation section 111a Fundamental wave / harmonic generation section 113 Background sound generation section 113a Voice section spectrum generation unit 113b Weighting processing unit 114 Compensation signal generation section 114a Combination ratio determination section 121 Input / Output Control Unit 2. Communication equipment
Claims
1. An audio compensation program for compensating for a missing portion of an audio signal, a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal; a compensation signal inserting means for inserting the compensation signal into the missing portion; to make the computer function, The background sound generating means generates the background sound signal based on the sound signal acquired as the ambient sound when a ratio between an amplitude of a signal portion of a voice section in the sound signal acquired as the ambient sound and an amplitude of a signal portion other than the voice section is equal to or smaller than a predetermined value. A voice compensation program characterized by:
2. An audio compensation program for compensating for a missing portion of an audio signal, a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal by combining the harmonic sound signal and the background sound signal or by connecting the harmonic sound signal and the background sound signal; a compensation signal inserting means for inserting the compensation signal into the missing portion; to make the computer function, When the time length of the missing portion is equal to or longer than a predetermined length, the compensation signal generating means generates the compensation signal by synthesizing the harmonic sound signal and the background sound signal using a synthesis ratio that is a monotonically increasing function of the ratio between the amplitude of the signal portion of the voice section in the signal portion before the missing portion and the amplitude of the signal portion other than the voice section. A voice compensation program characterized by:
3. 3. The audio compensation program according to claim 1, wherein the harmonic sound generating means determines the fundamental frequency of the signal portion before the missing portion, or acquires the fundamental frequency as the preset frequency condition, and generates the harmonic sound signal having a fundamental frequency component related to the fundamental frequency and a harmonic frequency component related to a harmonic frequency of the fundamental frequency.
4. 3. The audio compensation program according to claim 1, wherein the harmonic sound generating means generates a harmonic sound signal whose amplitude initially takes a value determined based on the amplitude of the signal portion preceding the missing portion and then decreases over time.
5. the compensation signal includes the harmonic sound signal at least at the beginning of the signal, The compensation signal inserting means inserts the compensation signal in a manner that connects it to the signal portion immediately before the missing portion.
3. The voice compensation program according to claim 1 or 2.
6. 3. The audio compensation program according to claim 1, wherein the background sound generation means generates the background sound signal by performing weighting processing and / or filtering processing having frequency dependency corresponding to the amplitude spectrum of the signal portion of the audio section on the audio signal acquired as the ambient sound when a ratio between the amplitude of the signal portion of the audio section and the amplitude of the signal portion that is not the audio section is equal to or smaller than a predetermined value, on the audio signal that is not the audio section, or on the audio signal portion that is not the audio section in the signal portion preceding the missing portion.
7. An audio compensation program for compensating for a missing portion of an audio signal, a harmonic sound generating means for generating a harmonic sound signal having a specific frequency component based on a signal portion preceding the missing portion in the audio signal or based on a pre-prepared acoustic signal; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal when a ratio of the amplitude of a signal portion in a voice section in the signal portion before the missing portion to the amplitude of a signal portion other than the voice section is lower than a predetermined threshold value; a compensation signal inserting means for inserting the compensation signal into the missing portion; A voice compensation program characterized by causing a computer to function as follows.
8. An audio compensation device for compensating for a missing portion of an audio signal, comprising: a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal; a compensation signal inserting means for inserting the compensation signal into the missing portion; and The background sound generating means generates the background sound signal based on the sound signal acquired as the ambient sound when a ratio between an amplitude of a signal portion of a voice section in the sound signal acquired as the ambient sound and an amplitude of a signal portion other than the voice section is equal to or smaller than a predetermined value. A voice compensation device characterized by:
9. 1. An audio compensation system for compensating for a missing portion of an audio signal, comprising: a harmonic sound generating means for generating a harmonic sound signal, which is an acoustic signal having a specific frequency component, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; a background sound generating means for generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; a compensation signal generating means for generating a compensation signal from the harmonic sound signal and the background sound signal; a compensation signal inserting means for inserting the compensation signal into the missing portion; and The background sound generating means generates the background sound signal based on the sound signal acquired as the ambient sound when a ratio between an amplitude of a signal portion of a voice section in the sound signal acquired as the ambient sound and an amplitude of a signal portion other than the voice section is equal to or smaller than a predetermined value. A voice compensation system comprising:
10. 1. An audio compensation method for compensating for a missing portion of an audio signal, comprising: generating a harmonic sound signal, which is an acoustic signal having specific frequency components, based on a signal portion preceding the missing portion in the audio signal or based on a preset frequency condition; A step of generating a background sound signal, which is an audio signal having an amplitude spectrum that peaks in a predetermined low frequency band, based on an audio signal prepared in advance or acquired as ambient sound, or based on a signal portion preceding the missing portion in the audio signal; generating a compensation signal from the harmonic sound signal and the background sound signal; inserting the compensation signal into the missing portion; and In the step of generating the background sound signal, when a ratio between an amplitude of a signal portion of a voice section in the sound signal acquired as the ambient sound and an amplitude of a signal portion other than the voice section is equal to or less than a predetermined value, the background sound signal is generated based on the sound signal acquired as the ambient sound.
10. A computer-implemented method for speech compensation, comprising:
Citation Information
Patent Citations
Voice encoding / decoding system with background sound reproducing function
JP1990288520A
Voice reproducer
JP1998111699A
Burst frame error handling
JP2017525985A
Concealing Audio Artifacts
US20110082575A1