Time window generation device, method and program
The time window generation device and method adaptively generate appropriate windows for acoustic signals, enhancing frequency component separation by balancing frequency resolution and dynamic range.
Patent Information
- Application Number
- JP2023578328
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-07
- Publication Date
- 2025-11-26
- Estimated Expiration
- 2042-02-07
AI Technical Summary
Existing acoustic signal processing systems lack the ability to generate an appropriate time window that balances frequency resolution and dynamic range, leading to suboptimal separation of adjacent frequency components.
A time window generation device and method that includes signal clipping, analysis and synthesis window generation, frequency and time domain conversion, and a learning unit to adaptively generate and refine window parameters based on ground truth data.
Enables the generation of tailored time windows that enhance frequency component separation, improving the accuracy of acoustic signal processing.
Smart Images

Figure 0007775899000001 
Figure 0007775899000002 
Figure 0007775899000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for processing sound signals such as voice. [Background technology]
[0002] A method using short-time Fourier transform is used to analyze acoustic signals in real time in terms of time frequency. In this method, a time window is used to extract a signal of a certain length and treat it as a periodic signal. The frequency resolution and dynamic range of the time window are determined according to its shape. Here, there is a trade-off between frequency resolution and dynamic range. Therefore, improving the separation performance of adjacent frequency components can result in overlooking small components.
[0003] Since there is no time window that is effective for all signals, it is necessary to use a time window that is appropriate for the situation.
[0004] In conventional acoustic signal processing systems, a specific window function is fixed and used, or a plurality of window functions prepared in advance are switched and used (for example, see Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Yuma Koizumi, "Sound source enhancement and phase control based on deep learning," Journal of the Acoustical Society of Japan, Vol. 75, No. 3, pp. 156-163, 2019 Summary of the Invention [Problem to be solved by the invention]
[0006] However, there has been no technology available to date for generating an appropriate time window.
[0007] An object of the present invention is to provide a time window generation device, method, and program for generating an appropriate time window. [Means for solving the problem]
[0008] A time window generation device according to one aspect of the present invention includes a signal clipping unit that generates a clipped signal by clipping a sound signal to a predetermined length; an analysis window generation unit that generates an analysis window using the clipped signal and an analysis window model determined by analysis window model parameters; a synthesis window generation unit that generates a synthesis window using at least the clipped signal or using the analysis window; a frequency domain conversion unit that converts the clipped signal into a frequency domain using the analysis window to generate a frequency domain signal; a signal processing unit that performs predetermined processing on the frequency domain signal to generate a processed frequency domain signal; a time domain conversion unit that converts the processed frequency domain signal into a time domain using the synthesis window to generate a time domain signal; and a learning unit that learns at least the analysis window model parameters using the time domain signal and ground truth data corresponding to the time domain signal. The predetermined processing is a speech signal enhancement processing, and the correct answer data is an original speech signal without noise. [Effects of the Invention]
[0009] An appropriate time window can be generated. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of the functional configuration of a time window generating device. [Figure 2] FIG. 2 is a diagram showing an example of a processing procedure of the time window generation method. [Figure 3] FIG. 3 is a diagram illustrating an example of a functional configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0011] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described in detail below with reference to the accompanying drawings, in which like reference numerals are used to designate like components having the same functions, and redundant description will be omitted.
[0012] [Time window generation device and method] As shown in FIG. 1, the time window generating device includes, for example, a signal extracting unit 1, an analysis window generating unit 2, a synthesis window generating unit 3, a frequency domain transforming unit 4, a signal processing unit 5, a time domain transforming unit 6, and a learning unit .
[0013] The time window generation method is realized, for example, by each component of the time window generation device performing the processes from step S1 to step S7 described below and shown in FIG.
[0014] Here, the time window means at least one of an analysis window and a synthesis window.
[0015] Each component of the time window generating device will be described below.
[0016] <Signal extraction section 1> A sound signal is input to the signal extractor 1.
[0017] The signal clipping unit 1 generates a clipped signal by clipping the sound signal to a predetermined length (step S1).
[0018] The generated extracted signal is output to the analysis window generating unit 2 and the synthesis window generating unit 3.
[0019] <Analysis window generation unit 2> The extracted signal is input to the analysis window generation unit 2. Furthermore, the analysis window generation unit 2 also receives analysis window model parameters generated by the learning unit 7, which will be described later.
[0020] The analysis window generator 2 generates an analysis window using the extracted signal and an analysis window model determined by the analysis window model parameters (step S2). The generated analysis window is output to the frequency domain transformer 4.
[0021] The analysis window model parameters are, for example, analysis window model parameters generated by the learning unit 7, which will be described later. If this is the first processing by the analysis window generation unit 2 and there are no analysis window model parameters generated by the learning unit 7, the analysis window generation unit 2 uses predetermined analysis window model parameters. In this case, the analysis window generation unit 2 may generate and output a predetermined analysis window.
[0022] <Synthetic window generator 3> The extracted signal is input to the synthesis window generation unit 3. Furthermore, the synthesis window generation unit 3 is input with synthesis window model parameters generated by the learning unit 7, which will be described later.
[0023] The synthesis window generator 3 generates a synthesis window using the extracted signal and a synthesis window model determined by the synthesis window model parameters (step S3). The generated synthesis window is output to the time domain transformer 6.
[0024] The synthesis window model parameters are, for example, synthesis window model parameters generated by the learning unit 7, which will be described later. If this is the first processing by the synthesis window generation unit 3 and there are no synthesis window model parameters generated by the learning unit 7, the synthesis window generation unit 3 uses predetermined synthesis window model parameters. In this case, the synthesis window generation unit 3 may generate and output a predetermined synthesis window.
[0025] <Frequency domain transform section 4> The frequency domain transform unit 4 receives the extracted signal and the analysis window.
[0026] The frequency domain transform unit 4 transforms the extracted signal into the frequency domain using an analysis window to generate a frequency domain signal (step S4). The generated frequency domain signal is output to the signal processing unit 5.
[0027] The frequency domain transform unit 4 performs a transform into the frequency domain using a technique such as a short-time Fourier transform.
[0028] <Signal processing unit 5> The signal processing unit 5 receives a frequency domain signal.
[0029] The signal processing unit 5 performs a predetermined process on the frequency domain signal to generate a processed frequency domain signal (step S5). The generated processed frequency domain signal is output to the time domain transform unit 6.
[0030] Examples of the predetermined processing include at least one of processing for enhancing a predetermined signal such as a voice (in other words, processing for enhancing a voice signal, processing for suppressing noise), and classification processing for classifying noise or the like.
[0031] When the predetermined processing is speech signal enhancement processing, the signal processing unit 5 estimates a speech enhancement filter using, for example, the frequency domain signal. Then, the signal processing unit 5 multiplies the estimated speech enhancement filter by the frequency domain signal to generate a frequency domain signal in which the speech is enhanced. This frequency domain signal is an example of a processed frequency domain signal.
[0032] The predetermined processing may include noise classification processing. In this case, the signal processing unit 5 uses the frequency domain signal to estimate the noise contained in the frequency domain signal. The estimated noise label is output to the learning unit 7.
[0033] <Time domain transform unit 6> The time domain transform unit 6 receives the processed frequency domain signal and the synthesis window.
[0034] The time domain transform unit 6 generates a time domain signal by transforming the processed frequency domain signal into the time domain using a synthesis window (step S6). The generated time domain signal is output to the learning unit 7. The generated time domain signal may be output from the time window generating device as a result of predetermined processing by the signal processing unit 5.
[0035] The frequency domain transform unit 4 performs a transform into the time domain using a technique such as an inverse short-time Fourier transform.
[0036] <Study Section 7> A time domain signal is input to the learning unit 7. Also, correct answer data corresponding to the time domain signal is input to the learning unit 7.
[0037] The learning unit 7 uses the time domain signal and the correct data corresponding to the time domain signal to learn the analysis window parameters and the synthesis window parameters (step S7).
[0038] The analysis window parameters and the synthesis window parameters are learned by, for example, gradient descent.
[0039] When the predetermined processing in the signal processing unit 5 is speech signal enhancement processing, the original sound signal without noise becomes the correct data. In this case, the analysis window parameters and the synthesis window parameters are learned using a gradient method or the like so as to minimize the mean squared error obtained by averaging the squares of the differences between the time domain signal and the original sound signal.
[0040] When the predetermined processing in the signal processing unit 5 includes noise classification processing, the true noise label is input as supervised data to the learning unit 7. When the predetermined processing in the signal processing unit 5 includes noise classification processing, the noise label estimated by the predetermined processing in the signal processing unit 5 is further input to the learning unit 7. In this case, the learning unit 7 may learn the analysis window parameter and the synthesis window parameter using the true noise label and the estimated noise label in addition to the time-domain signal and the supervised data corresponding to the time-domain signal.
[0041] The processes from step S1 to step S7 described above may be repeated as appropriate.
[0042] The learning unit 7 learns the analysis window parameters and synthesis window parameters based on the time domain signal, which is a signal that has been subjected to predetermined processing in the signal processing unit 5 and converted into the time domain. For this reason, it can be said that the analysis window parameters and synthesis window parameters are learned taking into consideration the predetermined processing in the signal processing unit 5, which is processing subsequent to the analysis window generation unit 2 and the synthesis window generation unit 3. In this way, by learning the analysis window parameters and synthesis window parameters taking into consideration the processing subsequent to the analysis window parameters, it is possible to generate time windows more appropriately than before.
[0043] [Variations] The above describes the embodiments of the present invention, but the specific configuration is not limited to these embodiments, and it goes without saying that even if design changes are made as appropriate within the scope of the present invention, they are still included in the present invention.
[0044] The various processes described in the embodiments may not only be executed in chronological order according to the order described, but may also be executed in parallel or individually depending on the processing capabilities of the devices executing the processes or as necessary.
[0045] For example, data may be exchanged directly between the components of the time window generating device, or may be exchanged via a storage unit (not shown).
[0046] The synthesis window can be generated from the analysis window. Therefore, the synthesis window generation unit 3 may generate the synthesis window using the analysis window generated by the analysis window generation unit 2. That is, the synthesis window generation unit 3 may generate the synthesis window using at least the extracted signal or the analysis window.
[0047] In this case, the learning unit 7 does not need to learn the synthesis window model parameters, but only needs to learn at least the analysis window model parameters using the time domain signal and the correct data corresponding to the time domain signal.
[0048] Not only the processing in the analysis window generation unit 2 and the synthesis window generation unit 3, but also the predetermined processing in the signal processing unit 5 may be implemented using deep learning.
[0049] That is, the predetermined processing in the signal processing unit 5 may be processing to generate a processed frequency domain signal using a model determined by the frequency domain signal and the model parameters. In this case, the learning unit 7 may further learn the model parameters using at least the time domain signal and the ground truth data corresponding to the time domain signal.
[0050] Furthermore, the predetermined processing in the signal processing unit 5 may include processing for estimating noise labels using a model determined by the frequency domain signal and the noise estimation model parameters. In this case, the learning unit 7 may further learn the noise estimation model parameters using the true noise labels input to the learning unit 7 and the noise labels estimated by the signal processing unit 5 in addition to the time domain signal and the ground truth data corresponding to the time domain signal.
[0051] [Programs, recording media] The processing of each unit of each of the above-mentioned devices may be realized by a computer, in which case the processing content of the functions that each device should have is described by a program. Then, by loading this program into storage unit 1020 of computer 1000 shown in Fig. 3 and operating arithmetic processing unit 1010, input unit 1030, output unit 1040, display unit 1060, etc., various processing functions of each of the above-mentioned devices are realized on the computer.
[0052] The program describing the processing contents can be recorded on a computer-readable recording medium, such as a non-transitory recording medium, specifically a magnetic recording device, an optical disk, or the like.
[0053] The program may be distributed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to another computer via a network, thereby distributing the program.
[0054] A computer that executes such a program, for example, first stores the program recorded on a portable recording medium or transferred from a server computer in its own non-transitory storage device, auxiliary storage unit 1050. Then, when executing a process, the computer loads the program stored in auxiliary storage unit 1050, its own non-transitory storage device, into storage unit 1020 and executes processing in accordance with the loaded program. Alternatively, as another form of execution of this program, the computer may load the program directly from a portable recording medium into storage unit 1020 and execute processing in accordance with the program. Furthermore, each time a program is transferred from a server computer to this computer, the computer may execute processing in accordance with the received program. Alternatively, the server computer may not transfer the program to this computer, but may instead execute the processing function by issuing an execution instruction and obtaining the results, thereby executing the above-described processing through a so-called ASP (Application Service Provider) type service. Note that the program in this embodiment includes information used for processing by a computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that define computer processing).
[0055] In this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing may be realized by hardware. For example, the signal extractor 1, the analysis window generator 2, the synthesis window generator 3, the frequency domain transformer 4, the signal processor 5, the time domain transformer 6, and the learning unit 7 may be configured by a processing circuit.
[0056] It goes without saying that other modifications are possible without departing from the spirit of the present invention.
Claims
1. a signal cutout unit that cuts out a sound signal to a predetermined length to generate a cutout signal; an analysis window generation unit that generates an analysis window using the extracted signal and an analysis window model determined by analysis window model parameters; a synthesis window generator that generates a synthesis window using at least the extracted signal or the analysis window; a frequency domain transform unit that generates a frequency domain signal by transforming the extracted signal into a frequency domain using the analysis window; a signal processing unit that performs predetermined processing on the frequency domain signal to generate a processed frequency domain signal; a time domain transform unit that generates a time domain signal by transforming the processed frequency domain signal into a time domain signal using the synthesis window; a learning unit that learns at least the analysis window model parameters using the time domain signal and ground truth data corresponding to the time domain signal; Including, the predetermined processing is a speech signal enhancement processing, The correct answer data is a noise-free original sound signal, the synthesis window generation unit generates a synthesis window using the extracted signal and a synthesis window model determined by a synthesis window model parameter; the learning unit further learns the synthesis window model parameters using the time-domain signal and ground truth data corresponding to the time-domain signal. Time window generator.
2. 2. The time window generating device of claim 1, the predetermined processing is processing for generating a processed frequency domain signal using the frequency domain signal and a model determined by model parameters; the learning unit further learns the model parameters by using at least the time-domain signal and ground truth data corresponding to the time-domain signal. Time window generator.
3. a signal cutting step in which a signal cutting unit cuts out a sound signal to a predetermined length to generate a cut-out signal; an analysis window generation step in which an analysis window generation unit generates an analysis window using the extracted signal and an analysis window model determined by analysis window model parameters; a synthesis window generation step in which a synthesis window generation unit generates a synthesis window using at least the extracted signal or the analysis window; a frequency domain transform step in which a frequency domain transform unit transforms the extracted signal into a frequency domain using the analysis window to generate a frequency domain signal; a signal processing step in which a signal processing unit performs predetermined processing on the frequency domain signal to generate a processed frequency domain signal; a time domain transform step in which a time domain transform unit transforms the processed frequency domain signal into a time domain signal using the synthesis window; a learning step in which a learning unit learns at least the analysis window model parameters using the time domain signal and ground truth data corresponding to the time domain signal; Including, the predetermined processing is a speech signal enhancement processing, The correct answer data is a noise-free original sound signal, In the synthesis window generating step, a synthesis window is generated using the extracted signal and a synthesis window model determined by a synthesis window model parameter; In the learning step, the synthesis window model parameters are further learned using the time-domain signal and ground truth data corresponding to the time-domain signal. Time window generation method.
4. A program for causing a computer to function as each unit of the time window generating device according to claim 1 or 2.
Citation Information
Patent Citations
Background noise eliminating device
JP1998003297A
Voice processing device, voice processing method, and computer program for voice processing
JP2015049354A
Method of operation of a hearing aid system and hearing aid system
JP2016537891A
Systems and methods for source signal separation
JP2017102488A
Sound source enhancement device, sound source enhancement learning device, sound source enhancement method, program
JP2020030373A