Method and system for training neural network

Through the end-to-end multi-task denoising framework, combining SDR and PESQ loss functions, the total loss function is optimized using convolutional neural networks, which solves the metric and spectrum mismatch problems in spectrum masking estimation, and improves the SDR and PESQ performance of speech denoising.

CN111435462BActive Publication Date: 2025-08-22SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911114742.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-25
Filing Date
2019-11-14
Publication Date
2025-08-22
Estimated Expiration
2039-11-14

AI Technical Summary

Technical Problem

Existing spectral masking estimation methods have metric mismatch and spectrum mismatch problems in signal distortion ratio (SDR) and perceived speech quality evaluation (PESQ) optimization, resulting in performance losses.

Method used

Using an end-to-end multi-task denoising framework, combining the loss function of signal distortion ratio (SDR) and perceived speech quality evaluation (PESQ), trained using a convolutional neural network through short-term Fourier transform and inverse transform, optimized the total loss function to maximize SDR and PESQ simultaneously.

Benefits of technology

It effectively avoids measurement mismatch and spectrum mismatch, significantly improves the performance of PESQ and SDR, and achieves better voice denoising effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111435462B_ABST
    Figure CN111435462B_ABST
Patent Text Reader

Abstract

Disclosed herein is a method and system for training a neural network. According to one embodiment, the method includes receiving a noisy signal, generating a denoised output signal, determining a signal-to-distortion ratio (SDR) loss function based on the denoised output signal, determining a perceptual speech quality assessment (PESQ) loss function based on the denoised output signal, and optimizing a total loss function based on the perceptual speech quality assessment loss function and the signal-to-distortion ratio loss function.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] [priority]

[0002] This application is based on and claims priority to the U.S. provisional patent application filed in the U.S. Patent and Trademark Office on January 11, 2019 and granted serial number 62 / 791,421, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates generally to signal processing and, more particularly, to signal-to-distortion ratio (SDR) and perceptual speech quality (PESQ) optimization. Background Art

[0004] Recently, supervised learning using deep neural networks has achieved substantial improvements in speech enhancement. A key difference from typical statistical approaches is that no prior assumptions about the signal model are necessary. For example, Wiener filters typically assume a Gaussian distribution for speech or noise models, which is often incorrect in real-world settings. In contrast, neural networks learn to denoise speech simply by referring to the mapping from noisy speech to clean speech in the training data.

[0005] Spectra mask estimation involves predicting the time-frequency spectra mask (i.e., the ratio between the clean spectrum and the noisy spectrum). An ideal binary mask (IBM) has been previously proposed for training labels, where the IBM is one or zero depending on the corresponding signal-to-noise ratio (SNR). Ideal ratio mask (IRM) and ideal magnitude mask (IAM) provide soft mask labels to overcome the coarse mapping of IBM. Both methods show improvements over IBM due to better label resolution. In addition, a phase sensitive mask (PSM) that takes into account clean phase and noisy phase has been previously proposed. PSM does not compensate for noisy phase, but by referring to the ratio of clean phase to noisy phase, PSM provides a better spectra amplitude label for mask estimation.

[0006] However, this approach has two key issues: metric mismatch and spectral mismatch. Spectral mask estimation typically minimizes the mean square error (MSE) between the clean spectrum amplitude and the estimated spectrum amplitude, which is suboptimal in terms of maximizing the signal distortion ratio (SDR) or perceptual evaluation of speech quality (PESQ) due to metric mismatch. For example, it is often observed that although the spectral mean square error is reduced, the SDR or PESQ often deteriorates. The second spectral mismatch problem arises from the estimation in the spectral domain. Generally speaking, due to the short-time Fourier transform (STFT) and inverse short-time Fourier transform (ISTFT) operations, any arbitrary modification of the spectral signal cannot be fully recovered, which is also known as STFT inconsistency. For example, the denoised spectrum generally does not match the spectral amplitude of the recovered waveform. Therefore, the denoised spectrum amplitude cannot be fully reflected in the reconstructed output, which may lead to substantial performance loss. Summary of the Invention

[0007] According to one embodiment, a method for training a neural network includes: receiving a noisy signal; generating a denoised output signal; determining a signal-to-distortion ratio (SDR) loss function based on the denoised output signal; determining a perceptual speech quality assessment (PESQ) loss function based on the denoised output signal; and optimizing a total loss function based on the PESQ loss function and the SDR loss function.

[0008] According to one embodiment, a system for training a neural network includes a memory and a processor, wherein the processor is configured to: receive a noisy signal; generate a denoised output signal; determine an SDR loss function based on the denoised output signal; determine a PESQ loss function based on the denoised output signal; and optimize a total loss function based on the PESQ loss function and the SDR loss function.

[0009] According to one embodiment, a method of training a neural network includes receiving a noisy signal; generating a denoised output signal; and determining a PESQ loss function based on the denoised output signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent upon reading the following detailed description in conjunction with the accompanying drawings, in which:

[0011] Figure 1 is a diagram of a denoising system according to an embodiment.

[0012] Figure 2 is a diagram illustrating the amplitude spectrum optimization problem.

[0013] Figure 3 is a flowchart of a method for training a neural network to maximize SDR and PESQ according to an embodiment.

[0014] Figure 4 is a diagram of an optimized network system according to an embodiment.

[0015] Figure 5 is a diagram of a system for determining a PESQ loss function, according to an embodiment.

[0016] Figure 6 is a block diagram of electronic devices in a network environment according to one embodiment. DETAILED DESCRIPTION

[0017] Hereinafter, embodiments of the present disclosure are described in detail with reference to the accompanying drawings. It should be noted that the same elements will be indicated by the same reference numerals even though they are shown in different drawings. In the following description, specific details such as detailed configuration and components are provided only to help fully understand the embodiments of the present disclosure. Therefore, it will be obvious to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. In addition, for the sake of clarity and brevity, descriptions of well-known functions and configurations are omitted. The terms described below are defined in view of the functions in the present disclosure and may vary depending on the user, the user's intention or custom. Therefore, the definitions of these terms should be determined based on the content of this specification throughout.

[0018] The present disclosure may have various modifications and various embodiments, and the embodiments thereof are described in detail below with reference to the accompanying drawings. However, it should be understood that the present disclosure is not limited to the embodiments described, but includes all modifications, equivalent forms and alternative forms within the scope of the present disclosure.

[0019] Although terms including ordinal numbers such as "first" and "second" may be used to describe various elements, structural elements are not limited by these terms. These terms are only used to distinguish between individual elements. For example, without departing from the scope of this disclosure, a "first structural element" may be referred to as a "second structural element." Similarly, a "second structural element" may also be referred to as a "first structural element." The term "and / or" used herein includes any and all combinations of one or more associated items.

[0020] The terms used herein are only used to illustrate various embodiments of the present disclosure and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In the present disclosure, it should be understood that the term "include" or "have" indicates the presence of a feature, number, step, operation, structural element, component or combination thereof, and does not exclude the presence or possibility of adding one or more other features, numbers, steps, operations, structural elements, components or combinations thereof.

[0021] Unless otherwise defined, all terms used herein have the same meaning as understood by those skilled in the art to which the present disclosure pertains. For example, terms defined in commonly used dictionaries should be interpreted as having the same meaning as in the context of the relevant technical field, and should not be interpreted as having an idealized or overly formal meaning unless clearly defined in the present disclosure.

[0022] The electronic device according to one embodiment may be one of various types of electronic devices. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer, a portable multimedia device, a portable medical device, a camera, a wearable device, or a household appliance. According to one embodiment of the present disclosure, the electronic device is not limited to the above-mentioned electronic devices.

[0023] The terms used in this disclosure are not intended to limit the disclosure but are intended to include various changes, equivalents or alternative forms to the corresponding embodiments. With respect to the description of the drawings, similar reference numerals may be used to refer to similar or related elements. Unless the relevant context clearly indicates otherwise, the singular form of a noun corresponding to an item may include one or more things. Each of the phrases used herein, such as "A or B", "at least one of A and B", "at least one of A or B", "A, B or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include all possible combinations of the items enumerated together with the corresponding one of the phrases. For example, "the first (1)" used herein stTerms such as “first”, “second”, etc. may be used to distinguish a corresponding component from another component and are not intended to limit the components in other respects (e.g., importance or order). The present invention intends that if an element (e.g., a first element) is referred to as being “coupled”, “coupled to”, “connected to” or “connected to” another element (e.g., a second element), with or without the term “operably” or “communicatively”, it means that the element may be coupled to the other element directly (e.g., in a wired manner), wirelessly, or through a third element.

[0024] The term "module" as used herein may include units implemented in the form of hardware, software, or firmware, and may be used interchangeably with other terms such as "logic," "logic block," "component," and "circuitry." A module may be a single integral component suitable for performing one or more functions, or the smallest unit or component of the single integral component. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0025] The present system and method provide end-to-end multi-task denoising for joint signal-to-distortion ratio (SDR) and perceptual speech quality (PESQ) optimization. The framework includes a loss function that is fully correlated with PESQ and SDR, which effectively avoids the metric mismatch problem. The loss function can be divided into two terms: SDR loss and PESQ loss. The SDR loss uses a scale-invariant SDR as the loss function, as shown in Equation (1).

[0026]

[0027] The PESQ loss is designed as an approximate symmetric and asymmetric disturbance of PESQ. During training, the approximation term is minimized to optimize PESQ.

[0028] In addition, after the inverse short-time Fourier transform (ISTFT), supervised learning is performed on the denoised time-domain speech. Unlike spectral mask estimation, loss minimization is not performed on the spectral domain. Therefore, the framework does not have spectrum mismatch. The present system and method for end-to-end multi-task denoising significantly improves PESQ and SDR. The present system and method performs optimization after the ISTFT domain to use the scale-invariant SDR metric as the loss function and the modified PESQ metric combined with the SDR to solve the spectrum mismatch, so as to jointly optimize PESQ and SDR.

[0029] Figure 1 1 is a diagram of a denoising system according to an embodiment. System 100 includes a short-time Fourier transform (STFT) block 102, an absolute value (ABS) block 104, a phase extraction block 106, a denoiser block 108 including a neural network 110, and an I S T F T block 112. u (n) is modeled as equation (2):

[0030] y u (n) = x u (n)+n u (n) (2)

[0031] Where u is the utterance index, n is the time index, and x u (n) is clean speech, and n u (n) is the noise signal. Then the noisy input signal y u (n) Grouping to produce As shown in equation (3):

[0032]

[0033] Where m is the frame index, Δ is the size of the frame shift, K is the size of each frame, and w(k) is the window function. For example, K is 1024 and Δ is 256, so that each frame has 75% overlap with the next frame.

[0034] After STFT 102, is fed into two separate paths. For the upper path, the The magnitude of is obtained and passed to the denoiser block 108 for spectral amplitude denoising. Phase, and use The phase of the complex spectrum with the denoised spectrum amplitude is synthesized After ISTFT112, the reconstructed time domain is denoised and output ISTFT 112 is composed of inverse fast Fourier transform (IFFT), windowing, and overlap-add operations.

[0035] Figure 2 2 is a diagram illustrating an amplitude spectrum optimization problem. System 200 includes an STFT block 204, a denoiser block 206, a Griffin-Lim (GL) ISTFT block 208, and an STFT block 210. Generally, there is an amplitude spectrum mismatch 212 after the ISTFT block 208. The denoiser output Not with Match, where is the reconstructed signal Use GL ISTFT to find To minimize the MSE of the two complex spectra. Due to this STFT inconsistency, the estimated Unable to fully reflect the real spectrum middle.

[0036] Furthermore, the MSE of the amplitude spectrum is not optimal in SDR and PESQ because MSE treats all time-frequency bins equally. SDR and PESQ apply nonlinear mappings to the time or spectrum signals to produce unequally weighted error averages. Therefore, it is necessary to use the correct loss function to minimize SDR and PESQ.

[0037] Figure 3 is a flow chart 300 of a method for training a neural network to maximize SDR and PESQ according to an embodiment. At 302, the system receives a noisy signal, and at 304, the system generates a denoised output signal, as described above with reference to Figure 1 and Figure 2 As stated.

[0038] At 306 , the system determines an SDR loss function based on the denoised output signal. Figure 44 is a diagram of an optimization network system 400 according to an embodiment. The system 400 is based on a convolutional neural network bi-directional long-short term memory (CNN-BLSTM) network 404. A CNN-BLSTM is an example of a denoiser network and other types of networks, such as CNN-based denoising autoencoders. The CNN-BLSTM 404 includes three convolutional layers 406, three bidirectional LSTMs (BLSTMs) 408 (where the double-headed arrows represent bidirectional scans in the BLSTM 408), and a fully-connected (FC) layer 409. Amplitude of a noisy spectrum with 513 frequency bins and an 11-frame context window 402 is fed into a convolutional layer 406 with a kernel size of 5x5. Dilated convolutions are applied to the second and third layers with dilation rates of 2 and 4. Dilated convolutions are applied only to the frequency dimension because the BLSTM 408 will learn temporal correlations. The output 410 of the CNN-BLSTM is also a time-frequency mask. It can be called the product of the CNN-BLSTM masked output and the noisy spectrum amplitude 402.

[0039] The estimated spectrum is converted by GL ISTFT 412 Transformed into As shown in equations (4) and (5).

[0040]

[0041]

[0042] The system 400 generates a denoised output 414, which is then used to determine an SDR loss function 416. is a time domain signal, SDR can be directly optimized using the SDR loss function, as shown in equations (6) and (7):

[0043]

[0044]

[0045] These include α u As part of the training factor. SDRis the average SDR of the batch utterances, so there is no metric mismatch with the SDR metric. The optimization does not have the spectrum mismatch problem because the optimization is performed after the GL transformation. As long as the CNN-BLSTM 404 is well trained so that L SDR Minimize, and the mismatch between the denoiser spectral output and the STFT of the reconstructed signal becomes insignificant.

[0046] At 308, the system determines a PESQ loss function 418 based on the denoised output signal 414. The SDR loss function 416 and the PESQ loss function 418 will be used later to optimize the total loss function 420. The end-to-end training maximizes the SDR by reconstructing the time domain signal from the GL transform. Although the loss function L SDR It may be optimal in terms of maximizing the SDR metric, but since the frame disturbance metric defined in PESQ may not necessarily decrease with lower L SDR Therefore, PESQ still suffers from a metric mismatch. For example, if SDR improves significantly in the high-frequency region but degrades slightly in the lower-frequency portion, the overall SDR may enhance the signal, but PESQ may degrade due to the higher weighting of the lower-frequency bands. Therefore, as with SDR optimization, it is best to directly maximize the PESQ metric to avoid metric mismatch.

[0047] Figure 5 is a diagram of a system 500 for determining a PESQ loss function according to an embodiment. Relative to a conventional PESQ loss function determination system, system 500 has several modifications to implement back-propagation and remove unnecessary operations to reduce complexity. First, because the time evolution of infinite impulse response (IIR) filters is so deep (more than hundreds of thousands), IIR filters are removed, and therefore back-propagation is not feasible. Second, because delay-aligned data is prepared for training, the delay adjustment routine is removed. Finally, bad-interval iterations are removed. PESQ improves metric calculations by detecting bad intervals of frames and updating the metric within these periods. Removing this operation has no significant impact on PESQ as long as the training clean data and noisy data pairs are time-aligned.

[0048] The system 500 receives a denoised signal 502 and a noisy signal 504 and performs level alignment 506 on the denoised signal 502 and the noisy signal 504. The average power of the denoised signal and the noisy signal within the range of 300 Hz and 3 kHz is aligned to a predefined value of 10 7The IIR filter uses the frequency response of the intermediate reference system (IRS) receiver characteristics to model the handset listening environment.

[0049] The system 500 performs an STFT 508 on the level aligned signals on both paths of the system 500 and then applies a Bark spectrum frequency analysis 510 to the linear spectrum input signals on both paths of the system 500. The Bark spectrum analysis 510 finds the average of the linear scale frequency bins according to the Bark scale mapping. Higher frequency bins are averaged over a larger number of bins, which effectively gives them a lower weight. The mapped Bark spectrum power can be formulated as shown in Equation (8):

[0050]

[0051] Among them I i is the start of the linear frequency band number of the i-th Bark spectrum, is the STFT spectrum magnitude of the denoised signal 502, and is the i-th Barker spectrum power of the denoised signal 502 . is the Bark spectral power of the noisy signal and can also be found in equation (8).

[0052] The system 500 then performs time-frequency equalization (TF Equal) 512 on the Barker spectrum power on the two paths of the system 500. Each Barker spectrum of the denoised signal 502 is first compensated by the average power ratio between the denoised Barker spectrum and the noisy Barker spectrum, as shown in equation (9):

[0053]

[0054] in And c is a constant. and is the silence mask, and they become 1 only when the corresponding Barker spectrum power exceeds the threshold. After frequency equalization, as shown in equations (10), (11), and (12), the short-term gain variation of the noisy Barker spectrum is compensated for each frame:

[0055]

[0056]

[0057]

[0058] in And c2 is a constant.

[0059] The system 500 then performs loudness mapping 514 on the compensated signals on both paths of the system 500. The power density is converted to the Sone loudness scale using Zwicker's law, as shown in equation (13):

[0060]

[0061] Among them, P 0,i is the absolute hearing threshold, S i is the loudness scaling factor, and r is the Zwick power and x can be c (denoised signal) or n (noisy signal).

[0062] The system 500 then performs disturbance processing 516 on the mapped signals on both paths of the system 500. The raw disturbance metric is the difference between the denoised loudness density and the noisy loudness density, and is then further processed as shown in Equations (14) and (15).

[0063]

[0064]

[0065] If the absolute difference between the denoised loudness density and the noisy loudness density is less than 0.25 of the minimum of the two densities, the original perturbation becomes zero. The symmetric frame perturbation is then calculated using the L2 norm operation, as shown in Equation (16):

[0066]

[0067] where w i is the predefined weight of the Barker spectrum band. For asymmetric frame perturbation, the original perturbation is weighted by the ratio between the noisy spectrum power with saturation and threshold and the denoised spectrum power, as shown in Equations (17), (18), and (19).

[0068]

[0069]

[0070]

[0071] The system 500 then determines the PESQ loss function L by aggregation 518 of the perturbations. PESQ 520. L PESQ 520 can be found by averaging the two-step frame perturbations, as shown in Equations (20), (21), (22), and (23):

[0072]

[0073]

[0074]

[0075]

[0076] in,

[0077] L PESQ 520 can be determined as shown in equation (24).

[0078] L PESQ =4.5-0.1d sym -0.0309d asym (twenty four)

[0079] At 310, the system optimizes the total loss function based on the SDR loss function and the PESQ loss function. PESQ With L SDR To minimize, the loss function combines the two, as shown in equation (25):

[0080]

[0081] Where α is a hyperparameter. Table 1 shows the performance comparison between different schemes, where SDR-PESQ is the denoising method disclosed in this paper.

[0082] Table 1

[0083]

[0084] Figure 6 FIG is a block diagram of an electronic device 601 in a network environment 600 according to one embodiment. Figure 6, the electronic device 601 in the network environment 600 can communicate with the electronic device 602 through a first network 698 (e.g., a short-range wireless communication network), or can communicate with the electronic device 604 or the server 608 through a second network 699 (e.g., a long-range wireless communication network). The electronic device 601 can communicate with the electronic device 604 through the server 608. The electronic device 601 may include a processor 620, a memory 630, an input device 650, an audio output device 655, a display device 660, an audio module 670, a sensor module 676, an interface 677, a haptic module 679, a camera module 680, a power management module 688, a battery 689, a communication module 690, a subscriber identification module (SIM) 696, or an antenna module 697. In one embodiment, at least one of these components (e.g., the display device 660 or the camera module 680) may be omitted from the electronic device 601, or one or more other components may be added to the electronic device 601. In one embodiment, some of the components described above may be implemented as a single integrated circuit (IC). For example, the sensor module 676 (e.g., a fingerprint sensor, an iris sensor, or an illuminance sensor) may be embedded in the display device 660 (e.g., a display).

[0085] The processor 620 may execute, for example, software (e.g., program 640) to control at least one other component (e.g., hardware component or software component) of the electronic device 601 coupled to the processor 620, and may perform various data processing or calculations. As at least part of the data processing or calculation, the processor 620 may load commands or data received from another component (e.g., sensor module 676 or communication module 690) into the volatile memory 632, process the commands or data stored in the volatile memory 632, and store the resulting data in the non-volatile memory 634. The processor 620 may include a main processor 621 (e.g., a central processing unit (CPU) or an application processor (AP)) and an auxiliary processor 623 (e.g., a graphics processing unit (GPU), an image signal processor (ISP), a sensor hub processor (SHP), or a communication processor (CP)) that can operate independently of or in conjunction with the main processor 621. Additionally or alternatively, the auxiliary processor 623 may be adapted to consume less power than the main processor 621, or to perform specific functions. The auxiliary processor 623 may be implemented separately from the main processor 621 or as part of the main processor 621.

[0086] When the main processor 621 is in an inactive state (e.g., a sleep state), the auxiliary processor 623 may replace the main processor 621 to control at least some of the functions or states related to at least one of the components of the electronic device 601 (e.g., the display device 660, the sensor module 676, or the communication module 690); or when the main processor 621 is in an active state (e.g., executing an application), the auxiliary processor 623 may control at least some of the above functions or states together with the main processor 621. According to one embodiment, the auxiliary processor 623 (e.g., an image signal processor or a communication processor) may be implemented as part of another component (e.g., the camera module 680 or the communication module 690) that is functionally related to the auxiliary processor 623.

[0087] The memory 630 may store various data used by at least one component of the electronic device 601 (e.g., the processor 620 or the sensor module 676). The various data may include, for example, software (e.g., the program 640) and input data or output data for commands associated with the software. The memory 630 may include a volatile memory 632 or a non-volatile memory 634.

[0088] The program 640 may be stored in the memory 630 as software, and may include, for example, an operating system (OS) 642 , middleware 644 , or an application 646 .

[0089] The input device 650 may receive commands or data from outside the electronic device 601 (eg, a user) to be used by other components of the electronic device 601 (eg, the processor 620). The input device 650 may include, for example, a microphone, a mouse, or a keyboard.

[0090] The sound output device 655 can output sound signals to the outside of the electronic device 601. The sound output device 655 may include, for example, a speaker or a receiver. The speaker can be used for general purposes (e.g., playing multimedia or recording), and the receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0091] The display device 660 can visually provide information to the outside of the electronic device 601 (e.g., a user). The display device 660 may include, for example, a display, a hologram device, or a projector, and a control circuit system for controlling a corresponding one of the display, the hologram device, and the projector. According to one embodiment, the display device 660 may include a touch circuit system suitable for detecting a touch, or a sensor circuit system (e.g., a pressure sensor) suitable for measuring the strength of a force caused by a touch.

[0092] The audio module 670 can convert sound into electrical signals and convert electrical signals into sound. According to one embodiment, the audio module 670 can obtain sound through the input device 650, or output sound through the sound output device 655 or through headphones of the external electronic device 602 directly (e.g., wired) or wirelessly coupled to the electronic device 601.

[0093] The sensor module 676 can detect the operating state (e.g., power or temperature) of the electronic device 601 or the environmental state (e.g., user state) outside the electronic device 601, and then generate an electrical signal or data value corresponding to the detected state. The sensor module 676 may include, for example, a gesture sensor, a gyroscope sensor, an atmospheric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or a brightness sensor.

[0094] The interface 677 may support one or more prescribed protocols to be used to couple the electronic device 601 directly (e.g., in a wired manner) or wirelessly with the external electronic device 602. According to one embodiment, the interface 677 may include, for example, a high-definition multimedia interface (HDMI), a universal serial bus (USB) interface, a secure digital (SD) card interface, or an audio interface.

[0095] The connection terminal 678 may include a connection structure, and the electronic device 601 can be physically connected to the external electronic device 602 through the connection structure. According to one embodiment, the connection terminal 678 may include, for example, an HDMI connection structure, a USB connection structure, an SD card connection structure, or an audio connection structure (e.g., a headphone connection structure).

[0096] The haptic module 679 may convert the electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be recognized by the user through tactile sensation or kinesthetic sensation. According to one embodiment, the haptic module 679 may include, for example, a motor, a piezoelectric element, or an electrical stimulator.

[0097] The camera module 680 may capture still images or moving images. According to one embodiment, the camera module 680 may include one or more lenses, image sensors, image signal processors, or flashes.

[0098] The power management module 688 may manage power supplied to the electronic device 601. The power management module 688 may be implemented as, for example, at least a portion of a power management integrated circuit (PMIC).

[0099] The battery 689 may supply power to at least one component of the electronic device 601. According to one embodiment, the battery 689 may include, for example, a non-rechargeable primary cell, a rechargeable secondary cell, or a fuel cell.

[0100] The communication module 690 can support establishing a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device 601 and an external electronic device (e.g., electronic device 602, electronic device 604, or server 608) and performing communication through the established communication channel. The communication module 690 may include one or more communication processors that can operate independently of the processor 620 (e.g., AP) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module 690 may include a wireless communication module 692 (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module 694 (e.g., a local area network (LAN) communication module or a power line communication (PLC) module). A corresponding one of these communication modules can communicate with the user via a first network 698 (e.g., a short-range communication network, such as Bluetooth). TM, wireless-fidelity (Wi-Fi) direct or Infrared Data Association (IrDA) standard) or a second network 699 (for example, a long-distance communication network such as a cellular network, the Internet, or a computer network (for example, a LAN or a wide area network (WAN)))) to communicate with external electronic devices. These various types of communication modules can be implemented as a single component (for example, a single integrated circuit) or can be implemented as multiple components separated from each other (for example, multiple integrated circuits). The wireless communication module 692 can use user information (for example, an international mobile subscriber identity (IMSI)) stored in the user identification module 696 to identify and authenticate the electronic device 601 in the communication network (for example, the first network 698 or the second network 699).

[0101] The antenna module 697 can transmit or receive signals or power to or from the outside of the electronic device 601 (e.g., an external electronic device). According to one embodiment, the antenna module 697 may include one or more antennas, and, for example, the communication module 690 (e.g., the wireless communication module 692) may select at least one antenna suitable for the communication scheme used in the communication network (e.g., the first network 698 or the second network 699) from the one or more antennas. Signals or power can then be transmitted or received between the communication module 690 and the external electronic device via the selected at least one antenna.

[0102] At least some of the above components may be coupled to each other and signals (e.g., commands or data) may be sent between the at least some of the components via an inter-peripheral communication scheme (e.g., a bus, general purpose input and output (GPIO), a serial peripheral interface (SPI), or a mobile industry processor interface (MIPI)).

[0103] According to one embodiment, commands or data can be sent or received between electronic device 601 and external electronic device 604 via server 608 coupled to second network 699. Each of electronic device 602 and electronic device 604 can be of the same or different type as electronic device 601. All or some operations that would otherwise be performed at electronic device 601 can be performed at one or more of external electronic device 602, external electronic device 604, or external electronic device 608. For example, if electronic device 601 is to perform a function or service automatically or in response to a request from a user or another device, electronic device 601 can request one or more external electronic devices to perform at least a portion of the function or service instead of or in addition to performing the function or service. Upon receiving the request, the one or more external electronic devices can perform the at least a portion of the requested function or service, or perform other functions or other services related to the request, and transmit the results of the execution to electronic device 601. The electronic device 601 may provide the result as at least part of a reply to the request with or without further processing the result. To this end, for example, cloud computing, distributed computing, or client-server computing technology may be used.

[0104] One embodiment may be implemented as software (e.g., program 640) comprising one or more instructions stored in a storage medium (e.g., internal memory 636 or external memory 638) readable by a machine (e.g., electronic device 601). For example, the processor of the electronic device 601 may call at least one of the one or more instructions stored in the storage medium and execute the at least one instruction with or without one or more other components controlled by the processor. Thus, the machine may be operable to perform at least one function according to the at least one instruction called. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), but this term does not distinguish between a situation where data is stored in a storage medium in a semi-permanent manner and a situation where data is temporarily stored in a storage medium.

[0105] According to one embodiment, the method of the present disclosure may be included in and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read only memory (CD-ROM)) or through an application store (e.g., a Play Store). TM(Play Store TM )) online distribution (e.g., download or upload), or directly between two user devices (e.g., smartphones). If distributed online, at least a portion of the computer program product may be temporarily generated or at least temporarily stored in a machine-readable storage medium (e.g., a memory of a manufacturer's server, a server of an app store, or a relay server).

[0106] According to one embodiment, each of the above-mentioned components (e.g., a module or a program) may include a single entity or multiple entities. One or more of the above-mentioned components may be omitted, or one or more other components may be added. Alternatively or additionally, multiple components (e.g., modules or programs) may be integrated into a single component. In this case, the integrated component may still implement the one or more functions of each of the multiple components in the same or similar manner as the manner in which the corresponding one of the multiple components implements the one or more functions before integration. The operations implemented by a module, a program or another component may be performed sequentially, in parallel, repeatedly or heuristically, or one or more of the operations may be performed in a different order or omitted, or one or more other operations may be added.

[0107] Although specific embodiments of the present disclosure have been described in the detailed description of the present disclosure, the present disclosure may be modified in various forms without departing from the scope of the present disclosure. Therefore, the scope of the present disclosure should not be determined based solely on the embodiments described, but should be determined based on the appended claims and their equivalents.

Claims

1. A method for training a neural network, comprising: receiving a noisy signal and performing a short-time Fourier transform on the noisy signal; generating a denoised output signal from the noisy signal and performing an inverse short-time Fourier transform on the denoised output signal; determining a signal-to-distortion ratio loss function based on the denoised output signal after the inverse short-time Fourier transform, wherein the signal-to-distortion ratio loss function minimizes a spectral mismatch of the signal-to-distortion ratio using a scale-invariant signal-to-distortion ratio metric; determining a perceptual speech quality assessment loss function based on the denoised output signal after the inverse short time Fourier transform, wherein the perceptual speech quality assessment loss function approximates symmetric and asymmetric perturbations of the perceptual speech quality assessment to minimize a metric mismatch of the approximate perceptual speech quality assessment; and A total loss function is optimized based on the perceptual speech quality assessment loss function and the signal-to-distortion ratio loss function. 2 . The method according to claim 1 , wherein the perceptual speech quality assessment loss function is further determined based on the noisy signal. 3 . The method of claim 2 , wherein determining the perceptual speech quality assessment loss function further comprises performing level alignment on the noisy signal and the denoised output signal. 4 . The method of claim 2 , wherein determining the perceptual speech quality assessment loss function further comprises applying Bark spectrum frequencies of the noisy signal and the denoised output signal. 5 . The method of claim 4 , wherein determining the perceptual speech quality assessment loss function further comprises performing time-frequency equalization on the applied Bark spectral frequencies of the noisy signal and the denoised output signal. The method of claim 2 , wherein determining the perceptual speech quality assessment loss function further comprises performing loudness mapping.

7. The method of claim 2, wherein determining the perceptual speech quality assessment loss function further comprises performing a perturbation process.

8. The method of claim 1, wherein the total loss function is optimized as the sum of the product of the signal-to-distortion ratio loss function and the perceptual speech quality assessment loss function and a hyperparameter.

9. A system for training a neural network, comprising: Memory; as well as A processor configured to: receiving a noisy signal and performing a short-time Fourier transform on the noisy signal; generating a denoised output signal from the noisy signal and performing an inverse short-time Fourier transform on the denoised output signal; determining a signal-to-distortion ratio loss function based on the denoised output signal after the inverse short-time Fourier transform, wherein the signal-to-distortion ratio loss function minimizes a spectral mismatch of the signal-to-distortion ratio using a scale-invariant signal-to-distortion ratio metric; determining a perceptual speech quality assessment loss function based on the denoised output signal after the inverse short time Fourier transform, wherein the perceptual speech quality assessment loss function approximates symmetric and asymmetric perturbations of the perceptual speech quality assessment to minimize a metric mismatch of the approximate perceptual speech quality assessment; and A total loss function is optimized based on the perceptual speech quality assessment loss function and the signal-to-distortion ratio loss function.

10. The system of claim 9, wherein the perceptual speech quality assessment loss function is further determined based on the noisy signal. 11 . The system of claim 10 , wherein the processor is further configured to determine the perceptual speech quality assessment loss function by performing level alignment on the noisy signal and the denoised output signal.

12. The system of claim 10, wherein the processor is further configured to determine the perceptual speech quality assessment loss function by applying Bark spectrum frequencies of the noisy signal and the denoised output signal.

13. The system of claim 12, wherein the processor is further configured to determine the perceptual speech quality assessment loss function by performing time-frequency equalization on the applied Bark spectrum frequencies of the noisy signal and the denoised output signal.

14. The system of claim 10, wherein the processor is further configured to determine the perceptual speech quality assessment loss function by performing loudness mapping.

15. The system of claim 10, wherein the processor is further configured to determine the perceptual speech quality assessment loss function by performing a perturbation process.

16. The system of claim 9, wherein the total loss function is optimized as the sum of the product of the signal-to-distortion ratio loss function and the perceptual speech quality assessment loss function and a hyperparameter.

17. A method for training a neural network, comprising: receiving a noisy signal and performing a short-time Fourier transform on the noisy signal; generating a denoised output signal from the noisy signal and performing an inverse short-time Fourier transform on the denoised output signal; as well as A perceptual speech quality assessment loss function is determined based on the denoised output signal after the inverse short time Fourier transform, wherein the perceptual speech quality assessment loss function approximates symmetric and asymmetric perturbations of the perceptual speech quality assessment to minimize a metric mismatch of the approximate perceptual speech quality assessment.

18. The method of claim 17, wherein the perceptual speech quality assessment loss function is further determined based on the noisy signal.

19. The method of claim 18, wherein determining the perceptual speech quality assessment loss function further comprises performing level alignment on the noisy signal and the denoised output signal.

20. The method of claim 18, wherein determining the perceptual speech quality assessment loss function further comprises performing a perturbation process.

Citation Information

Patent Citations

  • Method and system for suppressing noise in speech signals in hearing aids and speech communication devices

    US20170032803A1

  • Neural decoding of attentional selection in multi-speaker environments

    WO2017218492A1