Audio signal processing method and electronic device for removing near-field voice
The audio signal processing method addresses the challenge of near-field speech interference by dividing audio signals into frequency bands and using neural networks to extract far components, resulting in improved audio quality by removing unwanted near-field speech.
Patent Information
- Application Number
- PCT/KR2024/096731
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-11
- Filing Date
- 2024-12-11
- Publication Date
- 2025-08-21
AI Technical Summary
Existing audio signal processing technologies struggle to effectively remove near-field speech, which often interferes with desired audio signals, degrading the quality of recorded audio.
An audio signal processing method that divides input audio signals into low, mid, and high frequency bands, determines the presence of near-field speech in the low frequency band, and extracts far components from these bands to generate an output audio signal, using neural networks for efficient separation.
Effectively removes near-field speech, enhancing the quality of recorded audio by isolating and preserving far-field components, thereby improving the clarity of desired audio signals.
Smart Images

Figure KR2024096731_21082025_PF_FP_ABST
Abstract
Description
Audio signal processing method and electronic device for removing near-field voice
[0001] The present disclosure relates to an audio signal processing method and an electronic device for removing near-field speech.
[0002] Recent efforts have been made to improve the quality of audio signals, including adjusting the timbre of the audio signal and removing noise. Unintended sounds can be introduced into audio signals, and removing them can improve the quality of the audio signal.
[0003] For example, parents often make noise to attract their children's attention when filming with a mobile device. Removing unintended noise from the final audio signal, thereby capturing only the desired audio signal, can improve audio quality.
[0004] An audio signal processing method for removing a near voice according to one embodiment of the present disclosure may include a step of obtaining an input audio signal. An audio signal processing method for removing a near voice according to one embodiment of the present disclosure may include a step of dividing the input audio signal into a low frequency band signal, a mid frequency band signal, and a high frequency band signal. The method may include a step of determining whether a near voice is included in the input audio signal based on the low frequency band signal. If the input audio signal includes a near voice, the method may include a step of extracting a far component from the low frequency band signal. The method may include a step of generating an output audio signal using the far component extracted from the low frequency band signal, the mid frequency band signal, and the high frequency band signal.
[0005] An electronic device for processing audio signals for removing near-field voices according to one embodiment of the present disclosure may include a memory storing a program for processing an audio signal or at least one instruction, and at least one processor. The electronic device may obtain an input audio signal by the at least one processor executing the program stored in the memory or at least one instruction. The electronic device may obtain an input audio signal by the at least one processor executing the program stored in the memory or at least one instruction. The electronic device (700) may divide the input audio signal into a low frequency band signal, a mid frequency band signal, and a high frequency band signal by the at least one processor executing the program stored in the memory or at least one instruction. The electronic device may determine whether a near voice is included in the input audio signal based on the low frequency band signal by the at least one processor executing the program stored in the memory or at least one instruction. The electronic device can extract a far component from the low-band signal if the input audio signal includes a near-field voice by executing a program or at least one instruction stored in the memory by the at least one processor. The electronic device can generate an output audio signal by using the far component extracted from the low-band signal, the mid-band signal, and the high-band signal by executing a program or at least one instruction stored in the memory by the at least one processor.
[0006] According to one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for executing any one of the aforementioned and hereinafter described methods for operating an electronic device and / or an electronic device can be provided.
[0007] Figure 1 is a diagram illustrating an example in which a user's voice is acquired as an audio signal.
[0008] FIG. 2 is a diagram for explaining the overall operation of an audio signal processing method according to one embodiment of the present disclosure.
[0009] FIG. 3 is a diagram for explaining the operation of a low-band signal processing module according to one embodiment of the present disclosure.
[0010] FIG. 4 is a diagram for explaining the operation of a mid-band signal processing module according to one embodiment of the present disclosure.
[0011] FIG. 5 is a diagram for explaining the operation of a frequency band merging module according to one embodiment of the present disclosure.
[0012] FIG. 6 is a diagram for explaining an operation of performing a calculation through a neural network according to one embodiment of the present disclosure.
[0013] FIG. 7 is a diagram illustrating a configuration of an audio signal processing electronic device for removing near-field speech according to one embodiment of the present disclosure.
[0014] FIG. 8 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0015] FIG. 9 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0016] FIG. 10 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0017] FIG. 11 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0018] FIG. 12 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0019] FIG. 13 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0020] FIG. 14 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0021] FIG. 15 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0022] FIG. 16 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0023] In describing this disclosure, descriptions of technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure will be omitted. This is to avoid obscuring the gist of this disclosure by omitting unnecessary explanations and to convey it more clearly. Furthermore, the terms described below are defined based on their functions in this disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification as a whole.
[0024] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.
[0025] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure of the present disclosure is complete and to fully inform those skilled in the art of the present disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. In addition, when describing an embodiment of the present disclosure, if a detailed description of a related function or configuration is determined to unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, the terms described below are terms defined in consideration of the functions of the present disclosure and may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification.
[0026] In one embodiment, each block of the flowchart diagrams and combinations of the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.
[0027] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.
[0028] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In addition, in one embodiment, the '~ unit' may include one or more processors.
[0029] When a part of the specification is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part," "module," etc., used throughout the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.
[0030] Additionally, the description 'at least one of A, B, and C' means that it can be any one of 'A', 'B', 'C', 'A and B', 'A and C', 'B and C', and 'A, B, and C'.
[0031] It should be understood that the blocks and combinations of flowcharts in each flowchart can be performed by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.
[0032] Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are to be understood to include plural referents. Thus, for example, the description "a component surface" may also include reference to one or more such surfaces.
[0033] All functions or operations described in this document may be performed by a single processor or a combination of processors. A single processor or a combination of processors is a circuitry that performs processing, and may include circuitry such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).
[0034] The artificial intelligence-related functions according to the present disclosure are operated via a processor and memory. The processor may be comprised of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a Digital Signal Processor (DSP), a graphics-only processor such as a GPU or a Vision Processing Unit (VPU), or an artificial intelligence-only processor such as an NPU. One or more processors control the processing of input data according to predefined operating rules or artificial intelligence models stored in memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0035] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is trained using a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0036] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.
[0037] Embodiments of the present disclosure relate to an audio signal processing method for removing near-field speech. Before describing specific embodiments, the meanings of terms frequently used in the present disclosure are defined.
[0038] The term 'input audio signal (10)' may refer to an original audio signal input to a photographing device when photographing an image. In embodiments of the present disclosure, the electronic device may remove a near-field sound from the input audio signal (10). In the present disclosure, the term 'mix signal' may refer to the input audio signal (10). In addition, the term 'mix signal spectrogram (212)' in the present disclosure may refer to a spectrogram for the input audio signal (10). The term 'mix signal' may refer to an input signal, an input sound signal, a mixed sound signal, etc., and may refer to a sound signal including both near sound and far sound. According to one embodiment, the input audio signal (10) may be an audio signal extracted from an image including an audio signal.
[0039] The 'low frequency band signal (222)' may refer to a signal corresponding to a low frequency band among signals obtained by dividing an input audio signal (10) into multiple frequency bands. The low frequency signal (222) may include both a low frequency component of a near-field sound and a low frequency component of a far-field sound. In the present disclosure, 'mix low spectrogram' may refer to a spectrogram for the low frequency signal (222). In one embodiment, the low frequency may include a frequency range of 0 to 8 kHz. In the present disclosure, the low frequency may also be described as a 'first frequency band'. The low frequency signal (222) may include an audio signal, a sound signal, etc. that can be heard in daily life. The low frequency signal (222) may include an audio signal that largely includes a voiced sound signal among a user's voice signal.
[0040] The term 'mid-band signal (224)' may refer to a signal corresponding to the mid-band among signals obtained by dividing an input audio signal (10) into multiple frequency bands. The mid-band signal (224) may include both mid-band components of near-field sounds and mid-band components of far-field sounds. In the present disclosure, 'mix mid spectrogram' may refer to a spectrogram for the mid-band signal (224). In one embodiment, the mid-band may include a frequency range of 8 kHz to 16 kHz. In the present disclosure, the mid-band may also be described as a 'second frequency band'. The mid-band signal (224) may include an audio signal that largely includes an unvoiced sound proportion among a user's voice signal.
[0041] The term 'high frequency band signal (226)' may refer to a signal corresponding to a high band among signals obtained by dividing an input audio signal (10) into multiple frequency bands. The high frequency band signal (226) may include both high frequency components of near-field sound and high frequency components of far-field sound. In the present disclosure, 'mix high spectrogram' may refer to a spectrogram for the high frequency band signal (226). In one embodiment, the high frequency band may include a frequency range of 16 kHz to 20 kHz. In one embodiment, the high frequency band may include a frequency range of 16 kHz to 24 kHz. In the present disclosure, the high frequency band may also be described as a 'third frequency band'. The high frequency band signal (226) may hardly include a user's voice signal. In the present disclosure, the high frequency band signal (226) may be divided from the input audio signal (10) to preserve the band of the input audio signal (10). Additionally, in the present disclosure, the high-band signal (226) may be divided to maintain the quality of the input audio signal (10).
[0042] "Near voice" may refer to a voice received at a close range. For example, the voice of a user (e.g., a person operating a video recording device, a person assisting in filming, etc.) located at a close range (e.g., 0.5 m) from a location where a video is being filmed (e.g., a location where a video recording device is located) may be referred to as near voice. In this case, the standard for determining whether a video is near (e.g., 0.5 m or 1 m) may be set in various ways depending on the situation or need. In the present disclosure, "far" may refer to a distance other than the close range.
[0043] The term "near component" may refer to a voice component of a user at a close range. The term "far component" may include an acoustic component generated by a subject at a far range. In the present disclosure, the far component may include a background acoustic component at a far range.
[0044] The near-field component extracted from the low-band signal can be referred to as the "low-band near-field component (322)". The far-field component extracted from the low-band signal can be referred to as the "low-band far-field component." The far-field component extracted from the mid-band signal can be referred to as the "mid-band far-field component." The far-field component extracted from the high-band signal can be referred to as the "high-band far-field component."
[0045] The spectrogram for the near-field component (322) extracted from the low-band signal may be referred to as a 'near low spectrogram'. The spectrogram for the far-field component (232) extracted from the low-band signal may be referred to as a 'far low spectrogram'. The spectrogram for the far-field component (242) extracted from the mid-band signal may be referred to as a 'far mid spectrogram'. The spectrogram for the far-field component (522) obtained from the high-band signal may be referred to as a 'far high spectrogram'.
[0046] The term 'output audio signal (20)' may refer to an audio signal from which a near-field sound is removed from an input audio signal (10). The output audio signal (20) may be an audio signal including all of the far-field components (522) of the low-band (232), the mid-band (242), and the high-band signal. In the present disclosure, the output audio signal (20) may refer to an audio signal including only the far-field sound of the input audio signal (10). In addition, in the present disclosure, 'far spectrogram' may refer to a spectrogram for the output audio signal (20). In one embodiment, if the input audio signal (10) does not include a near-field sound, the output audio signal (20) may be the same signal as the input audio signal (10).
[0047] Figure 1 is a diagram illustrating an example in which a user's voice is acquired as an audio signal.
[0048] An electronic device according to one embodiment of the present disclosure may include a smartphone, a tablet personal computer, a mobile phone, a video phone, a desktop PC, a camera, or a wearable device (e.g., smart glasses, a head-mounted device (HMD)). However, the present invention is not limited thereto, and the electronic device according to one embodiment may include various forms of electronic devices having an audio signal processing function.
[0049] According to one embodiment of the present disclosure, the electronic device (700) may be a device having both a photographing function and an audio signal processing function, such as a smartphone. For example, if the electronic device (700) is a smartphone, the electronic device (700) may remove nearby audio from a video captured by a camera provided in the electronic device (700).
[0050] Alternatively, according to one embodiment of the present disclosure, the electronic device (700) may be a device having only an audio signal processing function and may perform a process of removing a nearby voice from a video received from an external device. For example, the electronic device (700) may be a PC and the electronic device (700) may remove a nearby voice from a video received from a camera. Alternatively, for example, the electronic device (700) may be a cloud server and the electronic device (700) may remove a nearby voice from a video received from a smartphone and then transmit the processed video back to the smartphone or store it in a database.
[0051] In one embodiment, when a user (110) is photographing a subject (120), the user (110) may emit a voice (e.g., “Look here”) to focus the attention of the subject (120) or attract the attention of the subject (120). For example, when the user (110) is a parent and the subject (120) is a young child, the user (110) may emit a voice when photographing the subject (120). In this case, the user’s voice may be captured as an audio signal by the electronic device together with the sound generated from the subject (120) and the sound generated from the background (130). However, the user (110) may not intend for the user’s voice to be captured as an audio signal.
[0052] Hereinafter, a method and device for processing audio signals for removing near-field voices, removing unintended user voices and obtaining only sounds generated from the subject of the shooting and sounds generated from the background, will be described.
[0053] The present disclosure relates to an audio signal processing method for removing a near voice. An electronic device according to an embodiment of the present disclosure obtains an input audio signal to remove a user's voice signal from the audio signal, divides the input audio signal into a low frequency band signal, a mid frequency band signal, and a high frequency band signal, determines whether a near voice is included in the input audio signal based on the low frequency band signal, and if the near voice is included in the input audio signal, extracts a far component from the low frequency band signal, and generates an output audio signal (20) using the far component, mid frequency band signal, and high frequency band signal extracted from the low frequency band signal. Hereinafter, the overall process of processing an audio signal will be described first, and the process of extracting an audio signal in each frequency domain will be described in detail.
[0054] FIGS. 2 to 5 are drawings for explaining detailed configurations classified based on functions or roles included in an electronic device (700) according to one embodiment of the present disclosure. The detailed configurations (210 to 250, 310 to 350, 410 to 430, and 510 to 530) illustrated in FIGS. 2 to 5 may be software configurations implemented by a processor (730) of the electronic device (700) executing a program stored in a memory (720), or may be virtual configurations for which no matching hardware device actually exists. In other words, the operations performed by the processor (730) of the electronic device (700) by executing the program stored in the memory (720) may be classified into a plurality of groups by function or purpose, and the entities performing the operations included in each classified group may be expressed by the detailed components (210 to 250, 310 to 350, 410 to 430, and 510 to 530) of FIGS. 2 to 5. Accordingly, the operations described as being performed by the detailed components (210 to 250, 310 to 350, 410 to 430, and 510 to 530) illustrated in FIGS. 2 to 5 may be viewed as actually being performed by the processor (730) of the electronic device (700) by executing the program stored in the memory (720).
[0055] FIG. 2 is a diagram for explaining the overall operation of an audio signal processing method according to one embodiment of the present disclosure.
[0056] Referring to FIG. 2, the input audio signal acquisition module (210) is a module for performing an operation of receiving an input audio signal (10). The input audio signal (10) may include both near-field sound and far-field sound. In one embodiment, the near-field sound may include a user's voice signal. The input audio signal acquisition module (210) may extract an audio signal from an image including an audio signal to acquire the input audio signal (10). The image including the audio signal may be acquired through various paths. For example, the image may be an image captured by a photographing device. For example, the image may be an image downloaded through a communication network.
[0057] The input audio signal acquisition module (210) can convert the acquired input audio signal (10) from the time domain to the frequency domain. For example, the input audio signal acquisition module (210) can convert the input audio signal (10) to the frequency domain using Fourier transform. In one embodiment, a short-time Fourier transform (STFT) can be used in the process of converting the input audio signal to the frequency domain. The input audio signal acquisition module (210) can convert the input audio signal (10) from a waveform form to a spectrogram form.
[0058] The frequency band division module (220) is a module for performing an operation of dividing an input audio signal (10) in a frequency domain into a plurality of bands. The frequency band division module (220) can divide the input audio signal (10) in the frequency domain into a plurality of bands. In one embodiment, the frequency band division module (220) can divide the frequency domain of the input audio signal (10) into a low band (a first frequency band), a mid band (a second frequency band), and a high band (a third frequency band). In one embodiment, the low band may include a frequency range of 0 to 8 kHz. In one embodiment, the mid band may include a frequency range of 8 kHz to 16 kHz. In one embodiment, the high band may include a frequency range of 16 kHz to 20 kHz. In one embodiment, the high band may include a frequency range of 16 kHz to 24 kHz.
[0059] The low-band signal (222) split in the frequency band division module (220) can be input to the frequency band merging module (250) after the near-field component is removed by passing through the low-band signal processing module (230). The mid-band signal (224) split in the frequency band division module (220) can be input to the frequency band merging module (250) after the far-field component (242) is strengthened by passing through the mid-band signal processing module (240). The high-band signal (226) split in the frequency band division module (220) can be input to the frequency band merging module (250) as is.
[0060] The frequency band division module (220) divides the frequency range of the input audio signal (10) into multiple parts so that it can process signals in each frequency band, thereby increasing the computational efficiency for processing the input audio signal (10). In addition, the quality of the audio signal for removing near-field voices can be improved.
[0061] The low-band signal processing module (230) can determine whether a near-field voice is included in the input audio signal (10) based on the low-band signal. When the low-band signal (222) is input, the low-band signal processing module (230) can extract a far-field component (232) from the low-band signal. When the input audio signal (10) includes a near-field voice, the low-band signal processing module (230) can extract a far-field component (232) from the low-band signal. When the low-band signal (222) is input, the low-band signal processing module (230) can extract a low-band latent variable (234) from the low-band signal. The low-band latent variable (234) can be input to the mid-band signal processing module (240) and used for mid-band signal processing.
[0062] The mid-band signal processing module (240) can extract a distant component (242) from the mid-band signal when a mid-band signal (224) is input. The mid-band signal processing module (240) can utilize a low-band latent variable (234) to extract the distant component (242) from the mid-band signal. The mid-band signal processing module (240) can use the low-band latent variable (234) to obtain frame information containing a near-field voice from the input audio signal (10).
[0063] The user's voice is mostly contained in the low-band below 8 kHz. In other words, the user's voice energy is mostly concentrated in the low-band below 8 kHz. Therefore, the electronic device (700) can determine whether the input audio signal (10) includes the user's voice even when using only the low-band signal (222). Based on the low-band latent variable (234), the electronic device (700) can determine a frame including a close-range voice in the low-band signal (222). Based on the low-band latent variable (234), the electronic device (700) can obtain frame information including a close-range voice in the low-band signal (222). Since the user's voice energy is concentrated in the low-band, the electronic device (700) can determine a frame including a close-range voice in the input audio signal (10) by using a frame including a close-range voice in the low-band.
[0064] The electronic device (700) can determine a frame including a near-field voice through the low-band latent variable (234) in the low band and can also refer to this when processing the mid-band signal (224) in the mid-band. In other words, the mid-band signal processing module (240) can use the low-band latent variable (234). The mid-band signal processing module (240) can obtain frame information including a mid-band far-field component (242) based on the frame information including a near-field voice obtained through the low-band latent variable (234). Since the mid-band signal processing module (240) uses the low-band latent variable (234), it can determine a frame including a near-field voice without processing a mid-band signal or a high-band signal. Therefore, the amount of computation of the electronic device (700) can be reduced. The electronic device (700) can perform efficient calculations by using the low-band latent variable (234).
[0065] The frequency band merging module (250) can generate an output audio signal (20) using a low-band distant component (232) extracted from a low-band signal in a low-band signal processing module (230), a mid-band distant component (242) extracted from a mid-band signal in a mid-band signal processing module (240), a mid-band signal (224), and a high-band signal (226). The frequency band merging module (250) can generate an output audio signal (20) using a distant component (232) extracted from a low-band signal, a mid-band signal (224), and a high-band signal (226).
[0066] FIG. 3 is a diagram for explaining the operation of a low-band signal processing module (230) according to one embodiment of the present disclosure.
[0067] The low-band signal processing module (230) is a module for removing a near-field component from a low-band signal (222). The low-band signal processing module (230) can determine whether a near-field voice is included in the input audio signal (10) based on the low-band signal (222). If the near-field voice is included in the input audio signal (10), the low-band signal processing module (230) can extract a far-field component (232) from the low-band signal. When a low-band signal (222) is input, the low-band signal processing module (230) can extract a low-band latent variable (234) from the low-band signal.
[0068] The low-band signal processing module (230) may include a low-band encoder (310), a low-band near-field decoder (320), an energy calculation module (330), a near-field component determination module (340), and a low-band far-field decoder (350). When a low-band signal (222) is input to the low-band signal processing module (230), the low-band signal processing module (230) may output a low-band latent variable (234) and a low-band far-field component (232). The low-band signal processing module (230) may separate the low-band far-field component (232) and the low-band near-field component (322) from the low-band signal (222). In one embodiment, the low-band signal processing module (230) may use the input audio signal (10) as the output audio signal (20) when a low-band near-field component (322) does not exist.
[0069] The low-band encoder (310) can output a low-band latent variable (234) when a low-band signal (222) is input. The low-band signal (222) can exist in an entangled form in which a low-band far-field component (232) and a low-band near-field component (322) are difficult to separate.
[0070] Through the low-band latent variable (234) output by the low-band encoder (310), a frame of the input audio signal (10) including a low-band near-field component (322) among a plurality of frames of the input audio signal (10) can be determined. The low-band latent variable (234) can include tagging information of the low-band near-field component (322) of the input audio signal (10). Through the low-band latent variable (234), a frame of the input audio signal (10) including a low-band far-field component (232) among a plurality of frames of the input audio signal (10) can be determined. The low-band latent variable (234) can include tagging information of the low-band far-field component (232) of the input audio signal (10). The low-band latent variable (234) can include information indicating an active region for a frame in which a near-field component exists.
[0071] When a low-band signal (222) is input, the low-band encoder (310) can encode the low-band signal (222) into a form that is easy to interpret by the low-band remote decoder (350). The low-band latent variable (234) output by the low-band encoder (310) can be described as a low-band encoded vector. The low-band latent variable (234) output by the low-band encoder (310) can be referred to as a 'low-band hidden representation'.
[0072] The low-bandwidth encoder (310) may utilize a neural network model. In one embodiment, the low-bandwidth encoder (310) may utilize a deep neural network model (DNN).
[0073] The low-band near-field decoder (320) can output a low-band near-field component (322) when a low-band signal (222) and a low-band latent variable (234) are input. The low-band near-field decoder (320) can extract the low-band near-field component (322) from the low-band signal (222). Since a frame including the near-field component among multiple frames of the low-band signal (222) can be determined through the low-band latent variable (234), the low-band near-field decoder (320) uses the low-band latent variable (234) when extracting the low-band near-field component (322) from the low-band signal (222). Since it is possible to determine a frame including a distant component among multiple frames of a low-band signal (222) through a low-band latent variable (234), the low-band near-field decoder (320) uses the low-band latent variable (234) when extracting a low-band distant component (232) from a low-band signal (222).
[0074] The low-bandwidth near-field decoder (320) may utilize a neural network model. In one embodiment, the low-bandwidth near-field decoder (320) may utilize a deep neural network model (DNN).
[0075] In one embodiment, the low-band near-field decoder (320) can generate a spectral mask of a low-band signal. The low-band near-field decoder (320) can operate in a manner that compensates for the spectral mask of the low-band signal using complex components of the low-band signal.
[0076] In one embodiment, the low-band near-field decoder (320) can generate a spectral mask of the low-band signal using the low-band latent variable (234) and the low-band signal (222). The low-band near-field decoder (320) can obtain the magnitude of the near-field component using the spectral mask of the low-band signal. The low-band near-field decoder (320) can add the complex component of the low-band signal to the spectral mask of the low-band signal using the low-band latent variable (234). The low-band near-field decoder (320) can extract the low-band near-field component (322) by compensating the complex component of the low-band signal to the spectral mask of the low-band signal.
[0077] The energy calculation module (330) can calculate the energy (332) of the low-band near-field component (322) when the low-band near-field component (322) is input. The energy calculation module (330) outputs a value obtained by calculating the energy (332) of the low-band near-field component (322). The energy calculation module (330) can calculate the amplitude of each frequency component constituting the low-band signal. The energy calculation module (330) can calculate the energy of each frequency component through the square of the amplitude of each frequency component. The energy calculation module (330) can calculate the energy (332) of the low-band near-field component (322) by adding up the energies of each frequency component.
[0078] The near-field component judgment module (340) receives the energy (332) of the low-band near-field component (322). The near-field component judgment module (340) determines that the energy (332) of the low-band near-field component (322) is greater than a preset energy threshold ( ) is greater than or equal to, it can be determined that a low-band near-field component (322) exists. If the near-field component determination module (340) determines that a low-band near-field component (322) exists, it can output an activation signal (342). If it is determined that a low-band near-field component (322) exists, the low-band far-field decoder (350) is activated through the activation signal (342).
[0079] The near-field component judgment module (340) determines whether the energy (332) of the low-band near-field component (322) exceeds a preset energy threshold ( ), it can be determined that the low-band near-field component (322) does not exist. If the near-field component determination module (340) determines that the low-band near-field component (322) does not exist, it can not output the activation signal (342). If the near-field component determination module (340) determines that the low-band near-field component (322) does not exist, it can use the input audio signal as the output audio signal (20).
[0080] By operating the near-field component determination module (340), the presence or absence of a near-field component can be determined early, thereby increasing computational efficiency. By determining the presence or absence of a near-field component early and using the input audio signal (10) as the output audio signal (20), distortion of the output audio signal (20) can be minimized.
[0081] The low-band remote decoder (350) can output a low-band remote component (232) when a low-band signal (222) and a low-band latent variable (234) are input. The low-band remote decoder (350) can be activated through an activation signal (342). The low-band remote decoder (350) can extract the low-band remote component (232) from the low-band signal (222). Since a frame of the low-band signal including the remote component (232) among multiple frames of the low-band signal can be determined through the low-band latent variable (234), the low-band remote decoder (350) uses the low-band latent variable (234) when extracting the low-band remote component from the low-band signal.
[0082] The low-bandwidth remote decoder (350) may utilize a neural network model. In one embodiment, the low-bandwidth remote decoder (350) may utilize a deep neural network model (DNN).
[0083] FIG. 4 is a diagram for explaining the operation of a mid-band signal processing module (240) according to one embodiment of the present disclosure.
[0084] The mid-band signal processing module (240) is a module for removing a near-field component from a mid-band signal (224). The mid-band signal processing module (240) can extract a far-field component (242) from the mid-band signal (224). The mid-band signal processing module (240) can extract a far-field component (242) from the mid-band signal by using a low-band latent variable (234) output from the low-band signal processing module (230) and a mid-band signal (224) divided by the frequency band division module (220).
[0085] The mid-band signal processing module (240) may include a mid-band encoder (410), a mid-band latent variable operation module (420), and a mid-band remote decoder (430). When a mid-band signal (224) is input to the mid-band signal processing module (240), the mid-band signal processing module (240) may output a mid-band remote component (242). The mid-band signal processing module (240) may separate the mid-band remote component (242) and the mid-band near-field component from the mid-band signal (224).
[0086] The mid-band encoder (410) can output a mid-band latent variable (412) when a mid-band signal (224) is input. The mid-band signal (224) can exist in an entangled form in which the mid-band far-field component (242) and the mid-band near-field component are difficult to separate.
[0087] Through the mid-band latent variable (412) output by the mid-band encoder (410), a frame of the input audio signal (10) including a mid-band near-field component can be determined among a plurality of frames of the input audio signal (10). Through the mid-band latent variable (412), a frame of the input audio signal (10) including a mid-band far-field component (242) can be determined among a plurality of frames of the input audio signal (10). The mid-band latent variable (412) can include tagging information of the mid-band far-field component of the input audio signal (10). The mid-band latent variable (412) can include information indicating an active region for a frame in which a mid-band near-field component exists.
[0088] The mid-band encoder (410) can encode the mid-band signal (224) into a form that is easy to interpret by the mid-band remote decoder (430) when the mid-band signal (224) is input. The mid-band latent variable (412) output by the mid-band encoder (410) can be described as a mid-band encoded vector. The mid-band latent variable (412) output by the mid-band encoder (410) can be described as a 'mid-band hidden representation'.
[0089] The mid-band encoder (410) may utilize a neural network model. In one embodiment, the mid-band encoder (410) may utilize a deep neural network model (DNN).
[0090] The mid-band latent variable (412) can be used to determine a frame of the input audio signal (10) including a mid-band near-field component among multiple frames of the input audio signal (10), but the low-band latent variable (234), which is an output of the low-band signal processing module (230), can also be used. Since most of the user's voice, excluding unvoiced sounds, is included in the low-band, even when the low-band latent variable (234) is used, the frame of the input audio signal (10) including a mid-band near-field component among multiple frames of the input audio signal (10) can be determined. Through this, the mid-band near-field component can be reliably removed from a frame in which a low-band near-field component (322) exists, thereby obtaining a mid-band far-field component (242) of better quality.
[0091] The mid-band latent variable operation module (420) can operate the low-band latent variable (234) and the mid-band latent variable. The mid-band latent variable operation module (420) can combine the low-band latent variable (234) and the mid-band latent variable (412) to output a combined latent variable (422). In one embodiment, the mid-band latent variable operation module (420) can perform an operation by adding, concatenating, or using adaptive normalization on the low-band latent variable (234) and the mid-band latent variable (412).
[0092] The mid-band latent variable calculation module (420) can operate to obtain a mid-band far-field component (242) of better quality by performing calculations using the low-band latent variable (234). For example, when a close-range speech includes an unvoiced sound, the unvoiced sound of the close-range speech has characteristics similar to noise in the far-field. When the mid-band latent variable calculation module (420) does not perform calculations using the low-band latent variable (234) and only uses the mid-band latent variable (412), a close-range speech including an unvoiced sound may remain in the mid-band far-field component (242) obtained as a result of mid-band signal processing. However, if the low-band latent variable (234) is used together with the mid-band latent variable (412) for mid-band signal processing, i.e., if the combined latent variable (422) is used, the mid-band far-field decoder (430) can obtain the mid-band far-field component (242) with the unvoiced sound removed.
[0093] The mid-band remote decoder (430) can output a mid-band remote component (242) when a mid-band signal (224) and a combined latent variable (422) are input. The mid-band remote decoder (430) can extract the mid-band remote component (242) from the mid-band signal (224) using the combined latent variable (422).
[0094] The mid-band remote decoder (430) may utilize a neural network model. In one embodiment, the mid-band remote decoder (430) may utilize a deep neural network model (DNN).
[0095] FIG. 5 is a diagram for explaining the operation of a frequency band merging module (250) according to one embodiment of the present disclosure.
[0096] The frequency band merging module (250) is a module for generating a full-band far-field audio signal (20) using a low-band far-field component (232), a mid-band far-field component (242), and a high-band signal (226). The frequency band merging module (250) can merge multiple far-field audio signals. The frequency band merging module (250) can include an energy ratio calculation module (510), a high-band far-field component acquisition module (520), and a far-field component merging module (530).
[0097] The energy ratio calculation module (510) can obtain the energy ratio (512) of the mid-band signal (224) and the mid-band distant component (242) when the mid-band signal (224) and the mid-band distant component (242) are input. The energy ratio calculation module (510) can calculate the energy of the mid-band signal (224) when the mid-band signal (224) is input. The energy ratio calculation module (510) can calculate the energy of the mid-band distant component (242) when the mid-band distant component (242) is input. The energy ratio calculation module (510) can obtain the energy ratio (512) of the mid-band signal (224) and the mid-band distant component (242).
[0098] The high-band far-field component acquisition module (520) can obtain the high-band far-field component (522) by multiplying the energy ratio (512) of the mid-band signal (224) and the mid-band far-field component (242) by the high-band signal (226). By obtaining the high-band far-field component (522) using the energy ratio (512) of the mid-band signal (224) and the mid-band far-field component (242), the degree of distortion occurring from the high-band signal can be reduced.
[0099] The far-field component merging module (530) can merge low-band far-field components (232), mid-band far-field components (242), and high-band far-field components (522). The far-field component merging module (530) can concatenate the low-band far-field components (232), mid-band far-field components (242), and high-band far-field components (522) along the frequency axis.
[0100] The frequency band merging module (250) can perform an inverse Fourier transform on the concatenated frequency remote audio signal. In one embodiment, the frequency band merging module (250) can perform an inverse short-time Fourier transform (Inverse STFT) on the concatenated frequency remote audio signal. The frequency band merging module (250) can obtain an output audio signal (20) through the inverse Fourier transform. The frequency band merging module (250) can convert the output audio signal (20) from a spectrogram form to a waveform form. The frequency band merging module (250) can generate the output audio signal (20).
[0101] FIG. 6 is a diagram for explaining an operation of performing a calculation through a neural network according to one embodiment of the present disclosure.
[0102] In one embodiment of the present disclosure, the operation of processing at least one audio signal may be performed using artificial intelligence (AI) technology that performs calculations through a neural network.
[0103] A neural network (620) can be trained by receiving training data. Then, the trained neural network (620) can receive input data (610) as an input terminal (630), and the output terminal (650) can perform an operation to analyze the input data (610) and output output data (660), which is the desired result. The operation through the neural network can be performed through a hidden layer (640). In FIG. 6, for convenience, the hidden layer (640) is simplified and illustrated as being formed as a single layer, but the hidden layer (640) can be formed as a plurality of layers.
[0104] In one embodiment, the low-band encoder (310) and the low-band near-field decoder (320) may be a pair of neural networks. In one embodiment, the low-band encoder (310) and the low-band far-field decoder (350) may be a pair of neural networks. In one embodiment, the mid-band encoder (410) and the mid-band far-field decoder (430) may be a pair of neural networks.
[0105] In one embodiment, the encoder and decoder of the present disclosure can use echo information as training data. In one embodiment, the encoder and decoder of the present disclosure can learn distance information through echo information. The encoder and decoder of the present disclosure can determine near and far distances by learning the distance information. The encoder and decoder of the present disclosure can distinguish near and far components by learning the distance information. The encoder and decoder of the present disclosure can separate near and far components by learning the distance information.
[0106] FIG. 7 is a diagram illustrating a configuration of an audio signal processing electronic device for removing near-field speech according to one embodiment of the present disclosure.
[0107] An audio signal processing electronic device (700) for removing near-field voice may include at least one input / output interface (710), a memory (720) for storing at least one instruction, and at least one processor (730) for executing at least one instruction stored in the memory (720). The audio signal processing electronic device (700) for removing near-field voice may include a communication interface (740) for communicating with a server.
[0108] The memory (720) can store various data, programs, or applications for driving and controlling the audio signal processing electronic device (700) according to one embodiment of the present disclosure. The memory (720) can include, for example, a non-volatile memory including at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), a ROM (Read-Only Memory), and an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), and a volatile memory such as a RAM (Random Access Memory) or an SRAM (Static Random Access Memory).
[0109] The memory (720) may store instructions, data structures, and program codes that can be read by the processor (730). In the following embodiments, the processor (730) may be implemented by executing instructions or codes of a program stored in the memory (720).
[0110] The programs stored in the memory (720) can be classified into a plurality of modules according to their functions, and may include, for example, an input audio signal acquisition module (210), a frequency band division module (220), a low-band signal processing module (230), a mid-band signal processing module (240), and a frequency band merging module (250).
[0111] However, the instructions or codes of the program stored in the memory (720) are not limited thereto. For example, the input audio signal acquisition module (210), the frequency band division module (220), the low-band signal processing module (230), the mid-band signal processing module (240), and the frequency band merging module (250) are configurations for one embodiment, and the configuration of the modules included in the memory (720) may be integrated, added, or omitted according to the specifications of the audio signal processing electronic device (700) actually implemented. That is, two or more modules may be combined into one module, or one module may be divided into two or more modules and configured.
[0112] The input audio signal acquisition module (210), frequency band division module (220), low-band signal processing module (230), mid-band signal processing module (240), and frequency band merging module (250) stored in the memory (720) refer to a unit that processes a function or operation performed by the processor (730), and this can be implemented as software such as instructions, algorithms, data structures, or program codes.
[0113] The processor (730) may be configured with hardware components that perform arithmetic, logic, and input / output operations and signal processing. The processor (730) may be configured with at least one of, for example, a central processing unit (CPU), a microprocessor, a graphic processing unit (GRAP), an application specific integrated circuits (ASICs), a digital signal processor (DSPs), a digital signal processing device (DSPDs), a programmable logic device (PLDs), and a field programmable gate array (FPGAs), but is not limited thereto.
[0114] According to one embodiment of the present disclosure, there may be one or more processors (730). When there is one or more processors (730), the operations of the present disclosure may be performed by one or more processors individually or collectively executing instructions and / or programs stored in the memory (130). When a method according to one embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one processor (730) or by multiple processors (730).
[0115] According to one embodiment of the present disclosure, one or more processors may be implemented as a single-core processor or as a multi-core processor. If a method according to one embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single core or by multiple cores included in one or more processors.
[0116] The communication interface (740) can communicate with an external device or server through at least one wired or wireless communication network by the processor (730).
[0117] A communication interface (740) according to one embodiment may include at least one short-range communication module that performs communication according to a communication standard such as Bluetooth, Wi-Fi, BLE (Bluetooth Low Energy), NFC / RFID, Wi-Fi Direct, UWB, or ZIGBEE, and a long-range communication module that performs communication with a server for supporting long-range communication according to a long-range communication standard. The long-range communication module may perform communication through a communication network according to a 3G, 4G, and / or 5G communication standard, or a network for Internet communication.
[0118] In one embodiment, even when there are two or more users, the electronic device (700) according to one embodiment of the present disclosure may obtain an input audio signal (10) to remove a user's voice signal from the audio signal, divide the input audio signal (10) into a low frequency band signal, a mid frequency band signal, and a high frequency band signal, determine whether a near voice is included in the input audio signal (10) based on the low frequency band signal, and if the near voice is included in the input audio signal (10), extract a far component (232) from the low frequency band signal, and generate an output audio signal (20) using the far component (232), the mid frequency band signal (224), and the high frequency band signal (226) extracted from the low frequency band signal.
[0119] In embodiments of the present disclosure, the electronic device (700) can process the audio signal to remove the user's (photographer's) voice from the audio signal of a video by removing voices generated within a short distance from the shooting location. However, if the subject of the video is also located within a short distance from the shooting location, the problem of the subject's voice also being removed may arise. Embodiments for resolving this problem of the subject's voice being unintentionally removed are presented below.
[0120] According to one embodiment of the present disclosure, the electronic device (700) can selectively remove only the voice of the user (photographer) from among the near-field voices based on the directional information of the sound (voice) included in the input audio signal. If the input audio signal includes the near-field voices of multiple speakers and the input audio signal also includes directional information corresponding to each near-field voice, the electronic device can determine the speaker corresponding to each near-field voice based on the directional information.
[0121] For example, if a user (photographer) takes a video using the rear camera of a smartphone, and the smartphone is equipped with a directional microphone that can distinguish between sounds received from the front and sounds received from the rear, the sound received from the front will be the user's (photographer's) voice, and the sound received from the rear will be the voice of the subject of the video. Accordingly, an electronic device (e.g., smartphone) can analyze the input audio signal acquired through the directional microphone to distinguish between the user's (photographer's) voice and the subject's voice among near-field voices, and selectively remove only the user's (photographer's) voice.
[0122] Alternatively, according to one embodiment of the present disclosure, the electronic device may use speaker separation technology to distinguish between the user's (photographer's) voice and the subject's voice among the near-field voices included in the input audio signal, and selectively remove only the user's (photographer's) voice. The method by which the electronic device determines which of the separated voices is the user's (photographer's) voice may be implemented in various ways. For example, the electronic device may determine the user's (photographer's) voice by comparing it with a voice recorded in advance by the user's (photographer's) voice. Alternatively, for example, the electronic device may analyze the meaning of the user's (photographer's) voice through voice recognition, and determine that it is the user's (photographer's) voice if it relates to a request or instruction related to photography.
[0123] FIG. 8 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0124] In step 810, the audio signal processing electronic device (700) can obtain an input audio signal (10).
[0125] In step 820, the audio signal processing electronic device (700) can divide the input audio signal (10) into a low frequency band signal (222), a mid frequency band signal (224), and a high frequency band signal (226).
[0126] In step 830, the audio signal processing electronic device (700) can determine whether a near voice is included in the input audio signal (10) based on the low-band signal.
[0127] In step 840, the audio signal processing electronic device (700) can extract a far component (232) from a low-band signal (222) if the input audio signal (10) includes a near-field voice.
[0128] In step 850, the audio signal processing electronic device (700) can generate an output audio signal (20) using the far-field component (232), the mid-band signal (224), and the high-band signal (226) extracted from the low-band signal (222).
[0129] FIG. 9 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0130] Steps 820, 830 and 840 of FIG. 9 correspond to steps 820, 830 and 840 of FIG. 8, respectively.
[0131] In step 910, the audio signal processing electronic device (700) can extract a near component (322) from the low-band signal (222).
[0132] At step 920, the audio signal processing electronic device (700) can produce energy (332) of a near-field component extracted from a low-band signal (222).
[0133] In step 930, the audio signal processing electronic device (700) can determine that the input audio signal (10) contains a near-field voice if the produced energy is greater than a preset threshold value.
[0134] FIG. 10 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0135] Steps 830, 840 and 850 of FIG. 10 correspond to steps 830, 840 and 850 of FIG. 8, respectively.
[0136] In step 1010, the audio signal processing electronic device (700) can obtain a low frequency band latent variable (234) by passing the low frequency band signal (222) through a low frequency band encoder (310).
[0137] In step 1020, the audio signal processing electronic device (700) can obtain the far-field component (232) by passing the low-band latent variable (234) to the low-band far-field decoder (350).
[0138] FIG. 11 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0139] Steps 840 and 850 of FIG. 11 correspond to steps 840 and 850 of FIG. 8, respectively.
[0140] In step 1110, the audio signal processing electronic device (700) can extract a far-field component (242) from the mid-band signal (224).
[0141] In step 1120, the audio signal processing electronic device (700) can merge the far-field component (232) extracted from the low-band signal, the far-field component (242) extracted from the mid-band signal, the mid-band signal (224), and the high-band signal (232).
[0142] FIG. 12 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0143] Steps 840, 1110 and 1120 of FIG. 12 correspond to step 840 of FIG. 8, and steps 1110 and 1120 of FIG. 11, respectively.
[0144] In step 1210, the audio signal processing electronic device (700) can obtain the mid-band latent variable (412) by passing the mid-band signal (224) through the mid-band encoder (410).
[0145] At step 1220, the audio signal processing electronics (700) can combine the mid-band latent variable (412) and the low-band latent variable (234).
[0146] At step 1230, the audio signal processing electronic device (700) can obtain the far-field component (242) by passing the combined latent variable (422) through the mid-band far-field decoder (430).
[0147] FIG. 13 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0148] Steps 820, 910 and 920 of FIG. 13 correspond to step 820 of FIG. 8 and step 910 and 920 of FIG. 9, respectively.
[0149] In step 1310, the audio signal processing electronic device (700) can obtain a low frequency band latent variable (234) by passing the low frequency band signal (222) through a low frequency band encoder (310).
[0150] In step 1320, the audio signal processing electronic device (700) can obtain a near-field component (322) from the low-band signal (222) by passing the low-band potential variable (234) and the low-band signal (222) through a low-band near-field decoder (320).
[0151] FIG. 14 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0152] Steps 1110 and 1120 of FIG. 14 correspond to steps 1110 and 1120 of FIG. 11, respectively.
[0153] In step 1410, the audio signal processing electronic device (700) can calculate the energy of the mid-band signal (224) and the energy of the far-field component (242) extracted from the mid-band signal.
[0154] In step 1420, the audio signal processing electronic device (700) can calculate the energy ratio (512) of the produced energy and the produced distant component.
[0155] In step 1430, the audio signal processing electronic device (700) can obtain a high-band far-field component (522) based on the calculated energy ratio (512).
[0156] In step 1440, the audio signal processing electronic device (700) can merge the output audio signal (20) using the far-field component (232) extracted from the low-band signal, the far-field component (242) extracted from the mid-band signal, and the high-band far-field component (522).
[0157] FIG. 15 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0158] Steps 810 and 820 of FIG. 15 correspond to steps 810 and 820 of FIG. 8, respectively.
[0159] In step 1510, the audio signal processing electronic device (700) can convert the input audio signal (10) from a waveform form to a spectrogram form using a Short Time Fourier Transform (STFT).
[0160] FIG. 16 is a flowchart illustrating an audio signal processing method according to one embodiment of the present disclosure.
[0161] Steps 840 and 850 of FIG. 16 correspond to steps 840 and 850 of FIG. 8, respectively.
[0162] In step 1610, the audio signal processing electronic device (700) can convert the output audio signal (20) from a spectrogram form to a waveform form using Inverse Short Time Fourier Transform (Inverse STFT).
[0163] An audio signal processing electronic device (700) for removing near-field speech according to one embodiment of the present disclosure may include a memory (720) in which a program or at least one instruction for processing an audio signal is stored, and at least one processor (730).
[0164] The electronic device (700) can obtain an input audio signal (10) by having at least one processor (730) execute a program or at least one instruction stored in the memory (720).
[0165] The electronic device (700) can divide the input audio signal (10) into a low frequency band signal (222), a mid frequency band signal (224), and a high frequency band signal (226) by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0166] The electronic device (700) can determine whether a near voice is included in the input audio signal (10) based on the low-band signal (222) by executing a program or at least one instruction stored in the memory (720) by at least one processor (730).
[0167] The electronic device (700) can extract a far component (232) from the low-band signal (222) if the input audio signal (10) includes a near-field voice by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0168] The electronic device (700) can generate an output audio signal (20) by using the remote component (232), the mid-band signal (224), and the high-band signal (226) extracted from the low-band signal (222) by executing the program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0169] The electronic device (700) can extract a near component (322) from the low-band signal (222) by having at least one processor (730) execute a program or at least one instruction stored in the memory (720).
[0170] The electronic device (700) can calculate the energy of the near-field component (322) extracted from the low-band signal (222) by having at least one processor (730) execute a program or at least one instruction stored in the memory (720).
[0171] The electronic device (700) can determine that the input audio signal (10) includes a near-field voice if the calculated energy is greater than a preset threshold value by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0172] The electronic device (700) can obtain a low frequency band latent variable (234) by passing the low frequency band signal (222) through a low frequency band encoder (310) by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0173] The electronic device (700) can obtain the remote component (232) from the low-band signal (222) by passing the low-band latent variable (234) to the low-band remote decoder (350) by executing the program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0174] The electronic device (700) can extract a remote component (242) from the mid-band signal (224) by having at least one processor (730) execute a program or at least one instruction stored in the memory (720).
[0175] The electronic device (700) can generate the output audio signal (20) by merging the remote component (232) extracted from the low-band signal (222), the remote component (242) extracted from the mid-band signal (224), and the high-band signal (226) by executing the program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0176] The electronic device (700) can obtain a mid-frequency band latent variable (412) by passing the mid-band signal (224) through a mid-frequency band encoder (410) by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0177] The electronic device (700) can combine the mid-band latent variable (412) and the low-band latent variable (234) by having the at least one processor (730) execute a program or at least one instruction stored in the memory (720).
[0178] The electronic device (700) can obtain the remote component (242) from the mid-band signal (224) by passing the combined latent variable (442) to the mid-band remote decoder (430) by executing the program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0179] The electronic device (700) can obtain a low frequency band latent variable (234) by passing the low frequency band signal (222) through a low frequency band encoder (310) by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0180] The electronic device (700) can obtain a near-field component (322) from the low-band signal (222) by passing the low-band potential variable (234) and the low-band signal (222) through a low-band near-field decoder (320) by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0181] The electronic device (700) can calculate the energy of the mid-band signal (224) and the energy of the distant component (242) extracted from the mid-band signal (224) by executing the program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0182] The electronic device (700) can calculate the energy ratio (512) of the calculated energy and the calculated remote component by having at least one processor (730) execute a program or at least one instruction stored in the memory (720).
[0183] The electronic device (700) can obtain the high-bandwidth far-field component (522) based on the calculated energy ratio (512) by executing the program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0184] The electronic device (700) can generate an output audio signal (20) by merging a remote component (232) extracted from the low-band signal (222), a remote component (242) extracted from the mid-band signal (224), and a high-band remote component (522) by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0185] The electronic device (700) can convert the input audio signal (10) from a waveform form to a spectrogram form using a Short Time Fourier Transform (STFT) by executing a program or at least one instruction stored in the memory (720) by at least one processor (730).
[0186] The electronic device (700) can convert the output audio signal (20) from a spectrogram form to a waveform form using Inverse Short Time Fourier Transform (Inverse STFT) by executing a program or at least one instruction stored in the memory (720) by at least one processor (730).
[0187] The electronic device (700) can use the input audio signal (10) as an output audio signal (20) if the input audio signal (10) does not include a near-field voice by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0188] The electronic device (700) can determine a frame including a distant component (232) among a plurality of frames of the low-band signal (222) based on the low-band potential variable (234) by executing a program or at least one instruction stored in the memory (720) by the at least one processor (730).
[0189] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining an input audio signal (10).
[0190] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of dividing the input audio signal (10) into a low frequency band signal (222), a mid frequency band signal (224), and a high frequency band signal (226).
[0191] An audio signal processing method for removing a near voice according to one embodiment of the present disclosure may include a step of determining whether a near voice is included in the input audio signal (10) based on the low-band signal (222).
[0192] An audio signal processing method for removing a near-field voice according to one embodiment of the present disclosure may include a step of extracting a far component (232) from the low-band signal (222) if the near-field voice is included in the input audio signal (10).
[0193] An audio signal processing method for removing a near-field voice according to one embodiment of the present disclosure may include a step of generating an output audio signal (20) using a far-field component (232) extracted from the low-band signal (222), the mid-band signal (224), and the high-band signal (226).
[0194] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of extracting a near component (322) from the low-band signal (222).
[0195] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of calculating energy of a near-field component (322) extracted from the low-band signal (222).
[0196] An audio signal processing method for removing a near-field voice according to one embodiment of the present disclosure may include a step of determining that a near-field voice is included in the input audio signal (10) if the calculated energy is greater than a preset threshold value.
[0197] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining a low frequency band latent variable (234) by passing the low-band signal (222) through a low frequency band encoder (310).
[0198] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining a far-field component (232) from the low-band signal (222) by passing the low-band latent variable (234) through a low-band far-field decoder (350).
[0199] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of extracting a far-field component (242) from the mid-band signal (224).
[0200] An audio signal processing method for removing a near-field voice according to one embodiment of the present disclosure may include a step of generating the output audio signal (20) by merging a far-field component (232) extracted from the low-band signal (222), a far-field component (242) extracted from the mid-band signal (224), and the high-band signal (226).
[0201] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining a mid-frequency band latent variable (412) by passing the mid-band signal (224) through a mid-frequency band encoder (410).
[0202] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of combining the mid-band latent variable (412) and the low-band latent variable (234).
[0203] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining a far-field component (242) from the mid-band signal (224) by passing the combined latent variable (442) through a mid-band far-field decoder (430).
[0204] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining a low frequency band latent variable (234) by passing the low-band signal (222) through a low frequency band encoder (310).
[0205] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining a near-field component (322) from the low-band signal (222) by passing the low-band latent variable (234) and the low-band signal (222) through a low-band near-field decoder (320).
[0206] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of calculating energy of the mid-band signal (224) and energy of a far-field component (242) extracted from the mid-band signal (224).
[0207] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of calculating an energy ratio (512) of the calculated energy and the calculated far-field component.
[0208] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of obtaining the high-band far-field component (522) based on the calculated energy ratio (512).
[0209] An audio signal processing method for removing a near-field voice according to one embodiment of the present disclosure may include a step of generating an output audio signal (20) by merging a far-field component (232) extracted from the low-band signal (222), a far-field component (242) extracted from the mid-band signal (224), and the high-band far-field component (522).
[0210] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of converting the input audio signal (10) from a waveform form to a spectrogram form using a Short Time Fourier Transform (STFT).
[0211] An audio signal processing method for removing near-field speech according to one embodiment of the present disclosure may include a step of converting the output audio signal (20) from a spectrogram form to a waveform form using an Inverse Short Time Fourier Transform (Inverse STFT).
[0212] An audio signal processing method for removing a near-field voice according to one embodiment of the present disclosure may include a step of using the input audio signal (10) as the output audio signal (20) if the near-field voice is not included in the input audio signal (10).
[0213] An audio signal processing method for removing a near-field voice according to one embodiment of the present disclosure may include a step of determining a frame including a far-field component (232) among a plurality of frames of the low-band signal (222) based on the low-band latent variable (234).
[0214] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may represent one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive (HDD), compact disc (CD), digital video disc (DVD), or various types of memory (720).
[0215] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.
[0216] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0217] A processor may include various processing circuits and / or multiple processors. For example, the term “processor” as used herein, including in the claims, may include various processing circuits, including at least one processor. One or more processors in at least one processor may be configured to perform various functions described herein, individually and / or collectively, in a distributed fashion. As used herein, “processor,” “at least one processor,” and “one or more processors” may be configured to perform multiple functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, at least one processor may include a combination of processors that perform various of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.
[0218] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.
[0219] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. A method for processing an audio signal to remove close-range voice, Step of obtaining an input audio signal (10); A step of dividing the above input audio signal (10) into a low frequency band signal (222), a mid frequency band signal (224), and a high frequency band signal (226); A step of determining whether a near voice is included in the input audio signal (10) based on the low-band signal (222); If the input audio signal (10) includes a near-field voice, a step of extracting a far component (232) from the low-band signal (222); and A method comprising the step of generating an output audio signal (20) by using the long-range component (232) extracted from the low-band signal (222), the mid-band signal (224), and the high-band signal (226).
2. In paragraph 1, The above judging step is, A step of extracting a near component (322) from the above low-band signal (222); A step of calculating the energy of a near-field component (322) extracted from the above low-band signal (222); and A method characterized by including a step of determining that a near-field voice is included in the input audio signal (10) if the above-described energy is greater than a preset threshold value.
3. In either paragraph 1 or paragraph 2, The step of extracting the long-range component (232) from the above low-band signal (222) is as follows: A step of obtaining a low frequency band latent variable (234) by passing the low frequency band signal (222) through a low frequency band encoder (310); and A method characterized by comprising a step of obtaining a distant component (232) from the low-band signal (222) by passing the low-band latent variable (234) through a low-band distant decoder (350).
4. In paragraph 3, The step of generating the above output audio signal (20) is: A step of extracting a long-range component (242) from the above mid-band signal (224); A method characterized by comprising a step of generating the output audio signal (20) by merging a distant component (232) extracted from the low-band signal (222), a distant component (242) extracted from the mid-band signal (224), and the high-band signal (226).
5. In paragraph 4, The step of extracting the long-range component (242) from the above mid-band signal (224) is as follows: A step of obtaining a mid-frequency band latent variable (412) by passing the mid-band signal (224) through a mid-frequency band encoder (410); A step of combining the above mid-band latent variable (412) and the above low-band latent variable (234); A method characterized by comprising a step of obtaining a distant component (242) from the intermediate band signal (224) by passing the combined latent variable (442) through a intermediate band distant decoder (430).
6. In any one of paragraphs 2 to 5, The step of extracting the near-field component (322) from the low-band signal (222) is as follows: A step of obtaining a low frequency band latent variable (234) by passing the low frequency band signal (222) through a low frequency band encoder (310); and A method characterized by comprising a step of obtaining a near-field component (322) from the low-band signal (222) by passing the low-band latent variable (234) and the low-band signal (222) through a low-band near-field decoder (320).
7. In either of paragraphs 4 or 5, The step of generating the above output audio signal (20) is: A step of calculating the energy of the above mid-band signal (224) and the energy of the distant component (242) extracted from the above mid-band signal (224); A step of calculating the energy ratio (512) of the above-mentioned calculated energy and the above-mentioned calculated long-distance component; A step of obtaining a high-band far-field component (522) based on the above-described energy ratio (512); A method characterized by comprising a step of generating the output audio signal (20) by merging the distant component (232) extracted from the low-band signal (222), the distant component (242) extracted from the mid-band signal (224), and the high-band distant component (522).
8. In an audio signal processing electronic device (700) for removing close-range voice, Memory (720) for storing instructions; and At least one processor (730) comprising processing circuitry, The electronic device (700) is configured such that the instructions are individually or collectively executed by the at least one processor (730). Acquire the input audio signal (10), The above input audio signal (10) is divided into a low frequency band signal (222), a mid frequency band signal (224), and a high frequency band signal (226), Based on the above low-band signal (222), it is determined whether the input audio signal (10) includes a near voice, If the input audio signal (10) above includes a near-field voice, the far component (232) is extracted from the low-band signal (222), Generating an output audio signal (20) by using the long-range component (232) extracted from the low-band signal (222), the mid-band signal (224), and the high-band signal (226). Electronic device (700).
9. In paragraph 8, The electronic device (700) is configured such that the instructions are individually or collectively executed by the at least one processor (730). Extracting the near component (322) from the above low-band signal (222), Calculate the energy of the near-field component (322) extracted from the above low-band signal (222), An electronic device (700) that determines that the input audio signal (10) contains a near-field voice if the energy produced above is greater than a preset threshold value.
10. In either of paragraphs 8 or 9, The electronic device (700) is configured such that the instructions are individually or collectively executed by the at least one processor (730). By passing the above low-band signal (222) through a low-band encoder (310), a low-band latent variable (234) is obtained, An electronic device (700) that obtains the remote component (232) by passing the low-band latent variable (234) through a low-band remote decoder (350).
11. In paragraph 10, The electronic device (700) is configured such that the instructions are individually or collectively executed by the at least one processor (730). Extracting the far-field component (242) from the above mid-band signal (224), An electronic device (700) that generates the output audio signal (20) by merging a distant component (232) extracted from the low-band signal (222), a distant component (242) extracted from the mid-band signal (224), and the high-band signal (226).
12. In paragraph 11, The electronic device (700) is configured such that the instructions are individually or collectively executed by the at least one processor (730). By passing the above mid-band signal (224) through a mid-band encoder (410), a mid-band latent variable (412) is obtained, An electronic device (700) that obtains the remote component (242) by combining the mid-band latent variable (412) and the low-band latent variable (234) and passing the combined latent variable (442) through a mid-band remote decoder (430).
13. In any one of paragraphs 9 to 12, The electronic device (700) is configured such that the instructions are individually or collectively executed by the at least one processor (730). By passing the above low-band signal (222) through a low-band encoder (310), a low-band latent variable (234) is obtained, An electronic device (700) that obtains the near-field component (322) by passing the low-band potential variable (234) and the low-band signal (222) through a low-band near-field decoder (320).
14. In either of paragraphs 11 or 12, The electronic device (700) is configured such that the instructions are individually or collectively executed by the at least one processor (730). An electronic device (700) that calculates the energy of the mid-band signal (224) and the energy of the distant component (242) extracted from the mid-band signal (224), calculates an energy ratio (512) of the calculated energy and the distant component, obtains a high-band distant component (522) based on the calculated energy ratio (512), and generates an output audio signal (20) by merging the distant component (232) extracted from the low-band signal (222), the distant component (242) extracted from the mid-band signal (224), and the high-band distant component (522).
15. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 7 on a computer.
Citation Information
Patent Citations
Apparatus and method for robust speech recognition of speaker distance character
KR100855592B1
System Of Foreign Language Training Based On Chatbot And Method Thereof
KR1020220017366A
Apparatus and method for enhancing speaker feature based on deep neural network that selectively compensates for distant utterances
KR102340359B1
Manufacturing method of marshmallow maximizing the content of lactic acid bacteria using lactic acid bacteria instantaneous deposition technique and marshmallow manufactured using the same
KR102365680B1
Voice recognition audio system and method
KR102491417B1