Electronic device, method, and non-transitory computer-readable recording medium for acoustic beamforming
By employing AI models to process audio signals in the latent and frequency domains, the electronic device achieves enhanced acoustic beamforming performance with a small number of microphones, addressing the limitations of conventional methods in terms of cost and space.
Patent Information
- Application Number
- PCT/KR2024/012110
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-08-14
- Publication Date
- 2025-06-12
AI Technical Summary
Conventional acoustic beamforming methods require a large number of microphones to achieve improved performance, which increases manufacturing costs and requires more space, making it challenging to enhance beamforming performance with a small number of microphones.
The use of artificial intelligence (AI) models to perform acoustic beamforming, where the electronic device encodes time-domain input audio signals into latent domain signals, identifies masks for AI-based beamforming, and applies these masks to the audio signals in the frequency domain to adjust beamforming directions and beamwidths.
This approach allows for improved beamforming performance with a small number of microphones, reducing costs and space requirements while enhancing the ability to selectively acquire sound source signals and reduce noise.
Smart Images

Figure KR2024012110_12062025_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transitory computer-readable recording medium for acoustic beamforming
[0001] The following descriptions relate to electronic devices, methods, and non-transitory computer-readable recording media for acoustic beamforming.
[0002] Acoustic beamforming can refer to a technology for acquiring a signal from a sound source in a specific direction among multiple sound sources around an electronic device through a microphone array, or for removing / reducing noise other than sound sources in a specific direction.
[0003] For example, an electronic device can selectively acquire a sound source signal input from a direction of interest by utilizing time delay and / or phase delay depending on the distance between microphones in a microphone array.
[0004] An electronic device is disclosed. The electronic device may include a plurality of microphones, a processor, and a memory storing instructions. The instructions, when executed by the processor, may cause the electronic device to encode a time-domain input audio signal received through the plurality of microphones into an input audio signal in a latent domain. The instructions, when executed by the processor, may cause the electronic device to identify a mask for performing AI-based beamforming on the input audio signal based on the input audio signal in the latent domain. The instructions, when executed by the processor, may cause the electronic device to obtain an AI beamforming input audio signal in which at least one of a beamforming direction or a beamforming beamwidth of the input audio signal is applied to the input audio signal by applying the mask to the audio signal in the frequency domain.
[0005] An electronic device is disclosed. The electronic device may include a processor, and a memory storing a sound source database (DB), a channel impulse response DB for a plurality of microphones, a channel noise DB, and instructions. The instructions, when executed by the processor, may cause the electronic device to select at least one sound signal from the sound source DB. The instructions, when executed by the processor, may cause the electronic device to generate a first audio signal including a direct sound signal and a reflected sound signal for the at least one sound signal based on the channel impulse response DB. The instructions, when executed by the processor, may cause the electronic device to generate an input audio signal in which noise selected from the channel noise DB and the first audio signal are combined. The instructions, when executed by the processor, may cause the electronic device to generate an output audio signal based on the input audio signal by reducing the noise of the input audio signal and applying at least one of a beamforming beamwidth or a beamforming direction to the first audio signal of the input audio signal. The instructions, when executed by the processor, may cause the electronic device to generate an AI-modulated output audio signal based on applying an AI model to the input audio signal. The instructions, when executed by the processor, may cause the electronic device to train the AI model based on reducing a difference between the AI-modulated output audio signal and the output audio signal.
[0006] A method is disclosed. The method may include an operation of encoding an input audio signal in a time domain received through the plurality of microphones into an input audio signal in a latent domain. The method may include an operation of identifying a mask for performing AI-based beamforming on the input audio signal based on the input audio signal in the latent domain. The method may include an operation of obtaining an AI beamforming input audio signal in which at least one of a beamforming direction or a beamforming beamwidth of the input audio signal is applied to the input audio signal by applying the mask to the audio signal in the frequency domain.
[0007] A method is disclosed. The method can be performed by an electronic device including a memory that stores a sound source database (DB), a channel impulse response DB for a plurality of microphones, and a channel noise DB. The method can include an operation of selecting at least one sound signal from the sound source DB. The method can include an operation of generating a first audio signal including a direct sound signal and a reflected sound signal for the at least one sound signal based on the channel impulse response DB (333). The method can include an operation of generating an input audio signal in which noise selected from the channel noise DB (335) and the first audio signal are combined. The method can include an operation of generating an output audio signal based on the input audio signal by reducing the noise of the input audio signal and applying at least one of a beamforming beamwidth or a beamforming direction to the first audio signal of the input audio signal. The method can include an operation of generating an AI-controlled output audio signal based on applying an AI (artificial intelligence) model (410) to the input audio signal. The method may include an operation of training the AI model (410) based on reducing a difference between the AI-controlled output audio signal and the output audio signal.
[0008] A computer-readable recording medium is disclosed. The computer-readable recording medium has computer-executable instructions stored thereon, which, when executed by a computer, cause the computer to encode a time-domain input audio signal (501) received through the plurality of microphones (271, 273, 275) into an input audio signal of a latent domain, identify a mask for performing AI-based beamforming on the input audio signal based on the input audio signal of the latent domain, and apply the mask to the audio signal of the frequency domain, thereby obtaining an AI beamforming input audio signal in which at least one of a beamforming direction or a beamforming beamwidth of the input audio signal (501) is applied to the input audio signal.
[0009] A computer-readable recording medium is disclosed. The computer-readable recording medium causes, when computer-executable instructions stored in the computer-readable recording medium are executed by the computer, the computer to select at least one sound signal from the sound source DB (331), generate a first audio signal including a direct sound signal and a reflected sound signal for the at least one sound signal based on the channel impulse response DB (333), generate an input audio signal in which noise selected from the channel noise DB (335) and the first audio signal are combined, reduce the noise of the input audio signal, and generate an output audio signal based on the input audio signal by applying at least one of a beamforming beamwidth or a beamforming direction to the first audio signal of the input audio signal, and generate an AI-controlled output audio signal based on applying an AI (artificial intelligence) model (410) to the input audio signal, and train the AI model (410) based on reducing a difference between the AI-controlled output audio signal and the output audio signal.
[0010] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0011] FIG. 2 is a block diagram of an electronic device according to one embodiment.
[0012] FIG. 3 is a diagram illustrating an operation for generating an input audio signal and an output audio signal in an electronic device according to one embodiment.
[0013] FIG. 4 is a diagram illustrating an operation for training an AI (artificial intelligence) model based on an input audio signal and an output audio signal in an electronic device according to one embodiment.
[0014] FIG. 5 is a diagram illustrating an operation for acoustic beamforming an audio signal in an electronic device according to one embodiment.
[0015] Figure 6 is a drawing illustrating a three-dimensional space around an electronic device.
[0016] FIG. 7 is a drawing illustrating an operation of outputting an audio signal formed by acoustic beamforming through an electronic device.
[0017] FIG. 8A is a flowchart illustrating an operation of an electronic device obtaining an input audio signal according to one embodiment.
[0018] Figure 8b is a diagram illustrating the path of an audio signal acquired by an electronic device.
[0019] Figure 8c is a graph illustrating an audio signal obtained from an electronic device along a path of the audio signal.
[0020] FIG. 9 is a flowchart illustrating an operation of an electronic device obtaining an output audio signal according to one embodiment.
[0021] FIG. 10 is a flowchart illustrating an operation of an electronic device for acoustic beamforming an audio signal according to one embodiment.
[0022] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.
[0023] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0024] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0025] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0026] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0027] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0028] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0029] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0030] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0031] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0032] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0033] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0034] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0035] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0036] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0038] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0039] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for realizing eMBB, a loss coverage (e.g., 664 dB or less) for realizing mMTC, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 6 ms or less for round trip) for realizing URLLC.
[0041] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0043] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0045] FIG. 2 is a block diagram of an electronic device according to one embodiment.
[0046] Referring to FIG. 2, the electronic device (101) may include a processor (120), a memory (130), a plurality of speakers (251, 253, 255), a display (260), and / or a plurality of microphones (mics or microphones) (271, 273, 275).
[0047] In one embodiment, the processor (120) may be used to execute the operations of the electronic device (101) exemplified in the descriptions of FIGS. 3 to 10. For example, the processor (120) may include at least a portion of the processor (120) of FIG. 1 or may correspond to at least a portion of the processor (120) of FIG. 1. For example, the processor (120) may include one or more processors, including an application processor (AP) and / or a communication processor (CP). For example, the processor (120) may be implemented as a single chip, such as a system on chip (SoC), or may be implemented as multiple chips. For example, the processor (120) may be implemented as a single integrated circuit or may be implemented as multiple integrated circuits. For example, the processor (120) may be arranged in a distributed manner within the electronic device (101).
[0048] In one embodiment, the memory (130) may (at least temporarily) store instructions for executing operations of the electronic device (101) exemplified in the descriptions of FIGS. 3 to 10. The instructions may be executed by the processor (120). The instructions may be included in one or more programs stored in the memory (130). For example, the memory (130) may include at least a portion of the memory (130) of FIG. 1 (or at least a portion of the non-volatile memory (134)) or may correspond to at least a portion of the memory (130) of FIG. 1 (or at least a portion of the non-volatile memory (134)). For example, the memory (130) may include a main memory (e.g., a random access memory (RAM)) within the electronic device (101), a register for the processor (120), a cache for the processor (120), a register for the communication circuit (390), a buffer (or soft buffer) for the communication circuit (390), and / or an auxiliary memory (e.g., a hard disk drive (HDD) or a solid state drive (SSD)) of the electronic device (101). For example, the memory (130) may be implemented as a single chip or may be implemented as multiple chips. For example, the memory (130) may be implemented as one integrated circuit or may be implemented as multiple integrated circuits. For example, the memory (130) may be arranged in a distributed manner within the electronic device (101).
[0049] In one embodiment, the plurality of speakers (251, 253, 255) may output audio signals to the outside. For example, the plurality of speakers (251, 253, 255) may include at least a portion of the audio module (170) and / or the audio output module (155) of FIG. 1, or may correspond to at least a portion of the audio module (170) and / or the audio output module (155) of FIG. 1.
[0050] In one embodiment, the display (260) can display visual content. For example, the display (260) can include at least a portion of the display module (160) of FIG. 1 or can correspond to at least a portion of the display module (160) of FIG. 1.
[0051] In one embodiment, a plurality of microphones (271, 273, 275) may be used to acquire (or receive) audio signals corresponding to sounds acquired from the outside. For example, the plurality of microphones (271, 273, 275) may include at least a portion of the audio module (170) and / or the audio output module (155) of FIG. 1, or may correspond to at least a portion of the audio module (170) and / or the audio output module (155) of FIG. 1. For example, the plurality of microphones (271, 273, 275) may include a dynamic microphone, a condenser microphone, and / or a piezo microphone. For example, the plurality of microphones (271, 273, 275) may be referred to as a microphone array. For example, the plurality of microphones (271, 273, 275) may be arranged at different locations within the electronic device (101). For example, the plurality of microphones (271, 273, 275) may acquire audio signals that are time-delayed and / or phase-delayed differently from a sound signal generated from a sound source in a three-dimensional space. Conventional acoustic beamforming may be implemented by applying predetermined weights / delays to signals from each microphone depending on the intended beamforming direction.
[0052] The beamforming performance may be improved as the number of multiple microphones (271, 273, 275) increases. However, as more microphones (271, 273, 275) are mounted (or arranged) on the electronic device (101), the manufacturing cost may increase. In addition, increased space is required as the number of microphones increases. Therefore, a method for improving beamforming performance while mounting (or arranging) a small number (e.g., three) of multiple microphones (271, 273, 275) on the electronic device (101) may be required. In other words, there may be an advantage if improved beamforming performance can be achieved without increasing the number of microphones.
[0053] Below, examples are described for improving beamforming performance using a small number of multiple microphones (271, 273, 275). In particular, approaches utilizing AI models to perform acoustic beamforming are provided, along with approaches for generating training data to train the AI models.
[0054] FIG. 3 is a diagram illustrating operations for generating input audio signals and output audio signals in an electronic device according to one embodiment. In particular, FIG. 3 illustrates an approach for generating artificial input and output audio signals for training an AI model that implements acoustic beamforming.
[0055] Figure 3 can be explained with reference to Figures 1 and 2.
[0056] The sound source database (DB) (331), the multi-channel impulse response DB (333), and the multi-channel noise DB (335) illustrated in FIG. 3 may be stored in the memory (130). The sound source DB (331), the multi-channel impulse response DB (333), and the multi-channel noise DB (335) may be accessible by the processor (120).
[0057] In one embodiment, the sound source DB (331) may include a plurality of different sound signals without noise (or background noise), and the sound signals may include any type of sound signals such as, for example, voices, animal noises, machine noises, and music. The sound signals may be considered as original sound / noise-free sound signals as opposed to sound signals received through a noisy environment. The sound signals of the sound source DB (331) may also be referred to as, for example, sound sources, sound source signals, original sound, original sound signals, clean sound, clean sound signals, noise-free sound, or noise-free sound signals.
[0058] In one embodiment, the multi-channel (or channel) impulse response DB (333) may include audio signals (or impulse responses, channel impulse responses, channel responses, channel values / vectors / matrices, etc.) measured at each of a plurality of (i.e., array) microphones (271, 273, 275) when a sound signal from a (static or dynamic) sound source in a three-dimensional space to an electronic device (101) is directly transmitted to the microphones (without reflection). In one embodiment, the position of the sound source in a three-dimensional space is, when the electronic device (101) is taken as the origin, the azimuth and elevation angle can be expressed by. The audio signal (or impulse response) of the multi-channel impulse response DB (333) is azimuthally relative to the electronic device (101) in an anechoic room environment. Range (e.g. 0 to 360 degrees) and elevation angle A sound signal (e.g., a sine sweep signal) generated (or created) through an external speaker located within a range (e.g., -90 to 90 degrees) can be obtained by having multiple microphones (271, 273, 275) measure (or acquire) (or record). However, the present invention is not limited thereto. For example, an audio signal (or impulse response) of a multi-channel impulse response DB (333) can be calculated through an acoustic simulator.
[0059] In one embodiment, the multi-channel (or channel) noise DB (335) may include spatially uncorrelated noise (or background noise) signals. In one embodiment, the noise (or background noise) signals of the multi-channel noise DB (335) may include diffuse noise (or later reverberations) and / or microphone self-noise. In one embodiment, the noise (or background noise) signals of the multi-channel noise DB (335) may be signals that are not related to sound signals (or speech signals) of the sound source DB (331). For example, the noise (or background noise) signals may be distinguished from signals along a direct path of the sound signals (or speech signals) of the sound source DB (331) or signals along a reflected path (e.g., early reflections). In one embodiment, early reflections may be sound that reaches the electronic device (101) after sound (or direct sound) along the direct path is reflected (a small number of times) by a reflector (e.g., a wall, a ceiling, or a floor) in a three-dimensional space. In one embodiment, late reverberation may be sound that reaches the electronic device (101) after sound (or direct sound) along the direct path is reflected (a large number of times) by a reflector (e.g., a wall, a ceiling, or a floor) in a three-dimensional space, from multiple reflections. Multi-channel noise may be generated statistically and / or acquired via microphones.
[0060] The parameter generator (310) and audio signal generator (320) illustrated in Fig. 3 can be stored as programs in the memory (130). The sound source DB (331), the multi-channel impulse response DB (333), and the multi-channel noise DB (335) can include instructions executable by the processor (120).
[0061] In one embodiment, the parameter generator (310) can generate parameters for adjusting (i.e., steering, or changing) a beam associated with a generated (i.e., simulated, synthesized, artificial) audio signal (an output audio signal (345) generated by the audio signal generator (320)). The parameter generator (310) can also be used to generate parameters when actual beamforming is performed, as described with respect to FIG. 5. In one embodiment, the parameters can include parameters for adjusting the direction and / or width of the beam. For example, the direction of the beam can be an azimuth angle relative to the electronic device (101). Range (e.g. 0 to 360 degrees) and elevation angle may include a direction within a range (e.g., -90 to 90 degrees). For example, the width (or angle) of a beam may represent the width (or angle) of the beam's direction (or orientation). For example, the beam width may be normalized to be between 0 and 1.
[0062] In one embodiment, the parameter generator (310) may generate a parameter for adjusting (i.e., adjusting, or changing) the level of background noise in the generated audio signal. In one embodiment, the parameter may have a value between 0 and 1 to indicate the level of background noise in the generated audio signal.
[0063] In one embodiment, the parameter generator (310) can randomly generate parameters related to the beam of the generated audio signal and parameters related to background noise within a specified range. For example, the parameter generator (310) can generate parameters related to the azimuth Range (e.g. 0 to 360 degrees) and elevation angle The parameter generator (310) can randomly generate a parameter related to direction within a range (e.g., -90 to 90 degrees). For example, the parameter generator (310) can randomly generate a parameter related to beam width within a range of 0 to 1. For example, the parameter generator (310) can randomly generate a parameter related to background noise within a range of 0 to 1.
[0064] In one embodiment, the audio signal generator (320) can generate audio signals (i.e., input audio signal (341) and output audio signal (345)) using one or more of a sound source DB (331), a multi-channel impulse response DB (333), a multi-channel noise DB (335), or a parameter generator (310).
[0065] Below, the operation of the audio signal generator (320) to generate an input audio signal (341) is described.
[0066] In one embodiment, the audio signal generator (320) can generate an input audio signal by simulating / calculating an audio signal as if it were received through microphones from one or more sound sources in a three-dimensional space, using a sound source DB (331) and a multi-channel impulse response DB (333).
[0067] In one embodiment, the audio signal generator (320) may generate a sound signal (hereinafter, referred to as a direct sound or direct sound signal) along a direct path by using a sound signal selected from a sound source DB (331) and an impulse response in a three-dimensional space selected from a multi-channel impulse response DB (333). In one embodiment, one or more reflection signals (e.g., early reflections and / or late reverberation) for the direct sound (hereinafter, referred to as reflected sounds) may be generated through one or more impulse responses to the selected sound signal. In one embodiment, a specified time delay and / or scaling factor may be applied to one or more reflection signals.
[0068] In one embodiment, the audio signal generator (320) combines (i.e., synthesizes, convolves, or mixes) a direct sound and a reflected sound to generate a noise-free input audio signal (which may also be referred to as an intermediate input audio signal, a first input audio signal, etc.). In one embodiment, the audio signal generator (320) may combine (i.e., synthesizes, convolves, or mixes) the noise-free input audio signal with one or more direct sounds and one or more reflected sounds for each of the direct sounds to generate a single audio signal. In examples, only the direct sounds may be combined.
[0069] In one embodiment, the audio signal generator (320) can generate an input audio signal (341) by combining (i.e., synthesizing, convolving, or mixing) a noise-free input audio signal with (one) noise selected from a multi-channel noise DB (335). In one embodiment, the audio signal generator (320) can generate the input audio signal (341) by combining a noise-free input audio signal and noise according to a specified signal to noise ratio (SNR).
[0070] In one embodiment, the audio signal generator (320) can generate an input audio signal (341) as in mathematical expression 1 below.
[0071]
[0072] In mathematical expression 1, y m (t) may represent an input audio signal (341) acquired at time point t through the mth microphone among M microphones. s n (t) can represent the nth sound signal among N sound signals in the sound source DB (331). h m (r n , t) is the position (r) of the nth sound signal at time t in the multichannel impulse response DB (333). n) can represent the impulse response (or room impulse response) from the mth microphone to the mth microphone. v m (t) can represent the noise (or background noise) acquired at time point t through the mth microphone in the multi-channel noise DB (335).
[0073] Below, the operation of the audio signal generator (320) to generate an output audio signal (345) is described.
[0074] In one embodiment, the audio signal generator (320) can generate an output audio signal by simulating / calculating an audio signal as if it were received through microphones from one or more sound sources in a three-dimensional space through beamforming using a sound source DB (331) and a multi-channel impulse response DB (333).
[0075] In one embodiment, the audio signal generator (320) may generate a sound signal along a direct path (hereinafter, referred to as a direct sound or direct sound signal) using a sound signal selected from a sound source DB (331) and an impulse response in a three-dimensional space selected from a multi-channel impulse response DB (333). In one embodiment, one or more reflection signals (e.g., early reflections and / or late reverberation) for the direct sound (hereinafter, referred to as reflected sounds) may be generated through one or more impulse responses to the selected sound signal. In one embodiment, a specified time delay and / or scaling factor may be applied to one or more reflection signals. In one embodiment, the direct sound and one or more reflected sounds for the output audio signal (345) may be the same as the direct sound and one or more reflected sounds for the input audio signal (341).
[0076] In one embodiment, the audio signal generator (320) can adjust the direct sound and one or more reflections for the direct sound based on beamforming parameters generated (or selected) using the parameter generator (310). In one embodiment, the audio signal generator (320) can identify a spatial gain that causes the audio signal to have a direction and a beamwidth corresponding to the beamforming parameters. In one embodiment, the audio signal generator (320) can adjust the direct sound and one or more reflections for the direct sound by multiplying the direct sound and one or more reflections for the direct sound by the spatial gain. In one embodiment, the spatial gain can be set to have a high spatial gain for selected locations (or directions) and beamwidths, and a low spatial gain for other locations (or directions) and beamwidths.
[0077] In one embodiment, the audio signal generator (320) can generate a noise-free output audio signal (which may also be referred to as an intermediate output audio signal) by combining (i.e., synthesizing, convolving, or mixing) one or more direct sounds and reflections, each of which has been multiplied by a spatial gain, into one audio signal. In one embodiment, the audio signal generator (320) can generate a noise-free output audio signal by combining (i.e., synthesizing, convolving, or mixing) one or more direct sounds to which a spatial gain has been applied (or multiplied) and one or more reflections for each of the direct sounds.
[0078] In one embodiment, the audio signal generator (320) can adjust the size of noise selected from the DB (335) based on parameters related to noise generated (or selected) using the parameter generator (310). The noise selected from the multi-channel noise DB (335) for the output audio signal (345) can be the same as the noise selected from the multi-channel noise DB (335) for the input audio signal (341).
[0079] In one embodiment, the audio signal generator (320) can identify (or determine) a parameter as in Equation 2 below using the parameter generator (310).
[0080]
[0081] In mathematical expression 2, d can represent a parameter generated (or selected) using a parameter generator (310). is the azimuth based on the electronic device (101) and elevation angle It can represent a unit vector (or a parameter related to direction) based on . can represent a parameter related to the beam width within the range of 0 to 1. g v can represent a parameter related to background noise within the range of 0 to 1.
[0082] In one embodiment, the audio signal generator (320) can generate an output audio signal (345) by combining (i.e., synthesizing, convolving, or mixing) a noise-free output audio signal with scaled noise.
[0083] In one embodiment, the audio signal generator (320) can generate an output audio signal (345) as in mathematical expression 3 below.
[0084]
[0085] In mathematical expression 3, z m,d (t) can represent the output audio signal (345) at time point t through the mth microphone among M microphones adjusted based on the parameter (d). f(h m (r n , t), d) is the position (r) of the nth sound signal at time t in the multi-channel impulse response DB (333) adjusted based on the parameter (d). n) impulse response from the mth microphone (h m (r n , t)) can be represented. s n (t) can represent the nth sound signal among N sound signals in the sound source DB (331). g v may represent a parameter related to background noise within the range of 0 to 1. v(t) may represent noise (or background noise) acquired at time point t in the multi-channel noise DB (335).
[0086] As described above, the output audio signal (345) may be an audio signal in which the beam is adjusted (or steered, changed) and / or noise is reduced from the input audio signal (341) (i.e., in the form of an input audio signal to which beamforming has been applied). As described below, the input audio signal and the output audio signal may be used as training data for training an AI model that performs acoustic beamforming. Although the generation of a single input audio signal and a single output audio signal has been described, any number of different pairs of input audio signals and output audio signals may be generated by selecting different signals from different databases and combining signals.
[0087] FIG. 4 is a diagram illustrating an operation for training an AI (artificial intelligence) model based on an input audio signal and an output audio signal in an electronic device according to one embodiment, wherein the AI model is for implementing acoustic beamforming.
[0088] Figure 4 can be explained with reference to Figures 1, 2, and 3.
[0089] The Fourier transformer (401), the operator (403), and the inverse Fourier transformer (405) illustrated in FIG. 4 may be stored as programs in the memory (130). The Fourier transformer (401), the operator (403), and the inverse Fourier transformer (405) may include instructions executable by the processor (120), but are not limited thereto. The Fourier transformer (401), the operator (403), and the inverse Fourier transformer (405) may be implemented as hardware including circuits.
[0090] In one embodiment, the Fourier transformer (401) can convert a time-domain input audio signal (341) into a frequency-domain signal. In one embodiment, the Fourier transformer (401) can convert a time-domain input audio signal (341) into a frequency-domain signal based on a specified Fourier transform algorithm (e.g., short time Fourier transform (STFT)).
[0091] In one embodiment, the Fourier transformer (401) may convert a time-domain input audio signal (341) into a frequency-domain signal for each of a plurality of sample signals of the input audio signal separated by a specified first window length. In one embodiment, two consecutive sample signals may overlap each other by an offset. In one embodiment, the time length (or offset) over which two consecutive sample signals overlap may be half of the first window length. For example, when the input audio signal (341) is separated into N (N is a natural number) sample signals of a specified first window length, the signal of the front half of the k (k is a natural number less than or equal to N-1)-th sample signal may overlap with the signal of the back half of the k-1-th sample signal, and the signal of the back half of the k-th sample signal may overlap with the signal of the front half of the k+1-th sample signal. However, the present invention is not limited thereto. The degree of overlap may be set differently.
[0092] The AI (artificial intelligence) model (410), loss measurer (470), and AI model learner (480) illustrated in FIG. 4 may be stored as programs in the memory (130). The AI model (410), loss measurer (470), and AI model learner (480) may include instructions executable by the processor (120). In one embodiment, the AI model (410) may include an encoder (420), a mask generator (430), and a domain transformer (440).
[0093] In one embodiment, the encoder (420) can convert an input audio signal (341) in the time domain into an audio signal in the latent domain. In one embodiment, the encoder (420) can convert an input audio signal (341) in the time domain into an audio signal in the latent domain using a specified number of layers (or convolution layers). In one embodiment, the latent domain can be a domain different from the time domain or the frequency domain.
[0094] In one embodiment, the encoder (420) may convert a plurality of time-domain sample signals separated by a specified second window length into audio signals in the latent domain. In one embodiment, the second window length may be shorter than the first window length. In particular, in order to preserve more spatial information (or spatial features) of the input audio signal (341), the length of the second window may be shorter than the length of the first window. Accordingly, the length (or frame length) of the sample signal converted from the time domain to the latent domain through the encoder (420) may be shorter than the length (or frame length) of the sample signal converted from the time domain to the frequency domain through the Fourier transformer (401). The spatial information may include information such as the location of the sound source of the sample signal, time delay, and / or phase delay. In other words, the sampling rate in the latent domain is higher than the sampling rate in the frequency domain.
[0095] In one embodiment, the mask generator (430) may include a plurality of parameters related to a neural network having a structure based on an encoder and a decoder, such as a transformer. In one embodiment, the mask generator (430) may include parameters for driving a neural network such as a convolutional neural network (CNN), a recurrent neural network (RNN), a temporal convolutional network (TCN), a feedforward neural network (FNN), and / or a long short-term memory (LSTM), but is not limited thereto. In one embodiment, the mask generator (430) may include a bi-directional model based on learning for an encoder (e.g., bidirectional encoder representations from transformers (BERT)), or an auto-encoding model (e.g., a diffusion model). In one embodiment, the mask generator (430) may include an auto-regressor model (e.g., a generative pre-trained transformer (GPT)) based on learning about the decoder. In one embodiment, the mask generator (430) may include a sequence-to-sequence model (e.g., stable diffusion, DALL-E 2) based on learning about the encoder and decoder.
[0096] In one embodiment, the mask generator (430) can generate a mask (or filter) of the latent region based on the audio signal of the latent region and the parameter (450). In one embodiment, the mask generator (430) can generate a mask (or filter) of the latent region for a plurality of sample signals of the latent region based on the plurality of audio signals of the latent region and the parameter (450). In one embodiment, the parameter (450) can be a beam-related parameter and / or a noise-related parameter obtained (or generated) from the parameter generator (310), such as the parameter(s) described in relation to Equation 2.
[0097] In one embodiment, the domain converter (440) can convert a mask (or filter) of a latent domain generated by the mask generator (430) into a mask (or filter) of a frequency domain. In one embodiment, the domain converter (440) can convert a mask (or filter) for a plurality of sample signals generated by the mask generator (430) into a mask (or filter) of a frequency domain.
[0098] In one embodiment, the domain converter (440) may convert a mask (or filter) of a frequency domain by adding (or concatenating) a number of masks (or filters) of a latent domain generated by the mask generator (430) corresponding to the ratio between the length (or frame length) of a sample signal converted from the time domain to the latent domain through the encoder (420) and the length (or frame length) of a sample signal converted from the time domain to the frequency domain through the Fourier transformer (401) (i.e., so that the sample length based on the masks output by the AI model corresponds to the sample length based on the Fourier transform). In one embodiment, the domain converter (440) may convert into a mask (or filter) of the frequency domain by adding (or concatenating) a number of masks (or filters) of the latent domain generated by the mask generator (430) corresponding to the ratio between the length of the specified second window and the length of the specified first window and the degree of overlap (or the size of the offset) of the specified second window within the specified first window.
[0099] In one embodiment, the operator (403) may apply an output result (or mask) of an AI (artificial intelligence) model (410) to an input audio signal (341) in a frequency domain. In one embodiment, the operator (403) may generate (or obtain) an audio signal (i.e., a frequency domain AI adjustment output signal) in which a beam of the input audio signal (341) is adjusted (i.e., steered or changed) and noise is reduced, based on the output result (or mask) of the AI model (410).
[0100] In one embodiment, the operator (403) may generate (or obtain) a plurality of sample signals in which a beam of the input audio signal (341) is adjusted (i.e., steered or changed) and noise is reduced based on the output result (or mask) of the AI model (410). In one embodiment, the output results (or masks) of the AI model (410) applied to the plurality of sample signals may be different from each other. For example, a first output result (or first mask) applied to a first sample signal may have a different value from a second output result (or second mask) applied to a second sample signal.
[0101] In one embodiment, the inverse Fourier transformer (405) can transform a frequency-domain audio signal (i.e., a frequency-domain AI conditioning output signal) into a time-domain output signal (460) (i.e., a time-domain AI conditioning output signal). In one embodiment, the Fourier transformer (401) can transform the frequency-domain audio signal into the time-domain output signal (460) based on a specified inverse Fourier transform algorithm (e.g., an inverse STFT (ISTFT)). In one embodiment, the inverse Fourier transformer (405) can generate the time-domain output signal (460) based on a plurality of sample signals whose beams are adjusted (or steered) (or changed) and whose noise is reduced.
[0102] In one embodiment, the loss meter (470) can measure the loss (i.e., difference, error, etc.) between the output signal (460) and the output audio signal (345). In one embodiment, the loss between the output signal (460) and the output audio signal (345) can be based on the signal to noise ratio (SNR) (or the scale invariant signal to distortion ratio (SI-SDR) (or the mean square error).
[0103] In one embodiment, the AI model learner (480) may train the AI model (410) based on the measured loss. In one embodiment, the AI model learner (480) may train the AI model (410) based on a backpropagation algorithm. For example, the AI model learner (480) may update the parameters of the AI model (410) so that the measured loss is less than or equal to a reference loss. For example, the AI model learner (480) may update the parameters of the encoder (420), the mask generator (430), and / or the domain transformer (440) of the AI model (410) so that the measured loss is less than or equal to a reference loss.
[0104] According to one embodiment, the electronic device (101) can train an AI model (410) based on the generated input audio signal (341) and output audio signal (345).
[0105] According to one embodiment, the electronic device (101) can train an AI model (410) that generates the same mask regardless of the audio channels. Accordingly, the electronic device (101) can train the AI model (410) to generate a mask in a specified direction (e.g., a specified azimuth) The specified elevation angle A stereo audio signal in which spatial cues are maintained can be obtained for a sound source.
[0106] According to one embodiment, the electronic device (101) has an adjustable direction (e.g., a specified azimuth) The specified elevation angle And / or parameters related to noise can be used to train the AI model (410). Accordingly, the electronic device (101) can form a sharper (or narrower beamwidth) frequency invariant beam compared to conventional acoustic beamforming. The operation described with reference to FIG. 4 can be performed by an electronic device that can subsequently perform AI beamforming, or can be performed separately from an electronic device that performs actual beamforming and then provided to the electronic device. In other words, the training can be performed by any suitable device, and is not limited to being performed by an electronic device that performs beamforming for actual audio signals received through microphones of the electronic device.
[0107] FIG. 5 is a diagram illustrating an operation for acoustic beamforming an audio signal in an electronic device according to one embodiment. FIG. 6 is a diagram illustrating a three-dimensional space surrounding an electronic device. FIG. 7 is a diagram illustrating an operation for outputting an audio signal formed through acoustic beamforming through an electronic device.
[0108] Figure 5 can be explained with reference to Figures 1, 2, 3, and 4.
[0109] The Fourier transformer (401), the operator (403), and the inverse Fourier transformer (405) illustrated in FIG. 5 may correspond to the Fourier transformer (401), the operator (403), and the inverse Fourier transformer (405) of FIG. 4, respectively. The AI model (410), the loss measurer (470), and the AI model learner (480) illustrated in FIG. 5 may correspond to the AI model (410), the loss measurer (470), and the AI model learner (480) of FIG. 4, respectively. In one embodiment, the AI model (410) of FIG. 5 may be an AI model (410) learned through the learning operation described with reference to FIG. 4. In the descriptions of FIGS. 5, 6, and 7, an audio signal (501) received through a microphone(s) may be referred to as an input audio signal, a received audio signal, a received input audio signal, a real input audio signal, or a real audio signal. The audio signal (505) may be referred to as an AI-controlled reception audio signal, an AI-controlled output audio signal, an AI beamforming output audio signal, or an AI beamforming audio signal.
[0110] In one embodiment, the Fourier transformer (401) may convert a time-domain audio signal (501) (i.e., a time-domain received audio signal) into a frequency-domain audio signal (i.e., a frequency-domain received audio signal). In one embodiment, the Fourier transformer (401) may convert the time-domain audio signal (501) into a frequency-domain audio signal based on a specified Fourier transform algorithm (e.g., STFT). In one embodiment, the audio signal (501) may include audio signals generated from different sound sources. For example, referring to FIG. 6, the audio signal (501) may be transmitted in different directions (e.g., azimuth) in three-dimensional space. and / or elevation angle It may include audio signals generated from sound sources (610, 620, 630) in the direction according to the direction.
[0111] In one embodiment, the Fourier transformer (401) can transform a time domain audio signal (501) into a frequency domain for each of a plurality of sample signals separated by a specified first window length.
[0112] In one embodiment, the encoder (420) can convert a time-domain audio signal (501) into a latent domain. In one embodiment, the encoder (420) can convert the time-domain audio signal (501) into a latent domain using a specified number of layers (or convolution layers). In one embodiment, the encoder (420) can convert a plurality of time-domain sample signals separated by a specified second window length into audio signals in the latent domain. Two consecutive sample signals can overlap each other by an offset. In one embodiment, the time length (or offset) over which two consecutive sample signals overlap can be half of the second window length. However, the present invention is not limited thereto. In one embodiment, the second window length can be shorter than the first window length.
[0113] In one embodiment, the mask generator (430) may generate a mask (or filter) of a latent region based on an audio signal (501) in the form of a latent region and a parameter (550). In one embodiment, the mask generator (430) may generate a mask (or filter) of a latent region for a plurality of sample signals of the latent region based on a plurality of audio signals of the latent region and the parameter (550).
[0114] In one embodiment, the parameter (550) may be a parameter related to a beam identified by an input and / or a parameter related to noise. For example, the parameter (550) may be identified based on a user input to the electronic device (101). For example, the parameter (550) may be identified based on a user input to the electronic device (101) selecting at least one sound source from among sound sources in a three-dimensional space. In one embodiment, the position (or direction) of each of the sound sources in the three-dimensional space may be identified by the electronic device (101) by a time delay and / or phase delay of an audio signal acquired by the plurality of microphones (271, 273, 275). The position of the identified position relative to the electronic device may then be used to determine the parameter (550).
[0115] For example, the parameter (550) may be identified based on a user input for selecting at least one object from among one or more objects included in an image displayed through the display (260) of the electronic device (101). For example, referring to FIG. 7, while playing an image related to an audio signal (710) including multi-channel audio signals (711, 715) for a human voice and a dog sound, the electronic device (101) may identify parameters for beamforming for the human voice and / or the dog sound based on a user input for selecting the human voice and / or the dog sound. For example, based on the user input for selecting the human voice, the electronic device (101) may identify parameters for beamforming the audio signal (710) including multi-channel audio signals (711, 715) into an audio signal (720) including multi-channel audio signals (721, 725) for only the human voice. For example, based on a user input for selecting a dog, the electronic device (101) can identify parameters for beamforming an audio signal (710) including multi-channel audio signals (711, 715) into an audio signal (730) including multi-channel audio signals (731, 735) for only the sound of a dog. For example, based on a user input for selecting a person and a dog, the electronic device (101) can identify parameters for beamforming an audio signal (710) including multi-channel audio signals (711, 715) into an audio signal (740) including multi-channel audio signals (741, 745) for only the sound of a human voice and a dog. In other words, the position of the selected object (e.g., a dog or a person) is identified relative to the electronic device, and beamforming parameters (e.g., beam angles, etc.) are then identified based on this relative position. However, the present invention is not limited thereto.
[0116] In one embodiment, the domain converter (440) can convert a mask (or filter) of a latent domain generated by the mask generator (430) into a mask (or filter) of a frequency domain. In one embodiment, the domain converter (440) can convert a mask (or filter) for a plurality of sample signals generated by the mask generator (430) into a mask (or filter) of a frequency domain.
[0117] In one embodiment, the domain converter (440) can convert into a mask (or filter) of the frequency domain by adding (or concatenating) a number of masks (or filters) of the latent domain generated by the mask generator (430) corresponding to the ratio between the length (or frame length) of the sample signal converted from the time domain to the latent domain through the encoder (420) and the length (or frame length) of the sample signal converted from the time domain to the frequency domain through the Fourier transformer (401). In one embodiment, the domain converter (440) can convert into a mask (or filter) of the frequency domain by adding (or concatenating) a number of masks (or filters) of the latent domain generated by the mask generator (430) corresponding to the ratio between the length of the specified second window and the length of the specified first window and the degree of overlap (or magnitude of the offset) of the specified second window within the specified first window.
[0118] In one embodiment, the operator (403) may apply an output result (or mask) of the AI model (410) to an audio signal (501) in a frequency domain. In one embodiment, the operator (403) may generate (or obtain) an audio signal (i.e., an AI beamforming output audio signal) in a frequency domain in which a beam of the audio signal (501) is adjusted (or steered) (or changed) and noise is reduced, based on the output result (or mask) of the AI model (410). In one embodiment, the operator (403) may generate (or obtain) a plurality of sample signals in which a beam of the audio signal (501) is adjusted and noise is reduced, based on the output result (or mask) of the AI model (410).
[0119] In one embodiment, the inverse Fourier transformer (405) can transform a frequency-domain audio signal (i.e., a frequency-domain AI-conditioned received audio signal) into a time-domain audio signal (505) (i.e., a time-domain AI-conditioned received audio signal). In one embodiment, the Fourier transformer (401) can transform the frequency-domain audio signal into the time-domain audio signal (505) based on a specified inverse Fourier transform algorithm (e.g., ISTFT). In one embodiment, the inverse Fourier transformer (405) can generate the time-domain audio signal (505) based on a plurality of sample signals whose beams are adjusted (or steered) (or changed) and whose noise is reduced.
[0120] According to one embodiment, the electronic device (101) can convert an audio signal (501) (i.e., a received audio signal) obtained through a small number of microphones (271, 273, 275) into a beamformed audio signal (505) (i.e., an AI-controlled received audio signal) through an AI model (410).
[0121] In one embodiment, the electronic device (101) applies the same mask generated through the AI model (410) to the audio signal regardless of the audio channels, thereby detecting a specified direction (e.g., a specified azimuth). The specified elevation angle A stereo audio signal in which spatial cues are maintained can be obtained for a sound source.
[0122] According to one embodiment, the electronic device (101) is configured to have an adjustable direction (e.g., a specified azimuth) through user input. The specified elevation angle and / or parameters related to noise can be used to form sharper (or narrower beamwidth) frequency invariant beams.
[0123] FIG. 8A is a flowchart illustrating an operation of an electronic device acquiring an input audio signal in a manner similar to that described with reference to Equation 1, according to one embodiment. FIG. 8B is a diagram illustrating a path of an audio signal acquired by the electronic device. FIG. 8C is a graph illustrating an audio signal acquired by the electronic device along a path of the audio signal.
[0124] FIG. 8a, FIG. 8b, and FIG. 8c can be explained with reference to FIG. 1, FIG. 2, and FIG. 3.
[0125] Referring to FIG. 8A, in operation 810, the electronic device (101) may perform random sampling. For example, the electronic device (101) may randomly sample one or more sound sources among a plurality of sound sources stored in a sound source DB (331). For example, the electronic device (101) may randomly sample one or more impulse responses among a plurality of impulse responses stored in a multi-channel impulse response DB (333). For example, the electronic device (101) may randomly sample one noise among a plurality of noises stored in a multi-channel noise DB (335).
[0126] In operation 820, the electronic device (101) may set a path of a randomly sampled audio signal. In particular, the electronic device (101) may set a direct path and / or one or more reflection paths for each of one or more sound sources based on paths randomly selected from a multi-channel impulse response DB (333). For example, the reflection path may be a path of an audio signal to the electronic device (101) from at least a slightly different location from the location of the sound source along the direct path. In one embodiment, the location of the sound source for the direct path and the location of the sound source for the reflection path are determined by an azimuth angle relative to the electronic device (101). , and elevation angle can be expressed as
[0127] In one embodiment, the electronic device (101) may generate a sound signal along a direct path (hereinafter, referred to as direct sound) and reflection signals (e.g., early reflections and / or late reverberation) along one or more reflection paths for the direct sound (hereinafter, referred to as reflected sounds). In one embodiment, the electronic device (101) may generate an audio signal (i.e., a noise-free input audio signal) including the direct sound and one or more reflections for the direct sound. Referring to FIG. 8B, the audio signal may include a direct sound (840) along a direct path from a sound source (801) to the electronic device (101), and reflections (851, 855, 859) along paths reflected through reflectors (e.g., a wall, a ceiling, a floor, or an obstacle).
[0128] In one embodiment, the electronic device (101) may apply a specified time delay and / or scaling factor to each reflection signal of one or more reflection paths. Referring to FIG. 8C, reflections (850) of an audio signal may be expressed as being acquired after a specified delayed time from the direct sound (840) to the plurality of microphones (271, 273, 275). Referring to FIG. 8C, the magnitude of the reflections (850) of the audio signal may decrease over time. In FIG. 8C, reflections acquired within a specified time from the time at which the direct sound (840) was acquired may be referred to as early reflections. In FIG. 8C, reflections acquired after a specified time from the time at which the direct sound (840) was acquired may be referred to as later reverberations.
[0129] In operation 830, the electronic device (101) may obtain an input audio signal by adding background noise to a routed audio signal. In one embodiment, the electronic device (101) may generate an input audio signal (341) by combining (i.e., synthesizing, convolving, or mixing) noise randomly sampled from a multi-channel noise DB (335) with an audio signal in which direct sound and reflected sound are combined. In one embodiment, the electronic device (101) may generate the input audio signal (341) by combining noise and noise-free input audio signals according to a specified signal-to-noise ratio (SNR). The approach of generating input audio signals based on the random sampling described above increases the number of input audio signals generated from the sound source database (331) and the multi-channel impulse response DB (333), thereby increasing the size of learning data to be generated.
[0130] FIG. 9 is a flowchart illustrating an operation of an electronic device to obtain an output audio signal in a manner similar to that described with reference to Equation 3, according to one embodiment.
[0131] FIG. 9 can be described with reference to FIGS. 1, 2, and 3. Operations 810 and 820 of FIG. 9 may correspond to operations 810 and 820 of FIG. 8a, respectively.
[0132] Referring to FIG. 9, in operation 810, the electronic device (101) can perform random sampling. In operation 820, the electronic device (101) can set a path for the randomly sampled audio signal.
[0133] In operation 910, the electronic device (101) can adjust the routed audio signal based on different gains for each path. In one embodiment, the electronic device (101) can adjust the direct sound and one or more reflections for the direct sound based on parameters associated with a beam of an audio signal generated (or selected) using the parameter generator (310) to form a noise-free output audio signal. In one embodiment, the electronic device (101) can identify a spatial gain that causes the audio signal to have a direction and a beamwidth corresponding to the parameters associated with the beam. In one embodiment, the electronic device (101) can adjust the direct sound and one or more reflections for the direct sound by multiplying the spatial gain by the direct sound and one or more reflections for the direct sound. In one embodiment, the spatial gain can be set to have a high spatial gain for selected locations (or directions) and beam widths, and a low spatial gain for other locations (or directions) and beam widths.
[0134] In operation 920, the electronic device (101) may obtain an output audio signal by adding adjusted background noise to an audio signal whose gain is adjusted. In one embodiment, the electronic device (101) may determine that the noise in operation 920 may be the same as the noise in operation 830 of FIG. 8A based on a parameter related to noise generated (or selected) using the parameter generator (310). For example, the noise in operation 920 and the noise in operation 830 of FIG. 8A may be noise selected from the multi-channel noise DB (335). For example, the noise in operation 920 may be noise adjusted based on a parameter related to noise in operation 830.
[0135] FIG. 10 is a flowchart illustrating an operation of an electronic device for acoustic beamforming an audio signal (i.e., a received audio signal) according to one embodiment.
[0136] Figure 10 can be explained with reference to Figures 1, 2, and 5.
[0137] Referring to FIG. 10, in operation 1001, the electronic device (101) may identify an input for beamforming of an audio signal. For example, the electronic device (101) may identify a user input for selecting at least one sound source among sound sources in a three-dimensional space. For example, the electronic device (101) may identify a user input for selecting at least one object among one or more objects included in an image displayed through the display (260) of the electronic device (101). In one embodiment, the object may be a sound source (or a speaker). However, the present invention is not limited thereto.
[0138] In one embodiment, the electronic device (101) can identify a user input for noise control. For example, the electronic device (101) can identify a user input for noise control through a controllable object (or visual object) (or bar) in an image displayed through the display (260) of the electronic device (101).
[0139] Depending on the embodiment, operation 1001 may not be performed. For example, the electronic device (101) may determine whether to perform beamforming of an audio signal without input. For example, the electronic device (101) may determine to perform beamforming when the audio signal includes multiple sound sources. However, the present invention is not limited thereto.
[0140] In operation 1010, the electronic device (101) can encode an audio signal into an audio signal in a latent region. The electronic device (101) can encode an audio signal in a time domain into an audio signal in a latent region.
[0141] In one embodiment, the electronic device (101) can convert a time-domain audio signal into a latent-domain audio signal using a specified number of layers (or convolution layers). In one embodiment, the electronic device (101) can convert a plurality of time-domain sample signals into latent-domain audio signals.
[0142] In operation 1020, the electronic device (101) can identify a mask based on an audio signal of a potential area.
[0143] In one embodiment, the electronic device (101) may generate a mask (or filter) of a latent region based on audio signals and parameters of the latent region. In one embodiment, the electronic device (101) may generate a mask (or filter) of a latent region for a plurality of sample signals of the latent region based on a plurality of audio signals and parameters of the latent region.
[0144] In one embodiment, the parameter may be a parameter associated with a beam identified by a user input and / or a parameter associated with noise. For example, the parameter may be identified based on a user input to the electronic device (101). For example, the parameter may be identified based on a user input to the electronic device (101) selecting at least one sound source among sound sources in a three-dimensional space and parameters corresponding to a beam directed toward the at least one sound source. For example, the parameter may be identified based on a user input to the electronic device (101) selecting a degree of noise control.
[0145] In operation 1030, the electronic device (101) can obtain an audio signal with a changed direction or beam width by applying a mask to the audio signal.
[0146] In one embodiment, the electronic device (101) can apply a mask to an audio signal in the frequency domain. In one embodiment, the electronic device (101) can adjust (or steer) (or change) a beam of the audio signal based on the mask, and generate (or obtain) an audio signal with reduced noise.
[0147] In one embodiment, the electronic device (101) can obtain a frequency domain mask through a specified number of masks of latent regions. In one embodiment, the specified number can correspond to a ratio between the length (or frame length) of a sample signal of the latent region and the length (or frame length) of a sample signal converted into the frequency domain. In one embodiment, the electronic device (101) can generate (or obtain) an audio signal in which a beam of an audio signal is adjusted (or steered) (or changed) and noise is reduced based on the masks of the frequency domain.
[0148] According to one embodiment, an electronic device (101) may include a plurality of microphones (271, 273, 275), a processor (120), and a memory (130) storing instructions. The instructions, when executed by the processor (120), may cause the electronic device (101) to encode a time-domain input audio signal (501) obtained through the plurality of microphones (271, 273, 275) into an audio signal of a latent domain. The instructions, when executed by the processor (120), may cause the electronic device (101) to identify a mask for an audio signal of a frequency domain of the input audio signal (501) based on the audio signal of the latent domain. The above instructions, when executed by the processor (120), may cause the electronic device (101) to obtain an output audio signal (505) in a time domain in which at least one of a direction of the input audio signal (501) or a beamwidth of the input audio signal (501) is changed based on applying the mask to the audio signal in the frequency domain.
[0149] The instructions, when executed by the processor (120), may cause the electronic device (101) to obtain an input for determining the direction or the beam width. The instructions, when executed by the processor (120), may cause the electronic device (101) to identify the mask through an artificial intelligence (AI) model (410) based on the input and an audio signal of the potential area.
[0150] The above instructions, when executed by the processor (120), may cause the electronic device (101) to obtain the mask of the frequency domain based on an output mask of the latent domain of the AI model (410). The above instructions, when executed by the processor (120), may cause the electronic device (101) to obtain the output audio signal (505) based on applying the mask of the frequency domain to an audio signal of the frequency domain.
[0151] The first frame length of the audio signal of the potential region may be shorter than the second frame length of the audio signal of the frequency region. The instructions, when executed by the processor (120), may cause the electronic device (101) to obtain the mask of the frequency region based on a number of output masks of the AI model (410) corresponding to a ratio between the second frame length and the first frame length.
[0152] The electronic device (101) may include a camera module (180). The instructions, when executed by the processor (120), may cause the electronic device (101) to obtain an input for selecting at least one object included in an image recorded through the camera module (180). The instructions, when executed by the processor (120), may cause the electronic device (101) to obtain a first parameter for the direction and a second parameter for the beam width based on the input. The instructions, when executed by the processor (120), may cause the electronic device (101) to obtain the mask by inputting the first parameter, the second parameter, and an audio signal of the potential region into the AI model (410).
[0153] The above AI model (410) can be trained by an input audio signal and an output audio signal generated based on a parameter indicating a beam width and a parameter indicating a direction.
[0154] The above AI model (410) can be trained by the input audio signal and the output audio signal generated based on a parameter indicating the degree of noise control.
[0155] The audio signal in the frequency domain may include sub-audio signals in the frequency domain corresponding to each of a plurality of channels. The instructions, when executed by the processor (120), may cause the electronic device (101) to obtain the output audio signal (505) including output sub-audio signals corresponding to each of the plurality of channels based on applying the mask to each of the sub-audio signals.
[0156] The electronic device (101) may include a plurality of speakers (251, 253, 255). The instructions, when executed by the processor (120), may cause the electronic device (101) to output the output audio signal (505) using the plurality of speakers (251, 253, 255).
[0157] The above instructions, when executed by the processor (120), may cause the electronic device (101) to obtain the output audio signal (505) in which the noise of the input audio signal (501) is adjusted based on applying the mask to the audio signal in the frequency domain.
[0158] As described above, the electronic device (101) may include a plurality of microphones (271, 273, 275), a processor (120), and a memory (130) storing instructions. The instructions, when executed by the processor (120), may cause the electronic device (101) to obtain, through an artificial intelligence (AI) model (410), an output audio signal in which at least one of a direction or a beamwidth is changed from an input audio signal (501) obtained through the plurality of microphones (271, 273, 275). The AI model (410) may be trained by a training output audio signal generated based on at least one parameter of a training input audio signal and a first parameter indicating a beamwidth of the audio, or a second parameter indicating a direction of the audio.
[0159] The input audio signal for training may include a direct path sound source, a plurality of early reflection signals for the sound source, a plurality of late reverberation signals for the sound source, and noise. The output audio signal for training may be an audio signal obtained by applying a gain according to at least one parameter to the input audio signal.
[0160] The above AI model (410) can be trained by the training output audio signal generated based on the third parameter indicating the degree of noise control.
[0161] The above AI model (410) can be trained so that the difference between the output signal for the input audio signal of the AI model (410) and the output audio signal is reduced.
[0162] The above difference may be a scale invariant signal-to-distortion ratio (SI-SDR) (or mean square error) between the output signal and the training output audio signal.
[0163] As described above, the electronic device (101) may include a processor (120), and a memory (130) storing a sound source database (DB) 331, a multi-channel impulse response DB (333), a multi-channel noise DB (335), and instructions. The instructions, when executed by the processor (120), may cause the electronic device (101) to select at least one sound signal from the sound source DB (331). The instructions, when executed by the processor (120), may cause the electronic device (101) to generate an audio signal including direct sound and reflected sound for at least one sound signal based on the multi-channel impulse response DB (333). The instructions, when executed by the processor (120), may cause the electronic device (101) to generate an input audio signal in which noise selected based on the multi-channel noise DB (335) and the audio signal are synthesized. The instructions, when executed by the processor (120), may cause the electronic device (101) to generate an output audio signal in which at least one of a direction of the audio signal or a beamwidth of the input audio signal is changed and the noise is reduced. The instructions, when executed by the processor (120), may cause the electronic device (101) to train the artificial intelligence (AI) model (410) so that a difference between an output signal for the input audio signal and the output audio signal using the AI model (410) is reduced.
[0164] For example, the output audio signal may have the noise of the input audio signal adjusted based on a parameter indicating the degree of noise adjustment.
[0165] For example, the difference may be the mean square error between the output signal and the training output audio signal.
[0166] As described above, the method can be performed by an electronic device (101) including a plurality of microphones (271, 273, 275). The method can include an operation of encoding a first audio signal (501) of a time domain obtained through the plurality of microphones (271, 273, 275) into a second audio signal of a latent domain. The method can include an operation of identifying a mask for a third audio signal of a frequency domain of the first audio signal (501) based on the audio signal of the latent domain. The method can include an operation of obtaining an output audio signal (505) of a time domain in which at least one of a direction of the first audio signal (501) or a beamwidth of the first audio signal (501) is changed based on applying the mask to the audio signal of the frequency domain.
[0167] The method may include an operation of obtaining an input for determining the direction and the beam width. The method may include an operation of identifying the mask through an artificial intelligence (AI) model (410) based on the input and an audio signal of the potential region.
[0168] The audio signal in the frequency domain includes sub-audio signals in the frequency domain corresponding to each of a plurality of channels, and the method may include an operation of obtaining the output audio signal (505) including output sub-audio signals corresponding to each of the plurality of channels based on applying the mask to each of the sub-audio signals.
[0169] The electronic device (101) may include a plurality of speakers (251, 253, 255). The method may include an operation of outputting the output audio signal (505) using the plurality of speakers (251, 253, 255).
[0170] The method may include an operation of obtaining an output audio signal (505) in which noise of the input audio signal (501) is adjusted based on applying the mask to the audio signal in the frequency domain.
[0171] As described above, the method can be performed by an electronic device (101) including a plurality of microphones (271, 273, 275). The method can include an operation of obtaining, through an artificial intelligence (AI) model (410), an output audio signal in which at least one of a direction or a beam width is changed from an input audio signal (501) obtained through the plurality of microphones (271, 273, 275). The AI model (410) can be trained by a training output audio signal generated based on at least one parameter of a training input audio signal and a first parameter indicating a beam width of the audio or a second parameter indicating a direction of the audio.
[0172] As described above, the method can be performed by an electronic device (101) including a memory (130) storing a sound source database (DB) 331, a multi-channel impulse response DB (333), and a multi-channel noise DB (335). The method can include an operation of selecting at least one sound signal from the sound source DB (331). The method can include an operation of generating an audio signal including direct sound and reflected sound for the at least one sound signal based on the multi-channel impulse response DB (333). The method can include an operation of generating an input audio signal in which the selected noise and the audio signal are synthesized based on the multi-channel noise DB (335). The method can include an operation of generating an output audio signal in which at least one of a direction of the audio signal or a beamwidth of the input audio signal is changed and the noise is reduced. The above method may include an operation of training the AI (artificial intelligence) model (410) so that the difference between the output signal for the input audio signal and the output audio signal is reduced using the AI model (410).
[0173] As described above, a non-transitory computer readable storage medium can store a program including instructions. The instructions, when executed by a processor (120) of an electronic device (101) including a plurality of microphones (271, 273, 275), can cause the electronic device (101) to encode a first audio signal (501) in a time domain obtained through the plurality of microphones (271, 273, 275) into a second audio signal in a latent domain. The instructions, when executed by the processor (120), can cause the electronic device (101) to identify a mask for a third audio signal in a frequency domain of the first audio signal (501) based on the audio signal in the latent domain. The above instructions, when executed by the processor (120), may cause the electronic device (101) to obtain an output audio signal (505) in a time domain in which at least one of a direction of the first audio signal (501) or a beamwidth of the first audio signal (501) is changed based on applying the mask to the audio signal in the frequency domain.
[0174] As described above, a non-transitory computer-readable recording medium can store a program including instructions. When executed by a processor (120) of an electronic device (101) including a plurality of microphones (271, 273, 275), the instructions can cause the electronic device (101) to obtain, through an artificial intelligence (AI) model (410), an output audio signal in which at least one of a direction or a beam width is changed from an input audio signal (501) obtained through the plurality of microphones (271, 273, 275). The AI model (410) can be trained by a training output audio signal generated based on at least one parameter of a training input audio signal and a first parameter indicating a beam width of the audio, or a second parameter indicating a direction of the audio.
[0175] As described above, a non-transitory computer readable storage medium can store a program including instructions. The instructions, when executed by a processor (120) of an electronic device (101) including a sound source database (DB) 331, a multi-channel impulse response DB (333), a multi-channel noise DB (335), and a memory (130) storing instructions, can cause the electronic device (101) to select at least one sound signal from the sound source DB (331). The instructions, when executed by the processor (120), can cause the electronic device (101) to generate an audio signal including a direct sound and a reflected sound for the at least one sound signal based on the multi-channel impulse response DB (333). The instructions, when executed by the processor (120), may cause the electronic device (101) to generate an input audio signal in which noise selected based on the multi-channel noise DB (335) and the audio signal are synthesized. The instructions, when executed by the processor (120), may cause the electronic device (101) to generate an output audio signal in which at least one of a direction of the audio signal or a beamwidth of the input audio signal is changed and the noise is reduced. The instructions, when executed by the processor (120), may cause the electronic device (101) to train the artificial intelligence (AI) model (410) so that a difference between an output signal for the input audio signal and the output audio signal using the AI model (410) is reduced.
[0176] As described above, the method can be performed by an electronic device (101) including a plurality of microphones (271, 273, 275). The method can include an operation of encoding a time-domain input audio signal (501) received through the plurality of microphones (271, 273, 275) into an input audio signal of a latent region, and an time-domain input audio signal (501) acquired through the plurality of microphones (271, 273, 275) into an audio signal of a latent region. The method can include an operation of identifying a mask for performing AI-based beamforming on the input audio signal based on the input audio signal of the latent region, and an operation of identifying a mask for an audio signal of a frequency domain for the input audio signal (501). The method may include an operation of obtaining an output audio signal (505) in a time domain in which at least one of a direction of the input audio signal (501) or a beamwidth of the input audio signal (501) is changed based on applying the mask to the audio signal in the frequency domain, by applying the mask to the audio signal in the frequency domain, such that at least one of a beamforming direction or a beamforming beamwidth of the input audio signal (501) is changed.
[0177] The method may include an operation of obtaining an input for determining the direction and the beam width, the beamforming input parameter including the beamforming direction or the beamforming beam width. The method may include an operation of identifying the mask based on the beamforming input parameter and the input audio signal of the potential region, through an AI (artificial intelligence) model (410), based on the input and the audio signal of the potential region.
[0178] The electronic device (101) may include sub-audio signals in the frequency domain corresponding to each of a plurality of channels in the input audio signal of the frequency domain. The method may include an operation of obtaining the AI beamforming input output audio signal (505) including sub-audio signals in the time domain corresponding to each of the plurality of channels based on applying the mask to each of the sub-audio signals.
[0179] The electronic device (101) may include a plurality of speakers (251, 253, 255). The method may include an operation of outputting the output audio signal (505) using the plurality of speakers (251, 253, 255).
[0180] The method may include an operation of obtaining the AI beamforming input audio signal (501) in which the noise of the input audio signal (501) is adjusted based on applying the mask to the input audio signal in the frequency domain, and the output audio signal (505) in which the noise of the first input audio signal (501) is adjusted based on applying the mask to the audio signal in the frequency domain.
[0181] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0182] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0183] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0184] The various embodiments described herein may be combined in any suitable combination, unless otherwise stated or incompatible. Embodiments that include various features described herein are not limited thereto, and features may be eliminated or introduced, unless otherwise stated.
[0185] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0186] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., by download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0187] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101), processor (120), and It includes a sound source database (DB, 331), a channel impulse response DB (333) for multiple microphones, a channel noise DB (335) and a memory (130) storing instructions, and when the instructions are executed by the processor (120), the electronic device (101) In the above sound source DB (331), at least one sound signal is selected, Based on the above channel impulse response DB (333), a first audio signal including a direct sound signal and a reflected sound signal for at least one sound signal is generated, Generate an input audio signal by combining the noise selected from the above channel noise DB (335) and the first audio signal, reducing the noise of the input audio signal, and generating an output audio signal based on the input audio signal by applying at least one of a beamforming beamwidth or a beamforming direction to the first audio signal of the input audio signal; Based on applying the AI (artificial intelligence) model (410) to the input audio signal, an AI-controlled output audio signal is generated, Based on reducing the difference between the above AI controlled output audio signal and the above output audio signal, causing the AI model (410) to be trained, Electronic devices.
2. In claim 1, The above output audio signal is a signal in which the noise of the input audio signal is adjusted based on a parameter indicating the degree of adjustment of the noise. Electronic devices.
3. In claim 1 or 2, The above difference is the mean square error between the output audio signal and the AI controlled output audio signal. Electronic devices.
4. In any of the preceding claims, Applying the AI model (410) to the input audio signal includes applying a mask generated by the AI model to the input audio signal. Electronic devices.
5. In claim 4, Generating the AI-modulated output audio signal comprises applying the AI model to the input audio signal in the frequency domain. Electronic devices.
6. In claim 5, The above mask is generated in the time domain and the frequency domain and other latent domains by the AI model based on the input audio signal. Electronic devices.
7. In claim 6, The sampling rate of the above potential region is higher than the sampling rate of the above frequency region. Electronic devices.
8. In any of the preceding claims, Including a plurality of microphones (271, 273, 275), The above instructions, when executed by the processor (120), cause the electronic device (101) to: By applying the learned AI model (410) to the input signal (501) received through the plurality of microphones (271, 273, 275), an AI beamforming input signal is obtained. Electronic devices.
9. In the electronic device (101), Multiple microphones (271, 273, 275), processor (120), and It includes a memory (130) that stores instructions, and when the instructions are executed by the processor (120), the electronic device (101) Encode the time domain input audio signal (501) received through the above-mentioned plurality of microphones (271, 273, 275) into a potential domain input audio signal, Based on the input audio signal of the above potential region, a mask for performing AI-based beamforming on the input audio signal is identified, By applying the mask to the audio signal in the frequency domain, at least one of the beamforming direction or the beamforming beamwidth of the input audio signal (501) is caused to obtain an AI beamforming input audio signal applied to the input audio signal. Electronic devices.
10. In claim 9, The above instructions, when executed by the processor (120), cause the electronic device (101) to: Obtaining beamforming input parameters including the beamforming direction or the beamforming beamwidth, causing the mask to be identified based on the beamforming input parameters and the input audio signal of the latent region. Electronic devices.
11. In claim 10, The above instructions, when executed by the processor (120), cause the electronic device (101) to: causing said mask to be converted to said frequency domain for application to said input audio signal in said frequency domain; Electronic devices.
12. In claim 11, The first frame length of the input audio signal in the potential region is shorter than the second frame length of the input audio signal in the frequency domain, The above instructions, when executed by the processor (120), cause the electronic device (101) to: causing the mask to be converted to the frequency domain based on the number of masks corresponding to the ratio between the second frame length and the first frame length, Electronic devices.
13. In claim 10, Contains a camera module (180), The above instructions, when executed by the processor (120), cause the electronic device (101) to: Obtaining an input for selecting at least one object included in the video recorded through the above camera module (180), Based on the above input, the beamforming input parameters are obtained, Based on the above beamforming input parameters, causing the mask of the potential region to be identified, Electronic devices.
14. In any one of claims 9 to 13, The above instructions, when executed by the processor (120), cause the electronic device (101) to: To apply the above mask, the above audio input signal is converted to the above frequency domain, Causing the AI beamforming input audio signal in the frequency domain to be converted into the time domain. Electronic devices.
15. In any one of claims 9 to 14, The above AI model (410) is trained by an artificial input audio signal and an artificial output audio signal generated by applying beamforming parameters indicating beam width and direction to the artificial input audio signal. Electronic devices.
Citation Information
Patent Citations
Centrifugal turbo compressor
KR1020220060075A
Ultra-fine hole processing method
KR1020250018675A
Method for providing improved convenience to caregivers based on smart calendars and devices thereof
KR102454289B1
Eco-friendly paint and manufacturing method thereof
KR102576920B1