Wearable device for detecting speech, and control method therefor

The wearable device enhances voice detection accuracy by using multiple microphones to calculate magnitude and energy ratios, which are then processed through an artificial neural network to identify voices in noisy environments.

WO2025121793A1PCT designated stage expired Publication Date: 2025-06-12SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019245
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-29
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Conventional voice detection technologies in wearable devices, such as wireless earphones, struggle with accuracy in environments with low signal-to-noise ratio (SNR), particularly due to weak voice power and strong noise, including energy echoes and wind noise.

Method used

A wearable device equipped with multiple microphones and a processor that calculates a magnitude spectrum and energy ratio of audio signals, using these features to determine through a trained artificial neural network whether a voice is present in the audio signal.

Benefits of technology

The solution provides improved accuracy in voice detection compared to conventional technologies, effectively handling noise-robust environments and improving performance in challenging SNR conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019245_12062025_PF_FP_ABST
    Figure KR2024019245_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A wearable device for detecting speech is disclosed. The wearable device according to one embodiment of the present document comprises: a first microphone; a second microphone arranged at a position differing from that of the first microphone; at least one processor; and a memory, wherein the memory can store instructions that, when executed, cause the at least one processor to: acquire a magnitude feature of a first audio signal acquired by the first microphone; acquire the energy ratio between the first audio signal and a second audio signal acquired by the second microphone; and, on the basis of the acquired magnitude feature and the energy ratio, determine, through a trained artificial neural network, whether speech is included in the first audio signal or the second audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Wearable device for detecting voice and method for controlling the same

[0001] This document relates to a wearable device for detecting voice and a method for controlling the same.

[0002] The variety of services and additional features offered through wearable devices, such as wireless earphones (e.g., Samsung® Buds®), is steadily increasing. To enhance the utility of these electronic devices and satisfy the diverse needs of users, telecommunications service providers and electronic device manufacturers are competitively developing electronic devices that offer a variety of features and differentiate themselves from competitors. Consequently, the various functions offered through wearable devices are also becoming increasingly sophisticated.

[0003] Conventional voice detection technologies are largely rule-based, extracting hand-crafted features and comparing them to a threshold. However, the performance of these simply designed voice detectors often degrades dramatically in environments with weak speech or relatively strong noise, i.e., low signal-to-noise ratio (SNR). In particular, "energy," one of the most critical voice detection features, is known to generate high false alarms when the noise signal characteristics are non-stationary. Voice detection as preprocessing for various applications, such as call solutions using wireless earbuds, requires even more careful consideration. Unlike mobile phones, where the user's voice is directly transmitted to the microphone from a close distance, wireless earbuds are located farther from the mouth, potentially resulting in a poorer SNR. Additionally, the short distance between the speaker and the microphone exposes the wearable device to strong energy echoes, and other noises such as wind noise can act as factors that make voice detection difficult for the wearable device.

[0004] According to one embodiment of the present document, a wearable device may be provided that performs a noise-resistant voice detection function or operation by determining whether a voice is included in an audio signal acquired by a microphone of the wearable device based on a magnitude spectrum (in other words, a magnitude feature or an energy spectrum) and an energy ratio of an audio signal acquired by a microphone of the wearable device.

[0005] A wearable device according to one embodiment of the present document can provide a wearable device with improved accuracy in voice detection compared to conventional technologies by determining whether a microphone is shielded and calculating a magnitude spectrum and energy ratio based on the determination result.

[0006] A wearable device according to one embodiment of the present document includes a first microphone, a second microphone disposed at a different location from the first microphone, at least one processor, and a memory, wherein the memory may store instructions that, when executed, cause the at least one processor to obtain a magnitude feature of a first audio signal obtained by the first microphone, obtain an energy ratio of the first audio signal and a second audio signal obtained by the second microphone, and determine, based on the obtained magnitude feature and the energy ratio, through a trained artificial neural network, whether a voice is included in the first audio signal or the second audio signal.

[0007] A method for controlling a wearable device according to one embodiment of the present document may include an operation of acquiring a magnitude feature of a first audio signal acquired by a first microphone of the wearable device, an operation of acquiring an energy ratio of the first audio signal and a second audio signal acquired by a second microphone of the wearable device, and an operation of determining, based on the acquired magnitude feature and the energy ratio, through a trained artificial neural network, whether a voice is included in the first audio signal or the second audio signal.

[0008] A computer-readable non-transitory recording medium storing instructions for controlling a wearable device according to one embodiment of the present document may store instructions for obtaining a magnitude feature of a first audio signal acquired by a first microphone of the wearable device, obtaining an energy ratio of the first audio signal and a second audio signal acquired by a second microphone of the wearable device, and determining, based on the obtained magnitude feature and the energy ratio, through a trained artificial neural network, whether a voice is included in the first audio signal or the second audio signal.

[0009] According to one embodiment of the present document, a wearable device may be provided that performs a noise-resistant voice detection function or operation by determining whether a voice is included in an audio signal acquired by a microphone of the wearable device based on a magnitude spectrum (in other words, a magnitude feature or an energy spectrum) and an energy ratio of an audio signal acquired by a microphone of the wearable device.

[0010] A wearable device according to one embodiment of the present document can provide a wearable device with improved accuracy in voice detection compared to conventional technologies by determining whether a microphone is shielded and calculating a magnitude spectrum and energy ratio based on the determination result.

[0011] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments of the present document.

[0012] FIG. 2 is an exemplary drawing for explaining a connection relationship between an electronic device and a wearable device according to one embodiment of the present document.

[0013] FIG. 3 is an exemplary drawing for explaining the configuration of a first wearable device according to one embodiment of the present document.

[0014] FIG. 4A is an exemplary drawing for explaining a function or operation of a wearable device according to one embodiment of the present document to perform voice detection based on characteristics of audio signals acquired through microphones of the wearable device.

[0015] FIG. 4b is an exemplary drawing for explaining a function or operation of a wearable device according to one embodiment of the present document to train an artificial neural network.

[0016] FIG. 5 is an exemplary diagram for explaining a function or operation of a wearable device according to one embodiment of the present document to perform voice detection through a learned artificial neural network based on magnitude features and energy ratio features of audio signals.

[0017] FIG. 6 is an exemplary drawing for explaining a function or operation of calculating an energy ratio based on whether microphones of the wearable device are shielded, according to an embodiment of the present document.

[0018] FIG. 7 is an exemplary diagram for explaining a function or operation of a wearable device according to one embodiment of the present document to process magnitude features and energy ratio features input to a learned artificial neural network.

[0019] FIG. 8 is an exemplary diagram for explaining a function or operation of a wearable device according to one embodiment of the present document to determine an operation ratio of a learned artificial neural network based on a signal-to-noise ratio of an audio signal input through a microphone of the wearable device.

[0020] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.

[0021] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0022] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0023] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0024] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0025] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0026] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0027] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0028] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0029] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0030] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0031] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0032] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0033] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0034] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0035] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0036] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0037] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0038] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0039] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0040] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0041] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0042] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0043] FIG. 2 is an exemplary drawing for explaining a connection relationship between an electronic device (101) and a wearable device (200) (e.g., electronic device (101)) according to one embodiment of the present document.

[0044] Referring to FIG. 2, an electronic device (101) according to one embodiment of the present document can be wirelessly connected to a wearable device (200). The electronic device (101) according to one embodiment of the present document can include a smart phone as illustrated in FIG. 2, but can also be implemented as various types of devices (e.g., a notebook computer including a standard notebook, an ultrabook, a netbook, and a tabbook, a laptop computer, a tablet computer, or a desktop computer).

[0045] A wearable device (200) according to one embodiment of the present document may be implemented as a wireless earphone as illustrated in FIG. 2, but may also be implemented as various types of devices (e.g., a smart watch, a head-mounted display device, or devices for measuring biosignals (e.g., an electrocardiogram patch)). According to one embodiment of the present document, the wearable device (200) may include a pair of devices (e.g., a first wearable device (202) and a second wearable device (204)). According to one embodiment of the present document, the pair of devices (e.g., the first wearable device (202) and the second wearable device (204)) may be implemented to include identical or similar configurations (e.g., the configurations described in FIG. 3).

[0046] According to one embodiment of the present document, the electronic device (101) and the wearable device (200) can establish a communication connection with each other and transmit and / or receive data with each other. For example, the electronic device (101) and the wearable device (200) can establish a communication connection with each other using D2D communication such as Wi-Fi direct or Bluetooth (e.g., using a communication circuit that supports the corresponding communication method), but are not limited thereto and can use various types of communication (e.g., a communication method such as Wi-Fi using an access point (AP), a cellular communication method using a base station, or a wired communication method).

[0047] According to one embodiment of the present document, when the wearable device (200) is a wireless earphone, the electronic device (101) can establish a communication connection with one device (e.g., a main earphone) among a pair of devices (e.g., a first wearable device (202) and a second wearable device (204)), but can also establish a communication connection with both devices (e.g., a first wearable device (202) and a second wearable device (204)). According to one embodiment of the present document, the electronic device (101) can communicate with another device (e.g., a sub earphone) through one device (e.g., a main earphone) among a pair of devices (e.g., a first wearable device (202) and a second wearable device (204)).

[0048] According to one embodiment of the present document, when the wearable device (200) is a wireless earphone, a pair of devices (e.g., a first wearable device (202) and a second wearable device (204)) can establish a communication connection with each other and transmit and / or receive data (e.g., audio data and / or control data) with each other. The communication connection can be established with each other using D2D communication such as Wi-Fi direct or Bluetooth (e.g., using a communication circuit that supports the communication) as described above, but is not limited thereto.

[0049] According to one embodiment of the present document, one of a pair of devices (e.g., a first wearable device (202) and a second wearable device (204)) becomes a main device (or primary device, or main device), the other device becomes a sub-device (or secondary device), and the main device can transmit data to the sub-device. For example, when a pair of devices (e.g., a first wearable device (202) and a second wearable device (204)) establish a communication connection with each other, one of the pair of devices (e.g., a first wearable device (202) and a second wearable device (204)) can be randomly selected as a main device, and the other device can be selected as a sub-device.

[0050] According to one embodiment of the present document, when a pair of devices (e.g., a first wearable device (202) and a second wearable device (204)) establish a communication connection with each other, a device that is first detected to be worn (e.g., a value indicating wearing is detected using a sensor for detecting wearing (e.g., a proximity sensor, a touch sensor, a 6-axis tilt sensor, or a 9-axis sensor)) may be selected as the main device, and the remaining devices may be selected as sub-devices. According to one embodiment of the present document, the main device may transmit data received from the electronic device (101) to the sub-device. For example, the first wearable device (202), which is the main device, may output audio to a speaker based on audio data received from the electronic device (101), and may also transmit the audio data to the second wearable device (204), which is the sub-device. In one embodiment, the sub-device may receive audio data transmitted from the electronic device (101) to the main device through sniffing, based on connection information provided from the main device.

[0051] According to one embodiment of the present document, the first wearable device (202), which is the main device, can transmit data (e.g., audio data or control data) received from the second wearable device (204), which is the sub device, to the electronic device (101). For example, when a touch event occurs in the second wearable device (204), which is the sub device, control data including information about the generated touch event can be transmitted to the electronic device (101) by the first wearable device (202), which is the main device. However, according to one embodiment of the present document, the sub device and the electronic device (101) establish a communication connection with each other, and thus, transmission and / or reception of data may be directly performed between the sub device and the electronic device (101).

[0052] According to one embodiment of the present document, the first wearable device (202) and the second wearable device (204) may be connected to an electronic device (210) by wire and / or wirelessly, the electronic device (210) having one or more storage spaces having a size and shape corresponding to the first wearable device (202) and the second wearable device (204). In one embodiment, the electronic device (210) may be an earphone case or a cradle device for storing and charging the first wearable device (202) and the second wearable device (204). In one embodiment, the electronic device (210) may establish a communication connection with the first wearable device (202) and the second wearable device (204), and transmit and / or receive data with each other. According to one embodiment of the present document, the electronic device (210) may establish a communication connection with the electronic device (101), and transmit and / or receive data with each other. For example, the electronic device (210) may transmit information about the status of the first wearable device (202) and the second wearable device (204) (e.g., whether they are operating) and / or the status of the electronic device (210) (e.g., whether the cover is open or closed, or whether the first wearable device (202) and / or the second wearable device (204) is stored) to the electronic device (101).

[0053] For convenience of explanation, the following description is given of a case where the wearable device (200) is a pair of wearable devices (202, 204), but the following description can be substantially equally applied to various types of wearable devices (200) (e.g., smart watches, head-mounted display devices, or devices for measuring biosignals).

[0054] FIG. 3 is an exemplary drawing for explaining the configuration of a first wearable device (202) according to one embodiment of the present document.

[0055] According to one embodiment of the present document, the first wearable device (202) may be a main earbud that can be connected to an electronic device (101) (e.g., a smartphone) as shown in FIG. 1, and the second wearable device (204) may be a sub earbud that can be connected to the main earbud.

[0056] According to one embodiment of the present document, the first wearable device (202) may include components identical to or similar to at least one of the components (e.g., modules) of the electronic device (101) illustrated in FIG. 1. According to one embodiment of the present document, the first wearable device (202) may include at least one of a communication circuit (320) (e.g., a communication module (190) of FIG. 1), an input device (330) (e.g., an input module (150) of FIG. 1), a sensor (340) (e.g., a sensor module (176) of FIG. 1), an audio processing module (350) (e.g., an audio module (170) of FIG. 1), a memory (390) (e.g., a memory (130) of FIG. 1), a power management module (360) (e.g., a power management module (188) of FIG. 1), a battery (370) (e.g., a battery (189) of FIG. 1), an interface (380) (e.g., an interface (177) of FIG. 1), and a processor (310) (e.g., a processor (120) of FIG. 1).

[0057] According to one embodiment of the present document, the communication circuit (320) may include at least one of a wireless communication module (e.g., a Bluetooth communication module, a cellular communication module, a wireless-fidelity (Wi-Fi) communication module, a near field communication (NFC) communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module, or a power line communication (PLC) communication module). According to one embodiment of the present document, the communication circuit (320) may communicate directly or indirectly with at least one of an electronic device (101) (e.g., a smart phone), a case device (210) (e.g., a charging device such as a cradle), or a second wearable device (204) (e.g., a sub-wearable device) through a first network (e.g., the first network (198) of FIG. 1) using the at least one communication module included therein. According to one embodiment of the present document, the second wearable device (204) may be a component of the ear wearable device (200) configured as a pair with the first wearable device (202). According to one embodiment of the present document, the communication circuit (320) may operate independently from the processor (310) and may include one or more communication processors that support wired or wireless communication. According to one embodiment of the present document, the communication circuit (320) may be connected to one or more antennas that may transmit or receive signals or information to or from another electronic device (e.g., at least one of the electronic device (101), the case device (210), or the second wearable device (204).According to one embodiment of the present document, at least one antenna suitable for a communication method used in a communication network, such as a first network (e.g., the first network (198) of FIG. 1) or a second network (e.g., the second network (199) of FIG. 2), may be selected from among the plurality of antennas, for example, by a communication circuit (320). A signal or information may be transmitted or received between the communication circuit (320) and another electronic device via the selected at least one antenna.

[0058] According to one embodiment of the present document, the input device (330) may be configured to generate various input signals that may be used in the operation of the first wearable device (202). According to one embodiment of the present document, the input device (330) may include at least one of a touchpad, a touch panel, or a button. According to one embodiment of the present document, the touchpad may recognize a touch input in at least one of, for example, a capacitive, a resistive, an infrared, or an ultrasonic manner. According to one embodiment of the present document, when a capacitive touchpad is provided, physical contact or proximity recognition may be possible. The touchpad may further include a tactile layer. A touchpad including a tactile layer may provide a tactile response to a user. The button may include, for example, a physical button or an optical key.

[0059] According to one embodiment of the present document, the input device (330) can generate a user input regarding turning on or off the first wearable device (202). According to one embodiment of the present document, the input device (330) can receive a user input for a communication connection between the first wearable device (202) and the second wearable device (204). According to one embodiment of the present document, the input device (330) can receive a user input related to audio data (or audio content). For example, the user input can be related to a function of starting playback, pausing playback, stopping playback, adjusting playback speed, adjusting playback volume, or muting audio data. According to one embodiment of the present document, the operation of the first wearable device (202) can be controlled by various gestures by the user, such as tapping or swiping up and down on a surface on which a touch pad is installed. According to one embodiment of the present document, the input device (330) may receive a user input that initiates pairing between the first wearable device (202) and the electronic device (101). For example, in response to the user input, the processor (310) may operate the first wearable device (202) in a pairing mode (e.g., an inquiry scan mode or a BLE advertising scan mode).

[0060] According to various embodiments, the sensor (340) may identify the location or operating state of the first wearable device (202), or identify whether the cover (302) of the case device (210) is in an open or closed state. According to one embodiment of the present document, the sensor (340) may convert the measured or identified information into an electrical signal. According to one embodiment of the present document, the sensor (340) may include, for example, at least one of a magnetic sensor, an acceleration sensor, a gyro sensor, a geomagnetic sensor, a proximity sensor, a gesture sensor, a grip sensor, or a biometric sensor. In one embodiment, the sensor (340) may further include an optical sensor. The optical sensor may include a light emitting unit (e.g., a light emitting diode (LED)) that outputs light of at least one wavelength band. The optical sensor may include a light receiving unit (e.g., a photodiode) that receives light of one or more wavelength bands scattered or reflected from an object and generates an electrical signal.

[0061] According to one embodiment of the present document, the audio processing module (350) can support an audio data collection function and can reproduce the collected audio data. According to one embodiment of the present document, the audio processing module (350) can include an audio decoder (not shown) and a D / A converter (not shown). According to one embodiment of the present document, the audio decoder can convert audio data stored in the memory (390) or received from the electronic device (101) through the communication circuit (320) into a digital audio signal. According to one embodiment of the present document, the D / A converter can convert the digital audio signal converted by the audio decoder into an analog audio signal. According to one embodiment of the present document, the audio decoder can convert audio data received from the electronic device (101) through the communication circuit (320) and stored in the memory (490) into a digital audio signal. According to one embodiment of the present document, the speaker can output an analog audio signal converted by the D / A converter. According to one embodiment of the present document, the audio processing module (350) may include an A / D converter (not shown). The A / D converter may convert an analog voice signal transmitted through a microphone into a digital voice signal. According to one embodiment of the present document, the audio processing module (350) may reproduce various audio data set in the operating operation of the first wearable device (202). For example, the processor (310) may be configured to detect a state in which the first wearable device (202) is coupled to or separated from the user's ear through the sensor (340), and reproduce audio data related to sound effects or guidance sounds through the audio processing module (350). According to one embodiment of the present document, the output of the sound effects or guidance sounds may be omitted according to user settings or designer intentions.

[0062] A wearable device (e.g., a first wearable device (202)) according to one embodiment of the present document may include at least one microphone (352). A wearable device (e.g., a first wearable device (202)) according to one embodiment of the present document may include an internal microphone (e.g., a first microphone (510)) disposed at a first location of the wearable device (e.g., a first wearable device (202)), a first external microphone (e.g., a second microphone (520)) disposed at a second location of the wearable device (e.g., a first wearable device (202)), and / or a second external microphone (e.g., a third microphone (530)) disposed at a third location of the wearable device (e.g., a first wearable device (202)). According to one embodiment of the present document, the internal microphone (e.g., the first microphone (510)) may be positioned closer to the user's body than other microphones when the first wearable device (202) is worn on the user's ear. According to one embodiment of the present document, the first external microphone (e.g., the second microphone (520)) may be designated or selected by the user. According to one embodiment of the present document, the wearable device (200) may include three or more external microphones.

[0063] According to one embodiment of the present document, the memory (390) can store various data used by at least one component (e.g., the processor (310) or the sensor (340)) of the first wearable device (202). According to one embodiment of the present document, the data can include, for example, input data or output data for software and commands related thereto. According to one embodiment, the data can include audio data received from the electronic device (101), information on the cover state (e.g., open state or closed state) of the case device (210), location information of the second wearable device (204) received from the case device (210) or the second wearable device (204), or role information and information on at least one parameter required for connection with the second wearable device (204). The memory (390) can include a volatile memory or a non-volatile memory.

[0064] According to one embodiment of the present document, the power management module (360) can manage power supplied to the first wearable device (202). According to one embodiment of the present document, the power management module (360) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC). According to one embodiment of the present document, the power management module (360) can include a battery charging module. According to one embodiment of the present document, when another electronic device (e.g., the electronic device (101), the case device (210), the second wearable device (204), or one of the other electronic devices) is electrically connected (wirelessly or by wire) to the first wearable device (202), the power management module (360) can receive power from the other electronic device and charge the battery (370).

[0065] According to one embodiment of the present document, the battery (370) can power at least one component of the first wearable device (202). According to one embodiment of the present document, the battery (370) can include, for example, a rechargeable battery. According to one embodiment of the present document, when the first wearable device (202) is mounted within the case device (210), the first wearable device (202) can charge the battery (370) to a predetermined charge level and then turn on the power of the first wearable device (202) or turn on at least a portion of the communication circuit (310).

[0066] According to one embodiment of the present document, the interface (380) can support one or more designated protocols that can be used for the first wearable device (202) to be directly (e.g., wired) connected to the electronic device (101), the case device (210), the second wearable device (204), or another electronic device. According to one embodiment of the present document, the interface (380) can include, for example, at least one of a high definition multimedia interface (HDMI), a USB interface, an SD card interface, a power line communication (PLC) interface, or an audio interface. According to one embodiment of the present document, the interface (380) can include at least one connection port for forming a physical connection with the case device (210). According to one embodiment of the present document, the processor (310) can control, for example, by executing software, at least one other component (e.g., hardware or software component) of the first wearable device (202) connected to the processor (310), and can perform various data processing or calculations. According to one embodiment of the present document, as at least a part of data processing or calculation, the processor (310) may load a command or data received from another component (e.g., a sensor (340) or a communication circuit (320)) into the memory (390), process the command or data stored in the memory (390), and store the resulting data in the memory (390). According to one embodiment of the present document, the processor (310) may determine whether an electrical connection is formed between the first wearable device (202) and the case device (210) through the sensor (340) or the interface (380).According to one embodiment of the present document, when an electrical connection is formed between the first wearable device (202) and the case device (210), the processor (310) can receive location information of the second wearable device (204) from the case device (210). According to one embodiment of the present document, the processor (310) can identify whether the cover (302) of the case device (210) is in an open or closed state by recognizing a magnetic body installed in the case device (210) through a magnetic sensor included in the sensor (340). According to one embodiment of the present document, the processor (410) can determine whether an electrical connection is formed between the first wearable device (202) and the case device (210) by recognizing a state in which a connection port included in the interface (480) is in contact with an electrical contact of the case device (210). According to one embodiment of the present document, the processor (310) may form a communication connection with the electronic device (101) through the communication circuit (320), and may receive data (e.g., audio data) from the electronic device (101) through the formed communication connection. According to one embodiment of the present document, the processor (310) may transmit data received from the electronic device (101) through the communication circuit (320) to the second wearable device (204).

[0067] According to one embodiment of the present document, a second wearable device (204) configured as a pair with a first wearable device (202) may include components identical or similar to those included in the first wearable device (202) and may perform all or part of the operations of the first wearable device (202) described in the drawings described below.

[0068] FIG. 4A is an exemplary diagram for explaining a function or operation of a wearable device (e.g., a first wearable device (202)) according to an embodiment of the present document performing voice detection based on characteristics of audio signals acquired through microphones (e.g., a first microphone (510), a second microphone (520) and / or a third microphone (530)) of the wearable device (e.g., the first wearable device (202)). FIG. 5 is an exemplary diagram for explaining a function or operation of a wearable device (e.g., the first wearable device (202)) according to an embodiment of the present document performing voice detection through a learned artificial neural network based on magnitude characteristics and energy ratio characteristics of audio signals.

[0069] Referring to FIGS. 4A and 5 , a wearable device (200) (e.g., a first wearable device (202)) according to an embodiment of the present document may, in operation 410, acquire a magnitude feature of a first audio signal acquired by a first microphone (510). The term “magnitude feature” according to an embodiment of the present document may be referred to alternatively / interchangeably with the term “energy spectrum.” The wearable device (200) according to an embodiment of the present document may acquire signals in a frequency domain by performing a short-time Fourier transform on audio signals acquired by the first microphone (510), the second microphone (520), and / or the third microphone (530). The wearable device (200) according to an embodiment of the present document may acquire signals in a frequency domain by using the following mathematical expression 1.

[0070]

[0071] In mathematical expression 1, i can mean the microphone index, can mean a short-time Fourier transform. In mathematical expression 1, may each mean a frame index and a frequency index. A wearable device (200) according to one embodiment of the present document may perform a short-time Fourier transform according to a specified time interval (e.g., 20 ms).

[0072] A wearable device (200) according to one embodiment of the present document can obtain a magnitude spectrum for a signal of a first microphone (510) (e.g., an internal microphone) among signals in a frequency domain converted by Equation 1, using Equation 2 below. In a wearable device (200) in which a plurality of microphones are arranged, when the wearable device is worn by a user, an audio signal obtained by a microphone (e.g., the first microphone (510)) arranged closer to a part of the user's body tends to be obtained with a high signal-to-noise ratio, so a wearable device (200) according to one embodiment of the present document can obtain (or calculate) a magnitude spectrum for an audio signal obtained by the first microphone (510) among a plurality of microphones. The wearable device (200) according to one embodiment of the present document can acquire a magnitude spectrum for a designated frequency range (e.g., a frequency range of 4 kHz or less) in which the user's voice is likely to be mainly distributed, when acquiring a magnitude spectrum. The designated frequency range according to one embodiment of the present document may be designated in advance (e.g., at the time of manufacturing the wearable device (200)) or may be designated by the user. The wearable device (200) according to one embodiment of the present document can at least temporarily store the acquired magnitude spectrum as a magnitude feature in the memory (390).

[0073]

[0074] In mathematical expression 2, can mean the magnitude spectrum, can mean the result obtained by mathematical expression 1, or in other words, the result obtained by short-time Fourier transform. By the absolute value symbol in mathematical expression 2, the result obtained as a complex value can be converted to a real value. In mathematical expression 2, may mean a frequency bin index, and a smaller k value may mean a lower frequency. The magnitude spectrum obtained according to one embodiment of the present document may have a value composed of multiple dimensions. The wearable device (200) according to one embodiment of the present document may use the magnitude spectrum obtained by mathematical expression 2 when performing a voice detection function or operation, or may perform a voice detection function or operation by using the power spectrum or log power spectrum obtained by mathematical expressions 3 and 4 below, respectively.

[0075]

[0076]

[0077] According to one embodiment of the present document, a wearable device (200) (e.g., a first wearable device (202)) may determine whether microphones (e.g., a second microphone (520) and a third microphone (530)) are shielded in operation 420. According to one embodiment of the present document, a wearable device (200) may determine whether microphones (e.g., a second microphone (520) and a third microphone (530)) are shielded based on the algorithm illustrated in FIG. 6. According to one embodiment of the present document, a wearable device (200) may obtain (e.g., determine or calculate) an energy ratio feature based on whether the microphones are shielded, as described below.

[0078] FIG. 6 is an exemplary drawing for explaining a function or operation of calculating an energy ratio based on whether microphones (e.g., a second microphone (520) and a third microphone (530)) of a wearable device (e.g., a first wearable device (202)) are shielded, according to one embodiment of the present document, of a wearable device (200).

[0079] Referring to FIG. 6, a wearable device (200) according to an embodiment of the present document can identify the energy of each signal of a second microphone (520) and a third microphone (530) in operation 610. A wearable device (200) according to an embodiment of the present document can identify (e.g., obtain or calculate) the energy of each signal of the second microphone (520) and the third microphone (530) based on the following mathematical expression 5.

[0080]

[0081] In mathematical equation 5 can mean the energy of an audio signal acquired by a microphone, can be a constant.

[0082] The wearable device (200) according to one embodiment of the present document may determine, in operation 620, whether the maximum value of energy exceeds a first threshold value. For example, if the acquired energy value is 100 for the second microphone (520) and 0 for the third microphone (530) (e.g., if the second microphone (520) is in an open state and the third microphone (530) is in a shielded state), the wearable device (200) according to one embodiment of the present document may identify the maximum value of energy as 100. Here, for example, if the first threshold value is specified as 50, the wearable device (200) according to one embodiment of the present document may determine that the maximum value of energy exceeds the first threshold value.

[0083] According to an embodiment of the present document, the wearable device (200) may identify the difference in energy in operation 630 when it is determined that the maximum value of energy exceeds the first threshold value in operation 620 (e.g., operation 620 - Yes). According to an embodiment of the present document, the wearable device (200) may determine whether the difference in energy in operation 630 exceeds the second threshold value in operation 650. According to an embodiment of the present document, the wearable device (200) may calculate the difference between the energy value of the audio signal acquired by the third microphone (530) and the energy value of the audio signal acquired by the second microphone (520). For example, according to an embodiment of the present document, the wearable device (200) may identify (e.g., determine or calculate) the energy difference as -100 in operation 630. A wearable device (200) according to one embodiment of the present document may determine, for example, that the energy difference does not exceed the second threshold value in operation 650, when the second threshold value is set to 30.

[0084] According to an embodiment of the present document, the wearable device (200) may calculate an energy ratio using an audio signal of the first external microphone (e.g., the second microphone (520)) because both the first external microphone (e.g., the second microphone (520)) and the second external microphone (e.g., the third microphone (530)) are not shielded when it is determined in operation 650 that the energy difference exceeds the second threshold value (operation 650-Yes). According to an embodiment of the present document, the wearable device (200) may calculate an energy ratio using an audio signal of the third microphone (530) because the second external microphone (e.g., the third microphone (530)) is shielded when it is determined in operation 650 that the energy difference does not exceed the second threshold value (operation 650-No). A wearable device (200) according to one embodiment of the present document can determine, through such an algorithm or process, whether a microphone among a plurality of external microphones is shielded by a part of the user's body or by other factors.

[0085] In one embodiment of the present document, the wearable device (200) may use a dummy value as an energy ratio to be described later, since in operation 640, if the maximum value of the energy does not exceed the first threshold value (e.g., operation 620-No), all external microphones are shielded. Since the first microphone (510) is in closer contact with a part of the user's body (e.g., an ear) than the second microphone (520) and the third microphone (530) when the wearable device (200) is worn by the user, the possibility of being shielded by a part of the user's body is low, and thus the wearable device (200) according to one embodiment of the present document may determine whether only the external microphones are shielded. However, the wearable device (200) according to one embodiment of the present document may also determine whether the first microphone (510) is shielded. In this case, the wearable device (200) according to one embodiment of the present document may determine whether the first microphone (510) is shielded by determining whether the energy of the audio signal acquired by the first microphone (510) exceeds a threshold value. In this case, the wearable device (200) according to one embodiment of the present document may use a dummy value as an energy ratio described below when the first microphone (510) is shielded.

[0086] Returning to FIG. 4A again, the wearable device (200) (e.g., the first wearable device (202)) according to one embodiment of the present document may, at operation 430, obtain an energy ratio of the first audio signal and the second audio signal acquired by the second microphone (520) based on the determination of whether shielding is present. The wearable device (200) according to one embodiment of the present document may obtain (e.g., determine or calculate) the energy ratio of the audio signal acquired by the first microphone (510) and the audio signal acquired by the second microphone (520) or the third microphone (530) using the following mathematical expression 6. The wearable device (200) according to one embodiment of the present document may at least temporarily store the obtained energy ratio as an energy ratio feature in the memory (390).

[0087]

[0088] In mathematical expression 6, may refer to the energy of the first audio signal acquired by the first microphone. In mathematical expression 6, may refer to the energy of a second audio signal acquired by an unshielded external microphone (e.g., a second microphone (520) or a third microphone (530)). The energy ratio according to one embodiment of the present document may include a one-dimensional value.

[0089] According to one embodiment of the present document, a wearable device (200) (e.g., a first wearable device (202)) may determine, in operation 440, whether a voice is included in a first audio signal or a second audio signal through a learned artificial neural network based on the magnitude feature obtained according to operation 410 and the energy ratio feature obtained according to operation 430. According to one embodiment of the present document, a wearable device (200) may use the magnitude feature obtained according to operation 410 and the energy ratio feature obtained according to operation 430 as inputs of a learned artificial neural network (e.g., DNN). According to one embodiment of the present document, a wearable device (200) may normalize the magnitude feature obtained according to operation 410 and the energy ratio feature obtained according to operation 430 according to the following mathematical expression 7 or mathematical expression 8, and input them as input values ​​of the learned artificial neural network. Normalization according to mathematical expression 7 below may mean maximum-minimum normalization, and normalization according to mathematical expression 8 may mean Z-score normalization.

[0090]

[0091]

[0092] In mathematical expression 7, can mean the acquired magnitude feature and energy ratio feature, and the regularization parameter may be predetermined or may include a personalized value for the user. For example, the wearable device (200) according to one embodiment of the present document, if it is identified that the user is using the wearable device (200) in a low-noise environment (e.g., a state of high signal-to-noise ratio), may identify the magnitude feature and the energy ratio feature for a specified time period (e.g., 5 seconds) or less, and may calculate a personalized normalization parameter for the same and store it in advance in the memory (390). For example, the wearable device (200) according to one embodiment of the present document may obtain the maximum and minimum values ​​of the magnitude features and the energy ratio features for a specified time period (e.g., 5 seconds), respectively. and can be stored at least temporarily in the memory (390) as a personalized normalization parameter. In addition, the wearable device (200) according to one embodiment of the present document stores the average value and standard deviation value of the magnitude features and energy ratio features acquired for a specified time interval (e.g., 5 seconds), respectively. can be determined and stored at least temporarily in memory (390) as a personalized normalization parameter.

[0093] A wearable device (200) according to one embodiment of the present document can input normalized magnitude features and energy ratio features as input values ​​of a trained artificial neural network. Since the input values ​​input to the trained artificial neural network according to one embodiment of the present document are input in frame units (e.g., frames divided into 20 ms time sections) distinguished by a short-time Fourier transform, the normalized magnitude features can be input to the artificial neural network as multi-dimensional values, and the normalized energy ratio features can be input to the artificial neural network as one-dimensional values. The artificial neural network according to one embodiment of the present document can be trained according to the algorithm or process illustrated in FIG. 4b.

[0094] Referring to FIG. 4B, the wearable device (200) according to one embodiment of the present document can acquire the magnitude feature of the first audio signal acquired by the first microphone (510) in operation 405. The wearable device (200) according to one embodiment of the present document can receive an audio signal prepared for learning of an artificial neural network. The wearable device (200) according to one embodiment of the present document can acquire signals in the frequency domain by performing a short-time Fourier transform on audio signals acquired by the first microphone (510), the second microphone (520), and / or the third microphone (530). The wearable device (200) according to one embodiment of the present document can acquire signals in the frequency domain by using the above-described mathematical expression 1. A wearable device (200) according to one embodiment of the present document can obtain a magnitude characteristic of a first audio signal obtained by a first microphone using mathematical expression 2.

[0095] According to one embodiment of the present document, the wearable device (200) can obtain an energy ratio of a first audio signal and a second audio signal obtained by a second microphone (520) in operation 415. According to one embodiment of the present document, the wearable device (200) can obtain the energy ratio using mathematical expression 6 based on the second audio signal obtained by the second microphone (520). According to one embodiment of the present document, in the process of training an artificial neural network, a function or operation for determining whether there is shielding may be omitted.

[0096] The wearable device (200) according to one embodiment of the present document can train an artificial neural network based on the acquired magnitude feature and energy ratio feature in operation 425. The wearable device (200) according to one embodiment of the present document can use the magnitude feature acquired in operation 405 and the energy ratio feature acquired in operation 415 as inputs of a trained artificial neural network (e.g., DNN). The wearable device (200) according to one embodiment of the present document can normalize the magnitude feature acquired in operation 405 and the energy ratio feature acquired in operation 415 according to the above-described mathematical expression 7 or mathematical expression 8, and input them as input values ​​of the trained artificial neural network. According to one embodiment of the present document, in the learning step of the artificial neural network, a normalization parameter is applied to the learning data. Normalization can be performed by calculating . The wearable device (200) according to one embodiment of the present document can output a value between 0 and 1 by taking normalized input features as input. According to one embodiment of the present document, in the learning process of the artificial neural network, the wearable device (200) can update the hyperparameters so that the output value and the predefined deep learning target value match each other. In this case, the wearable device (200) according to one embodiment of the present document can use a binary cross entropy function, etc., as a loss function.

[0097] Returning to FIG. 4A, the wearable device (200) according to one embodiment of the present document may, at operation 440, determine whether a voice is included in the first audio signal or the second audio signal through a learned artificial neural network based on the acquired magnitude feature and energy ratio. The wearable device (200) according to one embodiment of the present document may output an output value between 0 and 1 through the learned artificial neural network, as exemplarily illustrated in FIG. 7.

[0098] FIG. 7 is an exemplary diagram for explaining a function or operation of a wearable device (200) according to one embodiment of the present document to process magnitude features and energy ratio features input to a learned artificial neural network. Referring to FIG. 7, the artificial neural network according to one embodiment of the present document may include at least one of a batch normalization module (or layer) (710), an encoder (720), a gated recurrent unit (GRU) (730), a decoder (740), a linear module (750), a sigmoid module (760), and an average module (770). The encoder (720) according to one embodiment of the present document may include a plurality of convolutional layers. The decoder (740) according to one embodiment of the present document may include a plurality of transposed convolutional layers. According to one embodiment of the present document, a wearable device (200) may perform batch normalization on magnitude features and energy ratio features through a batch normalization module (710). According to one embodiment of the present document, a wearable device (200) may encode batch-normalized magnitude features through an encoder (720). According to one embodiment of the present document, encoding may not be performed on batch-normalized energy ratio features. According to one embodiment of the present document, encoded magnitude features may be input to a GRU (730). According to one embodiment of the present document, unencoded magnitude features may be input to a decoder (740) without passing through the GRU (730) (e.g., skip connection operation). The skip connection operation according to one embodiment of the present document may be omitted. According to one embodiment of the present document, a GRU (730) may identify correlations between frames based on a time axis.According to one embodiment of the present document, the GRU (730) can identify a correlation between frames by using information about a previous frame of an input frame stored in a wearable device (e.g., memory (390)). According to one embodiment of the present document, the GRU (730) can output an output value that combines a magnitude feature and an energy ratio feature to the decoder (740). The decoder (740) according to one embodiment of the present document can decode an input value input from the GRU (730). The linear module (750) according to one embodiment of the present document can convert the decoded high-dimensional result into a designated dimension (e.g., 64 dimensions). The sigmoid module (760) according to one embodiment of the present document can map the value converted into the designated dimension into values ​​between 0 and 1. The average module (770) according to one embodiment of the present document can output a value between 0 and 1 by averaging values ​​that are doubled to values ​​between 0 and 1. The wearable device (200) according to one embodiment of the present document can determine whether at least one audio signal among the acquired audio signals (e.g., the first audio signal or the second audio signal) contains a voice based on the output value. For example, the wearable device (200) according to one embodiment of the present document can determine that a voice is present when the output value is about 0.5 or more. The wearable device (200) according to one embodiment of the present document can determine that a voice is absent when the output value is less than about 0.5. The wearable device (200) according to one embodiment of the present document can determine that a voice is present when the output value is about 0.6 or more. According to one embodiment of the present document, the wearable device (200) may determine that there is a voice presence when the output value is less than about 0.4. According to one embodiment of the present document, the wearable device (200) may determine that there is a judgment pending when the output value is less than about 0.6 and greater than about 0.4. Such threshold values ​​(e.g., about 0.5) can be adaptively determined depending on the user's environment (e.g., noise level).

[0099] The wearable device (200) according to one embodiment of the present document may apply a pruning algorithm or process based on the signal-to-noise ratio of an audio signal acquired by an unshielded microphone (e.g., the second microphone (520)). FIG. 8 is an exemplary diagram for explaining a function or operation of the wearable device (200) according to one embodiment of the present document to determine the computational ratio of a trained artificial neural network based on the signal-to-noise ratio of an audio signal input through a microphone of the wearable device (200). The wearable device (200) according to one embodiment of the present document may apply 100% of the network computation in an environment with high noise (e.g., low signal-to-noise ratio). The wearable device (200) according to one embodiment of the present document may apply 70% of the network computation in an environment with moderate noise (e.g., medium signal-to-noise ratio). According to one embodiment of the present document, a wearable device (200) can apply a network operation of 50% in an environment with low noise (e.g., high signal-to-noise ratio). According to one embodiment of the present document, operations may not be performed for pruned operations. According to one embodiment of the present document, a wearable device (200) can apply a pruning algorithm to each layer or module (e.g., encoder (720)) and can exclude them from the operation in order of decreasing weight.

[0100] According to one embodiment of the present document, at least some of the various embodiments of the present document described above may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or a server) operatively connected to a wearable device (200). For example, the wearable device (200) may be implemented as a function or operation in which audio data is transmitted to an electronic device (e.g., the electronic device (101) of FIG. 1 or a server) and voice is detected through the electronic device that receives the audio data.

[0101] According to one embodiment of the present document, a wearable device (e.g., a first wearable device (202)) includes a first microphone (510), a second microphone (520) disposed at a different location from the first microphone, at least one processor (310), and a memory (390), wherein the memory may store instructions that, when executed, cause the at least one processor to obtain a magnitude feature of a first audio signal obtained by the first microphone (510), obtain an energy ratio of the first audio signal and a second audio signal obtained by the second microphone, and determine, based on the obtained magnitude feature and the energy ratio, through a trained artificial neural network, whether a voice is included in the first audio signal or the second audio signal.

[0102] According to one embodiment of the present document, the wearable device may further include a third microphone (530), and the instructions may further include instructions for determining whether the first microphone (510), the second microphone (520), and the third microphone (530) are shielded based on energy of audio data received by each of the first microphone (510), the second microphone (520), and the third microphone (530).

[0103] According to one embodiment of the present document, the instructions may further include instructions for obtaining the energy ratio based on an audio signal obtained from an unshielded microphone based on the determination of whether the microphone is shielded.

[0104] According to one embodiment of the present document, the instructions may further include an instruction to perform a short-time Fourier transform on audio signals acquired by the first microphone (510), the second microphone (520), and the third microphone (530) respectively to acquire the magnitude feature.

[0105] According to one embodiment of the present document, the instructions may further include an instruction to determine a designated dummy value as the energy ratio when both the second microphone (520) and the third microphone (530) are shielded.

[0106] According to one embodiment of the present document, the instructions may further include instructions for determining an operation ratio of the learned artificial neural network based on a noise intensity of an audio signal acquired by the second microphone (520).

[0107] According to one embodiment of the present document, the instructions may further include instructions to obtain the energy ratio based on an audio signal of the second microphone when both the second microphone (520) and the third microphone (530) are not shielded.

[0108] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0109] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0110] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0111] Various embodiments of the present document may be implemented as software (e.g., a program (2540)) including one or more instructions stored in a storage medium (e.g., an internal memory (2536) or an external memory (2538)) readable by a machine (e.g., an electronic device (2501)). For example, a processor of the machine (e.g., an electronic device (2501)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0112] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0113] According to one embodiment of the present document, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment of the present document, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment of the present document, operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In wearable devices, Microphone 1, A second microphone positioned at a different location from the first microphone; at least one processor, and comprising a memory, said memory being configured to cause at least one processor, when executed; Obtaining a magnitude feature of a first audio signal obtained by the first microphone, Obtaining an energy ratio of the first audio signal and the second audio signal obtained by the second microphone, A wearable device characterized by storing instructions for determining whether a voice is included in the first audio signal or the second audio signal through a learned artificial neural network based on the acquired magnitude feature and the energy ratio.

2. In paragraph 1, The wearable device further comprises a third microphone, A wearable device, characterized in that the instructions further include instructions for determining whether the first microphone, the second microphone, and the third microphone are shielded based on energy of audio data received by each of the first microphone, the second microphone, and the third microphone.

3. In paragraph 1 or 2, A wearable device, characterized in that the instructions further include instructions for obtaining the energy ratio based on an audio signal obtained from an unshielded microphone, based on the determination of whether or not there is shielding.

4. In any one of paragraphs 1 to 3, A wearable device, characterized in that the instructions further include instructions to perform a short-time Fourier transform on audio signals respectively acquired by the first microphone, the second microphone, and the third microphone to acquire the magnitude feature.

5. In any one of paragraphs 1 to 4, A wearable device, characterized in that the instructions further include instructions for determining a pre-specified dummy value as the energy ratio when both the second microphone and the third microphone are shielded.

6. In any one of paragraphs 1 to 5, A wearable device, characterized in that the instructions further include instructions for determining an operation ratio of the learned artificial neural network based on a noise intensity of an audio signal acquired by the second microphone.

7. In any one of paragraphs 1 to 6, A wearable device, characterized in that the instructions further include instructions for obtaining the energy ratio based on an audio signal of the second microphone when both the second microphone and the third microphone are not shielded.

8. In the method, An operation for obtaining a magnitude feature of a first audio signal obtained by a first microphone of a wearable device, An operation of obtaining an energy ratio of the first audio signal and the second audio signal obtained by the second microphone of the wearable device, and A method characterized by including an operation of determining whether a voice is included in the first audio signal or the second audio signal through a learned artificial neural network based on the acquired magnitude feature and the energy ratio.

9. In paragraph 8, The wearable device further comprises a third microphone, The method is characterized in that it further includes an operation of determining whether the first microphone, the second microphone, and the third microphone are shielded based on energy of audio data received by each of the first microphone, the second microphone, and the third microphone.

10. In a computer-readable non-transitory storage medium, the non-transitory storage medium stores instructions executable by a processor, the instructions, when executed, causing the processor to: Obtaining a magnitude feature of a first audio signal acquired by a first microphone of a wearable device, Obtaining an energy ratio of the first audio signal and the second audio signal obtained by the second microphone of the wearable device, A non-transitory recording medium characterized by storing instructions for determining whether a voice is included in the first audio signal or the second audio signal through a learned artificial neural network based on the acquired magnitude feature and the energy ratio.

11. In paragraph 10, The wearable device further comprises a third microphone, A non-transitory recording medium, characterized in that the instructions further include instructions for determining whether the first microphone, the second microphone, and the third microphone are shielded based on energy of audio data received by each of the first microphone, the second microphone, and the third microphone.

12. In clause 10 or 11, A non-transitory recording medium characterized in that the above instructions further include instructions for obtaining the energy ratio based on an audio signal obtained from an unshielded microphone based on the judgment of whether or not there is shielding.

13. In any one of paragraphs 10 to 12, A non-transitory recording medium, characterized in that the instructions further include instructions to perform a short-time Fourier transform on audio signals respectively acquired by the first microphone, the second microphone, and the third microphone to acquire the magnitude feature.

14. In any one of paragraphs 10 to 13, A non-transitory recording medium, characterized in that the instructions further include instructions for determining a pre-specified dummy value as the energy ratio when both the second microphone and the third microphone are shielded.

15. In any one of paragraphs 10 to 14, A non-transitory recording medium, characterized in that the instructions further include instructions for determining an operation ratio of the learned artificial neural network based on the noise intensity of the audio signal acquired by the second microphone.

Citation Information

Patent Citations

  • Method for detecting input using audio signal and apparatus thereof

    KR1020180101937A

  • Pediococcus pentosaceus strain, and vesicles from thereof and anti-inflammation and anti-bacteria uses of thereof

    KR102351148B1

  • A porous scaffold comprising a collagen and a polycarprolacton for regenerating the periodontal complex having improved healing characteristics, and method for preparing the same

    KR102458881B1

  • Electronic device and method for classifying voice and noise thereof

    KR102468148B1

  • Hearing aid method for in-situ occlusion effect and directly transmitted sound measurement

    US20090129619A1