Electronic device and method for performing voice recognition

By analyzing the low-frequency response of a speech signal to determine user proximity, the electronic device ensures efficient and secure voice recognition by activating only when within a preset distance, addressing issues of unintended activation and resource wastage.

WO2026034869A1PCT designated stage Publication Date: 2026-02-12SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/011052
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-07-25
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing voice recognition systems in electronic devices face challenges in energy efficiency, user convenience, and security, particularly in activating voice recognition based on wake-up words, as they often lead to unintended activation and wastage of system resources due to proximity issues.

Method used

An electronic device analyzes the low-frequency response of a speech signal to identify a wake-up word and determines the distance between the user and the device, performing actions only when within a preset distance, thereby reducing unintended voice recognition and conserving resources.

Benefits of technology

This approach enhances energy efficiency and security by ensuring voice recognition is performed only when the user is within a specified range, preventing unnecessary activation and resource wastage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025011052_12022026_PF_FP_ABST
    Figure KR2025011052_12022026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an electronic device and method for performing voice recognition. An electronic device and method for performing voice recognition may comprise the steps of: receiving a voice signal from a user; analyzing a frequency response of the received voice signal to obtain a feature vector of a wake-up word and a low-frequency response less than a preset frequency in the frequency response; comparing the obtained low-frequency response with a low-frequency response less than a preset frequency pre-stored in a memory of the electronic device, on the basis of identifying whether the obtained feature vector of the wake-up word corresponds to a feature vector of a wake-up word pre-stored in the memory of the electronic device; determining whether a distance between the user and the electronic device is within a preset distance on the basis of a result of the comparison; and performing an operation corresponding to the voice signal on the basis of a result of the determination.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method for performing voice recognition

[0001] An electronic device and method for performing speech recognition are disclosed. Specifically, an electronic device and method for performing speech recognition based on a low frequency response of a speech signal are disclosed.

[0002] Voice recognition technology based on wake-up words is crucial for the energy efficiency, user convenience, and security of electronic devices performing voice recognition. Electronic devices performing voice recognition can activate voice recognition for the device they wish to control and input commands through wake-up words.

[0003] According to one aspect of the present disclosure, a method for an electronic device to perform voice recognition may be provided. In one embodiment, the method may include receiving a voice signal from a user. In one embodiment, the method may include analyzing a frequency response of the received voice signal to obtain a feature vector of a wake-up word and a low-frequency response below a preset frequency among the frequency responses. In one embodiment, the method may include comparing the obtained low-frequency response with a low-frequency response below a preset frequency pre-stored in a memory based on identifying whether the obtained feature vector of the wake-up word corresponds to a feature vector of the wake-up word pre-stored in a memory of the electronic device. In one embodiment, the method may include determining whether a distance between the user and the electronic device is within a preset distance based on a result of the comparison. In one embodiment, the method may include performing an action corresponding to the voice signal based on the result of the determination.

[0004] According to one aspect of the present disclosure, an electronic device for performing voice recognition is disclosed. In one embodiment, the electronic device may include a microphone, a memory storing one or more instructions, and at least one processor operably coupled to the memory and including a processing circuit. In one embodiment, the at least one processor, alone or in cooperation, executes the instructions so that the electronic device can receive a voice signal from a user. In one embodiment, the at least one processor, alone or in cooperation, executes the instructions so that the electronic device analyzes a frequency response of the received voice signal to obtain a feature vector of a wake-up word and a low-frequency response below a preset frequency among the frequency responses. In one embodiment, the at least one processor, alone or in cooperation, executes the instructions so that the electronic device can compare the obtained low-frequency response with a low-frequency response pre-stored in the memory based on identifying whether the obtained feature vector of the wake-up word corresponds to a feature vector of the wake-up word pre-stored in the memory of the electronic device. In one embodiment, at least one processor, either alone or in cooperation, executes instructions, thereby enabling the electronic device to determine whether the distance between the user and the electronic device is within a preset distance based on the comparison result. In one embodiment, at least one processor, either alone or in cooperation, executes instructions, thereby enabling the electronic device to perform an action corresponding to the voice signal based on the determination result.

[0005] According to one aspect of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing an operation of an electronic device, one of the methods described above and below.

[0006] FIG. 1 is a drawing schematically illustrating the operation of electronic devices according to one embodiment of the present disclosure.

[0007] FIG. 2A is a drawing for explaining a directional microphone according to one embodiment of the present disclosure.

[0008] FIG. 2b is a diagram for explaining a frequency response according to the distance between a sound source and a directional microphone according to one embodiment of the present disclosure.

[0009] FIG. 3 is a flowchart illustrating a method for an electronic device to perform voice recognition according to an embodiment of the present disclosure.

[0010] FIG. 4A is a diagram illustrating a plurality of peaks detected below a first frequency of a low frequency response according to one embodiment of the present disclosure.

[0011] FIG. 4b is a diagram illustrating a plurality of peaks detected below a second frequency of a low frequency response according to one embodiment of the present disclosure.

[0012] FIG. 5 is a diagram for explaining adjusting the gain values ​​of multiple peaks according to one embodiment of the present disclosure.

[0013] FIG. 6 is a diagram for explaining an operation of an electronic device according to one embodiment of the present disclosure to obtain information about a preset distance.

[0014] FIG. 7 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.

[0015] Figure 8 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.

[0016] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0017] Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are to be understood to include plural referents. Thus, for example, the description "a constituent surface" may also include reference to one or more of such surfaces.

[0018] Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by one of ordinary skill in the art described in this disclosure.

[0019] When a part of this disclosure is said to "include" a component, this does not exclude other components, unless otherwise specifically stated, but rather implies the inclusion of other components. Furthermore, terms such as "part," "module," and the like described herein refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.

[0020] As used herein, the expression "configured to" can be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean something is "specifically designed to" in terms of hardware. Instead, in some contexts, the expression "a system configured to" can mean that the system is "capable of" in conjunction with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" can mean a dedicated processor for performing the operations (e.g., an embedded processor), or a general-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in memory.

[0021] It should be understood that the blocks and combinations of flowcharts in each of the flowcharts in this disclosure can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.

[0022] All functions or operations described in the present disclosure may be processed by a single processor or a combination of processors. A single processor or a combination of processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).

[0023] It will be appreciated that each block of the processing flow diagrams described in the present disclosure and combinations of the flow diagrams can be performed by computer program instructions. These computer program instructions can be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, such that the instructions executed by the processor of the computer or other programmable data processing equipment create a means for performing the functions described in the flow diagram block(s). These computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to implement the functions in a specific manner, such that the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes an instruction means for performing the functions described in the flow diagram block(s). Since the computer program instructions may be installed on a computer or other programmable data processing device, a series of operational steps may be performed on the computer or other programmable data processing device to create a computer-executable process, and the instructions that cause the computer or other programmable data processing device to perform the steps for performing the functions described in the flowchart block(s) may also provide steps for performing the functions described in the flowchart block(s).

[0024] Each block described herein may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). It should also be noted that in some alternative implementation examples, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in reverse order, depending on their respective functions.

[0025] One or more processors according to the present disclosure may include at least one of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), and a Neural Processing Unit (NPU). The one or more processors may be implemented in the form of an integrated system-on-a-chip (SoC) including one or more electronic components. Each of the one or more processors may also be implemented as separate hardware (H / W).

[0026] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-specific processor). However, the embodiments of the present disclosure are not limited thereto.

[0027] One or more processors according to the present disclosure may be implemented as a single core processor or as a multicore processor.

[0028] When a method according to one embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one core or may be performed by multiple cores included in one or more processors.

[0029] In the present disclosure, an 'artificial intelligence (AI) model (or model)' may refer to a set of functions or algorithms that are set to perform a desired characteristic (or purpose) by being learned using a plurality of learning data by a learning algorithm. Examples of learning algorithms include, but are not necessarily limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. In one embodiment, the artificial intelligence model may be stored in the memory of an electronic device. However, the present invention is not limited thereto, and the artificial intelligence model may be stored in an external server, and the electronic device may transmit data input to the artificial intelligence model to the server and receive data output from the artificial intelligence model from the server.

[0030] In the present disclosure, the 'artificial intelligence model' may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values, and can perform neural network operations through operations between the operation results of the previous layer and the plurality of weights. The plurality of weights of the plurality of neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the plurality of weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Examples of models including a plurality of neural network layers include, but are not limited to, a deep neural network (DNN), a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), and deep Q-networks.

[0031] In the present disclosure, a 'head-mounted display device' may refer to an augmented reality device capable of expressing augmented reality, a virtual reality device capable of expressing virtual reality, or a mixed reality device capable of expressing mixed reality. In one embodiment, the head-mounted display device may include a shape of glasses worn on the user's face or a shape of a helmet worn on the user's head, but is not necessarily limited to the examples described above.

[0032] In the present disclosure, the term "wake-up word" may refer to a word or phrase that activates a voice recognition system by recognizing a specific command in the voice recognition system. The wake-up word allows the voice recognition system to perform more complex voice recognition and data processing by utilizing more system resources after the voice recognition system is in an activated mode (or an activated state) than in a standby mode (or a standby state) prior to the activated mode.

[0033] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the description have been omitted for clarity of explanation, and similar reference numerals have been used throughout the specification to designate similar parts.

[0034] The present disclosure will be described below with reference to the attached drawings.

[0035] FIG. 1 is a drawing schematically illustrating the operation of electronic devices according to one embodiment of the present disclosure.

[0036] Referring to FIG. 1, the first to fourth electronic devices (101, 102, 103, 104) can receive a voice signal spoken by a user (1). In one embodiment, the voice signal received from the user can include a wake-up word for activating voice recognition of the first to fourth electronic devices (101, 102, 103, 104). In one embodiment, each of the first to fourth electronic devices (101, 102, 103, 104) can correspond to the electronic device (1000) of FIG. 7, which will be described later, and a specific method by which the electronic device (1000) performs voice recognition will be described later with reference to the drawings.

[0037] In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) may analyze a frequency response of a voice signal received from a user. In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) may analyze the frequency response of the received voice signal, thereby extracting a feature vector of a wake-up word included in the received voice signal, and identifying whether the extracted wake-up word corresponds to a feature vector of a pre-stored wake-up word.

[0038] In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) may receive a voice signal for registering a wake-up word from a user in order to store a feature vector of the wake-up word. In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) may analyze a frequency response of a voice signal for registering the wake-up word, thereby storing the extracted feature vector of the wake-up word in the first to fourth electronic devices (101, 102, 103, 104). In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) may use the stored feature vector of the wake-up word for voice recognition to be performed later.

[0039] In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) may perform speech recognition based on the received speech signal (or wake-up word) if it is identified that the feature vector of the wake-up word obtained from the received speech signal corresponds to the wake-up word stored in the first to fourth electronic devices (101, 102, 103, 104). In one embodiment, speech recognition based on a speech signal including the wake-up word may include the first to fourth electronic devices (101, 102, 103, 104) receiving the speech signal performing an operation corresponding to the received speech signal. In one embodiment, the operation corresponding to the received speech signal may include activating a speech recognition system by the wake-up word. In one embodiment, the action corresponding to the received speech signal may include performing an action or function corresponding to a command included in the received speech signal along with a wake-up word, or included in a speech signal received subsequent to the received speech signal.

[0040] For example, when a first user (1) wearing a first electronic device (101) utters the wake-up word “Hi Bixby,” the first electronic device (101) can extract a feature vector of “Hi Bixby” from the voice signal uttered by the first user (1) and identify whether the extracted feature vector corresponds to a feature vector corresponding to “Hi Bixby” pre-stored in the first electronic device (101). When the first electronic device (101) identifies that the feature vector of “Hi Bixby” obtained from the voice signal uttered by the user corresponds to the feature vector of “Hi Bixby” pre-stored, the first electronic device (101) can output “Yes, please say”, which is a voice signal indicating that the voice recognition system of the first electronic device (101) is activated. However, it is not necessarily limited to this, and if the voice signal spoken by the user includes a command other than “Hi Bixby,” the first electronic device (101) may perform an action or function corresponding to the included command without outputting “Yes, please say so.”

[0041] In one embodiment, when the first to fourth electronic devices (101, 102, 103, 104) are all located in the same space, the second electronic device (102), the third electronic device (103), and the fourth electronic device (104) may also receive a voice signal uttered by the first user (1) to activate the voice recognition system of the first electronic device (101). In one embodiment, the wake-up words of the second electronic device (102), the third electronic device (103), and the fourth electronic device (104) may include words or phrases that are identical or similar to the wake-up words stored in the first electronic device (101). In this case, even though the first user (1) utters a wake-up word to perform voice recognition for the first electronic device (101), the unintended second electronic device (102), the third electronic device (103), and the fourth electronic device (104) may perform voice recognition based on the wake-up word uttered by the user.

[0042] For example, since the second electronic device (102) has stored the feature vector of the wake-up word spoken by the first user (1), the feature vector of the wake-up word spoken by the first user (1) to perform voice recognition for the first electronic device (101) can be identified as corresponding to the feature vector pre-stored in the second electronic device (102). In addition, even though the third electronic device (103) and the fourth electronic device (104) have stored the feature vectors of the wake-up words spoken by the second user (2) and the third user (3), respectively, if the accuracy of voice recognition for the wake-up words of the third electronic device (103) and the fourth electronic device (104) is low, the feature vector of the wake-up word spoken by the first user (1) can be identified as corresponding to the feature vector of the pre-stored wake-up word.

[0043] In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) may acquire a low frequency response lower than a preset frequency among frequency responses of a received voice signal, and compare the acquired low frequency response with a low frequency response of a pre-stored voice signal. Here, the low frequency response of the pre-stored voice signal may include a low frequency response lower than the preset frequency among frequency responses of a voice signal (a voice signal for registering a wake-up word) from which a feature vector of a pre-stored wake-up word is acquired.

[0044] In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) can determine whether the distance between the user (1) and the electronic device (1000) is within a preset distance based on the comparison result. In one embodiment, the first to fourth electronic devices (101, 102, 103, 104) can perform an operation corresponding to the received voice signal if it is determined that the distance between the user (1) and the first to fourth electronic devices (101, 102, 103, 104) is within the preset distance. In other words, even if the feature vector of the wake-up word acquired from the received voice signal is identified as corresponding to the pre-stored feature vector, the first to fourth electronic devices (101, 102, 103, and 104) may perform an operation corresponding to the received voice signal only when the distance from the user (1) is within a preset distance, and may not perform the operation corresponding to the received voice signal when the distance from the user (1) is further than the preset distance. Here, the preset distance may correspond to the distance between the sound source of the received voice signal (e.g., a voice signal for registering the wake-up word) to acquire the feature vector of the pre-stored wake-up word and the electronic device in which the feature vector of the wake-up word is stored. However, the present invention is not necessarily limited thereto, and the preset distance may include a distance determined based on a user input selecting the preset distance.

[0045] For example, the first electronic device (101) and the third electronic device (103) may include a head-mounted display device. In one embodiment, the head-mounted display device is an electronic device that operates while worn on the user's face, and thus the preset distance may be 20 cm. In addition, the second electronic device (102) and the fourth external electronic device (104) may include a smartphone. In one embodiment, the smartphone is an electronic device that operates while held in the user's hand, and thus the voice recognition limit distance may be 50 cm. Accordingly, even if the first user (1) utters “Hi Bixby” as in FIG. 1, only the first electronic device (101) located within a preset distance may perform voice recognition based on “Hi Bixby” uttered by the first user (1), and the second electronic device (102), the third electronic device (103), and the fourth electronic device (104) located further away than the preset distance may not perform voice recognition based on “Hi Bixby.”

[0046] Thus, according to one embodiment of the present disclosure, voice recognition based on the wake-up word can be performed by taking into account the distance between the user uttering the wake-up word and the electronic device performing voice recognition. This prevents unintended voice recognition from being performed, and by not activating the voice recognition system unnecessarily, system resources of the electronic device performing voice recognition can be prevented from being wasted.

[0047] FIG. 2A is a drawing for explaining a directional microphone according to one embodiment of the present disclosure.

[0048] In one embodiment, the electronic device may include a directional microphone (200). The directional microphone (200) of FIG. 2A may correspond to the microphone (1100) described later in FIG. 7. The directional microphone (200) may include a microphone designed to intensively receive sounds coming from a specific direction and to receive less sounds coming from other directions except for the specific direction.

[0049] In one embodiment, the directional microphone (200) may include a diaphragm (201) that detects sound pressure of sound generated from a sound source (210). The diaphragm (201) is a thin membrane that vibrates due to sound waves located inside the directional microphone (200), and the directional microphone (200) can obtain information about sound by converting the vibration of the diaphragm (201) into an electrical signal.

[0050] In one embodiment, the directional microphone (200) may include a sound inlet (202) for receiving sound generated from a sound source. The sound inlet (202) may include holes through which sound generated from the sound source (210) passes to reach the diaphragm (201). For example, the sound inlet (202) may include a first sound inlet (202-1) located in front of the directional microphone (200) and a second sound inlet (202-2) located behind the directional microphone (200).

[0051] In one embodiment, when the sound source (210) generates sound, the generated sound may pass through both holes of the sound inlet (202) to reach the diaphragm (201). For example, some of the sound may pass through the first sound inlet (202-1) to first reach the diaphragm (201), and some of the sound may pass through the second sound inlet (202-2) to later reach the diaphragm (201).

[0052] In one embodiment, when a sound source (210) is positioned in front of a directional microphone (200) facing the sound inlet (202) and sound is generated, there is a phase difference due to time delay between the sound passing through the first sound inlet (202-1) and the sound passing through the second sound inlet (202-2), which causes movement of the diaphragm to be formed, enabling sound capture. In one embodiment, when a sound source (210) is positioned at the side of a directional microphone (200) not facing the sound inlet (202) and sound is generated, there is no phase difference due to time delay between the sound passing through the first sound inlet (202-1) and the sound passing through the second sound inlet (202-2), which combine in phases close to each other, so that the sounds can be canceled out.

[0053] In this way, the directional microphone (200) can intensively receive only sounds coming from a specific direction by receiving or canceling out sounds passing through the sound inlet (202).

[0054] FIG. 2b is a diagram for explaining a frequency response according to the distance between a sound source and a directional microphone according to one embodiment of the present disclosure.

[0055] In one embodiment, sounds received through a directional microphone may be received with higher intensity as the frequency of the sound decreases. In other words, sounds received through a directional microphone may be received with greater intensity as a lower frequency sound than as a higher frequency sound even if the sounds have the same intensity. This phenomenon may be referred to as the proximity effect of the directional microphone. This is because relatively low-frequency sounds have longer wavelengths than high-frequency sounds, and as the wavelength of the sound decreases, the sound energy per unit area decreases in proportion to the square of the distance due to diffraction. When a low-frequency sound source is positioned close to the directional microphone (200), the pressure difference between sounds arriving through the holes on both sides of the sound inlet (202) becomes large, so low-frequency sounds are received louder than high-frequency sounds.

[0056] Referring to FIG. 2B, the frequency response (250) may include first to fifth graphs (251, 252, 253, 254, 255). The first to fifth graphs (251, 252, 253, 254, 255) correspond to the distances between the sound source and the directional microphone (200) being '0.1 m', '0.5 m', '1 m', '3 m', and '5 m', respectively, and may represent the intensity of sound detected through the directional microphone (200) when sounds generated at the same intensity have different frequencies. Here, the horizontal axis of the frequency response (250) may represent the frequency in the frequency response, and the vertical axis may represent the relative gain value for the minimum intensity of the frequency response or the intensity of the frequency response corresponding to 10,000 Hz. However, it is not necessarily limited to the above-described example, and the vertical axis may also represent the relative gain value for the intensity of the frequency response of a preset frequency other than 10000 Hz.

[0057] Referring to the first graph (251), it can be confirmed that sounds with a frequency lower than 100 Hz are perceived as being 15 dB louder than sounds with a frequency higher than 10,000 Hz. That is, low-frequency sounds can be perceived as louder than high-frequency sounds due to the proximity effect of the directional microphone (200). In addition, referring to the second graph (252) and the third graph (253), it can be confirmed that as the distance between the sound source and the directional microphone (200) increases, the degree to which low-frequency sounds are perceived as louder than high-frequency sounds decreases. This is because the greater the distance between the sound source and the directional microphone (200), the greater the sound intensity decreases before the sound reaches the directional microphone (200), and therefore, the influence of the proximity effect becomes smaller at relatively far distances.

[0058] In other words, the greater the distance between the sound source and the directional microphone (200), the lower the amplification level of the low-frequency voice signal compared to the high-frequency voice signal may be as the voice signal received through the directional microphone (200).

[0059] In this way, an electronic device according to one embodiment of the present disclosure can determine whether the distance between a sound source and an electronic device is within a preset distance based on a change in the degree to which the intensity of a low-frequency voice signal is received according to the distance between the sound source and the electronic device.

[0060] FIG. 3 is a flowchart illustrating a method for an electronic device to perform voice recognition according to an embodiment of the present disclosure.

[0061] Referring to FIG. 3, the operation of the electronic device will be schematically described, and a detailed description of each operation will be described with reference to the drawings that follow. In addition, the operations of the electronic device described in the present disclosure can be understood as the operations of the first to fourth electronic devices (101, 102, 103, 104) illustrated in FIG. 1, the electronic device (1000) and the processor (1700) of the electronic device (1000) illustrated in FIG. 7, and the server (2000) and the processor (2300) of the server (2000) illustrated in FIG. 8. For the convenience of the description of the invention below, the electronic device (1000) illustrated in FIG. 7 will be described as performing the operations.

[0062] In step S310, the electronic device (1000) may receive a voice signal from a user. In one embodiment, the electronic device (1000) may include a microphone for receiving the voice signal. In one embodiment, the microphone may include a directional microphone. The electronic device (1000) may receive the voice signal through the microphone and analyze the received voice signal.

[0063] In one embodiment, a voice signal received from a user may include a wake-up word. Here, including a wake-up word may mean that voice signal components corresponding to the wake-up word are included in the received voice signal. In other words, the voice signal received from the user may include a voice signal for activating voice recognition based on a wake-up word pre-stored in the electronic device (1000) or for causing the electronic device (1000) to perform an operation corresponding to the voice signal. In one embodiment, the wake-up word included in the received voice signal may include a wake-up word pre-stored in the electronic device (1000). For example, if the wake-up word "Hi Bixby" is stored in the electronic device (1000), the received voice signal may include "Hi Bixby" spoken by the user.

[0064] In one embodiment, a voice signal received from a user may include a command related to voice recognition. Here, including a command related to voice recognition may mean that voice signal components corresponding to the command are included in the received voice signal. In one embodiment, the command related to voice recognition may include a plurality of preset words or sentences corresponding to operations that the electronic device (1000) can perform. For example, the command related to voice recognition may include "Set a 4-minute timer" corresponding to an operation in which the electronic device (1000) executes an application that sets a timer and sets the timer to 4 minutes.

[0065] In one embodiment, a voice signal received from a user may include both a wake-up word and a command related to voice recognition. For example, the received voice signal may include "Hi Bixby, set a 4-minute timer." In this case, the electronic device (1000) analyzes the received voice signal to identify the wake-up word "Hi Bixby" and the command related to voice recognition "set a 4-minute timer," respectively, and performs an operation corresponding to the voice signal, which will be described later, based on the identified wake-up word and the command related to voice recognition.

[0066] At step S320, the electronic device (1000) can analyze the frequency response of the received voice signal to obtain a feature vector of the wakeup word and a low frequency response below a preset frequency among the frequency responses.

[0067] In one embodiment, the electronic device (1000) may analyze the frequency response of the received voice signal. In one embodiment, analyzing the frequency response may mean converting a time-domain signal into a frequency domain and obtaining data necessary for performing voice recognition based on the voice signal converted into the frequency domain. For example, analyzing the frequency response may include digital conversion of the received voice signal, noise removal, frame segmentation that divides the voice signal into sections of a certain length, windowing that smoothes the boundaries of the signal, fast Fourier transform (FFT) that converts a time-domain signal into a frequency-domain signal, application of a Mel filterbank to perform a Mel scale that reflects human hearing characteristics, and extraction of Mel-Frequency Cepstral Coefficients (MFCC), but is not necessarily limited to the examples described above.

[0068] In one embodiment, the electronic device (1000) can obtain a feature vector of a wake-up word by analyzing a received frequency response. In one embodiment, the feature vector of the wake-up word may include, but is not necessarily limited to, MFCCs obtained by analyzing a frequency response of a voice signal.

[0069] In one embodiment, the electronic device (1000) may determine the preset frequency based on the characteristics of the directional microphone of the electronic device (1000). In one embodiment, the characteristics of the directional microphone may include information on the maximum frequency at which the proximity effect occurs. In other words, the characteristics of the directional microphone may mean a characteristic at which the maximum frequency at which the difference in the relative gain value of the frequency response is greater than or equal to a threshold value when the distance between the sound source and the directional microphone is a first distance and a second distance is determined.

[0070] In one embodiment, the electronic device (1000) may determine a preset frequency based on a maximum frequency determined according to the characteristics of the directional microphone. For example, referring to FIG. 2B, it can be confirmed that the difference in the relative gain values ​​of the first graph (251) and the second graph (252) in the frequency response (250) is from 1000 Hz or less to an exemplary threshold value of 2 dB or more. In this case, the electronic device (1000) may determine the preset frequency as 1000 Hz, or may determine a frequency within a certain range based on 1000 Hz as the preset frequency. However, the present invention is not necessarily limited to the above-described example, and the preset frequency may mean any frequency set by the user or the manufacturer of the electronic device (1000).

[0071] In one embodiment, the electronic device (1000) can obtain a low frequency response by extracting only the frequency response below a preset frequency from the frequency response of the received voice signal. For example, the electronic device (1000) can convert the received voice signal into a frequency response and extract only the frequency response below 1000 Hz from the converted frequency response as the low frequency response.

[0072] In one embodiment, the feature vector of the wakeup word may include a feature vector obtained from a low frequency response below a preset frequency among the frequency responses of the received voice signal. In one embodiment, the electronic device (1000) may obtain the feature vector of the wakeup word by applying a Mel Filterbank and extracting MFCCs for the low frequency response. Accordingly, when the electronic device (1000) performs voice recognition, the resistance to noise included in sounds above the preset frequency may be increased, and the computational cost required for voice recognition may be reduced.

[0073] At step S330, the electronic device (1000) can compare the acquired low frequency response with the low frequency response pre-stored in the memory based on identifying whether the feature vector of the acquired wake-up word corresponds to the feature vector of the wake-up word pre-stored in the memory of the electronic device.

[0074] In one embodiment, the electronic device (1000) may receive a voice signal for registering a wake-up word from a user. In one embodiment, the voice signal for registering the wake-up word may include a wake-up word that can be registered in the electronic device (1000). For example, the wake-up words that can be registered in the electronic device (1000) may include “Hi Bixby” and “Bixby,” and the voice signal received from the user may include one of them. In one embodiment, the electronic device (1000) may analyze a frequency response of the received voice signal to extract a low-frequency response below a preset frequency from among the feature vector and frequency response of the wake-up word, and store the extracted feature vector and low-frequency response of the wake-up word in a memory of the electronic device (1000).

[0075] In one embodiment, a voice signal for registering a wake-up word may be received from a sound source located a preset distance away from the electronic device (1000). In one embodiment, the electronic device (1000) may analyze a frequency response of a voice signal received from a sound source located a preset distance away from the electronic device (1000) to obtain a feature vector and a low frequency response of the wake-up word, and store the obtained feature vector and the low frequency response of the wake-up word in a memory of the electronic device (1000). Here, the preset distance may correspond to the preset distance of step 340, which will be described later.

[0076] In one embodiment, the electronic device (1000) can identify whether the feature vector of the acquired wake-up word corresponds to the feature vector of the pre-stored wake-up word. In one embodiment, the electronic device (1000) can calculate the similarity between the feature vector of the acquired wake-up word and the feature vector of the pre-stored wake-up word. For example, the electronic device (1000) can calculate the similarity between the feature vector of the wake-up word and the feature vector of the pre-stored wake-up word based on various algorithms that calculate similarity between different vectors, such as the Dynamic Time Warping (DTW) algorithm, Cosine Similarity, Euclidean Distance, etc. In one embodiment, the electronic device (1000) can identify that the feature vector of the acquired wake-up word corresponds to the feature vector of the pre-stored wake-up word if the calculated similarity exceeds a threshold value. In one embodiment, the electronic device (1000) can identify that the feature vector of the wake-up word does not correspond to the feature vector of the pre-stored wake-up word if the calculated similarity is below a threshold value.

[0077] In one embodiment, the electronic device (1000) can detect a plurality of peaks from each of the acquired low frequency response and the stored low frequency response.

[0078] In one embodiment, the electronic device (1000) can compare a plurality of peaks of the acquired low frequency response with a plurality of peaks of the stored low frequency response. In one embodiment, the plurality of peaks can be indexed or labeled in descending order of frequency. In one embodiment, the electronic device (1000) can compare a plurality of peaks by comparing peaks having the same index or the same label in each frequency response among the plurality of peaks of the acquired low frequency response and the stored low frequency response, respectively.

[0079] In one embodiment, the plurality of peaks may include harmonic frequencies of the speech signal. In one embodiment, the harmonic frequencies may include the fundamental frequency of the speech signal and harmonics corresponding to integer multiples of the fundamental frequency. However, the present invention is not necessarily limited thereto, and the plurality of peaks may include formant frequencies formed by vowel sounds in the speech signal, resonant frequencies which are frequency components amplified by the structural characteristics of the vocal tract, or frequencies at which the gain value exhibits a maximum value for each preset frequency range.

[0080] In one embodiment, the electronic device (1000) may obtain a user input including information about a preset distance. In one embodiment, the information about the preset distance may include information about the distance between the electronic device (1000) and the sound source of the voice signal that serves as the basis for obtaining the feature vector of the pre-stored low frequency response and the pre-stored wake-up word. For example, when registering a wake-up word from a user, the electronic device (1000) may obtain a user input that the distance between the user who uttered the voice signal and the electronic device (1000) is '50 cm', and determine the preset distance as '50 cm'.

[0081] In one embodiment, the electronic device (1000) may determine a second frequency lower than a first frequency, which is a preset frequency, based on information about a preset distance. In one embodiment, the determined second frequency may correspond to the first frequency, which is a preset frequency. In other words, the electronic device (1000) may determine a second frequency lower than the first frequency, which is a preset frequency, and perform an operation based on the preset frequency of the present disclosure based on the determined second frequency instead of the first frequency. In one embodiment, the electronic device (1000) may detect peaks lower than the second frequency as a plurality of peaks from each of the acquired low frequency response and the pre-stored low frequency response. In one embodiment, the electronic device (1000) may determine the second frequency to be lower as the preset distance becomes shorter. Specific details regarding the first frequency and the second frequency will be described again below with reference to FIGS. 4A and 4B.

[0082] In one embodiment, the electronic device (1000) can adjust the gain values ​​of multiple peaks of a pre-stored frequency response based on information about a preset distance. A detailed description of the method for adjusting the gain values ​​of multiple peaks will be described again below with reference to FIG. 5.

[0083] In step S340, the electronic device (1000) can determine whether the distance between the user and the electronic device (1000) is within a preset distance based on the comparison result. Here, the comparison result may refer to the result of comparing the acquired low frequency response performed in step S330 with the pre-stored low frequency response.

[0084] In one embodiment, the electronic device (1000) may calculate the difference between the gain values ​​of a first peak and a second peak among a plurality of peaks of the stored low frequency response. In one embodiment, the electronic device (1000) may calculate the difference between the gain values ​​of a third peak corresponding to the first peak and a fourth peak corresponding to the second peak among the plurality of peaks of the acquired low frequency response. In one embodiment, if the difference between the gain values ​​of the first peak and the second peak is greater than the difference between the gain values ​​of the third peak and the fourth peak, the electronic device (1000) may identify that the distance between the user and the electronic device (1000) is within a preset distance. For example, if the difference between the gain values ​​of the first peak and the second peak is 5 dB, and the difference between the gain values ​​of the third peak and the fourth peak is 2.5 dB, the electronic device (1000) may identify that the distance between the user and the electronic device (1000) is greater than or equal to the preset distance.

[0085] In one embodiment, the first peak may include a peak having a lowest frequency among a plurality of peaks of the pre-stored low frequency response. In one embodiment, the second peak may include a peak having a second lowest frequency among a plurality of peaks of the pre-stored low frequency response. For example, when the plurality of peaks include harmonic frequency components of a speech signal, the first peak may be a fundamental frequency component in the pre-stored low frequency response, and the second peak may be a first harmonic component in the pre-stored low frequency response. In this case, the third peak may be a fundamental frequency component of the acquired low frequency response, and the fourth peak may be a first harmonic component of the acquired low frequency response.

[0086] In one embodiment, the electronic device (1000) may calculate an average of the differences in the gain values ​​of adjacent peaks of the acquired low frequency response and may calculate an average of the differences in the gain values ​​of adjacent peaks of the pre-stored low frequency response. In one embodiment, the electronic device (1000) may identify that the distance between the user and the electronic device (1000) is within a preset distance if the average of the differences in the gain values ​​of adjacent peaks of the low frequency response is greater than the average of the differences in the gain values ​​of adjacent peaks of the pre-stored low frequency response.

[0087] FIG. 4A is a diagram illustrating a plurality of peaks detected below a first frequency of a low frequency response according to one embodiment of the present disclosure.

[0088] Referring to FIG. 4A, the frequency response may include a first low frequency response (410) and a second low frequency response (420). In one embodiment, the first low frequency response (410) may be a low frequency response pre-stored in the electronic device (1000), and the second low frequency response (420) may be a low frequency response acquired from a voice signal on which voice recognition is performed. In the frequency response of the present disclosure, a gain value may mean a relative gain value for a frequency component with the lowest intensity.

[0089] In one embodiment, the first low frequency response (410) and the second low frequency response (420) may be low frequency responses lower than a first frequency (Fa), which is a preset frequency among the frequency responses. For example, the first frequency (Fa) may be 1000 Hz, but is not necessarily limited thereto.

[0090] In one embodiment, the electronic device (1000) can detect a plurality of peaks from each of the first low frequency response (410) and the second low frequency response (420). In one embodiment, the plurality of peaks can include harmonic frequency components of a voice signal. For example, the fundamental frequency (F0) of the voice signal of the first low frequency response (410) and the voice signal of the second low frequency response (420) can be 120 Hz. Accordingly, the harmonic frequencies can be the fundamental frequency (F0) and the frequencies (F1, F2, F3, F4, F5, F6, F7) of first to seventh harmonics corresponding to integer multiples of the fundamental frequency. The electronic device (1000) can detect components corresponding to the harmonic frequencies as a plurality of peaks from each of the first low frequency response (410) and the second low frequency response (420).

[0091] In one embodiment, the electronic device (1000) may compare a plurality of peaks of the first low frequency response (410) with a plurality of peaks of the second low frequency response (420), and, based on the comparison result, determine whether a distance between a user who has uttered a voice signal of the second low frequency response (420) and the electronic device (1000) is within a preset distance. Here, the preset distance may correspond to a distance between a sound source of the voice signal of the first low frequency response (410) and the electronic device (1000) when the electronic device (1000) obtains the voice signal of the first low frequency response (410).

[0092] FIG. 4b is a diagram illustrating a plurality of peaks detected below a second frequency of a low frequency response according to one embodiment of the present disclosure.

[0093] Referring to FIG. 4b, the frequency response may include a first low frequency response (410) and a second low frequency response (420). The first low frequency response (410) and the second low frequency response (420) of FIG. 4b may correspond to the first low frequency response (410) and the second low frequency response (420) of FIG. 4a.

[0094] In one embodiment, the electronic device (1000) may calculate a difference (d1) between gain values ​​of a first peak (411) and a second peak (412) among a plurality of peaks of the first low frequency response (410). In one embodiment, the first peak (411) may be a peak having the lowest frequency among the plurality of peaks of the first low frequency response (410), and the second peak (412) may be a peak having the second lowest frequency among the plurality of peaks of the first low frequency response (410). For example, when the plurality of peaks include harmonic frequency components of a voice signal, a component of a fundamental frequency (F0), which is a peak having the lowest frequency among the plurality of peaks of the first low frequency response (410), may be determined as the first peak (411), and a component of a first harmonic frequency (F1), which is a peak having the second lowest frequency, may be determined as the second peak (412).

[0095] In one embodiment, the electronic device (1000) may calculate a difference (d2) between gain values ​​of a third peak (421) corresponding to a first peak (411) and a fourth peak (422) corresponding to a second peak (412) among a plurality of peaks of the second low frequency response (420). In one embodiment, the third peak (421) may be a peak having the lowest frequency among the plurality of peaks of the second low frequency response (420), and the fourth peak (422) may be a peak having the second lowest frequency among the plurality of peaks of the second low frequency response (420). For example, when the plurality of peaks include harmonic frequency components of a voice signal, a component of a fundamental frequency (F0), which is a peak having the lowest frequency among the plurality of peaks of the second low frequency response (420), may be determined as the third peak (421), and a component of a first harmonic frequency (F1), which is a peak having the second lowest frequency, may be determined as the fourth peak (422).

[0096] In one embodiment, the electronic device (1000) can identify that the distance between the user and the electronic device (1000) is within a preset distance if the difference (d1) between the gain values ​​of the first peak (411) and the second peak (412) is greater than the difference (d2) between the gain values ​​of the third peak (421) and the fourth peak (422). Here, the preset distance may correspond to the distance between the sound source of the voice signal of the first low frequency response (410) and the electronic device (1000).

[0097] For example, since the difference (d1) between the gain values ​​of the first peak (411) and the second peak (412) is 6 dB and the difference between the gain values ​​of the third peak (421) and the fourth peak (422) is 3 dB, the electronic device (1000) can identify that the distance between the user (or the location of the sound source) who uttered the voice signal of the second low frequency response (420) and the electronic device (1000) is within a preset distance, which is the distance between the user (or the location of the sound source) who uttered the voice signal of the first low frequency response (410) and the electronic device (1000). However, the present invention is not necessarily limited to the above-described example, and if the difference between the gain values ​​of any peaks other than the first peak (411) and the second peak (412) among the plurality of peaks of the first low frequency response (410) is greater than the difference between the gain values ​​of the peaks of the second low frequency response (420) corresponding to the corresponding peaks, the electronic device (1000) may identify that the distance between the user and the electronic device (1000) is within the preset distance. As another example, the electronic device (1000) may identify that the distance between the user and the electronic device (1000) is within the preset distance if the average of the differences between the gain values ​​of adjacent peaks among the plurality of peaks of the first low frequency response (410) is greater than the average of the differences between the gain values ​​of adjacent peaks among the plurality of peaks of the second low frequency response (420).

[0098] This is because the above-described examples mean that the phenomenon of the first low frequency response (410) receiving sound louder as the frequency decreases (proximity effect of a directional microphone) is greater than that of the second low frequency response (420).

[0099] In one embodiment, the electronic device (1000) may obtain a user input including information about a preset distance. In one embodiment, the electronic device (1000) may determine a second frequency (Fb) lower than a first frequency (Fa), which is a preset frequency, based on the information about the preset distance. In one embodiment, the electronic device (1000) may detect peaks lower than the second frequency (Fb) from each of the first low frequency response (410) and the second low frequency response (420) as a plurality of peaks. In one embodiment, the electronic device (1000) may determine that the second frequency (Fb) increases as the preset distance decreases.

[0100] For example, the electronic device (1000) may obtain a user input including information that the distance between the sound source of the voice signal of the first low frequency response (410) and the electronic device (1000) is '50 cm'. In one embodiment, the electronic device (1000) may store a distance-frequency table in which frequencies are mapped by distance section in the electronic device (1000). The electronic device (1000) may determine 500 Hz corresponding to '50 cm' as the second frequency (Fb) based on the pre-stored distance-frequency table. However, the present invention is not necessarily limited thereto, and the second frequency may be determined based on a formula or various algorithms that determine a frequency based on distance information. In one embodiment, the electronic device (1000) can detect frequency components of the fundamental frequency (F0) and first to third harmonic frequencies (F1, F2, F3) below the second frequency (Fb) as multiple peaks in each of the first low frequency response (410) and the second low frequency response (420).

[0101] In one embodiment, since the phenomenon of loud sounds being perceived as lower in frequency decreases as the distance between the sound source and the electronic device (1000) increases, the importance of low-frequency peak values, at which the phenomenon can be observed relatively better, increases. In one embodiment, the electronic device (1000) may further adjust the second frequency (Fb), which serves as a reference for detecting a plurality of peaks, to be lower based on information about a preset distance acquired based on a user input. Accordingly, it is possible to more accurately identify whether the distance between the user who uttered the voice signal of the second low-frequency response (420) and the electronic device (1000) is within the preset distance. In addition, since the number of multiple peaks to be compared is reduced, the computational cost required for the electronic device (1000) to perform voice recognition may be reduced.

[0102] FIG. 5 is a diagram for explaining adjusting the gain values ​​of multiple peaks according to one embodiment of the present disclosure.

[0103] Referring to FIG. 5, the frequency response may include a first low frequency response (510), a second low frequency response (520), and a third low frequency response (530). Here, the first low frequency response (510) may be a low frequency response pre-stored in the electronic device (1000), and the second low frequency response (520) may be a low frequency response obtained from a voice signal on which voice recognition is performed.

[0104] In one embodiment, the third low frequency response (530) may be a low frequency response corresponding to a preset distance. In one embodiment, the electronic device (1000) may store a low frequency response corresponding to each of a plurality of distance sections. In one embodiment, when the electronic device (1000) obtains a user input including information about the preset distance, the electronic device (1000) may obtain a low frequency response corresponding to a distance section including the preset distance based on the information about the preset distance. For example, when information that the preset distance is '50 cm' is obtained, the electronic device (1000) may obtain a third low frequency response (530) corresponding to a distance section including '50 cm' among low frequency responses corresponding to each of the plurality of pre-stored distance sections.

[0105] In one embodiment, the electronic device (1000) can adjust gain values ​​of a plurality of peaks of a pre-stored low frequency response based on information about a preset distance. In one embodiment, the electronic device (1000) can calculate a difference between the frequencies of the plurality of peaks of the pre-stored low frequency response and the frequencies of the plurality of peaks of the low frequency response obtained from a received voice signal. In one embodiment, the electronic device (1000) can obtain correction values ​​of the plurality of peaks of the pre-stored low frequency response based on the difference between the low frequency response corresponding to the preset distance and the calculated frequency. In one embodiment, the electronic device (1000) can adjust the gain values ​​of the plurality of pre-stored peaks by applying the obtained correction value of the gain to the gain values ​​of the plurality of peaks of the pre-stored low frequency response.

[0106] For example, the electronic device (1000) can obtain a third low frequency response (530) corresponding to a preset distance. Here, the frequencies of the first peak (511) and the second peak of the first low frequency response (510) may be F0_a and F1_a, respectively, and the frequencies of the first peak (521) and the second peak (522) of the second low frequency response (520) may be F0_b and F1_b, respectively. The electronic device (1000) can obtain a change in the gain value of the component (531) of F0_b with respect to the gain value of the component (532) of F0_a of the third low frequency response (530) as a correction value of the first peak (511) of the first low frequency response (510). Additionally, the electronic device (1000) can obtain the change in the gain value of the component (533) of F1_b with respect to the gain value of the component (534) of F1_a of the third low frequency response (530) as a correction value of the second peak (512) of the first low frequency response (510).

[0107] The electronic device (1000) can correct the gain values ​​of the first peak (511) and the second peak (512) by adding the respective correction values ​​obtained to the gain values ​​of the first peak (511) and the second peak (512) of the first low frequency response (510). In this case, the difference between the gain value of the first peak (511) and the gain value of the second peak (512) of the first low frequency response (510) may become larger than before as the gain value adjustment is performed.

[0108] In other words, before the gain value is adjusted, the difference between the gain values ​​of the first peak (511) and the second peak (512) of the first low frequency response (510) is smaller than the difference between the gain values ​​of the first peak (521) and the second peak (522) of the second low frequency response (520), but the difference between the gain values ​​of the first peak (511) and the second peak (512) of the first low frequency response (510) with the adjusted gain value may become larger than the difference between the gain values ​​of the first peak (521) and the second peak (522) of the second low frequency response (520).

[0109] In this way, the electronic device (1000) according to one embodiment of the present disclosure can more accurately determine whether the distance between the user who uttered the voice signal and the electronic device (1000) is within the preset distance by utilizing information about the preset distance obtained from the user.

[0110] FIG. 6 is a diagram for explaining an operation of an electronic device according to one embodiment of the present disclosure to obtain information about a preset distance.

[0111] In one embodiment, the electronic device (1000) may obtain user input that includes information about a preset distance.

[0112] For example, the electronic device (1000) may display a voice recognition UI (610). The voice recognition UI (610) may include a toggle (620) for determining whether to activate voice recognition, a first check box (631) and a second check box (632) for selecting a wake-up word, and a sub-UI (640) for entering a distance at which the wake-up word is recorded.

[0113] In one embodiment, the electronic device (1000) can obtain a user input on whether to activate voice recognition through the toggle (620). For example, when the electronic device (1000) obtains an input signal in which the user touches the toggle (620), the electronic device (1000) can change the activation status of voice recognition. In one embodiment, the electronic device (1000) can save power consumption by not receiving a voice signal when voice recognition is deactivated, or by not analyzing the frequency response of the received voice signal even if it is received. In one embodiment, the electronic device (1000) can perform various operations and functions related to voice recognition described in the present disclosure when voice recognition is activated.

[0114] In one embodiment, the electronic device (1000) may register a wake-up word based on a user input selecting one of the first check box (631) and the second check box (632). For example, when the electronic device (1000) obtains an input signal in which the user touches the first check box (631), the electronic device (1000) may extract and store a feature vector of 'Hi Bixby' from a voice signal received from the user. As another example, when the electronic device (1000) obtains an input signal in which the user touches the second check box (632), the electronic device (1000) may extract and store a feature vector of 'Bixby' from a voice signal received from the user.

[0115] In one embodiment, the electronic device (1000) may determine a wake-up word to be extracted from a feature vector based on a user input selecting one of the first check box (631) and the second check box (632). For example, the electronic device (1000) may store a feature vector of 'Hi Bixby' and a feature vector of 'Bixby' in advance. In this case, the electronic device (1000) may determine a feature vector of a wake-up word extracted from a voice signal on which voice recognition is performed based on a user input selecting one of the first check box (631) and the second check box (632) and a stored feature vector to be compared with the extracted wake-up word. For example, when the electronic device (1000) obtains an input signal in which the user touches the first check box (631), the electronic device can extract a feature vector of 'Hi Bixby' from the received voice signal and compare the extracted feature vector of 'Hi Bixby' with the feature vector of 'Hi Bixby' among the feature vectors of the stored wake-up word.

[0116] In one embodiment, the electronic device (1000) can obtain a user input including information about a preset distance through a sub-UI (640) that inputs a distance at which a wake-up word is recorded. For example, the sub-UI (640) can include a slider (641) and a handle (642). The electronic device (1000) can obtain a touch input signal (750) that causes the user to adjust the position of the handle (642) on the slider (641), and obtain '50 cm', which is a preset distance corresponding to the position of the handle (642), as information about the preset distance. However, the present invention is not necessarily limited to the above-described example, and the information about the preset distance can include a distance corresponding to a number directly input by the user, or a distance selected from among a plurality of preset distances based on the user input.

[0117] FIG. 7 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.

[0118] Referring to FIG. 7, the electronic device (1000) may include a microphone (1100), a memory (1200), a display (1300), a communication interface (1400), an input interface (1500), an output interface (1600), and a processor (1700). The microphone (1100), the memory (1200), the display (1300), the communication interface (1400), the input interface (1500), the output interface (1600), and the processor (1700) may each be electrically and / or physically connected to each other.

[0119] The components illustrated in FIG. 7 are merely in accordance with one embodiment of the present disclosure, and the components included in the electronic device (1000) are not limited to those illustrated in FIG. 7. The electronic device (1000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 7, and may further include components not illustrated in FIG. 7.

[0120] The microphone (1100) may be a component that converts sound into an electrical signal. The microphone (1100) may include a diaphragm and an electronic circuit for converting sound into an electrical signal. When the diaphragm vibrates due to sound waves, the electronic circuit may convert the vibration of the diaphragm into an electrical signal. The electrical signal generated by converting the sound may be transmitted to the processor (1700). In one embodiment, the microphone (1100) may include a directional microphone. In one embodiment, the microphone (1100) may include a directional system composed of a plurality of omnidirectional microphones. In one embodiment, the directional microphone may include a unidirectional microphone and a bidirectional microphone. However, the present invention is not limited thereto, and the microphone (1100) may also include a shotgun microphone. In one embodiment, the microphone (1100) may receive a voice signal from a user. In one embodiment, the voice signal received by the microphone (1100) may include a voice signal for registering a wake-up word in the electronic device (1000). In one embodiment, the voice signal received by the microphone (1100) may include a voice signal for activating voice recognition of the electronic device (1000) or performing voice recognition. In one embodiment, the voice signal received by the microphone (1100) may include at least one of a wake-up word and a command related to voice recognition.

[0121] The memory (1200) may store instructions or program codes for performing functions or operations of the electronic device (1000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (1200) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.

[0122] In one embodiment, the memory (1200) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a mask ROM, a flash ROM, a hard disk drive (HDD), or a solid state drive (SSD). The memory (1200) may not exist separately and may be configured to be included in the processor (1700). The memory (1200) may be configured as a volatile memory, a nonvolatile memory, or a combination of a volatile memory and a nonvolatile memory. A program or at least one instruction for performing operations according to embodiments described below may be stored in the memory (1200). The memory (1200) may also provide stored data to the processor (1700) at the request of the processor (1700).

[0123] In one embodiment, the memory (1200) may store a registered wake-up word (or a feature vector of the registered wake-up word) and a low frequency response. Here, the low frequency response may include a low frequency response below a preset frequency among the frequency responses of a voice signal that utters the registered wake-up word. In one embodiment, the processor (1700) may obtain the wake-up word and the low frequency response from a voice signal received through the microphone (1100), and transmit the obtained wake-up word and the low frequency response to the memory (1200). In one embodiment, the registered wake-up word and the low frequency response may be received from an external electronic device through the communication interface (1400) and stored in the memory (1200). However, the present invention is not limited to the above-described example, and the memory (1200) may further include various data necessary to perform the operations and functions of the electronic device (1000) disclosed in the present disclosure.

[0124] The display (1300) is a component for displaying images and / or videos. In one embodiment, the display (1300) may be configured as a physical device including at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display. In one embodiment, the display (1300) may display a UI for receiving a voice signal for registering a wake-up word from a user based on a signal received from the processor (1700). In one embodiment, the display (1300) may display a UI for receiving a user input including information about a preset distance based on the signal received from the processor (1700). Here, the preset distance may correspond to a distance between a user who has uttered a voice signal for registering a wake-up word and the electronic device (1000). However, it is not necessarily limited to the above-described example, and the display (1300) may display various UIs necessary for the electronic device (1000) to perform voice recognition or to display the results of voice recognition based on the signal received from the processor (1700).

[0125] In one embodiment, when the electronic device (1000) includes a head-mounted display device, the display (1300) may display a stereo image. In one embodiment, the stereo image may include a left image and a right image. In one embodiment, the display (1300) may include a first display corresponding to the left image (or the user's left eye) and a second display corresponding to the right image (or the user's right eye). In one embodiment, the display (1000) displays the left image through the first display and the right image through the second display, thereby allowing the user to experience a three-dimensional effect of the stereo image.

[0126] The communication interface (1400) is a component for the electronic device (1000) to communicate with an external electronic device. In one embodiment, the communication interface (1400) may perform data communication between the electronic device (1000) and the external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.

[0127] In one embodiment, the communication interface (1400) may receive a voice signal from an external electronic device. In one embodiment, the voice signal received through the communication interface (1400) may include a voice signal for registering a wake-up word, a voice signal for activating voice recognition of the electronic device (1000), or for performing an operation corresponding to the voice signal. In one embodiment, the communication interface (1400) may receive a wake-up word (or a feature vector of the wake-up word) and a low-frequency response obtained based on the voice signal received from the external electronic device. The wake-up word and the low-frequency response received through the communication interface (1400) may be transmitted and stored in the memory (1200), or transmitted to the processor (1700) and used to perform various operations and functions of the present disclosure.

[0128] The input interface (1500) is a component for receiving various user inputs. In one embodiment, the input interface (1500) may include a touch panel, a physical button, a microphone, etc. In one embodiment, information input through the input interface (1500) may be provided to the processor (1700). In one embodiment, a user input for determining whether to activate voice recognition of the electronic device (1000) may be obtained through the input interface (1500). In one embodiment, a user input for determining a wake-up word from which a feature vector is extracted may be obtained through the input interface (1500). In one embodiment, a user input including information on a preset distance may be obtained through the input interface (1500). However, the present invention is not limited to the above-described examples, and various data for performing operations and functions of the electronic device (1000) may be input through the input interface (1500).

[0129] The output interface (1600) is a component for the electronic device (1000) to provide various information to the user. In one embodiment, the electronic device (1000) may include a speaker, which is a component that outputs sound, or an indicator indicating that voice recognition based on a wake-up word has been activated. In one embodiment, the output interface (1600) may output voice corresponding to text displayed through the user interface based on a signal received from the processor (1700). In one embodiment, the output interface (1600) may output a voice signal including a wake-up word that can be registered in the electronic device (1000) based on a signal received from the processor (1700). In one embodiment, the output interface (1600) may output a voice signal indicating that voice recognition of the electronic device (1000) has been activated based on a signal received from the processor (1700). However, it is not necessarily limited to the above-described examples, and the output interface (1600) can output various voices or information for performing the operations and functions of the electronic device (1000) described in the present disclosure.

[0130] The processor (1700) can control the overall operations of the electronic device (1000). In one embodiment, the processor (1700) can include multiple processors. In one embodiment, at least one processor (1700) can perform the operations and functions of the electronic device (1000) described in the present disclosure by executing one or more instructions of a program stored in the memory (1200).

[0131] The processor (1700) may be configured as at least one of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), an Application Processor, a Neural Processing Unit, or an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto.

[0132] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first and second operations may be performed by a first processor and the third operation may be performed by a second processor. However, the embodiments of the present disclosure are not limited thereto.

[0133] One or more processors according to the present disclosure may be implemented as a single-core processor or a multi-core processor. If a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single core or by multiple cores included in one or more processors.

[0134] In one embodiment, at least one processor (1700) may receive a voice signal from a user by executing at least one command. In one embodiment, at least one processor (1700) may analyze a frequency response of the received voice signal by executing at least one command to obtain a low frequency response below a preset frequency among the feature vector and frequency response of the wake-up word. In one embodiment, at least one processor (1700) may compare the obtained low frequency response with the low frequency response pre-stored in the memory (1200) based on identifying whether the obtained feature vector of the wake-up word corresponds to a feature vector of the wake-up word pre-stored in the memory (1200) of the electronic device by executing at least one command. In one embodiment, at least one processor (1700) may determine whether the distance between the user and the electronic device is within the preset distance based on the comparison result by executing at least one command. In one embodiment, at least one processor (1700) can perform an operation corresponding to a voice signal based on a determination result by executing at least one instruction.

[0135] In one embodiment, the stored low frequency response can be obtained by analyzing the frequency response of a voice signal received from a sound source located a preset distance away from the electronic device.

[0136] In one embodiment, at least one processor (1700) may compare the acquired low frequency response with a pre-stored low frequency response by executing at least one instruction. In one embodiment, at least one processor (1700) may detect a plurality of peaks from each of the acquired low frequency response and the pre-stored low frequency response by executing at least one instruction. In one embodiment, at least one processor (1700) may compare a plurality of peaks of the acquired low frequency response with a plurality of peaks of the pre-stored low frequency response by executing at least one instruction.

[0137] In one embodiment, the plurality of peaks may comprise high frequency components of the speech signal.

[0138] In one embodiment, at least one processor (1700) may calculate a difference between gain values ​​of a first peak and a second peak among a plurality of peaks of a pre-stored low frequency response by executing at least one instruction. In one embodiment, at least one processor (1700) may calculate a difference between gain values ​​of a third peak corresponding to the first peak and a fourth peak corresponding to the second peak among a plurality of peaks of the acquired low frequency response by executing at least one instruction. In one embodiment, at least one processor (1700) may identify a distance between a user and an electronic device as being within a preset distance if a difference between gain values ​​of the first peak and the second peak is greater than a difference between gain values ​​of the third peak and the fourth peak by executing at least one instruction.

[0139] In one embodiment, the first peak may include a peak having the lowest frequency among the plurality of peaks of the stored low frequency response. In one embodiment, the second peak may include a peak having the second lowest frequency among the plurality of peaks of the stored low frequency response.

[0140] In one embodiment, at least one processor (1700) may obtain a user input including information about a preset distance by executing at least one command. In one embodiment, at least one processor (1700) may determine a second frequency lower than a first frequency, which is a preset frequency, based on the information about the preset distance by executing at least one command. In one embodiment, at least one processor (1700) may detect peaks lower than the second frequency as a plurality of peaks from each of the obtained low frequency response and the pre-stored low frequency response by executing at least one command.

[0141] In one embodiment, at least one processor (1700) may determine, by executing at least one instruction, that the second frequency is lowered as the preset distance becomes shorter.

[0142] In one embodiment, at least one processor (1700) can adjust gain values ​​of a plurality of peaks of a pre-stored low frequency response based on information about a preset distance by executing at least one instruction.

[0143] However, the present invention is not limited to the examples described above, and at least one processor (1700) can perform the operations and functions of the electronic device (1000) described in the present disclosure.

[0144] Figure 8 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.

[0145] Referring to FIG. 8, the server (2000) may include a memory (2100), a communication interface (2200), and a processor (2300). Each may be electrically and / or physically connected to each other.

[0146] The components illustrated in FIG. 8 are merely in accordance with one embodiment of the present disclosure, and the components included in the server (2000) are not limited to those illustrated in FIG. 8. The server (2000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 8, and may further include components not illustrated in FIG. 8.

[0147] The memory (2100) may store instructions or program codes for performing functions or operations of the server (2000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (2100) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or assembler.

[0148] In one embodiment, the memory (2100) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a mask ROM, a flash ROM, etc.), a hard disk drive (HDD), or a solid state drive (SSD).

[0149] In one embodiment, the memory (2100) may store a registered wake-up word (or a feature vector of the registered wake-up word) and a low frequency response. Here, the low frequency response may include a low frequency response below a preset frequency among the frequency responses of a voice signal uttering the registered wake-up word. In one embodiment, the wake-up word and low frequency response stored in the memory (2100) may be obtained based on the server (2000) receiving a voice signal from an external electronic device and the received voice signal. However, the present invention is not necessarily limited to the above-described example, and the memory (2100) may further include various data necessary for performing the operations and functions of the server (2000) described in the present disclosure.

[0150] The communication interface (2200) is a component for the server (2000) to communicate with an external electronic device. In one embodiment, the communication interface (2200) may perform data communication between the server (2000) and the external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.

[0151] In one embodiment, the communication interface (2200) can receive a voice signal from an external electronic device. In one embodiment, the voice signal received through the communication interface (2200) can include a voice signal for registering a wake-up word, a voice signal for activating voice recognition of the external electronic device, or a voice signal for performing an operation corresponding to the voice signal. In one embodiment, the communication interface (2200) can receive a wake-up word (or a feature vector of the wake-up word) and a low frequency response obtained based on the voice signal received from the external electronic device. In one embodiment, the communication interface (2200) can transmit a signal for activating voice recognition of the external electronic device or data indicating an operation corresponding to the received voice signal to the external electronic device.

[0152] The processor (2300) can control the overall operations of the server (2000). In one embodiment, the processor (2300) can include a plurality of processors. In one embodiment, at least one processor (2300) can perform the operations and functions of the electronic device (1000) described in the present disclosure by executing one or more instructions of a program stored in the memory (2100). In one embodiment, at least one processor (2300) can receive a voice signal from a user by executing at least one instruction. In one embodiment, at least one processor (2300) can analyze the frequency response of the received voice signal by executing at least one instruction to obtain a low frequency response below a preset frequency among the feature vector and frequency response of the wake-up word. In one embodiment, at least one processor (2300) may compare the acquired low frequency response with the low frequency response pre-stored in the memory (2100) based on identifying whether the feature vector of the acquired wake-up word corresponds to the feature vector of the wake-up word pre-stored in the memory (2100) of the electronic device by executing at least one instruction. In one embodiment, at least one processor (2300) may determine whether the distance between the user and the electronic device is within a preset distance based on the comparison result by executing at least one instruction. In one embodiment, at least one processor (2300) may perform an operation corresponding to the voice signal based on the determination result by executing at least one instruction. Since the operation and function of the processor (2300) correspond to the operation and function of the processor (1700) of the electronic device (1000) described in FIG. 7, a redundant description will be omitted.

[0153] In this way, according to one embodiment of the present disclosure, the server (2000) can receive a voice signal from an external electronic device through a communication interface (2200), and perform voice recognition on the received voice signal based on a feature vector of a wake-up word and a low frequency response stored in a memory (2100) of the server (2000). The server (2000) can transmit the result of the voice recognition on the voice signal to the external electronic device through the communication interface (2200).

[0154] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.

[0155] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0156] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0157] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. In a method for an electronic device (1000) to perform voice recognition, Step of receiving a voice signal from a user (S310); A step (S320) of analyzing the frequency response of the received voice signal to obtain a feature vector of a wakeup word and a low frequency response below a preset frequency among the frequency responses; A step (S330) of comparing the acquired low frequency response with a low frequency response below the preset frequency stored in the memory based on identifying whether the feature vector of the acquired wake-up word corresponds to the feature vector of the wake-up word stored in the memory of the electronic device; A step (S340) of determining whether the distance between the user and the electronic device is within a preset distance based on the comparison result; and A method comprising a step (S350) of performing an action corresponding to the voice signal based on the judgment result.

2. In paragraph 1, The above stored low frequency response is, A method obtained by analyzing the frequency response of a voice signal received from a sound source located at a preset distance from the electronic device.

3. In any one of paragraphs 1 and 2, The step of comparing the obtained low frequency response with the stored low frequency response is as follows: A step of detecting a plurality of peaks from each of the acquired low frequency response and the stored low frequency response; A method comprising: a step of comparing a plurality of peaks of the acquired low frequency response with a plurality of peaks of the stored low frequency response.

4. In paragraph 3, The step of determining whether the distance between the electronic devices is within the preset distance is: A step of calculating the difference between the gain values ​​of a first peak and a second peak among a plurality of peaks of the stored low frequency response; A step of calculating the difference between the gain values ​​of a third peak corresponding to the first peak and a fourth peak corresponding to the second peak among the plurality of peaks of the acquired low frequency response; and A method comprising: a step of identifying that the distance between the user and the electronic device is within the preset distance if the difference between the gain values ​​of the first peak and the second peak is greater than the difference between the gain values ​​of the third peak and the fourth peak.

5. In any one of paragraphs 3 to 4, A step of obtaining user input including information about the preset distance; The step of detecting the above multiple peaks is: A step of determining a second frequency lower than the first frequency, which is the preset frequency, based on information about the preset distance; and A method further comprising: a step of detecting peaks below the second frequency from each of the acquired low frequency response and the stored low frequency response as the plurality of peaks.

6. In paragraph 5, The step of determining the second frequency is: A method comprising: a step of determining that the second frequency is lowered as the preset distance becomes shorter; 7. In any one of paragraphs 5 to 6, A method further comprising: a step of adjusting gain values ​​of a plurality of peaks of the stored low frequency response based on information about the preset distance.

8. In an electronic device (1000) that performs voice recognition, Microphone (1100); A memory (1200) storing one or more instructions; and At least one processor (1700) operably coupled to the memory (1200) and including a processing circuit; The electronic device (1000) executes the instructions by at least one processor (1700) alone or in cooperation with each other, Receive a voice signal from the user, By analyzing the frequency response of the received voice signal, a feature vector of a wake-up word and a low frequency response below a preset frequency among the frequency responses are obtained, Based on identifying whether the feature vector of the acquired wake-up word corresponds to the feature vector of the wake-up word pre-stored in the memory of the electronic device, comparing the acquired low frequency response with the low frequency response pre-stored in the memory, Based on the comparison result, it is determined whether the distance between the user and the electronic device is within a preset distance, An electronic device that performs an action corresponding to the voice signal based on the judgment result.

9. In paragraph 8, The above stored low frequency response is, An electronic device obtained by analyzing the frequency response of a voice signal received from a sound source located at a preset distance from the electronic device.

10. In any one of paragraphs 8 to 9, The electronic device, by which at least one processor alone or in cooperation executes the instructions, Compare the obtained low frequency response with the stored low frequency response, Detecting a plurality of peaks from each of the acquired low frequency response and the stored low frequency response, An electronic device that compares a plurality of peaks of the acquired low frequency response with a plurality of peaks of the stored low frequency response.

11. In paragraph 10, The electronic device, by which at least one processor alone or in cooperation executes the instructions, Calculate the difference between the gain values ​​of the first peak and the second peak among the plurality of peaks of the above-mentioned stored low frequency response, Calculate the difference between the gain values ​​of the third peak corresponding to the first peak and the fourth peak corresponding to the second peak among the plurality of peaks of the obtained low frequency response, An electronic device that identifies the distance between the user and the electronic device as being within the preset distance if the difference between the gain values ​​of the first peak and the second peak is greater than the difference between the gain values ​​of the third peak and the fourth peak.

12. In any one of paragraphs 10 to 11, The electronic device, by which at least one processor alone or in cooperation executes the instructions, Obtaining user input containing information about the above preset distance, Based on the information about the above preset distance, a second frequency lower than the first frequency, which is the preset frequency, is determined, An electronic device that detects peaks below the second frequency as the plurality of peaks from each of the acquired low frequency response and the stored low frequency response.

13. In paragraph 12, The electronic device, by which at least one processor alone or in cooperation executes the instructions, An electronic device that determines that the second frequency is lowered as the preset distance becomes shorter.

14. In any one of paragraphs 12 to 13, The electronic device, by which at least one processor alone or in cooperation executes the instructions, An electronic device that adjusts gain values ​​of a plurality of peaks of the stored low frequency response based on information about the preset distance.

15. A computer-readable recording medium recording a program for executing the method of any one of clauses 1 to 7 on a computer.

Citation Information

Patent Citations

  • Audio source proximity estimation using sensor array for noise reduction

    KR1020110090940A

  • IoT Weighing device for resource circulation

    KR102703797B1

  • Speaker distance detection apparatus using microphone array and speech input / output apparatus

    US20040141418A1

  • Microphone proximity detection

    US20100080084A1

  • Distance-based automatic gain control and proximity-effect compensation

    US20140112483A1