Speech recognition method and apparatus, and recording medium
The voice recognition method uses IMU and proximity sensors to activate microphones based on device motion and proximity, ensuring fast and accurate voice recognition with user-controlled termination, addressing inconvenient triggers and misrecognition issues.
Patent Information
- Application Number
- PCT/KR2025/006292
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-25
- Filing Date
- 2025-05-09
- Publication Date
- 2026-01-29
AI Technical Summary
Existing voice recognition systems require inconvenient and time-consuming triggers for initiation, are prone to non-recognition or misrecognition, and lack user-controlled termination.
A voice recognition method utilizing IMU and proximity sensors to activate microphones only when a device is lifted and brought close to the user's mouth, analyzing voice properties to confirm proximity and user intent, and terminating recognition based on sensor signals.
Enables fast, accurate voice recognition by eliminating unnecessary triggers, reducing misrecognition, and allowing user-controlled termination.
Smart Images

Figure KR2025006292_29012026_PF_FP_ABST
Abstract
Description
Speech recognition method, device and recording medium
[0001] The present disclosure relates to a voice recognition method, device, and recording medium, and more particularly, to a method, device, and recording medium for recognizing voice using an IMU sensor signal and proximity voice detection.
[0002] As interest in user interfaces grows and voice processing technology advances, the number of IT devices with built-in voice recognition capabilities is increasing. Voice recognition is widely used in consumer devices such as AI speakers, smartphones, and smartwatches.
[0003] A trigger is essential for initiating voice recognition on a user device. For example, in devices like AI speakers, where the microphone for voice input is always active, voice recognition can be triggered by a specific wake command. For example, voice recognition supported by a smartphone's OS can be triggered by entering a wake command (e.g., "Siri," "Hey Google") or by pressing a button on the smartphone. For third-party apps like smartphones, voice recognition can be initiated after unlocking the smartphone, launching the app, and selecting the voice recognition icon.
[0004] Initiating voice recognition on a user device requires inputting a trigger, such as a call command, recognizing and confirming it, or separately launching the relevant app and inputting the voice recognition trigger. This process is not only inconvenient for the user, but can also be time-consuming. This is especially true when launching a third-party smartphone app through voice recognition.
[0005] Furthermore, when voice recognition is initiated via conventional triggers, non-recognition or misrecognition of the voice can often occur. Furthermore, there is no separate means to terminate voice recognition after it has begun, which can lead to problems such as voice recognition being terminated or continuing contrary to the user's intent.
[0006] The present disclosure is intended to solve the problems of the above-described prior art.
[0007] The present disclosure aims to provide a method and device capable of quickly performing voice recognition by omitting unnecessary steps for starting voice recognition.
[0008] In addition, the present disclosure aims to provide a voice recognition method and device capable of suppressing non-recognition and misrecognition that may occur during the voice recognition process.
[0009] In addition, the present disclosure also aims to provide a voice input method and device that can terminate voice recognition according to a user's intention.
[0010] A voice input method according to one embodiment of the present disclosure includes the steps of receiving a sensor signal from an IMU sensor of a user device, activating a microphone of the user device when it is determined that the user device is lifted or decelerates rapidly to a stop based on the sensor signal, receiving a voice signal input through the microphone, analyzing the voice signal to determine whether a voice input into the microphone is a proximity voice, and performing voice recognition when the voice input into the microphone is a proximity voice.
[0011] According to one embodiment of the present disclosure, in the step of determining whether a voice is a proximity voice, it is possible to determine whether a voice is a proximity voice by referring to the properties of a voice input through a microphone.
[0012] According to one embodiment of the present disclosure, in the step of determining whether a voice is a proximity voice, the speech distance can be estimated from the properties and sound pressure of the voice input through the microphone by referring to the correlation between the speech distance and sound pressure according to the properties of the voice.
[0013] According to one embodiment of the present disclosure, in the step of determining whether or not it is a proximity voice, it can be determined that the voice input through the microphone is a proximity voice if the sound pressure of the voice input through the microphone is higher than a threshold sound pressure.
[0014] According to one embodiment of the present disclosure, in the step of determining whether a voice is a proximity voice, feature information according to proximity speech can be detected from the voice input through the microphone to determine whether the voice is a proximity voice. Here, the feature information according to proximity speech can include at least one of a pop sound, a breathing sound, a body contact sound, a proximity effect, and an absorption phenomenon.
[0015] According to one embodiment of the present disclosure, in the step of receiving a voice signal, a plurality of voice signals input through a plurality of microphones of a user device can be received, and in the step of determining whether it is a proximity voice, the plurality of voice signals can be analyzed to evaluate the similarity between the plurality of voice signals and determine whether it is a proximity voice based on the similarity.
[0016] According to one embodiment of the present disclosure, in the step of determining whether a voice is a proximity voice, it is possible to determine whether a voice input through a microphone is a voice of a pre-registered user, and if it is not a voice of a pre-registered user, it is possible to determine that it is not a proximity voice.
[0017] According to one embodiment of the present disclosure, in the step of receiving a sensor signal, a proximity sensor signal may be further received from a proximity sensor of the user device. In addition, in the step of activating a microphone, if it is determined that the user device has rapidly decelerated and stopped while simultaneously approaching an object, the microphone of the user device may be activated based on a sensor signal received from an IMU sensor of the user device and a proximity sensor signal received from a proximity sensor of the user device.
[0018] A voice recognition method according to one embodiment of the present disclosure may further include a step of disabling a microphone of the user device. In this step, the microphone of the user device may be disabled if a sensor signal received from an IMU sensor of the user device is analyzed to determine that the user device is descending or moving rapidly.
[0019] According to one embodiment of the present disclosure, in the step of receiving a sensor signal, a proximity sensor signal may be further received from a proximity sensor of the user device, and in the step of disabling a microphone of the user device, if it is determined that the user device is rapidly accelerating and moving away from an object based on a sensor signal received from an IMU sensor of the user device and a proximity sensor signal received from a proximity sensor of the user device, the microphone of the user device may be disabled.
[0020] A voice recognition device according to one embodiment of the present disclosure includes a sensor signal receiving unit that receives a sensor signal from an IMU sensor of a user device, a microphone control unit that controls activation and deactivation of a microphone of the user device based on the sensor signal, a voice signal receiving unit that receives a voice signal input through a microphone, a proximity voice determination unit that analyzes the voice signal to determine whether a voice input into the microphone is a proximity voice, and a voice recognition unit that performs voice recognition when the voice input into the microphone is a proximity voice.
[0021] In addition, other methods for implementing the present disclosure, other devices, and computer-readable recording media for executing the above methods may be further provided.
[0022] According to one embodiment of the present disclosure, fast and accurate voice recognition can be achieved by performing voice recognition through detection of an IMU sensor signal and a proximity voice.
[0023] In addition, according to one embodiment of the present disclosure, since activation of the microphone is controlled through an IMU sensor signal and voice recognition is performed only when the voice input through the microphone is a proximity voice, the problem of non-recognition or misrecognition of voice can be suppressed.
[0024] In addition, according to one embodiment of the present disclosure, voice recognition can be terminated by disabling the microphone through an IMU sensor signal or the like, thereby controlling the termination of voice recognition in accordance with the user's intention.
[0025] FIG. 1 is a diagram exemplarily showing a voice input performed through a user device according to one embodiment of the present disclosure.
[0026] FIG. 2 is a diagram exemplarily showing the configuration of a user device for voice input according to one embodiment of the present disclosure.
[0027] FIG. 3 is a drawing exemplarily showing the functional configuration of a voice recognition device according to one embodiment of the present disclosure.
[0028] FIG. 4 is a diagram exemplarily showing how voice input is performed in various user devices according to one embodiment of the present disclosure.
[0029] FIG. 5 is a flowchart showing a voice recognition process according to one embodiment of the present disclosure.
[0030] [Explanation of symbols]
[0031] 100: User device
[0032] 200: Voice input device
[0033] 201: Sensor section
[0034] 203: Mike
[0035] 205: Communications Department
[0036] 207: Processor
[0037] 301: Sensor signal receiving unit
[0038] 303: Microphone control unit
[0039] 305: Voice signal receiver
[0040] 307: Proximity voice discrimination unit
[0041] 309: Voice recognition unit
[0042] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. Hereinafter, specific descriptions of previously known functions and configurations will be omitted if deemed likely to unnecessarily obscure the gist of the present disclosure. Furthermore, it should be noted that the following description relates only to one embodiment of the present disclosure and that the present disclosure is not limited thereto.
[0043] The terminology used in this disclosure is merely used to describe specific embodiments and is not intended to limit the present disclosure. For example, a component expressed in the singular should be understood to include plural components unless the context clearly indicates only a singular meaning. The term "and / or" used in this disclosure should be understood to encompass any and all possible combinations of one or more of the listed items. The terms "comprise," "include," or "have" used in this disclosure are intended to specify only the presence of a feature, number, step, motion, component, or combination thereof described in this disclosure, and the use of such terms does not exclude the presence or addition of one or more other features, numbers, steps, motions, components, or combinations thereof.
[0044] In the embodiments of the present disclosure, a "module" or "part" refers to a functional part that performs at least one function or motion, and may be implemented by hardware or software, or a combination of hardware and software. Furthermore, a plurality of "modules" or "parts" may be integrated into at least one software module and implemented by at least one processor, excluding "modules" or "parts" that need to be implemented by specific hardware.
[0045] Additionally, all terms used in this disclosure, including technical or scientific terms, unless otherwise defined, have the same meaning as commonly understood by those of ordinary skill in the art to which this disclosure pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with the contextual meaning of the relevant technology, and should not be interpreted in an unduly restrictive or expansive manner unless explicitly defined otherwise in this disclosure.
[0046] FIG. 1 is a diagram illustrating an example of a voice input performed through a user device according to one embodiment of the present disclosure. FIG. 1 illustrates a case where the user device performing the voice input is a smartphone.
[0047] Referring to FIG. 1, in one embodiment of the present disclosure, voice input is achieved by the user lifting the user device (100) and bringing it close to the mouth, followed by a proximity speech sound. In other words, the user's voice input is achieved through the process of the user lifting the user device (100) and bringing it close to the mouth, and the process of the user speaking at a location close to the user device (100).
[0048] According to one embodiment of the present disclosure, when the user device (100) is lifted or brought close to the user's mouth, the microphone of the user device (100) is activated, and if the voice input through the activated microphone is a voice by proximity speech, i.e., a proximity speech, it is determined that the voice is the voice intended by the user and voice recognition is performed. In this way, according to one embodiment of the present disclosure, voice recognition is performed only when the user device (100) is lifted or brought close to the user's mouth and a voice input is by proximity speech, thereby preventing non-recognition or misrecognition of the voice.
[0049] FIG. 2 is a diagram exemplarily showing the configuration of a user device for voice input according to one embodiment of the present disclosure.
[0050] A user device (200) according to one embodiment of the present disclosure is a device for a user to input voice, and may be, for example, a smartphone. However, the user device (200) is not limited thereto, and the user device (200) is a digital device equipped with a memory means and a microprocessor to provide computational capabilities, and may be a wearable device such as a smart ring, a smart watch, or a smart band, or an electronic device such as a smart pad, an AI speaker, or a PDA.
[0051] Referring to FIG. 2, the user device (200) may include a sensor unit (201), a microphone (203), a memory (205), a communication unit (207), and a processor (209). The components illustrated in FIG. 2 do not reflect all functions of the user device (200) and are not essential, so the user device (200) may include more or fewer components than the illustrated components.
[0052] According to one embodiment of the present disclosure, the sensor unit (201) of the user device (200) can perform a function of detecting an operating state or an external environmental state of the user device (200). In one embodiment, the sensor unit (201) can perform a function of detecting movement of the user device (200). To this end, the sensor unit (201) can include an IMU sensor. The IMU sensor of the sensor unit (201) can detect movement of the user device (200) (e.g., change in height, acceleration, deceleration, etc.), and the sensor unit (201) can generate a sensor signal regarding the movement of the user device (200) detected by the IMU sensor. In one embodiment, the sensor unit (201) further includes a proximity sensor to generate a sensor signal regarding distance or proximity information between the user device (200) and an object.
[0053] According to one embodiment of the present disclosure, the microphone (203) of the user device (200) is configured for voice input and can perform a function of detecting the user's voice and generating a voice signal. In order to prevent noise other than voice for voice recognition from being input and recognized through the microphone (203), the microphone (203) can be activated or deactivated when a predetermined condition is satisfied, as described below.
[0054] In one embodiment, the user device (200) may include a plurality of microphones (203). For example, the user device (200) may include a first microphone and a second microphone, and the first microphone and the second microphone may generate a first voice signal and a second voice signal, respectively, from voice input from the same sound source.
[0055] According to one embodiment of the present disclosure, the memory (205) of the user device (200) may store data used by components of the user device (200). For example, the memory (205) may store an operating system (OS) for operating the user device (200), and in addition, a software program or application for operating at least one of the sensor unit (201), the microphone (203), the communication unit (207), and the processor (209) may be stored. The memory (205) may include at least one of a volatile memory and a non-volatile memory.
[0056] According to one embodiment of the present disclosure, the communication unit (207) of the user device (200) may function to enable the user device (200) to communicate with an external communication network. For example, the communication unit (207) may include a communication module for performing a known communication method such as WiFi communication, WiFi-Direct communication, Long Term Evolution (LTE) communication, 5G communication, Bluetooth communication (including Bluetooth Low Energy (BLE) communication), infrared communication, or ultrasonic communication.
[0057] According to one embodiment of the present disclosure, the processor (209) of the user device (200) may perform a function of controlling the overall operation of the user device (200). In one embodiment, the processor (209) may control at least one other component (e.g., hardware or software component) of the user device (200) and may perform various data processing or calculations. For example, the processor (209) may be electrically connected to components of the user device (200), such as a sensor unit (201), a microphone (203), a memory (205), and a communication unit (207), and may control operations thereof.
[0058] A voice recognition device according to one embodiment of the present disclosure performs a function of determining and recognizing a voice intended to be input by a user from a voice input from a user device (200).
[0059] A voice recognition device according to one embodiment of the present disclosure may be included in a user device (200), or at least some of its components may be included in the user device (200). For example, the voice recognition device may be included in the user device (200) in the form of an application. Such an application may be downloaded from an external application distribution server (not shown). Here, at least a portion of the application may be replaced with a hardware device or firmware device that can perform functions substantially identical to or equivalent to the application, as needed.
[0060] Alternatively, the voice recognition device may be implemented in an external device capable of communicating with the user device (200) via a communication network or a predetermined processor. The external device equipped with the voice recognition device may be a server system, and the server system equipped with the voice recognition device may communicate with the user device (200) via a communication network.
[0061] FIG. 3 is a drawing exemplarily showing the functional configuration of a voice recognition device according to one embodiment of the present disclosure. Hereinafter, a device and method for recognizing voice input into a user device will be specifically described with reference to FIG. 3.
[0062] Referring to FIG. 3, a voice recognition device (300) according to an embodiment of the present disclosure may include a sensor signal receiving unit (301), a microphone control unit (303), a voice signal receiving unit (305), a proximity voice determination unit (307), and a voice recognition unit (309). The components illustrated in FIG. 3 do not reflect all functions of the voice recognition device (300) and are not essential, so the voice recognition device (300) may include more or fewer components than the illustrated components.
[0063] A sensor signal receiving unit (301) of a voice recognition device (300) according to one embodiment of the present disclosure can perform a function of receiving a sensor signal from a sensor unit (201) of a user device (200).
[0064] In one embodiment, the sensor signal receiving unit (301) may receive an IMU sensor signal from an IMU sensor of the user device (200). The IMU sensor signal received by the sensor signal receiving unit (301) may include information about movement of the user device (200), including information about height changes, acceleration, and deceleration of the user device (200).
[0065] In one embodiment, the sensor signal receiving unit (301) may receive a proximity sensor signal from a proximity sensor of the sensor unit (201) of the user device (200). The proximity sensor signal received by the sensor signal receiving unit (301) may include proximity information or distance information between the user device (200) and an object (e.g., the user's mouth or lips).
[0066] A microphone control unit (303) of a voice recognition device (300) according to one embodiment of the present disclosure may perform a function of controlling activation and deactivation of a microphone (203) of a user device (200) based on a sensor signal. In one embodiment, the microphone control unit (303) may control activation and deactivation of the microphone based on an IMU sensor signal and / or a proximity sensor signal received by a sensor signal receiving unit (301).
[0067] For example, the microphone control unit (303) may activate the microphone (203) when it is determined through an IMU sensor signal that the user device (200) is being lifted while the microphone (203) of the user device (200) is inactive. In addition, the microphone control unit (303) may activate the microphone (203) when it is determined through an IMU sensor signal that the user device (200) is rapidly decelerating and stopping while the microphone (203) of the user device (200) is inactive. At this time, the microphone control unit (303) may consider whether the user device (200) is approaching an object (e.g., the user's mouth or lips) based on a proximity sensor signal, and may activate the microphone (203) when the user device (200) is approaching the object while rapidly decelerating.
[0068] For another example, the microphone control unit (303) may deactivate the microphone (203) when it is determined through an IMU sensor signal that the user device (200) is being lowered while the microphone (203) of the user device (200) is activated. In addition, the microphone control unit (303) may deactivate the microphone (203) when it is determined through an IMU sensor signal that the user device (200) is moving rapidly while the microphone (203) of the user device (200) is activated. At this time, the microphone control unit (303) may consider whether the user device (200) is moving away from an object (e.g., the user's mouth or lips) based on a proximity sensor signal, and may deactivate the microphone (203) of the user device (200) when the user device (200) is moving rapidly while moving away from the object.
[0069] In contrast, the microphone control unit (303) can deactivate the microphone (203) of the user device (200) even when a preset command for terminating voice recognition is input while the microphone (203) of the user device (200) is activated.
[0070] In this way, since the microphone control unit (303) activates the microphone (203) based on the IMU sensor signal, voice recognition becomes possible by lifting the user device (200) or bringing it in front of the mouth, so that voice input becomes possible immediately without a separate trigger for voice recognition.
[0071] According to one embodiment of the present disclosure, the microphone control unit (303) controls the activation of the microphone (203) based on an IMU sensor signal, and activates the microphone (203) in all cases where a user's voice input is expected, thereby preventing the problem of non-recognition caused by the user's voice not being input. In addition, voice or noise input without the user's explicit intention of voice input is not input through the microphone (203), thereby resolving the problem of misrecognition.
[0072] In addition, power can be efficiently utilized by activating the microphone only when there is a possibility of voice input into the microphone (203) of the user device (200) (e.g., when the user device (200) is lifted or when the user device (200) is brought close to the mouth). In addition, the start as well as the end of voice input can be explicitly controlled based on an IMU sensor signal, a proximity sensor signal, or a user's command.
[0073] Meanwhile, if the user device (200) is equipped with multiple microphones, the microphone control unit (303) can control the activation and deactivation of these multiple microphones simultaneously. For example, the microphone control unit (303) can control the activation or deactivation of the first microphone and the second microphone simultaneously.
[0074] A voice signal receiving unit (305) of a voice recognition device (300) according to one embodiment of the present disclosure may perform a function of receiving a voice signal from a microphone (203) of a user device (200). In one embodiment, the voice signal receiving unit (305) may receive a voice signal regarding a voice input through a microphone (203) when the microphone (203) is activated by a microphone control unit (303). At this time, the voice signal receiving unit (305) may also receive information regarding the volume or sound pressure of the voice.
[0075] When the user device (200) is equipped with multiple microphones, the voice signal receiving unit (305) can receive voice signals through each of the multiple microphones. For example, the voice signal receiving unit (305) can receive a first voice signal and a second voice signal through a first microphone and a second microphone, respectively.
[0076] The proximity voice determination unit (307) of the voice recognition device (300) according to one embodiment of the present disclosure can perform a function of analyzing a voice signal to determine whether a voice input into the microphone (203) is a proximity voice.
[0077] According to one embodiment of the present disclosure, the proximity voice determination unit (307) can classify the properties of an input voice and determine whether it is a proximity voice based on the properties.
[0078] Voice is composed of a combination of vowels and consonants, and the resulting sounds encompass a wide frequency spectrum. These vocal properties can vary depending on how the voice is produced.
[0079] For example, a normal human voice produces sound by vibrating the vocal cords. Therefore, if the voice is too soft, the vocal cords will not vibrate properly, making it difficult to produce a clear sound. Consequently, the volume cannot be lowered below a certain level. In contrast, in a whisper, vowels are primarily produced by airflow, while consonants are pronounced without vocal cord vibration. Consequently, the high-frequency range of a whisper is relatively less emphasized in the voice spectrum than in a normal voice, resulting in a lower, softer tone. Shouting, on the other hand, is a sound produced by applying significant pressure to the vocal cords, utilizing much stronger vocal cord vibration and airflow than in a normal voice. The voice spectrum of a shout is broader than that of a normal voice, and is characterized by increased energy in the low-frequency and high-frequency ranges. This gives a shout a powerful and penetrating tone.
[0080] Thus, whispering is characterized by low volume, low pitch, a spectral shape with decreasing energy in the low-frequency band, periodicity, and low vocal effort, resulting in a steep spectral slope and high attenuation. On the other hand, shouting is characterized by high volume, high pitch, a spectral shape with increasing energy in the high-frequency band, non-periodic characteristics, and high vocal effort, resulting in a gentle spectral slope (i.e., high high-frequency energy) and low attenuation. Normal speech is characterized by an intermediate level of volume, pitch, vocal effort, spectral slope, and attenuation, and a clear periodicity.
[0081] In one embodiment, the proximity voice discriminator (307) can classify the properties of a voice based on these differences between a normal voice, a whisper, and a scream. Specifically, the proximity voice discriminator (307) can classify a voice as a whisper if the ratio of the volume of a high frequency band (e.g., a frequency band of 1 kHz or higher) to the volume of a low frequency band (e.g., a frequency band of 500 Hz or lower) in the acquired voice signal is below a predetermined value. Conversely, if the ratio of the volume of a high frequency band to the volume of a low frequency band exceeds a predetermined value, the voice can be classified as a normal voice. Meanwhile, a scream emphasizes the volume of a high frequency band compared to a normal voice, and at the same time, the volume of a low frequency band also increases compared to a normal voice, showing an increase in volume across the entire frequency band. Therefore, it is possible to determine whether a voice corresponds to a scream by comparing the spectral shape together with the ratio of the volume of a high frequency band to the volume of a low frequency band with the characteristics of a scream voice signal.
[0082] In one embodiment, the proximity voice discrimination unit (307) may classify voice attributes using an artificial intelligence-based classification model. For example, the proximity voice discrimination unit (307) may classify voice attributes using a voice attribute classification model learned based on the spectra of multiple accumulated input voice signals.
[0083] According to one embodiment of the present disclosure, the proximity voice determination unit (307) can determine whether an input voice is a proximity voice based on the voice properties and sound pressure.
[0084] In one embodiment, the proximity voice determination unit (307) calculates the speech distance of a voice by referring to the correlation between the speech distance and sound pressure according to the voice properties, and can determine whether it is a proximity voice based on this. For example, if the voice property is classified as a whisper and the sound pressure is 62 dBA, the speech distance can be estimated as 5 cm by referring to the relationship between the sound pressure in a whisper and the speech distance, and since the estimated speech distance is less than a predetermined distance (e.g., 10 cm), the voice can be determined as a proximity voice.
[0085] In one embodiment, the proximity voice determination unit (307) can determine a voice as a proximity voice if the sound pressure of the voice is higher than a preset threshold sound pressure. For example, based on a speech distance of 5 cm, the threshold sound pressures for whispering and normal voice properties can be set to 62 dBA and 68 dBA, respectively, and if the sound pressure of the voice is higher than the threshold sound pressure according to the voice property, the voice can be determined as a proximity voice.
[0086] In one embodiment, the proximity voice determination unit (307) may obtain distance information from a proximity sensor signal received through the sensor signal reception unit (301), and may determine whether a voice is a proximity voice by referring to the speech distance and distance information calculated based on the voice properties and sound pressure. For example, if the voice property is classified as a normal voice and the sound pressure is 62 dB, the speech distance may be estimated to be 10 cm by referring to the relationship between the sound pressure and speech distance in a normal voice. At this time, if the distance information obtained from the proximity sensor signal is 11 cm and the speech distance and the distance information are similar, the voice may be determined to be a proximity voice. On the other hand, if the distance information obtained from the proximity sensor signal is 25 cm and the speech distance and the distance information are not similar, the voice may be determined to be one that the user did not intend to recognize or other noise.
[0087] According to one embodiment of the present disclosure, the proximity voice determination unit (307) may determine whether a voice is a proximity voice by analyzing a characteristic phenomenon according to a proximity speech.
[0088] In one embodiment, the proximity voice determination unit (307) can analyze a voice signal and determine that the voice is a proximity voice if a popping sound, a so-called pop sound, is included in the voice. When a user brings his / her mouth close to the microphone (203) of the user device (200) and speaks, a pop sound may be generated when the air exhaled while speaking hits the microphone (203). The proximity voice determination unit (307) can detect a pop sound in the voice signal by analyzing the voice signal in the frequency domain or time domain, and determine that the voice is a proximity voice if the pop sound is included.
[0089] Meanwhile, various characteristics other than pop sounds may be expressed depending on the close-range speech. For example, the proximity effect may occur, in which the sensitivity to low sounds increases as the distance between the sound source and the microphone (203) gets closer, and breathing sounds such as breath or sniffing may be mixed in with the voice, or sounds caused by contact between body parts (e.g., the sound of a hand or finger touching the face) may be mixed in, and as the microphone (203) gets closer to the facial skin, sounds may be absorbed by the skin, reducing background noise.
[0090] In one embodiment, the proximity voice determination unit (307) may detect feature information according to such close-range speech from a voice signal in addition to a pop sound to determine whether it is a proximity voice. For example, the proximity voice determination unit (307) may detect feature information associated with the proximity effect by calculating the ratio of the size of the low-pitched and mid-pitched ranges and calculating the relative size of the low-pitched range through analysis of the frequency domain of the voice signal. Alternatively, the proximity voice determination unit (307) may detect feature information associated with breath sounds by comparing the waveform features or frequency features of the voice signal with the waveform features or frequency features of a preset reference breath sound. Alternatively, the proximity voice determination unit (307) may detect feature information associated with sounds caused by contact between body parts by comparing the waveform features or frequency features of the voice signal with the waveform features or frequency features of sounds caused by contact between body parts. Alternatively, the proximity sound discrimination unit (307) may detect feature information associated with sound absorption by the user's skin by analyzing the degree of reduction in background noise level.
[0091] According to one embodiment of the present disclosure, the proximity voice determination unit (307) may determine whether a proximity voice is received based on a plurality of voice signals received from a plurality of microphones.
[0092] In the case of proximity speech due to close-range speech, the relative difference between the voice input paths through multiple microphones, for example, a first microphone and a second microphone, is greater than that of non-proximity speech, and therefore, the sound pressure level difference between the first voice signal from the first microphone and the second voice signal from the second microphone appears large. Conversely, in the case of non-proximity speech, the sound pressure level difference between the first voice signal from the first microphone and the second voice signal from the second microphone appears small.
[0093] In one embodiment, the proximity voice determination unit (307) analyzes voice signals from multiple microphones to evaluate the similarity between the voice signals from the multiple microphones, and determines whether the voice is a proximity voice based on the similarity. At this time, the similarity may be evaluated with reference to at least one of the direction of the sound source and the relative difference in the voice input paths from the sound source to the multiple microphones.
[0094] For example, if the similarity between voice signals from multiple microphones is greater than a preset threshold, it can be determined whether the voice is a proximity voice. As another example, the proximity voice determination unit (307) analyzes voice signals from multiple microphones, compares the sound pressure between the voice signals from the multiple microphones, and determines that the voice is a proximity voice if the difference in sound pressure exceeds a preset threshold.
[0095] According to one embodiment of the present disclosure, the proximity voice determination unit (307) can determine whether an input voice is the voice of a pre-registered user. The proximity voice determination unit (307) can determine whether the input voice matches the voice of the pre-registered user through frequency domain analysis of the voice signal, and if the input voice is determined to be the voice of the pre-registered user, the voice can be determined to be a proximity voice. If the input voice is determined not to be the voice of the pre-registered user, the voice can be determined to be noise that is not a target of voice recognition.
[0096] The voice recognition unit (309) of the voice recognition device (300) according to one embodiment of the present disclosure can perform voice recognition when it is determined that the input voice is a proximity voice. The voice recognition unit (309) can perform voice recognition by determining that the user intends to input voice only when it is determined that the voice input through the microphone (203) is a proximity voice. That is, even if the microphone (203) is activated and voice is input, voice recognition can be performed only when the input voice is a proximity voice in order to prevent misrecognition of a voice or noise not intended by the user.
[0097] In one embodiment, the speech recognition unit (309) may perform speech recognition using a known STT (speech-to-text) model. For example, the speech recognition unit (309) may perform speech recognition by preprocessing, encoding, and decoding a voice signal and then outputting text.
[0098] Meanwhile, in the previous embodiment, the case where the user device (200) is a smartphone was described, but various user devices other than a smartphone can be used.
[0099] FIG. 4 is a diagram exemplifying voice input on various user devices according to one embodiment of the present disclosure. Referring to FIG. 4 , the present disclosure can be applied to voice recognition when inputting voice on a smartwatch (400a) or a smart ring (400b) in addition to a smartphone, and can also be applied to various other user devices.
[0100] Figure 5 is a flowchart illustrating a voice recognition process according to one embodiment of the present disclosure. Each step of the voice recognition method according to this embodiment does not necessarily have to be performed in the order shown, nor does it imply that each step is essential. In other words, it should be understood that each step of the voice recognition method according to this embodiment may be performed in a different order than shown, and some steps may be omitted or other steps may be added.
[0101] Referring to FIG. 5, in step S501, the voice recognition device (300) may receive a sensor signal. In one embodiment, the voice recognition device (300) may receive a sensor signal from an IMU sensor of the user device (200). The IMU sensor signal may include movement information including information about height changes, acceleration, and deceleration of the user device (200). In addition, the voice recognition device (300) may receive a proximity sensor signal from a proximity sensor of the user device (200). The proximity sensor signal may include proximity information or distance information between the user device (200) and an object (e.g., the user's mouth or lips).
[0102] In step (S503), the voice recognition device (300) may activate the microphone (203) of the user device (200) based on the sensor signal. In one embodiment, the voice recognition device (300) may activate the microphone (203) of the user device (200) when it is determined that the user device (200) is being lifted based on the IMU sensor signal. In addition, the voice recognition device (300) may activate the microphone (203) of the user device (200) when it is determined that the user device (200) is rapidly decelerating and stopping based on the IMU sensor signal. At this time, the voice recognition device (300) may also consider whether the user device (200) is approaching an object (e.g., the user's mouth or lips) based on the proximity sensor signal, and may activate the microphone (203) when the user device (200) is approaching the object while rapidly decelerating.
[0103] In this way, according to one embodiment of the present disclosure, a state in which voice recognition is possible is established by an action of lifting the user device (200) or an action of bringing the user device (200) to the mouth, thereby enabling voice input immediately without a separate trigger for voice recognition.
[0104] In addition, since the microphone (203) is activated in all cases where the user's voice input is expected, errors in which the user's voice is not input and thus not recognized can be prevented in advance. In addition, since the microphone (203) is activated only when the user device (200) is lifted or suddenly decelerates to a stop, it is possible to prevent the input of ambient voices, distant voices, and other noises that the user did not clearly intend for voice recognition. Furthermore, the selective activation of the microphone (203) enables efficient use of power through a low-power design.
[0105] In step (S505), the voice recognition device (300) can receive a voice signal input through the microphone (203) of the user device (200). In one embodiment, the voice recognition device (300) can also receive information about the volume or sound pressure of the voice through the microphone (203).
[0106] In step (S507), the voice recognition device (300) can analyze the received voice signal to determine whether the voice input into the microphone (203) of the user device (200) is a proximity voice.
[0107] In one embodiment, the voice recognition device (300) can classify the properties of an input voice and determine whether it is a proximity voice based on the classified properties. For example, the voice recognition device (300) can classify the properties of the input voice into one of a normal voice, a whisper, and a shout through analysis of a voice signal (e.g., frequency band analysis, etc.), and estimate the speech distance by referring to the correlation between the speech distance and the sound pressure according to the voice properties, thereby determining whether it is a proximity voice. In another example, if the sound pressure of the voice is higher than a preset threshold sound pressure, the voice can be determined as a proximity voice. In this case, the threshold sound pressure can be preset according to the properties of the voice, and whether it is a proximity voice can be determined based on the threshold sound pressure according to the properties of the voice.
[0108] In one embodiment, the voice recognition device (300) can determine whether a voice is a proximity voice by analyzing characteristic phenomena associated with proximity speech. For example, the voice signal can be analyzed to determine whether a voice is a proximity voice based on whether characteristic information associated with proximity speech, such as popping sounds, breathing sounds, body contact sounds, proximity effects, and sound absorption phenomena, is detected.
[0109] In one embodiment, the voice recognition device (300) can determine whether an input voice is a proximity voice based on whether it is the voice of a registered user. For example, if the input voice is determined through analysis of the voice signal not to be the voice of a registered user, the voice can be determined to be noise and not a target for voice recognition.
[0110] According to one embodiment of the present disclosure, when the user device (200) is lifted or decelerated suddenly to a stop, i.e., when there is a possibility that the user may input voice, the microphone (203) is always activated to input voice. However, in this case, even when the user device (200) is lifted or decelerated suddenly to a stop without intending to input voice, the microphone (203) may be activated, resulting in unintended voice input.
[0111] In one embodiment of the present disclosure, even if voice input is performed while the microphone (203) is activated, voice recognition is performed only when the input voice is determined to be a proximity voice as described above, thereby preventing misrecognition due to unintended voice input.
[0112] In step (S509), the voice recognition device (300) can perform voice recognition if it is determined that the input voice is a proximity voice.
[0113] Additionally, according to one embodiment of the present disclosure, the voice recognition device (300) may deactivate the microphone (203) of the user device (200) based on a sensor signal. In one embodiment, the voice recognition device (300) may deactivate the microphone (203) of the user device (200) when it is determined that the user device (200) is moving away from an object (e.g., the user's mouth or lips) while lowering or rapidly accelerating based on an IMU sensor signal. In another embodiment, the voice recognition device (300) may deactivate the microphone (203) of the user device (200) when it is determined that the user device (200) is moving away from an object (e.g., the user's mouth or lips) based on a proximity sensor signal. Alternatively, the voice recognition device (300) may deactivate the microphone (203) of the user device (200) even when a preset command for terminating voice recognition is input.
[0114] In this way, a simple action of the user device (200) can deactivate the microphone (203) and automatically terminate voice input. Furthermore, even if the user pauses speech while inputting voice, creating a gap, voice input is not terminated contrary to the user's intention, thereby preventing unintended non-recognition problems.
[0115] According to the voice recognition method according to one embodiment of the present disclosure described above, voice recognition can be performed without non-recognition or misrecognition of voices by recognizing a motion of lifting a user device or bringing it close to the mouth without a separate trigger for voice recognition and determining whether it is a proximity voice. Specifically, by broadly recognizing the motion of lifting the user device or bringing it close to the mouth to activate the microphone and switch to a state where voice input is possible, the non-recognition of the user's voice can be prevented in advance. In addition, by targeting only proximity voices among voices input while the microphone is activated for voice recognition, the problem of misrecognition in which a voice that the user did not intend is recognized can be prevented.
[0116] According to a voice recognition method according to one embodiment of the present disclosure, not only the operating system of a user device but also third-party apps can be easily operated. Conventionally, to operate the operating system of a user device, a call command had to be entered or a button on the user device had to be pressed followed by voice input. To operate a third-party app, the user device had to be unlocked, the app had to be launched, a voice recognition icon had to be selected, and then voice input had to be performed. However, according to one embodiment of the present disclosure, the operating system of the user device and third-party apps can be operated simply by lifting the user device and inputting a voice in proximity.
[0117] The embodiments of the present disclosure described above may be implemented in the form of program instructions that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either singly or in combination. The program instructions recorded on the computer-readable recording medium may be specially designed and configured for the present disclosure or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. Hardware devices may be modified into one or more software modules to perform processing according to the present disclosure, and vice versa.
[0118] Although the present disclosure has been described above with specific details such as specific components and limited examples, the above examples are provided only to help a more general understanding of the present disclosure, and the present disclosure is not limited thereto, and those skilled in the art to which the present disclosure pertains may attempt various modifications and variations from this description. For example, some or all of the functions of the voice recognition device (300) may be distributed to an external device such as a user device (200) or a server. In addition, appropriate results may be achieved even if the described techniques are performed in a different order from the described method, or the described components are combined in a different form from the described method, or are replaced or substituted by other components or equivalents.
[0119] Therefore, the spirit of the present disclosure should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the claims described below, as well as the claims, are considered to fall within the scope of the spirit of the present disclosure.
Claims
1. As a voice recognition method, A step of receiving a sensor signal from an IMU sensor of a user device, A step of activating the microphone of the user device when it is determined that the user device is lifted or decelerates rapidly to stop based on the sensor signal; A step of receiving a voice signal input through the above microphone, A step of analyzing the above voice signal to determine whether the voice input into the microphone is a proximity voice; and A step for performing voice recognition when the voice input to the above microphone is a proximity voice. A speech recognition method comprising:
2. In paragraph 1, A voice recognition method in which, in the step of determining whether the above-mentioned proximity voice is present, it is determined whether the voice is present by referring to the properties of the voice input through the microphone.
3. In paragraph 2, A voice recognition method in which, in the step of determining whether the above-mentioned proximity voice is a proximity voice, the speech distance is estimated from the properties and sound pressure of the voice input through the microphone by referring to the correlation between the speech distance and sound pressure according to the properties of the voice.
4. In paragraph 1, A voice recognition method in which, in the step of determining whether the voice is a proximity voice, the voice input through the microphone is determined to be a proximity voice if the sound pressure of the voice input through the microphone is higher than a threshold sound pressure.
5. In paragraph 1, In the step of determining whether the above-mentioned proximity voice is present, feature information according to close-range speech is detected from the voice input through the microphone to determine whether it is the proximity voice. A speech recognition method, wherein the characteristic information according to the above-mentioned close-range speech includes at least one of a pop sound, a breath sound, a body contact sound, a proximity effect, and an absorption phenomenon.
6. In paragraph 1, In the step of receiving the above voice signal, a plurality of voice signals input through a plurality of microphones of the user device are received, A voice recognition method, wherein in the step of determining whether the above-mentioned voice is a close-up voice, the plurality of voice signals are analyzed to evaluate the similarity between the plurality of voice signals, and whether the voice is a close-up voice is determined based on the similarity.
7. In paragraph 1, A voice recognition method in which, in the step of determining whether the voice is a proximity voice, it is determined whether the voice input through the microphone is a voice of a pre-registered user, and if it is not a voice of the pre-registered user, it is determined that it is not a proximity voice.
8. In paragraph 1, In the step of receiving the above sensor signal, a proximity sensor signal is further received from the proximity sensor of the user device, A voice recognition method in which, in the step of activating the microphone, the microphone of the user device is activated when it is determined that the user device has rapidly decelerated and stopped and is simultaneously approaching an object, based on a sensor signal received from an IMU sensor of the user device and a proximity sensor signal received from a proximity sensor of the user device.
9. In paragraph 1, Further comprising the step of disabling the microphone of the user device; A voice recognition method in which, in the step of disabling the microphone of the user device, a sensor signal received from an IMU sensor of the user device is analyzed to determine that the user device is moving downwards or rapidly accelerating, and the microphone of the user device is disabled.
10. In paragraph 9, In the step of receiving the above sensor signal, a proximity sensor signal is further received from the proximity sensor of the user device, A voice recognition method in which, in the step of disabling the microphone of the user device, the microphone of the user device is disabled when it is determined that the user device is moving rapidly and moving away from an object at the same time, based on a sensor signal received from an IMU sensor of the user device and a proximity sensor signal received from a proximity sensor of the user device.
11. A computer-readable recording medium for executing the method according to paragraph 1.
12. As a voice recognition device, A sensor signal receiving unit that receives a sensor signal from an IMU sensor of a user device; A microphone control unit that controls activation and deactivation of the microphone of the user device based on the sensor signal; A voice signal receiving unit that receives a voice signal input through the above microphone, A proximity voice determination unit that analyzes the above voice signal to determine whether the voice input to the microphone is a proximity voice; and A voice recognition unit that performs voice recognition when the voice input to the above microphone is a proximity voice. A voice recognition device comprising:
Citation Information
Patent Citations
Mobile terminal and control method for mobile terminal
KR101981316B1
Remote control apparatus for inputting user voice and method thereof
KR1020150040445A
Cloud edge computing policy server, and control method thereof
KR1020220033251A
Method and system for detecting surface damage on conveyor belts
KR102698495B1
KR20240067114A