Method, apparatus, and recording medium for near-field voice input
Patent Information
- Application Number
- PCT/KR2024/017883
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2024-11-12
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional voice recognition systems require users to press a button or speak a trigger word to initiate voice input, which is inconvenient and prone to misrecognition due to ambient noise, especially when the user and device are slightly separated.
A proximity voice input method that determines whether a voice is close-range by analyzing the voice's properties, such as sound pressure and frequency spectrum, and uses a proximity sensor to confirm the distance, eliminating the need for button presses or trigger words.
Enables quick and accurate voice input by recognizing only close-range voices, minimizing ambient noise interference and preventing misrecognition.
Smart Images

Figure KR2024017883_02102025_PF_FP_ABST
Abstract
Description
Proximity voice input method, device and recording medium
[0001] The present disclosure relates to a proximity voice input method, device, and recording medium, and more particularly, to a method, device, and recording medium for inputting proximity voice by determining proximity voice input.
[0002] With the recent rise in interest in user interfaces and advancements in voice processing technology, the number of IT devices equipped with voice recognition capabilities is increasing. For example, smartphones, smartwatches, smart TVs, and smart refrigerators, which can recognize a user's voice and perform requested actions, are becoming widespread.
[0003] However, according to conventional voice recognition techniques, the user had to specify the point in time when voice input began by pressing a button or entering a predetermined trigger word before starting voice input. The former method of pressing a button inevitably caused inconvenience because the user could not perform voice input if they did not have free use of their hands. Furthermore, the latter method of speaking a predetermined trigger word made it difficult to specify the point in time when voice input was initiated due to various noises, such as the voices of others occurring in the same space, even if the voice recognition device and the user were only slightly separated. In addition, even if the user spoke the predetermined trigger word, to assure the user that voice input had begun, feedback using sound or light had to be provided before the user could begin voice input. This limitation inevitably led to a considerable amount of time being required from the very beginning of voice input.
[0004] Accordingly, the inventor of the present invention proposes a technology for a proximity voice input method that determines whether a voice detected by a device is a proximity voice resulting from a close-range speech, without pressing a button or inputting a trigger word, and determines whether the input voice is a target of voice recognition based on this.
[0005] The present disclosure is intended to solve the problems of the above-described prior art, and its purpose is to enable a user to input voice quickly by omitting unnecessary steps for starting voice input.
[0006] In addition, the present disclosure aims to improve the accuracy of voice recognition by minimizing the influence of ambient noise during voice input and preventing misrecognition.
[0007] In addition, the purpose of the present disclosure is to improve the accuracy of voice recognition by determining whether a voice is a close-range voice based on the properties of the voice, such as a normal voice or a whisper, during the voice input process.
[0008] A method for determining input of a proximity voice for voice recognition according to one embodiment of the present disclosure includes the steps of detecting a voice from a microphone of a proximity voice input device to obtain a voice signal and a sound pressure of the voice, the step of analyzing a frequency spectrum of the voice signal to classify an attribute of the voice, the step of calculating a first distance, which is an utterance distance of the voice, based on the attribute of the voice and the sound pressure of the voice, and the step of determining whether the voice is a proximity voice due to a close-range utterance based on the first distance.
[0009] According to one embodiment of the present disclosure, in the step of classifying the properties of a voice, the voice can be classified into one of a whisper and a normal voice.
[0010] According to one embodiment of the present disclosure, in the step of determining whether a voice is a proximity voice, if the sound pressure of the voice is greater than a threshold sound pressure preset according to the properties of the voice, the voice may be determined to be a proximity voice.
[0011] According to one embodiment of the present disclosure, if a voice is classified as a whisper and the sound pressure of the voice is greater than a first threshold sound pressure, it may be determined as a proximity voice, and if a voice is classified as a normal voice and the sound pressure of the voice is greater than a second threshold sound pressure that is greater than the first threshold sound pressure, it may be determined as a proximity voice.
[0012] According to one embodiment of the present disclosure, in the step of calculating the first distance, the first distance can be calculated by referring to the correlation between the speech distance and sound pressure according to the properties of the voice.
[0013] According to one embodiment of the present disclosure, the correlation between the utterance distance and the sound pressure may be reflected by a weight according to the spectrum of the voice signal.
[0014] A proximity voice input method according to one embodiment of the present disclosure further includes, before the step of determining whether or not it is a proximity voice, a step of calculating a second distance, which is a distance between a proximity voice input device and an object detected by a proximity sensor of the proximity voice input device, and in the step of determining whether or not it is a proximity voice, it is possible to determine whether or not it is a proximity voice based on a similarity between the first distance and the second distance.
[0015] A proximity voice input method according to one embodiment of the present disclosure further includes, before the step of determining whether it is a proximity voice, a step of determining the similarity between a first distance and a second distance, and in the step of determining whether it is a proximity voice, it is possible to determine whether it is a proximity voice based on the similarity between the first distance and the second distance.
[0016] A proximity voice input device for voice recognition according to one embodiment of the present disclosure includes a microphone and a processor that detects voice and acquires a voice signal. Here, the processor analyzes the frequency spectrum of the voice signal to classify the voice properties, calculates a first distance, which is the speech distance of the voice signal, based on the voice properties and the sound pressure of the voice, and determines whether the voice is a proximity voice resulting from close-range speech based on the first distance.
[0017] In addition, other methods for implementing the present disclosure, other devices, and recording media for recording a computer program for executing the method are further provided.
[0018] According to one embodiment of the present disclosure, an effect is achieved whereby a user can quickly input voice by omitting unnecessary steps for starting voice input.
[0019] In addition, according to one embodiment of the present disclosure, only voices input in proximity to the device are recognized as input for voice recognition, thereby minimizing the influence of ambient noise and preventing misrecognition, thereby improving the accuracy of voice recognition.
[0020] In addition, according to one embodiment of the present disclosure, the accuracy of voice recognition can be improved by estimating the speech distance of a voice signal based on the properties of the voice and the sound pressure of the voice, and determining whether it is a close voice based on the distance.
[0021] FIG. 1 is a diagram schematically illustrating an overall system environment for proximity voice input according to one embodiment of the present disclosure.
[0022] FIG. 2 is a drawing exemplarily showing an example of inputting a proximity voice into a proximity voice input device according to one embodiment of the present disclosure.
[0023] FIG. 3 is a diagram showing the configuration of a proximity voice input device according to one embodiment of the present disclosure.
[0024] FIG. 4 is a functional block diagram schematically illustrating the functional configuration of a proximity voice input device according to one embodiment of the present disclosure.
[0025] FIG. 5 is a drawing showing a proximity voice input device according to another embodiment of the present disclosure.
[0026] FIG. 6 is a flowchart illustrating a method for determining input of a proximity voice according to one embodiment of the present disclosure.
[0027] [Explanation of symbols]
[0028] 110: Proximity voice input device
[0029] 120: Communications network
[0030] 130: User terminal
[0031] 140: Server
[0032] 150: External device
[0033] 301: Mike
[0034] 303: Sensor section
[0035] 305: Communications Department
[0036] 307: Processor
[0037] 401: Voice Attribute Classification
[0038] 403: Distance Calculation Unit
[0039] 405: Proximity voice discrimination unit
[0040] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. Hereinafter, specific descriptions of previously known functions and configurations will be omitted if deemed likely to unnecessarily obscure the gist of the present disclosure. Furthermore, it should be noted that the following description relates only to one embodiment of the present disclosure and that the present disclosure is not limited thereto.
[0041] The terminology used in this disclosure is merely used to describe specific embodiments and is not intended to be limiting of the present disclosure. For example, a singular element should be understood to include plural elements unless the context clearly dictates otherwise. The term "and / or" used in this disclosure should be understood to encompass any and all possible combinations of one or more of the listed items. The terms "comprises" or "has" as used in this disclosure are intended to specify that a feature, number, step, motion, component, part, or combination thereof described in this disclosure is present, but the use of such terms does not exclude the presence or addition of one or more other features, numbers, steps, motions, components, parts, or combinations thereof.
[0042] In the embodiments of the present disclosure, a "module" or "part" refers to a functional part that performs at least one function or motion, and may be implemented by hardware or software, or a combination of hardware and software. Furthermore, a plurality of "modules" or "parts" may be integrated into at least one software module and implemented by at least one processor, excluding "modules" or "parts" that need to be implemented by specific hardware.
[0043] Additionally, unless otherwise defined, all terms used in this disclosure, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which this disclosure pertains. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with the contextual meaning of the relevant technology, and should not be interpreted in an unduly limiting or expansive manner unless explicitly defined otherwise in this disclosure.
[0044] FIG. 1 is a diagram schematically illustrating an overall system environment for proximity voice input according to one embodiment of the present disclosure.
[0045] Referring to FIG. 1, the overall system environment (100) for proximity voice input may include a proximity voice input device (110), a communication network (120), a user terminal (130), and a server (140).
[0046] A proximity voice input device (110) according to one embodiment of the present disclosure may be an electronic device equipped with voice input and communication functions. The proximity voice input device (110) may be a portable device that a user can carry or a wearable device that a user can wear. In one embodiment, the proximity voice input device (110) may be a ring-shaped smart ring.
[0047] According to one embodiment of the present disclosure, a proximity voice input device (110) can communicate with a user terminal (130) via a communication network (120). In one embodiment, the proximity voice input device (110) can receive a user input such as a button press or voice, and a signal regarding the user input received by the proximity voice input device (110) can be transmitted to the user terminal (130) via the communication network (120).
[0048] A communication network (120) according to one embodiment of the present disclosure may include any wired or wireless communication network, such as a TCP / IP communication network. According to one embodiment of the present disclosure, the communication network (120) may be configured as a Local Area Network (LAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), the Internet, etc., but the present disclosure is not limited thereto. For example, the communication network (120) may be a wireless data communication network that implements, at least in part, a conventional communication method such as WiFi communication, WiFi-Direct communication, Long Term Evolution (LTE) communication, 5G communication, Bluetooth communication (including Bluetooth Low Energy (BLE) communication), infrared communication, ultrasonic communication, etc. As another example, the communication network (120) may be an optical communication network that implements, at least in part, a conventional communication method such as LiFi (Light Fidelity).
[0049] A user terminal (130) according to one embodiment of the present disclosure is a device capable of communicating with a proximity voice input device (110) via a communication network (120). The user terminal (130) is a digital device equipped with a memory means and a microprocessor to provide computing capabilities, and may be, for example, a smartphone. The user terminal (130) of the present disclosure is not limited to that illustrated, and the user terminal (130) may be a smart pad, a notebook computer, a personal digital assistant (PDA), a desktop computer, or any other device capable of achieving the purpose of the present disclosure may be used as the user terminal (130) of the present disclosure.
[0050] According to one embodiment of the present disclosure, a user terminal (130) can receive a signal regarding a user input from a proximity voice input device (110) via a communication network (120) and perform a function according to the user input. For example, when a proximity voice is detected by the proximity voice input device (110), the corresponding voice signal is transmitted to the user terminal (130) via the communication network (120), and the user terminal (130) can generate a response signal or a control signal based on the voice recognized from the voice signal and the user's intention.
[0051] A server (140) according to one embodiment of the present disclosure is a server system capable of communicating with a user terminal (130) via a communication network (120). The server (140) may be an independent physical server or may be a virtual server such as a cloud server.
[0052] According to one embodiment of the present disclosure, the server (140) is an artificial intelligence server that receives a voice signal from a user terminal (130), performs voice recognition on the voice signal, and then generates feedback. In one embodiment, the server (140) provides the generated feedback to the user terminal (130), and the user terminal (130) can generate a response signal and / or a control signal by referring to the feedback.
[0053] According to one embodiment of the present disclosure, a user terminal (130) can generate a response signal based on a recognized voice and transmit it to an external device (150). In one embodiment, the external device (150) to which the response signal is transmitted may be an earphone, a headset, or a speaker. For example, the user terminal (130) can transmit the response signal to an earphone worn by the user and respond to the voice input by the user through the earphone.
[0054] According to one embodiment of the present disclosure, a user terminal (130) can generate a control signal based on recognized voice and transmit the control signal to an external device (150). In one embodiment, the external device (150) to which the control signal is transmitted may be a smart TV, an air conditioner, a refrigerator, or other home appliance, a smart car, or the like. For example, the user terminal (130) can generate a control signal based on a user's intention or request identified from a voice signal and transmit the control signal to a smart TV, and the smart TV can be controlled (power on / off, channel or volume change, etc.) based on the control signal.
[0055] Meanwhile, the external device (150) to which the response signal and / or control signal is transmitted is not limited to the examples above, and should be understood to include any electronic device capable of providing an auditory or visual response to the user or any electronic device having a wireless communication function and capable of being controlled according to the user's intention.
[0056] In the illustrated embodiment, the proximity voice input device (110) communicates with the user terminal (130), and the user terminal (130) communicates with the server (140). However, this may be implemented differently. For example, the proximity voice input device (110) may communicate directly with the server (140). In this case, the proximity voice input device (110) may transmit a voice signal detected through the communication network (120) to the server (140), and the server (140) may recognize voice from the voice signal and then generate and transmit a response signal and / or a control signal. As another example, the user terminal (130) may function as an on-device AI, and may directly recognize voice from the voice signal and generate a response signal and / or a control signal without communicating with the server (140). As another example, the proximity voice input device (110) can communicate directly with another IoT device, such as an AI speaker, rather than a user terminal (130) or a server (140).
[0057] In addition, the proximity voice input device (110) is configured to perform operations such as voice recognition, generation of response signals and / or control signals, etc., so that the operations can be processed in the proximity voice input device (110) instead of being performed in an external device such as a user terminal (130) or a server (140).
[0058] FIG. 2 is a drawing exemplarily showing an example of inputting a proximity voice into a proximity voice input device according to one embodiment of the present disclosure.
[0059] The proximity voice input device (110) according to one embodiment of the present disclosure can determine that a proximity voice input is intended by the user if it is a proximity voice generated by close-range speech. Accordingly, as illustrated in FIG. 2, in order to input a voice into the proximity voice input device (110) (e.g., a smart ring) according to one embodiment of the present disclosure, the proximity voice input device (110) is first brought close to the mouth (or lips).
[0060] In this way, when speaking with the mouth close to the proximity voice input device (110), characteristic phenomena due to speaking at a close distance may occur. For example, a proximity effect may occur, in which the sensitivity to low-frequency bands increases as the distance between the sound source and the microphone gets closer. In addition, a phenomenon in which breathing sounds such as breath or sniffing are mixed in with the voice may occur. In addition, in the process of placing the proximity voice input device (110) near the lips, sounds due to contact between body parts (e.g., the sound of a hand or finger touching the face) may be generated, and as the proximity voice input device (110) gets closer to the facial skin, sounds may be absorbed by the skin, reducing background noise.
[0061] A proximity voice input device (110) according to one embodiment of the present disclosure is configured to perform a function of determining whether a voice signal is a proximity voice intended for voice recognition based on this phenomenon.
[0062] FIG. 3 is a diagram showing the configuration of a proximity voice input device according to one embodiment of the present disclosure.
[0063] Referring to FIG. 3, a proximity voice input device (110) according to one embodiment of the present disclosure may include a microphone (301), a sensor unit (303), a communication unit (305), and a processor (307).
[0064] The microphone (301) of the proximity voice input device (110) according to one embodiment of the present disclosure is for voice input and can perform a function of detecting a user's voice and generating a voice signal. The microphone (301) can be installed on one side of the proximity voice input device (110) and configured to detect an external voice. In one embodiment, the microphone (301) can perform a function of detecting the sound pressure of an input voice.
[0065] According to one embodiment of the present disclosure, the microphone (301) may be activated when a predetermined condition is met to minimize noise input. In one embodiment, the microphone (301) may be activated when the proximity voice input device (110) is ready to input proximity voice. For example, the microphone (301) may be activated when the proximity voice input device (110) is determined to be in proximity to an object (e.g., a user's mouth).
[0066] The sensor unit (303) of the proximity voice input device (110) according to one embodiment of the present disclosure is for detecting whether the proximity voice input device (110) is in proximity and / or moving, and can perform a function of generating a sensing signal regarding proximity information and / or movement information.
[0067] According to one embodiment of the present disclosure, the sensor unit (303) may include a proximity sensor. The proximity sensor may detect a distance between a proximity voice input device (110) and an object, and the sensor unit (303) may generate a sensing signal based on proximity information of the proximity voice input device (110) detected by the proximity sensor. The proximity sensor may be at least one of an optical sensor, a photoelectric sensor, an ultrasonic sensor, an inductive sensor, a capacitive sensor, a resistive sensor, an eddy current sensor, an infrared sensor, and a magnetic sensor.
[0068] According to one embodiment of the present disclosure, the sensor unit (303) may include an IMU sensor. The IMU sensor may detect movement (e.g., change in height) of the proximity voice input device (110), and the sensor unit (303) may generate a sensing signal regarding movement information of the proximity voice input device (110) detected by the IMU sensor.
[0069] According to one embodiment of the present disclosure, the sound pressure of a voice input from a microphone (301) can be detected, but alternatively, the sensor unit (303) may include a sound pressure sensor. For example, the sensor unit (303) may include at least one of a pressure sensor, a wind pressure sensor, and a noise sensor as a sound pressure sensor for detecting sound pressure.
[0070] The processor (307) of the proximity voice input device (110) according to one embodiment of the present disclosure may perform a function of receiving and processing at least one of a voice signal related to a voice input from a microphone (301) and a sensing signal from a sensor unit (303). In one embodiment, the voice signal from the microphone (301) and the sensing signal from the sensor unit (303) may be transmitted to the processor (307) for processing and transmitted to a user terminal (130) or a server (140) via a communication unit (305).
[0071] According to one embodiment of the present disclosure, the processor (307) can determine whether a voice signal from the microphone (301) is a proximity voice based on a sensing signal. Specifically, the processor (307) can determine whether a voice input from the microphone (301) is a proximity voice based on at least one of a spectrum of the voice signal, a sound pressure of the voice, proximity information of the proximity voice input device (110), and movement information. Here, the proximity voice refers to a voice that a user intentionally inputs into the proximity voice input device (110), and a voice input through the microphone (301) that is determined not to be a proximity voice can be classified as noise.
[0072] Meanwhile, although the processor (307) is exemplified as being provided in the proximity voice input device (110) in the present embodiment, it should be understood that at least some or all of the functions of the processor (307) may be performed in an external electronic device, for example, a user terminal (130) or a server (140). Conversely, it should be understood that at least some or all of the functions of the user terminal (130) or the server (140) may be performed in the processor (307). For example, the processor (307) may be configured to perform at least some or all of the functions of voice recognition, generation and processing of a response signal and / or a control signal according to a voice input mode, etc.
[0073] The communication unit (305) of the proximity voice input device (110) according to one embodiment of the present disclosure may function to enable the proximity voice input device (110) to communicate with an external communication network. For example, the communication unit (305) may perform a function of transmitting a button input signal or a signal regarding a voice input mode determined by the button input signal and a voice signal to a user terminal (130) or a server (140).
[0074] A proximity voice input device (110) according to one embodiment of the present disclosure may further include a storage unit (not shown). According to one embodiment of the present disclosure, the storage unit may store device information, user information, etc. of the proximity voice input device (110), and may also store voices input to the proximity voice input device (110).
[0075] FIG. 4 is a functional block diagram schematically illustrating the functional configuration of a proximity voice input device according to one embodiment of the present disclosure. Specifically, FIG. 4 is a functional block diagram illustrating the functional configuration of a processor of the proximity voice input device.
[0076] Referring to FIG. 4, the processor (307) of the proximity voice input device (110) according to one embodiment of the present disclosure may include a voice attribute classification unit (401), a distance calculation unit (403), and a proximity voice determination unit (405).
[0077] Meanwhile, the components illustrated in FIG. 4 do not reflect all functions of the processor (307) of the proximity voice input device (110), nor are they essential, so the processor (307) of the proximity voice input device (110) may include more or fewer components than the illustrated components.
[0078] According to one embodiment of the present disclosure, at least some of the voice attribute classification unit (401), the distance calculation unit (403), and the proximity voice determination unit (405) may be program modules that communicate with an external system (not shown) via the communication unit (305). These program modules may be included in the proximity voice input device (110) in the form of an operating system, an application program module, or other program modules, and may be physically stored in various known memory devices. In addition, these program modules may be stored in a remote memory device that can communicate with the proximity voice input device (110). Meanwhile, these program modules include, but are not limited to, routines, subroutines, programs, objects, components, data structures, etc. that perform specific tasks or execute specific abstract data types, which will be described later according to the present disclosure.
[0079] A voice attribute classification unit (401) according to one embodiment of the present disclosure can perform a function of analyzing voice attributes by analyzing the frequency spectrum of a voice signal.
[0080] Voice is composed of a combination of vowels and consonants, and the resulting sounds encompass a wide frequency spectrum. These vocal properties can vary depending on how the voice is produced.
[0081] For example, a human whisper, or whispering, is produced without vocal cord vibration, and therefore differs from normal speech in terms of frequency and timbre. Specifically, in whispering, vowels are primarily produced by airflow, while consonants are pronounced without vocal cord vibration. This results in a whispered voice spectrum with relatively less emphasis on the high frequencies than normal speech, which contributes to a whispered voice having a lower, softer tone. On the other hand, normal speech is produced using vocal cords. Since normal speech produces sound by vocal cord vibration, if the vocal cords try to speak too softly, the vocal cords will not vibrate properly, making it difficult to produce a clear sound. Consequently, normal speech cannot be lowered below a certain volume.
[0082] A voice attribute classification unit (401) according to one embodiment of the present disclosure can classify voice attributes based on these differences between whispering and normal voice.
[0083] In one embodiment, the voice attribute classification unit (401) may classify the voice as a whisper if the ratio of the volume of a high frequency band (e.g., a frequency band of 1 kHz or higher) to the volume of a low frequency band (e.g., a frequency band of 500 Hz or lower) in the acquired voice signal is lower than a reference value. Conversely, if the ratio of the volume of a high frequency band to the volume of a low frequency band exceeds a predetermined reference value, the voice may be classified as a normal voice. Here, the reference value may be a value set in advance. For example, the reference value may be set in the range of 0.1 to 0.5.
[0084] In another embodiment, the voice attribute classification unit (401) may classify voice attributes using an artificial intelligence model based on machine learning. For example, the voice attribute classification unit (401) may perform machine learning based on the spectra of multiple accumulated input voice signals, and classify voice attributes using a voice attribute classification model generated based on the machine learning.
[0085] A distance calculation unit (403) according to one embodiment of the present disclosure may perform a function of calculating the speech distance of a voice signal based on voice properties and sound pressure. In one embodiment, the distance calculation unit (403) may calculate the speech distance of a voice by referring to the correlation between the speech distance and sound pressure according to the voice properties. In this specification, the speech distance of a voice calculated based on the voice properties and sound pressure of the voice is defined as a first distance.
[0086] Table 1 exemplarily shows a reference table regarding the correlation between the speech distance and sound pressure according to the properties of the voice according to one embodiment of the present disclosure.
[0087]
[0088]
[0089] In Table 1, weighting refers to differentiating sensitivity by frequency in acquiring sound pressure. A-weighting (dBA) is a weighting that reflects the sensitivity of the human ear, and it gives a higher weight to the middle frequency components in the range of 500 Hz to 8 kHz, where the human ear is most sensitive, thereby minimizing the influence of low and high frequencies on the resulting measurement values. On the other hand, C-weighting (dBC) is a weighting that reacts sensitively to low and high frequencies, and Z-weighting (dBZ) is a weighting that reacts equally to all frequencies without being biased toward a specific frequency.
[0090] According to one embodiment of the present disclosure, when the voice is classified as a whisper in the voice attribute classification unit (401) and the sound pressure of the voice is 62 dBA (i.e., the sound pressure with A-weighting applied is 62 dB), the distance calculation unit (403) can calculate the first distance, which is the speech distance, as 5 cm by referring to the correlation between the speech distance and the sound pressure regarding the voice attribute in the above-mentioned reference table. In the same manner, when the voice is classified as a normal voice in the voice attribute classification unit (401) and the sound pressure of the voice is 62 dBA, the distance calculation unit (403) can calculate the first distance, which is the speech distance, as 10 cm by referring to the correlation between the speech distance and the sound pressure regarding the voice attribute in the above-mentioned reference table.
[0091] A distance calculation unit (403) according to one embodiment of the present disclosure may perform a function of calculating a distance between a proximity voice input device (110) and an object based on data acquired from a proximity sensor provided in a sensor unit (303) of the proximity voice input device (110). In this specification, the distance between a proximity voice input device (110) and an object acquired through the proximity sensor is defined as a second distance in order to distinguish it from a first distance, which is an utterance distance.
[0092] According to one embodiment of the present disclosure, the second distance calculated through the distance calculation unit (403), i.e., the distance between the actual proximity voice input device (110) and the object, can be used to determine whether the input voice is the user's input voice.
[0093] A proximity voice determination unit (405) according to one embodiment of the present disclosure can perform a function of determining whether a voice is a proximity voice resulting from a close-range speech, based on a first distance calculated based on the voice's properties and the voice's sound pressure.
[0094] In one embodiment, the proximity voice determination unit (405) may determine a voice as a proximity voice if the calculated first distance is less than or equal to a preset distance. For example, the proximity voice determination unit (405) may determine a voice as a proximity voice if the calculated first distance is less than or equal to 5 cm.
[0095] In another embodiment, the proximity voice determination unit (405) may determine a voice as a proximity voice if the sound pressure of the detected voice is higher than a preset threshold sound pressure. For example, the proximity voice determination unit (405) may set the sound pressure when the utterance distance is 5 cm based on the A-weight in the above-mentioned reference table as the threshold sound pressure, and determine the voice as a proximity voice if the sound pressure of the detected voice is higher than the threshold sound pressure.
[0096] Referring to Table 1, when the distance at which a whisper occurs is set to 5 cm, the threshold sound pressure level can be set to 62 dBA. In other words, this means that a whispered voice within 5 cm must generate a sound pressure level of 62 dBA or higher, and if the sound pressure level is lower than 62 dBA, it can be assumed that the voice was spoken from a distance, that is, a voice that the user did not intend to recognize, or that it is caused by noise.
[0097] Similarly, for normal speech, when the distance at which proximity speech occurs is set to 5 cm, the threshold sound pressure level can be set to 68 dBA. In other words, for normal speech, a proximity speech within 5 cm must generate a sound pressure level of 68 dBA or higher, and a sound pressure level lower than 68 dBA can be assumed to be a speech that the user did not intend to recognize or is caused by noise.
[0098] According to one embodiment of the present disclosure, the proximity voice determination unit (405) may perform a function of determining whether a voice is a proximity voice resulting from a close-range speech, based on a first distance calculated based on the properties of the voice and the sound pressure of the voice, and a second distance calculated based on data acquired through a proximity sensor.
[0099] In one embodiment, the proximity voice determination unit (405) can compare the first distance and the second distance to determine the similarity, and determine whether it is a proximity voice based on the similarity.
[0100] For example, if the voice property is classified as a whisper and the sound pressure is 62 dBA, the first distance calculated by referring to Table 1 is 5 cm. At this time, if the second distance calculated based on the measurement from the proximity sensor is 4 cm, the proximity voice determination unit (507) can determine the voice signal as a proximity voice spoken by the user since the first distance and the second distance are similar.
[0101] For another example, if the voice property is classified as a whisper and the sound pressure is 48 dBA, the first distance calculated by referring to Table 1 is 25 cm. In this case, if the second distance calculated based on the measurement from the proximity sensor is 3 cm, the proximity voice determination unit (405) can determine that the voice signal is not a proximity voice because the first distance and the second distance are not similar. In this case, if the first distance is greater than the second distance, the voice can be determined to be one that the user did not intend to recognize or other noise.
[0102] According to one embodiment of the present disclosure, the proximity voice determination unit (405) may refer to change information of the distance and / or position of the proximity voice determination unit (405) when determining whether or not a proximity voice is present. The proximity voice determination unit (405) may calculate a change in the distance between the proximity voice input device (110) and an object (e.g., the user's mouth or lips) measured through the sensor unit (303), that is, the second distance, and even if the second distance satisfies the condition of the proximity voice (e.g., within 5 cm), only when the change information of the second distance satisfies a predetermined condition may it be determined that the second distance finally satisfies the proximity voice criterion.
[0103] In one embodiment, the proximity voice determination unit (405) may determine that a voice meets the criteria for a proximity voice when a voice is input after the proximity voice input device (110) gradually approaches an object and approaches within a preset distance (e.g., 5 cm). Alternatively, the proximity voice determination unit (405) may determine that a voice meets the criteria for a proximity voice when a voice is input after the proximity voice input device (110) approaches within a preset distance (e.g., 5 cm) from an object and the distance is maintained for a predetermined period of time or longer. Accordingly, the proximity voice determination unit (405) may determine whether a voice is a proximity voice by referring to whether a change in the distance of the proximity voice input device (110) satisfies a predetermined condition, along with the similarity between the first distance and the second distance.
[0104] In another embodiment, the proximity voice determination unit (405) may determine that a voice meets the criteria for a proximity voice when a voice is input after the proximity voice input device (110) moves upward and approaches an object within a preset distance (e.g., 5 cm). Alternatively, the proximity voice determination unit (405) may determine that a voice meets the criteria for a proximity voice when a voice is input after the proximity voice input device (110) moves upward and then decelerates and stops and is positioned within a preset distance (e.g., 5 cm) from an object. Accordingly, the proximity voice determination unit (405) may determine whether a voice is a proximity voice by referring to whether a change in the distance and position of the proximity voice input device (110) satisfies a predetermined condition, along with the similarity between the first distance and the second distance.
[0105] In the above, the proximity voice input device (110) is described as a ring-shaped smart ring, but the proximity voice input device may be implemented in other forms.
[0106] FIG. 5 is a diagram illustrating a proximity voice input device according to another embodiment of the present disclosure. Referring to FIG. 5, the proximity voice input device may have a shape other than a ring shape, such as a remote control (500a) equipped with a button and a microphone, a stylus pen (500b), a smart watch (500c), a smartphone (500d), etc. As illustrated, the proximity voice input devices (500a, 500b, 500c, 500d) are configured to enable voice input (proximity voice input) by having microphones (501a, 501b, 501c, 501d). In addition, the proximity voice input device according to the present disclosure may be implemented in various other forms than those illustrated as long as it has a microphone capable of voice input.
[0107] FIG. 6 is an exemplary flowchart illustrating a method for determining proximity voice input in a proximity voice input device according to one embodiment of the present disclosure. Each step of the proximity voice input method according to this embodiment does not necessarily have to be performed in the order illustrated, nor does it imply that each step is essential. That is, it should be understood that each step of the proximity voice input method according to this embodiment may be performed in a different order than illustrated, and some steps may be omitted or other steps may be added.
[0108] In step (S601), a voice can be detected from a microphone of a proximity voice input device to obtain a voice signal and sound pressure of the voice.
[0109] In step (S603), the proximity voice input device can analyze the frequency spectrum of the voice signal to classify the properties of the voice. Specifically, the frequency spectrum of the voice signal can be analyzed to classify the input voice as either a whisper or a normal voice. In one embodiment, if the ratio of the volume of a high frequency band (e.g., a frequency band of 1 kHz or more) to the volume of a low frequency band (e.g., a frequency band of 500 Hz or less) in the acquired voice signal is below a reference value, the voice can be classified as a whisper, and conversely, if the ratio of the volume of a high frequency band to the volume of a low frequency band exceeds a predetermined reference value, the voice can be classified as a normal voice.
[0110] In step (S605), the proximity voice input device can calculate a first distance, which is a speech distance, based on the voice properties and the sound pressure of the voice. In one embodiment, the proximity voice input device can calculate the first distance by referring to the correlation between the speech distance and the sound pressure according to the voice properties.
[0111] In step (S607), the proximity voice input device can determine whether the voice is a proximity voice resulting from a close-range speech based on the first distance.
[0112] In one embodiment, the proximity voice input device may determine the voice as a proximity voice if the first distance produced is less than or equal to a preset distance (e.g., 5 cm). In another embodiment, the proximity voice input device may determine the voice as a proximity voice if the sound pressure of the detected voice is greater than or equal to a preset threshold sound pressure. Here, the threshold sound pressure may be set by referring to the correlation between the speech distance and sound pressure according to the properties of the voice, and may be set differently depending on the properties of the voice. For example, the threshold sound pressure may be set to 62 dBA for a whisper and to 68 dBA for a normal voice.
[0113] In step (S607), the proximity voice input device may determine whether the voice is a proximity voice resulting from a close-range speech, based on the first distance and the second distance, which is the distance between the proximity voice input device and the object (e.g., the user's mouth or lips). To this end, in step (S605), the proximity voice input device may calculate the second distance, which is the distance between the proximity voice input device and the object, based on data acquired through the sensor unit.
[0114] In one embodiment, the proximity voice input device can determine the similarity between a first distance and a second distance, and determine whether the proximity voice is present based on the similarity between the first distance and the second distance.
[0115] Through the above-described device and method, it is possible to accurately determine a proximity voice intended for voice recognition regardless of the properties of the input voice (i.e., whisper or normal voice), and to process a voice signal and perform a response and / or control in accordance with the user's intention.
[0116] Although the present disclosure has been described above with specific details such as specific components and limited examples, the above examples are provided only to help a more general understanding of the present disclosure, and the present disclosure is not limited thereto, and those skilled in the art to which the present disclosure pertains may attempt various modifications and variations from this description. For example, some or all of the functions of the processor (307) of the proximity voice input device (110) may be distributed to the user terminal (130) or the server (140). In addition, appropriate results may be achieved even if the described techniques are performed in a different order from the described method, or the described components are combined in a different form from the described method, or are replaced or substituted by other components or equivalents.
[0117] Meanwhile, the embodiments according to the present disclosure described above may be implemented in the form of program commands that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the computer-readable recording medium may be specially designed and configured for the present disclosure or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. Hardware devices may be changed into one or more software modules to perform processing according to the present disclosure, and vice versa.
[0118] In this way, the idea of the present disclosure should not be limited to the embodiments described above, and all things that are modified equally or equivalently to the claims described below as well as the claims are considered to fall within the scope of the idea of the present disclosure.
Claims
1. A method for determining input of a proximity voice for voice recognition, A step of detecting a voice from a microphone of a proximity voice input device to obtain a voice signal and sound pressure of the voice, A step of classifying the properties of the voice by analyzing the frequency spectrum of the voice signal, A step of calculating a first distance, which is an utterance distance of the voice, based on the properties of the voice and the sound pressure of the voice; and A step of determining whether the voice is a proximity voice caused by close-range speech based on the first distance. Including method.
2. In paragraph 1, A method for classifying the properties of the above voice, wherein the above voice is classified into one of a whisper and a normal voice.
3. In paragraph 2, In the step of determining whether the above-mentioned voice is a proximity voice, a method of determining that the voice is a proximity voice if the sound pressure of the voice is greater than a threshold sound pressure preset according to the properties of the voice.
4. In paragraph 3, If the above voice is classified as a whisper, if the sound pressure of the voice is greater than the first threshold sound pressure, it is judged as a close voice, A method for determining a proximity voice when the sound pressure of the voice is greater than a second threshold sound pressure that is greater than a first threshold sound pressure, when the voice is classified as a normal voice.
5. In paragraph 1, In the step of calculating the first distance, the method calculates the first distance by referring to the correlation between the speech distance and sound pressure according to the properties of the voice.
6. In paragraph 5, A method in which the correlation between the above-mentioned utterance distance and sound pressure is reflected by a weight according to the spectrum of the above-mentioned voice signal.
7. In paragraph 1, Before the step of determining whether the proximity voice is a proximity voice, the step of calculating a second distance, which is a distance between the proximity voice input device and an object detected through a proximity sensor of the proximity voice input device, is further included. A method in which, in the step of determining whether the above-mentioned proximity voice is present, whether the above-mentioned proximity voice is present is determined based on the first distance and the second distance.
8. In paragraph 7, Before the step of determining whether the above-mentioned proximity voice is present, a step of determining the similarity between the first distance and the second distance is further included. A method in which, in the step of determining whether the above-mentioned proximity voice is a proximity voice, whether the above-mentioned proximity voice is a proximity voice is determined based on the similarity between the first distance and the second distance.
9. A recording medium recording a computer program for executing the method according to paragraph 1.
10. As a proximity voice input device for voice recognition, A microphone that detects voice and acquires voice signals; processor Including, The processor analyzes the frequency spectrum of the voice signal to classify the properties of the voice, calculates a first distance, which is the speech distance of the voice signal, based on the properties of the voice and the sound pressure of the voice, and determines whether the voice is a proximity voice due to close-range speech based on the first distance. Proximity voice input device.