Intelligent wearable device and voice interaction method

By setting up a vibration detection module in the smart wearable device and combining it with a processing module to determine the tissue vibration when the wearer speaks, the problem of misrecognition in voice interaction is solved and the user experience is improved.

CN120612937APending Publication Date: 2025-09-09WEIFANG GOERTEK MICROELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510763264.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing smart wearable devices are prone to misrecognition problems during voice interaction, especially when multiple people use the same type of device and the same wake-up word. The voice information of the person next to them is picked up by the microphone, resulting in false wake-up or misrecognition.

Method used

A vibration detection module is set in the smart wearable device, directly or through an intermediary, at the part of the human body that vibrates when the wearer speaks. Combined with the processing module, it is determined whether the interactive content in the external voice is synchronized with the vibration of the human body tissue, ensuring that the interactive content is only carried out when the wearer speaks.

Benefits of technology

The vibration detection module determines whether the external voice is made by the wearer, reducing misidentification and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612937A_ABST
    Figure CN120612937A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wearable devices, and discloses an intelligent wearable device and a voice recognition method, the intelligent wearable device comprises a vibration detection module and a processing module, and the processing module is connected with the vibration detection module. The vibration detection module is directly arranged at or through a medium at the part where the human tissue vibrates when the wearer makes a sound; the processing module is used for obtaining interaction content in the external voice when the preset interaction voice exists in the collected external voice; the processing module is also used for judging whether the vibration detection module detects the vibration of the human tissue or not when the interaction content is obtained; and the processing module is also used for performing voice interaction according to the interaction content when the vibration of the human tissue is detected. Compared with the existing phenomenon that misrecognition easily occurs, whether the external voice is sent by the wearer or not can be judged through the vibration detection module, the misrecognition phenomenon is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of wearable devices, and in particular to a smart wearable device and a voice interaction method. Background Art

[0002] In smart wearable products, voice interaction function design has become increasingly common. Its function is generally achieved by using a microphone to pick up the wearer's voice information for voice processing and analysis to complete the corresponding control.

[0003] However, because human voices are transmitted through air vibrations, the microphones inevitably pick up the voices of people around the wearer, leading to false wakeups or misidentification. This means that someone speaking nearby could potentially control someone else's device, especially when multiple people use the same device and use the same wake-up word. Therefore, reducing this misidentification problem is an urgent issue. Summary of the Invention

[0004] The main purpose of this application is to provide a smart wearable device and a voice interaction method, aiming to solve the existing technical problem of how to reduce the phenomenon of misrecognition.

[0005] To achieve the above-mentioned object, the present application provides a smart wearable device, comprising: a vibration detection module and a processing module, wherein the processing module is connected to the vibration detection module, and the vibration detection module is directly or through an intermediary disposed at a part of the human body where vibration occurs when the wearer speaks;

[0006] The processing module is configured to obtain the interactive content in the collected external voice when there is a preset interactive voice in the collected external voice;

[0007] The processing module is further configured to determine whether the vibration detection module detects vibration of the human tissue when obtaining the interactive content;

[0008] The processing module is further configured to perform voice interaction according to the interaction content when vibration of the human tissue is detected.

[0009] In one embodiment, the smart wearable device further includes: a voice acquisition module and a voice recognition module;

[0010] The speech recognition module is connected to the speech acquisition module and the processing module;

[0011] The voice collection module is used to collect external voices;

[0012] The speech recognition module is used to recognize the external speech and determine whether the external speech contains a preset interactive speech according to the recognition result;

[0013] The speech recognition module is further configured to obtain the interactive content in the external speech according to the recognition result when the preset interactive speech exists in the external speech, and transmit the interactive content to the processing module.

[0014] In one embodiment, the smart wearable device further includes: a logic module;

[0015] The logic module is connected to the speech recognition module, the vibration detection module and the processing module respectively;

[0016] The speech recognition module is further configured to transmit a generated interrupt signal to the logic module when the preset interactive speech is present in the external speech;

[0017] The vibration detection module is configured to transmit a generated vibration signal to the logic module when vibration of the human tissue is detected;

[0018] The logic module is used to transmit the generated valid signal to the processing module when receiving the interrupt signal and the vibration signal, so that the processing module determines that the vibration detection module detects the vibration of the human tissue when obtaining the interactive content.

[0019] In one embodiment, the smart wearable device further includes: a comparison module;

[0020] The comparison module is connected to the vibration detection module and the logic module respectively;

[0021] The comparison module is configured to transmit the generated pulse signal to the logic module when the signal value of the vibration signal is higher than a preset signal threshold;

[0022] The logic module is further configured to transmit a generated valid signal to the processing module when receiving the interrupt signal and the pulse signal.

[0023] In one embodiment, the smart wearable device further includes: a pulse adjustment module;

[0024] The pulse adjustment module is connected to the comparison module and the logic module respectively;

[0025] The pulse adjustment module is used to adjust the pulse width of the pulse signal to a preset pulse width and transmit the adjusted pulse signal to the logic module. The preset pulse width is greater than the pulse width of the interrupt signal.

[0026] In one embodiment, the processing module is further connected to the voice acquisition module;

[0027] The processing module is further configured to, when detecting the vibration of the human tissue, determine a first signal sequence corresponding to the external speech, filter the first signal sequence based on a time period during which the interactive content appears in the external speech, and determine a first starting time and a first ending time of the filtered first signal sequence;

[0028] The processing module is further configured to determine a second signal sequence corresponding to the vibration signal, and determine a second starting time and a second ending time of the second signal sequence;

[0029] The processing module is also used to determine a first time difference based on the first start time and the second start time, determine a second time difference based on the first end time and the second end time, and perform voice interaction according to the interaction content when the first time difference and the second time difference meet the corresponding preset time difference conditions.

[0030] In one embodiment, the processing module is further connected to the voice acquisition module;

[0031] The speech recognition module is further configured to perform spectrum analysis on the speech corresponding to the interactive content in the external speech to obtain a first frequency component;

[0032] The processing module is further configured to, when detecting vibration of the human tissue, determine a second signal sequence corresponding to the vibration signal, and perform spectrum analysis on the second signal sequence to obtain a second frequency component;

[0033] The processing module is further configured to determine a frequency correlation between the first frequency component and the second frequency component, and perform voice interaction according to the interaction content when the frequency correlation meets a preset correlation condition.

[0034] In one embodiment, the smart wearable device further includes: a sight line detection module;

[0035] The sight line detection module is connected to the processing module;

[0036] The sight line detection module is used to collect eye information of the wearer and transmit the eye information to the processing module;

[0037] The processing module is further configured to determine the wearer's current gaze direction based on the eyeball information, and when vibration of the human tissue is detected, determine the expected gaze direction based on the interaction content;

[0038] The processing module is further configured to perform voice interaction according to the interaction content when a direction difference between the current gaze direction and the expected gaze direction satisfies a preset direction difference condition.

[0039] In one embodiment, the processing module is further configured to obtain a current display interface and determine an expected gaze content based on the interaction content;

[0040] The processing module is further configured to determine an expected gaze direction according to the expected gaze content when the expected gaze content exists on the current display interface;

[0041] The processing module is further configured to determine the expected gaze direction according to the current display interface when the expected gaze content does not exist on the current display interface.

[0042] In addition, to achieve the above-mentioned purpose, the present application also provides a voice interaction method, which is applied to a smart wearable device provided with a vibration detection module, wherein the vibration detection module is directly provided or provided through an intermediary at a part of the human body that vibrates when the wearer speaks;

[0043] The method comprises:

[0044] When the collected external voice contains a preset interactive voice, obtaining the interactive content in the external voice;

[0045] When obtaining the interactive content, determining whether the vibration detection module detects vibration of the human tissue;

[0046] When vibration of the human tissue is detected, voice interaction is performed according to the interaction content.

[0047] The present application provides an intelligent wearable device and a voice recognition method, wherein the intelligent wearable device includes: a vibration detection module and a processing module, wherein the processing module is connected to the vibration detection module, and the vibration detection module is directly or through an intermediary arranged at a part of the human body where the wearer's human tissue vibrates when speaking; the processing module is used to obtain the interactive content in the collected external voice when there is a preset interactive voice; the processing module is also used to determine whether the vibration detection module detects the vibration of the human tissue when obtaining the interactive content; the processing module is also used to perform voice interaction according to the interactive content when the vibration of the human tissue is detected.

[0048] Since the smart wearable device in this application can be provided with a vibration module, and the vibration module is provided at the part of the human body that vibrates when the wearer speaks. In actual use, when there is a preset interactive voice in the collected external voice, the interactive content in the external voice can be obtained first, and when the interactive content is obtained, it is determined whether the vibration detection module detects vibration. If it is detected, it can be said that the wearer is speaking at this time, and then voice interaction can be performed based on the interactive content. Compared with the existing phenomenon that is prone to misidentification, in this application, the vibration detection module can be used to determine whether the external voice is issued by the wearer, thereby reducing the misidentification phenomenon and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0051] Figure 1 This is a structural block diagram of the first embodiment of the smart wearable device of this application;

[0052] Figure 2 This is a processing logic diagram of the first embodiment of the smart wearable device of this application;

[0053] Figure 3 This is a structural block diagram of the second embodiment of the smart wearable device of this application;

[0054] Figure 4 This is a structural block diagram of the third embodiment of the smart wearable device of this application;

[0055] Figure 5 This is a flow chart of the first embodiment of the voice interaction method of the present application.

[0056] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0057] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0058] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0059] It should be noted that all directional indications in the embodiments of the present application (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0060] In addition, the descriptions of "first", "second", etc. in this application are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0061] It is understandable that voice interaction function design has become increasingly common in smart wearable products. Its function is generally realized by using a microphone to pick up the wearer's voice information for voice processing and analysis to complete the corresponding control.

[0062] However, because human voices are transmitted through air vibrations, the microphones inevitably pick up the voices of people around the wearer, leading to false wakeups or misidentification. This means that someone speaking nearby could potentially control someone else's device, especially when multiple people use the same device and use the same wake-up word. Therefore, reducing this misidentification problem is an urgent issue.

[0063] Therefore, in order to solve the above-mentioned defects, this embodiment provides a smart wearable device. Since the smart wearable device in this embodiment can be provided with a vibration module, and the vibration module is arranged at the part of the human body that vibrates when the wearer speaks. Furthermore, in actual use, when there is a preset interactive voice in the collected external voice, the interactive content in the external voice can be obtained first, and when the interactive content is obtained, it is determined whether the vibration detection module detects vibration. If it is detected, it can be said that the wearer is speaking at this time, and then voice interaction can be carried out according to the interactive content. Compared with the existing phenomenon that is prone to misidentification, this embodiment can use the vibration detection module to determine whether the external voice is uttered by the wearer, thereby reducing the misidentification phenomenon and improving the user experience.

[0064] For ease of understanding, the following Figures 1 to 4 The smart wearable device provided in the embodiments of the present application is specifically introduced.

[0065] Reference Figure 1 , Figure 1 This is a structural diagram of the first embodiment of the smart wearable device of this application, as shown in FIG. Figure 1 As shown, in this embodiment, the smart wearable device includes: a vibration detection module and a processing module, the processing module is connected to the vibration detection module, and the vibration detection module is directly or through a medium located at the part where the human body tissue vibrates when the wearer speaks.

[0066] It should be noted that the smart wearable device in this embodiment can be any device worn by a user, such as a smart bracelet, smart glasses, etc., and this embodiment does not limit this. The vibration detection module can be a module for detecting vibration, such as a vibration sensor. The processing module can be any module with logic processing and program execution functions, such as a microcontroller unit (MCU), and this embodiment does not limit this.

[0067] It should also be noted that when the wearer makes a sound, some of the wearer's body tissues will vibrate, such as the nose wings and ears. Therefore, in order to detect whether the wearer has made a sound in this embodiment, based on this principle, the vibration detection module can be directly installed or installed through an intermediary at the part of the wearer's body tissue that vibrates when the wearer makes a sound. When it is directly installed, it can directly contact the body tissue. When it is installed through an intermediary, it can contact the body tissue by utilizing the characteristics of the intermediary. The intermediary can be an object that can fix the vibration detection module in contact with the body tissue, such as adhesive or a carrier worn on the wearer. The specific configuration of the above-mentioned intermediary can be based on the actual product used, and this embodiment does not limit this.

[0068] For example, when the smart wearable device is smart glasses, it can be set near the nose pad of the smart glasses so that when the wearer wears the smart glasses, the vibration detection module can detect whether the wearer's nose bridge vibrates, or it can be set at the end of the temple of the smart glasses so that when the wearer wears the smart glasses, the vibration detection module can detect whether the wearer's ears (similar to bone conduction) vibrate.

[0069] As another implementation, when the smart wearable device is a smart bracelet, the vibration detection module can be separated from the smart wearable device. That is, the vibration detection module can be an independent component and can be placed in the part of the human body where the wearer's body tissue vibrates when the wearer speaks. For example, it can be independently placed in the wearer's ear in a form similar to bone conduction headphones, which can also achieve the above-mentioned vibration collection. Of course, other forms of vibration collection can also be used, and this embodiment does not limit this.

[0070] like Figure 1 As shown, in actual use, the processing module is used to obtain the interactive content in the external voice when there is a preset interactive voice in the collected external voice.

[0071] It is understandable that the above-mentioned external voice can be a sound outside the smart wearable device. The above-mentioned preset interactive voice can be a voice supported by the smart wearable device for voice interaction, such as the wake-up word of the smart wearable device or a specific control command (such as "cut song", "pause", "play", etc.). The above-mentioned interactive content can be the specific semantic content of the wake-up word or control command, which can be understood as the specific content to be executed by the smart device.

[0072] It is also understandable that in order to collect the above-mentioned external voice in this embodiment, Figure 1 As shown, the above-mentioned smart wearable device may be provided with a voice acquisition module, and the voice acquisition module may be any module with a voice acquisition function, such as an acoustic sensor, etc., and this embodiment does not impose any limitation on this.

[0073] In actual use, the voice collection module can transmit the collected external voice to the processing module. The processing module can recognize the external voice and determine whether the above-mentioned preset interactive voice is contained therein. If so, the specific interactive content in the external voice can be obtained; if not, no subsequent response will be made.

[0074] Continue as Figure 1 As shown, the processing module is further configured to determine whether the vibration detection module detects vibration of the human tissue when obtaining the interactive content;

[0075] The processing module is further configured to perform voice interaction according to the interaction content when vibration of the human tissue is detected.

[0076] When the processing module obtains the interaction content, it indicates that the external voice contains voice that the smart wearable device can interact with. It then needs to determine whether the voice is produced by the wearer. The processing module then uses the vibration detection module to determine whether the wearer's body tissue is vibrating. If no vibration is generated, it can be determined that the external voice is not produced by the wearer, and the processing module will not respond. If vibration is generated, it can be determined that the external voice is produced by the wearer, and the processing module can then conduct voice interaction based on the interaction content.

[0077] In a specific implementation, when the collected external speech contains a preset interactive voice, the interactive content of the external speech can be first obtained. Once the interactive content is obtained, the vibration detection module can be used to determine whether vibration has been detected. If vibration is detected, it can be determined that the wearer is speaking, and voice interaction can then be performed based on the interactive content. Compared to the existing system that is prone to misidentification, this embodiment uses the vibration detection module to determine whether the external speech is from the wearer, reducing misidentification and improving the user experience.

[0078] Further, in order to recognize the speech signal, continue as follows Figure 1 As shown, the smart wearable device in this embodiment may further include: a voice recognition module;

[0079] The speech recognition module is connected to the speech acquisition module and the processing module;

[0080] The voice collection module is used to collect external voices;

[0081] The speech recognition module is used to recognize the external speech and determine whether the external speech contains a preset interactive speech according to the recognition result;

[0082] The speech recognition module is further configured to obtain the interactive content in the external speech according to the recognition result when the preset interactive speech exists in the external speech, and transmit the interactive content to the processing module.

[0083] It should be understood that the voice collection module can be any module with a voice collection function, such as a microphone, etc. The voice recognition module can be any module with a voice recognition function, such as a voice recognition chip, etc., and this embodiment does not impose any limitation on this.

[0084] like Figure 1 As shown, in this embodiment, the voice recognition module can be connected to the voice collection module and the processing module respectively. Figure 2 , Figure 2 This is a processing logic diagram in the first embodiment of the smart wearable device of this application. Figure 2As shown, in actual use, the voice collection module can collect external voice (i.e. Figure 2 The microphone picks up the sound) and transmits it to the speech recognition module, which then recognizes the external speech (i.e. Figure 2 Then, the recognition result can be used to determine whether there is a preset interactive voice in the external voice; if so, the specific interactive content can be determined according to the recognition result, and the interactive content can be transmitted to the processing module (i.e. Figure 2 If it does not exist, no subsequent response will be made.

[0085] In this embodiment, the smart wearable device may be equipped with a vibration module, and this module is located in the area of ​​the human body that vibrates when the wearer speaks. Furthermore, in actual use, when the collected external voice contains a preset interactive voice, the interactive content in the external voice can be obtained first. When the interactive content is obtained, it is determined whether the vibration detection module detects vibration. If it is detected, it can be determined that the wearer is speaking at this time, and voice interaction can then be carried out based on the interactive content. Compared to the existing phenomenon that is prone to misidentification, this embodiment can use the vibration detection module to determine whether the external voice is uttered by the wearer, reducing misidentification and improving the user experience.

[0086] Reference Figure 3 , Figure 3 This is a structural block diagram of the second embodiment of the smart wearable device of this application, as shown in FIG. Figure 3 As shown, in order to enable the processing module to determine whether the vibration detection module detects vibration when receiving the interactive content, in this embodiment, the smart wearable device further includes: a logic module;

[0087] The logic module is connected to the speech recognition module, the vibration detection module and the processing module respectively.

[0088] It should be noted that the above-mentioned logic module can be a module for implementing output when all inputs meet the requirements (such as high level), such as an AND gate, etc. Of course, it can also be a module for implementing other types of functions, and this embodiment does not limit this.

[0089] The speech recognition module is further configured to transmit a generated interrupt signal to the logic module when the preset interactive speech is present in the external speech;

[0090] The vibration detection module is configured to transmit a generated vibration signal to the logic module when vibration of the human tissue is detected;

[0091] The logic module is used to transmit the generated valid signal to the processing module when receiving the interrupt signal and the vibration signal, so that the processing module determines that the vibration detection module detects the vibration of the human tissue when obtaining the interactive content.

[0092] It is understandable that if Figure 3 As shown, the above logic module can be connected to the voice recognition module, the vibration detection module and the processing module respectively, and the processing module can also be directly connected to the voice recognition module.

[0093] It is also understood that the interrupt signal may be a signal indicating the presence of a preset interactive voice in the external voice. The vibration signal may be a signal indicating that the vibration detection module has detected human tissue vibration. The valid signal may be a signal indicating that both the interrupt signal and the vibration signal have been received.

[0094] In actual use, the speech recognition module can be directly connected to the processing module via the Inter-Integrated Circuit Bus (I2C). Then the speech recognition module can directly transmit the obtained interactive content to the processing module (i.e. Figure 3 I2C), and the voice recognition module can generate an interrupt signal (i.e. Figure 3 and transmits the interrupt signal to the logic module.

[0095] In this embodiment, the logic module is also connected to the vibration detection module. When the vibration detection module detects vibrations in human tissue, it generates a vibration signal and transmits it to the logic judgment module. At this time, since both the interrupt signal and the vibration signal are at a high level, the logic module can generate the aforementioned valid signal upon receiving the interrupt signal and the vibration signal and transmit it to the processing module. When the processing module receives this valid signal, it can determine that the external voice is the voice of the wearer. Then, when the processing module receives the interactive content, it can determine that the vibration detection module has detected vibrations in human tissue, and thus perform voice interaction based on the interactive content.

[0096] Furthermore, considering that when the wearer generates slight vibration but does not make a sound, the vibration detection module will also generate the above-mentioned vibration signal, so in order to improve the accuracy of the detection, continue as follows Figure 3 As shown, in this embodiment, the smart wearable device further includes: a comparison module;

[0097] The comparison module is connected to the vibration detection module and the logic module respectively;

[0098] The comparison module is configured to transmit the generated pulse signal to the logic module when the signal value of the vibration signal is higher than a preset signal threshold;

[0099] The logic module is further configured to transmit a generated valid signal to the processing module when receiving the interrupt signal and the pulse signal.

[0100] It should be understood that the comparison module can be any module for implementing the comparison function. Figure 3 As shown, in this embodiment, a comparator can be used for illustration, that is, the positive input terminal of the comparator can be connected to the vibration detection module, and the negative input terminal of the comparator can be connected to any device for generating a reference voltage (ie Figure 3 A module for determining a preset signal threshold value is connected to the reference voltage in the middle, and the preset signal can be set according to actual conditions, and this embodiment does not limit this.

[0101] It should also be understood that the above-mentioned pulse signal may be a signal used to represent a vibration signal generated by the wearer's voice.

[0102] In actual use, when the vibration detection module detects vibration, the generated vibration signal can be transmitted to the comparison module, and the comparison module can compare the signal value of the vibration signal with the preset signal threshold (i.e. Figure 2 When the signal value of the vibration signal is not higher than the preset signal threshold, it can be indicated that the vibration is not caused by the wearer's voice, and no pulse signal is transmitted to the logic module; when the signal value of the vibration signal is higher than the preset signal threshold, it can be indicated that the vibration is caused by the wearer's voice, and the comparison module can output a pulse signal to the logic module (i.e. Figure 2 Output signal). When the logic module receives the interrupt signal, it can determine whether the pulse signal is received. When the pulse signal is not received, no valid signal is output to the processing module. When the pulse signal is received, a valid signal can be output to the processing module (i.e. Figure 2 The speech recognition is effective and the effective signal is output).

[0103] Furthermore, in order to ensure that the interrupt signal received by the logic module can be fully identified, in this embodiment, continue as follows Figure 3 As shown, the smart wearable device further includes: a pulse adjustment module;

[0104] The pulse adjustment module is connected to the comparison module and the logic module respectively;

[0105] The pulse adjustment module is used to adjust the pulse width of the pulse signal to a preset pulse width and transmit the adjusted pulse signal to the logic module. The preset pulse width is greater than the pulse width of the interrupt signal.

[0106] It should be noted that the pulse adjustment module may be a module for adjusting the width of the pulse signal, such as a pulse width modulation (PWM) chip, etc., and this embodiment does not impose any limitation on this.

[0107] It should also be noted that, in this embodiment, the preset pulse width may be greater than the pulse width of the interrupt signal. The pulse width of the interrupt signal may be preset.

[0108] In actual use, after the comparison module outputs a pulse signal, it can be transmitted to the pulse adjustment module. The pulse adjustment module can adjust the pulse width of the pulse signal to be greater than the pulse width of the interrupt signal, and then transmit the adjusted pulse signal to the logic module. When the logic module receives the adjusted pulse signal, it can ensure that the interrupt signal is completely received, and then output a valid signal to the processing module.

[0109] Reference Figure 4 , Figure 4 This is a block diagram of the third embodiment of the smart wearable device of this application. Considering that when the external voice is generated by other external sources (such as other passers-by), and the wearer also generates other types of vibrations (but not the wearer generating the external voice), the vibration detection module may also output a vibration signal, causing the subsequent logic module to output a valid signal, resulting in misrecognition. Therefore, in this embodiment, if Figure 4 As shown, the processing module is also connected to the voice acquisition module;

[0110] The processing module is also used to determine a first signal sequence corresponding to the external voice when the vibration of the human tissue is detected, filter the first signal sequence according to the time period when the interactive content appears in the external voice, and determine the first starting time and the first ending time of the filtered first signal sequence.

[0111] It should be noted that the first signal sequence may be a digital signal sequence corresponding to the external voice obtained by the voice acquisition module. The first start time may be the time when the interactive content begins in the first signal sequence, and the first end time may be the time when the interactive content ends in the first signal sequence.

[0112] In actual use, the voice collection module can not only transmit the first signal sequence corresponding to the collected external voice to the voice recognition module for recognition, obtain the above-mentioned interactive content and determine whether there is a preset interactive voice, but also directly transmit the first signal sequence to the processing module; at the same time, the voice recognition module can also identify the time period when the interactive content appears in the external voice based on the first signal sequence, that is, identify the specific time period when the interactive content appears in the external voice, obtain the above-mentioned appearance time period, and transmit the appearance time period to the processing module;

[0113] After receiving the first signal sequence and the occurrence period, the processing module can filter the first signal sequence according to the occurrence period, determine the start time and end time of the interactive content in the first signal sequence, and use the start time as the above-mentioned first starting time and the end time as the above-mentioned first ending time.

[0114] The processing module is further configured to determine a second signal sequence corresponding to the vibration signal, and determine a second starting time and a second ending time of the second signal sequence.

[0115] It is understood that the second signal sequence may be a signal sequence corresponding to the vibrations collected by the vibration detection module. The second starting time may be the time when the wearer's vibration begins in the second signal sequence, and the second ending time may be the time when the wearer's vibration ends in the second signal sequence.

[0116] In this embodiment, the processing module can also be directly connected to the vibration detection module. In actual use, the vibration detection module can not only transmit the second signal sequence generated when vibration is collected to the logic module as the vibration signal, but also directly transmit this second signal sequence to the processing module. The processing module can determine the start and end times of the vibration signal, and use the start time as the second start time and the end time as the second end time.

[0117] It should be emphasized that, in order to ensure the synchronization of the moments during the subsequent time comparison, in this embodiment, the above modules can be connected to the same clock module for clock synchronization.

[0118] The processing module is also used to determine a first time difference based on the first start time and the second start time, determine a second time difference based on the first end time and the second end time, and perform voice interaction according to the interaction content when the first time difference and the second time difference meet the corresponding preset time difference conditions.

[0119] It should be understood that the first time difference may be a difference value obtained by subtracting the second starting time from the first starting time, and the second time difference may be a difference value obtained by subtracting the second ending time from the first ending time.

[0120] It should also be understood that the above-mentioned preset time difference condition may be a condition that the time difference is lower than a preset time difference threshold. Specifically, the preset time difference threshold may be set according to actual conditions, and this embodiment does not impose any limitation thereto.

[0121] Because the propagation paths of vibration and speech differ in their generation mechanisms when the wearer emits external speech that meets the preset interaction requirements, the vibrations of human tissue are detected slightly before the speech. Based on this principle, in this embodiment, after obtaining the first and second time differences, the processing module can determine whether the first and second time differences are less than corresponding preset time difference thresholds. If the first and / or second time differences are not less than the thresholds, it can be determined that the external speech was not emitted by the wearer, but simply coincidentally the wearer is vibrating. The processing module then does not proceed with subsequent speech interaction based on the interaction content. If both the first and second time differences are less than the corresponding preset time difference thresholds, it can be determined that the external speech was emitted by the wearer. The processing module can then proceed with speech interaction based on the interaction content. This improves the accuracy of the interaction.

[0122] Furthermore, in order to ensure that the external audio is emitted by the wearer, in this embodiment, the processing module is also connected to the voice collection module;

[0123] The speech recognition module is further configured to perform spectrum analysis on the speech corresponding to the interactive content in the external speech to obtain a first frequency component.

[0124] The processing module is further configured to, when detecting vibration of the human tissue, determine a second signal sequence corresponding to the vibration signal, and perform spectrum analysis on the second signal sequence to obtain a second frequency component;

[0125] It should be noted that the first frequency component mentioned above may be the frequency component of the interactive content in the external speech, for example, including but not limited to: fundamental frequency, harmonics, etc. This embodiment is not limited to this. The second frequency component mentioned above may be the frequency component of the second signal sequence corresponding to the vibration signal, for example, including but not limited to: fundamental frequency, harmonics, etc. This embodiment is not limited to this.

[0126] In actual use, the voice recognition module in this embodiment can perform spectral analysis on the interactive content part of the received external voice, for example, using Fourier transform, etc., to obtain the above-mentioned first frequency component; at the same time, the processing module can perform spectral analysis on the second signal sequence when receiving the second signal sequence, or it can be Fourier transform, to obtain the second frequency component.

[0127] The processing module is further configured to determine a frequency correlation between the first frequency component and the second frequency component, and perform voice interaction according to the interaction content when the frequency correlation meets a preset correlation condition.

[0128] It is understood that the frequency correlation may be the degree of overlap or similarity between the first frequency component and the second frequency component. The preset correlation condition may be a condition that the frequency correlation satisfies a preset correlation threshold. The preset correlation threshold may be set based on actual conditions and is not limited in this embodiment.

[0129] In actual use, after the voice recognition module obtains the first frequency component, it can transmit the first frequency component to the processing module. After receiving the first and second frequency components, the processing module can determine the similarity between the first and second frequency components as the frequency correlation, and determine whether the frequency correlation meets the preset correlation requirements. If it does not meet the requirements, it can be determined that the voice was not generated by the wearer at this time, and the processing module will not proceed with the subsequent voice interaction based on the interaction content. If it meets the requirements, it can be determined that the voice was generated by the wearer at this time, and the processing module will proceed with the voice interaction based on the interaction content.

[0130] Further, continue as Figure 4 As shown, in order to further ensure the accuracy of recognition, in this embodiment, the above-mentioned smart wearable device further includes: a sight detection module;

[0131] The sight line detection module is connected to the processing module;

[0132] The sight line detection module is used to collect eye information of the wearer and transmit the eye information to the processing module.

[0133] It should be noted that the gaze detection module may be any module for collecting the wearer's eye information, such as a camera, an eye tracking module, etc., and this embodiment does not limit this. The eye information may be information about the wearer's eye gaze characteristics.

[0134] As an implementation method, when the smart wearable device is a pair of smart glasses, the gaze detection module can be installed on the frame of the smart glasses or elsewhere. This gaze detection module can collect the wearer's eye information and transmit the eye information to the processing module. Similarly, if the smart wearable device is a smart bracelet, a camera or the like can be installed on the smart bracelet to collect eye information.

[0135] The processing module is further configured to determine the wearer's current gaze direction based on the eyeball information, and when vibration of the human tissue is detected, determine the expected gaze direction based on the interaction content;

[0136] The processing module is further configured to perform voice interaction according to the interaction content when a direction difference between the current gaze direction and the expected gaze direction satisfies a preset direction difference condition.

[0137] It is understood that the current gaze direction may be the direction in which the wearer's eyes are currently looking. The expected gaze direction may be the direction in which the wearer should be looking when emitting the interactive content. Specifically, the processing module may determine a user requirement based on the interactive content, and determine the direction in which the wearer should be looking based on the user requirement as the expected gaze direction.

[0138] For example, in this embodiment, the smart wearable device may have a display function. When the wearer has a voice interaction need, they generally look at the corresponding position in the content displayed by the smart wearable device. For example, when the wearer needs to check the weather, they generally say "Hello Xiao X (the wake-up word for the smart wearable device), what's the weather like today?" while looking at the corresponding weather icon in the content displayed by the smart wearable device. Therefore, the processing module can determine that the wearer needs to check the weather based on the interaction content (i.e., the above-mentioned "What's the weather like today?"), and then the corresponding expected gaze direction is the direction in which the wearer's eyes are looking at the corresponding weather icon in the content displayed by the smart wearable device.

[0139] It should be emphasized that the above-mentioned preset direction difference condition may be a condition that is less than a preset direction difference threshold. Specifically, the preset direction difference threshold may be set according to actual conditions, and this embodiment does not impose any limitation on this.

[0140] In actual use, after the processing module receives eye information, it can first determine the wearer's current gaze direction based on the eye information. When the vibration detection module detects human tissue vibration, it can determine the corresponding expected gaze direction based on the obtained interaction content. The processing module can then determine the direction difference between the current gaze direction and the expected gaze direction, and determine whether the direction difference is less than a preset direction difference threshold. If it is less than, it can be determined that the preset direction difference condition is met, indicating that the external audio is generated by the wearer, and the processor can then perform voice interaction based on the interaction content. If it is not less than, it can be determined that the preset direction difference condition is not met, indicating that the wearer is not looking in the corresponding direction and the external audio is not generated by the wearer. The processor can then not perform subsequent voice interaction, thereby reducing the occurrence of misidentification.

[0141] Furthermore, in order to determine the expected gaze direction corresponding to the interactive content, in this embodiment, the processing module is further configured to obtain the current display interface and determine the expected gaze content according to the interactive content;

[0142] The processing module is further configured to determine an expected gaze direction according to the expected gaze content when the expected gaze content exists on the current display interface;

[0143] The processing module is further configured to determine the expected gaze direction according to the current display interface when the expected gaze content does not exist on the current display interface.

[0144] It should be noted that the current display interface may be the interface currently displayed on the display screen of the smart wearable device. The expected gaze content may be the content that should exist in the display interface of the smart wearable device corresponding to the user demand for the interactive content, such as an icon, window, etc.

[0145] In actual use, after obtaining the current gaze direction and detecting the vibration of human tissue, the above-mentioned processing module can first obtain the current display interface of the smart wearable device, determine the user needs based on the interactive content, and determine the corresponding expected gaze content (such as the above-mentioned weather icon) based on the user needs.

[0146] After determining the expected gaze content, taking into account the situations where the expected gaze content may or may not exist on the current display interface, when the processing module determines that the expected gaze content exists on the current display interface, the relative direction of the expected gaze content and the wearer's eyeball can be used as the above-mentioned expected gaze direction; if it is determined that the expected gaze content does not exist on the current display interface, the direction of the entire current display interface and the wearer's eyeball can be used as the above-mentioned expected gaze direction.

[0147] For example, if the current display interface of the smart bracelet is not lit (i.e., not awakened), the processing module can determine that there is no expected gaze content (i.e., there is no weather icon) in the current display interface (i.e., the interface in the dormant state), and then the processing module can use the entire current display interface, that is, the direction between the wearer's eyeball and the entire display screen as the above-mentioned expected gaze direction. As long as the wearer looks at the entire display screen, it can be determined that the direction difference between the wearer's current gaze direction and the expected gaze direction meets the preset direction difference condition.

[0148] For example, if the current display interface of a smart bracelet is the main interface (i.e., the working interface after waking up), and the weather icon is located on the main interface, the processing module can determine that the expected gaze content exists in the current display interface (i.e., the weather icon does not exist). Then, the processing module can use the direction between the wearer's eyeball and the weather icon as the expected gaze direction. When the wearer looks at the weather icon, it can be determined that the direction difference between the wearer's current gaze direction and the expected gaze direction meets the preset direction difference condition. This improves the accuracy of recognition.

[0149] Reference Figure 5 , Figure 5 This is a flow chart of the first embodiment of the voice interaction method of this application, as shown in FIG. Figure 5 As shown, an embodiment of the present application further provides a voice interaction method, which is applied to a smart wearable device provided with a vibration detection module, wherein the vibration detection module is directly provided or provided through an intermediary at a part of the human body that vibrates when the wearer speaks;

[0150] The method comprises:

[0151] Step S10: When there is a preset interactive voice in the collected external voice, obtaining the interactive content in the external voice;

[0152] Step S20: when obtaining the interactive content, determining whether the vibration detection module detects vibration of the human tissue;

[0153] Step S30: When the vibration of the human tissue is detected, voice interaction is performed according to the interaction content.

[0154] It should be noted that the execution subject of the above method in this embodiment can be any device with data processing, program running and voice interaction functions, such as the smart wearable device in the above embodiments, specifically the processing module in the above smart wearable device.

[0155] It should also be noted that the specific structure and implementation of the above-mentioned smart wearable device in this embodiment can refer to the description of the above-mentioned smart wearable device embodiment, and this embodiment is not limited to this.

[0156] In this embodiment, the smart wearable device may be equipped with a vibration module, and this module is located in the area of ​​the human body that vibrates when the wearer speaks. Furthermore, in actual use, when the collected external voice contains a preset interactive voice, the interactive content in the external voice can be obtained first. When the interactive content is obtained, it is determined whether the vibration detection module detects vibration. If it is detected, it can be determined that the wearer is speaking at this time, and voice interaction can then be carried out based on the interactive content. Compared to the existing phenomenon that is prone to misidentification, this embodiment can use the vibration detection module to determine whether the external voice is uttered by the wearer, reducing misidentification and improving the user experience.

[0157] Furthermore, the smart wearable device further includes: a voice acquisition module and a voice recognition module;

[0158] When the collected external voice contains a preset interactive voice, before the step of obtaining the interactive content in the external voice, the method further includes:

[0159] Collecting external voices through the voice collection module;

[0160] Recognizing the external voice by the voice recognition module, and determining whether there is a preset interactive voice in the external voice according to the recognition result;

[0161] When the preset interactive voice exists in the external voice, the voice recognition module obtains the interactive content in the external voice according to the recognition result.

[0162] Furthermore, the smart wearable device further includes: a logic module;

[0163] The logic module is connected to the speech recognition module, the vibration detection module and the processing module respectively;

[0164] The step of determining whether the vibration detection module detects the vibration of the human tissue comprises:

[0165] When the preset interactive voice exists in the external voice, the voice recognition module transmits the generated interrupt signal to the logic module;

[0166] When the vibration of the human tissue is detected by the vibration detection module, the generated vibration signal is transmitted to the logic module;

[0167] When obtaining the interactive content, if a valid signal is received, it is determined that the vibration detection module detects the vibration of the human tissue, and the valid signal is generated by the logic module when the interrupt signal and the vibration signal are received.

[0168] Furthermore, the smart wearable device further includes: a comparison module;

[0169] The comparison module is connected to the vibration detection module and the logic module respectively;

[0170] Before the step of determining that the vibration detection module detects the vibration of the human tissue if a valid signal is received when obtaining the interactive content, the method further includes:

[0171] When the signal value of the vibration signal is higher than a preset signal threshold, the comparison module transmits the generated pulse signal to the logic module;

[0172] The logic module generates a valid signal when receiving the interrupt signal and the pulse signal.

[0173] Furthermore, the smart wearable device further comprises: a pulse adjustment module;

[0174] The pulse adjustment module is connected to the comparison module and the logic module respectively;

[0175] After the step of generating a valid signal by the logic module upon receiving the interrupt signal and the pulse signal, the method further includes:

[0176] The pulse width of the pulse signal is adjusted to a preset pulse width by the pulse adjustment module, and the adjusted pulse signal is transmitted to the logic module. The preset pulse width is greater than the pulse width of the interrupt signal.

[0177] Furthermore, the processing module is also connected to the voice collection module;

[0178] The step of performing voice interaction according to the interaction content when the vibration of the human tissue is detected includes:

[0179] When the human tissue vibration is detected, determining a first signal sequence corresponding to the external voice, filtering the first signal sequence according to a time period during which the interactive content appears in the external voice, and determining a first start time and a first end time of the filtered first signal sequence;

[0180] determining a second signal sequence corresponding to the vibration signal, and determining a second starting time and a second ending time of the second signal sequence;

[0181] A first time difference is determined based on the first start time and the second start time, and a second time difference is determined based on the first end time and the second end time. When the first time difference and the second time difference meet the corresponding preset time difference conditions, voice interaction is performed according to the interaction content.

[0182] Furthermore, the processing module is also connected to the voice collection module;

[0183] The step of performing voice interaction according to the interaction content when the vibration of the human tissue is detected includes:

[0184] Performing spectrum analysis on the speech corresponding to the interactive content in the external speech by the speech recognition module to obtain a first frequency component;

[0185] When the vibration of the human tissue is detected, determining a second signal sequence corresponding to the vibration signal, and performing spectrum analysis on the second signal sequence to obtain a second frequency component;

[0186] Determine a frequency correlation between the first frequency component and the second frequency component, and perform voice interaction according to the interaction content when the frequency correlation meets a preset correlation condition.

[0187] Furthermore, the smart wearable device further includes: a sight line detection module;

[0188] The sight line detection module is connected to the processing module;

[0189] The step of performing voice interaction according to the interaction content when the vibration of the human tissue is detected includes:

[0190] Collecting the wearer's eye information through the sight line detection module;

[0191] determining a current gaze direction of the wearer based on the eyeball information, and determining an expected gaze direction based on the interaction content when vibration of the human tissue is detected;

[0192] When a direction difference between the current gaze direction and the expected gaze direction satisfies a preset direction difference condition, voice interaction is performed according to the interaction content.

[0193] Furthermore, the step of determining the expected gaze direction according to the interaction content includes:

[0194] Obtaining the current display interface and determining the expected gaze content based on the interaction content;

[0195] When the expected gaze content exists in the current display interface, determining the expected gaze direction according to the expected gaze content;

[0196] When the expected gaze content does not exist in the current display interface, the expected gaze direction is determined according to the current display interface.

[0197] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A smart wearable device, characterized in that: The smart wearable device includes: a vibration detection module and a processing module, the processing module is connected to the vibration detection module, and the vibration detection module is directly or through an intermediary provided at a part of the human body where vibration occurs when the wearer speaks; The processing module is configured to obtain the interactive content in the collected external voice when there is a preset interactive voice in the collected external voice; The processing module is further configured to determine whether the vibration detection module detects vibration of the human tissue when obtaining the interactive content; The processing module is further configured to perform voice interaction according to the interaction content when vibration of the human tissue is detected.

2. The smart wearable device according to claim 1, wherein: The smart wearable device also includes: a voice acquisition module and a voice recognition module; The speech recognition module is connected to the speech acquisition module and the processing module; The voice collection module is used to collect external voices; The speech recognition module is used to recognize the external speech and determine whether the external speech contains a preset interactive speech according to the recognition result; The speech recognition module is further configured to obtain the interactive content in the external speech according to the recognition result when the preset interactive speech exists in the external speech, and transmit the interactive content to the processing module.

3. The smart wearable device according to claim 2, wherein: The smart wearable device further includes: a logic module; The logic module is connected to the speech recognition module, the vibration detection module and the processing module respectively; The speech recognition module is further configured to transmit a generated interrupt signal to the logic module when the preset interactive speech is present in the external speech; The vibration detection module is configured to transmit a generated vibration signal to the logic module when vibration of the human tissue is detected; The logic module is used to transmit the generated valid signal to the processing module when receiving the interrupt signal and the vibration signal, so that the processing module determines that the vibration detection module detects the vibration of the human tissue when obtaining the interactive content.

4. The smart wearable device according to claim 3, wherein: The smart wearable device further includes: a comparison module; The comparison module is connected to the vibration detection module and the logic module respectively; The comparison module is configured to transmit the generated pulse signal to the logic module when the signal value of the vibration signal is higher than a preset signal threshold; The logic module is further configured to transmit a generated valid signal to the processing module when receiving the interrupt signal and the pulse signal.

5. The smart wearable device according to claim 4, wherein: The smart wearable device further includes: a pulse adjustment module; The pulse adjustment module is connected to the comparison module and the logic module respectively; The pulse adjustment module is used to adjust the pulse width of the pulse signal to a preset pulse width and transmit the adjusted pulse signal to the logic module. The preset pulse width is greater than the pulse width of the interrupt signal.

6. The smart wearable device according to claim 3, wherein: The processing module is also connected to the voice acquisition module; The processing module is further configured to, when detecting the vibration of the human tissue, determine a first signal sequence corresponding to the external speech, filter the first signal sequence based on a time period during which the interactive content appears in the external speech, and determine a first starting time and a first ending time of the filtered first signal sequence; The processing module is further configured to determine a second signal sequence corresponding to the vibration signal, and determine a second starting time and a second ending time of the second signal sequence; The processing module is also used to determine a first time difference based on the first start time and the second start time, determine a second time difference based on the first end time and the second end time, and perform voice interaction according to the interaction content when the first time difference and the second time difference meet the corresponding preset time difference conditions.

7. The smart wearable device according to claim 3, wherein: The processing module is also connected to the voice acquisition module; The speech recognition module is further configured to perform spectrum analysis on the speech corresponding to the interactive content in the external speech to obtain a first frequency component; The processing module is further configured to, when detecting vibration of the human tissue, determine a second signal sequence corresponding to the vibration signal, and perform spectrum analysis on the second signal sequence to obtain a second frequency component; The processing module is further configured to determine a frequency correlation between the first frequency component and the second frequency component, and perform voice interaction according to the interaction content when the frequency correlation meets a preset correlation condition.

8. The smart wearable device according to claim 1, wherein: The smart wearable device further includes: a sight line detection module; The sight line detection module is connected to the processing module; The sight line detection module is used to collect eye information of the wearer and transmit the eye information to the processing module; The processing module is further configured to determine the wearer's current gaze direction based on the eyeball information, and when vibration of the human tissue is detected, determine the expected gaze direction based on the interaction content; The processing module is further configured to perform voice interaction according to the interaction content when a direction difference between the current gaze direction and the expected gaze direction satisfies a preset direction difference condition.

9. The smart wearable device according to claim 8, wherein: The processing module is further configured to obtain a current display interface and determine an expected gaze content based on the interaction content; The processing module is further configured to determine an expected gaze direction according to the expected gaze content when the expected gaze content exists on the current display interface; The processing module is further configured to determine the expected gaze direction according to the current display interface when the expected gaze content does not exist on the current display interface.

10. A voice interaction method, characterized in that: The method is applied to a smart wearable device provided with a vibration detection module, wherein the vibration detection module is directly or through an intermediary provided at a part of the human body where vibration occurs when the wearer speaks; The method comprises: When the collected external voice contains a preset interactive voice, obtaining the interactive content in the external voice; When obtaining the interactive content, determining whether the vibration detection module detects vibration of the human tissue; When vibration of the human tissue is detected, voice interaction is performed according to the interaction content.