A method, system and device for preventing false wake-up of a smart voice device for voice interaction

By setting recognition features on smart voice devices and comparing the wake-up recognition features of wake-up commands, the problem of smart voice devices having difficulty recognizing user identities is solved, enabling safe operation during voice interaction and avoiding accidental operation.

CN115798473BActive Publication Date: 2026-02-17E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211640979.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-08-22
Filing Date
2022-12-20
Publication Date
2026-02-17
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

Existing smart voice devices have difficulty identifying whether the operation commands are issued by the user themselves, which can easily lead to misoperation, especially in video calls, and cause security incidents.

Method used

By setting recognition features on each smart voice device, the wake-up recognition features of the wake-up command are obtained and compared with the device's own recognition features. Only when they match are the operation executed; otherwise, the device remains silent.

Benefits of technology

This effectively avoids accidental operation caused by wake-up commands from non-users, improves the security of smart voice devices during voice interaction, and prevents security incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798473B_ABST
    Figure CN115798473B_ABST
Patent Text Reader

Abstract

The application relates to a false wake-up prevention method, system and device of an intelligent voice device for voice interaction, which is applied to video voice interaction of at least two intelligent voice devices corresponding to users, and each intelligent voice device is provided with an identification feature for identification. The method compares the identification feature of the intelligent voice device itself with the wake-up identification feature extracted from the received wake-up instruction, and only when the wake-up identification feature is consistent with the identification feature of the intelligent voice device, the intelligent voice device can execute corresponding operations according to the wake-up instruction, the intelligent voice device can avoid executing operations due to wake-up instructions that are not for the intelligent voice device in the voice interaction process, and the occurrence of safety accidents is avoided, and the technical problem that the existing intelligent voice device is difficult to identify whether an operation instruction is issued by the user itself and false operations exist is solved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202211006289.9, filed on August 22, 2022, entitled "A method, system and device for preventing accidental wake-up of a smart voice device with voice interaction". Technical Field

[0002] This application relates to the field of data transmission technology, and in particular to a method, system and device for preventing accidental wake-up of a smart voice device with voice interaction. Background Technology

[0003] With the development of artificial intelligence, various smart devices are becoming necessities in people's daily lives. In recent years, in particular, the gradual improvement of voice recognition technology has made human-computer interaction possible, and the smart speaker industry is growing exponentially. Smart speakers can do many things through voice control, especially after smart speakers are interconnected with smart home devices, allowing control of other connected devices such as lights, air conditioners, and televisions.

[0004] However, smart speakers lack effective control over user identification. A common problem arises when users A and B are video calling through a smart speaker; the speaker is easily misoperated by the other user's voice commands, causing significant inconvenience to both. For example, if user A commands to wake up their speaker, user B's speaker will also receive this command and be activated. In this situation, regardless of any other voice commands from user A, user B's speaker is likely to perform the same action. This misoperation of the speaker can easily lead to safety hazards, such as accidental activation or deactivation of smart devices like smart sockets, or misoperation of gas controls. Summary of the Invention

[0005] This application provides a method, system, and device for preventing accidental wake-up of intelligent voice devices with voice interaction, which solves the technical problem that existing intelligent voice devices have difficulty in identifying whether the operation command is issued by the user, resulting in accidental operation.

[0006] To achieve the above objectives, the embodiments of this application provide the following technical solutions:

[0007] A method for preventing accidental wake-up of intelligent voice devices with voice interaction is applied to video and voice interaction between at least two intelligent voice devices corresponding to users. Each of the intelligent voice devices is equipped with recognition features for identification. The method for preventing accidental wake-up of intelligent voice devices includes the following steps:

[0008] Acquire a first recognition feature of a first intelligent voice device and a first user corresponding to the first recognition feature, and acquire a second recognition feature of a second intelligent voice device and a second user corresponding to the second recognition feature;

[0009] When the first user and the second user interact with each other through their respective smart voice devices, the first user or the second user issues a wake-up command to control the operation of the corresponding smart voice device. The first smart voice device and the second smart voice device receive the wake-up command and perform feature extraction on the received wake-up command to obtain the wake-up recognition feature corresponding to the wake-up command.

[0010] The wake-up recognition feature is determined to be consistent with the first recognition feature or the second recognition feature. Only when the smart voice device corresponding to the first recognition feature or the second recognition feature that is consistent with the wake-up recognition feature is woken up can the smart voice device perform an operation according to the wake-up command.

[0011] Preferably, after the first user and the second user interact via their respective smart voice devices, the method for preventing accidental wake-up of the smart voice devices during the voice interaction includes: transmitting the first identification feature to the second smart voice device along with the video call and transmitting the second identification feature to the first smart voice device along with the video call.

[0012] Preferably, the method for preventing accidental wake-up of the intelligent voice device in voice interaction includes: if the intelligent voice device corresponding to the first identification feature or the second identification feature that is inconsistent with the wake-up identification feature will not be woken up, the intelligent voice device remains silent.

[0013] Preferably, the first identification feature, the second identification feature, and the wake-up identification feature are all white noise audio, and the white noise audio contains feature values ​​that identify its audio.

[0014] Preferably, the intelligent voice device is equipped with a feature generation module, a feature activation module, a feature playback module, a video and voice call module, a voice recognition module, and a feature matching module;

[0015] The feature generation module is used to generate recognition audio with recognition features corresponding to the user of the smart voice device;

[0016] The feature activation module is used to activate the feature playback module when the user speaks to the smart voice device;

[0017] The feature playback module is used to play the identified audio;

[0018] The video and voice call module is used for remote video call interaction;

[0019] The speech recognition module is used to recognize and extract wake-up recognition features from the video speech of the smart voice device receiving a wake-up command.

[0020] The feature matching module is used to match the wake-up recognition features with the recognition features of the smart voice device to determine whether to wake up or remain silent based on the wake-up command received by the smart voice device.

[0021] This application also provides a voice interaction-based intelligent voice device anti-mistake wake-up system, which is applied to video voice interaction between at least two intelligent voice devices corresponding to users. Each of the intelligent voice devices is equipped with recognition features for identification. The intelligent voice device anti-mistake wake-up system includes: a data acquisition unit, an interaction extraction unit, and a recognition wake-up unit.

[0022] The data acquisition unit is used to acquire a first recognition feature of a first intelligent voice device and a first user corresponding to the first recognition feature, and to acquire a second recognition feature of a second intelligent voice device and a second user corresponding to the second recognition feature.

[0023] The interaction extraction unit is used to enable the first user and the second user to interact via voice, and the first user or the second user to issue a wake-up command to control the operation of the corresponding smart voice device. The first smart voice device and the second smart voice device receive the wake-up command and perform feature extraction on the received wake-up command to obtain the wake-up recognition feature corresponding to the wake-up command.

[0024] The recognition and wake-up unit is used to determine whether it is consistent with the first recognition feature or the second recognition feature based on the wake-up recognition feature. Only when the smart voice device corresponding to the first recognition feature or the second recognition feature that is consistent with the wake-up recognition feature is woken up can the smart voice device perform an operation according to the wake-up command.

[0025] Preferably, the recognition and wake-up unit is further configured to keep the smart voice device silent if the smart voice device corresponding to the first recognition feature or the second recognition feature that is inconsistent with the wake-up recognition feature will not be woken up.

[0026] Preferably, the voice interaction intelligent voice device anti-mistake wake-up system includes a feature transmission unit, which is used to transmit the first identification feature to the second intelligent voice device along with the video call and the second identification feature to the first intelligent voice device along with the video call after the first user and the second user have interacted via voice.

[0027] Preferably, the first identification feature, the second identification feature, and the wake-up identification feature are all white noise audio, and the white noise audio contains feature values ​​that identify its audio.

[0028] This application also provides a terminal device, including a processor and a memory;

[0029] The memory is used to store program code and transmit the program code to the processor;

[0030] The processor is used to execute the above-described method for preventing accidental wake-up of intelligent voice devices based on instructions in the program code.

[0031] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: The method, system, and device for preventing accidental wake-up of intelligent voice devices with voice interaction include obtaining a first identification feature of a first intelligent voice device and a first user corresponding to the first identification feature, and obtaining a second identification feature of a second intelligent voice device and a second user corresponding to the second identification feature; when the first user and the second user interact with each other through their respective intelligent voice devices, the first user or the second user issues a wake-up command to control the operation of the corresponding intelligent voice device, the first intelligent voice device and the second intelligent voice device receive the wake-up command and perform feature extraction on the received wake-up command to obtain a wake-up identification feature corresponding to the wake-up command; it is determined whether the wake-up identification feature is consistent with the first identification feature or the second identification feature, and only when the intelligent voice device corresponding to the first identification feature or the second identification feature consistent with the wake-up identification feature is woken up, the intelligent voice device can perform an operation according to the wake-up command. This method for preventing accidental wake-up of intelligent voice devices in voice interaction compares the intelligent voice device's own recognition features with the wake-up recognition features extracted from the received wake-up command. Only when the wake-up recognition features match the intelligent voice device's recognition features can the intelligent voice device execute the corresponding operation according to the wake-up command. This prevents the intelligent voice device from executing operations due to non-wake-up commands during voice interaction, thus avoiding safety incidents. It solves the technical problem of existing intelligent voice devices having difficulty in recognizing whether the operation command is issued by the user, which leads to accidental operation. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1This is a flowchart illustrating the steps of the method for preventing accidental wake-up of a smart voice device with voice interaction as described in an embodiment of this application.

[0034] Figure 2 This is a framework diagram of the method for preventing accidental wake-up of a smart voice device with voice interaction as described in the embodiments of this application;

[0035] Figure 3 This is a framework diagram of a method for preventing accidental wake-up of a smart voice device with voice interaction, as described in another embodiment of this application.

[0036] Figure 4 This is a framework diagram of the intelligent voice device anti-false wake-up system for voice interaction described in the embodiments of this application. Detailed Implementation

[0037] To make the inventive objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0038] This application proposes a method, system, and device for preventing accidental wake-up of intelligent voice devices with voice interaction. During a video call between two intelligent voice devices, the first intelligent voice device automatically plays audio with a first recognition feature and simultaneously remotely sends the audio to the second intelligent voice device. Simultaneously, the second intelligent voice device performs the same operation and transmits its own second recognition feature to the first intelligent voice device. Alternatively, when a user of the first intelligent voice device wakes up their own device via voice, the second intelligent voice device will also receive the user's voice because the first and second intelligent voice devices are in a normal video call. The voice recognition modules of both the first and second intelligent voice devices each acquire existing voice wake-up commands, extract features, and further recognize the commands. When a "wake-up" command is recognized, it is determined whether the wake-up recognition feature in the audio contained in this "wake-up" frame belongs to the first or second intelligent voice device. Only when the wake-up recognition feature belongs to its own intelligent voice device will the intelligent voice device be woken up for voice control operations. This avoids the problem of accidental operation of the intelligent voice device by the other party during a call. This technology addresses the technical problem of existing smart voice devices struggling to identify whether an operation command was issued by the user, leading to potential misoperations.

[0039] Example 1:

[0040] Figure 1This is a flowchart illustrating the steps of the method for preventing accidental wake-up of a smart voice device with voice interaction as described in an embodiment of this application.

[0041] like Figure 1 As shown, this application provides a method for preventing accidental wake-up of intelligent voice devices, applied to video and voice interaction between at least two intelligent voice devices corresponding to users. Each intelligent voice device is equipped with recognition features for identification. The method for preventing accidental wake-up of intelligent voice devices includes the following steps:

[0042] S10. Obtain the first recognition feature of the first intelligent voice device and the first user corresponding to the first recognition feature, and obtain the second recognition feature of the second intelligent voice device and the second user corresponding to the second recognition feature.

[0043] It should be noted that in step S10, each smart voice device has its own wake-up recognition features. In this embodiment, the recognition features of each pair of smart voice devices and the user using the device are obtained to provide basic data for subsequent wake-up. The smart voice devices can be smart office products such as smart speakers, televisions, voice notepads, conference stenographers, and smart set-top boxes; or smart education products such as smart voice reading lamps, smart voice learning machines, and smart voice translation pens.

[0044] S20. When the first user and the second user interact with each other through their respective smart voice devices, the first user or the second user issues a wake-up command to control the operation of the corresponding smart voice device. The first smart voice device and the second smart voice device receive the wake-up command and extract features from the received wake-up command to obtain wake-up recognition features corresponding to the wake-up command.

[0045] It should be noted that in step S20, after the first user and the second user conduct a video call through their respective smart voice devices, the first user or the second user issues a voice control command to their own smart voice device. In this embodiment, after each smart voice device receives the wake-up command voice, it performs feature extraction on the wake-up command voice to obtain the wake-up recognition feature corresponding to the wake-up command voice. This provides a basis for the smart voice device to identify whether the wake-up command was issued to itself, avoiding the smart voice device directly executing the wake-up command and resulting in no operation. The first recognition feature, the second recognition feature, and the wake-up recognition feature can all be white noise audio, which contains feature values ​​that can identify the audio. The wake-up command includes a wake-up command with the name of the smart voice device and a control command for the smart voice device to perform an operation.

[0046] S30. Determine whether the wake-up recognition feature matches the first or second recognition feature. Only when the smart voice device corresponding to the first or second recognition feature that matches the wake-up recognition feature is woken up can the smart voice device perform operations according to the wake-up command.

[0047] It should be noted that in step S30, each smart voice device compares the extracted wake-up recognition features with its own recognition features to determine whether the wake-up recognition features are consistent with the smart voice device's recognition features. Only when the wake-up recognition features are consistent with the smart voice device's recognition features can the smart voice device execute the corresponding operation according to the wake-up command. This prevents users from performing operations due to non-wake-up commands during voice interaction, thus avoiding safety incidents.

[0048] The method for preventing accidental wake-up of a smart voice device with voice interaction provided in this application includes obtaining a first identification feature of a first smart voice device and a first user corresponding to the first identification feature, and obtaining a second identification feature of a second smart voice device and a second user corresponding to the second identification feature; when the first user and the second user interact with each other through their respective smart voice devices, the first user or the second user issues a wake-up command to control the operation of the corresponding smart voice device; the first smart voice device and the second smart voice device receive the wake-up command and perform feature extraction on the received wake-up command to obtain a wake-up identification feature corresponding to the wake-up command; it is determined whether the wake-up identification feature is consistent with the first identification feature or the second identification feature; only when the smart voice device corresponding to the first identification feature or the second identification feature that is consistent with the wake-up identification feature is woken up can the smart voice device perform an operation according to the wake-up command. This method for preventing accidental wake-up of intelligent voice devices in voice interaction compares the intelligent voice device's own recognition features with the wake-up recognition features extracted from the received wake-up command. Only when the wake-up recognition features match the intelligent voice device's recognition features can the intelligent voice device execute the corresponding operation according to the wake-up command. This prevents the intelligent voice device from executing operations due to non-wake-up commands during voice interaction, thus avoiding safety incidents. It solves the technical problem of existing intelligent voice devices having difficulty in recognizing whether the operation command is issued by the user, which leads to accidental operation.

[0049] Figure 2 This is a framework diagram of the method for preventing accidental wake-up of a smart voice device with voice interaction as described in the embodiments of this application.

[0050] In one embodiment of this application, the method for preventing accidental wake-up of the intelligent voice device for voice interaction includes: if the intelligent voice device corresponding to a first identification feature or a second identification feature that is inconsistent with the wake-up identification feature will not be woken up, the intelligent voice device will remain in place.

[0051] It should be noted that if the wake-up recognition feature is inconsistent with the recognition feature of the smart voice device, the smart voice device will remain silent and will not perform any corresponding operations based on the wake-up command. This avoids the occurrence of security incidents caused by performing operations due to an incorrect wake-up command. In this embodiment, the method for preventing accidental wake-up of smart voice devices in voice interaction can be applied to one-to-one interaction between two smart voice devices, or to multiple users using a single smart voice device simultaneously, since the same smart voice device is being used, meaning the white noise belongs to the same smart voice device.

[0052] In one embodiment of this application, after a first user and a second user interact via voice through their respective smart voice devices, the method for preventing accidental wake-up of the smart voice devices during the voice interaction includes: transmitting a first identification feature to the second smart voice device along with the video call and transmitting a second identification feature to the first smart voice device along with the video call.

[0053] It should be noted that, in this voice interaction method for preventing accidental wake-up of smart voice devices, after the first user and the second user engage in voice interaction through their respective smart voice devices, their respective smart voice devices will play audio with recognition characteristics, which is transmitted to each other's smart voice devices along with the first user and the second user, ensuring normal video call interaction between the first user and the second user through their respective smart voice devices. For example, audio with recognition characteristics from the first smart voice device is transmitted to and received by the second smart voice device along with the first user's call; similarly, audio with recognition characteristics from the second smart voice device is transmitted to and received by the first smart voice device along with the second user's call.

[0054] In this embodiment of the application, a video call interaction between a first intelligent voice device A and a second intelligent voice device B is used as an example to illustrate the method for preventing accidental wake-up of intelligent voice devices in this voice interaction. Figure 2As shown, the first intelligent voice device A has its own audio with a first recognition feature (such as white noise 1 feature), and the second intelligent voice device B has its own audio with a second recognition feature (such as white noise 2 feature). After video call interaction, the first intelligent voice device A and the second intelligent voice device B transmit their respective recognition features to each other. The first user (i.e., user A) sends a wake-up command with white noise 1 audio to the first intelligent voice device A. This wake-up command is also transmitted to the second intelligent voice device B during the call between the first user and the second user. Both the first intelligent voice device A and the second intelligent voice device B extract features from the received wake-up command with white noise 1 audio to obtain wake-up recognition features. The wake-up recognition features of the first intelligent voice device A and the second intelligent voice device B are compared with their respective recognition features to see if they match. If the wake-up recognition features match the white noise 1 feature of the first intelligent voice device A, the first intelligent voice device A is woken up and performs the operation corresponding to the wake-up command. If the wake-up recognition features do not match the white noise 2 feature of the second intelligent voice device B, the second intelligent voice device B remains silent and does not perform the operation according to the wake-up command.

[0055] In one embodiment of this application, the smart voice device is provided with a feature generation module, a feature activation module, a feature playback module, a video and voice call module, a voice recognition module, and a feature matching module.

[0056] In this embodiment of the application, the feature generation module can be used to generate recognition audio with recognition features corresponding to the user of the smart voice device.

[0057] It should be noted that the audio to be recognized can be white noise. The recognition audio with recognition features generated by the feature generation module on the smart voice device can be generated locally by the smart voice device or downloaded as white noise audio with recognition feature values ​​corresponding to the user and having a certain degree of distinguishability. The white noise audio can be adjusted to a frequency that is inaudible to the human ear to avoid affecting the quality of the call voice; at the same time, it ensures that the white noise of the communication device has different feature values.

[0058] In this embodiment, the feature activation module is used to activate the feature playback module when the user speaks to the smart voice device.

[0059] It should be noted that during video and voice calls using a smart voice device, the feature playback module is activated when the user speaks to their own smart voice device. If the user's smart voice device is quiet or in a silent state, the feature playback module is deactivated.

[0060] In this embodiment, the feature playback module is used to play the recognition audio.

[0061] It should be noted that the feature playback module controls the smart voice device to play the recognition audio generated by the feature generation module, so that the recognition audio and the user's voice are stored by the smart voice device together.

[0062] In this embodiment, the video and voice call module is used for remote video call interaction.

[0063] It should be noted that intelligent voice devices have remote video call functionality through video and voice call modules, enabling remote exchange and storage of audio with recognition features.

[0064] In this embodiment, the speech recognition module identifies and extracts wake-up recognition features from the video speech of the smart voice device receiving a wake-up command.

[0065] It should be noted that the intelligent voice device uses a voice recognition module to recognize the wake-up command received by the intelligent voice device with recognition characteristics to perform operations such as device wake-up; and extracts all features contained in the audio (including white noise features).

[0066] In this embodiment, the feature matching module is used to match the wake-up recognition feature with the recognition feature of the smart voice device to determine whether to wake up or remain silent based on the wake-up command received by the smart voice device.

[0067] It should be noted that intelligent voice devices use a feature matching module to query and match which device the wake-up command belongs to based on wake-up recognition features, thus determining whether their own device is "wake-up" or "silent".

[0068] In this embodiment, the intelligent voice device achieves anti-false wake-up functionality through a feature generation module, a feature activation module, a feature playback module, a video and voice call module, a voice recognition module, and a feature matching module. This results in lower costs and eliminates the need for additional hardware or equipment. It also enriches the application scenarios of the intelligent voice device, supporting various call modes such as one-to-one and many-to-many calls, as well as complex scenarios where multiple users operate a single device. This intelligent voice device is more practical; regardless of network fluctuations or other uncontrollable factors causing issues such as intermittent sound, sound delays, sound without image, or audio-visual asynchrony, the performance remains unaffected because white noise is mixed with the human voice before transmission. The intelligent voice device is also more convenient to use, eliminating the need to set different wake-up words for each device individually.

[0069] Figure 3 This is a framework diagram of a method for preventing accidental wake-up of a smart voice device with voice interaction, as described in another embodiment of this application.

[0070] In this application embodiment, the method for preventing accidental wake-up of a smart voice device with anti-accidental wake-up in this voice interaction smart voice device can be illustrated by the following examples, such as... Figure 3As shown, user Xiao Wang is having a video call with smart speaker device B at Xiao Li's home via smart speaker device A. Both smart speakers are paired with other devices such as smart sockets and smart lights, and are controlled through their respective smart speakers. During the video call, Xiao Wang wakes up smart speaker device A with the voice command "Xiao Yi Xiao Yi" and then uses the voice command "Disconnect the socket power" to turn off the power to the smart socket. At the same time, smart speaker device B at Xiao Li's home also receives the "Xiao Yi Xiao Yi" wake-up voice command. After determining that smart speaker device B is in a "silent" state, it ignores this wake-up command, thus preventing its own smart socket from being accidentally "powered off". The process is as follows: The voice recognition module of the smart speaker continuously recognizes the sound received by the smart speaker and waits for the device to be woken up; the smart speaker obtains the white noise audio in the wake-up command through the feature generation module; Xiao Wang initiates a video call request with Xiao Li through smart speaker A. After Xiao Li answers, the video and voice call module remotely exchanges white noise audio representing the two smart speaker devices with their own recognition features; the smart speaker supports smart speaker A and smart speaker B to start a video call through the video and voice call module; during the call between Xiao Wang and Xiao Li, when the user speaks to their own smart speaker, its feature activation module activates the feature playback module; if the user is silent, the feature playback module is turned off; when Xiao Wang says "Xiao Yi Xiao Yi" to his smart speaker A, the feature playback module is activated, controlling the smart speaker. The speaker device plays white noise 1 audio, which is received and stored by smart speaker device A along with Xiao Wang's voice. At the same time, Xiao Li's smart speaker device B also plays the "Xiao Yi Xiao Yi" audio transmitted from smart speaker device A. The voice recognition modules of smart speaker devices A and B extract audio features (including white noise features) from the audio received by the smart speaker devices. Both smart speaker devices recognize the "Xiao Yi Xiao Yi" wake-up voice command. After Xiao Wang's smart speaker device A recognizes the "Xiao Yi Xiao Yi" wake-up command, it further extracts the white noise features of this audio segment. The feature matching module of smart speaker device A queries and matches the white noise 1 based on this feature value. Since white noise 1 belongs to its own smart speaker device A, smart speaker device A "wakes up" and waits for the voice command to be issued. Upon hearing that smart speaker device A had been activated, Xiao Wang continued to issue a voice command to the device, "Disconnect the power to the socket." The voice recognition module of smart speaker device A recognized the command and executed the "power off operation for the smart socket." At the same time, Xiao Li's smart speaker device B also recognized the "Xiao Yi Xiao Yi" wake-up voice command and further extracted the white noise feature of this audio segment. The feature matching module of smart speaker device B used this feature value to query and match white noise 1. White noise 1 belongs to its own smart speaker device A, not its own smart speaker device B, and determined that the command was not issued to smart speaker device B itself.Therefore, smart speaker device B remains "silent." When user Xiao Wang says "disconnect the power" via voice, although smart speaker device B also receives the sound, it will not perform any operation because the device is not awakened. If Xiao Li wakes up smart speaker device B, smart speaker device B recognizes the "wake up" voice command, extracts the white noise features contained in the frame containing the "wake up" command, and matches the white noise to itself using its own feature matching module. Smart speaker device B is then "awakened" and awaits the voice command. Although Xiao Wang's smart speaker device A also recognizes the "wake up" command, it finds that the white noise does not belong to itself when further determining its attribution, and thus remains "silent." If, at this time, Xiao Wang and Xiao Li are speaking simultaneously, resulting in two sets of white noise features, then both smart speakers A and B will remain or return to a "silent" state.

[0071] In one embodiment of this application, the method for preventing accidental wake-up of the intelligent voice device in voice interaction includes: if a first user and a second user simultaneously issue a wake-up command to control the operation of the corresponding intelligent voice device, neither the first intelligent voice device nor the second intelligent voice device will be woken up.

[0072] Example 2:

[0073] Figure 4 This is a framework diagram of the intelligent voice device anti-false wake-up system for voice interaction described in the embodiments of this application.

[0074] like Figure 4 As shown, this application also provides a voice interaction-based intelligent voice device anti-mistake wake-up system, which is applied to video voice interaction between at least two intelligent voice devices corresponding to users. Each intelligent voice device is equipped with recognition features for identification. The intelligent voice device anti-mistake wake-up system includes: a data acquisition unit 10, an interaction extraction unit 20, and a recognition wake-up unit 30.

[0075] The data acquisition unit 10 is used to acquire a first recognition feature of a first intelligent voice device and a first user corresponding to the first recognition feature, and to acquire a second recognition feature of a second intelligent voice device and a second user corresponding to the second recognition feature.

[0076] The interaction extraction unit 20 is used to enable the first user and the second user to interact via voice, and the first user or the second user to issue a wake-up command to control the operation of the corresponding smart voice device. The first smart voice device and the second smart voice device receive the wake-up command and perform feature extraction on the received wake-up command to obtain the wake-up recognition feature corresponding to the wake-up command.

[0077] The recognition and wake-up unit 30 is used to determine whether it is consistent with the first recognition feature or the second recognition feature based on the wake-up recognition feature. Only when the smart voice device corresponding to the first recognition feature or the second recognition feature that is consistent with the wake-up recognition feature is woken up can the smart voice device perform operations according to the wake-up command.

[0078] Among them, the first identification feature, the second identification feature, and the wake-up identification feature are all white noise audio, and the white noise audio contains feature values ​​that identify the audio.

[0079] In this embodiment, the wake-up recognition unit 30 is further configured to remain silent if the smart voice device corresponding to the first or second recognition feature that is inconsistent with the wake-up recognition feature will not be woken up.

[0080] In this embodiment of the application, the voice interaction intelligent voice device anti-mistake wake-up system includes a feature transmission unit. The feature transmission unit is used to transmit a first recognition feature to a second intelligent voice device along with a video call and to transmit a second recognition feature to a first intelligent voice device along with a video call after the first user and the second user have interacted via voice.

[0081] It should be noted that the content of the modules in Embodiment 2 corresponds to the steps in the method of Embodiment 1. The content of the steps in the method of Embodiment 1 has been described in detail in Embodiment 1, and the content of the modules in the system will not be described again in Embodiment 2.

[0082] Example 3:

[0083] This application also provides a terminal device, including a processor and a memory;

[0084] Memory is used to store program code and transfer the program code to the processor;

[0085] The processor is used to execute the above-mentioned method for preventing accidental wake-up of intelligent voice devices based on instructions in the program code.

[0086] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0087] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0089] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0090] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0091] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for false wake-up prevention of a smart voice device in voice interaction, applied to video voice interaction of at least two smart voice devices corresponding to a user, characterized in that, Each of the intelligent voice devices is provided with an identification feature for identification, and the intelligent voice device false wake-up prevention method comprises the following steps: obtaining a first identification feature of a first intelligent voice device and a first user corresponding to the first identification feature, and obtaining a second identification feature of a second intelligent voice device and a second user corresponding to the second identification feature; when the first user and the second user perform voice interaction through the respective intelligent voice devices, the first user or the second user issues a wake-up instruction for controlling the operation of the corresponding intelligent voice device, the first intelligent voice device and the second intelligent voice device receive the wake-up instruction and perform feature extraction on the received wake-up instruction to obtain a wake-up identification feature corresponding to the wake-up instruction; whether the first identification feature or the second identification feature is consistent with the wake-up identification feature is determined, and only the corresponding intelligent voice device with the first identification feature or the second identification feature consistent with the wake-up identification feature is woken up, and the intelligent voice device can execute the operation according to the wake-up instruction. 2.The method of claim 1, wherein, After the first user and the second user perform voice interaction through the respective intelligent voice devices, the intelligent voice device false wake-up prevention method of the voice interaction comprises the following steps: transmitting the first identification feature to the second intelligent voice device through a video call and transmitting the second identification feature to the first intelligent voice device through a video call. 3.The method of claim 1, wherein, comprises: if the corresponding intelligent voice device with the first identification feature or the second identification feature inconsistent with the wake-up identification feature is not woken up, the intelligent voice device remains silent. 4.The method of claim 1, wherein, The first identification feature, the second identification feature and the wake-up identification feature are all white noise audios, and the white noise audios contain feature values for identifying the audios. 5.The method of claim 1, wherein, The intelligent voice device is provided with a feature generation module, a feature activation module, a feature playing module, a video voice call module, a voice recognition module and a feature matching module; The feature generation module is configured to generate an identification audio with an identification feature corresponding to a user of the intelligent voice device; The feature activation module is configured to activate the feature playing module when the user speaks to the intelligent voice device; The feature playing module is configured to play the identification audio; The video voice call module is configured to perform remote video call interaction; The voice recognition module is configured to recognize and extract a wake-up identification feature in the video voice received by the intelligent voice device; The feature matching module is configured to match the wake-up identification feature with the identification feature of the intelligent voice device, and determine whether to wake up or remain silent according to the wake-up instruction received by the intelligent voice device.

6. A false wake-up prevention system for voice interaction of intelligent voice devices, applied to video voice interaction of at least two intelligent voice devices corresponding to users, characterized in that, Each of the intelligent voice devices is provided with an identification feature for identification, and the intelligent voice device false wake-up prevention system comprises a data acquisition unit, an interaction extraction unit and an identification wake-up unit; The data acquisition unit is configured to acquire a first identification feature of a first smart voice device and a first user corresponding to the first identification feature, and acquire a second identification feature of a second smart voice device and a second user corresponding to the second identification feature. The interaction extraction unit is configured to extract, through voice interaction between the first user and the second user, a wake-up instruction for controlling operation of a corresponding smart voice device, and extract, by the first smart voice device and the second smart voice device, a wake-up identification feature corresponding to the wake-up instruction. The wake-up identification unit is configured to determine, according to the wake-up identification feature, whether the first identification feature or the second identification feature is consistent, and only the corresponding smart voice device with the first identification feature or the second identification feature consistent with the wake-up identification feature can be woken up to perform operation according to the wake-up instruction.

7. The smart voice device false wake-up prevention system for voice interaction according to claim 6, wherein, The wake-up identification unit is further configured to keep silent the corresponding smart voice device with the first identification feature or the second identification feature inconsistent with the wake-up identification feature.

8. The smart voice device false wake up system for voice interaction according to claim 6, wherein, The feature transmission unit is configured to transmit, after the voice interaction between the first user and the second user, the first identification feature to the second smart voice device and the second identification feature to the first smart voice device through a video call.

9. The smart voice device false wake up system for voice interaction according to claim 6, wherein, The first identification feature, the second identification feature, and the wake-up identification feature are all white noise audios, and the white noise audios contain feature values for identifying the audios.

10. A terminal device, comprising: The processor and the memory are included. The memory is configured to store program code and transmit the program code to the processor. The processor is configured to execute the voice interaction smart voice device false wake-up prevention method according to instructions in the program code.

Citation Information

Patent Citations

  • Voice control method and mobile terminal

    CN106453859A

  • Method, device and system for preventing mistaken wake-up of voice interaction equipment and application method

    CN109473110A