Voice interaction anti-interference method, electronic device, storage medium and program product
Through the collaborative work of terminal devices and cloud servers within the distributed collaborative group, the problem of mutual interference when multiple devices wake up is solved, and the success rate of recognition of voice commands is improved.
Patent Information
- Application Number
- CN202111501502.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-12-09
AI Technical Summary
When multiple smart terminal devices are awakened in distributed manner, synchronization failure caused by poor network performance or hardware performance differences will result in multiple devices awakening at the same time, causing mutual interference and reducing the success rate of voice command recognition.
By receiving the audio signal and identifying the result on the first terminal device in the distributed collaborative group, it sends it to the cloud server. The cloud server issues a rejection identification instruction based on the identification result, so that the first terminal device does not respond to the audio signal of the second terminal device and avoids mutual interference.
The recognition success rate of smart terminal devices when receiving user voice commands is improved, and mutual interference between multiple devices is avoided.
Smart Images

Figure CN114283800B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic technology, and in particular to a voice interaction anti-interference method, electronic equipment, storage medium and program product. Background Art
[0002] With the continuous iteration of intelligent voice technology, its recognition rate has continued to improve. Voice recognition technology has attracted the attention of many IoT (Internet of Things) manufacturers. More and more home devices are using intelligent voice technology, and the demand and use cases for voice interaction are constantly expanding. Especially in the home, it provides a convenient and efficient way for people to interact with various smart devices. It not only frees up hands, but also extends the distance of human-computer communication, making human-computer interaction more natural.
[0003] In daily use, there may be multiple smart terminal devices in the same physical space, and distributed wake-up is adopted among these multiple smart terminal devices. Distributed wake-up means that only the device closest to the user is woken up, while other devices are not woken up. However, this distributed wake-up algorithm relies on the network. Once the network performance deteriorates, it is very likely that two or more devices will wake up at the same time. Specifically, the so-called simultaneous wake-up does not necessarily mean that multiple devices will wake up at the same time. In this case, the response word of the device that wakes up later may be recognized and read by the device that wakes up earlier. At the same time, the cloud responds to the response word of the device that wakes up later. At this time, the user has not actually interacted with the device. This causes interference between multiple smart terminal devices when receiving user voice commands, thereby reducing the success rate of the smart terminal devices in recognizing voice commands. Summary of the Invention
[0004] The present invention provides a voice interaction anti-interference method, electronic device, storage medium and program product, which are used to solve the defect in the prior art that when receiving user voice commands, multiple intelligent terminal devices interfere with each other, thereby reducing the success rate of the intelligent terminal devices in recognizing voice commands.
[0005] The present invention provides a voice interaction anti-interference method, which is applied to a first terminal device located in a distributed collaborative group, wherein the second terminal device and the first terminal device in the distributed collaborative group are both configured with a voice processing module, and the method includes: according to a voice wake-up instruction, starting a sound pickup operation and receiving a first audio signal; recognizing the received first audio signal to obtain a recognition result, and sending the recognition result to a cloud server; when the recognition result indicates that the first audio signal includes a target response word, when the recognition result indicates that the first audio signal includes a target response word, receiving a rejection recognition instruction made by the cloud server based on the recognition result, and not making an audio response to the first audio signal based on the rejection recognition instruction.
[0006] The present invention also provides a voice interaction anti-interference method, which is applied to a cloud server, including: receiving a recognition result sent by a first terminal device, wherein the recognition result is obtained based on the first terminal device recognizing a first audio signal; when the recognition result indicates that the first audio signal includes a target response word, sending a rejection recognition instruction to the first terminal device, so that the first terminal device does not make an audio response to the first audio signal based on the rejection recognition instruction.
[0007] The present invention also provides a voice interaction anti-interference device, including: a first receiving module, used to start the pickup operation and receive a first audio signal according to a voice wake-up instruction; a sending module, used to identify the received first audio signal to obtain a recognition result, and send the recognition result to a cloud server; a processing module, used to receive a rejection recognition instruction made by the cloud server based on the recognition result when the recognition result indicates that the first audio signal includes a target response word, and not make an audio response to the first audio signal based on the rejection recognition instruction.
[0008] The present invention also provides a voice interaction anti-interference device, including: a second receiving module, used to receive the recognition result sent by the first terminal device, and the recognition result is obtained based on the first terminal device's recognition of the first audio signal; a second sending module, used to send a rejection recognition instruction to the first terminal device when the recognition result indicates that the first audio signal includes a target response word, and the rejection recognition instruction is used to instruct the first terminal device not to make an audio response to the first audio signal.
[0009] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-described voice interaction anti-interference methods are implemented.
[0010] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described voice interaction anti-interference methods.
[0011] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned voice interaction anti-interference methods.
[0012] The voice interaction anti-interference method, electronic device, storage medium and program product provided by the present invention start the sound pickup operation according to the voice wake-up instruction, receive the first audio signal, recognize the received first audio signal to obtain a recognition result, and send the recognition result to the cloud server; the first terminal device will receive the rejection recognition instruction made by the cloud server based on the recognition result. As long as the rejection recognition instruction is received, the first terminal device will not make an audio response to the first audio signal including the target response word. Therefore, when the voice interaction functions of the first terminal device and the second terminal device are both activated, there is no mutual interference between the first terminal device and the second terminal device when receiving the user's voice instruction, thereby improving the success rate of the smart terminal device in recognizing the voice instruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 This is one of the flow charts of the voice interaction anti-interference method provided by the present invention;
[0015] Figure 2 This is the second flow chart of the voice interaction anti-interference method provided by the present invention;
[0016] Figure 3 This is the third flow chart of the voice interaction anti-interference method provided by the present invention;
[0017] Figure 4 This is the fourth flow chart of the voice interaction anti-interference method provided by the present invention;
[0018] Figure 5 This is the fifth flow chart of the voice interaction anti-interference method provided by the present invention;
[0019] Figure 6 This is the sixth flow chart of the voice interaction anti-interference method provided by the present invention;
[0020] Figure 7 This is the seventh flow chart of the voice interaction anti-interference method provided by the present invention;
[0021] Figure 8 This is the eighth flow chart of the voice interaction anti-interference method provided by the present invention;
[0022] Figure 9 A timing diagram showing a rejection recognition when a first terminal device and a second terminal device are woken up simultaneously in the present invention;
[0023] Figure 10 This is one of the structural diagrams of the voice interaction anti-interference device provided by the present invention;
[0024] Figure 11 This is the second structural diagram of the voice interaction anti-interference device provided by the present invention;
[0025] Figure 12 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0026] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0027] In the home, due to the diverse range of smart devices, it's common to have multiple smart voice interaction devices in a home, and distributed wakeup is a growing trend. However, distributed wakeup can fail, for example when network performance is affected by factors like the electromagnetic environment and climatic conditions, leading to multiple smart devices waking up simultaneously. Alternatively, performance differences between multiple smart devices can cause score calculations to take varying amounts of time, preventing them from completing the calculations synchronously. This can lead to decision failures and multiple devices waking up simultaneously.
[0028] In the above two scenarios, multiple smart terminal devices will enter the voice interaction process at the same time. On a macro level, multiple smart terminal devices will respond to users at the same time. On a micro level, they actually respond in sequence. The smart terminal device that responds first will input the response words of the smart terminal device that responds later as normal interaction content into the voice recognition process of the smart terminal device that responds first, and broadcast the corresponding reply content to the smart terminal device that responds later. When receiving user voice commands, there is mutual interference between the first terminal device and the second terminal device, thereby reducing the success rate of the smart terminal device in recognizing voice commands.
[0029] In view of this, the present invention provides a voice interaction anti-interference method, electronic device, storage medium and program product. The inventive concept of the present invention is that, in the case of distributed wake-up failure, that is, when the voice interaction functions of the first terminal device and the second terminal device are both activated, when the first terminal device receives the first audio signal sent by the second terminal device, it recognizes the received first audio signal to obtain a recognition result, and sends the recognition result to the cloud server; the first terminal device will receive a rejection recognition instruction made by the cloud server based on the recognition result, so that based on the rejection recognition instruction, the first terminal device does not make an audio response to the first audio signal sent by the second terminal device. Therefore, when the voice interaction functions of the first terminal device and the second terminal device are both activated, when receiving user voice commands, there is no mutual interference between the first terminal device and the second terminal device, thereby improving the success rate of the smart terminal device in recognizing voice commands.
[0030] The following combination Figure 1-Figure 7 The voice interaction anti-interference method of the present invention is described.
[0031] Please refer to Figure 1 The voice interaction anti-interference method of an embodiment of the present invention is applied to a first terminal device in a distributed collaborative group, wherein both a second terminal device and the first terminal device in the distributed collaborative group are configured with a voice processing module, including:
[0032] Step 100a: according to the voice wake-up instruction, start the sound pickup operation and receive the first audio signal;
[0033] It should be noted that the second terminal device and the first terminal device in the distributed collaborative group are both equipped with a voice processing module, and both can recognize the user's voice instructions and make corresponding responses to the user's voice instructions. The first terminal device and the second terminal device in the distributed collaborative group refer to two different terminal devices with the same wake-up words in the same physical space. For example, the wake-up words for the first terminal device and the second terminal device are both "Xiao A Xiao A". When the user calls out the wake-up word "Xiao A Xiao A" to the first terminal device and the second terminal device in the same physical space, when the competition fails due to poor network signals or hardware performance differences, the first terminal device first receives "Xiao A Xiao A" to wake up, and the second terminal device wakes up after the first terminal device.
[0034] For details, please refer to Figure 2 Step 100a: according to the voice wake-up instruction, start the sound pickup operation and receive the first audio signal, including:
[0035] Step 110a: according to the voice wake-up instruction, start the sound pickup operation and obtain the target response word sent by the second terminal device as the first audio signal;
[0036] The first terminal device starts a voice pickup operation in response to the user's voice wake-up command, and obtains the target response word sent by the second terminal device as the first audio signal. The target response word can be a response word such as "What do you want to say?", "Yes, I'm here?", "Here I am?", "Master, what are your instructions?", etc., indicating that the second terminal device is responding to the user's voice wake-up command.
[0037] Among them, when the network performance in the home is affected by multiple factors such as the electromagnetic environment and climatic conditions; or due to performance differences between multiple smart terminal devices, resulting in differences in the time consumption of score calculations, multiple smart terminal devices cannot complete the calculation synchronously, resulting in decision failure, distributed wake-up failure occurs between the first terminal device and the second terminal device. At this time, after the user calls out the wake-up word, the first terminal device and the second terminal device "wake up at the same time", and the first terminal device and the second terminal device both start the voice interaction function. It can be understood that although the first terminal device and the second terminal device wake up at the same time at the macro level. However, at the micro level, there is still a first and a later. This embodiment is explained by waking up the first terminal device first and waking up the second terminal device later.
[0038] When the first terminal device wakes up first, the first terminal device starts picking up audio first. Since the second terminal device wakes up later, the first audio signal sent by the second terminal device to the user will be received by the first terminal device. The first audio signal can be various response words used to respond to the user's wake-up. For example, the response word "What are your instructions?" sent by the second terminal device to the user will be picked up by the first terminal device in the pickup state. This embodiment of the application does not limit the response word.
[0039] Step 200a: Recognize the received first audio signal to obtain a recognition result, and send the recognition result to a cloud server;
[0040] The first terminal device recognizes the received first audio signal, obtains a recognition result, and sends the recognition result to the cloud server.
[0041] Step 300a: When the recognition result indicates that the first audio signal includes a target response word, receiving a rejection recognition instruction made by the cloud server based on the recognition result, and not making an audio response to the first audio signal based on the rejection recognition instruction.
[0042] After the cloud server receives the recognition result (i.e., the response word of the second terminal device), it makes a judgment based on the recognition result. When it is determined that the recognition result is a response word that refuses recognition, the first terminal device receives a rejection recognition instruction sent by the cloud server. Therefore, even if the first terminal device receives the response word of the second terminal device, it will not respond to the response word of the second terminal device. When the voice interaction functions of the first terminal device and the second terminal device are both activated, there is no mutual interference between the first terminal device and the second terminal device when receiving the user's voice command, thereby improving the success rate of the smart terminal device in recognizing the voice command.
[0043] For example, when the first terminal device receives the response word "What are your instructions?" sent by the second terminal device, it sends "What are your instructions?" to the cloud server. The cloud server issues a rejection recognition instruction to the first terminal device based on the response word "What are your instructions?" Based on the rejection recognition instruction, the first terminal device will not make an audio response to the response word "What are your instructions?". This avoids the situation where the first and second terminal devices interfere with each other, resulting in a reduced success rate of the smart terminal device in recognizing voice commands.
[0044] The following is an example of a typical example. Figure 9 , the first terminal device receives the wake-up message from "Xiao A Xiao A" and wakes up first and enters the voice pickup state. Due to poor network conditions, the first terminal device does not receive the wake-up decision from the cloud server and times out, waking up and responding with "What do you want?". In this way, the first terminal device wakes up first and enters the voice pickup state, and device 2 wakes up after a timeout and responds with "What do you want?". At this time, the first terminal device in the voice pickup state will pick up "What do you want?" and send it to the cloud server. The cloud server will recognize that this is a response word, not the actual content the user wants to interact with, and send a rejection recognition command to the first terminal device, thereby continuing to pick up voice. Until the user speaks a normal voice interaction command, the cloud server will send the normal recognition result to the first and second terminal devices.
[0045] Figure 9 SDK stands for software development kit, which refers to a software development kit, that is, a software development kit for the first terminal device and the second terminal device to recognize the audio signal into a response word.
[0046] It's worth noting that the first audio signal can also be a command issued by the user. When the user speaks a target response word such as "What do you want to say?", "Yes, I'm here," "Here I am," or "Master, what instructions do you have?", the first terminal device recognizes the first audio signal, obtains a recognition result, and sends the recognition result to the cloud server. The first terminal device then receives a recognition rejection instruction issued by the cloud server based on the recognition result.
[0047] The voice interaction anti-interference method provided in an embodiment of the present invention activates a sound pickup operation based on a voice wake-up command, receives a first audio signal, recognizes the received first audio signal, obtains a recognition result, and sends the recognition result to a cloud server. A first terminal device then receives a rejection recognition instruction issued by the cloud server based on the recognition result. Based on the rejection recognition instruction, the first terminal device does not provide an audio response to the first audio signal including the target response word. Therefore, when the voice interaction functions of both the first and second terminal devices are activated, there is no mutual interference between the first and second terminal devices when receiving user voice commands, thereby improving the success rate of the intelligent terminal device in recognizing voice commands.
[0048] In some possible embodiments, please refer to Figure 2 , after step 300a, further comprising:
[0049] Step 310a: Continue the sound pickup operation.
[0050] After the first terminal device does not make an audio response to the first audio signal based on the rejection recognition instruction, it continues to pick up the sound, thereby re-enabling the service of making an audio response to the user's voice interaction instruction, ensuring that the first terminal device can normally recognize the user's voice interaction instruction.
[0051] In some possible embodiments, please refer to Figure 3 , after step 300a, further comprising:
[0052] Step 320a: Determine, based on the current first conversation mode of the first terminal device, whether the current moment is within the target interaction time range of the current conversation mode;
[0053] The first terminal device's current first conversation mode is configured with a fixed interaction timer, which has a target interaction time range. This target interaction time range specifies the time period during which a user waits for a voice command to be issued during a valid voice interaction. For example, the target interaction time range can be set to 1 second. After the first terminal device receives a rejection recognition instruction from the cloud server, it determines whether the current moment is within the 1-second target interaction time range and then performs different operations based on the situation.
[0054] Step 330a: When the current moment is within the target interaction time range, switch the current first dialogue mode to the second dialogue mode, and continue the sound pickup operation based on the second dialogue mode.
[0055] If the first terminal device is currently within the target interaction time range, it continues to pick up audio. For example, if the current time is within the target interaction time range of 1 second, the current first conversation mode is switched to the second conversation mode. The first terminal device starts the next round of voice recognition in the second conversation mode. If the user issues a voice interaction command, the first terminal device sends the interaction audio to the cloud server. The cloud server returns a valid response result. If the first terminal device determines that the response result is valid, it ends the current conversation and broadcasts the response result.
[0056] The second dialogue mode is, for example, a continuous dialogue mode. The first terminal device continues to start the service of providing an audio response to the user's voice interaction instruction based on the second dialogue mode, ensuring that the first terminal device can normally recognize the user's voice interaction instruction.
[0057] In some other feasible embodiments, after step 320a, the following steps are further included:
[0058] Step 340a: When the current moment is not within the target interaction time range, a second audio signal is issued to instruct the user to issue a voice command.
[0059] When the current moment is not within the target interaction time range, the first terminal device sends a second audio signal for instructing the user to issue a voice command. The second audio signal can be various audio signals indicating that the user should re-issue the voice command. For example, the first terminal device receives a rejection recognition instruction sent by the cloud server, and determines whether the current moment exceeds the target interaction time range of 10s. If it exceeds 10s, the first terminal device broadcasts an audio signal of "didn't hear clearly" to inform the user to re-issue the voice command. Therefore, after the first terminal device refuses to make an audio response to the answer word of the second terminal device, when the current moment is not within the target interaction time range, the first terminal device sends a second audio signal to re-issue the voice command. The first terminal device re-prepare the service for audio response to the user's voice interaction instruction to ensure that the first terminal device can normally recognize the user's voice interaction instruction.
[0060] For other possible embodiments, please refer to Figure 4 , after step 310a, further comprising:
[0061] Step 350a: Acquire a second audio signal according to a user voice instruction;
[0062] After the first terminal device continues to pick up sound, it acquires a second audio signal based on the user's voice command. The second audio signal includes a keyword associated with the terminal device the user wishes to wake up. For example, if the first terminal device is a television and the second terminal device is a refrigerator, the user speaks the voice command "show me where the milk is." Since a refrigerator is a device used to store milk, milk is associated with the refrigerator. Therefore, the keyword "milk" in the second audio signal is associated with the refrigerator.
[0063] Step 360a: Send the second audio signal to the cloud;
[0064] The electronic device sends the second audio signal of the keyword associated with the terminal device to the cloud.
[0065] Step 370a: Receive a decision instruction generated by the cloud server according to the second audio signal, wherein the decision instruction is used to instruct to continue or terminate the sound pickup operation;
[0066] It should be noted that the cloud server stores a list of associated words corresponding to the terminal device. This list of associated words corresponding to the terminal device records information about associated words related to the terminal device. The cloud server performs voice analysis on the received second audio signal to obtain keywords. Based on the relationship between the keywords obtained from the voice analysis of the second audio signal and the preset list of associated words, a decision instruction is generated. The decision instruction is used to instruct the user to continue or terminate the sound pickup operation.
[0067] For example, the associated word list for refrigerator includes keywords related to refrigerator, such as milk, vegetables, meat, beverages, cooked food, eggs, etc. The associated word list for TV includes keywords related to TV, such as TV series, movies, variety shows, columns, movies, channels, etc.
[0068] Furthermore, step 370a, receiving a decision instruction generated by the cloud server according to the second audio signal, wherein the decision instruction is used to instruct to continue or terminate the sound pickup operation, includes:
[0069] Step 371a: Receive a continue sound pickup instruction generated by the cloud server when the correlation between the second audio signal and the first terminal device meets the requirements.
[0070] If the cloud server determines that the keyword obtained through semantic analysis of the second audio signal is associated with the first terminal device, that is, the keyword obtained through semantic analysis of the second audio signal is present in the associated word list of the first terminal device. For example, the keyword "milk" in the user's voice command "show the location of milk" is associated with the refrigerator. If the refrigerator is the first terminal device, then the cloud server generates a continue audio pickup instruction when the association between the keyword "milk" and the first terminal device "refrigerator" meets the requirements. The first terminal device receives the continue audio pickup instruction from the cloud server.
[0071] Step 380a: Execute sound pickup control according to the judgment instruction. Specifically including:
[0072] Step 381: Continue the sound pickup operation based on the continue sound pickup instruction.
[0073] The first terminal device continues to pick up audio based on the continue pickup instruction, thereby continuing to receive the user's voice instructions. By receiving the continue pickup instruction generated by the cloud server when the correlation between the second audio signal and the first terminal device meets the requirements, and when the cloud server confirms that the user's second audio signal is associated with the first terminal device, the first terminal device receives the continue pickup instruction from the cloud server to further pick up audio. This helps improve the accuracy of responding to the user's voice instructions by allowing the first terminal device, which the cloud server recognizes as being associated with the user's second audio signal, to further pick up audio from the user.
[0074] For other possible embodiments, please refer to Figure 5 , after step 300a, further comprising:
[0075] Step 400a: Send a third audio signal to the second terminal device to instruct the second terminal device to terminate sound pickup.
[0076] After not making an audio response to the first audio signal based on the rejection recognition instruction, the first terminal device sends a third audio signal to the second terminal device for instructing the second terminal device to terminate sound pickup.
[0077] It can be understood that after the first terminal device does not make an audio response to the first audio signal based on the rejection recognition instruction, both the first terminal device and the second terminal device can pick up the user's voice instruction.
[0078] In order to avoid the situation where both the first terminal device and the second terminal device respond to the user's voice command, i.e., the responses of the first terminal device and the second terminal device overlap, resulting in the user not being able to clearly hear the responses of the first terminal device and the second terminal device to the user, the first terminal device can send a third audio signal to the second terminal device to instruct the second terminal device to stop picking up the sound, thereby terminating the voice interaction function of the second terminal device, so that only the first terminal device is selected as the terminal device that responds to the user's voice command.
[0079] Among them, the third audio signal can be various response words that cause the second terminal device to stop picking up the sound. Such as "It's okay, thank you", "No need", "No", etc. For example, after the first terminal device receives the rejection recognition instruction issued by the cloud server, the first terminal device knows that the received first audio signal (such as "What are your instructions") is an invalid response word from the second terminal device. At this time, the first terminal device sends "No" to the second terminal device. After the second terminal device receives the audio signal "No" and recognizes it, the second terminal device immediately turns off its own voice interaction function. Therefore, at this time, only the first terminal device responds to the user's voice command.
[0080] It should be noted that the above-mentioned response words such as "No problem, thank you", "No need", and "No" can be preset in the memory of the first terminal device, or sent to the first terminal device by the cloud server. When the rejection recognition instruction is received, the first terminal device sends a third audio signal to the second terminal device to instruct the second terminal device to terminate the sound pickup.
[0081] After the first terminal device does not make an audio response to the first audio signal based on the rejection recognition instruction, the first terminal device sends a third audio signal to the second terminal device to instruct the second terminal device to stop picking up audio, causing the second terminal device to turn off the voice interaction function, thereby determining the first terminal device used for voice interaction with the user. This avoids the situation where both the first terminal device and the second terminal device respond to the user's voice instruction, i.e., the responses of the first terminal device and the second terminal device overlap, causing the user to unclearly hear the responses of the first terminal device and the second terminal device to the user, thereby improving the success rate of obtaining correct response content when multiple terminal devices are woken up at the same time.
[0082] Please refer to Figure 6 The voice interaction anti-interference method of an embodiment of the present invention is applied to a cloud server and includes:
[0083] Step 100b: receiving a recognition result sent by the first terminal device, where the recognition result is obtained based on the first terminal device recognizing the first audio signal;
[0084] The first audio signal is the target response word sent by the second terminal device, which is received by the first terminal device after the first terminal device activates the voice pickup operation in response to the voice wake-up command. Both the second terminal device and the first terminal device in the distributed collaborative group are equipped with a voice processing module, capable of recognizing the user's voice command and responding accordingly.
[0085] The cloud server receives a recognition result sent by the first terminal device, where the recognition result is obtained based on the first terminal device recognizing the first audio signal emitted by the second terminal device. That is, the recognition result is a response word emitted by the second terminal device. The recognition result may be a response word such as "Yes, I'm here," "Here I am," or "What can I do for you?"
[0086] Step 200b: When the recognition result indicates that the first audio signal includes a target response word, a rejection recognition instruction is sent to the first terminal device, where the rejection recognition instruction is used to instruct the first terminal device not to make an audio response to the first audio signal.
[0087] The cloud server's memory is pre-set with multiple target response words for rejection of recognition. When the cloud server recognizes that the recognition result indicates that the first audio signal includes the target response word, it sends a rejection of recognition instruction to the first terminal device, so that the first terminal device does not make an audio response to the first audio signal based on the rejection of recognition instruction. In this way, when the voice interaction functions of the first terminal device and the second terminal device are both activated, there is no mutual interference between the first terminal device and the second terminal device when receiving the user's voice command, thereby improving the success rate of the smart terminal device in recognizing the voice command.
[0088] It should be noted that the target response words include multiple response words used in daily life, such as "Yes, I'm here", "Master, I'm here", "I'm here", "What do you want me to do?", "What are your instructions?", "What does the master need me to do?", "What does the master need to know?" and other commonly used terminal device response words.
[0089] In some possible embodiments, please refer to Figure 7 Step 200b, in the case where the recognition result indicates that the first audio signal includes the target response word, sending a rejection recognition instruction to the first terminal device, comprises:
[0090] Step 210b: When the recognition result is the same as the target response word, a rejection recognition instruction is sent to the first terminal device.
[0091] When the recognition result is the same as the target response word pre-stored in the cloud server, that is, when the recognition result exactly matches the target response word pre-stored in the cloud server, the cloud server determines that the recognition result belongs to the target response word for rejection, and the cloud server sends a rejection recognition instruction to the first terminal device.
[0092] In some other feasible embodiments, please refer to Figure 8 , step 200b, when the recognition result indicates that the first audio signal includes a target response word, sending a rejection recognition instruction to the first terminal device includes:
[0093] Step 220b, when the similarity between the recognition result and the target response word exceeds a target threshold, send a rejection recognition instruction to the first terminal device.
[0094] In practice, the recognition result may not be exactly the same as the target response word. In some cases, the recognition result may have a few more or fewer words than the target response word. For example, the recognition result is "Mm-hmm, I'm here", while the target response word is "Mm, I'm here". At this time, there is a difference of one word between the recognition result and the target response word.
[0095] In this embodiment, when the similarity between the recognition result and the target response word exceeds the target threshold, a rejection recognition instruction is sent to the first terminal device. That is, in this embodiment, when the cloud server recognizes whether the received recognition result belongs to the target response word, as long as the similarity between the recognition result and the target response word exceeds the target threshold, the cloud server considers the recognition result to be an ineffective response word, and thus sends a rejection recognition instruction to the first terminal device for the first terminal device not to perform an audio response to the first audio signal based on the rejection recognition instruction.
[0096] Among them, the target threshold can be set according to the situation. For example, the target threshold is set to 90%. That is, when the similarity between the recognition result and the target response word exceeds 90%, the cloud server sends a rejection recognition instruction to the first terminal device, but the embodiments of the present application do not limit this.
[0097] It should be noted that the similarity between the recognition result and the target response word can be obtained through various text similarity calculation methods, such as through Word2Vec model, Doc2Vec model, etc.
[0098] By sending a rejection recognition instruction to the first terminal device when the similarity between the recognition result and the target response word exceeds the target threshold, it is avoided that when there is a slight difference between the recognition result and the target response word, the recognition result is recognized as a normal user voice interaction instruction, causing the first terminal device to respond to the first audio signal of the second terminal device, and mutual interference between the first terminal device and the second terminal device. It improves the success rate of intelligent terminal devices in recognizing voice instructions.
[0099] The voice interaction anti-interference device provided by the present invention is described below. The voice interaction anti-interference device described below and the voice interaction anti-interference method described above can be referenced to each other.
[0100] Please refer to Figure 10 The present invention also provides a voice interaction anti-interference device, including: a first receiving module 201, used to start the pickup operation and receive a first audio signal according to a voice wake-up instruction; a first sending module 202, used to identify the received first audio signal to obtain a recognition result, and send the recognition result to a cloud server; a processing module 203, used to receive a rejection recognition instruction made by the cloud server based on the recognition result when the recognition result indicates that the first audio signal includes a target response word, and not make an audio response to the first audio signal based on the rejection recognition instruction.
[0101] In the voice interaction anti-interference device provided by the present invention, a first terminal device does not provide an audio response to a first audio signal including a target response word. Consequently, when the voice interaction functions of both the first and second terminal devices are activated, there is no mutual interference between the first and second terminal devices when receiving user voice commands, thereby improving the success rate of the intelligent terminal device in recognizing voice commands.
[0102] As an optional embodiment, the first receiving module 201 is used to start the sound pickup operation according to the voice wake-up instruction, and obtain the target response word sent by the second terminal device as the first audio signal.
[0103] As an optional embodiment, the voice interaction anti-interference device further includes: a continuing sound pickup module, which is used to continue the sound pickup operation.
[0104] As an optional embodiment, the voice interaction anti-interference device also includes: a first judgment module, used to judge whether the current moment is within the target interaction time range of the current dialogue mode based on the current first dialogue mode of the first terminal device; a sound pickup module, used to switch the current first dialogue mode to the second dialogue mode when the current moment is within the target interaction time range, and continue the sound pickup operation based on the second dialogue mode.
[0105] As an optional embodiment, the voice interaction anti-interference device further includes: when the current moment is not within the target interaction time range, issuing a second audio signal for instructing the user to issue a voice command.
[0106] As an optional embodiment, the voice interaction anti-interference apparatus further includes: sending a third audio signal to the second terminal device for instructing the second terminal device to terminate voice pickup.
[0107] As an optional embodiment, the voice interaction anti-interference device also includes: a first acquisition module, used to obtain a second audio signal according to a user voice command; a second sending module, used to send the second audio signal to the cloud; an instruction generation module, used to receive a decision instruction generated by a cloud server according to the second audio signal, the decision instruction being used to instruct to continue or terminate the sound pickup operation; and a sound pickup control module, used to execute sound pickup control according to the decision instruction.
[0108] As an optional embodiment, the voice interaction anti-interference device also includes: an instruction generation module, specifically used to receive a continue pickup instruction generated by the cloud server when the correlation between the second audio signal and the first terminal device meets the requirements; a pickup control module, specifically used to continue the pickup operation based on the continue pickup instruction.
[0109] Please refer to Figure 11 The present invention also provides a voice interaction anti-interference device, including: a second receiving module 204, used to receive a recognition result sent by a first terminal device, wherein the recognition result is obtained based on the first terminal device recognizing the first audio signal; a second sending module 205, used to send a rejection recognition instruction to the first terminal device when the recognition result indicates that the first audio signal includes a target response word, wherein the rejection recognition instruction is used to instruct the first terminal device not to make an audio response to the first audio signal.
[0110] As an optional embodiment, the first audio signal is obtained by the first terminal device starting the sound pickup operation according to the voice wake-up instruction and obtaining the target response word sent from the second terminal device.
[0111] As an optional embodiment, the second sending module includes: a first sub-sending module, configured to send a rejection recognition instruction to the first terminal device if the recognition result is the same as the target response word;
[0112] As an optional embodiment, the second sending module includes: a second sub-sending module, which is used to send a rejection recognition instruction to the first terminal device when the similarity between the recognition result and the target response word exceeds a target threshold.
[0113] Figure 12 An example of a physical structure diagram of an electronic device is shown below. Figure 12 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the voice interaction anti-interference method described in the above embodiment.
[0114] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0115] It should be noted that the electronic device further includes a server. When the electronic device is a server, it is used to execute the voice interaction anti-interference method described in the above embodiment.
[0116] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voice interaction anti-interference method described in the above embodiment.
[0117] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the voice interaction anti-interference method described in the above embodiment.
[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A voice interaction anti-interference method, characterized in that: Applied to a first terminal device in a distributed collaborative group, where a second terminal device in the distributed collaborative group and the first terminal device are both configured with a voice processing module, the method includes: According to the voice wake-up instruction, the sound pickup operation is started to receive the first audio signal; Recognize the received first audio signal to obtain a recognition result, and send the recognition result to a cloud server; When the recognition result indicates that the first audio signal includes a target response word, a rejection recognition instruction made by the cloud server based on the recognition result is received, and no audio response is made to the first audio signal based on the rejection recognition instruction; wherein the target response word is the response word that the second terminal device responds to the voice wake-up instruction.
2. The voice interaction anti-interference method according to claim 1, characterized in that: After not making an audio response to the first audio signal based on the rejection of the recognition instruction, the method further includes: Continue with the pickup operation.
3. The voice interaction anti-interference method according to claim 1 or 2, characterized in that: Also includes: Determining, based on the current first conversation mode of the first terminal device, whether the current moment is within a target interaction time range of the current conversation mode; When the current moment is within the target interaction time range, the current first dialogue mode is switched to the second dialogue mode, and the sound pickup operation is continued based on the second dialogue mode.
4. The voice interaction anti-interference method according to claim 3, characterized in that: Also includes: When the current moment is not within the target interaction time range, a second audio signal is emitted to instruct the user to issue a voice command.
5. The voice interaction anti-interference method according to claim 2, characterized in that: After the sound pickup operation is continued, the method further includes: Acquiring a second audio signal according to a user voice command; sending the second audio signal to the cloud; receiving a decision instruction generated by the cloud server according to the second audio signal, wherein the decision instruction is used to instruct to continue or terminate the sound pickup operation; Sound pickup control is performed according to the decision instruction.
6. The voice interaction anti-interference method according to claim 5, characterized in that: The receiving a decision instruction generated by the cloud server according to the second audio signal includes: receiving a continuing sound pickup instruction generated by the cloud server when the correlation between the second audio signal and the first terminal device meets the requirement; The performing of the sound pickup control according to the decision instruction includes: The sound pickup operation is continued based on the continue sound pickup instruction.
7. A voice interaction anti-interference method, characterized in that: Applied to cloud servers, including: Receiving a recognition result sent by the first terminal device, where the recognition result is obtained based on recognition of the first audio signal by the first terminal device; When the recognition result indicates that the first audio signal includes a target response word, a rejection recognition instruction is sent to the first terminal device, and the rejection recognition instruction is used to instruct the first terminal device not to make an audio response to the first audio signal; wherein, the target response word is the response word of the second terminal device in response to the voice wake-up instruction.
8. The voice interaction anti-interference method according to claim 7, characterized in that: The sending a recognition rejection instruction to the first terminal device when the recognition result indicates that the first audio signal includes a target response word includes: If the recognition result is the same as the target response word, sending a recognition rejection instruction to the first terminal device; When the similarity between the recognition result and the target response word exceeds a target threshold, a recognition rejection instruction is sent to the first terminal device.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the voice interaction anti-interference method as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Speech recognition method, device and equipment and computer readable storage medium
CN112185388A