Wake-up-free voice interaction method, device and equipment and readable storage medium

Through voice detection and recognition combined with the confidence and effective command judgment of voice audio text, the cumbersome and limitations of existing voice interaction methods are solved, efficient wake-up-free voice interaction is achieved, and memory usage is reduced.

CN120260565APending Publication Date: 2025-07-04AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510388905.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the existing voice interaction methods, the wake-up voice interaction system is cumbersome and inefficient, while the wake-up voice interaction is limited and cannot effectively recognize multiple voice commands.

Method used

Through voice detection and recognition, the duration of the voice is judged. If it is greater than the time threshold, the detection is paused. If it is less than or equal to the time threshold, the wake-up voice interaction is performed, and the confidence of the voice audio text and the effective command judgment rules are used to interact.

Benefits of technology

On the premise of ensuring the efficiency of wake-up-free voice interaction, memory usage is reduced, invalid voice input is reduced, and the convenience and efficiency of voice interaction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260565A_ABST
    Figure CN120260565A_ABST
Patent Text Reader

Abstract

The invention discloses a wakeup-free voice interaction method, device and equipment and a readable storage medium, and relates to the technical field of artificial intelligence. Comprising the following steps: firstly, acquiring voice audio information, and carrying out human voice detection on the voice audio information to obtain a human voice detection result; based on the human voice detection result, carrying out human voice recognition on the voice audio information to obtain a voice audio text; if the duration of the human voice recognition is greater than a time threshold value, pausing human voice detection on the voice audio information, and after the pause time is over, performing human voice detection on the voice audio information again; and if the duration of the human voice recognition is less than or equal to a time threshold, performing wakeup-free voice interaction based on the voice audio text. According to the method provided by the invention, the memory occupation is greatly reduced on the premise of ensuring the wake-free voice interaction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a wake-up-free voice interaction method, device, equipment and readable storage medium. Background Art

[0002] With the development of artificial intelligence, more and more devices have begun to change from the original button control to voice control, such as cars, televisions, and air conditioners. Voice control greatly increases the control convenience and brings great convenience to people's lives.

[0003] Currently, voice interaction methods are generally divided into two types. One is to first wake up the voice interaction system through a wake-up word and then perform voice interaction. The other does not need to wake up the voice system first and directly performs wake-up-free voice interaction. However, both of these voice interaction methods have certain defects. The former needs to first wake up the voice interaction system through a wake-up word to perform voice interaction, and the process is a bit cumbersome, and the voice interaction efficiency is relatively low. The latter can currently only recognize a part of special voice commands, and the limitations of voice interaction are relatively large.

[0004] Therefore, in order to further increase the convenience of people's daily lives and the convenience of voice interaction, there is an urgent need for a wake-up-free voice interaction method that can overcome the above defects. Summary of the Invention

[0005] The purpose of the present invention is to provide a wake-up-free voice interaction method, device, equipment and readable storage medium. When the duration of human voice recognition is greater than the time threshold, it indicates that the human voice is a user's chat rather than for initiating voice interaction. At this time, the human voice detection is paused. Only when the duration of human voice recognition is less than or equal to the time threshold, wake-up-free voice interaction is performed. This pause can reduce the input of invalid voices and greatly reduce the memory occupancy on the premise of ensuring the efficiency of wake-up-free voice interaction.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a wake-up-free voice interaction method, and the method includes:

[0008] Obtain voice audio information, and perform human voice detection on the voice audio information to obtain a human voice detection result;

[0009] Based on the human voice detection result, perform human voice recognition on the voice audio information to obtain a voice audio text;

[0010] If the duration of the human voice recognition is greater than the time threshold, pause the human voice detection on the voice audio information. After the pause time ends, re-perform the human voice detection on the voice audio information;

[0011] If the duration of the voice recognition is less than or equal to the time threshold, perform wake-free voice interaction based on the voice audio text.

[0012] In some embodiments, performing wake-free voice interaction based on the voice audio text includes:

[0013] If the confidence of the voice audio text is greater than the confidence threshold, parse semantic information from the voice recognition result;

[0014] Perform wake-free voice interaction based on the semantic information.

[0015] In some embodiments, performing wake-free voice interaction based on the semantic information includes:

[0016] Based on a preset valid instruction determination rule, determine whether the semantic information contains a valid instruction;

[0017] If the semantic information contains a valid instruction, perform wake-free interaction based on the valid instruction.

[0018] In some embodiments, performing wake-free voice interaction based on the semantic information further includes:

[0019] If the semantic information does not contain a valid instruction, pause the voice detection of the voice audio information, and after the pause time ends, re-perform voice detection on the voice audio information.

[0020] In some embodiments, after performing wake-free interaction based on the valid instruction, the method further includes:

[0021] Based on the valid instruction, determine whether to enter a continuous conversation;

[0022] If it is necessary to enter a continuous conversation, obtain voice audio information again.

[0023] In some embodiments, the method further includes:

[0024] If the confidence of the voice audio text is less than or equal to the confidence threshold,

[0025] and / or if the semantic information does not contain a valid instruction,

[0026] and / or if it is not necessary to enter a continuous conversation, end the current wake-free voice interaction.

[0027] In a second aspect, the present invention further provides a wake-free voice interaction device, and the device includes:

[0028] A voice detection module, configured to obtain voice audio information, perform voice detection on the voice audio information, and obtain a voice detection result;

[0029] A voice recognition module, configured to perform voice recognition on the voice audio information based on the voice detection result, and obtain a voice audio text;

[0030] A detection pause module, configured to, if the duration of the voice recognition is greater than a time threshold, pause the voice detection on the voice audio information, and after the pause time ends, perform voice detection on the voice audio information again;

[0031] A voice interaction module, configured to, if the duration of the voice recognition is less than or equal to the time threshold, perform wake - free voice interaction based on the voice audio text.

[0032] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the wake - free voice interaction method provided in the first aspect is implemented.

[0033] In a fourth aspect, the present invention further provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the wake - free voice interaction method provided in the first aspect is implemented.

[0034] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the wake - free voice interaction method provided in the first aspect is implemented.

[0035] The beneficial effects of the present invention are as follows:

[0036] In the wake - free voice interaction method of the present invention, first, voice audio information is obtained, and voice detection is performed on the voice audio information to obtain a voice detection result; then, based on the voice detection result, voice recognition is performed on the voice audio information to obtain a voice audio text; if the duration of the voice recognition is greater than a time threshold, the voice detection on the voice audio information is paused, and after the pause time ends, voice detection on the voice audio information is performed again; if the duration of the voice recognition is less than or equal to the time threshold, wake - free voice interaction is performed based on the voice audio text. When the duration of the voice recognition is greater than the time threshold, it indicates that the voice is the user's chatting rather than for initiating voice interaction. At this time, the voice detection is paused. Only when the duration of the voice recognition is less than or equal to the time threshold, wake - free voice interaction is performed. This pause can reduce the input of invalid voices and greatly reduce the memory occupation on the premise of ensuring the efficiency of wake - free voice interaction.

[0037] The above description is only an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and implement it according to the content of the specification, the following describes in detail with reference to the preferred embodiments of the present invention and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic flowchart of a voice interaction method without wake-up according to an embodiment of the present invention;

[0039] Figure 2 It is a schematic flowchart of another voice interaction method without wake-up according to an embodiment of the present invention;

[0040] Figure 3 It is a schematic structural diagram of a voice interaction device without wake-up according to an embodiment of the present invention;

[0041] Figure 4 It is a schematic structural diagram of another voice interaction device without wake-up according to an embodiment of the present invention;

[0042] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] It should be noted that the references to "one embodiment", "embodiment", "example embodiment", etc. in this specification mean that the described embodiment may include specific features, structures or characteristics, but not every embodiment must include these specific features, structures or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining specific features, structures or characteristics in an embodiment, it is within the knowledge of those skilled in the art to combine such features, structures or characteristics into other embodiments whether or not there is an explicit description.

[0045] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0046] In some embodiments, as Figure 1 shown, a schematic flowchart of a voice interaction method without wake-up is provided:

[0047] S101, Obtain voice audio information, perform voice detection on the voice audio information, and obtain the voice detection result.

[0048] Among them, the voice audio information is the voice audio emitted in the environment collected by the microphone. The voice audio information may or may not include human voices, and voice detection is required to further confirm. The voice detection result includes: there is a human voice and there is no human voice.

[0049] Specifically, obtain the voice audio information in the environment, and input the voice audio information into the voice detection model. The voice detection model can then output the voice detection result. The voice detection result may be that the voice audio information includes a human voice or the voice audio information does not include a human voice. When the voice detection result shows that the voice audio information includes a human voice, continue with the subsequent steps. When the voice detection result shows that the voice audio information does not include a human voice, there is no need to perform the subsequent steps.

[0050] S102, Based on the voice detection result, perform voice recognition on the voice audio information to obtain the voice audio text.

[0051] Among them, voice recognition can recognize the content in the voice audio information and output the corresponding text. For example, when the user says "The weather is nice today", voice recognition will convert it into the text "The weather is nice today".

[0052] Specifically, if the voice detection result shows that the voice audio information includes a human voice, then input the voice audio information into the voice recognition model, and the voice recognition model will output the voice audio text.

[0053] S103, If the duration of voice recognition is greater than the time threshold, then pause the voice detection of the voice audio information. After the pause time ends, re-perform the voice detection on the voice audio information.

[0054] Exemplarily, if there is always a human voice in the surrounding environment, the system will always perform voice recognition. When the duration of voice recognition exceeds a certain length, it can be determined that the user is currently in a conversation rather than performing voice interaction, because usually during voice interaction, it is a single sentence rather than a long speech. Therefore, at this time, the voice detection of the voice audio information can be paused (for example, paused for 1000 milliseconds), which can reduce the excessive CPU resources occupied by invalid voice detection. After the pause time ends, re-perform the voice detection on the voice audio information.

[0055] S104, If the duration of voice recognition is less than or equal to the time threshold, then perform wake-free voice interaction based on the voice audio text.

[0056] Specifically, if the duration of voice recognition is less than or equal to the time threshold, it may be that the user needs to perform voice interaction at this time. Then, based on the voice audio text, wake-free voice interaction is performed.

[0057] Optionally, the method for performing wake-free voice interaction based on the voice audio text can also be: if the confidence of the voice audio text is greater than the confidence threshold, semantic information is parsed from the voice recognition result; based on the semantic information, wake-free voice interaction is performed.

[0058] Specifically, first judge the confidence of the voice audio text. Only when the confidence of the voice audio text is greater than the confidence threshold (for example, 0.6), the voice recognition result is input into the semantic parsing model to obtain semantic information, and then based on the semantic information, wake-free voice interaction is performed. When the confidence of the voice audio text is less than or equal to the confidence threshold, it means that voice interaction is not required, and this wake-free voice interaction is directly ended.

[0059] Optionally, the method for performing wake-free voice interaction based on semantic information can also be: based on a preset valid instruction determination rule, judge whether the semantic information contains a valid instruction; if the semantic information contains a valid instruction, perform wake-free interaction based on the valid instruction.

[0060] Among them, the preset valid instruction determination rules include skill whitelists, rejection rules, confidence, word count matching, etc.

[0061] Specifically, when it is determined based on the preset valid instruction determination rule that the semantic information contains a valid instruction, a response dialogue can be generated based on the valid instruction, or a response operation can be performed simultaneously.

[0062] Exemplarily, when applied to a vehicle scenario, the valid instruction may be to open the window. At this time, the operation of opening the window can be performed, and a corresponding dialogue "The window has been opened" can be generated.

[0063] Optionally, if the semantic information does not contain a valid instruction, the voice detection of the voice audio information is paused, and after the pause time ends, the voice detection of the voice audio information is restarted.

[0064] Specifically, when the semantic information does not contain a valid instruction, it may be used to have a conversation with relatively short sentences and voice interaction is not required. At this time, the voice detection of the voice audio information can still be paused (for example, paused for 500 milliseconds). When the pause time ends, the voice detection of the voice audio information is restarted. This can reduce the CPU occupancy of voice detection, voice recognition, and semantic parsing. When the semantic information does not contain a valid instruction, in addition to pausing the voice detection of the voice audio information, this wake-free voice interaction also needs to be ended.

[0065] Optionally, when the semantic information contains a valid instruction and after generating a response dialogue, it is possible to determine whether to enter a continuous dialogue based on the content of the valid instruction. For example, when the valid instruction is "play music", the subsequent response dialogue to be generated is "Okay, what song do you want to play?" At this time, it can be determined that a continuous dialogue is required, and at this time, it is necessary to obtain the voice audio information again and then conduct a continuous dialogue. When it is determined that a continuous dialogue is not required, the current voice interaction without wake-up can be directly ended.

[0066] In the voice interaction method without wake-up in the above embodiments, first, obtain the voice audio information, and perform voice detection on the voice audio information to obtain a voice detection result; then, based on the voice detection result, perform voice recognition on the voice audio information to obtain a voice audio text; if the duration of the voice recognition is greater than the time threshold, pause the voice detection on the voice audio information, and after the pause time ends, re-perform voice detection on the voice audio information; if the duration of the voice recognition is less than or equal to the time threshold, perform voice interaction without wake-up based on the voice audio text. When the duration of the voice recognition is greater than the time threshold, it indicates that this voice is the user's chat rather than for initiating a voice interaction. At this time, the voice detection is paused, and only when the duration of the voice recognition is less than or equal to the time threshold, the voice interaction without wake-up is performed. This pause can reduce the input of invalid voices and greatly reduce the memory occupancy on the premise of ensuring the efficiency of the voice interaction without wake-up.

[0067] To more comprehensively demonstrate the present solution, an optional way of the voice interaction method without wake-up in this embodiment is given, as Figure 2 shown below:

[0068] S201, Obtain the voice audio information, and perform voice detection on the voice audio information to obtain a voice detection result.

[0069] S202, Based on the voice detection result, perform voice recognition on the voice audio information to obtain a voice audio text.

[0070] S203, Determine whether the duration of the voice recognition is greater than the time threshold. If so, execute S204; if not, execute S205.

[0071] S204, Pause the voice detection on the voice audio information, and after the pause time ends, execute S201.

[0072] S205, Determine whether the confidence level of the voice audio text is greater than the confidence level threshold. If so, execute S206; if not, execute S210.

[0073] S206, Parse the semantic information from the voice recognition result.

[0074] S207. Based on a preset valid instruction determination rule, determine whether the semantic information contains a valid instruction. If yes, execute S208; if no, execute S204 and S210.

[0075] S208. Perform wake-up-free interaction based on the valid instruction.

[0076] S209. Based on the valid instruction, determine whether to enter a continuous conversation. If yes, re-execute S201; if no, execute S210.

[0077] S210. End the current wake-up-free voice interaction.

[0078] For the specific processes of the above S201 - S210, reference can be made to the description of the method embodiments above. Their implementation principles and technical effects are similar, and will not be elaborated here.

[0079] Based on the same inventive concept, an embodiment of the present application also provides a wake-up-free voice interaction device for implementing the above-mentioned wake-up-free voice interaction method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following wake-up-free voice interaction device can refer to the limitations on the wake-up-free voice interaction method above, and will not be elaborated here.

[0080] In one embodiment, as Figure 3 shown, a wake-up-free voice interaction device is provided, and this device includes:

[0081] A voice detection module 30, configured to obtain voice audio information and perform voice detection on the voice audio information to obtain a voice detection result;

[0082] A voice recognition module 31, configured to perform voice recognition on the voice audio information based on the voice detection result to obtain a voice audio text;

[0083] A detection pause module 32, configured to, if the duration of the voice recognition is greater than a time threshold, pause the voice detection on the voice audio information, and after the pause time ends, re-perform the voice detection on the voice audio information;

[0084] A voice interaction module 33, configured to, if the duration of the voice recognition is less than or equal to the time threshold, perform wake-up-free voice interaction based on the voice audio text.

[0085] In another embodiment, as Figure 4 shown, the above Figure 3 voice interaction module 33 includes:

[0086] A semantic parsing unit 330, configured to parse semantic information from the voice recognition result if the confidence of the voice audio text is greater than a confidence threshold.

[0087] A voice interaction unit 331, configured to perform wake-free voice interaction based on the semantic information.

[0088] In another embodiment, the above Figure 4 The voice interaction unit 331 is specifically configured to: determine whether the semantic information contains a valid instruction based on a preset valid instruction determination rule; if the semantic information contains a valid instruction, perform wake-free interaction based on the valid instruction.

[0089] Optionally, performing wake-free voice interaction based on the semantic information further includes: if the semantic information does not contain a valid instruction, pause the voice detection of the voice audio information, and after the pause time ends, re-perform voice detection on the voice audio information.

[0090] Optionally, after performing wake-free interaction based on the valid instruction, the method further includes: determining whether to enter a continuous conversation based on the valid instruction; if it is necessary to enter a continuous conversation, obtain voice audio information again.

[0091] In another embodiment, the above Figure 3 The wake-free voice interaction device is further specifically configured to: end the current wake-free voice interaction if the confidence of the voice audio text is less than or equal to the confidence threshold, and / or if the semantic information does not contain a valid instruction, and / or if it is not necessary to enter a continuous conversation.

[0092] An embodiment of the present application further provides an electronic device. In some embodiments, as shown in Figure 5 reference, the electronic device 700 includes an input unit 710, a memory 720, a processor 730, and an output unit 740. The memory 720 stores program instructions that can run on the processor 730, and the processor 730 can execute the wake-free voice interaction method and / or technical solution based on the foregoing embodiments by invoking the program instructions. The electronic device 700 can be a mobile terminal device such as a mobile phone or a computer.

[0093] In addition, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program for executing the wake-free voice interaction method. For example, computer program instructions, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. The program instructions for calling the method of the present application may be stored in a fixed or removable storage medium, and / or transmitted and / or stored in a storage medium running according to the program instructions through a data stream in a broadcast or other signal-bearing medium.

[0094] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented with program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0095] The technical features of the above embodiments can be arbitrarily integrated. For the sake of brevity of description, not all possible integrations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the integration of these technical features, it should be considered to be within the scope described in this specification.

[0096] The above embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A voice interaction method without wake-up, characterized in that, The method includes: Obtaining voice audio information and performing voice detection on the voice audio information to obtain a voice detection result; Based on the voice detection result, performing voice recognition on the voice audio information to obtain a voice audio text; If the duration of the voice recognition is greater than a time threshold, pause the voice detection on the voice audio information, and after the pause time ends, perform voice detection on the voice audio information again; If the duration of the voice recognition is less than or equal to the time threshold, perform wake-free voice interaction based on the voice audio text.

2. The voice interaction method without waking up as claimed in claim 1, wherein, Performing wake-free voice interaction based on the voice audio text includes: If the confidence level of the voice audio text is greater than a confidence level threshold, parse semantic information from the voice recognition result; Based on the semantic information, perform wake-free voice interaction.

3. The voice interaction method without wake-up as claimed in claim 2, wherein, Performing wake-free voice interaction based on the semantic information includes: Based on a preset valid instruction determination rule, determine whether the semantic information contains a valid instruction; If the semantic information contains a valid instruction, perform wake-free interaction based on the valid instruction.

4. The voice interaction method without wake-up as claimed in claim 3, wherein Performing wake-free voice interaction based on the semantic information further includes: If the semantic information does not contain a valid instruction, pause the voice detection on the voice audio information, and after the pause time ends, perform voice detection on the voice audio information again.

5. The voice interaction method without waking up as claimed in claim 3, wherein After performing wake-free interaction based on the valid instruction, the method further includes: Based on the valid instruction, determine whether to enter a continuous conversation; If it is necessary to enter a continuous conversation, obtain voice audio information again.

6. The voice interaction method without wake-up as claimed in any one of claims 1-5, characterized in that The method further includes: If the confidence level of the voice audio text is less than or equal to the confidence level threshold, and / or if the semantic information does not contain a valid instruction, and / or if it is not necessary to enter a continuous conversation, end the current wake-free voice interaction.

7. A voice interaction device without wake-up, characterized in that, The device includes: A voice detection module for obtaining voice audio information and performing voice detection on the voice audio information to obtain a voice detection result; A voice recognition module for performing voice recognition on the voice audio information based on the voice detection result to obtain a voice audio text; A detection pause module for, if the duration of the voice recognition is greater than a time threshold, pausing the voice detection on the voice audio information, and after the pause time ends, performing voice detection on the voice audio information again; A voice interaction module for, if the duration of the voice recognition is less than or equal to the time threshold, performing wake-free voice interaction based on the voice audio text.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the wake-free voice interaction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the wake-free voice interaction method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the wake-free voice interaction method according to any one of claims 1 to 6.