Wake-up-free voice interaction method, device and equipment and readable storage medium

By performing confidence and similarity calculations in the cloud, the problem that existing voice interaction systems are difficult to simultaneously improve end-to-end achievement rates and reduce system false alarm values ​​is solved, and more efficient voice interaction performance is achieved.

CN120089137APending Publication Date: 2025-06-03AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237546.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

It is difficult for existing voice interaction systems to increase end-to-end achievement rates and reduce system false alarm values ​​at the same time. There is a "seesaw" problem at the model level and cannot be effectively solved.

Method used

By calculating confidence and similarity in the cloud, local computing power and memory usage are reduced, and wake-up-free voice interaction is only performed when the preset threshold is met, ensuring the end-to-end achievement rate and reducing the system's false alarm value.

Benefits of technology

It realizes that while ensuring end-to-end achievement rates, the system false alarm value is reduced and the overall performance of the voice interaction system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089137A_ABST
    Figure CN120089137A_ABST
Patent Text Reader

Abstract

The invention discloses a wakeup-free voice interaction method, device and equipment and a readable storage medium, and relates to the technical field of artificial intelligence. Comprising the following steps: acquiring a voice interaction audio, and performing human voice recognition on the voice interaction audio to obtain a voice recognition text; performing semantic recognition on the voice recognition text to obtain semantic information and semantic slots of the voice recognition text; performing cloud confidence calculation based on the speech recognition text and the semantic information, and performing cloud similarity calculation based on the semantic slot to obtain a first confidence, a second confidence and a slot similarity; and if a wake-up half word is detected from the voice recognition text and the first confidence coefficient, the second confidence coefficient and the slot similarity are all greater than a preset threshold, performing wake-free voice interaction based on the semantic information and the semantic slot. According to the scheme, the false alarm value of the system is reduced under the condition of ensuring the end-to-end achievement rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a wake-up-free voice interaction method, device, equipment and readable storage medium. Background Art

[0002] With the development of artificial intelligence, most new energy vehicles are equipped with some artificial intelligence functions, and users can interact with the vehicle-mounted computer through voice to control the vehicle. However, in essence, they are all doing some relatively basic in-set operations (such as vehicle control commands, multimedia control, etc.) by combining recognition and semantics in related limited fields. Such voice interaction systems often start a thread outside the wake-up thread in the system to continuously and real-time run voice activity detection. After obtaining the effective human voice audio, it is input into the automatic speech recognition system to obtain text information. The text information is processed through natural language understanding to obtain semantic results. Finally, the subsequent operation information can be analyzed by obtaining the semantic results in the dialogue system.

[0003] However, it is difficult to simultaneously ensure a relatively good end-to-end achievement rate and reduce the system false alarm value at the performance level of the entire system by the above method. These two indicators belong to a "seesaw" problem at the model level, and the current technology cannot ensure reducing the system false alarm value while improving the end-to-end achievement rate. Therefore, there is an urgent need for a wake-up-free voice interaction method that can solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide a wake-up-free voice interaction method, device, equipment and readable storage medium. By calculating the confidence and similarity in the cloud, the calculation process does not occupy local computing power and local memory. Only when the first confidence, the second confidence and the slot similarity are all greater than the preset threshold, the wake-up-free voice interaction is performed based on the semantic information and semantic slots. The cloud calculation ensures the end-to-end achievement rate, so that the local computing power can focus on the processing of voice interaction audio, reduce the system false alarm value, and thus achieve reducing the system false alarm value while ensuring the end-to-end achievement rate.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a wake-up-free voice interaction method, and the method includes:

[0007] Obtain a voice interaction audio, and perform human voice recognition on the voice interaction audio to obtain a speech recognition text;

[0008] Perform semantic recognition on the speech recognition text to obtain semantic information and semantic slots of the speech recognition text;

[0009] Perform cloud confidence calculation based on the speech recognition text and semantic information, and perform cloud similarity calculation based on the semantic slots to obtain a first confidence level, a second confidence level, and a slot similarity;

[0010] If a wake-up half-word is detected from the speech recognition text, and the first confidence level, the second confidence level, and the slot similarity are all greater than a preset threshold, then perform wake-up-free speech interaction based on the semantic information and semantic slots.

[0011] In some embodiments, performing voice recognition on the speech interaction audio to obtain a speech recognition text includes:

[0012] Calculate the voice confidence level in the speech interaction audio;

[0013] If the voice confidence level is greater than the voice confidence threshold, input the speech interaction audio into an acoustic model for voice recognition to obtain a speech recognition text.

[0014] In some embodiments, performing semantic recognition on the speech recognition text to obtain the semantic information and semantic slots of the speech recognition text includes:

[0015] Input the speech recognition text into a semantic model to obtain the semantic information of the speech recognition text;

[0016] Based on the semantic information, determine the semantic slots in the speech recognition text.

[0017] In some embodiments, performing cloud confidence calculation based on the speech recognition text and semantic information, and performing cloud similarity calculation based on the semantic slots to obtain a first confidence level, a second confidence level, and a slot similarity includes:

[0018] Based on the speech recognition text, calculate the first confidence level of the acoustic model in the cloud;

[0019] Based on the semantic information, calculate the second confidence level of the semantic model in the cloud;

[0020] Based on the semantic slots, calculate the slot similarity of the semantic slots in the cloud.

[0021] In some embodiments, the preset threshold includes a first confidence threshold, a second confidence threshold, and a slot similarity threshold. The step of, if a wake-up half-word is detected from the speech recognition text, and the first confidence level, the second confidence level, and the slot similarity are all greater than the preset threshold, then performing wake-up-free speech interaction based on the semantic information and semantic slots includes:

[0022] Perform wake-up half-word detection on the speech recognition text;

[0023] If there is a wake-up partial word in the speech recognition text, and the first confidence level is greater than the first confidence level threshold, the second confidence level is greater than the second confidence level threshold, and the slot similarity is greater than the slot similarity threshold, then perform wake-up-free voice interaction based on the semantic information and semantic slots.

[0024] In some embodiments, after performing semantic recognition on the speech recognition text to obtain the semantic information and semantic slots of the speech recognition text, the method further includes:

[0025] Determine whether multi-round voice interaction is required based on the semantic information and semantic slots;

[0026] If so, obtain the voice interaction audio again.

[0027] In a second aspect, the present invention further provides a wake-up-free voice interaction device, which includes:

[0028] A voice recognition module, configured to obtain voice interaction audio and perform voice recognition on the voice interaction audio to obtain a speech recognition text;

[0029] A semantic recognition module, configured to perform semantic recognition on the speech recognition text to obtain the semantic information and semantic slots of the speech recognition text;

[0030] A cloud computing module, configured to perform cloud confidence calculation based on the speech recognition text and semantic information and perform cloud similarity calculation based on the semantic slots to obtain a first confidence level, a second confidence level, and a slot similarity;

[0031] A voice interaction module, configured to, if a wake-up partial word is detected from the speech recognition text, and the first confidence level, the second confidence level, and the slot similarity are all greater than a preset threshold, then perform wake-up-free voice interaction based on the semantic information and semantic slots.

[0032] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the wake-up-free voice interaction method provided in the first aspect is implemented.

[0033] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the wake-up-free voice interaction method provided in the first aspect is implemented.

[0034] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the wake-up-free voice interaction method provided in the first aspect is implemented.

[0035] The beneficial effects of the present invention are as follows:

[0036] In the solution of the present invention, first, the voice interaction audio is obtained, and voice recognition is performed on the voice interaction audio to obtain a voice recognition text; then semantic recognition is performed on the voice recognition text to obtain semantic information and semantic slots of the voice recognition text; then cloud confidence calculation is performed based on the voice recognition text and semantic information, and cloud similarity calculation is performed based on the semantic slots to obtain a first confidence, a second confidence, and a slot similarity; if a wake-up half-word is detected in the voice recognition text, and the first confidence, the second confidence, and the slot similarity are all greater than a preset threshold, then wake-up-free voice interaction is performed based on the semantic information and semantic slots. In the above solution, by performing confidence and similarity calculations in the cloud, the calculation process does not occupy local computing power and local memory. Only when the first confidence, the second confidence, and the slot similarity are all greater than the preset threshold, wake-up-free voice interaction is performed based on the semantic information and semantic slots. The cloud calculation ensures the end-to-end achievement rate, so that the local computing power can focus on the processing of the voice interaction audio, reducing the system false alarm value, and thus realizing the reduction of the system false alarm value while ensuring the end-to-end achievement rate.

[0037] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of the present invention and combines the accompanying drawings to describe in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic flowchart of a wake-up-free voice interaction method shown in an embodiment of the present invention;

[0039] Figure 2 It is a schematic flowchart of another wake-up-free voice interaction method shown in an embodiment of the present invention;

[0040] Figure 3 It is a schematic structural diagram of a wake-up-free voice interaction device shown in an embodiment of the present invention;

[0041] Figure 4 It is a schematic structural diagram of another wake-up-free voice interaction device shown in an embodiment of the present invention;

[0042] Figure 5 It is a schematic structural diagram of yet another wake-up-free voice interaction device shown in an embodiment of the present invention;

[0043] Figure 6 It is a schematic structural diagram of still another wake-up-free voice interaction device shown in an embodiment of the present invention;

[0044] Figure 7Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0045] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that the references to "one embodiment", "embodiment", "example embodiment", etc. in this specification mean that the described embodiment may include specific features, structures, or characteristics, but not every embodiment must include these specific features, structures, or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining embodiments to describe specific features, structures, or characteristics, whether or not there is an explicit description, it has been shown that it is within the knowledge of those skilled in the art to combine such features, structures, or characteristics into other embodiments.

[0047] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0048] In some embodiments, as Figure 1 shown, a flowchart of a voice interaction method without wake-up is provided, and the method includes:

[0049] S101, obtaining a voice interaction audio, and performing voice recognition on the voice interaction audio to obtain a voice recognition text.

[0050] Among them, the voice interaction audio is the audio collected for voice interaction, and the voice interaction audio is the audio collected from the environment, which may or may not contain human voices; the voice recognition text is the text contained in the voice interaction audio.

[0051] Specifically, the voice interaction audio can be collected by using a microphone, and a voice-to-text tool can be called to perform voice recognition processing on the voice interaction audio to obtain a voice recognition text.

[0052] Optionally, the method for obtaining the voice recognition text can also be: calculating the confidence of the human voice in the voice interaction audio; if the confidence of the human voice is greater than the human voice confidence threshold, inputting the voice interaction audio into an acoustic model for voice recognition to obtain a voice recognition text.

[0053] Among them, the confidence of the human voice is the ratio of the human voice to the non-human voice.

[0054] Specifically, relevant tools can be used to extract the human voice from the voice interaction audio and calculate the ratio of the human voice to the non-human voice, that is, the human voice confidence. The purpose of this process is to distinguish whether there is a human voice in the voice interaction audio. If the human voice confidence is greater than the human voice confidence threshold, it indicates that there is a human voice in the voice interaction audio. At this time, the voice interaction audio is input into the acoustic model for human voice recognition to obtain the speech recognition text. If the human voice confidence is not greater than the human voice confidence threshold, it indicates that there is no human voice in the voice interaction audio, and no subsequent processing is required at this time.

[0055] S102. Perform semantic recognition on the speech recognition text to obtain the semantic information and semantic slots of the speech recognition text.

[0056] Among them, semantic slots are an important task in natural language understanding (NLU), especially in applications such as speech recognition, dialogue systems, and information extraction. Semantic slots usually refer to specific information units that are concerned in a specific context. For example, in a dialogue about booking a flight, the slots may include the departure place, destination, date, etc.

[0057] Specifically, first perform word segmentation on the speech recognition text to obtain several segmented words, and perform part-of-speech tagging on these segmented words (such as nouns, verbs, adjectives, etc.). Then, according to the context semantics, determine the semantics of each segmented word, and combine the semantics of each segmented word to obtain the semantic information of the speech recognition text. The speech recognition text can be input into the semantic slot recognition model to obtain the semantic slots of the speech recognition text.

[0058] Optionally, the speech recognition text can also be input into the semantic model to obtain the semantic information of the speech recognition text; based on the semantic information, determine the semantic slots in the speech recognition text.

[0059] S103. Calculate the cloud confidence based on the speech recognition text and semantic information, and calculate the cloud similarity based on the semantic slots to obtain the first confidence, the second confidence, and the slot similarity.

[0060] Specifically, the speech recognition text and semantic information can be sent to the cloud, so that the cloud can calculate the confidence based on the speech recognition text to obtain the first confidence; calculate the confidence based on the semantic information to obtain the second confidence; calculate the similarity speed based on the semantic slots to obtain the slot similarity.

[0061] Optionally, the method for calculating the first confidence, the second confidence, and the slot similarity can also be: based on the speech recognition text, calculate the first confidence of the acoustic model in the cloud; based on the semantic information, calculate the second confidence of the semantic model in the cloud; based on the semantic slots, calculate the slot similarity of the semantic slots in the cloud.

[0062] Specifically, there are three detection links in the cloud, namely the cloud acoustic confidence module, which is used to calculate the first confidence of the acoustic model based on the speech recognition text; the cloud semantic confidence module, which is used to calculate the second confidence of the semantic model based on the semantic information; and the cloud dictionary matching module, which is used to calculate the slot similarity of the semantic slots based on the semantic slots.

[0063] S104. If a wake-up half-word is detected from the speech recognition text, and the first confidence, the second confidence, and the slot similarity are all greater than a preset threshold, then perform wake-up-free speech interaction based on the semantic information and the semantic slots.

[0064] Among them, the wake-up half-word is an incomplete wake-up word. Exemplarily, when waking up a system, the wake-up word is usually "XX, hello", etc., and at this time XX is the wake-up half-word.

[0065] Specifically, only when the speech recognition text contains a wake-up half-word (this setting is to avoid the voice interaction system being accidentally woken up, that is, to reduce the false alarm rate), and the first confidence, the second confidence, and the slot similarity are all greater than the preset threshold, then perform interaction with the user based on the semantic information and the semantic slots; when the speech recognition text does not contain a wake-up half-word, and / or, the first confidence, the second confidence, or the slot similarity is not greater than the preset threshold, then do not enter the voice interaction link.

[0066] Optionally, it can also be to detect the wake-up half-word in the speech recognition text; if there is a wake-up half-word in the speech recognition text, and the first confidence is greater than the first confidence threshold, the second confidence is greater than the second confidence threshold, and the slot similarity is greater than the slot similarity threshold, then perform wake-up-free speech interaction based on the semantic information and the semantic slots.

[0067] Specifically, relevant voice commands can be generated based on the semantic information and the relevant voice commands can be executed. It can also be to generate a reply voice based on the semantic information and the semantic slots, and feedback the reply voice to the user, and then obtain the user's voice interaction audio again.

[0068] For the method in the above embodiments, first obtain the voice interaction audio, perform voice recognition on the voice interaction audio to obtain the voice recognition text; then perform semantic recognition on the voice recognition text to obtain the semantic information and semantic slots of the voice recognition text; then calculate the cloud confidence based on the voice recognition text and semantic information and calculate the cloud similarity based on the semantic slots to obtain the first confidence, the second confidence, and the slot similarity; if a wake-up half-word is detected in the voice recognition text, and the first confidence, the second confidence, and the slot similarity are all greater than a preset threshold, then perform wake-up-free voice interaction based on the semantic information and semantic slots. In the above solution, by calculating the confidence and similarity in the cloud, the calculation process does not occupy local computing power and local memory. Only when the first confidence, the second confidence, and the slot similarity are all greater than the preset threshold, perform wake-up-free voice interaction based on the semantic information and semantic slots. The cloud calculation ensures the end-to-end achievement rate, so that the local computing power can focus on the processing of the voice interaction audio, reducing the system false alarm value, and thus realizing the reduction of the system false alarm value while ensuring the end-to-end achievement rate.

[0069] In another embodiment, there is also a case of multi-turn voice interaction. On the basis of the above embodiment, the wake-up-free voice interaction method further includes the following steps:

[0070] Judge whether multi-turn voice interaction is required based on the semantic information and semantic slots; if so, obtain the voice interaction audio again.

[0071] Exemplarily, when the semantic information is "play a song", and the semantic slot is a song at this time. From the current semantic information, it can be seen that the user does not specify what song to play. Therefore, it can be judged that multi-turn voice interaction is required. At this time, a reply voice can be generated according to the semantic information and semantic slots, such as "What song to play". After feeding the reply voice back to the user, obtain the voice interaction audio sent by the user again, and repeat the steps of S101-S104 above.

[0072] To more comprehensively demonstrate the present solution, this embodiment gives an alternative way of the wake-up-free voice interaction method, as Figure 2 shown:

[0073] S201, Obtain the voice interaction audio.

[0074] S202, Calculate the voice confidence in the voice interaction audio.

[0075] S203, If the voice confidence is greater than the voice confidence threshold, input the voice interaction audio into the acoustic model for voice recognition to obtain the voice recognition text.

[0076] S204. Input the speech recognition text into the semantic model to obtain the semantic information of the speech recognition text.

[0077] S205. Based on the semantic information, determine the semantic slots in the speech recognition text.

[0078] S206. Based on the speech recognition text, calculate the first confidence level of the acoustic model in the cloud.

[0079] S207. Based on the semantic information, calculate the second confidence level of the semantic model in the cloud.

[0080] S208. Based on the semantic slots, calculate the slot similarity of the semantic slots in the cloud.

[0081] S209. If a wake-up half-word is detected in the speech recognition text, and the first confidence level, the second confidence level, and the slot similarity are all greater than the preset threshold, then perform wake-up-free speech interaction based on the semantic information and the semantic slots.

[0082] Among them, the preset threshold includes a first confidence level threshold, a second confidence level threshold, and a slot similarity threshold.

[0083] Specifically, perform wake-up half-word detection on the speech recognition text; if there is a wake-up half-word in the speech recognition text, and the first confidence level is greater than the first confidence level threshold, the second confidence level is greater than the second confidence level threshold, and the slot similarity is greater than the slot similarity threshold, then perform wake-up-free speech interaction based on the semantic information and the semantic slots.

[0084] S210. Based on the semantic information and the semantic slots, determine whether multi-round speech interaction is required.

[0085] S211. If so, obtain the speech interaction audio again.

[0086] For the specific processes of the above S201 - S211, reference can be made to the description of the above method embodiments. Their implementation principles and technical effects are similar, and will not be elaborated here.

[0087] Based on the same inventive concept, an embodiment of the present application also provides a wake-up-free speech interaction device for implementing the above-mentioned wake-up-free speech interaction method. The implementation solution provided by this device for solving problems is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the following wake-up-free speech interaction device can refer to the limitations on the wake-up-free speech interaction method in the above text, and will not be elaborated here.

[0088] In one embodiment, as Figure 3 shown, a wake-up-free speech interaction device is provided. The device includes:

[0089] The voice recognition module 30 is used to obtain the voice interaction audio and perform voice recognition on the voice interaction audio to obtain the voice recognition text;

[0090] The semantic recognition module 31 is used to perform semantic recognition on the voice recognition text to obtain the semantic information and semantic slots of the voice recognition text;

[0091] The cloud computing module 32 is used to perform cloud confidence calculation based on the voice recognition text and semantic information and perform cloud similarity calculation based on the semantic slots to obtain the first confidence, the second confidence, and the slot similarity;

[0092] The voice interaction module 33 is used to, if a wake-up half-word is detected from the voice recognition text and the first confidence, the second confidence, and the slot similarity are all greater than a preset threshold, perform wake-up-free voice interaction based on the semantic information and semantic slots.

[0093] In another embodiment, as Figure 4 shown, the above Figure 3 voice recognition module 30 includes:

[0094] The confidence calculation unit 300 is used to calculate the voice confidence in the voice interaction audio;

[0095] The voice recognition unit 301 is used to, if the voice confidence is greater than the voice confidence threshold, input the voice interaction audio into the acoustic model for voice recognition to obtain the voice recognition text.

[0096] In another embodiment, as Figure 5 shown, the above Figure 3 semantic recognition module 31 includes:

[0097] The semantic acquisition unit 310 is used to input the voice recognition text into the semantic model to obtain the semantic information of the voice recognition text;

[0098] The slot determination unit 311 is used to determine the semantic slots in the voice recognition text based on the semantic information.

[0099] In another embodiment, as Figure 6 shown, the above Figure 3 cloud computing module 32 includes:

[0100] The first calculation unit 320 is used to calculate the first confidence of the acoustic model in the cloud based on the voice recognition text;

[0101] The second calculation unit 321 is used to calculate the second confidence of the semantic model in the cloud based on the semantic information;

[0102] A third computing unit 322, configured to calculate the slot similarity of the semantic slot in the cloud based on the semantic slot.

[0103] In another embodiment, the preset threshold includes a first confidence threshold, a second confidence threshold, and a slot similarity threshold. The voice interaction module 33 specifically: performs wake-up half-word detection on the voice recognition text; if there is a wake-up half-word in the voice recognition text, and the first confidence is greater than the first confidence threshold, the second confidence is greater than the second confidence threshold, and the slot similarity is greater than the slot similarity threshold, then performs wake-up-free voice interaction based on the semantic information and the semantic slot. Figure 3

[0104] In another embodiment, the wake-up-free voice interaction device is further specifically configured to: determine whether multi-round voice interaction is required based on the semantic information and the semantic slot; if so, obtain the voice interaction audio again.

[0105] Figure 7 An embodiment of the present application further provides an electronic device. In some embodiments, as shown in FIG. 700, the electronic device 700 includes an input unit 710, a memory 720, a processor 730, and an output unit 740. The memory 720 stores program instructions that can be run on the processor 730. The processor 730 can execute the wake-up-free voice interaction method and / or technical solution based on the foregoing embodiments by invoking the program instructions. The electronic device 700 can be a mobile terminal device such as a mobile phone or a computer.

[0106] In addition, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program for executing the wake-up-free voice interaction method. For example, computer program instructions, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. The program instructions for calling the method of the present application may be stored in a fixed or removable storage medium, and / or transmitted and / or stored in a storage medium running according to the program instructions through a data stream in a broadcast or other signal-bearing medium.

[0107] Obviously, those skilled in the art should understand that the foregoing modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.

[0108] The technical features of the above embodiments can be arbitrarily combined. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0109] The above embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A wake-up-free voice interaction method, characterized in that: The method comprises: Acquire voice interaction audio, and perform voice recognition on the voice interaction audio to obtain voice recognition text; Performing semantic recognition on the speech recognition text to obtain semantic information and semantic slots of the speech recognition text; Performing cloud confidence calculation based on the speech recognition text and semantic information, and performing cloud similarity calculation based on the semantic slot to obtain a first confidence, a second confidence, and a slot similarity; If a wake-up half word is detected from the speech recognition text, and the first confidence, the second confidence and the slot similarity are all greater than a preset threshold, a wake-up-free voice interaction is performed based on the semantic information and the semantic slot.

2. The wake-up-free voice interaction method according to claim 1, characterized in that: Performing voice recognition on the voice interaction audio to obtain voice recognition text includes: Calculating the human voice confidence in the voice interaction audio; If the human voice confidence is greater than the human voice confidence threshold, the voice interaction audio is input into the acoustic model for human voice recognition to obtain a voice recognition text.

3. The wake-up-free voice interaction method according to claim 2, characterized in that: Performing semantic recognition on the speech recognition text to obtain semantic information and semantic slots of the speech recognition text includes: Inputting the speech recognition text into a semantic model to obtain semantic information of the speech recognition text; Based on the semantic information, a semantic slot in the speech recognition text is determined.

4. The wake-up-free voice interaction method according to claim 3, characterized in that: Performing cloud confidence calculation based on the speech recognition text and semantic information, and performing cloud similarity calculation based on the semantic slot to obtain a first confidence, a second confidence, and a slot similarity, including: Based on the speech recognition text, calculating a first confidence of the acoustic model in the cloud; Based on the semantic information, calculating a second confidence level of the semantic model in the cloud; Based on the semantic slots, slot similarities of the semantic slots are calculated in the cloud.

5. The wake-up-free voice interaction method according to claim 1, characterized in that: The preset threshold includes a first confidence threshold, a second confidence threshold and a slot similarity threshold. If a wake-up half word is detected from the speech recognition text, and the first confidence, the second confidence and the slot similarity are all greater than the preset threshold, then wake-up-free voice interaction is performed based on the semantic information and the semantic slot, including: Performing wake-up half-word detection on the speech recognition text; If there is a wake-up half word in the speech recognition text, and the first confidence is greater than the first confidence threshold, the second confidence is greater than the second confidence threshold, and the slot similarity is greater than the slot similarity threshold, then wake-up-free voice interaction is performed based on the semantic information and the semantic slot.

6. The wake-up-free voice interaction method according to claim 1, characterized in that: The method further comprises: Determining whether multiple rounds of voice interaction are required based on the semantic information and the semantic slots; If yes, the voice interaction audio is obtained again.

7. A wake-up-free voice interaction device, characterized in that: The device comprises: A voice recognition module is used to obtain voice interaction audio and perform voice recognition on the voice interaction audio to obtain voice recognition text; A semantic recognition module, used for performing semantic recognition on the speech recognition text to obtain semantic information and semantic slots of the speech recognition text; A cloud computing module, used for performing cloud confidence calculation based on the speech recognition text and semantic information and performing cloud similarity calculation based on the semantic slot to obtain a first confidence, a second confidence and a slot similarity; A voice interaction module is used to perform wake-up-free voice interaction based on the semantic information and the semantic slot if a wake-up half-word is detected from the voice recognition text and the first confidence, the second confidence and the slot similarity are all greater than a preset threshold.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the wake-up-free voice interaction method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the wake-up-free voice interaction method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the wake-up-free voice interaction method according to any one of claims 1 to 6 is implemented.