A voice interaction method and device, electronic equipment and storage medium
By acquiring and analyzing historical voice interaction data, and selecting historical recognized text with high similarity or high priority for response, the problem of poor accuracy in voice interaction is solved, and more efficient response to user needs is achieved.
Patent Information
- Application Number
- CN202310268457.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-03-15
AI Technical Summary
In existing voice interaction technologies, the accuracy of voice interaction is poor, resulting in the inability to accurately respond to user needs.
By acquiring user confirmation text from historical voice interaction data, the matching status between the speech to be recognized and historical speech is determined, and the corresponding historical recognition text is selected for response based on similarity and priority, or speech recognition is performed in the case of low similarity to obtain the target text.
It improves the accuracy and efficiency of voice interaction, enabling it to meet user needs more quickly and reducing the tedious voice recognition process.
Smart Images

Figure CN116403578B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a voice interaction method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, intelligent voice technology has become widespread, and most devices with screens support voice interaction, allowing users to complete tasks quickly and efficiently. However, the accuracy of voice interaction is relatively poor. Summary of the Invention
[0003] The main technical problem addressed by this application is to provide a voice interaction method, device, electronic device, and storage medium that can improve the accuracy of voice interaction.
[0004] To address the aforementioned technical problems, this application provides a voice interaction method, comprising: acquiring a first voice to be recognized and several historical voice interaction data; wherein the historical voice interaction data includes historical recognized voice and historical recognized text, the historical recognized text being selected by the user from reference recognized text corresponding to the historical recognized voice, and the reference recognized text being obtained based on voice recognition of the historical recognized voice; determining the matching status between the first voice to be recognized and each historical recognized voice; and responding to the user based on the historical recognized text corresponding to the historical recognized voice that matches the first voice to be recognized.
[0005] The process of determining the matching status between the first speech to be recognized and each historical speech includes: determining the similarity between the first speech to be recognized and each historical speech; and responding to the user based on the historical recognized text corresponding to the historical recognized speech that matches the first speech to be recognized, including: responding to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity.
[0006] The process of responding to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity includes: responding to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity when the highest similarity is greater than or equal to a similarity threshold; and responding to the user based on the first target recognized text when the highest similarity is less than a similarity threshold.
[0007] The highest similarity historical recognized speech includes at least two, and the corresponding historical recognized texts are all different. Based on the historical recognized text corresponding to the highest similarity historical recognized speech, the user is responded to, including: determining the priority of the historical recognized text corresponding to each highest similarity historical recognized speech; wherein the priority of the historical recognized text is used to characterize the likelihood that the historical recognized text meets the user's needs; and the user is responded to based on the historical recognized text with the highest priority.
[0008] The process of determining the priority of the historical recognized text corresponding to the historical recognized speech with the highest similarity includes: obtaining the time when the historical recognized text corresponding to the highest similarity was last selected; and determining the priority of the historical recognized text corresponding to the highest similarity based at least on the time when the historical recognized text corresponding to the highest similarity was last selected.
[0009] The process of obtaining the most recent selection time of the historical recognition text corresponding to the highest similarity score includes: obtaining the first time when the historical recognition text corresponding to the highest similarity score was most recently selected on the current device and the second time when it was most recently selected on the associated device; and determining the priority of the historical recognition text corresponding to the highest similarity score based at least on the most recent selection time of the historical recognition text corresponding to the highest similarity score, including: determining the priority of the historical recognition text corresponding to the highest similarity score based at least on the most recent selection time of the historical recognition text corresponding to the highest similarity score.
[0010] Specifically, the priority of each historical recognition text corresponding to the highest similarity is determined based on at least the time when the historical recognition text corresponding to the highest similarity was most recently selected. This includes: obtaining the number of times each historical recognition text corresponding to the highest similarity was selected within a preset time period; and for each historical recognition text corresponding to the highest similarity, determining the priority of the historical recognition text corresponding to the highest similarity based on the number of times the historical recognition text corresponding to the highest similarity was selected within the preset time period and the time when it was most recently selected.
[0011] The voice interaction method further includes, after performing speech recognition on the first speech to be recognized in response to the highest similarity being less than the similarity threshold, and responding to the user based on the first target recognition text obtained from the recognition, acquiring a second speech to be recognized; performing speech recognition on the second speech to be recognized to obtain the speech intent of the second speech to be recognized and several initial recognition texts corresponding to the second speech to be recognized; in response to the second speech to be recognized being consistent with the first speech to be recognized, filtering the several initial recognition texts based on the speech intent to obtain a second target recognition text, and displaying the second target recognition text on the display interface; and in response to the user's selection of any second target recognition text, responding to the user based on the second target recognition text selected by the user.
[0012] The process of filtering several initial recognition texts based on speech intent to obtain the second target recognition text includes: obtaining the type of execution content corresponding to each initial recognition text; and removing initial recognition texts whose execution content type does not match the speech intent from several initial recognition texts to obtain the second target recognition text.
[0013] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a voice interaction device, which includes an acquisition module, a determination module, and a response module; the acquisition module is used to acquire a first voice to be recognized and several historical voice interaction data; wherein, the historical voice interaction data includes historical recognized voice and historical recognized text, the historical recognized text being selected by the user from reference recognized text corresponding to the historical recognized voice, and the reference recognized text being obtained based on voice recognition of the historical recognized voice; the determination module is used to determine the matching status between the first voice to be recognized and each historical recognized voice; the response module is used to respond to the user based on the historical recognized text corresponding to the historical recognized voice that matches the first voice to be recognized.
[0014] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide an electronic device, which includes a memory and a processor. The memory stores program instructions, and the processor executes the program instructions to implement the above-mentioned voice interaction method.
[0015] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium for storing program instructions that can be executed to implement the above-mentioned voice interaction method.
[0016] In the above technical solution, the historical recognized text corresponding to the historical recognized speech is selected by the user from the reference recognized text corresponding to the historical recognized speech. The reference recognized text is obtained based on speech recognition of the historical recognized speech. That is, the historical recognized text corresponding to the historical recognized speech is selected and confirmed by the user. In other words, when the user inputs the historical recognized speech, the content the user actually wants to trigger is the historical recognized text corresponding to the historical recognized speech. Therefore, responding to the user based on the historical recognized text corresponding to the historical recognized speech that matches the first speech to be recognized, on the one hand, greatly increases the probability of triggering the content that the user actually needs, improving the accuracy of voice interaction; on the other hand, it eliminates the need for cumbersome speech recognition, improving the efficiency of voice interaction and enabling faster response to user needs. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an embodiment of the voice interaction method provided in this application;
[0018] Figure 2 This is a flowchart illustrating an embodiment of the present application of responding to users based on the historically recognized text corresponding to the historically recognized speech with the highest similarity.
[0019] Figure 3 yes Figure 2 The flowchart of step S21 shown is a schematic diagram of one embodiment;
[0020] Figure 4 This is a flowchart illustrating another embodiment of the voice interaction method provided in this application;
[0021] Figure 5 This is a schematic diagram of an embodiment of the first voice interaction interface provided in this application;
[0022] Figure 6 This is a schematic diagram of an embodiment of the second voice interaction interface provided in this application;
[0023] Figure 7 This is a schematic diagram of the structure of an embodiment of the voice interaction device provided in this application;
[0024] Figure 8 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application;
[0025] Figure 9 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation
[0026] To make the purpose, technical solution and effects of this application clearer and more explicit, the following describes this application in further detail with reference to the accompanying drawings and embodiments.
[0027] It should be noted that if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the voice interaction method provided in this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily replace it. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, this embodiment includes:
[0029] Step S11: Obtain the first speech to be recognized and some historical speech interaction data.
[0030] The method in this embodiment aims to improve the accuracy and convenience of user interaction with screen-equipped products, thereby enhancing the user experience. Screen-equipped products, as described herein, include, but are not limited to, in-vehicle voice assistants, mobile phones, tablets, and home appliances. These products support voice interaction, allowing users to complete tasks quickly and efficiently. For example, using an in-vehicle voice assistant as an example, users can quickly play music, open the navigation system, etc., through voice interaction with the in-vehicle voice assistant.
[0031] In one embodiment, the first speech to be recognized can be obtained from local storage or cloud storage, or it can be acquired in real time by a speech acquisition device; no specific limitation is made here. The language, length, etc., of the first speech to be recognized are not limited and can be specifically set according to actual usage needs; for example, the language of the first speech to be recognized may be English, Chinese, Japanese, Sichuan dialect, or other local dialects.
[0032] Because the text obtained from speech recognition of the first speech to be recognized may be inaccurate, directly responding to the user based on this text would lead to inaccurate voice interaction. However, the historical text corresponding to previously recognized speech is selected by the user from reference texts corresponding to those historical speech. These reference texts are derived from speech recognition of historical speech; in other words, the historical text corresponding to a previous speech is selected and confirmed by the user. Therefore, when a similar speech to be recognized appears later, the user-selected and confirmed historical text is used as the text to respond to the user. Compared to directly responding based on the text obtained from speech recognition of the first speech to be recognized, this method offers higher accuracy and efficiency in voice interaction.
[0033] Therefore, in this embodiment, several historical voice interaction data will also be acquired. These historical voice interaction data include historical recognized speech and historical recognized text. The historical recognized text is selected by the user from the reference recognized text corresponding to the historical recognized speech, and the reference recognized text is obtained based on speech recognition of the historical recognized speech. For example, the historical recognized speech is "I want to hear Nan Que," and the reference recognized text corresponding to the historical recognized speech "I want to hear Nan Que" is "I want to hear Nan Que," "I want to hear Nan Que," and "I want to hear Nan Que." The recognized text selected by the user from the three reference recognized texts is "I want to hear Nan Que." Therefore, the historical recognized text corresponding to the historical recognized speech "I want to hear Nan Que" is the reference recognized text "I want to hear Nan Que."
[0034] In one implementation, historical voice interaction data can be obtained from local storage or cloud storage, without any specific limitation.
[0035] Step S12: Determine the matching status between the first speech to be recognized and each historical speech.
[0036] In this embodiment, the matching status between the first speech to be recognized and each historical recognized speech is determined. In one embodiment, the matching status between the first speech to be recognized and each historical recognized speech is determined by determining the similarity between the first speech to be recognized and each historical recognized speech. In a specific embodiment, algorithms such as Dynamic Time Warping (DTW) and SimHash can be used to determine the similarity between the first speech to be recognized and each historical recognized speech; no specific limitation is made here.
[0037] Step S13: Respond to the user based on the historical recognized text corresponding to the historical recognized speech that matches the first speech to be recognized.
[0038] In this embodiment, the user is responded to based on the historical recognized text corresponding to the historical recognized speech that matches the first speech to be recognized. The historical recognized text is selected by the user from reference recognized text corresponding to the historical recognized speech. The reference recognized text is obtained by performing speech recognition on the historical recognized speech; that is, the historical recognized text corresponding to the historical recognized speech is selected and confirmed by the user. In other words, when the user inputs the historical recognized speech, the content the user actually wants to trigger is the historical recognized text corresponding to that speech. Therefore, when a speech to be recognized that matches the historical recognized speech is subsequently received, it indicates that the speech to be recognized is similar to the matched historical recognized speech, and there is a high probability that the content the user actually wants to trigger when inputting the speech to be recognized is the historical recognized text corresponding to the matched historical recognized speech. Therefore, responding to the user based on the historical recognized text corresponding to the historical recognized speech that matches the first speech to be recognized, on the one hand, greatly increases the probability of triggering the content the user actually needs, improving the accuracy of voice interaction; on the other hand, it eliminates the need for cumbersome speech recognition, improving the efficiency of voice interaction and enabling faster responses to user needs.
[0039] In one embodiment, determining the matching status of the first speech to be recognized with each historical recognized speech specifically involves determining the similarity between the first speech to be recognized and each historical recognized speech. At this point, the user is responded to based on the historical recognized text corresponding to the historical recognized speech with the highest similarity. A higher similarity indicates a greater degree of similarity between the first speech to be recognized and the historical recognized speech, and a greater likelihood that the content of the first speech to be recognized and the historical recognized speech are consistent. Therefore, responding to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity is more likely to trigger the content that the user actually needs, thus improving the accuracy of voice interaction.
[0040] In one specific implementation, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating an embodiment of responding to a user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity, as provided in this application. The historical recognized speech with the highest similarity to the first speech to be recognized includes at least two such speeches, and their corresponding historical recognized texts are all different. The response to the user is based on the historical recognized text corresponding to the historical recognized speech with the highest similarity, specifically including the following sub-steps:
[0041] Step S21: Determine the priority of the historical recognized text corresponding to each historical recognized speech with the highest similarity.
[0042] The historical recognized speech that has the highest similarity to the first speech to be recognized includes at least two, and the corresponding historical recognized texts are different from each other. That is, there are multiple historical recognized speech that are highly similar to the first speech to be recognized, and the historical recognized texts selected and confirmed by the user are different. In this case, it is impossible to determine which historical recognized speech corresponds to which historical recognized text to select. This can more accurately trigger the content that the user really needs, thereby improving the accuracy of voice interaction.
[0043] Therefore, to improve the accuracy of voice interaction, this embodiment determines the priority of the historical recognized text corresponding to each historical recognized speech with the highest similarity. The priority of the historical recognized text is used to characterize the likelihood that the historical recognized text meets the user's needs. Since the priority of the historical recognized text indicates its likelihood of meeting the user's needs, determining the priority of the historical recognized text corresponding to each historical recognized speech with the highest similarity facilitates subsequent responses to the user based on recognized text that better matches the user's needs, more accurately hitting the user's needs, and improving the accuracy and convenience of voice interaction.
[0044] In one implementation, such as Figure 3 As shown, Figure 3 yes Figure 2 The flowchart shown in step S21 is a schematic diagram of an embodiment. It determines the priority of the historical recognized text corresponding to each historical recognized speech with the highest similarity, specifically including the following sub-steps:
[0045] Step S31: Obtain the time when the historical text with the highest similarity was last selected.
[0046] In this embodiment, the system obtains the most recent selection time of the historical recognized text corresponding to each historical recognized speech with the highest similarity. In a specific embodiment, the most recent selection time of the historical recognized text corresponding to each historical recognized speech with the highest similarity can be obtained from the user's personalized dictionary; specifically, the system stores the recognized text data previously selected by the user to form the user's personalized dictionary, and records the user's selection time synchronously when storing the recognized text.
[0047] In one implementation, the time when the historical recognized text corresponding to each historically recognized speech with the highest similarity is most recently selected is obtained, specifically the time when it is most recently selected on the current device. For example, taking a user's voice interaction with an in-vehicle voice assistant as an example, what is obtained is the time when the historically recognized text corresponding to each historically recognized speech with the highest similarity is most recently selected by the in-vehicle voice assistant.
[0048] In other implementations, the time when the historical recognized text corresponding to each historically recognized speech with the highest similarity was most recently selected is obtained. Specifically, this includes the first time it was most recently selected on the current device and the second time it was most recently selected on the associated device. This allows for a comprehensive consideration of both the time the historically recognized text corresponding to each historically recognized speech with the highest similarity was most recently selected on the current device and the time it was most recently selected on the associated device to determine the priority of each historically recognized text with the highest similarity. It should be noted that the associated device is a device capable of establishing a communication connection with the current device. After establishing a communication connection, the time when the historically recognized text corresponding to each historically recognized speech with the highest similarity was most recently selected on that associated device can be obtained. For example, taking an in-vehicle voice assistant as the current device and a mobile phone as the associated device: First, the first time when the historically recognized speech with the highest similarity was most recently selected on the in-vehicle voice assistant is obtained; second, after the in-vehicle voice assistant establishes a communication connection with the mobile phone, the second time when the historically recognized text corresponding to each historically recognized speech with the highest similarity was most recently selected on the mobile phone is obtained through the mobile phone.
[0049] Step S32: Determine the priority of each historical recognition text corresponding to the highest similarity based at least on the time when the historical recognition text corresponding to the highest similarity was most recently selected.
[0050] In this embodiment, the priority of each historically recognized text corresponding to the most similar historically recognized speech is determined at least based on the time when it was most recently selected. In one embodiment, the time of most recently selected is negatively correlated with the priority; that is, the earlier the time when the historically recognized text corresponding to the most similar historically recognized speech was most recently selected, the lower its priority and the less likely it is to meet the user's needs. Conversely, the later the time when the historically recognized text corresponding to the most similar historically recognized speech was most recently selected, the more likely the user is to select that historically recognized text again, thus giving it a higher priority. Of course, in other embodiments, the time of most recently selected can be positively correlated with the priority, and this is not specifically limited here.
[0051] In one embodiment, the time when the historical recognized text corresponding to each historical recognized speech with the highest similarity was most recently selected is obtained, specifically the time when it was most recently selected on the current device. At this time, the priority of each historical recognized text corresponding to the highest similarity is determined based at least on the first time when it was most recently selected on the current device. In other embodiments, the time when the historical recognized text corresponding to each historical recognized speech with the highest similarity was most recently selected is obtained, specifically the first time when it was most recently selected on the current device and the second time when it was most recently selected on the associated device. At this time, the priority of each historical recognized text corresponding to the highest similarity is determined based at least on the first and second times, that is, by comprehensively considering both the time when it was most recently selected on the current device and the time when it was most recently selected on the associated device, the priority of each historical recognized text corresponding to the highest similarity is determined.
[0052] In one embodiment, to improve the accuracy of determining the priority of the historical recognized texts corresponding to the historical recognized speech with the highest similarity, the priority can also be determined based on the number of times the texts were selected within a preset time period and the time of the most recent selection. Specifically, firstly, the number of times each historical recognized text corresponding to the highest similarity was selected within the preset time period is obtained; then, for each historical recognized text corresponding to the highest similarity, the priority is determined based on the number of times it was selected within the preset time period and the time of the most recent selection. In other words, for each historical recognized text corresponding to the highest similarity, the priority is determined by combining both the number of times it was selected and the time of the most recent selection.
[0053] Step S22: Respond to the user based on the historical recognized text with the highest priority.
[0054] In this implementation, the response to the user is based on the historically recognized text with the highest priority. Responding to the user based on the historically recognized text with the highest priority allows for a more accurate match with the user's needs, precisely triggering the content the user truly requires, thereby improving the accuracy of voice interaction.
[0055] In other specific implementations, there are at least two historical recognized voices that have the highest similarity to the first voice to be recognized; in this case, the user may also be responded to randomly based on the historical recognized text corresponding to the historical recognized voice with the highest similarity.
[0056] In other specific embodiments, at least two historical recognized voices have the highest similarity to the first voice to be recognized, and the historical recognized texts corresponding to these two highest-similar historical recognized voices are identical. In this case, the user is responded to based on the historical recognized text that corresponds to the historical recognized voice with the highest similarity. The historical recognized text that corresponds to the historical recognized voice with the highest similarity is the one that has been selected and confirmed by the user multiple times under that historical recognized voice. Therefore, when a voice to be recognized that matches this historical recognized voice is subsequently received, it indicates that the voice to be recognized is similar to the matched historical recognized voice, and there is a high probability that the content the user actually wants to trigger when inputting the voice to be recognized is the historical recognized text that corresponds to the historical recognized voice with the highest similarity. Therefore, responding to the user based on the historical recognized text that corresponds to the historical recognized voice with the highest similarity greatly increases the likelihood of triggering the content that the user actually needs, thus improving the accuracy of voice interaction.
[0057] Considering that the historical recognized speech with the highest similarity to the first speech to be recognized may actually have a relatively low similarity to the first speech to be recognized, responding to the user based on the historical recognized text corresponding to that historical recognized speech would likely trigger content that does not match the user's needs. Therefore, to improve the accuracy of voice interaction, in one embodiment, in response to a highest similarity greater than or equal to a similarity threshold, the user is responded to based on the historical recognized text corresponding to the historical recognized speech with the highest similarity; in response to a highest similarity less than the similarity threshold, speech recognition is performed on the first speech to be recognized, and the user is responded to based on the recognized first target text. In other words, when the similarity to the historical recognized speech with the highest similarity to the first speech to be recognized is greater than or equal to the similarity threshold, it indicates that the first speech to be recognized is highly similar to the speech with the highest similarity. Responding to the user based on its corresponding historical recognized text is highly likely to trigger the content that the user actually needs, thus improving the accuracy of voice interaction and the user experience. On the other hand, when the similarity to the historical recognized speech with the highest similarity to the first speech to be recognized is less than the similarity threshold, it indicates that the first speech to be recognized is relatively low in similarity to several historical recognized speech, and may be different from each of the historical recognized speech. In this case, it is necessary to perform speech recognition on the first speech to be recognized and respond to the user based on the first target recognized text obtained from the recognition.
[0058] Since the content executed in response to the first target recognition text may not match the user's expectations, upon receiving the second speech input from the user, execution is triggered to display the recognized text. This allows for rapid modification and correction should the executed content not match the user's expectations. Therefore, in one embodiment, as... Figure 4 As shown, Figure 4 This is a flowchart illustrating another embodiment of the voice interaction method provided in this application. After performing voice recognition on the first speech to be recognized in response to the highest similarity being less than a similarity threshold, and responding to the user based on the recognized first target text, the method further includes the following steps:
[0059] Step S41: Obtain the second speech to be recognized.
[0060] In one embodiment, the second speech to be recognized can be obtained from local storage or cloud storage. Of course, in other embodiments, the second speech to be recognized can also be acquired in real time using a speech acquisition device, and no specific limitation is made here.
[0061] Step S42: Perform speech recognition on the second speech to be recognized to obtain the speech intent of the second speech to be recognized and several initial recognition texts corresponding to the second speech to be recognized.
[0062] In this embodiment, speech recognition is performed on the second speech to be recognized to obtain several initial recognition texts corresponding to the second speech to be recognized; that is, data processing is performed on the second speech to be recognized, including recognition and speech understanding, to obtain several initial recognition texts corresponding to the second speech to be recognized. For example, taking the second speech to be recognized as "I want to hear the magpie" as an example, speech recognition is performed on the second speech to be recognized "I want to hear the magpie", and three initial recognition texts are obtained, namely "I want to hear the magpie", "I want to hear the sparrow", and "I want to hear the peacock".
[0063] In one embodiment, speech recognition of the second speech to be recognized can be performed using algorithms based on dynamic time warping, parametric Hidden Markov Models (HMMs), nonparametric Vector Quantization (VQ) models, or deep learning, to obtain several initial recognition texts corresponding to the second speech to be recognized.
[0064] Because the content to be executed by the initial recognition texts obtained from the speech recognition of the second speech to be recognized may not match the speech intent corresponding to the second speech to be recognized, the accuracy of the initial recognition texts is not high, and it is impossible to directly display the initial recognition texts for the user to select. Therefore, in this embodiment, the second speech to be recognized is also subjected to speech recognition to obtain the speech intent corresponding to the second speech to be recognized, so as to filter the initial recognition texts in the subsequent process, thereby improving the accuracy of subsequent voice interaction and improving the user experience. For example, taking the second speech to be recognized as "I want to listen to Nan Que", and the initial recognition texts as "I want to listen to Nan Que", "I want to listen to Nan Que", and "I want to listen to Nan Que" as examples; it can be seen from the second speech to be recognized that the user's speech intent is to listen to a song, but the content "Nan Que" in "I want to listen to Nan Que" is not a song, so displaying the initial recognition text "I want to listen to Nan Que" as the recognition text along with the other two initial recognition texts at the same time makes the recognition text obtained from the second speech to be recognized inaccurate, thus affecting the accuracy of voice interaction.
[0065] In one embodiment, the second speech to be recognized can be recognized based on rules, traditional machine learning algorithms (SVM) or deep learning algorithms (CNN, LSTM, RCNN, C-LSTM or FastText) to obtain the speech intent corresponding to the second speech to be recognized.
[0066] It should be noted that there is no limitation on the language of the second speech to be recognized; it can be set according to the actual needs of use; for example, English, Chinese, Japanese, Sichuan dialect, etc.
[0067] Step S43: In response to the second speech to be recognized being consistent with the first speech to be recognized, filter several initial recognition texts based on the speech intent to obtain the second target recognition text, and display the second target recognition text on the display interface.
[0068] In this embodiment, in response to the second voice to be recognized being the same as the first voice to be recognized, a number of initial recognized texts are filtered based on the voice intention to obtain a second target recognized text, and the second target recognized text is displayed on the display interface. That is to say, when the second voice to be recognized is the same as the first voice to be recognized, it may be that the currently executed content does not match the user's expectation. Therefore, the user inputs the second voice to be recognized, and then uses the voice intention corresponding to the second voice to be recognized to filter the obtained number of initial recognized texts, so as to filter out the initial recognized texts that do not match the voice intention, improving the accuracy of the obtained second target recognized text, that is, improving the matching degree between the second target recognized text and the user's needs. In addition, the second target recognized text will be displayed on the display interface, that is, the voice interaction interface will be presented, so that the user can more intuitively see the voice recognition result, facilitating the user to make a quick selection based on the needs subsequently (such as, manual selection or voice interaction selection, etc.) to trigger the content that the user actually needs, improving the accuracy and convenience of voice interaction, reducing the disturbance to the user and enhancing the user experience.
[0069] For example, as Figure 5 shown, Figure 5 is a schematic diagram of an embodiment of the first voice interaction interface provided by the present application, taking the second voice to be recognized as "I want to listen to Nan Que", and the initial recognized texts being "I want to listen to Nan Que", "I want to listen to Nan Que", and "I want to listen to Nan Que" as examples; it can be seen from the second voice to be recognized that the user's voice intention is to listen to a song, and the content "Nan Que" of "I want to listen to Nan Que" is not a song, so the initial recognized text "I want to listen to Nan Que" is removed, and the initial recognized texts "I want to listen to Nan Que" and "I want to listen to Nan Que" are displayed on the display interface.
[0070] In one embodiment, the initial recognized texts are filtered based on the matching situation between the type of the execution content corresponding to the initial recognized text and the voice intention. Specifically, the type of the execution content corresponding to each initial recognized text is obtained; from a number of initial recognized texts, the initial recognized texts whose type of execution content does not match the voice intention are removed to obtain a second target recognized text; for example, the voice intention is to listen to a song, and the type of the execution content "Nan Que" of the initial recognized text "I want to listen to Nan Que" is not a song, so the initial recognized text "I want to listen to Nan Que" is removed. Of course, in other embodiments, the initial recognized texts can also be filtered based on the voice intention by other means, which are not specifically limited herein.
[0071] Step S44: In response to the user's selection of any second target recognized text, respond to the user based on the second target recognized text selected by the user.
[0072] In this embodiment, in response to the user's selection of any second target recognition text, a response is given to the user based on the selected second target recognition text. In other words, the user can quickly select according to their needs on the display interface (manual selection or voice interaction, etc.) to quickly trigger the content they actually need, improving the accuracy and convenience of voice interaction.
[0073] For example, such as Figure 6 As shown, Figure 6 This is a schematic diagram of an embodiment of the second voice interaction interface provided in this application; taking the user selecting any second target recognition text through voice interaction as an example, the user voice inputs "first", the first option is "I want to listen to Nan Que", and at this time the song Nan Que is played.
[0074] In the above implementation, the historical recognized text corresponding to the historical recognized speech is selected by the user from the reference recognized text corresponding to the historical recognized speech. The reference recognized text is obtained based on speech recognition of the historical recognized speech. That is, the historical recognized text corresponding to the historical recognized speech is selected and confirmed by the user. In other words, when the user inputs the historical recognized speech, the content the user actually wants to trigger is the historical recognized text corresponding to the historical recognized speech. Therefore, responding to the user based on the historical recognized text corresponding to the historical recognized speech that matches the first speech to be recognized, on the one hand, greatly increases the probability of triggering the content that the user actually needs, improving the accuracy of voice interaction; on the other hand, it eliminates the need for cumbersome speech recognition, improving the efficiency of voice interaction and enabling faster response to user needs.
[0075] Please see Figure 7 , Figure 7 This is a schematic diagram of an embodiment of the voice interaction device provided in this application. The voice interaction device 70 includes an acquisition module 71, a determination module 72, and a response module 73. The acquisition module 71 is used to acquire a first voice to be recognized and a plurality of historical voice interaction data; wherein, the historical voice interaction data includes historical recognized voice and historical recognized text, the historical recognized text being selected by the user from the reference recognized text corresponding to the historical recognized voice, and the reference recognized text being obtained based on the voice recognition of the historical recognized voice; the determination module 72 is used to determine the matching status between the first voice to be recognized and each historical recognized voice; the response module 73 is used to respond to the user based on the historical recognized text corresponding to the historical recognized voice that matches the first voice to be recognized.
[0076] The determining module 72 is used to determine the matching status between the first speech to be recognized and each historical speech, specifically including: determining the similarity between the first speech to be recognized and each historical speech; the responding module 73 is used to respond to the user based on the historical recognized text corresponding to the historical recognized speech that matches the first speech to be recognized, specifically including: responding to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity.
[0077] The response module 73 is used to respond to the user based on the historical recognition text corresponding to the historical recognition speech with the highest similarity. Specifically, it includes: responding to the user based on the historical recognition text corresponding to the historical recognition speech with the highest similarity when the highest similarity is greater than or equal to the similarity threshold; and responding to the user based on the first target recognition text obtained when the highest similarity is less than the similarity threshold.
[0078] The aforementioned historical recognized speech with the highest similarity includes at least two, and the corresponding historical recognized texts are all different. The response module 73 is used to respond to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity, specifically including: determining the priority of the historical recognized text corresponding to each historical recognized speech with the highest similarity; wherein, the priority of the historical recognized text is used to characterize the possibility that the historical recognized text meets the user's needs; and responding to the user based on the historical recognized text with the highest priority.
[0079] The response module 73 is used to determine the priority of the historical recognized text corresponding to the historical recognized speech with the highest similarity. Specifically, it includes: obtaining the time when the historical recognized text corresponding to the highest similarity was last selected; and determining the priority of the historical recognized text corresponding to the highest similarity based at least on the time when the historical recognized text corresponding to the highest similarity was last selected.
[0080] The response module 73 is used to obtain the time when the historical recognition text corresponding to the highest similarity was last selected, specifically including: obtaining the first time when the historical recognition text corresponding to the highest similarity was last selected on the current device and the second time when it was last selected on the associated device; the response module 73 is used to determine the priority of the historical recognition text corresponding to the highest similarity based at least on the time when the historical recognition text corresponding to the highest similarity was last selected, specifically including: determining the priority of the historical recognition text corresponding to the highest similarity based at least on the first time and the second time corresponding to the historical recognition text corresponding to the highest similarity.
[0081] The response module 73 is used to determine the priority of each historical recognition text corresponding to the highest similarity, based at least on the time when the historical recognition text corresponding to the highest similarity was most recently selected. Specifically, it includes: obtaining the number of times each historical recognition text corresponding to the highest similarity was selected within a preset time period; and for each historical recognition text corresponding to the highest similarity, determining the priority of the historical recognition text corresponding to the highest similarity based on the number of times the historical recognition text corresponding to the highest similarity was selected within the preset time period and the time when it was most recently selected.
[0082] The voice interaction device 70 further includes a recognition module 74. The recognition module 74 is used to perform speech recognition on the first speech to be recognized in response to the highest similarity being less than the similarity threshold, and to respond to the user based on the first target recognition text obtained by recognition. Specifically, the recognition module 74 includes: acquiring a second speech to be recognized; performing speech recognition on the second speech to be recognized to obtain the speech intent of the second speech to be recognized and several initial recognition texts corresponding to the second speech to be recognized; in response to the second speech to be recognized being consistent with the first speech to be recognized, filtering the several initial recognition texts based on the speech intent to obtain the second target recognition text, and displaying the second target recognition text on the display interface; and in response to the user's selection of any second target recognition text, responding to the user based on the second target recognition text selected by the user.
[0083] The recognition module 74 is used to filter several initial recognition texts based on the voice intent to obtain the second target recognition text. Specifically, it includes: obtaining the type of execution content corresponding to each initial recognition text; removing initial recognition texts whose execution content type does not match the voice intent from several initial recognition texts to obtain the second target recognition text.
[0084] Please see Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the electronic device provided in this application. The electronic device 80 includes a memory 81 and a processor 82 coupled to each other. The processor 82 is used to execute program instructions stored in the memory 81 to implement the steps of any of the above-described voice interaction method embodiments. In a specific implementation scenario, the electronic device 80 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 80 may also include mobile devices such as laptops and tablets, which are not limited here.
[0085] Specifically, processor 82 controls itself and memory 81 to implement the steps of any of the above-described voice interaction method embodiments. Processor 82 may also be referred to as a CPU (Central Processing Unit). Processor 82 may be an integrated circuit chip with signal processing capabilities. Processor 82 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 82 may be implemented using integrated circuit chips.
[0086] Please see Figure 9 , Figure 9 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 90 of this embodiment stores program instructions 91, which, when executed, implement the methods provided by any embodiment of the voice interaction method and any non-conflicting combination thereof. The program instructions 91 can form a program file and be stored in the aforementioned computer-readable storage medium 90 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) can execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned computer-readable storage medium 90 includes various media capable of storing program code, such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.
[0087] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0088] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A voice interaction method, characterized in that, The method comprises: obtaining a first to-be-recognized speech and historical speech interaction data; wherein the historical speech interaction data comprises historical recognized speech and historical recognized text, the historical recognized text being selected by a user from reference recognized text corresponding to the historical recognized speech, the reference recognized text being obtained based on speech recognition of the historical recognized speech; determining matching conditions of the first to-be-recognized speech and each historical recognized speech; responding to the user based on historical recognized text corresponding to historical recognized speech matching the first to-be-recognized speech; wherein the historical recognized text corresponding to the historical recognized speech matching the first to-be-recognized speech is the historical recognized text with the highest priority among historical recognized text corresponding to at least two historical recognized speeches with the highest similarity to the first to-be-recognized speech, the priority of the historical recognized text corresponding to each of the historical recognized speeches with the highest similarity being determined based on at least the time when the historical recognized text corresponding to each of the historical recognized speeches with the highest similarity was last selected, and the historical recognized text corresponding to the at least two historical recognized speeches with the highest similarity being different from each other.
2. The method of claim 1, wherein, The determination of the matching conditions of the first to-be-recognized speech and each historical recognized speech comprises: determining the similarity of the first to-be-recognized speech and each historical recognized speech; The response to the user based on the historical recognized text corresponding to the historical recognized speech matching the first to-be-recognized speech comprises: responding to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity.
3. The method of claim 2, wherein, The response to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity comprises: in response to the highest similarity being greater than or equal to a similarity threshold, responding to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity; in response to the highest similarity being less than the similarity threshold, performing speech recognition on the first to-be-recognized speech and responding to the user based on the first target recognized text obtained by the speech recognition.
4. The method of claim 2, wherein, The response to the user based on the historical recognized text corresponding to the historical recognized speech with the highest similarity comprises: determining the priority of the historical recognized text corresponding to each historical recognized speech with the highest similarity; wherein the priority of the historical recognized text is used to represent the possibility of the historical recognized text meeting the user's demand; responding to the user based on the historical recognized text with the highest priority.
5. The method of claim 4, wherein, The determination of the priority of the historical recognized text corresponding to each historical recognized speech with the highest similarity comprises: obtaining the time when each historical recognized text corresponding to the highest similarity was last selected; determining the priority of each historical recognized text corresponding to the highest similarity based on at least the time when each historical recognized text corresponding to the highest similarity was last selected.
6. The method of claim 5, wherein, The obtaining of the time when each historical recognized text corresponding to the highest similarity was last selected comprises: obtaining the first time when each historical recognized text corresponding to the highest similarity was last selected on a current device and the second time when each historical recognized text corresponding to the highest similarity was last selected on an associated device; The priority of each of the historical recognized texts corresponding to the highest similarity is determined based at least on a time when each of the historical recognized texts corresponding to the highest similarity was last selected. The priority of each of the historical recognized texts corresponding to the highest similarity is determined based at least on a time when each of the historical recognized texts corresponding to the highest similarity was last selected.
7. The method of claim 5, wherein, The priority of each of the historical recognized texts corresponding to the highest similarity is determined based at least on a time when each of the historical recognized texts corresponding to the highest similarity was last selected. The number of times each of the historical recognized texts corresponding to the highest similarity is selected within a preset time period is obtained. The priority of each of the historical recognized texts corresponding to the highest similarity is determined based on the number of times each of the historical recognized texts corresponding to the highest similarity is selected within the preset time period and a time when each of the historical recognized texts corresponding to the highest similarity was last selected.
8. The method of claim 3, wherein, After the first to-be-recognized speech is recognized based on the highest similarity being less than the similarity threshold, and the user is responded to based on the first target recognized text obtained by the recognition, the method further includes: obtaining second to-be-recognized speech; recognizing the second to-be-recognized speech to obtain a speech intent of the second to-be-recognized speech and a plurality of initial recognized texts corresponding to the second to-be-recognized speech; in response to the second to-be-recognized speech being consistent with the first to-be-recognized speech, filtering the plurality of initial recognized texts based on the speech intent to obtain second target recognized texts, and displaying the second target recognized texts on a display interface; in response to a user selecting any of the second target recognized texts, responding to the user based on the second target recognized text selected by the user.
9. The method of claim 8, wherein, The filtering of the plurality of initial recognized texts based on the speech intent to obtain the second target recognized texts includes: obtaining a type of execution content corresponding to each of the initial recognized texts; from the plurality of initial recognized texts, eliminating initial recognized texts whose type of execution content does not match the speech intent to obtain the second target recognized texts.
10. A voice interaction device, characterized by The apparatus includes: an obtaining module configured to obtain first to-be-recognized speech and a plurality of historical voice interaction data; wherein the historical voice interaction data includes historical recognized speech and historical recognized text, the historical recognized text being selected by a user from reference recognized text corresponding to the historical recognized speech, the reference recognized text being obtained based on voice recognition of the historical recognized speech; a determining module configured to determine a matching condition of the first to-be-recognized speech and each of the historical recognized speech; and The response module is configured to respond to the user based on historical recognized text corresponding to the historical recognized speech matching the first to-be-recognized speech; wherein the historical recognized text corresponding to the historical recognized speech matching the first to-be-recognized speech is the historical recognized text with the highest priority among historical recognized texts corresponding to at least two historical recognized speeches having the highest similarity with the first to-be-recognized speech, and the priority of each historical recognized text corresponding to the historical recognized speech with the highest similarity is determined based on at least the time when each historical recognized text corresponding to the historical recognized speech with the highest similarity was last selected, and the historical recognized texts corresponding to the at least two historical recognized speeches with the highest similarity are different from each other.
11. An electronic device, comprising: The electronic device includes a memory and a processor, the memory stores program instructions, and the processor is configured to execute the program instructions to implement the voice interaction method according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store program instructions, and the program instructions can be executed to implement the voice interaction method according to any one of claims 1-9.
Citation Information
Patent Citations
Voice interaction method and device, terminal equipment and storage medium
CN112242143A
Speech recognition method and device, equipment and storage medium
CN112802474A
Voice recognition method and device, storage medium and electronic equipment
CN113539272A