Voice interaction method, device, electronic device, and computer-readable storage medium
By filtering the voice intention, removing mismatched initial recognition text, displaying the target recognition text and responding according to user selection, the problem of poor speech interaction accuracy is solved and higher speech interaction accuracy and convenience is achieved.
Patent Information
- Application Number
- CN202310268452.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-03-15
AI Technical Summary
Among the existing voice interaction technologies, the accuracy of voice interaction is poor, resulting in poor user experience.
By filtering the voice intention, the mismatched initial recognition text is eliminated, the target recognition text is displayed, and the response is carried out according to user selection, improving the accuracy and convenience of voice interaction.
It improves the accuracy and convenience of voice interaction, reduces user disturbance, and improves the user experience.
Smart Images

Figure CN116403577B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent voice technology, and in particular to a voice interaction method, device, electronic device, and computer-readable storage medium. Background Art
[0002] Currently, intelligent voice technology has become widely used. Almost all screen-based devices support voice interaction, allowing users to complete tasks quickly and efficiently. However, the accuracy of voice interaction is relatively poor. Summary of the Invention
[0003] The main technical problem solved by this application is to provide a voice interaction method, device, electronic device and computer-readable storage medium, which can improve the accuracy of voice interaction.
[0004] In order to solve the above technical problems, a technical solution adopted in this application is: to provide a voice interaction method, which includes: performing voice recognition on a first voice to be recognized, obtaining the voice intention of the first voice to be recognized and several first initial recognition texts corresponding to the first voice to be recognized; filtering the several first initial recognition texts based on the voice intention to obtain a first target recognition text, and displaying the first target recognition text on a display interface; in response to the user's selection of any first target recognition text, responding to the user based on the first target recognition text selected by the user.
[0005] Among them, filtering several first initial recognition texts based on the voice intention to obtain the first target recognition text includes: obtaining the type of execution content corresponding to each first initial recognition text; and eliminating the first initial recognition text whose execution content type does not match the voice intention from the several first initial recognition texts to obtain the first target recognition text.
[0006] In which, before recognizing the first voice to be recognized, a response is first made to the user based on the second target recognition text corresponding to the second voice to be recognized; before performing voice recognition on the first voice to be recognized and obtaining the voice intention of the first voice to be recognized and several first initial recognition texts corresponding to the first voice to be recognized, the voice interaction method also includes: detecting whether there is a first initial recognition text consistent with the second target recognition text; performing voice recognition on the first voice to be recognized and obtaining the voice intention of the first voice to be recognized and several first initial recognition texts corresponding to the first voice to be recognized, including: in response to the existence of the first initial recognition text consistent with the second target recognition text, performing voice recognition on the first voice to be recognized and obtaining the voice intention of the first voice to be recognized and several first initial recognition texts corresponding to the first voice to be recognized, and subsequent steps.
[0007] The voice interaction method further includes: in response to the absence of the first initial recognition text that is consistent with the second target recognition text, responding to the user based on the second target recognition text.
[0008] Among them, the recognition step of the second target recognition text includes: performing speech recognition on the second speech to be recognized to obtain several second initial recognition texts; determining the priority of each second initial recognition text, and taking the second initial recognition text with the highest priority as the second target recognition text; wherein, the priority of the second initial recognition text is used to represent the possibility that the second initial recognition text meets user needs.
[0009] Determining the priority of each second initial recognition text includes: obtaining the time when each second initial recognition text was last selected; and determining the priority of each second initial recognition text based at least on the time when each second initial recognition text was last selected.
[0010] Among them, the time of the most recent selection is negatively correlated with the priority.
[0011] Among them, the priority size of each second initial recognition text is determined at least based on the time when each second initial recognition text was last selected, including: obtaining the number of times each second initial recognition text is selected within a preset time period; for each second initial recognition text, based on the number of times the second initial recognition text is selected within the preset time period and the corresponding time when it was last selected, the priority size of the second initial recognition text is determined.
[0012] In order to solve the above technical problems, another technical solution adopted in this application is: providing a voice interaction device, which includes a recognition module, a filtering module and a response module; the recognition module is used to perform voice recognition on a first voice to be recognized, obtain the voice intention of the first voice to be recognized and several first initial recognition texts corresponding to the first voice to be recognized; the filtering module is used to filter the several first initial recognition texts based on the voice intention, obtain the first target recognition text and display the first target recognition text on the display interface; the response module is used to respond to the user's selection of any first target recognition text, and respond to the user based on the first target recognition text selected by the user.
[0013] In order to solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, which includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the above voice interaction method.
[0014] In order to solve the above technical problems, another technical solution adopted in this application is: providing a computer-readable storage medium, which is used to store program instructions, and the program instructions can be executed to implement the above-mentioned voice interaction method.
[0015] The above technical solution filters the first initial recognition texts based on the voice intent to obtain the first target recognition text, and displays the first target recognition text on the display interface. Therefore, the voice intent corresponding to the first speech to be recognized is used to filter the first initial recognition texts to eliminate the recognition texts that do not match the user's intent, thus streamlining the number of displayed first target recognition texts, making it easier for users to quickly select according to their needs and trigger the content they really need, thereby improving the accuracy and convenience of voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of an embodiment of the voice interaction method provided by this application;
[0017] Figure 2 This is a schematic diagram of an embodiment of the first voice interaction interface provided by this application;
[0018] Figure 3 This is a schematic diagram of an embodiment of a second voice interaction interface provided by this application;
[0019] Figure 4 This is a schematic diagram of an embodiment of a third voice interaction interface provided by this application;
[0020] Figure 5 This is a flow chart of an embodiment of the recognition step 1 of the second target recognition text provided by the present application;
[0021] Figure 6 yes Figure 5 The flowchart of step S52 is shown as an embodiment;
[0022] Figure 7 This is a structural diagram of an embodiment of a voice interaction device provided by the present application;
[0023] Figure 8 This is a structural diagram of an embodiment of an electronic device provided by the present application;
[0024] Figure 9 It is a structural diagram of an embodiment of a computer-readable storage medium provided by this application. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and effects of this application clearer and more specific, this application is further described in detail below with reference to the accompanying drawings and examples.
[0026] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0027] See also Figure 1 , Figure 1 It is a flow chart of an embodiment of the voice interaction method provided by this application. It should be noted that if there are substantially the same results, this embodiment does not Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment includes:
[0028] Step S11: performing speech recognition on a first speech to be recognized to obtain the speech intention of the first speech to be recognized and a plurality of first initial recognition texts corresponding to the first speech to be recognized.
[0029] The method of this embodiment is used to improve the accuracy and convenience of voice interaction between users and screen-based products, thereby enhancing the user's sense of security. The screen-based products described herein include, but are not limited to, in-car voice assistants, mobile phones, tablet computers, and home appliances. Screen-based products support voice interaction, which allows users to complete tasks quickly and efficiently.
[0030] In this embodiment, speech recognition is performed on the first speech to be recognized to obtain a plurality of first initial recognition texts corresponding to the first speech to be recognized; that is, data processing is performed on the first speech to be recognized, including recognition, speech comprehension, etc., to obtain a plurality of first initial recognition texts corresponding to the first speech to be recognized. For example, taking the first speech to be recognized as "I want to listen to the southern magpie" as an example; speech recognition is performed on the first speech to be recognized "I want to listen to the southern magpie", and three first initial recognition texts are obtained, namely "I want to listen to the southern magpie", "I want to listen to the southern magpie", and "I want to listen to the magpie".
[0031] In one embodiment, speech recognition can be performed on the first speech to be recognized using a dynamic time warping algorithm, a hidden Markov model (HMM) based on a parametric model, a vector quantization (VQ) based on a non-parametric model, or deep learning, to obtain several first initial recognition texts corresponding to the first speech to be recognized.
[0032] Because the content of the first initial recognition texts obtained by performing voice recognition on the first speech to be recognized may not match the speech intent corresponding to the first speech to be recognized, the accuracy of the first initial recognition texts is low, and the first initial recognition texts cannot be directly displayed for user selection. Therefore, in this embodiment, voice recognition is performed on the first speech to be recognized at the same time to obtain the speech intent corresponding to the first speech to be recognized, so as to facilitate subsequent filtering of the first initial recognition texts, thereby improving the accuracy of subsequent voice interaction and enhancing the user experience. For example, taking the first speech to be recognized as "I want to listen to the magpie", and the first initial recognition texts as "I want to listen to the magpie", "I want to listen to the magpie", and "I want to listen to the magpie" as an example; from the first speech to be recognized, it can be seen that the user's speech intent is to listen to a song, but the content "the magpie" in "I want to listen to the magpie" is not a song. Therefore, the first initial recognition text "I want to listen to the magpie" is displayed as the recognition text together with the other two first initial recognition texts, resulting in low accuracy of the recognition text obtained by the first speech to be recognized, thereby affecting the accuracy of the voice interaction.
[0033] In one embodiment, speech recognition can be performed on the first speech to be recognized based on rules, traditional machine learning algorithms (SVM) or deep learning algorithms (CNN, LSTM, RCNN, C-LSTM or FastText), etc., to obtain the speech intention corresponding to the first speech to be recognized.
[0034] It should be noted that the language of the first speech to be recognized is not limited and can be set according to actual needs; for example, English, Chinese, Japanese, Sichuan dialect and other local dialects.
[0035] Step S12: Filtering a plurality of first initial recognition texts based on the voice intent to obtain a first target recognition text, and displaying the first target recognition text on a display interface.
[0036] In this embodiment, a number of first initial recognition texts are filtered based on the voice intent to obtain a first target recognition text. That is to say, the first initial recognition texts obtained by recognition are filtered using the voice intent corresponding to the first voice to be recognized to filter out the first initial recognition texts that do not conform to the voice intent, thereby improving the accuracy of the obtained first target recognition text, that is, improving the matching of the first target recognition text with the user's needs. In addition, the first target recognition text will be displayed on the display interface, that is, the voice interaction interface will be presented so that the user can see the voice recognition results more intuitively, thereby facilitating the user to make quick selections based on subsequent needs (such as manual selection or voice interaction selection, etc.) to trigger the content that is really needed, thereby improving the accuracy and convenience of voice interaction, reducing disturbance to the user and improving the user experience.
[0037] For example, if Figure 2 As shown, Figure 2 This is a schematic diagram of an embodiment of the first voice interaction interface provided by the present application, taking the first voice to be recognized as "I want to listen to Nan Que", and the first initial recognition texts as "I want to listen to Nan Que", "I want to listen to Nan Que" and "I want to listen to Nan Que" as examples; it can be seen from the first voice to be recognized that the user's voice intention is to listen to songs, and the content of "I want to listen to Nan Que" "Nan Que" is not a song, so the first initial recognition text "I want to listen to Nan Que" is eliminated, and the first initial recognition text "I want to listen to Nan Que" and the first initial recognition text "I want to listen to Nan Que" are displayed on the display interface.
[0038] In one embodiment, the first initial recognition text is filtered out based on the match between the type of execution content corresponding to the first initial recognition text and the voice intent. Specifically, the type of execution content corresponding to each first initial recognition text is obtained; from a number of first initial recognition texts, the first initial recognition texts whose execution content type does not match the voice intent are eliminated to obtain the first target recognition text; for example, the voice intent is to listen to songs, and the type of execution content "Nan Que" of the first initial recognition text "I want to listen to Nan Que" is not a song, so the first initial recognition text "I want to listen to Nan Que" is eliminated. Of course, in other embodiments, the first initial recognition text can also be filtered out based on the voice intent in other ways, which are not specifically limited here.
[0039] During the voice interaction process, the system may also autonomously find and execute results that match the user's needs, and at this time, there is no need to provide the user with the first target recognition text displayed on the display interface, reducing the amount of calculation. Therefore, in one embodiment, before recognizing the first voice to be recognized, the user is responded to based on the second target recognition text corresponding to the second voice to be recognized; at this time, before performing voice recognition on the first voice to be recognized and obtaining the voice intent of the first voice to be recognized and several first initial recognition texts corresponding to the first voice to be recognized, it is also detected whether there is a first initial recognition text consistent with the second target recognition text; then, in response to the existence of the first initial recognition text consistent with the second target recognition text, voice recognition is performed on the first voice to be recognized to obtain the voice intent of the first voice to be recognized and several first initial recognition texts corresponding to the first voice to be recognized, as well as subsequent steps. That is to say, before recognizing the first voice to be recognized, the system will autonomously match, recognize and execute the second voice to be recognized, and after obtaining the first voice to be recognized, it will recognize the first voice to be recognized again to obtain several first initial recognition texts; then, if it is detected that there is a first initial recognition text that is consistent with the second target recognition text, it is considered that the first voice to be recognized is consistent with the second voice to be recognized. At this time, it may be that the currently executed content does not match the user's needs, so the voice intent of the first voice to be recognized will be recognized, and several first initial recognition texts will be filtered out based on the voice intent, and the first target recognition text obtained after filtering will be displayed on the display interface for the user to select, so as to trigger the content that is really needed and improve the accuracy and efficiency of voice interaction. In other words, the display of the recognition text will be triggered only when the response result does not meet the user's expectations, so that it can be quickly modified and corrected when the response result does not meet the user's expectations.
[0040] For example, if Figure 2 and Figure 3 As shown, Figure 3 This is a schematic diagram of an embodiment of the second voice interaction interface provided by this application; Figure 2 As shown, the user inputs the voice "I want to listen to Nan Que", and the system receives the second voice to be recognized "I want to listen to Nan Que"; the system autonomously matches and recognizes the second voice to be recognized "I want to listen to Nan Que" to obtain the second target recognition text "I want to listen to Nan Que" and executes "Play Song Nan Que"; Figure 3As shown, when the user finds that the song being played is "Nan Que" instead of "Nan Que", the user re-enters the voice "I want to listen to Nan Que". At this time, the system receives the first voice to be recognized "I want to listen to Nan Que" and recognizes it, and obtains three first initial recognition texts - "I want to listen to Nan Que", "I want to listen to Nan Que" and "I want to listen to Nan Que"; since the second target recognition text "I want to listen to Nan Que" is consistent with the first initial recognition text "I want to listen to Nan Que", it indicates that the first voice to be recognized is consistent with the second voice to be recognized, and the voice intention of the first voice to be recognized is determined at this time; then, the three first initial recognition texts are filtered according to the voice intention of the first voice to be recognized, and the first target recognition text "I want to listen to Nan Que" and the first target recognition text "I want to listen to Nan Que" are obtained, and are displayed on the display interface for users to choose according to their needs, so that users can trigger the content they really need later, thereby improving the accuracy and convenience of voice interaction.
[0041] In other embodiments, in response to the absence of the first initial recognition text that is consistent with the second target recognition text, the user is responded to based on the second target recognition text. That is, if it is detected that there is no first initial recognition text that is consistent with the second target recognition text, it is considered that the first speech to be recognized is inconsistent with the second speech to be recognized, and the current execution content will be continued. For example, if Figure 2 As shown, the user inputs the voice "I want to listen to lover", and the system receives the second voice to be recognized "I want to listen to lover"; the system autonomously matches and recognizes the second target recognition text "I want to listen to lover" according to the second voice to be recognized "I want to listen to lover" and executes "play song lover"; Figure 3 As shown, when the user finds that the song being played is "Lover", he re-enters the voice "I want to listen to Nan Que". At this time, the system receives the first voice to be recognized "I want to listen to Nan Que" and recognizes it, and obtains 3 first initial recognition texts - "I want to listen to Nan Que", "I want to listen to Nan Que" and "I want to listen to Nan Que"; since the second target recognition text "I want to listen to Lover" is inconsistent with each first initial recognition text, the second voice to be recognized is inconsistent with the first voice to be recognized, and the song "Lover" continues to play.
[0042] Step S13: In response to the user's selection of any first target recognition text, respond to the user based on the first target recognition text selected by the user.
[0043] In this embodiment, in response to the user's selection of any first target recognition text, the user is responded to based on the first target recognition text selected by the user. In other words, the user can quickly select (manually or through voice interaction) according to their needs on the display interface to quickly trigger the content they really need, thereby improving the accuracy and convenience of voice interaction.
[0044] For example, if Figure 4 As shown, Figure 4 It is a schematic diagram of an embodiment of the third voice interaction interface provided in this application; taking the example of a user selecting any first target recognition text through voice interaction, the user voice inputs "first", and the first option is "I want to listen to Nan Que", and the song Nan Que is played at this time.
[0045] In the above embodiment, the first target recognition text is obtained by filtering the first initial recognition texts based on the voice intent, and the first target recognition texts are displayed on the display interface. Therefore, the first target recognition texts are filtered using the voice intent corresponding to the first speech to be recognized to eliminate the recognition texts that do not match the user's intent, thus streamlining the number of displayed first target recognition texts, making it easier for users to quickly select according to their needs, triggering the content they really need, and improving the accuracy and convenience of voice interaction.
[0046] See also Figure 5 , Figure 5 It is a flow chart of an embodiment of the recognition step of the second target recognition text provided by this application. It should be noted that if there is substantially the same result, this embodiment does not use Figure 5 The process sequence shown is limited. Figure 5 As shown, this embodiment includes:
[0047] Step S51: performing speech recognition on a second speech to be recognized to obtain a plurality of second initial recognition texts.
[0048] In this embodiment, speech recognition is performed on the second speech to be recognized to obtain a plurality of second initial recognition texts. In other words, data processing is performed on the second speech to be recognized, including recognition and speech comprehension, to obtain a plurality of second initial recognition texts corresponding to the second speech to be recognized. For example, taking the second speech to be recognized as "I want to listen to the southern magpie" as an example; speech recognition is performed on the second speech to be recognized "I want to listen to the southern magpie", and three second initial recognition texts are obtained, namely "I want to listen to the southern magpie", "I want to listen to the southern magpie", and "I want to listen to the magpie".
[0049] In one embodiment, speech recognition can be performed on the second speech to be recognized using a dynamic time warping algorithm, a hidden Markov model (HMM) based on a parametric model, a vector quantization (VQ) based on a non-parametric model, or deep learning, to obtain several second initial recognition texts corresponding to the second speech to be recognized.
[0050] Step S52: Determine the priority of each second initial recognition text, and use the second initial recognition text with the highest priority as the second target recognition text.
[0051] In this embodiment, the priority of each second initial recognition text is determined, and the second initial recognition text with the highest priority is used as the second target recognition text; wherein the priority of the second initial recognition text is used to represent the likelihood that the second initial recognition text meets the user's needs. Since the priority of the second initial recognition text indicates the likelihood that it meets the user's needs, by determining the priority of each second initial recognition text and using the second initial recognition text with the highest priority as the second target recognition text, the second target recognition text is a result that better meets the user's needs. Therefore, subsequent responses to the user based on the second target recognition text can better match the user's needs, that is, more accurately meet the user's needs, thereby improving the accuracy and convenience of voice interaction.
[0052] In one embodiment, if Figure 6 As shown, Figure 6 yes Figure 5 The flowchart of step S52 of an embodiment is shown, which determines the priority of each second initial recognition text, specifically including the following sub-steps:
[0053] Step S61: Obtain the time when each second initial recognition text was last selected.
[0054] In this embodiment, the time when each second initial recognized text was most recently selected is obtained. In one specific embodiment, the time when each second initial recognized text was most recently selected can be obtained from the user's personalized vocabulary. Specifically, the system stores the recognized text data previously selected by the user to form the user's personalized vocabulary, and the time when the user selected the recognized text is simultaneously recorded when the recognized text is stored.
[0055] Step S62: Determine the priority of each second initially recognized text based at least on the time when each second initially recognized text was last selected.
[0056] In this embodiment, the priority of each second initially recognized text is determined based at least on the time when each second initially recognized text was last selected.
[0057] In one embodiment, the priority of each second initial recognition text is determined based on the time when each second initial recognition text was last selected. In a specific embodiment, the time of last selection is negatively correlated with the priority, that is, the earlier the second initial recognition text was last selected, the smaller its corresponding priority, and the less likely it is to meet user needs; the later the second initial recognition text was last selected, the more likely the user has recently selected the second initial recognition text, so the possibility of continuing to select the second initial recognition text in the future is relatively high, so the priority of this second initial recognition text is higher. Of course, in other specific embodiments, the time of last selection can also be set to be positively correlated with the priority, which is not specifically limited here.
[0058] In other embodiments, in order to improve the accuracy of the priority of each second initial recognition text, the priority can also be determined based on the number of times it is selected within a preset time period and the corresponding time of the most recent selection. Specifically, the number of times each second initial recognition text is selected within a preset time period is first obtained; then, for each second initial recognition text, the priority of the second initial recognition text is determined based on the number of times the second initial recognition text is selected within the preset time period and the corresponding time of the most recent selection. In other words, for each second initial recognition text, the priority of the second initial recognition text is determined based on the two dimensions of the number of times it is selected and the time of the most recent selection.
[0059] See also Figure 7 , Figure 7 It is a structural diagram of an embodiment of the voice interaction device provided in the present application. The voice interaction device 70 includes a recognition module 71, a filtering module 72 and a response module 73. The recognition module 71 is used to perform voice recognition on the first voice to be recognized, and obtain the voice intention of the first voice to be recognized and a plurality of first initial recognition texts corresponding to the first voice to be recognized; the filtering module 72 is used to filter the plurality of first initial recognition texts based on the voice intention, obtain the first target recognition text and display the first target recognition text on the display interface; the response module 73 is used to respond to the user's selection of any first target recognition text and respond to the user based on the first target recognition text selected by the user.
[0060] Among them, the filtering module 72 is used to filter several first initial recognition texts based on the voice intention to obtain the first target recognition text, specifically including: obtaining the type of execution content corresponding to each first initial recognition text; from several first initial recognition texts, eliminating the first initial recognition text whose execution content type does not match the voice intention to obtain the first target recognition text.
[0061] Among them, before recognizing the first voice to be recognized, the above-mentioned response to the user is first made based on the second target recognition text corresponding to the second voice to be recognized; the voice interaction device 70 also includes a detection module 74, and the detection module 74 is used to filter a number of first initial recognition texts based on the voice intention, obtain the first target recognition text and display the first target recognition text on the display interface, specifically including: detecting whether there is a first initial recognition text consistent with the second target recognition text; the recognition module 71 is used to perform voice recognition on the first voice to be recognized, obtain the voice intention of the first voice to be recognized and a number of first initial recognition texts corresponding to the first voice to be recognized, specifically including: in response to the existence of the first initial recognition text consistent with the second target recognition text, perform voice recognition on the first voice to be recognized, obtain the voice intention of the first voice to be recognized and a number of first initial recognition texts corresponding to the first voice to be recognized and subsequent steps.
[0062] The response module 73 is further configured to respond to the user based on the second target recognition text in response to the absence of the first initial recognition text that is consistent with the second target recognition text.
[0063] Among them, the recognition module 71 is also used to identify the second target recognition text. The recognition steps of the above-mentioned second target recognition text include: performing speech recognition on the second speech to be recognized to obtain several second initial recognition texts; determining the priority of each second initial recognition text, and taking the second initial recognition text with the highest priority as the second target recognition text; wherein, the priority of the second initial recognition text is used to represent the possibility that the second initial recognition text meets user needs.
[0064] Among them, the recognition module 71 is used to determine the priority size of each second initial recognition text, specifically including: obtaining the time when each second initial recognition text was last selected; and determining the priority size of each second initial recognition text based at least on the time when each second initial recognition text was last selected.
[0065] Among them, the time of the most recent selection is negatively correlated with the priority.
[0066] Among them, the recognition module 71 is used to determine the priority size of each second initial recognition text based at least on the time when each second initial recognition text was last selected, specifically including: obtaining the number of times each second initial recognition text is selected within a preset time period; for each second initial recognition text, determining the priority size of the second initial recognition text based on the number of times the second initial recognition text is selected within the preset time period and the corresponding time of last selection.
[0067] See also Figure 8 , Figure 8is a schematic diagram of the structure of an embodiment of an electronic device provided in this application. The electronic device 80 includes a memory 81 and a processor 82 coupled to each other. The processor 82 is configured to execute program instructions stored in the memory 81 to implement the steps of any of the above-described voice interaction method embodiments. In a specific implementation scenario, the electronic device 80 may include, but is not limited to, a microcomputer and a server. Furthermore, the electronic device 80 may also include mobile devices such as laptops and tablet computers, which are not limited here.
[0068] Specifically, the processor 82 is used to control itself and the memory 81 to implement the steps of any of the above-mentioned voice interaction method embodiments. The processor 82 can also be called a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with signal processing capabilities. The processor 82 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 82 can be implemented by an integrated circuit chip.
[0069] See also Figure 9 , Figure 9 : It is a structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 90 of the embodiment of the present application stores program instructions 91, and when the program instructions 91 are executed, the method provided by any embodiment of the voice interaction method of the present application and any non-conflicting combination is implemented. Among them, the program instructions 91 can form a program file and be stored in the above-mentioned computer-readable storage medium 90 in the form of a software product, so that a computer device (which can be a personal computer, server, or network device, etc.) executes all or part of the steps of the methods of each embodiment of the present application. The aforementioned computer-readable storage medium 90 includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0070] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0071] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A voice interaction method, characterized in that: The method comprises: responding to the user based on the second target recognition text corresponding to the second speech to be recognized; Performing speech recognition on a first speech to be recognized to obtain a plurality of first initial recognition texts corresponding to the first speech to be recognized, and detecting whether there is a first initial recognition text that is consistent with the second target recognition text; In response to the presence of the first initial recognition text that is consistent with the second target recognition text, performing speech intent recognition on the first speech to be recognized; Filtering the plurality of first initial recognition texts based on the speech intent of the first speech to be recognized to obtain a first target recognition text, and displaying the first target recognition text on a display interface; In response to the user's selection of any one of the first target recognition texts, a response is given to the user based on the first target recognition text selected by the user.
2. The method according to claim 1, characterized in that The filtering of the plurality of first initial recognition texts based on the speech intent to obtain a first target recognition text includes: Obtaining the type of execution content corresponding to each of the first initially recognized texts; From the plurality of first initial recognition texts, first initial recognition texts whose execution content types do not match the speech intention are eliminated to obtain the first target recognition text.
3. The method according to claim 1, characterized in that The method further comprises: In response to the absence of the first initial recognition text consistent with the second target recognition text, a response is given to the user based on the second target recognition text.
4. The method according to claim 1, wherein The step of identifying the second target recognition text includes: performing speech recognition on the second speech to be recognized to obtain a plurality of second initial recognition texts; Determine the priority of each second initial recognition text, and use the second initial recognition text with the highest priority as the second target recognition text; wherein the priority of the second initial recognition text is used to represent the possibility that the second initial recognition text meets user needs.
5. The method according to claim 4, characterized in that The determining the priority of each of the second initially recognized texts includes: Obtaining the time when each of the second initially recognized texts was last selected; The priority of each second initially recognized text is determined based at least on the time when each second initially recognized text was last selected.
6. The method according to claim 5, characterized in that The time of the most recent selection is negatively correlated with the priority level.
7. The method according to claim 5 or 6, characterized in that The determining the priority of each second initially recognized text based at least on the time when each second initially recognized text was last selected includes: Obtaining the number of times each of the second initially recognized texts is selected within a preset time period; For each second initially recognized text, the priority level of the second initially recognized text is determined based on the number of times the second initially recognized text is selected within a preset time period and the corresponding time of the most recent selection.
8. A voice interaction device, characterized in that: The device comprises: a recognition module configured to respond to a user based on a second target recognition text corresponding to a second speech to be recognized; perform speech recognition on a first speech to be recognized to obtain a plurality of first initial recognition texts corresponding to the first speech to be recognized, and detect whether there is a first initial recognition text that is consistent with the second target recognition text; and perform speech intent recognition on the first speech to be recognized in response to the presence of the first initial recognition text that is consistent with the second target recognition text; a filtering module, configured to filter the plurality of first initial recognition texts based on the speech intent of the first speech to be recognized, obtain a first target recognition text, and display the first target recognition text on a display interface; The response module is used to respond to the user's selection of any of the first target recognition texts and respond to the user based on the first target recognition text selected by the user.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the voice interaction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the voice interaction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice interaction method and device, terminal equipment and storage medium
CN112242143A
Voice recognition method and device, storage medium and electronic equipment
CN113539272A