A voice reply method, device, equipment and storage medium
By selecting the target dialogue strategy from the preset dialogue strategy library in the voice interaction system, and generating reply texts of different levels of detail, the problem of single reply text generation method in the existing system is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202310049675.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-01
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-02-01
AI Technical Summary
The replies text generation method in the existing voice interaction system is single, resulting in users receiving fixed replies, losing freshness, and having low user experience satisfaction.
By receiving voice commands, extracting voice text, and selecting the target dialogue strategy from the preset dialogue strategy library based on the voice text, generating reply texts of different levels of detail, and finally generating corresponding voice message replies.
It enriches the way of generating reply text, increases the fun of human-computer interaction, and improves user experience satisfaction.
Smart Images

Figure CN116246625B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice interaction, and in particular, to a voice reply method, device, equipment and storage medium. Background Art
[0002] Voice interaction is playing an increasingly important role in the field of intelligent cockpits. In current voice interaction systems, usually only one conversation strategy is stored. When the voice interaction system receives a voice command issued by a user, a corresponding reply can be generated according to the voice command and the conversation strategy. For voice commands with the same meaning but different voice contents, since the meanings of the commands are the same and there is only one conversation strategy stored in the voice interaction system, the voice interaction system often only gives a fixed reply for these commands. As users become more familiar with the voice interaction system, the fixed voice interaction conversation will quickly make users lose their freshness, and the user experience satisfaction is relatively low. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a voice reply method, device, equipment and storage medium to solve the above technical problems.
[0004] On the one hand, a voice reply method is provided. The method includes:
[0005] Receiving a voice command;
[0006] Extracting voice text from the voice command;
[0007] Selecting a target conversation strategy from a preset conversation strategy library according to the voice text; at least two different conversation strategies are included in the conversation strategy library, and different conversation strategies can generate reply texts with different levels of detail according to the voice text;
[0008] Determining a corresponding reply text according to the voice text and the target conversation strategy;
[0009] Generating a corresponding voice message based on the reply text to reply to the voice command.
[0010] In one of the embodiments, the step of selecting a target conversation strategy from a preset conversation strategy library according to the voice text includes:
[0011] Counting the number of words in the voice text to obtain a first command word count;
[0012] Selecting a target conversation strategy from a preset conversation strategy library according to the first command word count.
[0013] In one embodiment, the dialogue strategy library contains the correspondence between the word count range intervals and the dialogue strategies. Selecting a target dialogue strategy from a preset dialogue strategy library according to the word count of the first instruction includes:
[0014] Determine the target word count range interval to which the word count of the first instruction belongs among the word count range intervals;
[0015] Determine the target dialogue strategy corresponding to the target word count range interval according to the correspondence between the word count range interval and the dialogue strategy.
[0016] In one embodiment, determining the target dialogue strategy corresponding to the target word count range interval according to the correspondence between the word count range interval and the dialogue strategy includes:
[0017] When the word count of the first instruction is within the first word count range interval, use the dialogue strategy corresponding to the first word count range interval for generating a detailed response text as the target dialogue strategy; the first word count range interval is the range interval where the word count of the first instruction is greater than or equal to a preset word count threshold;
[0018] When the word count of the first instruction is within the second word count range interval, use the dialogue strategy corresponding to the second word count range interval for generating a simple response text as the target dialogue strategy; the second word count range interval is the range interval where the word count of the first instruction is less than the preset word count threshold.
[0019] In one embodiment, selecting a target dialogue strategy from a preset dialogue strategy library according to the number of characters of the first text includes:
[0020] Search for a standard instruction that matches the current intention from a preset standard instruction library according to the voice text;
[0021] Determine the word count of the second instruction of the standard instruction;
[0022] Calculate the ratio of the word count of the first instruction to the word count of the second instruction;
[0023] Select a target dialogue strategy from a preset dialogue strategy library according to the ratio.
[0024] In one embodiment, the dialogue strategy library contains the correspondence between the ratio range intervals and the dialogue strategies. Selecting a target dialogue strategy from a preset dialogue strategy library according to the ratio includes:
[0025] Determine the target ratio range interval to which the ratio belongs among the ratio range intervals;
[0026] Determine the target dialogue strategy corresponding to the target ratio range interval according to the corresponding relationship between the ratio range interval and the dialogue strategy.
[0027] In one embodiment, the determining the target dialogue strategy corresponding to the target ratio range interval according to the corresponding relationship between the ratio range interval and the dialogue strategy includes:
[0028] When the ratio is in the first ratio range interval, use the dialogue strategy corresponding to the first ratio range interval for generating a detailed reply text as the target dialogue strategy; the first ratio range interval is the range interval where the ratio is greater than or equal to the preset ratio threshold;
[0029] When the ratio is in the second ratio range interval, use the dialogue strategy corresponding to the second ratio range interval for generating a simple reply text as the target dialogue strategy; the second ratio range interval is the range interval where the ratio is less than the preset ratio threshold.
[0030] On the other hand, a voice reply device is provided, including:
[0031] A receiving module, configured to receive a voice command;
[0032] An extraction module, configured to extract voice text from the voice command;
[0033] A selection module, configured to select a target dialogue strategy from a preset dialogue strategy library according to the voice text; at least two different dialogue strategies are included in the dialogue strategy library, and different dialogue strategies can generate reply texts with different levels of detail according to the voice text;
[0034] A determination module, configured to determine a corresponding reply text according to the voice text and the target dialogue strategy;
[0035] A reply module, configured to generate a corresponding voice message based on the reply text to reply to the voice command.
[0036] On the other hand, the present application also provides an electronic device, including a processor and a memory, where a computer program is stored in the memory, and the processor executes the computer program to implement the method described in any one of the above.
[0037] On the other hand, the present application also provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program is executed by at least one processor, the method described in any one of the above is implemented.
[0038] The voice reply method, device, equipment and storage medium provided by this application extract voice text from a voice command when receiving the voice command, select a target conversation strategy from a preset conversation strategy library according to the voice text, then determine a corresponding reply text according to the voice text and the target conversation strategy, and generate a corresponding voice message based on the reply text to reply to the voice command. Since the target conversation strategy can be flexibly selected from the preset conversation strategy library according to the voice text in the voice command to generate the reply text, the problem that the reply text generation method in the existing solution is single, resulting in the user terminal always receiving a fixed reply, is solved. The generation method of the reply text is enriched, adding fun to the human-computer interaction and improving the satisfaction of the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flow chart of the voice reply method provided in Embodiment 1 of this application;
[0040] Figure 2 It is a schematic flow chart of selecting a target conversation strategy provided in Embodiment 1 of this application;
[0041] Figure 3 It is a schematic structural diagram of the voice reply device provided in Embodiment 2 of this application;
[0042] Figure 4 It is a schematic structural diagram of the electronic device provided in Embodiment 3 of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the purpose, technical solution and advantages of this application clearer, the following further details this application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.
[0044] Embodiment 1:
[0045] The embodiment of this application provides a voice reply method, which can be applied to any terminal device with voice interaction function. For example, it can be applied to the terminal device in a car, such as in a driving computer, which is integrated with a voice interaction system.
[0046] Please refer to Figure 1 As shown, the voice reply method provided by the embodiment of this application may include the following steps:
[0047] S11: Receive a voice command.
[0048] S12: Extract voice text from the voice command.
[0049] S13: Select a target dialogue strategy from a preset dialogue strategy library according to the speech text; the dialogue strategy library contains at least two different dialogue strategies, and different dialogue strategies can generate response texts with different levels of detail according to the speech text.
[0050] S14: Determine the corresponding response text according to the speech text and the target dialogue strategy.
[0051] S15: Generate a corresponding voice message based on the response text to reply to the voice command.
[0052] Next, the specific processes of the above steps will be described in detail.
[0053] In step S11, the terminal device can collect environmental sounds and recognize the voice command issued by the user. For example, the terminal device can collect audio information in the cockpit and recognize the voice command issued by the user from the collected audio information.
[0054] It should be noted that the dialogue strategy library can include multiple different dialogue strategies. For example, it can include 2, 3, 4 or even more different dialogue strategies. For the same speech text, or speech texts with the same semantics, the response texts generated by different dialogue strategies have different levels of detail.
[0055] In step S13, the level of detail of the voice command can be calculated based on the speech text, and then the corresponding target dialogue strategy can be selected from the dialogue strategy library according to the level of detail, so as to generate a response text with the corresponding level of detail.
[0056] Exemplarily, the number of words in the speech text can be counted to obtain the first command word count, and the target dialogue strategy can be selected from the preset dialogue strategy library according to the first command word count.
[0057] In an alternative embodiment, the level of detail of the voice command can be directly represented by the first command word count. The larger the first command word count, the more detailed the voice command, and the smaller the first command word count, the simpler the voice command. At this time, the dialogue strategy library can include the corresponding relationship between the word count range interval and the dialogue strategy. Step S13 includes: determining the target word count range interval to which the first command word count belongs in the word count range interval, and determining the target dialogue strategy corresponding to the target word count range interval according to the corresponding relationship between the word count range interval and the dialogue strategy.
[0058] Specifically, when there are two dialogue strategies in the dialogue strategy library, the dialogue strategies can be classified according to the detail level of the generated response text into: a dialogue strategy for generating a detailed response text and a dialogue strategy for generating a simple response text. When the first instruction word count is within the first word count range interval, the dialogue strategy corresponding to the first word count range interval for generating a detailed response text is used as the target dialogue strategy; the first word count range interval is the range interval where the first instruction word count is greater than or equal to the preset word count threshold.
[0059] When the first instruction word count is within the second word count range interval, the dialogue strategy corresponding to the second word count range interval for generating a simple response text is used as the target dialogue strategy; the second word count range interval is the range interval where the first instruction word count is less than the preset word count threshold.
[0060] In another alternative implementation, please refer to Figure 2 As shown, step S13 includes the following sub-steps:
[0061] S131: Search for the standard instruction that matches the current intention from the preset standard instruction library according to the speech text.
[0062] S132: Determine the second instruction word count of the standard instruction.
[0063] S133: Calculate the ratio of the first instruction word count to the second instruction word count.
[0064] S134: Select the target dialogue strategy from the preset dialogue strategy library according to the ratio.
[0065] In sub-step S131, semantic analysis can be performed on the speech text, and then the standard instruction that matches the current intention is searched from the standard instruction library. The second instruction word count of the standard instruction also refers to the word count of the standard instruction. Exemplarily, the corresponding relationship between each standard instruction and the second instruction word count can be stored in advance. When the standard instruction that matches the current intention is found, the second instruction word count of the standard instruction can be obtained according to this corresponding relationship. Of course, it is also possible to count the word count of the standard instruction after finding the standard instruction that matches the current intention to obtain the second instruction word count.
[0066] In this implementation, the detail level of the voice instruction is characterized by the ratio of the first instruction word count to the second instruction word count. The larger the ratio, the more detailed the voice instruction; the smaller the ratio, the simpler the voice instruction. At this time, the dialogue strategy library can include the corresponding relationship between the ratio range interval and the dialogue strategy. Sub-step S134 includes:
[0067] Determine the target ratio range interval to which the ratio belongs in the ratio range interval;
[0068] Determine the target dialogue strategy corresponding to the target ratio range interval according to the corresponding relationship between the ratio range interval and the dialogue strategy.
[0069] Similarly, when the dialogue strategy library includes a dialogue strategy for generating a detailed response text and a dialogue strategy for generating a simple response text, if the ratio is in the first ratio range interval, the dialogue strategy for generating a detailed response text corresponding to the first ratio range interval is used as the target dialogue strategy; the first ratio range interval is the range interval where the ratio is greater than or equal to the preset ratio threshold.
[0070] If the ratio is in the second ratio range interval, the dialogue strategy for generating a simple response text corresponding to the second ratio range interval is used as the target dialogue strategy; the second ratio range interval is the range interval where the ratio is less than the preset ratio threshold.
[0071] It can be understood that the corresponding relationship between the ratio range interval and the dialogue strategy can be flexibly set by the developer.
[0072] For example, the set corresponding relationship can be as shown in Table 1 below:
[0073] Table 1
[0074] Ratio range interval (0,a) [a, b) [b, c] Dialogue strategy First dialogue strategy Second dialogue strategy Third dialogue strategy
[0075] Among them, a < b < c, the first dialogue strategy, the second dialogue strategy, and the third dialogue strategy are for the same voice text, and the detailed degree of the determined response text increases in turn. The greater the detailed degree of the response text, the more detailed the corresponding text content.
[0076] When there is a corresponding relationship as shown in Table 1, if the ratio of the first instruction word count to the second instruction word count falls within the range of (0, a), the first dialogue strategy is used as the target dialogue strategy; if the ratio falls within the range of [a, b), the second dialogue strategy is used as the target dialogue strategy; if the ratio falls within the range of [b, c], the third dialogue strategy is used as the target dialogue strategy.
[0077] It should be noted that in an alternative embodiment, a standard instruction library can be preset, that is, the standard instruction library mentioned above. For each standard instruction, multiple response texts with different detailed degrees can be pre-stored. At this time, in step S14, the following steps can be included:
[0078] Search for the standard instruction that meets the current intention from the preset standard instruction library according to the voice text;
[0079] Select a response text from the multiple response texts with different detailed degrees corresponding to the standard instruction pre-stored according to the target dialogue strategy.
[0080] At this time, the target dialogue strategy is used to indicate which reply text with what level of detail should be selected to reply to the voice command. In this example, for each standard command, multiple reply texts with different levels of detail are pre-stored, and each dialogue strategy corresponds to a reply text. Continuing with the above Table 1 as an example for illustration. When the corresponding relationship shown in Table 1 is set, the first dialogue strategy is used to indicate that the reply text with the first level of detail corresponding to the standard command is selected currently, the second dialogue strategy is used to indicate that the reply text with the second level of detail corresponding to the standard command is selected currently, and the third dialogue strategy is used to indicate that the reply text with the third level of detail corresponding to the standard command is selected currently, where the first level of detail < the second level of detail < the third level of detail.
[0081] It should be noted that in another alternative embodiment, the dialogue strategies in the dialogue strategy library can also be an algorithm that can generate reply texts with corresponding levels of detail according to the voice text. In this case, there is no need to store the corresponding reply texts for each standard command. Specifically, the standard command that conforms to the current intention can be searched from the preset standard command library according to the voice text, and the reply text with the corresponding level of detail can be generated according to the target dialogue strategy and the standard command.
[0082] It should be understood that although Figure 1-2 the steps in the flowchart of Figure 1-2 are shown in sequence according to the indication of the arrows, these steps do not necessarily execute in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0083] The voice reply method provided by the embodiments of the present application can identify the voice text, compare the voice text with the standard command, judge the level of detail of the user's speech, and thus give a corresponding reply according to this level of detail. It can achieve that the more the user says, the more the reply content is, and the less the user says, the less the reply content is, adding fun to the human-computer interaction and improving the user experience.
[0084] Embodiment 2:
[0085] Based on the same inventive concept, the embodiments of the present application provide a voice reply device. Please refer to Figure 3As shown, it should be understood that the specific functions of the voice response device can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.
[0086] The voice reply device includes at least one software functional unit that can be stored in a memory in the form of software or firmware or fixed in the operating system of the device. Specifically, the voice reply device includes:
[0087] The receiving module 301 is used to receive voice instructions.
[0088] The extraction module 302 is used to extract voice text from the voice instruction.
[0089] The selection module 303 is used to select a target dialogue strategy from a preset dialogue strategy library according to the speech text; the dialogue strategy library contains at least 2 different dialogue strategies, and different dialogue strategies can generate response texts with different levels of detail according to the speech text.
[0090] The determination module 304 is used to determine the corresponding reply text according to the speech text and the target dialogue strategy.
[0091] The reply module 305 is used to generate a corresponding voice message based on the reply text to reply to the voice instruction.
[0092] It should be noted that, for the sake of brevity, the contents described in the above embodiments will not be repeated in this embodiment.
[0093] Embodiment three:
[0094] This embodiment provides an electronic device. Figure 4 As shown, the electronic device includes a processor 401 and a memory 402, the memory 402 stores a computer program, the processor 401 and the memory 402 communicate via a communication bus, and the processor 401 executes the computer program to implement the steps of the voice reply method in the above embodiment 1, which will not be repeated here. It can be understood that Figure 4 The structure shown is for illustration only. The electronic device may also include Figure 4 More or fewer components as shown, or with Figure 4 It should be noted that the electronic device in the embodiment of the present application can be arranged in a car.
[0095] The processor 401 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 401 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0096] The memory 402 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0097] This embodiment also provides a computer-readable storage medium, such as a floppy disk, optical disk, hard disk, flash memory, USB flash drive, SD card, MMC card, etc. One or more programs for implementing the above-mentioned various steps are stored in the computer storage medium. These one or more programs can be executed by one or more processors 301 to implement the steps of the voice reply method in the first embodiment above, which will not be elaborated here.
[0098] It should be noted that the diagrams provided in this embodiment only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation may be arbitrarily changed, and the component layout type may also be more complex. The structures, proportions, sizes, etc. shown in the diagrams of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical substance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention. At the same time, the terms such as "upper", "lower", "left", "right", "middle", and "one" cited in this specification are only for the convenience of clear narration and are not used to limit the scope under which the present invention can be implemented. The change or adjustment of their relative relationships, without substantial change in the technical content, should also be regarded as the scope under which the present invention can be implemented.
[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0100] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A voice reply method, characterized in that, it includes: Receiving a voice command; Extracting voice text from the voice command; Selecting a target dialogue strategy from a preset dialogue strategy library according to the voice text; The dialogue strategy library contains at least two different dialogue strategies, and different dialogue strategies can generate reply texts with different levels of detail according to the voice text; Determining a corresponding reply text according to the voice text and the target dialogue strategy; Generating a corresponding voice message based on the reply text to reply to the voice command; The selecting a target dialogue strategy from a preset dialogue strategy library according to the voice text includes: Counting the number of characters in the voice text to obtain a first command character count; Selecting a target dialogue strategy from a preset dialogue strategy library according to the first command character count.
2. The voice reply method according to claim 1, characterized in that, The dialogue strategy library contains the corresponding relationship between the character count range interval and the dialogue strategy, and the selecting a target dialogue strategy from a preset dialogue strategy library according to the first command character count includes: Determining the target character count range interval to which the first command character count belongs in the character count range interval; Determining a target dialogue strategy corresponding to the target character count range interval according to the corresponding relationship between the character count range interval and the dialogue strategy.
3. The voice reply method according to claim 2, characterized in that, The determining a target dialogue strategy corresponding to the target character count range interval according to the corresponding relationship between the character count range interval and the dialogue strategy includes: When the first command character count is in the first character count range interval, using the dialogue strategy corresponding to the first character count range interval for generating a detailed reply text as the target dialogue strategy; the first character count range interval is the range interval where the first command character count is greater than or equal to a preset character count threshold; When the first command character count is in the second character count range interval, using the dialogue strategy corresponding to the second character count range interval for generating a simple reply text as the target dialogue strategy; the second character count range interval is the range interval where the first command character count is less than the preset character count threshold.
4. The voice reply method according to claim 1, characterized in that, The selecting a target dialogue strategy from a preset dialogue strategy library according to the first command character count includes: Searching for a standard command that conforms to the current intention from a preset standard command library according to the voice text; Determining the second command character count of the standard command; Calculating the ratio of the first command character count to the second command character count; Selecting a target dialogue strategy from a preset dialogue strategy library according to the ratio.
5. The voice reply method according to claim 4, characterized in that, The dialogue strategy library contains the corresponding relationship between the ratio range interval and the dialogue strategy, and the selecting a target dialogue strategy from a preset dialogue strategy library according to the ratio includes: Determining the target ratio range interval to which the ratio belongs in the ratio range interval; Determining a target dialogue strategy corresponding to the target ratio range interval according to the corresponding relationship between the ratio range interval and the dialogue strategy.
6. The voice reply method according to claim 5, wherein, the step of determining a target dialogue strategy corresponding to the target ratio range interval according to the corresponding relationship between the ratio range interval and the dialogue strategy includes: when the ratio is in the first ratio range interval, using the dialogue strategy corresponding to the first ratio range interval for generating a detailed reply text as the target dialogue strategy; the first ratio range interval is the range interval where the ratio is greater than or equal to a preset ratio threshold; when the ratio is in the second ratio range interval, using the dialogue strategy corresponding to the second ratio range interval for generating a simple reply text as the target dialogue strategy; the second ratio range interval is the range interval where the ratio is less than the preset ratio threshold.
7. A voice reply device, wherein, it includes: a receiving module, configured to receive a voice command; an extraction module, configured to extract voice text from the voice command; a selection module, configured to count the number of words in the voice text to obtain a first command word count, and select a target dialogue strategy from a preset dialogue strategy library according to the first command word count; at least two different dialogue strategies are included in the dialogue strategy library, and different dialogue strategies can generate reply texts with different levels of detail according to the voice text; a determination module, configured to determine a corresponding reply text according to the voice text and the target dialogue strategy; a reply module, configured to generate a corresponding voice message based on the reply text to reply to the voice command.
8. An electronic device, wherein, it includes a processor and a memory, a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by at least one processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Model training method and device based on dialogue template
CN110008319A
Reply information output method and device, equipment and storage medium
CN114443814A