Voice interaction method, device, electronic device and storage medium

The voice interaction method of frequency comparison and volume comparison solves the problem of air conditioners having difficulty capturing user commands in noisy environments, improves accuracy and efficiency, reduces misoperations and saves energy.

CN113870851BActive Publication Date: 2025-09-16QINGDAO HAIER AIR CONDITIONER GENERAL CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111087391.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-16
Publication Date
2025-09-16
Estimated Expiration
2041-09-16

AI Technical Summary

Technical Problem

Existing air conditioners have difficulty accurately capturing user commands in noisy environments, especially when multiple people are speaking at the same time, and their voice interaction capabilities are insufficient.

Method used

By receiving and parsing the user's voice input information, using frequency comparison and volume comparison to ensure the accuracy of voice interaction, including wake-up word triggering, frequency range analysis, confirmation reminders and command confirmation processes, combined with time thresholds and termination command words to improve the accuracy of command acquisition.

Benefits of technology

It improves the accuracy and efficiency of air-conditioning voice interaction in noisy environments, reduces misoperation, saves energy and reduces emissions, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113870851B_ABST
    Figure CN113870851B_ABST
Patent Text Reader

Abstract

The present invention provides a voice interaction method, device, electronic device and storage medium, which relate to the field of air conditioning technology. The voice interaction method includes: receiving a first voice input information from a user, determining whether the first voice input information includes a preset wake-up word; determining whether the first voice input information includes a preset wake-up word, parsing the sound frequency of the first voice input information, and outputting a target sound frequency range; receiving a second voice input information from a user, parsing the sound frequency of the second voice input information; determining whether the sound frequency of the second voice input information is within the target sound frequency range, and parsing the second voice input information to obtain a first execution request information. The voice interaction method provided by the present invention effectively reduces the interference of other noises by comparing sound frequencies, ensures that the voice input information processed during voice interaction comes from the same user, and improves the accuracy of command acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air conditioning, and in particular to a voice interaction method, device, electronic device and storage medium. Background Art

[0002] Existing air conditioners all have voice interaction technology, allowing users to communicate with the air conditioner and provide operating instructions. In quiet environments, the air conditioner can generally quickly and accurately capture user instructions. However, in noisy environments, especially when multiple people are talking simultaneously, the air conditioner has difficulty capturing user instructions. Therefore, improving the air conditioner's ability to receive instructions in noisy environments is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0003] The present invention provides a voice interaction method, device, electronic device and storage medium, which are used to solve the problem that traditional air conditioners are difficult to obtain user instructions in a noisy environment.

[0004] To address the problems in the prior art, an embodiment of the present invention provides a voice interaction method, including:

[0005] Receive first voice input information from a user, and determine whether the first voice input information includes a preset wake-up word;

[0006] Determining that the first voice input information includes a preset wake-up word, analyzing the sound frequency of the first voice input information, and outputting a target sound frequency range;

[0007] receiving a second voice input message from a user, and analyzing a sound frequency of the second voice input message;

[0008] Determine that the sound frequency of the second voice input information is within the target sound frequency range, and parse the second voice input information to obtain first execution request information.

[0009] According to a voice interaction method provided by the present invention, after parsing the second voice input information to obtain the first execution request information, the method further includes:

[0010] Outputting a confirmation reminder of the first execution request information;

[0011] Receive user confirmation of the first execution request information, and execute the target instruction of the first execution request information.

[0012] A voice interaction method provided by the present invention further includes:

[0013] receiving a third voice input message from a user, and analyzing a sound frequency of the third voice input message;

[0014] Determining that the sound frequency of the third voice input information is within the target sound frequency range, and analyzing the volume of the second voice input information and the third voice input information;

[0015] It is determined that the volume of the third voice input information is greater than the volume of the second voice input information, and the third voice input information is parsed to obtain second execution request information.

[0016] According to a voice interaction method provided by the present invention, after parsing the third voice input information to obtain the second execution request information, the method further includes:

[0017] Outputting a confirmation reminder of the second execution request information;

[0018] Receive user confirmation of the second execution request information, and execute the target instruction of the second execution request information.

[0019] According to a voice interaction method provided by the present invention, the time difference between receiving the second voice input information and receiving the third voice input information does not exceed 1s.

[0020] According to a voice interaction method provided by the present invention, after analyzing the sound frequency of the first voice input information and outputting the target sound frequency range, the method further includes:

[0021] If it is determined that the input duration of the user's voice input information is greater than a preset input time threshold, the user's voice input information is stopped.

[0022] A voice interaction method provided by the present invention further includes:

[0023] receiving fourth voice input information from the user;

[0024] It is determined that the fourth voice input information includes a preset termination command word, and the receiving of the user's voice input information is stopped.

[0025] The present invention also provides a control device, comprising:

[0026] A voice receiving module, configured to receive a user's first voice input information and a second voice input information;

[0027] a frequency determination module, configured to determine the sound frequencies of the first voice input information and the second voice input information;

[0028] A first processing module, configured to determine whether the first voice input information contains a preset wake-up word;

[0029] A second processing module is configured to determine the volume of the first voice input information and the second voice input information; and

[0030] The third processing module is used to output execution request information according to the processing result of the second processing module.

[0031] The present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the voice interaction method as described above when executing the program.

[0032] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the voice interaction method as described in any one of the above items are implemented.

[0033] Air conditioners operate in a relatively noisy environment, especially when multiple people are talking at once. This makes it difficult for the air conditioner to capture user commands. In the voice interaction method provided by the present invention, a first voice input is used to trigger the voice interaction, and a second voice input is used to convey the command. By comparing the frequencies of the first and second voice inputs, the execution request information of the second voice input is parsed only if the frequency of the second voice input is similar to that of the first. This ensures that the voice interaction process is performed by the same user, improving the accuracy of command acquisition. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 This is one of the flow charts of the voice interaction method provided by the present invention;

[0036] Figure 2 It is a structural schematic diagram of the electronic device provided by the present invention.

[0037] Reference numerals:

[0038] 1: Processor; 2: Communication interface; 3: Memory;

[0039] 4: Communication bus. DETAILED DESCRIPTION

[0040] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0041] The following combination Figure 1-Figure 2 Describe the voice interaction method, device, electronic device and storage medium of the present invention. Figure 1 The present invention provides a voice interaction method, comprising the following steps:

[0042] S10: Receive first voice input information from a user, and determine whether the first voice input information includes a preset wake-up word;

[0043] S20: Determine whether the first voice input information includes a preset wake-up word, analyze the sound frequency of the first voice input information, and output a target sound frequency range;

[0044] S30, receiving second voice input information from the user, and analyzing the sound frequency of the second voice input information;

[0045] S40: Determine whether the sound frequency of the second voice input information is within the target sound frequency range, and parse the second voice input information to obtain first execution request information.

[0046] Existing air conditioners all have voice interaction technology, allowing users to communicate with the air conditioner and convey operating instructions. In quiet environments, the air conditioner can generally quickly and accurately grasp user instructions. However, in noisy environments, especially when multiple people are talking at the same time, the air conditioner has difficulty grasping user instructions.

[0047] To this end, the present invention provides a voice interaction method, in which the first voice input information is used to trigger the voice interaction. When the first voice input information contains a preset wake-up word, the voice interaction can be triggered. Generally speaking, the wake-up word can be set to a variety of words, such as "Xiaoyou", "Xiaoai", etc., which can be set before the air conditioner leaves the factory, or can be set by the user according to his or her own preferences after receiving the air conditioner. The second voice input information is used to convey instructions. It should be noted that after the voice interaction is triggered, if the voice interaction environment is too noisy and there are multiple human voices, it is difficult to capture the specific second voice input information. Therefore, in the technical solution provided by the present invention, when it is determined that the first voice input information contains the preset wake-up word, the sound frequency of the first voice input information will be parsed to obtain the target sound frequency range; after receiving the second voice input information, the sound frequency of the second voice input information will be parsed to see if it is within the above-mentioned target sound frequency. If so, the second voice input information will be parsed to obtain the first execution request information.

[0048] It should be noted that the frequency of each sound is different, and the frequency of each person's voice will be different. For example, the frequency of a man's voice is lower, the frequency of a woman's voice is higher, the frequency of an elderly person's voice is lower, and the frequency of a child's voice is higher. Therefore, by comparing the sound frequencies of the first voice input information and the second voice input information, if the sound frequencies of the two are similar, it can be considered that the first voice input information and the second voice input information are input by the same person, and then the instruction information carried by the second voice input information can be parsed. Other voice input information with a different frequency from the first voice input information will be ignored and no instruction acquisition processing will be performed. This can reduce the impact of other noises, ensure that the voice interaction process is carried out by the same person, ensure the accuracy of instruction acquisition, and improve the efficiency of voice interaction.

[0049] Furthermore, after S40, parsing the second voice input information to obtain the first execution request information, the method further includes:

[0050] S41: Output a confirmation reminder for the first execution request information;

[0051] S42: Receive user confirmation of the first execution request information, and execute the target instruction of the first execution request information.

[0052] Under normal circumstances, when the first execution request information carried by the second voice input information is obtained, confirmation will be made with the user. The user only needs to confirm yes or no to control the execution of the instruction related to the execution request information. For example, if the second voice input information is obtained carrying the instruction to "lower the air conditioner temperature to 25°C", the user will be asked "Do you want to lower the air conditioner temperature to 25°C?" If the user confirms "yes", the relevant operation will be performed. If the user confirms "no", other voice input information waiting for the user will be asked. Adding an inquiry operation between obtaining the execution request information and executing the instruction can ensure the accuracy of the instruction execution and prevent the occurrence of erroneous execution of the instruction.

[0053] Furthermore, the voice interaction method provided by the present invention further includes:

[0054] S50: Receive third voice input information from the user, and analyze the sound frequency of the third voice input information;

[0055] S60: Determine that the sound frequency of the third voice input information is within the target sound frequency range, and analyze the volume of the second voice input information and the third voice input information;

[0056] S70: Determine that the volume of the third voice input information is greater than the volume of the second voice input information, and parse the third voice input information to obtain second execution request information.

[0057] It should be noted that during voice interaction, there may be multiple voice input messages with similar sound frequencies. In this case, it is necessary to compare these voice input messages to obtain the correct voice input information. Therefore, in the technical solution provided by the present invention, if the sound frequency of the third voice input message is similar to that of the first voice input message and is also within the target sound frequency, it is necessary to compare the volume of the second voice input message and the third voice input message. If the volume of the third voice input message is greater than that of the second voice input message, it is necessary to obtain the second execution request information carried by the third voice input message, because compared with the second voice input message, the third voice input message is more likely to carry a valid user instruction.

[0058] Furthermore, after S70, parsing the third voice input information to obtain the second execution request information, the method further includes:

[0059] S71: Output a confirmation reminder of the second execution request information;

[0060] S72: Receive user confirmation of the second execution request information, and execute the target instruction of the second execution request information.

[0061] Similar to the previous process, after confirming that the volume of the third voice input information is greater than that of the second voice input information, the user will be asked for instructions related to the second execution request information carried by the third voice input information. It is understandable that the third voice input information has a higher volume and is closer to the air conditioner, and is more likely to be issued by the same user. However, in actual use, the possibility of two users of the same gender and speaking voice inputting voice input information at the same time is not very high, but the usage logic provided by the present invention for this situation can also further improve the recognition accuracy of voice interaction commands.

[0062] It should be noted that the time difference between receiving the second voice input information and receiving the third voice input information does not exceed 1s. It is understandable that within the effective time of 1s, the second voice input information and the third voice input information are received at the same time, and the input time difference between the two voice input information is very small. It is impossible to determine which one is the valid voice input information, then it is necessary to compare the two voice input information to confirm which one carries the execution request information. In other words, if the time to receive the third voice input information exceeds 1s, the third voice input information is most likely invalid voice input information, and there is no need to analyze the third voice input information. Of course, you can also choose time intervals of 2s, 3s, etc. to improve the accuracy of voice recognition, but it is best not to exceed 2s.

[0063] Furthermore, in conventional technologies, the voice interaction function will remain in the on state for a long time after being awakened, which is likely to cause erroneous operations and of course increase power consumption to a certain extent. Therefore, in the technical solution provided by the present invention, after S20, analyzing the sound frequency of the first voice input information and outputting the target sound frequency range, it also includes:

[0064] S21: Determine that the input duration of the user's voice input information is greater than a preset input time threshold, and stop receiving the user's voice input information.

[0065] It should be noted that the preset input time threshold can be 5-6 seconds. When the voice interaction function is awakened, if the user does not input subsequent voice input information in time, the voice wake-up function will be disabled and no further voice input information from the user will be received. If the user wants to continue inputting voice, the wake-up operation must be repeated. This can improve the accuracy of the voice interaction function and also play a role in energy conservation and emission reduction to a certain extent.

[0066] Involving the awakening of the voice interaction function will certainly involve shutting down the voice interaction function. In the technical solution provided by the present invention, the voice interaction method provided by the present invention also includes:

[0067] S80: Receive fourth voice input information from the user;

[0068] S90: Determine that the fourth voice input information includes a preset termination command word, and stop receiving the user's voice input information.

[0069] For example, when the user enters voice messages such as "turn off voice interaction" or "I don't need to perform any operations", the voice interaction function will be turned off and the user's subsequent voice input information will stop being received. Of course, a closing inquiry operation can be added before closing, which makes the voice interaction more intelligent and improves the user experience.

[0070] The control device provided by the present invention is described below. The control device described below and the voice interaction method described above can be referenced to each other.

[0071] The present invention provides a control device, comprising:

[0072] The present invention also provides a control device, comprising:

[0073] A voice receiving module, configured to receive a user's first voice input information and a second voice input information;

[0074] a frequency determination module, configured to determine the sound frequencies of the first voice input information and the second voice input information;

[0075] A first processing module, configured to determine whether the first voice input information contains a preset wake-up word;

[0076] A second processing module is configured to determine the volume of the first voice input information and the second voice input information; and

[0077] The third processing module is used to output execution request information according to the processing result of the second processing module.

[0078] The present invention also provides an electronic device, which uses the voice interaction method described in any one of the above. Figure 2 An example of a physical structure diagram of an electronic device is shown below. Figure 2 As shown, the electronic device may include: a processor 1, a communication interface 2, a memory 3, and a communication bus 4, wherein the processor 1, the communication interface 2, and the memory 3 communicate with each other via the communication bus 4. The processor 1 may call the logic instructions in the memory 3 to execute the voice interaction method, which includes:

[0079] S10: Receive first voice input information from a user, and determine whether the first voice input information includes a preset wake-up word;

[0080] S20: Determine whether the first voice input information includes a preset wake-up word, analyze the sound frequency of the first voice input information, and output a target sound frequency range;

[0081] S30, receiving second voice input information from the user, and analyzing the sound frequency of the second voice input information;

[0082] S40: Determine whether the sound frequency of the second voice input information is within the target sound frequency range, and parse the second voice input information to obtain first execution request information.

[0083] In addition, the logic instructions in the above-mentioned memory 3 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0084] On the other hand, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the voice interaction method provided above, the method comprising:

[0085] S10: Receive first voice input information from a user, and determine whether the first voice input information includes a preset wake-up word;

[0086] S20: Determine whether the first voice input information includes a preset wake-up word, analyze the sound frequency of the first voice input information, and output a target sound frequency range;

[0087] S30, receiving second voice input information from the user, and analyzing the sound frequency of the second voice input information;

[0088] S40: Determine whether the sound frequency of the second voice input information is within the target sound frequency range, and parse the second voice input information to obtain first execution request information.

[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A voice interaction method, characterized in that: include: Receive first voice input information from a user, and determine whether the first voice input information includes a preset wake-up word; Determining that the first voice input information includes a preset wake-up word, analyzing the sound frequency of the first voice input information, and outputting a target sound frequency range; receiving a second voice input message from a user, and analyzing a sound frequency of the second voice input message; determining that a sound frequency of the second voice input information is within the target sound frequency range, and parsing the second voice input information to obtain first execution request information; In response to receiving a third voice input message within a preset time interval after receiving the second voice input message, analyzing the sound frequency of the third voice input message; determining that a sound frequency of the third voice input information is within the target sound frequency range, analyzing the volume of the second voice input information and the third voice input information to determine, based on the volume, the voice input information carrying the valid user command from the second voice input information and the third voice input information; Determine that the volume of the third voice input information is greater than the volume of the second voice input information, parse the third voice input information to obtain second execution request information, and execute the target instruction based on the second execution request information.

2. The voice interaction method according to claim 1, characterized in that: After parsing the second voice input information to obtain the first execution request information, the method further includes: Outputting a confirmation reminder of the first execution request information; Receive user confirmation of the first execution request information, and execute the target instruction of the first execution request information.

3. The voice interaction method according to claim 1, wherein: After parsing the third voice input information to obtain the second execution request information, the method further includes: Outputting a confirmation reminder of the second execution request information; Receive user confirmation of the second execution request information, and execute the target instruction of the second execution request information.

4. The voice interaction method according to claim 1, wherein: The time difference between receiving the second voice input information and receiving the third voice input information does not exceed 1s.

5. The voice interaction method according to claim 1, wherein: After analyzing the sound frequency of the first voice input information and outputting the target sound frequency range, the method further includes: If it is determined that the input duration of the user's voice input information is greater than a preset input time threshold, the user's voice input information is stopped.

6. The voice interaction method according to claim 1, characterized in that: Also includes: receiving fourth voice input information from the user; It is determined that the fourth voice input information includes a preset termination command word, and the receiving of the user's voice input information is stopped.

7. A control device, characterized in that: include: A voice receiving module, configured to receive a user's first voice input information and a second voice input information; a frequency determination module, configured to determine the sound frequencies of the first voice input information and the second voice input information; A first processing module, configured to determine whether the first voice input information contains a preset wake-up word; a second processing module, configured to determine the volume of the first voice input information and the second voice input information; as well as, a third processing module, configured to output execution request information according to the processing result of the second processing module; Also includes: Determining that the first voice input information includes a preset wake-up word, analyzing the sound frequency of the first voice input information, and outputting a target sound frequency range; In response to receiving a third voice input message within a preset time interval after receiving the second voice input message, analyzing the sound frequency of the third voice input message; determining that a sound frequency of the third voice input information is within the target sound frequency range, analyzing the volume of the third voice input information, and determining, based on the volume, the voice input information carrying the valid user command from the second voice input information and the third voice input information; Determine that the volume of the third voice input information is greater than the volume of the second voice input information, parse the third voice input information to obtain second execution request information, and execute the target instruction based on the second execution request information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the voice interaction method according to any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the voice interaction method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Control device

    CN112687264A