Voice command response method and device, storage medium and electronic device
By adjusting the waiting time according to the network quality, the voice control delay problem of smart devices in poor network conditions is solved, and fast response is achieved in different network environments, improving the user experience.
Patent Information
- Application Number
- CN202310095496.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-31
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-01-31
AI Technical Summary
Smart devices cannot flexibly adjust the waiting time when the network is poor, resulting in longer voice control response time and affecting the user experience.
The waiting time is determined according to the network level of the current network quality. If the first voice recognition control instruction fed back by the set voice server is received within the waiting time, a response is made. Otherwise, the voice control information is recognized and responded through the preset voice recognition model.
By flexibly adjusting the waiting time, the smoothness of voice interaction and user experience are improved, ensuring that voice commands can be responded to quickly in any network conditions.
Smart Images

Figure CN116092498B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of smart home technology, and more specifically, to a method and device for responding to voice commands, a storage medium, and an electronic device. Background Art
[0002] To better utilize voice functions, smart devices offer two voice control modes: online and offline voice commands. When the device is offline, offline voice commands are used. When the device is online, to prevent online voice commands from timing out, both offline and online voice commands can be received. However, online voice commands have a richer vocabulary, so online voice commands are preferred. Because offline voice commands are provided by local algorithms, they are received before online voice commands. Therefore, after receiving an offline voice command, the device waits for a period of time. If the online voice command is not received after the timeout, the offline voice command is broadcast and executed. Otherwise, the online voice command is broadcast and executed.
[0003] However, in some scenarios, this voice interaction method will affect the user experience. For example, when the device is connected to the Internet but the network is poor, the online voice command is likely to time out, but the smart device still needs to wait until the timeout before executing the offline voice command, which increases the time from user voice control to device execution and affects the user's voice interaction experience.
[0004] Regarding the problem in related technologies that smart devices cannot flexibly determine the waiting time according to network conditions, no effective solution has been proposed. Summary of the Invention
[0005] The embodiments of the present application provide a method and device for responding to voice commands, a storage medium, and an electronic device to at least solve the problem in the related art that smart devices cannot flexibly determine the waiting time according to network conditions.
[0006] According to one embodiment of the embodiments of the present application, a method for responding to voice commands is provided, including: when voice control information is received, sending the voice control information to a set voice server for voice recognition, and determining a waiting time based on a network level of current network quality; if a first voice recognition control command fed back by the set voice server is received within the waiting time, responding to the first voice recognition control command; if the first voice recognition control command fed back by the set voice server is not received within the waiting time, recognizing the voice control information based on a preset voice recognition model, and obtaining and responding to a second voice recognition control command.
[0007] In an exemplary embodiment, the waiting time is determined based on the network level of the current network quality, including: determining a network level coefficient corresponding to the network level; determining the number of times a third voice recognition control instruction is responded to within a preset time period, and determining a first size relationship between the number and a first preset threshold, wherein the third voice recognition control instruction is a control instruction obtained by the preset voice recognition model recognizing historical voice control information; and determining the waiting time based on the first size relationship and the network level coefficient.
[0008] In an exemplary embodiment, the waiting period is determined based on the first size relationship and the network level coefficient, including: when the first size relationship indicates that the number of times is less than or equal to the first preset threshold, determining a second size relationship between the number of times and a second preset threshold, and determining the waiting period based on the second size relationship and the network level coefficient; when the first size relationship indicates that the number of times is greater than the first preset threshold, determining the waiting period as a target value.
[0009] In an exemplary embodiment, the waiting time is determined based on the second size relationship and the network level coefficient, including: when the second size relationship indicates that the number of times is greater than or equal to the second preset threshold, determining the default waiting time coefficient of the smart device, calculating the first product of the waiting time coefficient, the network level coefficient and the default waiting time, and determining the first product as the waiting time; when the second size relationship indicates that the number of times is greater than the first preset threshold, calculating the second product of the network level coefficient and the default waiting time, and determining the second product as the waiting time.
[0010] In an exemplary embodiment, before determining the waiting time based on the network level of the current network quality, the method also includes: determining whether to establish a network connection with the network distribution device; if a network connection is established with the network distribution device, determining the current network quality of the network connection established with the network distribution device; if a network connection is not established with the network distribution device, recognizing the voice control information based on a preset voice recognition model, and obtaining and responding to the second voice recognition control instruction.
[0011] In an exemplary embodiment, after determining the waiting time based on the network level of the current network quality, the method includes at least one of the following: periodically determining the current network quality and updating the waiting time based on the determined current network quality; in the case of a successful response to the first voice recognition control instruction, determining the current network quality and updating the waiting time based on the determined current network quality.
[0012] In an exemplary embodiment, after determining the waiting time based on the network level of the current network quality, the method further includes: if the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, adding a target value to the number of times the smart device responds to the third voice recognition control instruction within a preset time period; if the first voice recognition control instruction fed back by the set voice server is received within the waiting time, clearing the number of times the smart device responds to the third voice recognition control instruction within the preset time period.
[0013] According to another embodiment of the embodiment of the present application, a voice instruction response device is also provided, including: a determination module, which is used to send the voice control information to a set voice server for voice recognition when receiving voice control information, and determine the waiting time according to the network level of the current network quality; a response module, which is used to respond to the first voice recognition control instruction if the first voice recognition control instruction fed back by the set voice server is received within the waiting time; and an identification module, which is used to identify the voice control information based on a preset voice recognition model if the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, and obtain and respond to a second voice recognition control instruction.
[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for responding to voice instructions when running.
[0015] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned method for responding to voice commands through the computer program.
[0016] In an embodiment of the present application, when voice control information is received, the voice control information is sent to a set voice server for voice recognition, and the waiting time is determined according to the network level of the current network quality; if a first voice recognition control instruction fed back by the set voice server is received within the waiting time, the first voice recognition control instruction is responded to; if the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the voice control information is recognized based on a preset voice recognition model, and a second voice recognition control instruction is obtained and responded to; the above technical solution is adopted to solve the problem that the smart device cannot flexibly determine the waiting time according to the network situation. The embodiment of the present invention can determine the waiting time of the smart device by the current network quality of the smart device, and then flexibly determine the waiting time, and flexibly determine the first voice recognition control instruction fed back by the voice server or the second voice recognition control instruction recognized by the voice recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 This is a schematic diagram of the hardware environment of a method for responding to voice commands according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of a method for responding to a voice command according to an embodiment of the present application;
[0021] Figure 3 is a schematic diagram of a method for responding to a voice command according to an optional embodiment of the present application;
[0022] Figure 4 is a flow chart of a method for responding to a voice instruction according to an optional embodiment of the present application;
[0023] Figure 5 This is a structural block diagram of a device for responding to voice commands according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to one aspect of the embodiment of the present application, a method for responding to a voice command is provided. The method for responding to a voice command is widely used in smart home (Smart Home), smart home, smart home device ecology, smart home (IntelligenceHouse) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned method for responding to a voice command can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal device 102 and a server 104. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.
[0027] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, and a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal device 102 may be, but is not limited to, a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing machine, a smart dishwasher, a smart projection device, a smart TV, a smart clothes drying rack, smart curtains, smart audio and video, a smart socket, a smart speaker, a smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purifier, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.
[0028] In this embodiment, a method for responding to a voice command is provided, which is applied to a computer terminal. Figure 2 : is a flowchart of a method for responding to a voice command according to an embodiment of the present application, the process comprising the following steps:
[0029] Step S202, when the voice control information is received, the voice control information is sent to the set voice server for voice recognition, and the waiting time is determined according to the network level of the current network quality;
[0030] It should be noted that the current network quality can be understood as the received signal strength, where the unit of the received signal strength is dBm. Normally, -70dBm is the ideal signal strength, -70dBm to -80dBm is the medium signal strength, and less than -80dBm is the weak signal strength. The embodiment of the present invention can preset the value ranges of ideal signal strength, medium signal strength, and weak signal strength according to actual conditions.
[0031] It should be noted that received signal strength = transmit power + receive antenna gain + transmit antenna gain - cable and cable head attenuation - free space loss.
[0032] Step S204: If the first voice recognition control instruction fed back by the set voice server is received within the waiting time, respond to the first voice recognition control instruction;
[0033] Step S206: If the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the voice control information is recognized based on a preset voice recognition model, and a second voice recognition control instruction is obtained and responded to.
[0034] Through the above steps, when voice control information is received, the voice control information is sent to the set voice server for voice recognition, and the waiting time is determined according to the network level of the current network quality; if the first voice recognition control instruction fed back by the set voice server is received within the waiting time, the first voice recognition control instruction is responded to; if the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the voice control information is recognized based on the preset voice recognition model, and the second voice recognition control instruction is obtained and responded to, which solves the problem in the related art that the smart device cannot flexibly determine the waiting time according to the network situation. The embodiment of the present invention can determine the waiting time of the smart device by the current network quality of the smart device, and then flexibly determine the waiting time, and flexibly determine the response to the first voice recognition control instruction fed back by the voice server or the second voice recognition control instruction recognized by the voice recognition model.
[0035] In step S204, there are many ways to determine the waiting time based on the network level of the current network quality. An embodiment of the present invention provides an implementation method, specifically: determining the network level coefficient corresponding to the network level; determining the number of times a third voice recognition control instruction is responded to within a preset time period, and determining a first size relationship between the number and a first preset threshold, wherein the third voice recognition control instruction is a control instruction obtained by the preset voice recognition model recognizing historical voice control information; determining the waiting time based on the first size relationship and the network level coefficient.
[0036] It should be noted that network levels can be understood as: ideal signal strength level, medium signal strength level, weak signal strength level; and can also be understood as: first level, second level, ..., Nth level. When the first level is the highest level, the Nth level is the lowest level; when the first level is the lowest level, the Nth level is the highest level. In the embodiments of the present invention, the signal strength value ranges of different levels can be preset according to actual conditions.
[0037] For example, -70dBm is preset as ideal signal strength, -70dBm to -80dBm is medium signal strength, and less than -80dBm is weak signal strength. The network level coefficient corresponding to ideal signal strength is 1, the network level coefficient corresponding to medium signal strength is 0.8, and the network level coefficient corresponding to weak signal strength is 0.6. If the smart device determines that the current network quality is -75dBm, it determines that the current network quality is medium signal strength, and the corresponding network level coefficient is 0.8.
[0038] In an exemplary embodiment, determining the waiting time according to the first size relationship and the network level coefficient includes the following two cases:
[0039] 1) if the first size relationship indicates that the number of times is less than or equal to the first preset threshold, determining a second size relationship between the number of times and a second preset threshold, and determining the waiting time according to the second size relationship and the network level coefficient;
[0040] 2) When the first size relationship indicates that the number of times is greater than the first preset threshold, determine the waiting time as a target value.
[0041] That is to say, in case 1), it indicates that the online voice command occasionally times out. Therefore, further determining the waiting time of the smart device based on the second size relationship and the network level coefficient can further shorten the waiting time of the smart device. In case 2), it indicates that the online voice command times out multiple times in a row, which assists in determining that the network strength at this time is very poor, and directly executes the control instructions obtained by the preset voice recognition model to recognize the voice control information, thereby solving the problem in the related art that the smart device cannot flexibly determine the waiting time according to the network situation. The embodiment of the present invention can determine the waiting time by the current network quality, and thus can flexibly determine the waiting time, as well as flexibly determine the first voice recognition control instruction fed back by the voice server or the second voice recognition control instruction recognized by the voice recognition model.
[0042] In an exemplary embodiment, the waiting time is determined based on the second size relationship and the network level coefficient, including: when the second size relationship indicates that the number of times is greater than or equal to the second preset threshold, determining the default waiting time coefficient, calculating the first product of the waiting time coefficient, the network level coefficient and the default waiting time, and determining the first product as the waiting time; when the second size relationship indicates that the number of times is greater than the first preset threshold, calculating the second product of the network level coefficient and the default waiting time, and determining the second product as the waiting time.
[0043] For example, -70dBm is preset as the ideal signal strength, -70dBm to -80dBm is the medium signal strength, and less than -80dBm is the weak signal strength. The network level coefficient corresponding to the ideal signal strength is 1, the network level coefficient corresponding to the medium signal strength is 0.8, and the network level coefficient corresponding to the weak signal strength is 0.6. The second preset threshold is 10, and the waiting time coefficient is 0.8. When the smart device determines that the current network quality is -75dBm, it determines that the current network quality is medium signal strength, and the corresponding network level coefficient is 0.8; when the smart device continuously executes the preset voice recognition model to recognize the control instructions obtained by the voice control information 9 times, it determines that the waiting time = 0.8 * default waiting time. When the smart device continuously executes the preset voice recognition model to recognize the control instructions obtained by the voice control information 13 times, it determines that the waiting time = 0.8 * 0.8 * default waiting time.
[0044] That is, in an embodiment of the present invention, when the second size relationship indicates that the number of times is less than or equal to the second preset threshold, the waiting time is determined by multiplying the waiting time coefficient, the network level coefficient, and the default waiting time by the second product; when the second size relationship indicates that the number of times is greater than the second preset threshold, upon receiving voice control information, the voice control information is recognized based on a preset voice recognition model, and a second voice recognition control instruction is obtained and responded to, without waiting for the first voice recognition control instruction fed back by the preset voice server. This solves the problem in the related art that smart devices cannot flexibly determine the waiting time based on network conditions, and achieves the technical effect of flexibly determining the waiting time based on the current network quality and the number of control instructions obtained by executing the preset voice recognition model to recognize voice control information.
[0045] In an exemplary embodiment, before determining the waiting time based on the network level of the current network quality, it also includes: determining whether to establish a network connection with the network distribution device; when a network connection is established with the network distribution device, determining the current network quality of the network connection established with the network distribution device; when a network connection is not established with the network distribution device, recognizing the voice control information based on a preset voice recognition model, and obtaining and responding to the second voice recognition control instruction.
[0046] That is, after receiving an offline voice command, it is necessary to determine whether the current smart device is connected to the network. If not, the voice control information is immediately recognized based on the preset voice recognition model, the second voice recognition control instruction is obtained and responded to, and the current operation is terminated. If the network is already connected, it is necessary to continue to determine the current network quality and determine the waiting time based on the current network quality. Compared with the related art, which requires a preset waiting time regardless of whether the network is connected, the embodiment of the present invention can more quickly and flexibly determine whether to execute the control instruction obtained by the preset voice recognition model recognizing the voice control information.
[0047] In an exemplary embodiment, after determining the waiting time based on the network level of the current network quality, the method includes at least one of the following: periodically determining the current network quality and updating the waiting time based on the determined current network quality; in the case of a successful response to the first voice recognition control instruction, determining the current network quality and updating the waiting time based on the determined current network quality.
[0048] It should be noted that the embodiment of the present invention needs to trigger the detection of the current network quality. Therefore, it is necessary to determine the trigger conditions for triggering the detection of the current network quality, specifically: 1) periodically detect the current network quality, for example, set a periodic timing detection time, for example, detect the current network quality every 40 seconds, and update the waiting time according to the current network quality; 2) in the case of successfully responding to the first voice recognition control instruction, detect the current network quality, and update the waiting time according to the determined current network quality.
[0049] In an exemplary embodiment, after determining the waiting time based on the network level of the current network quality, it also includes: if the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the number of times the third voice recognition control instruction is responded to within the preset time period will be increased by a target value; if the first voice recognition control instruction fed back by the set voice server is received within the waiting time, the number of times the third voice recognition control instruction is responded to within the preset time period will be reset to zero.
[0050] It should be noted that the above target value can be understood as one.
[0051] In order to better understand the process of the above-mentioned voice instruction response method, the following describes the implementation method flow of the above-mentioned voice instruction response in combination with an optional embodiment, but is not used to limit the technical solution of the embodiment of the present application.
[0052] In this embodiment, a method for responding to a voice command is provided. Figure 4 is a flow chart of a method for responding to a voice instruction according to an optional embodiment of the present application, such as Figure 4As shown, the specific steps are as follows:
[0053] Step S401: Start;
[0054] Step S402: collecting network signal strength data (equivalent to the current network quality in the above embodiment);
[0055] Smart devices detect the current network status in real time and classify network signals. The worse the signal, the lower the level and the lower the network signal coefficient.
[0056] Step S403: receiving an offline voice command (equivalent to the second voice recognition control instruction in the above embodiment);
[0057] Receive offline voice commands. Offline voice commands are received before online voice commands (equivalent to the first voice recognition control instruction in the above embodiment). Therefore, it is necessary to determine whether to execute offline voice commands based on actual conditions. Figure 3 As shown in the figure below, because the offline voice command is provided by the local algorithm, it is received before the online voice command (time 1). If the interval between the two commands (interval) is greater than the set threshold (interval threshold), the online voice command times out and the offline voice command is executed. If the network signal is poor, the online voice command response is slow and likely to time out. In this case, the smart device still needs to wait for the set threshold time to expire before executing.
[0058] Step S404: Determine whether the smart device is connected to the network;
[0059] After receiving the offline voice command, it is necessary to determine whether the smart device is connected to the network. If not, step S412 is immediately executed; if it is connected, step S405 is executed;
[0060] Step S405: Detecting whether the voice is in an offline state. If the voice is in an offline state, executing step S406; otherwise, executing step S408;
[0061] Regardless of the current network status, a system wait time (T) must be set regularly. The purpose of this system wait time is to periodically detect the network status and determine which voice command to use next. This design is mainly for situations with complex network conditions. Initially, when the network status is poor, the device shortens the wait time (such as 0.8*T) or even directly uses offline voice commands to improve voice interaction response. However, at intervals (such as 30 seconds), a longer wait time (such as T) is set. If the online voice command still times out, the wait time is further shortened (0.8T). If the online voice command is received before the timeout, the system wait time (T) is restored.
[0062] Step S406: Set the device waiting time to the system waiting time;
[0063] Step S407: Check whether the number of consecutive executions of the offline voice command is greater than a maximum value M. If so, execute step S408; otherwise, execute step S412.
[0064] Step S408: Detect whether the number of consecutive executions of the offline voice command is greater than the expected number N for shortening the waiting time. If so, execute step S409; otherwise, execute step S410.
[0065] Step S409: determining a timeout period according to the network signal strength and the waiting time coefficient (timeout period = T * network signal strength coefficient * waiting time coefficient);
[0066] Step S410: determining a timeout period according to the network signal strength (timeout period = T * network signal strength coefficient);
[0067] Step S411: Start the timer.
[0068] Step S412: executing offline voice instructions;
[0069] Step S413: End.
[0070] In one embodiment, for example, the calculation method of the waiting time corresponding to each network state is shown in Table 1:
[0071] Table 1
[0072]
[0073] It should be noted that: the network level coefficient <= 1, the waiting time coefficient < 1
[0074] The system default waiting time is 1500 milliseconds, the network level coefficient is (Level 1: 0.8; Level 2: 0.9; Level 3: 0.95), the current network level is Level 2, the waiting time coefficient is 0.7, N is 10 times, M is 25 times, and the voice online detection time is 40 seconds.
[0075] If the online voice command instruction has not timed out, the waiting time is 1500 milliseconds;
[0076] If the online voice command occasionally times out 1-2 times (not exceeding N), the waiting time is 1500 milliseconds;
[0077] If online voice commands frequently time out, the waiting time is 1500 milliseconds before 10 consecutive timeouts (N), and the number of consecutive offline voice command executions is accumulated;
[0078] After 10 consecutive timeouts (N), the waiting time is adjusted to 945 milliseconds (system default waiting time * network level coefficient * waiting time coefficient), and the number of consecutive offline voice command executions is accumulated;
[0079] After 25 consecutive timeouts (M), the offline voice command is executed immediately after receiving it, without waiting for the online voice command, and the number of consecutive offline voice command executions is accumulated;
[0080] In any of the above cases, the online voice timeout period is set to 40 seconds every 40 seconds (the time it takes for the voice to be offline). If the online voice still times out, the current timeout period is still used and the number of consecutive offline voice command executions is accumulated. If the online voice command does not time out, the timeout period is restored to 1500 milliseconds (the system default waiting period) and the number of consecutive offline voice command executions is reset to zero.
[0081] It should be noted that as long as the online voice command has not timed out, the timeout period needs to be restored to 1500 milliseconds (the system default waiting time) and the number of consecutive offline voice command executions needs to be cleared.
[0082] It should be noted that the smart device in the embodiment of the present invention can send the voice control information to the set voice server for voice recognition at the same time as receiving the voice control information, and recognize the voice control information through the preset voice recognition model of the smart device; or it can first send the voice control information to the set voice server for voice recognition when receiving the voice control information. If the online voice command fed back by the set voice server is not received within the waiting time, the voice control information is recognized by the preset voice recognition model to obtain and respond to the offline voice command. The embodiment of the present invention is not limited to this.
[0083] Through the above embodiments, when the device is connected to the network but the network signal is poor, the network signal is detected and combined with the previous offline and online voice command usage, the time waiting for the online voice command after receiving the offline voice command is optimized, and the voice command response time is shortened when the network signal is poor. At the same time, the voice command response speed when the network signal is good is not affected, allowing users to experience smooth voice interaction in any network conditions, thereby improving the fluency of voice conversations and user experience.
[0084] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0085] Figure 5 is a structural block diagram of a device for responding to voice commands according to an embodiment of the present application; Figure 5 Shown, including:
[0086] The determination module 52 is configured to, upon receiving the voice control information, send the voice control information to a predetermined voice server for voice recognition and determine a waiting time based on a network level of the current network quality;
[0087] A response module 54 is configured to respond to the first voice recognition control instruction if the first voice recognition control instruction fed back by the set voice server is received within the waiting time;
[0088] The recognition module 56 is configured to recognize the voice control information based on a preset voice recognition model and obtain and respond to a second voice recognition control instruction if the first voice recognition control instruction fed back by the set voice server is not received within the waiting period.
[0089] Through the above-mentioned device, when voice control information is received, the voice control information is sent to the set voice server for voice recognition, and the waiting time is determined according to the network level of the current network quality; if the first voice recognition control instruction fed back by the set voice server is received within the waiting time, the first voice recognition control instruction is responded to; if the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the voice control information is recognized based on the preset voice recognition model, and the second voice recognition control instruction is obtained and responded to, which solves the problem in the related art that the smart device cannot flexibly determine the waiting time according to the network situation. The embodiment of the present invention can determine the waiting time according to the current network quality, and then flexibly determine the waiting time, and flexibly determine the response to the first voice recognition control instruction fed back by the voice server or the second voice recognition control instruction recognized by the voice recognition model.
[0090] In an exemplary embodiment, the determination module 52 is used to determine the network level coefficient corresponding to the network level; determine the number of times a third voice recognition control instruction is responded to within a preset time period, and determine a first size relationship between the number and a first preset threshold; and determine the waiting time based on the first size relationship and the network level coefficient.
[0091] In an exemplary embodiment, the determination module 52 is used to determine a second size relationship between the number of times and a second preset threshold when the first size relationship indicates that the number of times is less than or equal to the first preset threshold, and determine the waiting time based on the second size relationship and the network level coefficient; when the first size relationship indicates that the number of times is greater than the first preset threshold, determine the waiting time as a target value.
[0092] In an exemplary embodiment, the determination module 52 is used to determine the default waiting time coefficient when the second size relationship indicates that the number of times is greater than or equal to the second preset threshold, calculate the first product of the waiting time coefficient, the network level coefficient and the default waiting time, and determine the first product as the waiting time; when the second size relationship indicates that the number of times is greater than the first preset threshold, calculate the second product of the network level coefficient and the default waiting time, and determine the second product as the waiting time.
[0093] In an exemplary embodiment, the determination module 52 is used to determine whether a network connection is established with the network distribution device; if a network connection is established with the network distribution device, determine the current network quality of the network connection established with the network distribution device; if a network connection is not established with the network distribution device, recognize the voice control information based on a preset voice recognition model, and obtain and respond to the second voice recognition control instruction.
[0094] In an exemplary embodiment, the determination module 52 is used to do at least one of the following: periodically determine the current network quality and update the waiting time based on the determined current network quality; in the case of a successful response to the first voice recognition control instruction, determine the current network quality and update the waiting time based on the determined current network quality.
[0095] In an exemplary embodiment, the determination module 52 is used to add a target value to the number of times the smart device responds to the third voice recognition control instruction within a preset time period if the first voice recognition control instruction fed back by the set voice server is not received within the waiting period; and to clear the number of times the smart device responds to the third voice recognition control instruction within the preset time period if the first voice recognition control instruction fed back by the set voice server is received within the waiting period.
[0096] An embodiment of the present application further provides a storage medium, which includes a stored program, wherein the program executes any of the above methods when it is run.
[0097] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:
[0098] S1, when receiving voice control information, sending the voice control information to the set voice server for voice recognition, and determining the waiting time according to the network level of the current network quality;
[0099] S2, if the first voice recognition control instruction fed back by the set voice server is received within the waiting time, responding to the first voice recognition control instruction;
[0100] S3: If the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the voice control information is recognized based on a preset voice recognition model, and a second voice recognition control instruction is obtained and responded to.
[0101] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0102] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0103] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0104] S1, when receiving voice control information, sending the voice control information to the set voice server for voice recognition, and determining the waiting time according to the network level of the current network quality;
[0105] S2, if the first voice recognition control instruction fed back by the set voice server is received within the waiting time, responding to the first voice recognition control instruction;
[0106] S3: If the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the voice control information is recognized based on a preset voice recognition model, and a second voice recognition control instruction is obtained and responded to.
[0107] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.
[0108] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0109] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into separate integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0110] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for responding to a voice command, characterized in that: include: When receiving the voice control information, the voice control information is sent to the set voice server for voice recognition to determine the network level coefficient corresponding to the network level; Determining the number of times a third voice recognition control instruction is responded to within a preset time period, and determining a first size relationship between the number and a first preset threshold, wherein the third voice recognition control instruction is a control instruction obtained by a preset voice recognition model recognizing historical voice control information; if the first size relationship indicates that the number is less than or equal to the first preset threshold, determining a second size relationship between the number and a second preset threshold, and determining a waiting time based on the second size relationship and the network level coefficient; if the first size relationship indicates that the number is greater than the first preset threshold, determining the waiting time to be a target value; If the first voice recognition control instruction fed back by the set voice server is received within the waiting time, responding to the first voice recognition control instruction; If the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the voice control information is recognized based on a preset voice recognition model, and a second voice recognition control instruction is obtained and responded to.
2. The method for responding to a voice command according to claim 1, wherein: Determining the waiting time according to the second size relationship and the network level coefficient includes: If the second size relationship indicates that the number of times is greater than or equal to the second preset threshold, determining a default waiting time coefficient, calculating a first product of the waiting time coefficient, the network level coefficient, and the default waiting time, and determining the first product as the waiting time; When the second size relationship indicates that the number of times is greater than the first preset threshold, a second product of the network level coefficient and the default waiting time is calculated, and the second product is determined to be the waiting time.
3. The method for responding to a voice command according to claim 1, wherein: Before determining the waiting time according to the network level of the current network quality, the method further includes: Determine whether a network connection is established with the distribution network device; In case a network connection is established with the network distribution device, determining a current network quality of the network connection established with the network distribution device; In the case where a network connection is not established with the network distribution device, the voice control information is recognized based on a preset voice recognition model, and the second voice recognition control instruction is obtained and responded to.
4. The method for responding to a voice command according to claim 1, wherein: After determining the waiting time according to the network level of the current network quality, the method includes at least one of the following: Periodically determining the current network quality, and updating the waiting time according to the determined current network quality; In case of a successful response to the first voice recognition control instruction, the current network quality is determined, and the waiting time is updated according to the determined current network quality.
5. The method for responding to a voice command according to claim 1, wherein: After determining the waiting time according to the network level of the current network quality, the method further includes: If the first voice recognition control instruction fed back by the set voice server is not received within the waiting time, the number of responses to the third voice recognition control instruction within the preset time period is accumulated; If the first voice recognition control instruction fed back by the set voice server is received within the waiting time, the number of times of responding to the third voice recognition control instruction within the preset time period will be reset to zero.
6. A device for responding to voice commands, characterized in that: include: a determination module configured to, upon receiving voice control information, send the voice control information to a set voice server for voice recognition and determine a network level coefficient corresponding to the network level; Determining the number of times a third voice recognition control instruction is responded to within a preset time period, and determining a first size relationship between the number and a first preset threshold, wherein the third voice recognition control instruction is a control instruction obtained by a preset voice recognition model recognizing historical voice control information; determining a waiting time based on the first size relationship and the network level coefficient; if the first size relationship indicates that the number is less than or equal to the first preset threshold, determining a second size relationship between the number and a second preset threshold, and determining the waiting time based on the second size relationship and the network level coefficient; if the first size relationship indicates that the number is greater than the first preset threshold, determining the waiting time to be a target value; a response module, configured to respond to the first voice recognition control instruction if the first voice recognition control instruction fed back by the set voice server is received within the waiting time; The recognition module is used to recognize the voice control information based on a preset voice recognition model and obtain and respond to a second voice recognition control instruction if the first voice recognition control instruction fed back by the set voice server is not received within the waiting time.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the method according to any one of claims 1 to 5 is executed when the program is executed.
8. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 5 through the computer program.
Citation Information
Patent Citations
Online and offline dual-mode voice control method and system and storage medium
CN112735414A