Voice control method and device, electronic equipment and storage medium

By matching the target speech recognition model according to the device status and directly identifying the user's voice signal, the defects in the prior art need to wake up words are solved, fast and accurate voice interaction is achieved, and user experience is improved.

CN120199243APending Publication Date: 2025-06-24HEFEI MIDEA REFRIGERATOR CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311797595.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing voice control technology requires users to enter wake-up words first, resulting in complex and cumbersome interactions, reducing the human-computer dialogue experience.

Method used

By matching the target speech recognition model from the speech recognition model library based on the device status of the target device, the user's voice signal is directly recognized without inputting a wake-up word.

Benefits of technology

It realizes fast and accurate interaction without user inputting wake words, improving the accuracy and user experience of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199243A_ABST
    Figure CN120199243A_ABST
Patent Text Reader

Abstract

The invention provides a voice control method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: receiving a voice signal sent by a user; matching a target speech recognition model from a speech recognition model library according to the equipment state of the target equipment; inputting the voice signal into a target voice recognition model; and if it is determined that the target voice recognition model can recognize the voice signal, executing a control instruction corresponding to the control command word output by the target voice recognition model. According to the voice control method and device, the electronic equipment and the storage medium provided by the invention, user wakeup is not needed, the scene condition of the equipment state is effectively utilized, the corresponding voice recognition model can be matched according to different real-time states of the target equipment, the voice recognition model can recognize the voice signal in a targeted manner, and the user experience is improved. The man-machine interaction process is simplified, the voice recognition precision is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a voice control method, apparatus, electronic device, and storage medium. Background Art

[0002] In order to avoid accidental triggering of functions, devices supporting voice control on the current market basically require users to first input a wake-up word and then input a corresponding control command word before executing the corresponding command, such as playing music, asking about the weather, etc., that is, wake-up is required before each voice control. This control method results in complex interaction. Generally, the entire process takes several seconds to complete, and the dialogue interaction process is cumbersome, greatly reducing the human-computer dialogue experience.

[0003] Of course, there are also some devices supporting voice control. If there are relatively few command words, such as two or more, one of the command words will be directly used as the wake-up word, and the other command words will be used as control command words. At this time, a multi-wake-up word recognition model needs to be used for recognition. However, in the case of the same hardware computing power, because the multi-wake-up word recognition model needs to recognize two or more command words simultaneously, the recognition effect is worse than that of a single command word recognition model, and the probability of false wake-up will increase, etc. Summary of the Invention

[0004] The present invention provides a voice control method, apparatus, electronic device, and storage medium to solve the defect in the prior art that a wake-up word needs to be recognized first, and to achieve fast and accurate interaction without the need for users to input a wake-up word. While improving the recognition accuracy, the user experience can be greatly improved.

[0005] In a first aspect, the present invention provides a voice control method, including:

[0006] Receiving a voice signal sent by a user;

[0007] In response to the voice signal, matching a target voice recognition model from a voice recognition model library according to the device state of the target device;

[0008] Inputting the voice signal into the target voice recognition model;

[0009] If it is determined that the target voice recognition model can recognize the voice signal, executing a control instruction corresponding to the control command word output by the target voice recognition model.

[0010] According to the voice control method provided by the present invention, there are multiple types of device states;

[0011] For different device states, the target voice recognition models matched from the voice recognition model library are different.

[0012] A voice control method provided according to the present invention, wherein the device state includes a device door opening state and a device door closing state;

[0013] Responding to the voice signal, and matching a target voice recognition model from a voice recognition model library according to the device state of the target device, includes:

[0014] When it is determined that the device state is the device door opening state, matching a door closing command recognition model from the voice recognition model library as the target voice recognition model;

[0015] When it is determined that the device state is the device door closing state, matching a door opening command recognition model from the voice recognition model library as the target voice recognition model;

[0016] The command words that can be recognized by the door closing command recognition model are matched from a first command word library according to the signal characteristics of the voice signal;

[0017] The command words that can be recognized by the door opening command recognition model are matched from a second command word library according to the signal characteristics of the voice signal;

[0018] The first command word library is a set of command words that can be recognized by the door closing command recognition model set in advance, and the second command word library is a set of command words that can be recognized by the door opening command recognition model set in advance.

[0019] A voice control method provided according to the present invention, wherein the first command word library only includes door closing control command words, and the second command word library only includes door opening control command words;

[0020] The control instruction corresponding to the door closing control command word is to control the target device to close the door, and the control instruction corresponding to the door opening control command word is to control the target device to open the door.

[0021] A voice control method provided according to the present invention, wherein the device state further includes a human-machine distance state, and the human-machine distance state is used to represent the distance interval in which the distance between the user and the target device is located;

[0022] Responding to the voice signal, and matching a target voice recognition model from a voice recognition model library according to the device state of the target device, includes:

[0023] According to the distance interval determined by the human-machine distance state, matching a control command recognition model from the voice recognition model library as the target voice recognition model;

[0024] Different distance intervals correspond to different control command recognition models, and each control command recognition model can only recognize a preset control command word.

[0025] According to a voice control method provided by the present invention, the device state further includes a human-computer interaction state, and the human-computer interaction state is determined according to the interaction behavior between the user and the target device.

[0026] Responding to the voice signal and matching a target voice recognition model from the voice recognition model library according to the device state of the target device includes:

[0027] According to the interaction behavior determined by the human-computer interaction state, a control command recognition model is matched from the voice recognition model library as the target voice recognition model.

[0028] Different interaction behaviors correspond to different control command recognition models, and each control command recognition model can only recognize a preset control command word.

[0029] According to a voice control method provided by the present invention, before receiving the voice signal sent by the user, it further includes:

[0030] When it is determined that the device state of the target device meets a preset condition, the target device is switched to the voice control waiting mode.

[0031] A signal acquisition prompt is output to the user to prompt the user that the target device can currently accept voice signal control.

[0032] The signal acquisition prompt includes at least one of a voice prompt, a light prompt, and a screen display prompt.

[0033] The device state meeting the preset condition includes at least one of the device state being the device door opening state, the distance between the user and the target device being less than a preset distance threshold, a preset interaction behavior occurring between the user and the target device, and the acceleration of the reduction of the distance between the user and the target device being greater than a preset acceleration threshold.

[0034] According to a voice control method provided by the present invention, if it is determined that the target voice recognition model cannot recognize the voice signal, the user is prompted to input a wake-up voice signal.

[0035] After receiving the wake-up voice signal, a multi-semantic recognition model is matched from the voice recognition model library, and the user is prompted to input a control voice signal.

[0036] The control voice signal is input into the multi-semantic recognition model, and the control instruction corresponding to the control command word output by the multi-semantic recognition model is executed.

[0037] In a second aspect, the present invention further provides a voice control device, comprising:

[0038] a voice collection unit, configured to receive a voice signal sent by a user;

[0039] a model calling unit, configured to, in response to the voice signal, match a target voice recognition model from a voice recognition model library according to the device state of a target device;

[0040] a voice processing unit, configured to input the voice signal into the target voice recognition model;

[0041] a control execution unit, configured to execute a control instruction corresponding to a control command word output by the target voice recognition model after determining that the target voice recognition model can recognize the voice signal.

[0042] According to a voice control device provided by the present invention, the target device is a refrigerator.

[0043] In a third aspect, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the voice control method as described in any one of the above is implemented.

[0044] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the voice control method as described in any one of the above is implemented.

[0045] The voice control method, device, electronic device, and storage medium provided by the present invention do not require user wake-up, effectively utilize the scenario condition of the device state, can match a corresponding voice recognition model according to the different real-time states of the target device, and the voice recognition model will specifically recognize the voice signal, which not only simplifies the human-computer interaction process, improves the accuracy of voice recognition, but also enhances the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0047] Figure 1 is one of the schematic flowcharts of the voice control method provided by the present invention;

[0048] Figure 2It is the second schematic flowchart of the voice control method provided by the present invention;

[0049] Figure 3 It is the third schematic flowchart of the voice control method provided by the present invention;

[0050] Figure 4 It is the fourth schematic flowchart of the voice control method provided by the present invention;

[0051] Figure 5 It is the fifth schematic flowchart of the voice control method provided by the present invention;

[0052] Figure 6 It is the sixth schematic flowchart of the voice control method provided by the present invention;

[0053] Figure 7 It is the schematic structural diagram of the voice control device provided by the present invention;

[0054] Figure 8 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0055] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without any creative efforts shall fall within the protection scope of the present invention.

[0056] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. Unless otherwise clearly specified and defined, the terms "mount", "connect", "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0057] The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and do not limit the number of objects. For example, the first object can be one or multiple.

[0058] Voice control means that when the device is in the standby state, it receives a specific voice signal sent by the user and enters the working state to perform a certain operation. Currently, various devices generally have the voice control function. For example, smart phones, wearable devices, in-vehicle devices, smart home devices, etc. can all be controlled by voice.

[0059] In the existing voice control methods, it is mainly necessary to wake up through a wake-up word. When the device detects that the voice signal input by the user contains a preset wake-up word, it will start to enter the subsequent execution process. The main disadvantage of the current voice control method is that every time the user interacts with the device, it is necessary to first input the preset wake-up word by voice. In actual scenarios, most users do not want to input the wake-up word every time they talk to the device, because this way the dialogue interaction process is cumbersome and will reduce the human-machine dialogue experience.

[0060] In view of this, the present invention provides a new voice control method, device, electronic device and storage medium to overcome the defects of the prior art. The following will be described in detail in conjunction with Figures 1 - 8 for specific illustration.

[0061] Figure 1 is one of the flow diagrams of the voice control method provided by the present invention. As Figure 1 shown, the execution subject of the voice control method can be a device with a microphone for collecting voice signals. This device can be an electronic device such as a mobile phone, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, etc., or a smart home device, a smart household appliance, etc. For the convenience of more clearly and concisely elaborating the specific implementation manner and the principle of the embodiment of the voice control method, in the subsequent embodiments, a household refrigerator is generally used as an example of the device to describe the embodiments, which is not regarded as a specific limitation on the protection scope of the present invention.

[0062] Specifically, the voice control method provided by the present invention, when specifically implemented, includes but is not limited to the following steps:

[0063] Step 101: Receive a voice signal issued by a user, mainly by continuously collecting all voice signals within its pickup area through a microphone.

[0064] It is possible to determine whether to use it as a voice signal for the user to perform voice control according to whether the voice energy of the collected voice signal is greater than a preset energy threshold.

[0065] Other methods can also be adopted. For example, first identify the user identity of each voice signal, and then, according to the user identity, use the voice signals issued by some users with control permissions as the voice signals for subsequent voice control in the implementation of the present invention.

[0066] Taking a household refrigerator as an example, some family members are pre-granted the control permission to perform voice control on it. The voice samples of these people are pre-collected, and the timbre features of each voice sample are extracted and saved. The above timbre features can be stored in a cloud server according to the unique device identifier of the household refrigerator.

[0067] During the actual operation, the timbre features of each received voice signal are sequentially compared with the timbre features in the cloud server. If any timbre feature can be matched, it means that the user who issued the received voice signal is a user with control permissions, and the voice signal of this user can be used for subsequent voice control.

[0068] Of course, other methods can also be combined to identify whether the user grants control rights, such as combining with visual recognition technology, etc., which will not be elaborated one by one here.

[0069] Step 102: In response to the voice signal, match the target voice recognition model from the voice recognition model library according to the device state of the target device.

[0070] After the target device receives the voice signal input by the user, it will further confirm the current device state. Among them, the device state is some states that are closely related to device control, and can be determined according to the sensing data of the sensors set on the target device.

[0071] Taking a household refrigerator as an example, the device state can be the device door opening state, device door closing state, freezer temperature, freezer sterilization mode, current screen display mode, distance between the user and the household refrigerator, indoor temperature and humidity, etc., which can be set according to the scenario requirements of voice control.

[0072] The biggest difference from the prior art is that the voice control method provided by the present invention pre-designs and trains a voice recognition model for some device states of the target device respectively.

[0073] Still taking a household refrigerator as an example, a closing command recognition model is set for its device door opening state. This closing command recognition model can only recognize a limited number of voice signals related to closing control of the household refrigerator. Through this setting, in the voice control method provided by the present invention, recognizing the device door opening state serves as a wake-up word. That is, after recognizing that the household refrigerator is in the device door opening state, the closing command recognition model is directly retrieved.

[0074] For another example, by configuring a corresponding voice recognition model for the device state of the freezer sterilization mode, that is, when it is determined that the freezer sterilization mode of the household refrigerator is the ultraviolet sterilization mode, the corresponding target voice recognition model is automatically retrieved. This target voice recognition model can only recognize voice signals related to turning off or switching the ultraviolet sterilization mode, and cannot effectively recognize the voice signals that the above-mentioned closing command recognition model can recognize.

[0075] The voice recognition model library is constructed by all the above-mentioned pre-designed and trained voice recognition models, and the voice recognition model library can be stored in the target device. Of course, considering the requirements of hardware computing power, in the case of adopting this storage method, it can be selectively designed only for several scenarios with high voice control demand frequencies, and the voice recognition model library is constructed by the voice recognition models corresponding to the associated device states.

[0076] Of course, with the rapid development of current communication technologies, the data transmission speed is getting faster and faster. Then, the above-mentioned speech recognition model library can also be stored in a cloud server, so that the cloud server can match a target speech recognition model from the speech recognition model library according to the device state of the target device, and use the target speech recognition model to implement the recognition of speech signals. At this time, it is possible to pre-design and train the corresponding speech recognition models for the device states in more speech control demand scenarios.

[0077] It should be noted that in the present invention, after receiving the speech signal sent by the user, it can be used as a trigger condition to collect the device state of the target device, and then match the target speech recognition model from the speech recognition model library.

[0078] As another alternative implementation, it is also possible to obtain the device state in real time and pre-match the target speech recognition model from the speech recognition model library according to the device state. In this way, once the speech signal sent by the user is received, the speech signal can be quickly input into the target speech recognition model for recognition.

[0079] Step 103: Input the speech signal into the target speech recognition model.

[0080] By inputting the collected speech signal into the target speech recognition model matched by the current device state in real time, the recognition result of the target speech recognition model for the speech signal can be obtained.

[0081] Among them, the target speech recognition model can be obtained by training basic models such as a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), and a Hidden Markov Model (HMM).

[0082] It should be noted that in the speech control method provided by the present invention, the target speech recognition model used for speech signal recognition generally adopts a single speech recognition model, that is, a model that can only recognize a single or a limited number of speech signals. In this way, the architecture of the model will be simpler and more lightweight.

[0083] In addition, compared with ordinary speech recognition models that usually require large-scale training data to cover various changes in speech samples and more comprehensive language background knowledge, the training of the single speech recognition model adopted in the present invention mainly focuses on a single speech signal. This means that the training data set will be relatively small, and it can focus more on the specific features and patterns of the speech signal, with a more definite recognition object. Therefore, its recognition accuracy will be better, and it can effectively make up for the risk of misoperation in voice control without a wake-up word.

[0084] Step 104: If it is determined that the target speech recognition model can recognize the speech signal, then execute the control instruction corresponding to the control command word output by the target speech recognition model.

[0085] After the speech signal is input into the target speech recognition model, if it can be recognized by the target speech recognition model, it means that the speech signal is the set speech signal that can control the target device in the current device state. At this time, the corresponding control instruction can be generated according to the control command word output by the target speech recognition model, and the control instruction is sent to the controller of the target device for relevant control.

[0086] Of course, if the target speech recognition model cannot recognize and output a control command word after the speech signal is input into the target speech recognition model, it means that the speech signal does not meet the control requirements in the current device state, and no control instruction is executed at this time.

[0087] The voice control method provided by the present invention does not require the user to wake up, effectively utilizes the scenario condition of the device state, can match the corresponding speech recognition model according to the different real-time states of the target device, and the speech recognition model will specifically recognize the speech signal, which not only simplifies the human-computer interaction process, improves the accuracy of speech recognition, but also enhances the user experience.

[0088] As an alternative embodiment, in the present invention, multiple device states of the target device are selected, and a speech recognition model is designed and trained for each device state and stored in the speech recognition model library. The above speech recognition model generally adopts a single speech recognition model, that is, a model that can only recognize a single or a limited number of speech signals.

[0089] In this way, it is possible to respond to a single or a limited number of speech signals according to different device states of the target device, that is, in a device state, only the preset speech signal can be responded to achieve voice control.

[0090] Figure 2 It is the second flow diagram of the voice control method provided by the present invention. As an alternative embodiment, as Figure 2As shown, the device status includes the device door open status and the device door closed status;

[0091] Responding to the voice signal, according to the device status of the target device, matching the target voice recognition model from the voice recognition model library includes:

[0092] When determining that the device status is the device door open status, matching the door closing command recognition model from the voice recognition model library as the target voice recognition model;

[0093] When determining that the device status is the device door closed status, matching the door opening command recognition model from the voice recognition model library as the target voice recognition model;

[0094] The command words that can be recognized by the door closing command recognition model are matched from the first command word library according to the signal characteristics of the voice signal;

[0095] The command words that can be recognized by the door opening command recognition model are matched from the second command word library according to the signal characteristics of the voice signal;

[0096] The first command word library is a set of command words that are preset to be recognized by the door closing command recognition model, and the second command word library is a set of command words that are preset to be recognized by the door opening command recognition model.

[0097] Taking the target device as a household refrigerator as an example, the entire voice control process is implemented as follows:

[0098] First, collect and extract the voice signal output by the user through the microphone, and at the same time obtain the door opening and closing status of the household refrigerator according to the threshold sensor. Among them, the door opening and closing status of the household refrigerator is divided into the device door open status and the device door closed status.

[0099] Furthermore, match the corresponding voice recognition model from the voice recognition model library according to the door opening and closing status of the household refrigerator. When the household refrigerator is in the device door open status, the retrieved door closing command recognition model is used as the target voice recognition model; when the household refrigerator is in the device door closed status, the retrieved door opening command recognition model is used as the target voice recognition model.

[0100] Among them, one or more command words that can be recognized by the door closing command recognition model are preset, and these command words can be assembled into a word library, called the first command word library.

[0101] After retrieving the door closing command recognition model according to the door opening state of the device, the voice signal is input into the door closing command recognition model, and the command words in the first command word library are matched according to the signal characteristics of the voice signal by the door closing command recognition model. If there is a command word in the first command word library whose matching degree with the signal characteristics exceeds the preset matching threshold, then the command word is used as the control command word, and a corresponding control instruction is generated and sent to the controller of the target device to drive it to perform relevant actions.

[0102] Correspondingly, after retrieving the door opening command recognition model according to the door opening state of the device, the voice signal is input into the door opening command recognition model, and the signal characteristics extracted by the door opening command recognition model are matched with the command words in the second command word library. If there is a command word in the second command word library that matches the signal characteristics, then the command word is used as the control command word, and a corresponding control instruction is generated and sent to the controller of the target device to drive it to perform relevant actions.

[0103] The voice control method provided by the present invention pre-configures an independent voice recognition model for some specific device states of the target device respectively. During the actual operation process, it can retrieve the corresponding voice recognition model in real time according to the different device states, and is used to recognize a limited number of voice signals that each voice recognition model can recognize. For the voice signals that the voice recognition model cannot recognize, incorrect operations will not be performed, realizing voice operation without a wake-up word, and the user interaction is more convenient.

[0104] Moreover, by adopting the voice control method provided by the present invention, some sensing data of the device itself is fully utilized, and on the basis of maximizing the saving of hardware computing power, the operation of false wake-up is effectively reduced.

[0105] Based on the content of the above embodiments, as an optional embodiment, only door closing control command words are included in the first command word library, and only door opening control command words are included in the second command word library;

[0106] The control instruction corresponding to the door closing control command word is to control the target device to close the door, and the control instruction corresponding to the door opening control command word is to control the target device to open the door.

[0107] In this embodiment, the first command word library and the second command word library are defined. In the two sets of command words that can be recognized by the door closing command recognition model and the door opening command recognition model, there is only one control command word respectively, that is, the door closing command control word and the door opening command control word. That is to say, it is defined that both the door closing command recognition model and the door opening command recognition model are single voice recognition models.

[0108] The purpose of such a setting is that each device state corresponds to a different speech recognition model, and each speech recognition model can only recognize one control command word, which can minimize the probability of accidental wake-up to the greatest extent and ensure the safety of the target device and the user.

[0109] Figure 3 It is the third flow schematic diagram of the voice control method provided by the present invention. As Figure 3 shown, taking the target device as a household refrigerator as an example, during the actual operation, if a voice signal sent by the user is received and extracted, the door opening and closing state of the household refrigerator is obtained synchronously.

[0110] If it is determined that the household refrigerator is in the device closed state, the door opening command recognition model is retrieved from the speech recognition model library, and the voice signal is input into the door opening command recognition model for recognition. Then it is judged whether the command word recognized by the door opening command recognition model is the preset door opening control command word.

[0111] If so, a corresponding door opening control instruction is generated according to the door opening control command word and sent to the controller; if not, no control instruction is executed. After the controller executes the door opening control instruction, the door of the household refrigerator is automatically opened.

[0112] Correspondingly, if it is determined that the household refrigerator is in the device open state, the door closing command recognition model is retrieved from the speech recognition model library, and it is obtained whether the command word recognized by the door closing command recognition model for recognizing the voice signal is the preset door closing control command word.

[0113] If so, a corresponding door closing control instruction is generated according to the door closing control command word and sent to the controller; if not, no control instruction is executed. After the controller executes the door closing control instruction, the door of the household refrigerator is automatically closed.

[0114] For the voice control method provided by the present invention, by selecting a single speech recognition model as the speech recognition model corresponding to each device state, on the one hand, the single speech recognition model has a simpler and more lightweight architecture compared to the ordinary speech recognition model, and it can be applied to some terminal devices with insufficient computing power, which can reduce the selling price of these terminal devices.

[0115] On the other hand, compared with the ordinary speech recognition model which usually requires a large amount of training data to cover various changes in speech samples and more comprehensive language background knowledge, the training of the single speech recognition model adopted by the present invention mainly focuses on a single voice signal. This means that the training data set will be relatively small, and it can be more focused on the specific features and patterns of this voice signal, with a more definite recognition object. Therefore, its recognition accuracy will be better and the response speed will be faster, which can effectively make up for the risk of misoperation in voice control without a wake-up word.

[0116] Figure 4 It is the fourth flow schematic diagram of the voice control method provided by the present invention. As an alternative embodiment, the device state further includes a human-machine distance state, and the human-machine distance state is used to represent the distance range between the user and the target device.

[0117] Combined with Figure 4 As shown, the above-mentioned response to the voice signal to match the target voice recognition model from the voice recognition model library according to the device state of the target device mainly includes, but is not limited to, the following steps:

[0118] According to the distance range determined by the human-machine distance state, match the control command recognition model from the voice recognition model library as the target voice recognition model.

[0119] Among them, different distance ranges correspond to different control command recognition models, and the control command recognition model can only recognize a preset control command word.

[0120] For example, according to the distance between the user and the device, four distance ranges A-D are divided and labeled in sequence from near to far, and the interval span between each distance range can be the same or different. For example, with 1 meter as the interval span, the four distance ranges A-D are set as [0,1), [1,2), [2,3), and [3,∞) in sequence.

[0121] In addition, a different voice recognition model is designed and trained for each distance range in sequence as the control command recognition model.

[0122] Taking the target device as a household refrigerator as an example, the control command recognition model set for the A distance range is mainly used to detect whether the voice signal contains a command word for controlling the stop of ultraviolet sterilization; the control command recognition model set for the B distance range is mainly used to detect whether the voice signal contains a command word for controlling the screen display to light up; the control command recognition model set for the C distance range is mainly used to detect whether the voice signal contains a command word for controlling the screen display to turn off; the control command recognition model set for the D distance range is mainly used to detect whether the voice signal contains a command word for controlling the operation of ultraviolet sterilization, etc.

[0123] Through the above settings, the distance range to which the current distance between the user and the target device belongs can be determined according to the human-machine distance state collected by the sensor.

[0124] Combined with the above embodiment, assuming that the distance range is the A range, then the control command recognition model extracted from the voice recognition model library is mainly used to detect whether the voice signal contains a preset command word for controlling the stop of ultraviolet sterilization.

[0125] If it is included, corresponding control instructions can be generated according to the command word for controlling the stop of ultraviolet sterilization and sent to the controller to perform the corresponding operation of stopping ultraviolet sterilization. If it is not included, the voice signal of the user will not be responded to and no control instructions will be executed.

[0126] It should be noted that the logic for voice control of the device based on different human-machine distance states of the present invention can be combined with any of the above embodiments according to other device states, such as the door opening / closing state, etc., to comprehensively improve the accuracy of voice control, which will not be elaborated here one by one.

[0127] The voice control method provided by the present invention can call different voice recognition models corresponding to the device state according to different device states. In particular, using the human-machine distance state as a device state in voice control can improve the comprehensiveness and robustness of the control method.

[0128] Figure 5 It is the fifth flowchart of the voice control method provided by the present invention. As an alternative embodiment, the above device state further includes a human-machine interaction state, and the human-machine interaction state is determined according to the interaction behavior between the user and the target device;

[0129] Responding to the voice signal and matching a target voice recognition model from the voice recognition model library according to the device state of the target device includes:

[0130] According to the interaction behavior determined by the human-machine interaction state, a control command recognition model is matched from the voice recognition model library as the target voice recognition model.

[0131] Among them, different interaction behaviors correspond to different control command recognition models, and the control command recognition model can only recognize a preset control command word.

[0132] Taking the target device as a household refrigerator as an example, assuming that the human-machine interaction state mainly includes the user looking at or pointing to the refrigerator display screen, the user looking at or pointing to the refrigerator door handle, etc., the detection of the above user interaction state can be realized by means of cameras, sensors, etc. arranged on the target device, and by using computer vision and machine learning technologies, for example:

[0133] Capturing images related to the user's actions and postures through a camera or other sensors.

[0134] Using computer vision technology to process the above images to obtain the target position features of the user's gaze or arm pointing.

[0135] Using machine learning algorithms, a target recognition model is established, which can recognize objects such as the display screen of a household refrigerator and the refrigerator door handle.

[0136] According to the target recognition result, the human-computer interaction state of the user can be distinguished. For example, if the user is looking at or operating the refrigerator display screen, it can be inferred that the user currently wants to perform a human-computer interaction behavior with the display screen. At this time, the voice control model related to lighting up the display screen can be retrieved from the voice recognition model as the control command recognition model.

[0137] If it is included, the corresponding control instruction can be generated according to the command word for lighting up the display screen and sent to the controller to execute the corresponding operation of lighting up the display screen. If it is not included, then the user's voice signal is not responded to and no control instruction is executed.

[0138] It should be noted that the logic for determining voice control of the device based on different human-computer interaction behaviors provided by the present invention can be combined with any of the above embodiments according to other device states, such as the door opening state, etc., and the human-computer distance state, etc., to comprehensively improve the accuracy of voice control, which will not be elaborated here one by one.

[0139] Figure 6 is the sixth flow diagram of the voice control method provided by the present invention. As an alternative embodiment, as Figure 6 shown, before receiving the voice signal sent by the user, that is, before the target device officially runs, it further includes:

[0140] When it is determined that the device state of the target device meets a preset condition, the target device is switched to the voice control waiting mode;

[0141] A signal acquisition prompt is output to the user, which is used to prompt the user that the target device is currently controllable by voice signals.

[0142] Among them, the above signal acquisition prompt mainly includes but is not limited to at least one of voice prompt, light prompt, and display screen prompt.

[0143] In addition, the device state meeting the preset condition mainly includes but is not limited to at least one of the following: the device state is the device door opening state, the distance between the user and the target device is less than a preset distance threshold, a preset interaction behavior occurs between the user and the target device, and the acceleration of the distance reduction between the user and the target device is greater than a preset acceleration threshold.

[0144] Taking the target device as a household refrigerator as an example, assume that the device state meeting the preset condition is that the distance between the user and the target device is less than a preset distance threshold, and assume that the preset distance is 3 meters.

[0145] In this scenario, when the distance between the user and the household refrigerator detected by the distance sensor set on the household refrigerator is less than 3 meters, its control state is switched to the voice control waiting mode. In the voice control waiting mode, the voice signal of the user can be received, and voice control can be achieved according to any of the methods provided in the above embodiments.

[0146] After detecting that the user leaves this area, that is, the distance between the user and it is greater than 3 meters, the voice control waiting mode is turned off. At this time, the voice signal of the user will not be actively collected, and even less will the relevant operation instructions be responded to according to the voice signal.

[0147] With this setting, on the one hand, false wake-up can be effectively avoided. On the other hand, since the household refrigerator will not always be in the voice control waiting mode to collect the voice signal of the user, it can save computing power and energy consumption to a certain extent and extend the service life of the device.

[0148] In addition, after switching the household refrigerator to the voice control waiting mode, the user can be reminded that they have entered this mode by outputting a signal collection prompt, and the household refrigerator can be controlled by voice control.

[0149] For example, after switching the household refrigerator to the voice control waiting mode, the user can be informed by means of light indication. At the same time, different colors of lights can be used for different reminders according to the distance between the user and the household refrigerator. At the same time, a pre-recorded voice can be broadcast to the user by means of voice prompt while the light reminder is given, such as "Master, I'm at your service. Please give instructions."

[0150] The voice control method provided by the present invention uses the judgment of preset conditions as the wake-up condition for voice control, and combines the corresponding signal collection prompt to enhance the interaction with the user, not only reducing the probability of false wake-up, but also enhancing the user experience to a certain extent.

[0151] Based on the content of the above embodiments, as an alternative embodiment, if it is determined that the target voice recognition model cannot recognize the voice signal, the user is prompted to input a wake-up voice signal;

[0152] After receiving the wake-up voice signal, a multi-semantic recognition model is matched from the voice recognition model library, and the user is prompted to input a control voice signal;

[0153] The control voice signal is input into the multi-semantic recognition model, and the control instruction corresponding to the control command word output by the multi-semantic recognition model is executed.

[0154] In the above embodiments, the relevant situations where the target speech recognition model matched according to the device state can recognize speech signals are described. It is not excluded that there will still be situations where the target speech recognition model cannot recognize speech signals sometimes.

[0155] In order to enable the user to control the target device through voice control, the present invention can also continue with subsequent control in the form of asking the user questions or interacting with the user when the target speech recognition model fails to recognize the speech signal.

[0156] For example, output the pre-recorded voice "I have not verified your authority. Please enter our agreed wake-up word" through the speaker. After receiving the wake-up speech signal spoken by the user, call the multi-semantic recognition model and prompt the user to enter a command, such as "Okay, master. What can I do for you?".

[0157] After receiving the control speech signal re-entered by the user, use the above multi-semantic recognition model to perform semantic recognition on the user's wake-up speech signal at this time, and execute relevant actions according to the recognized control command words.

[0158] It should be noted that the above supplement for the situation where the target speech recognition model fails to recognize the speech signal can be implemented by using some existing technologies, and the present invention does not make specific limitations or excessive elaborations.

[0159] The voice control method provided by the present invention designs solutions respectively for whether the target speech recognition model can correctly recognize the speech signal, which can fully meet the user's voice control methods for the target device in various scenarios and can satisfy the user's usage experience.

[0160] Figure 7 It is a schematic structural diagram of the voice control device provided by the present invention, as Figure 7 shown, mainly including a voice acquisition unit 71, a model calling unit 72, a voice processing unit 73, and a control execution unit 74, where:

[0161] The voice acquisition unit 71 is mainly used to receive the voice signal sent by the user;

[0162] The model calling unit 72 is mainly used to respond to the voice signal and match the target speech recognition model from the speech recognition model library according to the device state of the target device;

[0163] The voice processing unit 73 is mainly used to input the voice signal into the target speech recognition model;

[0164] The control execution unit 74 is mainly used to execute the control instruction corresponding to the control command word output by the target speech recognition model after determining that the target speech recognition model can recognize the speech signal.

[0165] It should be noted that the speech control device provided in the embodiments of the present invention can execute the speech control method described in any of the above embodiments during specific operation, and this embodiment will not be elaborated herein.

[0166] The speech control device provided by the present invention does not require the user to wake up, effectively utilizes the scenario condition of the device state, can match the corresponding speech recognition model according to the real-time state of the target device, and this speech recognition model will specifically recognize the speech signal, which not only simplifies the process of human-computer interaction, improves the accuracy of speech recognition, but also enhances the user experience.

[0167] In the above embodiments, the target device is taken as a household refrigerator as an example for illustration. In fact, in the speech control method and device provided by the present invention, the target device can also be a refrigerator for other purposes. Of course, the target device is not limited to being a refrigerator, and can also be devices such as a television, a water heater, an air conditioner, and so on.

[0168] Figure 8 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 8 shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the speech control method, and this method includes: receiving a speech signal sent by a user; in response to the speech signal, matching a target speech recognition model from a speech recognition model library according to the device state of the target device; inputting the speech signal into the target speech recognition model; if it is determined that the target speech recognition model can recognize the speech signal, then execute the control instruction corresponding to the control command word output by the target speech recognition model.

[0169] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0170] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the voice control method provided in the above-mentioned various embodiments. The method includes: receiving a voice signal sent by a user; in response to the voice signal, matching a target voice recognition model from a voice recognition model library according to the device state of the target device; inputting the voice signal into the target voice recognition model; if it is determined that the target voice recognition model can recognize the voice signal, then execute the control instruction corresponding to the control command word output by the target voice recognition model.

[0171] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the voice control method provided in the above-mentioned various embodiments. The method includes: receiving a voice signal sent by a user; in response to the voice signal, matching a target voice recognition model from a voice recognition model library according to the device state of the target device; inputting the voice signal into the target voice recognition model; if it is determined that the target voice recognition model can recognize the voice signal, then execute the control instruction corresponding to the control command word output by the target voice recognition model.

[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voice control method, characterized in that, including: Receiving a voice signal sent by a user; In response to the voice signal, matching a target voice recognition model from a voice recognition model library according to the device state of the target device; Inputting the voice signal into the target voice recognition model; If it is determined that the target voice recognition model can recognize the voice signal, executing a control instruction corresponding to a control command word output by the target voice recognition model.

2. The voice control method according to claim 1, wherein There are multiple types of the device state; For different device states, the target voice recognition models matched from the voice recognition model library are different.

3. The voice control method according to claim 2, wherein The device state includes a device door opening state and a device door closing state; The step of in response to the voice signal, matching a target voice recognition model from a voice recognition model library according to the device state of the target device includes: When it is determined that the device state is the device door opening state, matching a door closing command recognition model from the voice recognition model library as the target voice recognition model; When it is determined that the device state is the device door closing state, matching a door opening command recognition model from the voice recognition model library as the target voice recognition model; Command words that can be recognized by the door closing command recognition model are matched from a first command word library according to the signal characteristics of the voice signal; Command words that can be recognized by the door opening command recognition model are matched from a second command word library according to the signal characteristics of the voice signal; The first command word library is a set of command words preset to be recognized by the door closing command recognition model, and the second command word library is a set of command words preset to be recognized by the door opening command recognition model.

4. The voice control method according to claim 3, wherein The first command word library only includes door closing control command words, and the second command word library only includes door opening control command words; The control instruction corresponding to the door closing control command word is to control the target device to close the door, and the control instruction corresponding to the door opening control command word is to control the target device to open the door.

5. The voice control method according to claim 3, wherein The device state further includes a human-machine distance state, and the human-machine distance state is used to represent the distance interval in which the distance between the user and the target device is located; The step of in response to the voice signal, matching a target voice recognition model from a voice recognition model library according to the device state of the target device includes: According to the distance interval determined by the human-machine distance state, matching a control command recognition model from the voice recognition model library as the target voice recognition model; Different distance intervals correspond to different control command recognition models, and the control command recognition model can only recognize a preset control command word.

6. The voice control method according to claim 3, wherein The device state further includes a human-machine interaction state, and the human-machine interaction state is determined according to the interaction behavior between the user and the target device; The step of in response to the voice signal, matching a target voice recognition model from a voice recognition model library according to the device state of the target device includes: According to the interaction behavior determined by the human-machine interaction state, matching a control command recognition model from the voice recognition model library as the target voice recognition model; Different interaction behaviors correspond to different control command recognition models, and each control command recognition model can only recognize a preset control command word.

7. The voice control method according to any one of claims 1-6, characterized in that Before receiving the voice signal sent by the user, it further includes: When it is determined that the device state of the target device meets a preset condition, switching the target device to the voice control waiting mode; Outputting a signal collection prompt to the user to prompt that the target device is currently acceptable for voice signal control; The signal collection prompt includes at least one of voice prompt, light prompt, and screen display prompt; The device state meeting the preset condition includes at least one of the device state being the device door open state, the distance between the user and the target device being less than a preset distance threshold, a preset interaction behavior occurring between the user and the target device, and the acceleration of the distance reduction between the user and the target device being greater than a preset acceleration threshold.

8. The voice control method according to any one of claims 1-6, characterized in that, If it is determined that the target voice recognition model cannot recognize the voice signal, prompting the user to input a wake-up voice signal; After receiving the wake-up voice signal, matching a multi-semantic recognition model from the voice recognition model library and prompting the user to input a control voice signal; Inputting the control voice signal into the multi-semantic recognition model and executing the control instruction corresponding to the control command word output by the multi-semantic recognition model.

9. A voice control device, characterized in that, It includes: A voice collection unit for receiving the voice signal sent by the user; A model calling unit for responding to the voice signal and matching a target voice recognition model from the voice recognition model library according to the device state of the target device; A voice processing unit for inputting the voice signal into the target voice recognition model; A control execution unit for executing the control instruction corresponding to the control command word output by the target voice recognition model after determining that the target voice recognition model can recognize the voice signal.

10. The voice control device according to claim 9, wherein The target device is a refrigerator.

11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the voice control method according to any one of claims 1 to 8.

12. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the voice control method according to any one of claims 1 to 8.