Voice control method, device and system, electronic equipment and storage medium
By combining voice monitoring, voiceprint and user status judgment, control instructions are generated and executed, the problem of misoperation in voice control is solved and the user experience is improved.
Patent Information
- Application Number
- CN202311850525.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
When controlling voice devices, ambient noise and daily conversations are prone to causing misoperation of the device, especially in voice control without wake-up words, the probability of misoperation is high and affects the user experience.
Only by obtaining voice data through voice monitoring, determining whether voice data includes control-related information, and complying with preset voiceprint standards and user status judgment standards, control instructions are generated and executed.
It significantly reduces the misoperation of equipment caused by environmental noise and daily conversations, improves the user experience, especially in the voice control of wake-up words, and reduces the probability of misoperation.
Smart Images

Figure CN120236567A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of device control, and particularly to a voice control method, apparatus, system, electronic device, and storage medium. Background Art
[0002] With the development of science and technology and urban construction, a large number of devices need to be installed in places such as office buildings, apartments, schools, shopping malls, etc., such as air conditioner indoor units. With the progress of intelligent technology, more and more devices can be controlled by voice.
[0003] It should be noted that the above introduction to the technical background is only for the convenience of clearly and completely explaining the technical solutions of this application and facilitating the understanding of those skilled in the art. It cannot be considered that the above technical solutions are well-known to those skilled in the art just because these solutions are described in the background art part of this application. Summary of the Invention
[0004] The inventors found that when controlling devices by voice, environmental noises (such as sounds generated by televisions, speakers, etc.) and daily conversations (such as muttering sounds when getting up at night, talking in sleep) are likely to cause misoperations of the devices, affecting the user experience. Especially in voice control without wake-up words, the probability of misoperation increases significantly.
[0005] To solve at least one of the above problems, embodiments of this application provide a voice control method, apparatus, system, electronic device, and storage medium.
[0006] According to a first aspect of the embodiments of this application, a voice control method is provided. The method includes: obtaining voice data through voice monitoring; generating a control instruction according to the voice data when the voice data includes control-related information and meets a preset voiceprint standard and a user state judgment standard; and executing the control instruction.
[0007] According to a second aspect of the embodiments of this application, a voice control system is provided. The system includes: a voice interaction device that obtains voice data through voice monitoring; a control device that generates a control instruction according to the voice data when the voice data includes control-related information and meets a preset voiceprint standard and a user state judgment standard; and an execution device that executes the control instruction.
[0008] According to a third aspect of the embodiments of this application, a voice control apparatus is provided. The apparatus includes: an obtaining module that obtains voice data through voice monitoring; and a processing module that generates a control instruction according to the voice data when the voice data includes control-related information and meets a preset voiceprint standard and a user state judgment standard.
[0009] According to a fourth aspect of the embodiments of the present application, an electronic device is provided, and the electronic device includes: a memory that stores a computer program; and a processor that implements the method described in the first aspect of the embodiments of the present application when executing the computer program.
[0010] According to a fifth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program implements the method described in the first aspect of the embodiments of the present application when executed by a processor.
[0011] One of the beneficial effects of the embodiments of the present application is that: voice data is obtained through voice monitoring, and when the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, a control instruction is generated according to the voice data and the control instruction is executed. In this way, the misoperation of the device caused by environmental noise or daily conversations can be reduced, and the user experience can be improved. For example, in voice control without a wake-up word, the probability of misoperation can be significantly reduced.
[0012] Referring to the following description and drawings, specific embodiments of the present application are disclosed in detail, indicating the ways in which the principles of the present application can be adopted. It should be understood that the embodiments of the present application are not limited in scope thereby. Within the spirit and terms of the appended claims, the embodiments of the present application include many changes, modifications, and equivalents.
[0013] The characteristic information described and illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, combined with the characteristic information in other embodiments, or replace the characteristic information in other embodiments.
[0014] It should be emphasized that the term "comprising / including" when used herein refers to the presence of characteristic information, the whole, steps, or components, but does not exclude the presence or addition of one or more other characteristic information, the whole, steps, or components. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Many aspects of the present application can be better understood with reference to the following drawings. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged or reduced. The elements and characteristic information described in one drawing or one embodiment of the present application can be combined with the elements and characteristic information shown in one or more other drawings or embodiments. In addition, in the drawings, like reference numerals represent corresponding components in several drawings and can be used to indicate corresponding components used in more than one embodiment.
[0016] In the drawings:
[0017] Figure 1 It is a schematic diagram of the voice control method according to an embodiment of the present application;
[0018] Figure 2 It is a schematic diagram of the method for generating a control instruction according to voice data in an embodiment of the present application;
[0019] Figure 3 It is another schematic diagram of the voice control method according to an embodiment of the present application;
[0020] Figure 4 It is another schematic diagram of the voice control method according to an embodiment of the present application;
[0021] Figure 5 It is a schematic diagram of the voice control system according to an embodiment of the present application;
[0022] Figure 6 It is a schematic diagram of the voice control device according to an embodiment of the present application;
[0023] Figure 7 It is a schematic diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0024] The preferred implementation manners of the present application will be described below with reference to the accompanying drawings.
[0025] Embodiment 1
[0026] Embodiment 1 of the present application provides a voice control method.
[0027] Figure 1 It is a schematic diagram of the voice control method according to an embodiment of the present application. As Figure 1 shown, the method includes:
[0028] Step 101: Obtain voice data through voice listening;
[0029] Step 102: Generate a control instruction according to the voice data when the voice data includes information related to control and meets the preset voiceprint standard and user state judgment standard; and
[0030] Step 103: Execute the control instruction.
[0031] According to the above embodiments, voice data is obtained through voice monitoring. When the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, a control instruction is generated based on the voice data and the control instruction is executed. In this way, misoperations of the device caused by environmental noise or daily conversations can be reduced, and the user experience can be improved. For example, in voice control with a wake-up word or without a wake-up word, by generating or executing a control instruction according to the voiceprint and user status corresponding to the voice data, the probability of misoperation can be significantly reduced.
[0032] In some embodiments, in step 101, voice monitoring can be performed through a voice interaction device to obtain voice data. The voice interaction device can be various devices capable of performing voice monitoring. For example, it can be at least one of the following: a central controller, a line controller, a remote controller, a user terminal, and a smart speaker, etc.
[0033] In some embodiments, the voice interaction device can monitor voices emitted from various sound sources. For example, the monitored voice can come from a user or a device capable of emitting a voice, such as a TV, a speaker, etc.
[0034] In some embodiments, the voice data obtained through voice monitoring can include various types of voice data. For example, instruction-type voice data or non-instruction-type voice data. Among them, the instruction-type voice data can refer to voice data that includes a voice command, such as "increase the temperature", "raise the temperature", "raise the temperature to 28 °C", etc. The non-instruction-type voice data can be conversational voice data, etc., such as "too cold", "too bright", etc.
[0035] In some embodiments, the voice data obtained through voice monitoring can include control-related information. This information can be explicit control information, such as "turn on the air conditioner", "increase the temperature", etc.; the present application is not limited thereto, and this information can also be control information representing the user's intention, such as "too hot", etc.
[0036] In some embodiments, voice monitoring can be started at a preset period. Voice monitoring is performed within the first time of the preset period, and voice monitoring is stopped within other times of the preset period. For example, voice monitoring is started once every 5 seconds, and each monitoring lasts for 2 seconds. By performing voice monitoring periodically, the power consumption of the voice interaction device can be reduced, and the privacy of the user can be protected. However, the present application is not limited thereto, and voice monitoring can also be performed in real time. Thus, voice data can be obtained reliably and completely, avoiding missing some voice data.
[0037] In some embodiments, different monitoring strategies can be implemented in different time periods. For example, voice monitoring is started at a preset period in a first time period; or voice monitoring is performed in real time in a second time period. By performing periodic or real-time voice monitoring at different times, the power consumption of the voice interaction device and the reliability and integrity of the voice data can be taken into account, further improving the user experience.
[0038] In some embodiments, the first time period or the second time period may be set in various ways, for example, the first time period or the second time period may be set according to at least one of time information, user usage information, or user input information.
[0039] For example, the first time period or the second time period is set by using time information, and the daytime time period (for example, 6:00 to 22:00) can be set as the second time period, and other times can be set as the first time period.
[0040] By setting the first time period or the second time period using time information, the time period can be set in a simple manner, which is easy to implement.
[0041] For another example, the first time period or the second time period is set based on the user's usage information, and the time period can be set based on the user's usage frequency. For example, the time period when the user's usage frequency is higher is set as the second time period (for example, the high-frequency usage stage), and the time period when the user's usage frequency is lower is set as the first time period (for example, the non-high-frequency usage stage).
[0042] The user's usage frequency can be represented by the frequency of generating or executing control instructions, or by the frequency of acquiring voice data. For example, the frequency of generating or executing control instructions in the first time period is less than the frequency of generating or executing control instructions in the second time period; or, the frequency of acquiring voice data in the first time period is less than the frequency of acquiring voice data in the second time period. The present application is not limited to this, and the user's usage frequency can also be represented by other methods.
[0043] By setting the first time period or the second time period based on the user's usage information, the time period can be set in a targeted manner according to the user's habits, thereby further improving the user experience.
[0044] For another example, the first time period or the second time period is set according to the user's input information, and the time period can be set according to the received user's input information, thereby meeting the user's personalized requirements and further improving the user experience.
[0045] In some embodiments, the first time period can be further divided into more than one interval, and different listening strategies are adopted in different intervals. For example, the first time period can include a first interval and a second interval. Within the first interval, voice listening is started at a first preset period; within the second interval, voice listening is started at a second preset period. Thus, the control of voice listening can be carried out more meticulously, the power consumption of the voice interaction device, as well as the reliability and integrity of voice data can be further taken into account, and the user experience can be further improved.
[0046] Taking the first time period as the time period when the user is sleeping (for example, 22:00 - 6:00) as an example, the first interval can be the time period when the user is in a light sleep state, for example, 22:00 - 23:00, 5:00 - 6:00; the second interval can be the time period when the user is in a deep sleep state, for example, 23:00 - 5:00.
[0047] Within the first interval and the second interval, listening can be carried out at different periods. For example, compared with the deep sleep state, the user is more likely to wake up from the light sleep state. Therefore, the possibility that the user issues a control voice within the first interval is higher, and the listening period (the first preset period) corresponding to the first interval can be shorter than the listening period (the second preset period) corresponding to the second interval. Thus, corresponding control instructions can be executed in a timely manner according to the user's voice data.
[0048] Among them, the setting method of the first interval or the second interval can be similar to the setting method of the first time period or the second time period. For example, it can be set according to at least one of time information, the user's usage information, or the user's input information.
[0049] Taking the setting according to the user's usage information as an example, the first frequency of generating or executing control instructions within the first interval can be lower than the second frequency of generating or executing control instructions within the second interval. Correspondingly, the first preset period can be greater than the second preset period.
[0050] In some embodiments, in step 102, the control device can determine whether the voice data includes control-related information, whether it meets the preset voiceprint standard, and the user status judgment standard, and perform corresponding actions according to the judgment results. For example, in the case where the voice data includes control-related information and meets the preset voiceprint standard and the user status judgment standard, a control instruction is generated according to the voice data; otherwise, reconfirmation is carried out or no control instruction is generated.
[0051] Among them, the control device can be, for example, at least one of the following: a server, a gateway device, a centralized controller, a line controller, a remote controller, and a control board, etc.
[0052] In some embodiments, the preset voiceprint standard can be set according to the user's voiceprint characteristics. Thus, user identity verification can be performed according to the preset voiceprint standard. For example, when the voice data conforms to the preset voiceprint standard, the voice data is the voice data of the user. Thus, environmental noise can be distinguished (for example, the sounds emitted by devices such as televisions and stereos, the voices of other people whose voiceprint characteristics are not recorded, etc.). Only when the acquired voice data belongs to a user whose voiceprint characteristics have been recorded, can the corresponding control instruction be generated or executed. Thus, misoperation of the device caused by environmental noise and the like can be avoided.
[0053] In some embodiments, the voiceprint characteristics used to set the preset voiceprint standard can be actively input by the user. Or, they can also be collected by the voice interaction device in daily life. For example, within a certain period of time, if the cumulative duration of the voice of the same user monitored by the voice interaction device is greater than the preset duration, or the cumulative number of times the voice of the same user is monitored is greater than the preset number of times, etc., the preset voiceprint standard can be set according to the voiceprint characteristics of the user.
[0054] In some embodiments, the preset voiceprint standard can be set according to the voiceprint characteristics of one or more than one user. For example, in the scenario of a smart home, the preset voiceprint standard can be set according to the respective voiceprint characteristics of multiple main family members. Thus, the voice data of each of the main family members conforms to the preset voiceprint standard, and thus each of the main family members can perform voice control. The specific method of setting the preset voiceprint standard according to the voiceprint characteristics can refer to the related technology and will not be elaborated here.
[0055] In some embodiments, the user state judgment standard can include at least one of the following: speech rate standard, semantic standard, intonation standard, volume standard, or timbre standard. Through this user state judgment standard, the state of the user who emits the voice can be judged.
[0056] In some embodiments, the state of the user can include: normal state and abnormal state. Among them, the normal state includes, for example, a waking state or an emotionally stable state, etc.; the abnormal state includes, for example, a non-waking state (such as a drowsy state or a sleeping state, etc.) or an emotionally unstable state (such as an angry state or a sad state, etc.).
[0057] When the voice data conforms to the user state judgment standard, the user is in a normal state. Thus, the control instruction can be generated or executed according to the voice data emitted by the user; when the voice data does not conform to the user state judgment standard, the user is in an abnormal state. In this state, the possibility of the user performing voice control is small, and the control instruction may not be generated or executed, thus avoiding misoperation.
[0058] Taking the awake state and the non-awake state as examples, an exemplary description of the method for determining the user state is given below. It can be understood that the following content is also applicable to other states.
[0059] For example, make a judgment based on the speech rate standard. The speech rate of a user in different states is usually different. For example, the speech rate in the awake state is moderate, and the speech rate in the non-awake state is slightly faster or slower. When the speech rate corresponding to the acquired speech data is within the preset range, it can be judged that the speech data meets the speech rate standard and the user is in the awake state; otherwise, it is judged that the speech data does not meet the speech rate standard and the user is in a drowsy or sleeping state.
[0060] For example, make a judgment based on the semantic standard. The speech uttered by a user in the awake state usually has a clear meaning, while the speech uttered in a drowsy or sleeping state usually does not have a clear meaning. For example, the muttering sound when getting up at night, or the dream talk during sleep, etc. are usually not complete sentences and have no clear meaning. By performing semantic analysis on the speech data, when the speech data has a clear meaning, it can be judged that the speech data meets the semantic standard and the user is in the awake state; otherwise, it is judged that the speech data does not meet the semantic standard and the user is in a drowsy or sleeping state.
[0061] For example, make a judgment based on the intonation standard. The intonation of the speech uttered by a user in different states is different. For example, the intonation of the speech uttered in the awake state can include imperative intonation, declarative intonation, interrogative intonation, and emphatic intonation, etc. For example, the imperative intonation usually has a high pitch at the beginning of a sentence; the declarative intonation has less pitch variation and is usually a medium-high pitch throughout the sentence; the interrogative intonation usually has a high pitch at the end of a sentence; the variation of the emphatic intonation is usually reflected in a certain stressed syllable in a word or phrase. By performing intonation analysis on the speech data, when the intonation of the speech data has the characteristics of imperative intonation, declarative intonation, interrogative intonation, or emphatic intonation, it can be judged that the speech data meets the intonation standard and the user is in the awake state; otherwise, it is judged that the speech data does not meet the intonation standard and the user is in a drowsy or sleeping state.
[0062] For example, make a judgment based on the volume standard. The volume of the speech uttered by a user in the awake state is usually slightly higher than the volume of the speech uttered in a drowsy or sleeping state. When the volume of the acquired speech data is greater than the second threshold, it can be judged that the speech data meets the volume standard and the user is in the awake state; otherwise, it is judged that the speech data does not meet the volume standard and the user is in a drowsy or sleeping state.
[0063] For example, make a judgment based on the timbre standard. The timbre of the speech uttered by a user in the awake state is usually clearer or brighter than the timbre of the speech uttered in a drowsy or sleeping state. It is possible to judge whether the user is in the awake state according to the waveform of the speech data.
[0064] For example, the user state can also be determined based on more than one combination of the above criteria. For example, when the voice data meets the speech rate criterion and also meets the semantic criterion, the user state is determined to be the awake state.
[0065] In some embodiments, when the voice data includes control-related information and meets the preset voiceprint criterion (i.e., the voice data is determined to be the user's voice data according to the preset voiceprint criterion) and the user state determination criterion (i.e., the user is determined to be in the normal state according to the user state determination criterion), a control instruction is generated or executed based on the voice data. Thus, in the process of voice control, control instructions are not generated or executed for all the acquired voice data, but the conditions for generating or executing control instructions are appropriately increased, that is, when the voice data includes control-related information, the voice data is the voice data of the preset user, and the user who emits the voice data is in the normal state, a control instruction is generated or executed based on the voice data, so as to avoid misoperations caused by environmental noise or daily conversations, etc., and improve the user experience.
[0066] Especially in the voice control without a wake-up word, since it is not necessary to include a preset word (i.e., the wake-up word) in the voice, in this context, it is easier for the device to be misoperated due to environmental noise or daily conversations. By using the above control method in the voice control without a wake-up word, the user experience can be further improved.
[0067] In some embodiments, in step 102, the control instruction can be generated in various ways. Figure 2 It is a schematic diagram of the method for generating a control instruction based on voice data in an embodiment of the present application. As Figure 2 shown, generating a control instruction based on voice data may include:
[0068] 201, generating text information based on the voice data; and
[0069] 202, generating a control instruction based on the text information.
[0070] In step 201, text information can be generated based on the voice data in various ways. For example, text information is generated based on a deep learning model or a traditional speech recognition algorithm, etc. The specific method for generating text information based on voice data can refer to related technologies and will not be elaborated here.
[0071] In step 202, a control instruction can be generated based on the text information in various ways. For example, using the keywords in the text information as an index, querying or matching the control instruction corresponding to the text information in the preset control instruction database.
[0072] The preset control instruction database may include the correspondence between keywords and control instructions. If there is a keyword in the control instruction database that is the same as or has a similarity exceeding a preset threshold with the query index, the control instruction corresponding to the keyword may be used as the control instruction for the text information; otherwise, no control instruction is generated.
[0073] This application is not limited thereto, and control instructions may also be directly generated based on voice data, that is, it is not necessary to convert voice data into text information. For example, voice data may be used as the input of a deep learning model, and control instructions may be used as the output of the deep learning model. The trained deep learning model directly generates corresponding control instructions based on features such as the waveform of the voice data.
[0074] In some embodiments, step 102 may be executed within a preset time period, and different strategies for generating control instructions may be executed in other time periods. For example, step 102 may include: within a first time period, when the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, control instructions are generated based on the voice data.
[0075] The setting method of the first time period may refer to the foregoing content. For example, the first time period may be a non-high-frequency usage stage. For example, during the night time period (e.g., 22:00 - 6:00), users are usually in a sleeping state. By appropriately raising the conditions for generating or executing control instructions in the manner of step 102, it is possible to avoid misoperations of the device caused by environmental noise or daily conversations (such as muttering when getting up at night, talking in sleep), and the user experience can be further improved. As described above, by generating control instructions in the manner of step 102 in voice control without a wake-up word, the user experience can be further improved.
[0076] Figure 3 is another schematic diagram of the voice control method according to an embodiment of the present application. In some embodiments, as Figure 3 shown, the voice control method may include:
[0077] Step 301, obtaining voice data through voice listening;
[0078] Step 302, within a first time period, when the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, generating control instructions based on the voice data;
[0079] Step 303, within a second time period different from the first time period, when the voice data includes control-related information, generating control instructions based on the voice data; and
[0080] Step 304, execute the control instruction.
[0081] Among them, the content of steps 301, 303, and 304 can refer to steps 101, 102, and 103, and will not be elaborated here.
[0082] In some embodiments, the setting method of the second time period can refer to the foregoing content. For example, the second time period can be a high-frequency usage stage.
[0083] In some embodiments, in step 303, within the second time period, the above-mentioned preset voiceprint standard or user status judgment standard may not be judged, and a control instruction can be generated according to the voice data when the voice data including information related to control is obtained.
[0084] Among them, the method of generating a control instruction according to the voice data can refer to the foregoing content. For example, text information can be generated according to the voice data, and a control instruction can be generated according to the text information. Or, the conversion of text information may not be performed, and a control instruction can be directly generated according to the voice data.
[0085] In some embodiments, when the voice data includes information related to control, conforms to the preset voiceprint standard, and the voice data does not conform to the user status judgment standard, reconfirmation can be performed. In this way, the control instruction can be executed more accurately, avoiding misoperation of the device. As Figure 1 shown, the voice control method may further include:
[0086] Step 105, when the voice data includes information related to control, conforms to the preset voiceprint standard, and the voice data does not conform to the user status judgment standard, send a prompt message to the user; and
[0087] Step 106, when the confirmation information of the user for the prompt message is received, generate and execute a control instruction according to the voice data.
[0088] For example, the user status judgment standard includes a speech rate standard, and it is judged that the user is in a non-awake state according to the speech rate standard. If the voice data includes information such as "turn on the air conditioner" or "too hot", a prompt message can be sent to the user, for example, asking the user if they want to turn on the air conditioner, etc. When a positive confirmation message from the user is received, the air conditioner device can be controlled to turn on according to the voice data.
[0089] Among them, in steps 105 and 106, the prompt message can be sent to the user or the confirmation information of the user can be received by means of voice, text, picture, video, etc.
[0090] In some embodiments, in the above control method, it is possible to set whether a wake word needs to be included or not in the voice data, or the specific content of the wake word can be set. For example, as Figure 1 shown, the voice control method may further include:
[0091] Step 107, setting whether a preset word or the content of the preset word needs to be included in the voice data according to the user's input information.
[0092] Thereby, it is possible to set the control method according to the user's habits based on the user's input information, which helps to improve the user experience.
[0093] In some embodiments, when the voice data needs to include a preset word, the control method may further include: generating a control instruction according to the voice data in the case where the preset word is included in the voice data. Among them, the method of generating a control instruction according to the voice data may refer to the foregoing content.
[0094] This application is not limited thereto. When the voice data needs to include a preset word, the control method may also include: in the case where the preset word is included in the voice data, and, in the case where the voice data includes information related to control and meets the preset voiceprint standard and the user status judgment standard, generating a control instruction according to the voice data. Among them, the specific content of the preset voiceprint standard and the user status judgment standard may refer to the foregoing content.
[0095] In some embodiments, in step 103, the control instruction may be executed by an execution device. Among them, the execution device may be, for example, at least one of the following: an air handling device and a smart device, etc.
[0096] The air handling device may include at least one of an indoor unit, a air supply device, a humidity control device, a floor heating device, and a sensor.
[0097] The smart device includes at least one of a smart socket, a smart lighting device, and a smart curtain.
[0098] In some embodiments, as Figure 1 shown, the voice control method may further include:
[0099] Step 108: Store the historical record of executing the control instruction.
[0100] By storing the historical record of executing the control instruction, the historical record can be used for the user to query or analyze the user's usage habits, etc., which is convenient for improving the accuracy and rationality of subsequent voice control and helps to further improve the user experience.
[0101] In some embodiments, the historical record may be stored according to time periods.
[0102] For example, the historical record may include a first historical record for a first time period and a second historical record for a second time period. Thus, it is convenient for users to query or analyze according to the time period.
[0103] The setting methods of the first time period and the second time period may refer to the foregoing content. For example, the frequency of executing the control instruction in the first time period is less than the frequency of executing the control instruction in the second time period.
[0104] Again, for example, the historical record may include a third historical record for a third time period and a fourth historical record for a fourth time period. The third time period and the fourth time period may be longer than the first time period or the second time period. For example, the third time period may be a time period with a higher frequency of generating or executing control instructions in a year, and the fourth time period may be a time period with a lower frequency of generating or executing control instructions in a year.
[0105] Taking the execution device as an air-conditioning device as an example, in July and August, the frequency of using the air conditioner is relatively high. Taking a year as a cycle, the historical records of executing control instructions in July and August of each year and the historical records of other times can be stored separately. Thus, it is convenient to query or refer to the historical records of the same period in previous years.
[0106] The historical record can also be stored in other ways. For example, it can be classified and stored according to the controlled object (i.e., the execution device), the content of the control instruction, or the user who issues the control instruction, etc. Thus, it can meet the various query or analysis needs of users.
[0107] In some embodiments, the historical record may be stored in at least one of the following: a voice interaction device that acquires voice data, a control device that generates a control instruction, or an execution device that executes the control instruction. For example, the historical record may be stored in at least one of a central controller, a line controller, a remote controller, a user terminal, a smart speaker, a server, a gateway device, a control board, an air handling device, and a smart device. The server may be a local server or a cloud server.
[0108] In some embodiments, the historical record may be stored in at least one of the following: a local device or a cloud device. The local device includes, for example, a voice interaction device that acquires voice data, a control device that generates a control instruction, or an execution device that executes the control instruction. The cloud device includes, for example, a cloud server, etc.
[0109] Figure 4 It is another schematic diagram of the voice control method according to an embodiment of the present application. As Figure 4 shown, the method may include:
[0110] Step 401, obtain the current time, which can be the time of the voice control system, the Internet time, etc.;
[0111] Step 402, determine whether the current time is the second time period, for example, the high-frequency usage time period; if the determination is yes, execute Step 403, otherwise execute Step 406;
[0112] Step 403, perform real-time voice monitoring to obtain voice data;
[0113] Step 404, generate a control instruction according to the voice data;
[0114] Step 405, execute the control instruction;
[0115] Step 406, perform voice monitoring at a preset period to obtain voice data;
[0116] Step 407, determine whether the voice data conforms to the preset voiceprint standard; if the determination is yes, execute Step 408, otherwise return to Step 401;
[0117] Step 408, determine whether the voice data conforms to the user status judgment standard; if the determination is yes, execute Step 409, otherwise execute Step 411;
[0118] Step 409, generate a control instruction according to the voice data;
[0119] Step 410, execute the control instruction;
[0120] Step 411, send a prompt message to the user;
[0121] Step 412, determine whether a confirmation message from the user for this prompt message is received; if the determination is yes, execute Step 409, otherwise return to Step 401.
[0122] It should be noted that the above drawings only schematically illustrate the embodiments of the present application, but the present application is not limited thereto. For example, the execution order between various operations can be appropriately adjusted. In addition, some other operations can be added or some of the operations can be reduced (for example, the operations or steps corresponding to the dashed boxes in the drawings). Those skilled in the art can make appropriate modifications according to the above content, not limited to the records in the above drawings.
[0123] Each of the above embodiments only exemplarily illustrates the embodiments of the present application, but the present application is not limited thereto, and appropriate modifications can also be made on the basis of each of the above embodiments. For example, each of the above embodiments can be used alone, or one or more of the above embodiments can be combined.
[0124] According to the above embodiments, voice data is obtained through voice monitoring. When the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, a control instruction is generated based on the voice data and the control instruction is executed. In this way, misoperations caused by environmental noise or daily conversations can be reduced, and the user experience can be improved.
[0125] For example, in voice control with a wake word:
[0126] After the voice interaction device is awakened, it is possible to monitor background noise during the monitoring stage, which may cause misoperations of the device. By executing the control method of the embodiments of the present application in voice control with a wake word, when the obtained voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, a control instruction is generated and executed. In this way, the conditions for generating or executing the control instruction can be appropriately improved, thereby reducing the probability of device misoperations and improving the user experience.
[0127] Another example is in voice control without a wake word:
[0128] During the high-frequency usage stage, voice monitoring is performed in real time to avoid missing the user's control instructions. Since voice monitoring is performed in real time and the voice data does not need to include a wake word, the probability of misoperations caused by environmental noise or daily conversations is relatively high. By executing the control method of the embodiments of the present application during the high-frequency usage stage, a control instruction is generated or executed according to the voiceprint corresponding to the obtained voice data and the corresponding user status, which can significantly reduce the probability of device misoperations and improve the user experience.
[0129] During the low-frequency usage stage, voice monitoring is performed periodically, which can reduce the power consumption of the voice interaction device and protect the user's privacy. Since background noise such as televisions and radios may be monitored during the monitoring period, there is also a risk of device misoperations during the low-frequency usage stage. Moreover, the low-frequency usage stage may be at night, and device misoperations are likely to affect the user's rest. By executing the control method of the embodiments of the present application during the low-frequency usage stage, the probability of device misoperations can be reduced and the user experience can be improved.
[0130] Embodiment 2
[0131] Embodiment 2 of the present application provides a voice control system. This system corresponds to the method described in Embodiment 1, and the specific content can be referred to the description in Embodiment 1.
[0132] Figure 5 is a schematic diagram of the voice control system of the embodiments of the present application. As Figure 5As shown, the voice control system 500 includes a voice interaction device 510, a control device 520, and an execution device 530.
[0133] The voice interaction device 510 obtains voice data through voice monitoring; the control device 520 generates a control instruction according to the voice data when the voice data includes control-related information and meets a preset voiceprint standard and a user status judgment standard; and the execution device 530 executes the control instruction.
[0134] In some embodiments, the voice interaction device 510 includes at least one of a central controller, a line controller, a remote controller, a user terminal, and a smart speaker.
[0135] In some embodiments, the control device 520 includes at least one of a server, a gateway device, a central controller, a line controller, a remote controller, and a control board.
[0136] In some embodiments, the voice interaction device 510 and the control device 520 can be the same device. For example, the functions implemented by the voice interaction device 510 in this application are implemented by the control device 520. In this case, the control device 520 has the functions of voice monitoring and voice broadcasting. For example, both the control device 520 and the voice interaction device 510 are central controllers.
[0137] In some embodiments, the execution device 530 includes at least one of an air handling device and a smart device.
[0138] In some embodiments, the air handling device includes at least one of an indoor unit, a ventilation device, a humidity control device, a floor heating device, and a sensor.
[0139] In some embodiments, the indoor unit is an indoor unit in an air conditioning system, and the air conditioning system can be a commercial air conditioning system or a household air conditioning system.
[0140] In some embodiments, the air conditioning system can include at least one set of outdoor units and at least one indoor unit connected to each set of outdoor units. That is to say, in the air conditioning system, one or more sets of outdoor units can be included, and each set of outdoor units includes at least one outdoor unit; for one set of outdoor units, the set of outdoor units is connected to at least one indoor unit.
[0141] For example, the air conditioning system includes one outdoor unit and one indoor unit connected to the outdoor unit.
[0142] For example, the air conditioning system includes one outdoor unit and at least two indoor units connected to the outdoor unit.
[0143] For example, the air conditioning system includes at least two outdoor units and at least two indoor units connected to the at least two outdoor units.
[0144] For example, a group of outdoor units and at least one indoor unit connected to the group of outdoor units form a refrigerant system, or multiple groups of outdoor units and at least one indoor unit respectively connected to the multiple groups of outdoor units form multiple refrigerant systems. Thus, the air conditioning system may include one refrigerant system or multiple refrigerant systems.
[0145] In some embodiments, the outdoor unit and the indoor unit can be air conditioning devices of various models, types, forms, and capacities. For example, the form of the indoor unit can be all-round air outlet, two-sided air outlet, duct machine, floor air outlet, or skirting board air outlet, etc.; the form of the outdoor unit can be single-fan top air outlet, double-fan top air outlet, single-fan front air outlet, or double-fan front air outlet, etc.
[0146] In some embodiments, the air supply device includes at least one of a fresh air treatment device, a total heat exchange device (for example, with or without an internal circulation function), and a ventilation device (for example, with or without an internal circulation function).
[0147] In some embodiments, the humidity control device has at least one of a humidifying function and a dehumidifying function. For example, the humidity control device is arranged on the output side of the air supply device to adjust the humidity of the air output by the air supply device. For example, it humidifies or dehumidifies the air output by the air supply device and introduces the humidified or dehumidified air into the indoor space. For example, the humidity control device has a heat exchanger, and the heat exchanger can operate as an evaporator.
[0148] In some embodiments, the floor heating device is connected to the outdoor unit through a refrigerant pipeline, and the floor heating device is connected to the floor heating pipeline in the indoor space through a floor heating water pipeline. For example, the outdoor unit and the floor heating device form an air source heat pump water heater, using the heat in the air and the operation of the compressor to heat water; the floor heating device may include a water heat exchange unit, and the refrigerant flowing out of the compressor of the outdoor unit enters the water heat exchange unit of the floor heating device through the refrigerant pipeline to exchange heat with water, thereby achieving the purpose of heating water, and introducing the heated water into the floor heating pipeline in the indoor space through the floor heating water pipeline to supply heat to the indoor space. In some embodiments, the outdoor unit and the floor heating device are separate or integrated.
[0149] In some embodiments, the sensor can be various types of sensors: temperature, humidity, pressure, time, etc.
[0150] For example, the sensor may include at least one of the following: a sensor disposed at an air inlet or an air outlet of an indoor unit of an air handling system, a sensor in a controller electrically connected to the indoor unit, a sensor disposed at an outdoor air inlet (intake air port, air inlet) or an outdoor exhaust outlet (exhaust port, air outlet) or an indoor air supply port (air supply port) or an indoor air return port (air return port) of a fresh air handling device of the air handling system, a sensor in a controller electrically connected to the indoor unit or the fresh air handling device, an external sensor (for example, the external sensor may be connected to the indoor unit or the fresh air handling device through a repeater).
[0151] In some embodiments, the intelligent device includes at least one of a smart socket, a smart lighting device, and a smart curtain.
[0152] This application is not limited thereto. The intelligent device may further include household appliances such as a television, a washing machine, a refrigerator, or a water heater, or office equipment such as a computer, a projector, a printer, a scanner, a fax machine, or a telephone, or industrial equipment related to industrial production, and so on.
[0153] For the specific processing procedures of the above various devices, reference may be made to the relevant records in Embodiment 1, which will not be elaborated here.
[0154] According to the above embodiments, voice data is obtained through voice monitoring. When the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, a control instruction is generated according to the voice data and the control instruction is executed. In this way, misoperations caused by environmental noise or daily conversations can be reduced, and the user experience can be improved. For example, in voice control with a wake-up word or without a wake-up word, by generating or executing a control instruction according to the voiceprint corresponding to the voice data and the user status, the probability of misoperation can be significantly reduced.
[0155] Embodiment 3
[0156] Embodiment 3 of this application provides a voice control device. This voice control device corresponds to the voice control method described in Embodiment 1, and the specific content can be referred to the records in Embodiment 1.
[0157] Figure 6 is a schematic diagram of the voice control device of the embodiment of this application. As Figure 6 shown, the device 600 includes an acquisition module 610 and a processing module 620.
[0158] The acquisition module 610 obtains voice data through voice monitoring. The processing module 620 generates a control instruction according to the voice data when the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard.
[0159] In some embodiments, the device may further include a communication module 630, which sends the generated control instructions to an execution device that executes the control instructions via a communication network.
[0160] For the specific functions of the above units, reference may be made to the relevant steps in Embodiment 1, which will not be elaborated here.
[0161] It should be noted that only the components or modules related to the present application are described above, but the present application is not limited thereto. The voice control device may further include other components or modules or omit some components or modules (for example, the components or modules corresponding to the dashed boxes in the drawings). For the specific content of these components or modules, reference may be made to the related art.
[0162] In addition, for simplicity, the above drawings only exemplarily show the connection relationships or signal directions between the various components or modules, but those skilled in the art should clearly understand that various related technologies such as bus connection can be adopted. The above-mentioned various components or modules can be implemented by hardware facilities such as a processor, a memory, a transmitter, a receiver, etc.; the implementation of the present application does not limit this.
[0163] According to the above embodiments, voice data is obtained through voice monitoring. When the voice data includes control-related information and meets the preset voiceprint standard and user status judgment standard, a control instruction is generated based on the voice data and the control instruction is executed. In this way, misoperations caused by environmental noise or daily conversations can be reduced, and the user experience can be improved. For example, in voice control with a wake-up word or without a wake-up word, by generating or executing a control instruction according to the voiceprint and user status corresponding to the voice data, the probability of misoperation can be significantly reduced.
[0164] Embodiment 4
[0165] Embodiment 4 of the present application provides an electronic device. The steps executed by the processor of the electronic device correspond to all or part of the steps of the voice control method described in Embodiment 1, and the specific content can be referred to the description in Embodiment 1.
[0166] Figure 7 is a schematic diagram of the composition of the electronic device according to the embodiment of the present application. As Figure 7 shown, the electronic device 700 may include a processor 710 and a memory 720; the memory 720 is coupled to the processor 710. It should be noted that this figure is exemplary; other types of structures can also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0167] In one embodiment, the processor 710 may be configured to: obtain voice data through voice listening; generate a control instruction according to the voice data when the voice data includes control-related information and meets a preset voiceprint standard and a user status judgment standard; and execute the control instruction.
[0168] As Figure 7 shown, the electronic device 700 may further include: a communication module 730, an input unit 740, a display 750, a speaker 760, a microphone 770, and a power supply 780. It should be noted that the electronic device 700 does not necessarily have to include Figure 7 all the components shown in Figure 7 ; in addition, the electronic device 700 may further include
[0169] As Figure 7 shown, the processor 710 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The processor 710 receives inputs and controls the operations of the various components of the electronic device 700.
[0170] Among them, the memory 720 may be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. It can store various data, and can also store programs for executing relevant information. And the processor 710 can execute the programs stored in the memory 720 to implement information storage or processing, etc. The functions of other components are similar to those in the prior art and will not be elaborated here. The various components of the electronic device 700 can be implemented by dedicated hardware, firmware, software, or a combination thereof without departing from the scope of the present invention.
[0171] According to the above embodiments, voice data is obtained through voice listening, and when the voice data includes control-related information and meets a preset voiceprint standard and a user status judgment standard, a control instruction is generated according to the voice data and the control instruction is executed. In this way, it is possible to reduce misoperations caused by environmental noise or daily conversations, etc., and improve the user experience. For example, in voice control with a wake-up word or without a wake-up word, by generating or executing a control instruction according to the voiceprint corresponding to the voice data and the user status, the probability of misoperation can be significantly reduced.
[0172] The embodiments of the present application further provide a computer-readable program, wherein when the program is executed, the program causes a computer to execute the voice control method described in the embodiments of the present application.
[0173] An embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer-readable program enables a computer to execute the voice control method described in the embodiment of the present application.
[0174] The devices and methods above the embodiments of the present application can be implemented by hardware or by a combination of hardware and software. The present application relates to such a computer-readable program that when executed by a logic component, can enable the logic component to implement the above-mentioned device or component, or enable the logic component to implement the above-mentioned various methods or steps.
[0175] An embodiment of the present application also relates to a storage medium for storing the above program, such as a hard disk, a magnetic disk, an optical disk, a DVD, a flash memory, etc.
[0176] It should be noted that the limitations of the steps involved in the present application do not, on the premise of not affecting the implementation of the specific solution, be regarded as limiting the order of the steps. The steps written in the front can be executed first, or can be executed later, or even can be executed simultaneously. As long as the solution can be implemented, it should be regarded as falling within the protection scope of the present application.
[0177] The present application has been described above in combination with specific embodiments, but those skilled in the art should understand that these descriptions are exemplary and do not limit the protection scope of the present application. Those skilled in the art can make various variations and modifications to the present application according to the spirit and principle of the present application, and these variations and modifications are also within the scope of the present application.
Claims
1. A voice control method, characterized in that, The method includes: obtaining voice data through voice monitoring; when the voice data includes control-related information and meets a preset voiceprint standard and a user status judgment standard, generating a control instruction according to the voice data; and executing the control instruction.
2. The method according to claim 1, wherein the voice monitoring is started at a preset period, voice monitoring is performed within the first time of the preset period, and voice monitoring is stopped within other times of the preset period; or the voice monitoring is started at a preset period within a first time period; or the voice monitoring is performed in real time within a second time period.
3. The method according to claim 2, wherein the first time period or the second time period is set according to at least one of time information, user usage information, or user input information; or the frequency of generating or executing the control instruction within the first time period is less than the frequency of generating or executing the control instruction within the second time period.
4. The method according to claim 2, wherein the first time period includes a first interval and a second interval, within the first interval, the voice monitoring is started at a first preset period; within the second interval, the voice monitoring is started at a second preset period; wherein, the first frequency of generating or executing the control instruction within the first interval is lower than the second frequency of generating or executing the control instruction within the second interval, and the first preset period is greater than the second preset period.
5. The method according to claim 1, wherein the preset voiceprint standard is set according to the user's voiceprint characteristics; or when the voice data meets the preset voiceprint standard, the voice data is the user's voice data.
6. The method according to claim 1, wherein the user status judgment standard includes at least one of the following: speech rate standard, semantic standard, intonation standard, volume standard, or timbre standard.
7. The method according to claim 1, characterized in that, When the voice data includes control-related information and meets a preset voiceprint standard and a user status judgment standard, generating a control instruction according to the voice data includes: when the voice data includes control-related information, determines that the voice data is the user's voice data according to the preset voiceprint standard, and determines that the user is in a normal state according to the user status judgment standard, generating a control instruction according to the voice data.
8. The method according to claim 1, wherein Generating a control instruction according to the voice data includes: generating text information according to the voice data; and generating the control instruction according to the text information.
9. The method according to claim 1, characterized in that, When the voice data includes control-related information and meets a preset voiceprint standard and a user status judgment standard, generating a control instruction according to the voice data includes: within a first time period, when the voice data includes control-related information and meets a preset voiceprint standard and a user status judgment standard, generating a control instruction according to the voice data.
10. The method according to claim 9, wherein The method further includes: In a second time period different from the first time period, when the voice data includes control-related information, a control instruction is generated according to the voice data.
11. The method according to claim 1, characterized in that, The method further includes: When the voice data includes control-related information, conforms to a preset voiceprint standard, and the voice data does not conform to the user status judgment standard, a prompt message is sent to the user; and When a confirmation message of the user for the prompt message is received, a control instruction is generated and executed according to the voice data.
12. The method according to claim 1, characterized in that, The method further includes: setting whether the voice data needs to include a preset word or the content of the preset word according to the input information of the user, or When the voice data needs to include a preset word, the method further includes: When the voice data includes a preset word, a control instruction is generated according to the voice data.
13. The method according to claim 1, characterized in that, The method further includes: Storing a historical record of executing the control instruction.
14. The method according to claim 13, characterized in that, The historical record includes a first historical record of a first time period and a second historical record of a second time period, wherein the frequency of executing the control instruction in the first time period is less than the frequency of executing the control instruction in the second time period; or The historical record is stored in at least one of the following: a local device or a cloud device; or, The historical record is stored in at least one of the following: a voice interaction device that acquires the voice data, a control device that generates the control instruction, or an execution device that executes the control instruction.
15. A voice control system, characterized in that, The system includes: A voice interaction device that acquires voice data through voice monitoring; A control device that generates a control instruction according to the voice data when the voice data includes control-related information and conforms to a preset voiceprint standard and a user status judgment standard; and An execution device that executes the control instruction.
16. The system according to claim 15, characterized in that, The voice interaction device includes at least one of a central controller, a line controller, a remote controller, a user terminal, and a smart speaker; or The control device includes at least one of a server, a gateway device, a central controller, a line controller, a remote controller, and a control board; or The execution device includes at least one of an air handling device and a smart device.
17. The system according to claim 16, characterized in that, The air handling device includes at least one of an indoor unit, a air supply device, a humidity control device, a floor heating device, and a sensor; or The smart device includes at least one of a smart socket, a smart lighting device, and a smart curtain.
18. A voice control device, characterized in that, The device includes: An acquisition module that acquires voice data through voice monitoring; and A processing module that generates a control instruction according to the voice data when the voice data includes control-related information and conforms to a preset voiceprint standard and a user status judgment standard.
19. An electronic device, characterized in that, The electronic device includes: A memory that stores a computer program; and A processor, which implements the method according to any one of claims 1-14 when executing the computer program.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program implements the method according to any one of claims 1-14 when executed by a processor.