Household equipment control method and device, electronic equipment and storage medium

By adjusting the cloud model to a local control model, and extracting voiceprint feature of voiceprint and determining user permissions in smart home devices, the slow response speed and security of voice control of smart home devices are solved, and equipment control with low latency and high security is achieved.

CN120143635APending Publication Date: 2025-06-13CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510142434.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The voice control of smart home devices relies on cloud computing, resulting in slow response speed of devices, affecting user experience, and problems such as user privacy data leakage and poor security.

Method used

By obtaining resource information of home equipment, adjusting the preset cloud model to obtain a local control model, extracting voiceprint features of voice patterns, determining user operation permissions, and using local control model to control the target equipment to realize local data processing.

Benefits of technology

Improves device response speed, reduces the possibility of user privacy data breaches, enhances security, and provides low latency and reliable services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143635A_ABST
    Figure CN120143635A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a control method of home equipment, which is applied to the field of smart home, and comprises the following steps: obtaining resource information of the home equipment, adjusting a preset cloud model according to the resource information, obtaining a local control model, obtaining voice information, extracting voiceprint feature information of the voice information, and obtaining the home equipment according to the voiceprint feature information. The method comprises the following steps: determining a user operation authority, determining a target control device for voice information under the condition that the user operation authority is a specified type, controlling the target control device according to the user operation authority by adopting a local control model, and locally processing all data, so that the data security is ensured, and the privacy is protected. And low-delay and reliable services are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart home, and particularly to a control method and device for home appliances, an electronic device, and a storage medium. Background Art

[0002] Smart home refers to a modern lifestyle in which various devices in a home are connected through Internet of Things technology and centrally controlled and managed through an intelligent system. It is applied to various smart home products, such as smart speakers, smart lighting systems, smart security systems, etc. The smart lighting system refers to automatically adjusting the brightness and color of lights according to time, light, or user habits. The smart security system includes smart door locks, cameras, door and window sensors, etc., which can monitor the home security status in real time and issue an alarm in case of an abnormality; Smart home appliances such as smart refrigerators, washing machines, air conditioners, etc. can be remotely controlled or run through an automated program to improve the living efficiency; Environmental monitoring devices refer to real-time monitoring of the home environment and automatic adjustment through temperature and humidity sensors, air quality detectors, etc. to create a comfortable living condition.

[0003] In the related art, smart home devices are usually connected to a cloud service platform to process speech recognition and natural language processing tasks through the powerful computing resources of the cloud. Some special devices also support integration with other smart home systems. The cloud also stores user data and device status information and provides a remote control function.

[0004] Although the smart home voice control solutions in the related art have achieved automation and intelligence in many aspects, there are still limitations. Since speech recognition and natural language processing tasks usually rely on cloud computing, the high latency of cloud operations will result in a slow response speed of the device, affecting the user experience. Summary of the Invention

[0005] In view of the above problems, a control method and device for home appliances, an electronic device, and a storage medium are provided to overcome or at least partially solve the above problems, including:

[0006] In the first aspect of the implementation of the present application, a control method for home appliances is first provided, and the method includes:

[0007] Obtain the resource information of the home appliance, and adjust a preset cloud model according to the resource information to obtain a local control model;

[0008] Obtain voice information, and extract the voiceprint feature information of the voice information;

[0009] Determine the user operation permission according to the voiceprint feature information;

[0010] When the user operation permission is of a specified type, determine the target control device for the voice information;

[0011] Use the local control model to control the target control device according to the user operation permission.

[0012] In an alternative embodiment of the present application, the adjusting the preset cloud model according to the resource information to obtain a local control model includes:

[0013] Obtain open-source Q&A data and historical Q&A data, and extract target Q&A data corresponding to the home appliances from the open-source Q&A data and historical Q&A data;

[0014] Use the target Q&A data and the resource information to adjust the preset cloud model to obtain the local control model.

[0015] In an alternative embodiment of the present application, the using the target Q&A data and the resource information to adjust the preset cloud model to obtain the local control model includes:

[0016] Divide the target Q&A data to obtain a training data set, a test data set, and a validation data set;

[0017] Use the training data set to train the preset cloud model to obtain a first control model;

[0018] Use the validation data set to evaluate the first control model to obtain an evaluation result;

[0019] Use the evaluation result to adjust the hyperparameters of the first control model to obtain a second control model;

[0020] Use the test data set to test the second control model, and when the test passes, determine the second control model as the local control model.

[0021] In an alternative embodiment of the present application, before obtaining the resource information of the home appliance, it further includes:

[0022] Collect positive sample audio segments and negative sample audio segments;

[0023] Label the positive sample audio segments and negative sample audio segments to obtain positive sample reference data and negative sample reference data;

[0024] Use the positive sample audio segments and the positive sample reference data, the negative sample audio segments and the negative sample reference data to train a preset wake-up model to obtain a target wake-up model.

[0025] In an alternative embodiment of the present application, before extracting the voiceprint feature information of the voice information, it includes:

[0026] Using the target wake-up model to detect whether there is a wake-up word in the voice information, and obtaining a detection result;

[0027] According to the detection result, determining whether to perform the step of extracting the voiceprint feature information of the voice information.

[0028] In an alternative embodiment of the present application, extracting the voiceprint feature information of the voice information includes:

[0029] Extracting the Mel cepstral coefficients of the voice information to obtain voiceprint feature information;

[0030] Determining the user operation permission according to the voiceprint feature information includes:

[0031] According to the voiceprint feature information, determining whether the voiceprint feature information matches the pre-stored feature information, and obtaining a matching result;

[0032] Determining the user operation permission according to the matching result.

[0033] In an alternative embodiment of the present application, after controlling the target control device according to the user operation permission, it includes:

[0034] Determining the interaction text information corresponding to the voice information;

[0035] Converting the interaction text information into interaction voice information and outputting the interaction voice information.

[0036] In the second aspect of the implementation of the present application, there is also provided a control device for a home device, characterized in that the device includes:

[0037] A resource acquisition module, configured to acquire the resource information of the home device, and adjust a preset cloud model according to the resource information to obtain a local control model;

[0038] A feature extraction module, configured to acquire voice information and extract the voiceprint feature information of the voice information;

[0039] A permission determination module, configured to determine the user operation permission according to the voiceprint feature information;

[0040] A device determination module, configured to determine the target control device for the voice information when the user operation permission is of a specified type;

[0041] The device control module is configured to control the target control device according to the user operation authority by using the local control model.

[0042] An embodiment of the present application further discloses an electronic device, which is characterized by including a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the control method of the home device as described above is implemented.

[0043] An embodiment of the present application further discloses a computer-readable storage medium, which is characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the control method of the home device as described above is implemented.

[0044] The embodiment of the present invention has the following advantages: All data is processed locally, providing low latency and reliable services while ensuring data security and protecting privacy. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 is a flowchart of the steps of a control method for a home device provided by an embodiment of the present invention;

[0047] Figure 2 is a flowchart of the steps of another control method for a home device provided by an embodiment of the present invention;

[0048] Figure 3 is a block diagram of the structure of a control device for a home device provided by an embodiment of the present invention. Detailed Embodiments

[0049] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0050] The smart home control solutions of the related technologies mainly rely on two core technologies: speech recognition technology and natural language processing technology. By using speech recognition technology and natural language processing technology, voice control of various smart home devices can be achieved. However, these devices usually need to be connected to a cloud service platform, and the cloud computing resources are used to process speech recognition and natural language processing tasks. Some special devices also support integration with other smart home systems, but this integration is usually limited and requires specific protocols or APIs (Application Programming Interfaces). Although the smart home control solutions of the related technologies have achieved automation and intelligence in some aspects, there are still limitations, such as high latency in cloud operations, easy leakage of user privacy data, and poor security.

[0051] Referring to Figure 1 , the step flowchart of a control method for a home device provided by some embodiments of the present invention is shown, which may specifically include the following steps:

[0052] Step 101: Obtain the resource information of the home device, and adjust the preset cloud model according to the resource information to obtain a local control model.

[0053] Among them, the home devices include smart lighting devices, security devices, environmental control devices, entertainment devices, household appliances, and energy management devices. Among them, the smart lighting devices can support functions such as remote control, dimming, color temperature adjustment, and timed switching. The security devices include smart door locks, smart cameras, smart doorbells, etc. The smart door locks support functions such as fingerprint, password, and remote unlocking. The smart cameras support functions such as real-time monitoring, motion detection, and video storage. The smart doorbells support functions such as video calls, motion detection, and remote viewing. The entertainment devices include, but are not limited to, smart TVs, smart speakers, home theater systems, etc. The smart TVs support voice control, streaming media playback, and linkage with other devices. The smart speakers support voice assistants, music playback, and device control. The home theater systems support multi-room audio playback and remote control. The household appliances include, but are not limited to, smart refrigerators, smart washing machines, smart floor sweepers, etc. Among them, the smart refrigerators support functions such as food ingredient management, recipe recommendation, and remote viewing of the internal situation. The smart washing machines support functions such as remote start, mode selection, and fault diagnosis. The smart floor sweeping robots support functions such as automatic cleaning, path planning, and remote control. These devices have improved the convenience, security, and comfort of family life through intelligent technologies.

[0054] A preset cloud model refers to a machine learning or artificial intelligence model deployed on a remote cloud server. Users access these models via the Internet, send input data (such as text, images, voice, etc.) to them, and the cloud model processes the data and returns the results. In related technologies, the data processing for home device control and question-and-answer interaction is all through the cloud model, resulting in a long delay, easy leakage of user privacy data, and poor security. In this application, by adjusting the preset cloud model, a local control model is obtained and deployed in the home device control system to process data locally, improve the data processing speed, reduce the possibility of privacy data leakage, and have high security.

[0055] However, since each family is equipped with different home devices, some families may have devices such as smart curtains, smart sweepers, and smart washing machines, while some families may have devices such as smart door locks and smart speakers, that is, the resource information of home devices in each family is different. Therefore, it is necessary to adjust the preset cloud model according to the resource information of home devices in each family to obtain a local control model.

[0056] In the specific implementation of the embodiments of this application, the preset cloud model can be the Qwen model. By adjusting the Qwen model according to the resource information of home devices, a local control model is obtained. That is, in this application, by adjusting the preset cloud model through the different home device resource information of each family, the obtained local control model can meet personalized needs and better serve users. Through the obtained local control model, all data is processed locally at the user end, providing low latency and reliable services while ensuring data security and protecting privacy.

[0057] Step 102: Obtain voice information and extract the voiceprint feature information of the voice information.

[0058] Among them, the voice information refers to the voice commands or voice data issued by the user when interacting with the smart home system through voice. These voice information are the key inputs for the smart home system to understand the user's intention and perform corresponding operations. The voiceprint feature information refers to the features that can uniquely identify the speaker extracted by analyzing the voice signal. Everyone's voice has unique features in terms of frequency, pitch, rhythm, etc. In this application, the voiceprint feature information is used for identity recognition to prevent illegal users from controlling home devices and improve the security of home device control.

[0059] Step 103: Determine the user operation authority according to the voiceprint feature information.

[0060] After extracting the voiceprint feature information, the user's operation authority can be determined through the following steps: comparing the voiceprint features of the current user with the pre-registered voiceprint templates.

[0061] For example, the voiceprint features when the user says "Turn on the lights in the living room" are matched with the voiceprint templates stored in the database. According to the voiceprint matching result, the user's identity is verified. If the match is successful, the user is confirmed as "Zhang San"; if the match fails, "Identity verification failed" is prompted. According to the user's identity, their operation permissions are checked. For another example, if the user is an "administrator", all operations are allowed. If the user is an "ordinary user", only partial operations are allowed (such as controlling the lights, but not modifying system settings). If the user's permission verification passes, the corresponding operation is executed. If the permission is insufficient or the identity verification fails, the request is rejected and the user is prompted.

[0062] In some embodiments of the present application, the user age information corresponding to the voice information can also be further determined based on the voiceprint feature information. Then, the user operation permissions are determined by combining the voiceprint feature information and the user age information. If the user is "Master A" and an adult, the operation of turning on the gas stove can be executed. However, if the user is "Master B" and a child under 12 years old, they are not allowed to execute the operation of turning on the gas stove, but can execute other operations such as controlling the lights, controlling the air conditioner, and opening the curtains.

[0063] Step 104: When the user operation permission is of a specified type, determine the target control device for the voice information.

[0064] Here, the user operation permission being of a specified type means that the user has successfully verified their identity and has operation permissions. For example, when the user says "Turn on the lights in the living room" or "Turn on the speaker", and it is determined that the user is a family member corresponding to the home device based on the voiceprint feature information, it is determined that the user has operation permissions. On this basis, the target control device for the voice information is determined. In some embodiments, the target control device for the voice information can be determined through voice recognition technology. For example, if the text information contained in the voice is recognized as "living room, lights", the target control device is determined as the "living room lights"; if the voice contains "air conditioner", "temperature", and "26 degrees", the target control device is determined as the "air conditioner".

[0065] Step 105: Use the local control model to control the target control device according to the user operation permission.

[0066] The local control model refers to a small and refined model that can be deployed locally after adjusting the preset cloud model (cloud large model). In specific implementation, the preset cloud model can be quantized and pruned, converting the model weights from high precision to lower precision, reducing storage requirements and computational volume, and removing unnecessary connections and neurons in the network. The local control model is more suitable for the home environment, can be personalized for different families, and the fine-tuned large model, that is, the local control model, has better recognition ability for voice control instructions and can better serve users.

[0067] After determining the user operation permissions and the target control device, a small and refined local control model is adopted to control the target control device. For example, turn on the living room light or adjust the living room air conditioner to 26 degrees.

[0068] In the embodiments of the present application, by obtaining the resource information of the home devices, adjusting the preset cloud model according to the resource information to obtain a local control model, obtaining voice information, and extracting the voiceprint feature information of the voice information, and determining the user operation permissions according to the voiceprint feature information. When the user operation permissions are of a specified type, determining the target control device for the voice information, and using the local control model to control the target control device according to the user operation permissions, and locally processing all data, while ensuring data security and protecting privacy, providing low-latency and reliable services.

[0069] Referring to Figure 2 , a step flowchart of another control method for home devices provided by an embodiment of the present invention is shown, which may specifically include the following steps:

[0070] Step 201: Obtain the resource information of the home devices.

[0071] In some embodiments of this embodiment, there are different home devices in each family, and these different home devices can be smart home devices of different brands, different ecosystems or different protocols. These devices are usually produced by different manufacturers and may use different communication protocols, control interfaces or cloud service platforms. To achieve unified control of these devices, the embodiments of the present application connect them to a unified smart home platform, such as Home Assistant (HA). Home Assistant supports connecting smart home devices from different platforms and achieving unified control. Users can no longer be limited to devices of a single brand and can freely choose the products that best suit their needs. In specific implementation, the resource information of the home devices in the family can be obtained through the HA platform.

[0072] In some embodiments of the present application, the following steps may be included before step 201:

[0073] S11: Collect positive sample audio segments and negative sample audio segments;

[0074] S12: Label the positive sample audio segments and negative sample audio segments to obtain positive sample reference data and negative sample reference data;

[0075] S13: Use the positive sample audio segments and the positive sample reference data, the negative sample audio segments and the negative sample reference data to train a preset wake-up model to obtain a target wake-up model.

[0076] Among them, the positive sample audio segment refers to the speech segment containing the wake-up word. In specific implementation, the microphone can be used to record the speech of the user saying the wake-up word, ensuring coverage of different speech rates, intonations, accents, and environments (such as quiet environment, noisy environment). For example, if the wake-up word is "ABC", the recorded positive sample audio segments can be "ABC, turn on the air conditioner", "ABC, turn off the TV", etc. said by the user in a noisy environment. The negative sample audio segment refers to the speech segment not containing the wake-up word. In specific implementation, background noise (such as street noise, home appliance noise) can be recorded, other words or phrases (such as "hello") can be recorded, and non-speech sounds (such as knocking sounds, music sounds) can be recorded.

[0077] Annotate each audio segment to obtain the reference data corresponding to the positive sample audio segment and the negative sample parameter data corresponding to the negative sample audio segment. For example, the positive sample audio segment can be annotated with 1 (indicating containing the wake-up word), and the negative sample audio segment can be annotated with 0 (indicating not containing the wake-up word).

[0078] Among them, the preset wake-up model is used to wake up the unified control platform of the home appliance during actual use. The preset wake-up model can be a deep learning model. Specifically, the deep learning model can be CNN (Convolutional Neural Network), which is suitable for processing the spectral features of audio; RNN / LSTM (Recurrent Neural Network), which is suitable for processing time series data; CRNN (Convolutional Recurrent Neural Network), which combines the advantages of CNN and RNN.

[0079] Use the positive sample audio segment and its corresponding positive sample reference data, the negative sample audio segment and its corresponding negative sample reference data to train the preset wake-up model, and a target wake-up model that can accurately recognize the wake-up word can be obtained.

[0080] Step 202: Obtain the open-source Q&A data and historical Q&A data, and extract the target Q&A data corresponding to the home appliance from the open-source Q&A data and historical Q&A data.

[0081] Step 203: Adjust the preset cloud model using the target Q&A data and the resource information to obtain the local control model.

[0082] In some embodiments of this embodiment, a high-quality Q&A dataset for home appliances can be constructed by extracting open-source Q&A data from open-source datasets or obtaining historical Q&A data from daily usage records. The open-source datasets include, but are not limited to, CoQA (Conversational Question Answering), SQuAD (Stanford Question Answering Dataset), and MS MARCO (Microsoft Machine Reading Comprehension). CoQA contains conversational Q&A data, SQuAD contains Q&A pairs, and MS MARCO contains real user questions and answers. Data can also be obtained from daily usage records, such as obtaining user logs of smart home appliances, or obtaining user feedback or customer service records. Interaction records of voice assistants can also be obtained, and user questions and corresponding answers can be extracted.

[0083] Extract target Q&A data corresponding to home appliances from the obtained open-source Q&A data and historical Q&A data. For example, when the home appliances include speakers and air conditioners, target Q&A data related to speakers and air conditioners can be obtained, specifically Q&A data pairs. For example, the user question "What music should be played when I'm in a good mood?" and the answer "Music by singer YY can be played on your XX speaker" can be obtained. Another example is that the user question "Can you help me set the air conditioner to 25 degrees?" and the answer "Sure, the air conditioner in your living room is now set to 25 degrees."

[0084] In some embodiments of this embodiment, tags can also be added to each Q&A pair, such as device type (lights, thermostats, security, etc.), and the intent of the question (such as control, query, setting) can be marked. The preset cloud model is adjusted using the obtained target Q&A data, corresponding tags, and the resource information obtained in the previous steps to obtain a local control model. The local control model is integrated into the HA platform. The cooperation between the local control model and the HA platform can not only control home appliances of various brands but also interact with users through voice for Q&A.

[0085] In some embodiments of this embodiment, a high-quality base large model can be selected as the preset cloud model. Considering that this model is only used in the local area network and the resources of users are usually limited, a model with less than 10B parameters can be selected to meet daily use and ensure more than 90% correctness. After constructing a high-quality Q&A dataset, data processing can also be performed on the dataset. For example, incorrect, duplicate, and irrelevant data can be removed, Q&A pairs that do not conform to grammar or logic can be deleted, duplicate Q&A pairs can be merged or deleted, and Q&A pairs irrelevant to home appliances can be deleted. The preset cloud model is adjusted using the processed dataset to reduce the amount of data processing and save computing costs.

[0086] In some other embodiments of this embodiment, a large-parameter model can also be used to synthesize data for the constructed high-quality dataset. A large-parameter model refers to a machine learning model with a very large number of parameters, especially a deep learning model. Such models usually have billions or even trillions of parameters and can capture complex patterns and relationships in the data. In this application, by using a large-parameter model to synthesize data for the constructed high-quality dataset, the data volume is further increased. The high-quality dataset is used as input to the large-parameter model, and the content and form of the output result are specified. A larger-scale dataset is quickly and efficiently generated through the capabilities of the large-parameter model. Then, the preset cloud model is adjusted using the generated larger-scale dataset, so that the local control model can improve the accuracy of Q&A interactions and enhance the user experience.

[0087] In some embodiments of this application, step 203 may include the following sub-steps:

[0088] Sub-step 11: Divide the target Q&A data to obtain a training dataset, a test dataset, and a validation dataset;

[0089] Sub-step 12: Train the preset cloud model using the training dataset to obtain a first control model;

[0090] Sub-step 13: Evaluate the first control model using the validation dataset to obtain an evaluation result;

[0091] Sub-step 14: Adjust the hyperparameters of the first control model using the evaluation result to obtain a second control model;

[0092] Sub-step 15: Test the second control model using the test dataset. If the test passes, determine the second control model as the local control model.

[0093] Among them, the training dataset is used for model training to adjust model parameters. The model learns patterns and regularities in the data through the training dataset. During the training process, the model adjusts its parameters to minimize the loss function. The training dataset is usually the largest part of the dataset (such as 70%-80%), and it needs to cover the diversity of the data to ensure that the model can generalize to unseen data.

[0094] During the training process, the validation dataset is used to monitor the performance of the model for tuning to prevent overfitting. By observing the performance on the validation dataset, the best hyperparameters can be selected. The test dataset is used for the final evaluation of the model performance, reflecting the generalization ability of the model. In specific implementations, the test dataset is strictly isolated and only used in the final evaluation to prevent the model from "peeking" at the test data during the training process.

[0095] In some embodiments of this embodiment, the target Q&A data can be divided into a training dataset, a validation dataset, and a test dataset to prepare data for model training. For example, the data can be divided into a training set (70%), a validation set (15%), and a test set (15%), or the data can be divided into a training set (70%), a validation set (20%), and a test set (10%). This application does not limit the proportion of data division.

[0096] First, use the obtained training dataset after division to train a preset cloud model. For example, convert the Q&A pairs into the model input format, and the model can also be fine-tuned using the training data. Save the trained model as a file, that is, generate the first control model. Second, use the validation dataset to evaluate the first control model to obtain an evaluation result. For example, calculate metrics such as the accuracy rate and F1 score of the model on the validation set. Determine the model performance based on the evaluation result and decide whether to adjust the hyperparameters. If adjustment is needed, adjust the hyperparameters of the first control model. In specific implementations, hyperparameters such as the learning rate, batch size, and number of model layers of the model can be adjusted, and use the adjusted hyperparameters to retrain the model to obtain the second control model. Third, use the test dataset to test the adjusted second control model. If the test passes, it is determined as the local control model. For example, if the test result meets the requirements (such as the accuracy rate > 90%), then determine the second control model as the local control model.

[0097] Step 204: Obtain voice information and extract the voiceprint feature information of the voice information.

[0098] In some embodiments of this application, before step 204, the following steps may be included:

[0099] S41: Use the target wake-up model to detect whether there is a wake-up word in the voice information to obtain a detection result.

[0100] In a specific implementation, the target wake-up model obtained by training using the steps of the foregoing steps S11 - S13 can be used to detect whether a wake-up word exists in the voice information. For example, if the wake-up word is "ABC", when the entire smart home HA platform is in the sleep mode and the voice information "ABC, please turn on the air conditioner" is received, the voice information can be first detected by the target wake-up model to obtain a detection result of "including the wake-up word"; for another example, in the sleep mode, when the voice information "open the curtain" is received, the voice information is detected by the target wake-up model to obtain a detection result of "not including the wake-up word".

[0101] S42: According to the detection result, determine whether to perform the step of extracting the voiceprint feature information of the voice information.

[0102] According to the detection result obtained in step S41, determine whether to perform the subsequent steps of extracting the voiceprint feature for identity verification. If the wake-up word is not included, operations such as voiceprint feature extraction and identity verification are not performed. Wake-up word detection is a lightweight task that only needs to detect specific keywords (such as "ABC"). Only after the wake-up word is detected will more complex voiceprint feature detection be started, thereby reducing unnecessary computational overhead. When the wake-up word is not detected, the smart home HA platform can maintain a low-power state and save energy. In addition, false triggers can be reduced, and since subsequent voiceprint feature extraction operations, identity authentication, operation permission determination, and operations on whether to continuously obtain voice information are only performed after the wake-up word is detected, the privacy protection of the user's daily conversations in the home environment is enhanced.

[0103] In some embodiments of the present application, "extracting the voiceprint feature information of the voice information" in step 204 may include the following sub-steps:

[0104] Sub-step 21: Extract the Mel cepstral coefficients of the voice information to obtain voiceprint feature information.

[0105] When the smart home HA platform is in the sleep mode and the wake-up word is detected, the voiceprint feature information can be obtained by extracting the Mel cepstral coefficients of the voice information. Among them, the Mel cepstral coefficients (MFCC, Mel-Frequency Cepstral Coefficients) is a voice feature extraction method based on the auditory characteristics of the human ear, which can capture the spectral characteristics of the voice signal. The specific implementation of the Mel cepstral coefficients is to perform a non-linear transformation (Mel scale) on the frequency distribution of the voice signal, which conforms to the auditory characteristics of the human ear, and the extracted feature vector has a low dimension and is suitable for use in machine learning models.

[0106] Step 205: Determine the user operation permission according to the voiceprint feature information.

[0107] In some embodiments of the present application, step 205 may include the following sub-steps:

[0108] Sub-step 31: According to the voiceprint feature information, determine whether the voiceprint feature information matches the pre-stored feature information to obtain a matching result;

[0109] Sub-step 32: Determine the user operation permission according to the matching result.

[0110] In some implementation manners of this embodiment, to determine whether the voiceprint feature information matches the pre-stored feature information, the voiceprint feature of the currently received voice information is compared one by one with the pre-stored voiceprint feature to determine whether the pre-stored voiceprint feature can be matched. In a specific implementation, the voiceprint features of each authorized user, such as MFCC feature vectors, can be pre-stored, and a similarity measurement method (such as cosine similarity, Euclidean distance) is used to calculate the matching degree between the current voiceprint feature and the pre-stored feature. If the matching degree exceeds the set threshold, it is considered a successful match; otherwise, the match fails.

[0111] After obtaining the matching result, it can be determined whether the user has the permission to perform a specific operation according to the voiceprint matching result. In a specific implementation, different operation permissions (such as administrator, ordinary user) can be assigned to each user first, and permission checking is performed through voiceprint matching. If the voiceprint match is successful, the operation is performed according to the user permission; if the match fails, the operation is rejected. The operation permission information and operation result can also be fed back to the user. For example, "operation successful" or "insufficient permission" is fed back to the user.

[0112] Step 206: When the user operation permission is of a specified type, determine the target control device for the voice information;

[0113] Step 207: Use the local control model to control the target control device according to the user operation permission.

[0114] Steps 206 - 207 are similar to the aforementioned steps 104 - 105 and will not be elaborated here.

[0115] Step 208: Determine the interactive text information corresponding to the voice information;

[0116] Step 209: Convert the interactive text information into interactive voice information and output the interactive voice information.

[0117] Among them, the interactive text information corresponding to the voice information refers to the response result obtained by the local control model querying the local database after obtaining the voice information. Therefore, the interactive text information can also be referred to as the text response result, and the response result is information in text format. The exchanged voice information refers to the voice information obtained by performing voice conversion on the information in text format.

[0118] In some implementation manners of this embodiment, the result returned by the local control model can be subjected to text-to-speech conversion to interact with the user in the form of voice, providing the user with a seamless voice interaction experience. Specifically, first, an efficient and natural text-to-speech (TTS) model is selected to implement the conversion of text to voice. For example, Google TTS can be used. Google TTS is simple and easy to use and supports multiple languages; Microsoft Azure TTS can also be used, which supports high-quality speech synthesis and custom voices. The result returned by the local control model is a text string, and the model is used to convert the text into a voice file or directly play it. In some implementation manners, the language setting of the TTS model can also be dynamically switched according to the language type of the voice information corresponding to the user input. For example, if the voice information input by the user is Chinese, the returned text is converted into Chinese through the TTS model; if the voice information input by the user is English, the returned text is converted into English through the TTS model, thereby providing a more natural interaction experience for the user. In some implementation manners, the Wyoming integration tool can also be used to create Wyoming configuration files for each service (speech recognition, text-to-speech, wake word detection), simplify the integration of services such as speech recognition, text-to-speech, and wake word detection through the Wyoming protocol, seamlessly integrate these services into a unified voice interaction system, implement the integration of functions such as speech recognition, text-to-speech, wake word detection, and identity confirmation on the user local, and a voice assistant can also be created in Home Assistant to implement voice control of smart home devices.

[0119] In the embodiment of this application, by obtaining the resource information of the home device, adjusting the preset cloud model according to the resource information to obtain the local control model, obtaining the voice information, and extracting the voiceprint feature information of the voice information, and determining the user operation permission according to the voiceprint feature information. When the user operation permission is of a specified type, determining the target control device for the voice information, using the local control model to control the target control device according to the user operation permission, and locally processing all data, while ensuring data security and protecting privacy, providing low-latency and reliable services. Saving computing resources through wake word detection and determining the user operation permission through the matching result of the voiceprint feature information, improving the security of home device control.

[0120] It should be noted that, for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be carried out in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0121] Referring to Figure 3 , a schematic structural diagram of a control device for a home appliance provided by an embodiment of the present invention is shown, which may specifically include the following modules:

[0122] A resource acquisition module 301, configured to acquire resource information of the home appliance, and adjust a preset cloud model according to the resource information to obtain a local control model;

[0123] A feature extraction module 302, configured to acquire voice information and extract voiceprint feature information of the voice information;

[0124] An authority determination module 303, configured to determine a user operation authority according to the voiceprint feature information;

[0125] An equipment determination module 304, configured to determine a target control device for the voice information when the user operation authority is of a specified type;

[0126] An equipment control module 305, configured to use the local control model to control the target control device according to the user operation authority.

[0127] In an alternative embodiment of the present application, the resource acquisition module 301 includes:

[0128] A Q&A data acquisition sub-module, configured to acquire open-source Q&A data and historical Q&A data, and extract target Q&A data corresponding to the home appliance from the open-source Q&A data and historical Q&A data;

[0129] A model adjustment sub-module, configured to adjust the preset cloud model by using the target Q&A data and the resource information to obtain the local control model.

[0130] In an alternative embodiment of the present application, the model adjustment sub-module includes:

[0131] A data division unit, configured to divide the target Q&A data to obtain a training data set, a test data set, and a validation data set;

[0132] A model training unit, configured to train the preset cloud model using the training data set to obtain a first control model;

[0133] A model evaluation unit, configured to evaluate the first control model using the verification data set to obtain an evaluation result;

[0134] A model adjustment unit, configured to adjust the hyperparameters of the first control model using the evaluation result to obtain a second control model;

[0135] A model testing unit, configured to test the second control model using the test data set, and in the case of passing the test, determine the second control model as the local control model.

[0136] In an alternative embodiment of the present application, the apparatus further includes:

[0137] An audio acquisition module, configured to acquire positive sample audio segments and negative sample audio segments;

[0138] An audio annotation module, configured to annotate the positive sample audio segments and negative sample audio segments to obtain positive sample reference data and negative sample reference data;

[0139] A wake-up model training module, configured to train a preset wake-up model using the positive sample audio segments and the positive sample reference data, and the negative sample audio segments and the negative sample reference data to obtain a target wake-up model.

[0140] In an alternative embodiment of the present application, the apparatus further includes:

[0141] A wake-up word detection module, configured to detect whether there is a wake-up word in the voice information using the target wake-up model to obtain a detection result;

[0142] A judgment module, configured to judge whether to perform the step of extracting the voiceprint feature information of the voice information according to the detection result.

[0143] In an alternative embodiment of the present application, the feature extraction module 302 includes:

[0144] A coefficient extraction sub-module, configured to extract the Mel cepstral coefficients of the voice information to obtain voiceprint feature information;

[0145] The permission determination module 303 includes:

[0146] A matching result determination sub-module, configured to judge whether the voiceprint feature information matches the pre-stored feature information according to the voiceprint feature information to obtain a matching result;

[0147] A permission determination sub-module, configured to determine the user operation permission according to the matching result.

[0148] In an alternative embodiment of the present application, the device further includes:

[0149] A text determination module, configured to determine the interactive text information corresponding to the voice information;

[0150] A voice conversion module, configured to convert the interactive text information into interactive voice information and output the interactive voice information.

[0151] An embodiment of the present invention further provides an electronic device, which may include a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the control method of the above home device is implemented.

[0152] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the control method of the above home device is implemented.

[0153] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment.

[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0155] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0156] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0157] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0160] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0161] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the above elements.

[0162] The above provides a detailed introduction to the control method and device for home appliances, electronic devices, and storage media. In this article, specific examples are used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for controlling a household appliance, characterized in that: The method comprises: Acquire resource information of the home device, and adjust the preset cloud model according to the resource information to obtain a local control model; Acquire voice information and extract voiceprint feature information of the voice information; Determining the user's operation authority based on the voiceprint feature information; In the case where the user operation authority is of a specified type, determining a target control device for the voice information; The local control model is adopted to control the target control device according to the user operation authority.

2. The method according to claim 1, characterized in that The step of adjusting the preset cloud model according to the resource information to obtain a local control model includes: Acquire open source question and answer data and historical question and answer data, and extract target question and answer data corresponding to the home device from the open source question and answer data and the historical question and answer data; The preset cloud model is adjusted using the target question and answer data and the resource information to obtain the local control model.

3. The method according to claim 2, characterized in that The step of adjusting the preset cloud model by using the target question-answer data and the resource information to obtain the local control model includes: Dividing the target question-answer data into a training data set, a test data set, and a validation data set; Using the training data set to train the preset cloud model to obtain a first control model; Using the verification data set to evaluate the first control model, to obtain an evaluation result; Using the evaluation result, adjusting the hyperparameters of the first control model to obtain a second control model; The second control model is tested using the test data set, and if the test passes, the second control model is determined as the local control model.

4. The method according to claim 1, characterized in that: Before obtaining the resource information of the home device, the method further includes: Collect positive sample audio clips and negative sample audio clips; Annotating the positive sample audio clips and the negative sample audio clips to obtain positive sample reference data and negative sample reference data; The preset wake-up model is trained using the positive sample audio segment and the positive sample reference data, the negative sample audio segment and the negative sample reference data to obtain a target wake-up model.

5. The method according to claim 4, characterized in that Before extracting the voiceprint feature information of the voice information, the method includes: Using the target wake-up model to detect whether there is a wake-up word in the voice information, and obtaining a detection result; Based on the detection result, it is determined whether it is necessary to execute the step of extracting the voiceprint feature information of the speech information.

6. The method according to claim 1, characterized in that The step of extracting the voiceprint feature information of the speech information includes: Extracting the Mel-frequency cepstral coefficients of the speech information to obtain voiceprint feature information; Determining the user operation authority based on the voiceprint feature information includes: According to the voiceprint feature information, determining whether the voiceprint feature information matches pre-stored feature information, and obtaining a matching result; The user operation authority is determined based on the matching result.

7. The method according to any one of claims 1 to 6, characterized in that: After the target control device is controlled according to the user operation authority, the method includes: Determining interactive text information corresponding to the voice information; The interactive text information is converted into interactive voice information, and the interactive voice information is output.

8. A control device for household appliances, characterized in that: The device comprises: A resource acquisition module is configured to acquire resource information of the home device, and adjust the preset cloud model according to the resource information to obtain a local control model; A feature extraction module is configured to obtain voice information and extract voiceprint feature information of the voice information; An authority determination module is configured to determine the user's operation authority based on the voiceprint feature information; a device determination module, configured to determine a target control device targeted by the voice information when the user operation authority is of a specified type; The device control module is configured to adopt the local control model to control the target control device according to the user operation authority.

9. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the control method of the household appliance as claimed in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the control method of the household device according to any one of claims 1 to 7 is implemented.