Intelligent household equipment adaptive switch control method and system based on deep learning
Through deep learning and machine learning technology, users' voiceprint aging trends, correct audio and record command habits, the recognition performance degradation caused by voiceprint aging in smart home voice control is solved, and the stability and accuracy of voice control are improved.
Patent Information
- Application Number
- CN202510486156.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing smart home voice control system, voiceprint aging phenomenon leads to a decrease in the timeliness and accuracy of user voice commands, affecting the stability and effectiveness of long-term use.
Deep learning technology is adopted to collect audio information through IoT sensors, extract voiceprint features after noise filtering, perform identity verification, and compare voiceprint aging features and trend analysis when verification fails, correct audio to improve recognition accuracy, combine machine learning to record user command habits, and dynamically adjust device control.
It reduces the interference of voiceprint aging, improves the stability and effectiveness of voice control, and enhances the accuracy of user identity verification and instruction issuance.
Smart Images

Figure CN120340485A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of smart home voice control technology, and more specifically, to a method and system for adaptive switch control of smart home devices based on deep learning. Background Art
[0002] Smart home refers to a system that interconnects various devices in the home through the Internet of Things technology, and uses mobile phones, voice assistants or automation systems to achieve remote control, status monitoring and intelligent scheduling, thereby improving living convenience and safety. In current technological applications, voice control has become one of the mainstream interaction methods of smart homes due to its outstanding convenience.
[0003] However, there are significant technical bottlenecks in the voiceprint verification process in voice control: the aging of voiceprints will seriously affect the timeliness and accuracy of command execution. The paper "Cross-age Voiceprint Data Collection System and Voiceprint Aging Research" published by the School of Communication and Information Engineering of Shanghai University in April 2024 confirmed through experiments that: as the time span increases, the similarity between the user's real-time voiceprint features and the registered features shows a significant downward trend. This research result shows that the existing voiceprint recognition system will gradually fail with the extension of usage time, which directly leads to the continuous attenuation of the long-term effectiveness of user voice commands. Therefore, how to overcome the recognition performance decline caused by voiceprint aging and ensure the stability and effectiveness of users' long-term use of voice control has become a key technical problem that needs to be solved in the field of smart homes. Summary of the invention
[0004] The present application provides a method and system for adaptive switch control of smart home devices based on deep learning, which can reduce the interference of voiceprint aging and improve the stability and effectiveness of voice control when users use voice to control smart home devices for a long time.
[0005] In a first aspect, the present application provides a method for adaptive switch control of smart home devices based on deep learning. The method can be executed by a network device, or can also be executed by a chip configured in the network device, and the present application does not limit this.
[0006] Specifically, the method includes:
[0007] Collect indoor real-time audio information through IoT smart sensors, filter the indoor real-time audio information for noise, and obtain the homeowner-controlled audio;
[0008] A deep neural network is used to extract the voiceprint features of the householder's control audio and use them for householder identity authentication. When the householder identity authentication fails, historical voiceprint features are obtained to compare voiceprint aging features to obtain voiceprint aging credibility;
[0009] When the voiceprint aging credibility is higher than a preset threshold, historical voiceprint features are obtained for aging trend extraction to obtain a voiceprint aging trend, and the householder control audio is corrected based on the voiceprint aging trend, and the householder identity verification is performed again on the corrected householder control audio;
[0010] Use a machine learning algorithm to record the operation instruction habits of the householder identity corresponding to the householder control audio, and determine the instruction control preference of this householder identity;
[0011] Perform speech text recognition on the corrected householder control audio, and perform on / off control of furniture and equipment based on the recognized text and the instruction control preference.
[0012] Combined with the first aspect, in some implementation manners of the first aspect, filtering the noise of the indoor real-time audio information to obtain the householder control audio specifically includes:
[0013] Obtain the real-time audio information, perform audio frame division on the real-time audio information to obtain a plurality of real-time audio frames;
[0014] Perform frequency domain transformation on each real-time audio frame respectively to obtain a plurality of real-time frequency domain frames;
[0015] Obtain background noise, perform noise spectrum extraction on the background noise to obtain a background noise spectrum;
[0016] Perform background noise filtering on a plurality of real-time frequency domain frames respectively based on the background noise spectrum, and perform inverse frequency domain transformation on the filtered plurality of real-time frequency domain frames to obtain the householder control audio.
[0017] Combined with the first aspect, in some implementation manners of the first aspect, using a deep neural network to extract the voiceprint features of the householder control audio specifically includes:
[0018] Convert the householder control audio into a mel cepstrum, where the mel cepstrum includes a plurality of mel frequency cepstral coefficients and corresponding acquisition times;
[0019] Input the mel cepstrum into a trained deep neural network model to perform forward inference on the mel cepstrum, and the output layer of the deep neural network model outputs a preset number of forward predicted mel frequency cepstral coefficients;
[0020] Compose a multi-dimensional feature vector from the initial mel cepstrum and the forward predicted mel frequency cepstral coefficients according to the time sequence as the voiceprint features of the householder control audio.
[0021] In combination with the first aspect, in some implementations of the first aspect, a machine learning algorithm is used to record the operation instruction habits of the household head identity corresponding to the household head control audio, and determining the instruction control preference of this household head identity specifically includes:
[0022] Obtain the household head identity corresponding to the household head control audio;
[0023] Record each voice control instruction corresponding to this household head identity, extract machine learning behavior features, and use a single-hidden layer neural network model as a deep learning algorithm to perform instruction control preference clustering on the voice control instruction and the machine learning behavior features, so as to obtain the instruction control preference corresponding to this household head identity.
[0024] In combination with the first aspect, in some implementations of the first aspect, the on / off control of furniture devices based on the recognized text and the instruction control preference specifically includes: extracting user instructions based on the recognized text to obtain user control instructions, obtaining the current moment, performing instruction matching on the instruction control preference based on the current moment and the user control instructions, adjusting the instruction intensity according to the matching result, and performing on / off control of home devices according to the adjusted user control instructions.
[0025] In combination with the first aspect, in some implementations of the first aspect, after the corrected household head control audio is verified for the household head identity again, when the household head identity verification fails again, it is determined that the household head control audio fails and the verification record is uploaded to the smart home cloud service center.
[0026] In combination with the first aspect, in some implementations of the first aspect, a multi-channel microphone array is used as the Internet of Things intelligent sensor.
[0027] In a second aspect, the present application provides an adaptive on / off control system for smart home devices based on deep learning, which includes a voice instruction recognition unit, and the voice instruction recognition unit includes:
[0028] An audio processing module, which is used to collect real-time indoor audio information through an Internet of Things intelligent sensor, filter out noise from the real-time indoor audio information, and obtain household head control audio;
[0029] A voice aging recognition module, which is used to extract the voiceprint features of the household head control audio by using a deep neural network and use them for household head identity verification. When the household head identity verification fails, obtain historical voiceprint features for voiceprint aging feature comparison to obtain the voiceprint aging credibility;
[0030] The voice aging recognition module is further configured to, when the voiceprint aging credibility is higher than a preset threshold, obtain historical voiceprint features for aging trend extraction to obtain a voiceprint aging trend, perform audio correction on the household owner control audio based on the voiceprint aging trend, and perform household owner identity verification again on the corrected household owner control audio;
[0031] The voice command recognition module is configured to record the operation command habits of the household owner identity corresponding to the household owner control audio by using a machine learning algorithm, and determine the command control preference of this household owner identity;
[0032] The voice command recognition module is further configured to perform voice text recognition on the corrected household owner control audio, and perform on / off control of furniture devices based on the recognized text and the command control preference.
[0033] In a third aspect, the present application provides a computer terminal device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned method for adaptively controlling the switching of smart home devices based on deep learning.
[0034] In a fourth aspect, the present application provides a computer-readable storage medium, which stores at least one computer program. The computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned method for adaptively controlling the switching of smart home devices based on deep learning.
[0035] The technical solution provided by the embodiments disclosed in the present application has the following beneficial effects:
[0036] In the method and system for adaptively controlling the switching of smart home devices based on deep learning provided by the present application, first, an Internet of Things intelligent sensor is used to collect real-time indoor audio information, and noise filtering is performed on the real-time indoor audio information to obtain household owner control audio; a deep neural network is used to extract the voiceprint features of the household owner control audio and is used for household owner identity verification. When the household owner identity verification fails, historical voiceprint features are obtained for voiceprint aging feature comparison to obtain the voiceprint aging credibility; when the voiceprint aging credibility is higher than a preset threshold, historical voiceprint features are obtained for aging trend extraction to obtain a voiceprint aging trend, audio correction is performed on the household owner control audio based on the voiceprint aging trend, and household owner identity verification is performed again on the corrected household owner control audio; a machine learning algorithm is used to record the operation command habits of the household owner identity corresponding to the household owner control audio, and determine the command control preference of this household owner identity; voice text recognition is performed on the corrected household owner control audio, and on / off control of furniture devices is performed based on the recognized text and the command control preference.
[0037] It can be seen that this application determines the credibility of voiceprint aging by analyzing the offset trend of historical voiceprint features over time in each dimension, thereby judging the possibility that the voiceprint change is natural aging, avoiding misidentification caused by aging, improving the fault tolerance and adaptability of the system to changes in the user's time dimension, using the voiceprint aging trend to correct the current voiceprint features, automatically compensating for the natural drift caused by the user's voiceprint aging, improving the effectiveness of user identity verification and command issuance, matching the text recognition result with the "user habits", dynamically adjusting the executed device commands, improving the depth of semantic understanding of voice control, and enhancing the accuracy and personalized experience of home devices under voice command control.
[0038] In summary, when the user uses voice control of smart home devices for a long time, this application can reduce the interference of voiceprint aging and improve the stability and effectiveness of voice control. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is an exemplary flowchart of a method for adaptive switch control of a smart home device based on deep learning shown in some embodiments of this application;
[0040] Figure 2 is a schematic diagram of the exemplary hardware and / or software of a voice command recognition unit shown in some embodiments of this application;
[0041] Figure 3 is a schematic diagram of the structure of a computer terminal device for implementing a method for adaptive switch control of a smart home device based on deep learning shown in some embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] This application uses an Internet of Things intelligent sensor to collect real-time indoor audio information, filters the noise of the real-time indoor audio information to obtain the householder's control audio; uses a deep neural network to extract the voiceprint features of the householder's control audio and uses them for householder identity verification. When the householder identity verification fails, historical voiceprint features are obtained for voiceprint aging feature comparison to obtain the credibility of voiceprint aging; when the credibility of voiceprint aging is higher than a preset threshold, historical voiceprint features are obtained for aging trend extraction to obtain the voiceprint aging trend, and the householder's control audio is corrected based on the voiceprint aging trend, and the householder identity verification is performed again on the corrected householder's control audio; uses a machine learning algorithm to record the operation instruction habits of the householder identity corresponding to the householder's control audio, and determines the instruction control preference of this householder identity; performs voice text recognition on the corrected householder's control audio, and based on the recognition text and instruction control preference, performs switch control of furniture devices, which can reduce the interference of voiceprint aging and improve the stability and effectiveness of voice control when the user uses voice control of smart home devices for a long time.
[0043] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. Refer to Figure 1 , which is an exemplary flowchart of a method for adaptive switch control of smart home devices based on deep learning according to some embodiments of the present application. The method 100 for adaptive switch control of smart home devices based on deep learning mainly includes the following steps:
[0044] In step S101, the indoor real-time audio information is collected through the Internet of Things intelligent sensor, and the noise of the indoor real-time audio information is filtered to obtain the household control audio.
[0045] Optionally, in some embodiments, a multi-channel microphone array can be used as the Internet of Things intelligent sensor. In some other embodiments, other devices or equipment capable of collecting audio information can also be used, and the present application does not limit this.
[0046] It should be noted that the household control audio is the control audio for home device control obtained by processing and eliminating background noise and enhancing the human voice from the audio collected in the indoor environment. Preferably, in some embodiments, filtering the noise of the indoor real-time audio information to obtain the household control audio specifically includes:
[0047] Obtain the real-time audio information, divide the real-time audio information into audio frames to obtain a plurality of real-time audio frames;
[0048] Perform frequency-domain transformation on each real-time audio frame respectively to obtain a plurality of real-time frequency-domain frames;
[0049] Obtain the background noise, extract the noise spectrum of the background noise to obtain the background noise spectrum;
[0050] Perform background noise filtering on the plurality of real-time frequency-domain frames respectively based on the background noise spectrum, and perform inverse frequency-domain transformation on the filtered plurality of real-time frequency-domain frames to obtain the household control audio.
[0051] When specifically implemented, in the process of dividing the real-time audio information into audio frames to obtain a plurality of real-time audio frames, the continuous audio signal can be divided into short-time segments frame by frame, and each frame is set to 20-30 milliseconds according to the preset denoising depth. In some other embodiments, a Hamming window can also be added as the window function to reduce the spectral leakage in the audio frames.
[0052] Optionally, in some embodiments, fast Fourier transform and inverse Fourier transform can be used for frequency-domain transformation. Among them, other methods or devices capable of converting time-domain signals into frequency-domain signals can also be used, and the present application does not limit this.
[0053] It should be noted that in the process of performing background noise filtering on multiple real-time frequency domain frames based on the background noise spectrum, after taking the amplitude spectrum of the background noise as the background noise spectrum, the amplitude spectrum of the real-time frequency domain frame is subtracted from the background noise spectrum, and the real-time frequency domain frame after subtracting the background noise spectrum is combined based on the original phase spectrum, and then each frame of audio is reconstructed through inverse Fourier transform to obtain the household owner control audio.
[0054] In step S102, a deep neural network is used to extract the voiceprint features of the household owner control audio and used for household owner identity verification. When the household owner identity verification fails, historical voiceprint features are obtained for voiceprint aging feature comparison to obtain the voiceprint aging credibility.
[0055] Optionally, in some embodiments, when the household owner identity verification is successful, speech text recognition is performed on the household owner control audio, and the switching control of furniture equipment is performed based on the recognized text and the instruction control preference.
[0056] It should be noted that the voiceprint features in the present application can be expressed by a voiceprint feature vector with a preset dimension. Each vector dimension of the voiceprint feature vector contains a Mel frequency cepstral coefficient. Specifically, in implementation, according to different household owner identity verification levels, the dimension of the voiceprint feature vector can be set to 192 dimensions, 216 dimensions or 256 dimensions. Optionally, in some embodiments, using a deep neural network to extract the voiceprint features of the household owner control audio specifically includes:
[0057] Converting the household owner control audio into a Mel cepstrum, where the Mel cepstrum contains multiple Mel frequency cepstral coefficients and corresponding acquisition times;
[0058] Inputting the Mel cepstrum into a trained deep neural network model to perform forward inference on the Mel cepstrum, and the output layer of the deep neural network model outputs a preset number of forward predicted Mel frequency cepstral coefficients;
[0059] Composing a multi-dimensional feature vector from the initial Mel cepstrum and the forward predicted Mel frequency cepstral coefficients according to time series as the voiceprint feature of the household owner control audio.
[0060] In specific implementation, the deep neural network model has multiple hidden layers, and each hidden layer has corresponding activation function neurons for non-linear transformation to extract features. Among them, the deep neural network model is trained with multiple long-segment human voice audio data and their corresponding voiceprint features. Part of the data of the long-segment human voice audio data is input into the deep neural network model for feature extraction and forward prediction, so as to output the complete voiceprint feature of this personal voice audio data. Furthermore, the prediction result is compared with the voiceprint feature corresponding to the actual long-segment human voice audio data, and the activation function parameters in the hidden layer of the deep neural network are optimized according to the comparison result and the corresponding loss function until the feature deviation between the voiceprint feature extracted by the deep neural network model and the actual voiceprint feature is lower than the preset threshold, indicating that the training of the deep neural network model is completed.
[0061] Optionally, in some embodiments, the voiceprint feature extracted in real time can be compared with the registered household head voiceprint template. The household head voiceprint template is a fixed vector generated by training the audio sample provided by the household head during the first verification, and usually includes features such as the user's timbre, accent, and pronunciation habits. In specific implementation, the cosine similarity between the voiceprint feature extracted in real time and the registered household head voiceprint template can be calculated. When the cosine similarity is higher than the preset threshold, it is determined that the household head identity verification is successful; otherwise, it is determined that the household head identity verification fails.
[0062] It should be noted that the voiceprint aging credibility is used to quantify the probability of the household head's voiceprint aging drift over time, and is determined according to the correlation of the feature deviation of the voiceprint feature corresponding to the household head in each dimension over time. It should be noted that, optionally, in some embodiments, obtaining historical voiceprint features for voiceprint aging feature comparison to obtain the voiceprint aging credibility specifically includes:
[0063] Obtain multiple historical voiceprint features and construct a historical voiceprint feature matrix based on the multiple historical voiceprint features;
[0064] Collect voiceprint deviation sequences for the historical voiceprint feature matrix to obtain multiple voiceprint deviation sequences;
[0065] Obtain the acquisition time difference sequence corresponding to multiple historical voiceprint features, and extract the voiceprint aging correlation for each voiceprint deviation sequence respectively based on the acquisition time difference sequence to obtain multiple voiceprint aging correlations;
[0066] Determine the voiceprint aging credibility based on the mean value of multiple voiceprint aging correlations.
[0067] It should be noted that the historical voiceprint features are the voiceprint features when the household head's identity verification is successful in the past period of time, and forward differences are performed based on the acquisition times of the respective historical voiceprint features to obtain an acquisition time difference sequence, where the element value in the acquisition time difference sequence is the time distance between the acquisition times of adjacent historical voiceprint features.
[0068] In specific implementation, for example, when using a 192-dimensional voiceprint feature vector as the voiceprint feature, each voiceprint feature vector is used as a row vector of the historical voiceprint feature matrix, and then different multiple historical voiceprint features are combined to form a historical voiceprint feature matrix. Among them, a voiceprint deviation sequence is formed by the vector value deviations of different voiceprint feature vectors in different dimensions according to time sequence. In specific implementation, the forward difference sequence of each column vector of the historical voiceprint feature matrix can be used as one of the voiceprint deviation sequences, and the number of the voiceprint deviation sequences is the same as the dimension of the voiceprint feature vector.
[0069] Optionally, in some embodiments, after normalizing each voiceprint deviation sequence based on the sequence maximum value, the Pearson correlation coefficient between the voiceprint deviation sequence and the acquisition time difference sequence is used as the voiceprint aging correlation degree corresponding to this voiceprint deviation sequence, and the correlation degree mean value of the voiceprint aging correlation degrees respectively corresponding to each voiceprint deviation sequence is used as the voiceprint aging credibility.
[0070] In step S103, when the voiceprint aging credibility is higher than a preset threshold, historical voiceprint features are obtained for aging trend extraction to obtain a voiceprint aging trend, and the household head control audio is corrected based on the voiceprint aging trend, and the household head identity verification is performed again on the corrected household head control audio.
[0071] Preferably, in some embodiments, obtaining historical voiceprint features for aging trend extraction to obtain a voiceprint aging trend specifically includes:
[0072] Obtaining the voiceprint deviation sequences respectively corresponding to different voiceprint feature dimensions in the historical voiceprint features;
[0073] Obtaining the acquisition time difference sequence corresponding to the historical voiceprint features, and performing aging amplitude extraction on the voiceprint deviation sequences respectively corresponding to different voiceprint feature dimensions based on the acquisition time difference sequence to obtain the aging amplitude values respectively corresponding to each voiceprint feature dimension;
[0074] Combining the aging amplitude values respectively corresponding to each voiceprint feature dimension according to the dimension order to form a voiceprint aging trend.
[0075] When specifically implemented, for any voiceprint feature dimension, the minimum change rate is determined as the aging amplitude value of this voiceprint feature dimension based on the acquisition time difference value in the acquisition time difference sequence and the corresponding voiceprint deviation value in the voiceprint deviation sequence (i.e., the difference of mel-frequency cepstral coefficients), where the change rate = voiceprint deviation value / corresponding acquisition time difference value.
[0076] Optionally, in some embodiments, specifically correcting the household owner control audio based on the voiceprint aging trend includes:
[0077] Obtain the historical audio interval time;
[0078] Perform spectrum conversion on the household owner control audio to obtain the mel cepstrum corresponding to the household owner control audio;
[0079] Obtain each voiceprint feature dimension and the corresponding aging amplitude value in the voiceprint aging trend, and determine the aging compensation ratio corresponding to each voiceprint feature dimension based on each voiceprint feature dimension, the corresponding aging amplitude value, and the historical audio interval time;
[0080] Based on the aging compensation ratio corresponding to each voiceprint feature dimension, correct the mel cepstrum corresponding to the household owner control audio, and perform inverse spectrum transformation to obtain the corrected household owner control audio.
[0081] Among them, in the process of determining the aging compensation ratio corresponding to each voiceprint feature dimension based on the minimum time interval between all historical voiceprint features and the current moment as the historical audio interval time, the aging compensation ratio corresponding to the voiceprint feature dimension = (1 / the aging amplitude value corresponding to this voiceprint feature dimension) × the ratio between the historical audio interval time and the standard interval time, where after proportionally correcting the mel-frequency cepstral coefficients of each frequency band in the mel cepstrum corresponding to the household owner control audio according to the aging compensation ratio and performing inverse spectrum transformation, the corrected household owner control audio is determined.
[0082] Optionally, in some embodiments, after verifying the identity of the household owner for the corrected household owner control audio again, when the verification of the identity of the household owner fails again, it is determined that the household owner control audio fails and the verification record is uploaded to the smart home cloud service center.
[0083] In step S104, a machine learning algorithm is used to record the operation instruction habits of the identity of the household owner corresponding to the household owner control audio, and the instruction control preference of this household owner identity is determined.
[0084] Preferably, in some embodiments, using a machine learning algorithm to record the operation instruction habits of the identity of the household owner corresponding to the household owner control audio and determine the instruction control preference of this household owner identity specifically includes:
[0085] Obtain the household head identity corresponding to the household head control audio;
[0086] Record each voice control instruction corresponding to this household head identity, extract machine learning behavior features, and use a single hidden layer neural network model as a deep learning algorithm to perform instruction control preference clustering on the voice control instruction and the machine learning behavior features, so as to obtain the instruction control preference corresponding to this household head identity.
[0087] Specifically, when implemented, the voice control instruction includes: household head identity: confirmed by voiceprint recognition; control audio: the original audio of what the user said; timestamp: the occurrence time (accurate to hours or minutes); recognized text: the text content after voice recognition of the audio; controlled device: the object of this control, such as "bedroom light", "air conditioner"; execution status: whether the execution is successful, the device response status. Among them, the machine learning behavior features include: time features: the time period when the control occurs, day of the week, whether it is a holiday, etc.; semantic features: instruction keywords (such as "turn on", "turn off"); user preferences: the usage frequency of household devices, instruction expression methods; voice style: fast or slow speech rate, emotional features (optional, combined with acoustic features).
[0088] Optionally, in some embodiments, the user operation records sorted by time (voice text, controlled device, time, etc.) are used as the model input of the single hidden layer neural network model. Among them, the user operation records can be stored in the form of data points. Each data point is composed of time, voice text (vectorized), controlled device, and control parameters. It is necessary to convert non-numerical inputs (such as voice text, device name) into vectors: text instruction → encoded with a custom instruction, controlled device / parameter → use one-hot encoding, time feature → discrete time period (morning, afternoon, evening) / periodic sin-cos encoding. Finally, the encodings are spliced into a fixed-length vector as the input of each time step to the single hidden layer neural network model. The single hidden layer of the single hidden layer neural network model performs preference learning on the instruction type, corresponding instruction time, and instruction intensity through a built-in activation function, and outputs the clustering result through the output layer of the single hidden layer neural network model as the instruction control preference corresponding to this household head identity.
[0089] In step S105, perform voice text recognition on the corrected household head control audio, and perform on-off control of furniture devices based on the recognized text and the instruction control preference.
[0090] Optionally, in some embodiments, a voice recognition model is used to perform voice text recognition on the corrected household head control audio to obtain the recognized text. Specifically, when implemented, a commercial voice recognition model in the prior art, such as the Baidu UNIT voice recognition model, can be used for voice text recognition. This application does not make any limitations in this regard.
[0091] Optionally, in some embodiments, the on / off control of the furniture device based on the recognized text and the instruction control preference specifically includes:
[0092] Extract user instructions based on the recognized text to obtain user control instructions, obtain the current moment, perform instruction matching on the instruction control preference based on the current moment and the user control instructions, adjust the instruction intensity according to the matching result, and perform on / off control of the home device according to the adjusted user control instructions.
[0093] Specifically, the instruction control preference includes the clustering results of multiple historical instructions of the household head. For example, the clustering result of the household head includes the instruction to turn on the lights from 6:00 pm to 6:30 pm, and the voice instruction is used to control and adjust the light brightness to the 80% gear (instruction intensity). According to this clustering result, the user instruction is filled with slots, so as to adjust the instruction intensity of the household head. When the user issues an instruction to turn on the lights from 6:00 pm to 6:30 pm, the instruction intensity is default adjusted to 80%, so as to realize the intention recognition and instruction slot filling of the user instruction.
[0094] In addition, on the other hand of the present application, in some embodiments, the present application provides an adaptive on / off control system for smart home devices based on deep learning. The device includes a voice instruction recognition unit. Refer to Figure 2 , this figure is a schematic diagram of the exemplary hardware and / or software of the voice instruction recognition unit shown according to some embodiments of the present application. The voice instruction recognition unit 200 includes: an audio processing module 201, a voice aging recognition module 202, and a voice instruction recognition module 203, which are described as follows:
[0095] The audio processing module 201 is used to collect real-time indoor audio information through an Internet of Things intelligent sensor, filter noise from the real-time indoor audio information, and obtain the household head control audio;
[0096] The voice aging recognition module 202 is used to extract the voiceprint features of the household head control audio by using a deep neural network and use them for household head identity verification. When the household head identity verification fails, historical voiceprint features are obtained for voiceprint aging feature comparison to obtain the voiceprint aging credibility;
[0097] The voice aging recognition module 202 is further used to, when the voiceprint aging credibility is higher than a preset threshold, obtain historical voiceprint features for aging trend extraction to obtain a voiceprint aging trend, perform audio correction on the household head control audio based on the voiceprint aging trend, and perform household head identity verification on the corrected household head control audio again;
[0098] The voice command recognition module 203 is used to record the operation instruction habits of the household head's identity corresponding to the household head control audio by using a machine learning algorithm, and determine the command control preference of this household head identity;
[0099] The voice command recognition module 203 is further used to perform voice text recognition on the corrected household head control audio, and perform on / off control of furniture devices based on the recognized text and the command control preference.
[0100] The above has introduced in detail an example of a method and system for adaptive on / off control of smart home devices based on deep learning provided by the embodiments of the present application. It can be understood that, in order to implement the above functions, the corresponding device includes the corresponding hardware structure and / or software module for executing each function.
[0101] Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function in the application is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraint conditions of the technical solution. Therefore, professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the present application.
[0102] In addition, the present application also provides a computer terminal device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned method for adaptive on / off control of smart home devices based on deep learning.
[0103] In some embodiments, refer to Figure 3 , this figure is a schematic structural diagram of a computer terminal device applying a method for adaptive on / off control of smart home devices based on deep learning according to some embodiments of the present application. The above-mentioned method for adaptive on / off control of smart home devices based on deep learning in the above embodiments can be implemented by Figure 3 The shown computer terminal device. The computer terminal device 300 includes at least one communication bus 301, a communication interface 302, a processor 303, and a memory 304.
[0104] The processor 303 can be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more for controlling the execution of the method for adaptive on / off control of smart home devices based on deep learning in the present application.
[0105] The communication bus 301 may include a path for transmitting information between the above components.
[0106] The memory 304 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 304 may exist independently and be connected to the processor 303 through the communication bus 301. The memory 304 may also be integrated with the processor 303.
[0107] Among them, the memory 304 is used to store the program code for executing the solution of this application and is controlled by the processor 303 for execution. The processor 303 is used to execute the program code stored in the memory 304. The program code may include one or more software modules. The determination of the voiceprint aging credibility in the above embodiments may be implemented by one or more software modules in the program code in the processor 303 and the memory 304.
[0108] The communication interface 302 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0109] Optionally, the above computer terminal device 300 may further include a power supply 305 for supplying power to various components or circuits in the real-time computer terminal device.
[0110] In a specific implementation, as an example, a computer terminal device may include multiple processors, and each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0111] The above computer terminal device may be a general-purpose computer terminal device or a special-purpose computer terminal device. In a specific implementation, the computer terminal device may be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer terminal device.
[0112] In addition, in other aspects of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores at least one computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned method for adaptively controlling the switch of a smart home device based on deep learning.
[0113] In summary, in the method and system for adaptively controlling the switch of a smart home device based on deep learning disclosed in the embodiments of the present application, first, an Internet of Things intelligent sensor is used to collect real-time indoor audio information, and the real-time indoor audio information is filtered for noise to obtain the householder's control audio; a deep neural network is used to extract the voiceprint features of the householder's control audio and used for householder identity verification. When the householder identity verification fails, historical voiceprint features are obtained for voiceprint aging feature comparison to obtain the voiceprint aging credibility; when the voiceprint aging credibility is higher than a preset threshold, historical voiceprint features are obtained for aging trend extraction to obtain the voiceprint aging trend, and the householder's control audio is corrected based on the voiceprint aging trend, and the householder identity verification is performed again on the corrected householder's control audio; a machine learning algorithm is used to record the operation instruction habits of the householder identity corresponding to the householder's control audio to determine the instruction control preference of this householder identity; the corrected householder's control audio is subjected to speech text recognition, and the switch control of the furniture device is performed based on the recognized text and the instruction control preference, which can reduce the interference of voiceprint aging and improve the stability and effectiveness of voice control when the user uses voice to control smart home devices for a long time.
[0114] The above are only embodiments of the present application, and common general technical solutions or characteristics in the solutions are not described in detail herein. It should be noted that for those skilled in the art, without departing from the technical solutions of the present application, several modifications and improvements can still be made, which should also be regarded as the protection scope of the present application, and these will not affect the implementation effect of the present application and the practicality of the patent.
[0115] The protection scope required by the present application shall be subject to the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. An adaptive switch control method for smart home devices based on deep learning, characterized in that Including: Collecting real-time indoor audio information through Internet of Things intelligent sensors, filtering noise from the real-time indoor audio information to obtain the householder's control audio; Using a deep neural network to extract the voiceprint features of the householder's control audio for householder identity verification. When the householder identity verification fails, obtaining historical voiceprint features for voiceprint aging feature comparison to obtain the voiceprint aging credibility; When the voiceprint aging credibility is higher than a preset threshold, obtaining historical voiceprint features for aging trend extraction to obtain the voiceprint aging trend, performing audio correction on the householder's control audio based on the voiceprint aging trend, and performing householder identity verification again on the corrected householder's control audio; Using a machine learning algorithm to record the operation instruction habits of the householder identity corresponding to the householder's control audio, and determining the instruction control preference of this householder identity; Performing speech text recognition on the corrected householder's control audio, and performing on / off control of furniture and equipment based on the recognized text and the instruction control preference.
2. The method according to claim 1, wherein Filtering noise from the real-time indoor audio information to obtain the householder's control audio specifically includes: Obtaining the real-time audio information, performing audio frame division on the real-time audio information to obtain multiple real-time audio frames; Performing frequency domain transformation on each real-time audio frame respectively to obtain multiple real-time frequency domain frames; Obtaining background noise, performing noise spectrum extraction on the background noise to obtain the background noise spectrum; Performing background noise filtering on multiple real-time frequency domain frames respectively based on the background noise spectrum, and performing inverse frequency domain transformation on the filtered multiple real-time frequency domain frames to obtain the householder's control audio.
3. The method according to claim 1, wherein Using a deep neural network to extract the voiceprint features of the householder's control audio specifically includes: Converting the householder's control audio into a Mel cepstrum, where the Mel cepstrum includes multiple Mel frequency cepstral coefficients and corresponding acquisition times; Inputting the Mel cepstrum into a trained deep neural network model to perform forward inference on the Mel cepstrum, and the output layer of the deep neural network model outputs a preset number of forward predicted Mel frequency cepstral coefficients; Composing a multi-dimensional feature vector from the initial Mel cepstrum and the forward predicted Mel frequency cepstral coefficients according to time sequence as the voiceprint features of the householder's control audio.
4. The method according to claim 1, wherein Using a machine learning algorithm to record the operation instruction habits of the householder identity corresponding to the householder's control audio, and determining the instruction control preference of this householder identity specifically includes: Obtaining the householder identity corresponding to the householder's control audio; Recording each voice control instruction corresponding to this householder identity, extracting machine learning behavior features, and using a single hidden layer neural network model as a deep learning algorithm to perform instruction control preference clustering on the voice control instruction and the machine learning behavior features to obtain the instruction control preference corresponding to this householder identity.
5. The method according to claim 1, wherein The on / off control of furniture equipment based on the recognized text and the instruction control preference specifically includes: extracting user instructions based on the recognized text to obtain user control instructions, acquiring the current moment, performing instruction matching on the instruction control preference based on the current moment and the user control instructions, adjusting the instruction intensity according to the matching result, and performing on / off control of home equipment according to the adjusted user control instructions.
6. The method according to claim 1, wherein After re-verifying the household head control audio after correction, when the verification of the household head's identity fails again, it is determined that the household head control audio is invalid and the verification record is uploaded to the smart home cloud service center.
7. The method according to claim 1, wherein A multi-channel microphone array is used as the Internet of Things intelligent sensor.
8. An adaptive switch control system for smart home devices based on deep learning, including a voice command recognition unit, characterized in that, The voice instruction recognition unit includes: An audio processing module for collecting real-time indoor audio information through the Internet of Things intelligent sensor, filtering out noise from the real-time indoor audio information to obtain the household head control audio; A voice aging recognition module for extracting the voiceprint features of the household head control audio using a deep neural network and using them for household head identity verification. When the household head identity verification fails, historical voiceprint features are obtained for voiceprint aging feature comparison to obtain the voiceprint aging credibility; The voice aging recognition module is further configured to, when the voiceprint aging credibility is higher than a preset threshold, obtain historical voiceprint features for aging trend extraction to obtain a voiceprint aging trend, perform audio correction on the household head control audio based on the voiceprint aging trend, and re-verify the identity of the household head for the corrected household head control audio; A voice instruction recognition module for recording the operation instruction habits of the household head corresponding to the household head control audio using a machine learning algorithm to determine the instruction control preference of this household head identity; The voice instruction recognition module is further configured to perform voice text recognition on the corrected household head control audio, and perform on / off control of furniture equipment based on the recognized text and the instruction control preference.
9. A computer terminal device, characterized in that, The computer terminal device includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute a method for adaptive on / off control of smart home devices based on deep learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing at least one computer program, characterized in that, The computer program is loaded and executed by the processor to implement the operations performed by a method for adaptive on / off control of smart home devices based on deep learning as described in any one of claims 1 to 7.