Home system awakening method based on artificial intelligence

The method uses frequency spectrum analysis and a target Hidden Markov Model to filter noise and enhance voice recognition in smart home systems, addressing environmental interference and improving wake-up reliability and user experience.

CN120321061AInactive Publication Date: 2025-07-15GUANGXI DIRICO INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510523626.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing smart home systems are easily disturbed by environmental noise during the wake-up process, resulting in failure in wake-up and poor user experience.

Method used

Through spectrum analysis, the environmental noise interference is removed, the target hidden Markov model is used to extract vocals, and the device type is recognized by the preset voice feature library to achieve accurate wake-up.

Benefits of technology

Effectively overcome environmental noise interference, improve wake-up success rate, ensure accurate execution of user instructions, and significantly improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321061A_ABST
    Figure CN120321061A_ABST
Patent Text Reader

Abstract

The invention provides a home system wake-up method based on artificial intelligence, and the method comprises the steps: carrying out the spectrum analysis of sound wake-up data in a home environment, and generating initial human voice data; the frequency spectrum analysis utilizes sound frequency and intensity characteristics to separate human voice from environment noise on a frequency spectrum level, and unmatched noise signals are preliminarily filtered out, so that interference of the environment noise on wake-up instruction identification is reduced. The target hidden Markov model carries out deep mining on the initial human voice data and extracts human voice signals based on the time sequence probability distribution characteristics of human voice, and even if the human voice is fuzzy or distorted due to the influence of noise, the wake-up instruction can be accurately identified and further purified. And carrying out equipment type identification on the target human voice data and a preset voice feature library, generating equipment type data, awakening corresponding intelligent equipment in the home environment according to the equipment type data, and generating home system awakening data. The false wakeup or wakeup failure of the equipment caused by noise interference is avoided, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart home, and particularly to a method for waking up a home system based on artificial intelligence. Background Art

[0002] A smart home system utilizes advanced computer technology, network communication technology, intelligent cloud control, integrated wiring technology, and medical electronics technology in accordance with the principles of ergonomics, integrating individual needs, and organically combining various subsystems related to home life. Through networked integrated intelligent control and management, a brand-new home life experience is realized. The characteristic of a smart home system is that it can achieve voice control. In order to achieve the purpose of energy conservation and prevent false triggering, the user needs to wake up the voice system first before interacting with the user's voice.

[0003] Currently, a large number of smart home products have emerged on the market. The core goal of smart home products is to continuously improve the user experience. Whether a good user experience can be created for users has a direct impact on brand competitiveness. However, there are prominent problems in the wake-up link of existing smart devices. In daily life, the ambient sound is complex and variable, such as the sound of TV programs, the sound of cars outside the window, the conversation sound of family members, etc., which will interfere with the voice wake-up function of smart devices. When the voice information received by the device is deviated, it is easy to cause wake-up failure, resulting in many inconveniences for users during use and a poor user experience. Summary of the Invention

[0004] In view of this, the present invention proposes a method for waking up a home system based on artificial intelligence. By using spectrum analysis to remove the interference of other noises in the home environment on the voice wake-up data, and extracting human voices through a target hidden Markov model to improve the accuracy of human voice recognition, thereby avoiding deviation of the voice information received by the device, enabling the device to respond in a timely manner, and improving the user experience.

[0005] The technical solution of the present invention is realized as follows:

[0006] A method for waking up a home system based on artificial intelligence, comprising the following steps:

[0007] Obtain voice wake-up data in the home environment, perform spectrum analysis on the voice wake-up data to generate initial human voice data;

[0008] Use a target hidden Markov model to extract human voices from the initial human voice data to generate target human voice data;

[0009] Perform device type recognition on the target human voice data and a preset voice feature library to generate device type data;

[0010] Wake up the corresponding intelligent device in the home environment according to the device type data, and generate home system wake-up data.

[0011] Preferably, the spectral analysis of the voice wake-up data to generate the initial voice data includes:

[0012] Perform short-time Fourier transform on the voice wake-up data to generate initial voice time-frequency data;

[0013] Use a preset time-frequency filter to filter the initial voice time-frequency data to generate target voice time-frequency data;

[0014] Perform inverse short-time Fourier transform on the target voice time-frequency data to generate initial voice data.

[0015] Preferably, the short-time Fourier transform of the voice wake-up data to generate the initial voice time-frequency data includes:

[0016] Use the first short-time Fourier transform formula to filter out noise from the continuous-time signal in the voice wake-up data to generate the first voice time-frequency data;

[0017] The first short-time Fourier transform formula is:

[0018]

[0019] Where X improved (τ, ω) represents the spectral value at time position τ and angular frequency ω; x(t) is the original continuous-time audio signal, that is, the continuous-time signal in the voice wake-up data; ω adaptive (t, τ) is the adaptive window function; e -jωt Is the complex exponential function in the continuous signal, used to convert the time-domain signal to the frequency domain, where j is the imaginary unit, ω is the angular frequency, and t is the time variable;

[0020] Use the second short-time Fourier transform formula to filter out noise from the discrete-time signal in the voice wake-up data to generate the second voice time-frequency data;

[0021] The second short-time Fourier transform formula is:

[0022]

[0023] Where X improved[m, k] is a two-dimensional array representing the spectral values at discrete time position m and discrete frequency index k; x[n] is the original discrete-time audio signal, i.e., the discrete-time signal in the voice wake-up data, where n represents the discrete-time index corresponding to the sampling moment of the continuous-time signal; ω[n - m] is a discrete window function used to window the discrete signal x[n]. is a complex exponential function in the discrete signal, used to transform the discrete-time domain signal to the discrete frequency domain, where j is the imaginary unit. is the frequency resolution, k is the discrete frequency index, n is the discrete-time index, N is the number of points of the fast Fourier transform; c[m, k] is the noise compensation factor.

[0024] Using the first voice time-frequency data and the second voice time-frequency data, the initial voice time-frequency data.

[0025] Preferably, the filtering the initial voice time-frequency data using a preset time-frequency filter to generate target voice time-frequency data includes:

[0026] Calculating the average value using the noise spectral data corresponding to the home environment to generate an energy distribution estimate value.

[0027] Calculating the signal-to-noise ratio estimate values of each time-frequency point corresponding to the voice wake-up data using the voice time-frequency data and the energy distribution estimate value respectively to generate a plurality of signal-to-noise ratio estimate values.

[0028] Calculating the filter coefficients corresponding to the time-frequency points respectively using the signal-to-noise ratio estimate values and a preset target signal-to-noise ratio to generate the filter coefficients corresponding to the time-frequency points.

[0029] Selecting the filter coefficients greater than a preset coefficient threshold to construct a time-frequency data set.

[0030] Performing a multiplication calculation on the time-frequency data in the time-frequency data set and the corresponding filter data respectively to generate target voice time-frequency data.

[0031] Preferably, before the step of extracting the human voice from the initial human voice data using the target hidden Markov model to generate target human voice data, it further includes:

[0032] Extracting features from the human voice data and non-human voice data corresponding to the home environment using Mel-frequency cepstral coefficients to generate audio feature data.

[0033] Initializing the initial hidden Markov model according to preset initialization data to generate an intermediate hidden Markov model.

[0034] Training the intermediate hidden Markov model using the expectation-maximization algorithm to generate a target hidden Markov model.

[0035] Preferably, the step of using the target Hidden Markov Model to extract human voice data from the initial human voice data to generate target human voice data includes:

[0036] Extracting features in the initial human voice data that are the same as the training features corresponding to the target Hidden Markov Model to generate a human voice feature sequence;

[0037] Performing state estimation on the human voice feature sequence by using the Viterbi algorithm through the target Hidden Markov Model to generate a state sequence;

[0038] Extracting human voice segments from the state sequence to generate human voice data;

[0039] Using spectral subtraction to remove background data from the human voice data to generate initial human voice data;

[0040] Using an adaptive filtering method to remove reverberation from the initial human voice data to generate intermediate human voice data;

[0041] Normalizing the volume of the intermediate human voice data to generate target human voice data.

[0042] Preferably, the step of performing device type recognition on the target human voice data with a preset voice feature library to generate device type data includes:

[0043] Screening out voiceprints in the preset voice feature library that match the target human voice data to generate candidate voiceprints;

[0044] Using a cosine similarity improvement algorithm to calculate the similarity between the target human voice data and the candidate voiceprints respectively to generate a target cosine similarity;

[0045] Performing device type recognition based on the target cosine similarity to generate device type data.

[0046] Preferably, the step of performing device type recognition based on the target cosine similarity to generate device type data includes:

[0047] Judging whether the cosine similarity is greater than a preset similarity threshold;

[0048] If so, the voiceprint verification passes, and device keyword recognition is performed on the target human voice data to generate device type data;

[0049] If not, the voiceprint verification fails, and a preset alarm is activated to play preset voice prompt data.

[0050] Preferably, the step of performing device keyword recognition on the target human voice data to generate device type data includes:

[0051] Input the target human voice data into a convolutional layer for convolution operation to generate an initial feature map;

[0052] Input the initial feature map into a pooling layer for downsampling to generate a target feature map;

[0053] Input the target feature map into a recurrent layer to capture sequential data and generate sequential features;

[0054] Input the sequential features into an attention layer for feature weighting to generate feature vectors;

[0055] Input the feature vectors into a fully connected layer for device type recognition to generate device type data.

[0056] Preferably, after the step of waking up the corresponding intelligent device in the home environment according to the device type data to generate home system wake-up data, the following steps are further included:

[0057] Divide the device response current values corresponding to the home system wake-up data according to a preset time period to generate multiple current value sets;

[0058] Calculate the mean values of the current value sets respectively to generate multiple current means;

[0059] Compare the current means with a preset mean threshold respectively to generate current comparison data;

[0060] Compare the device temperature change data in the device response data with a preset temperature threshold to generate temperature comparison data;

[0061] When any one of the current comparison data and the temperature comparison data is greater than the preset alarm data, send the alarm data to a preset alarm device.

[0062] Compared with the prior art, the beneficial effects of the present invention are:

[0063] A method for waking up a home system based on artificial intelligence according to the present invention collects voice wake-up data in the home environment, performs spectral analysis on the voice wake-up data to generate initial human voice data. The spectral analysis utilizes the characteristics of voice frequency and intensity to separate human voices from environmental noise at the spectral level, preliminarily filtering out mismatched noise signals and reducing the interference of environmental noise on the recognition of wake-up instructions. The target hidden Markov model deeply mines and extracts human voice signals based on the probability distribution characteristics of the time series of human voices. Even if the human voice is blurred or distorted due to noise interference, it can accurately identify and further purify the wake-up instructions. Comparing the processed target human voice data with a preset voice feature library can accurately identify the device type, avoiding incorrect wake-up or wake-up failure of the device caused by noise interference. Through this series of progressive processes, this technical solution can effectively overcome environmental noise interference, greatly improve the wake-up success rate of the home system, ensure the accurate execution of user instructions, significantly improve the user experience, and enable the smart home system to operate stably and efficiently in a noisy environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only the preferred embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0065] Figure 1 It is a flowchart of the steps of a method for waking up a home system based on artificial intelligence provided in Embodiment 1 of the present invention;

[0066] Figure 2 It is a flowchart of the steps of another method for waking up a home system based on artificial intelligence provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The embodiments of the present invention provide a method for waking up a home system based on artificial intelligence, which is used to solve the technical problem that the existing home system is easily interfered by environmental noise during the wake-up process, resulting in wake-up failure and poor user experience.

[0068] In order to make the object, features, and advantages of the present invention more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the following described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0069] Please refer to Figure 1, Figure 1 It is a flowchart of the steps of a method for waking up a home system based on artificial intelligence provided in the first embodiment of the present invention.

[0070] A method for waking up a home system based on artificial intelligence provided by the present invention includes:

[0071] Step 101: Obtain the voice wake-up data in the home environment, perform spectral analysis on the voice wake-up data, and generate initial human voice data.

[0072] Step 102: Use the target hidden Markov model to extract the human voice from the initial human voice data and generate target human voice data.

[0073] Step 103: Identify the device type by comparing the target human voice data with the preset voice feature library, and generate device type data.

[0074] Step 104: Wake up the corresponding intelligent device in the home environment according to the device type data, and generate home system wake-up data.

[0075] Among them, the voice wake-up data includes the complete waveform data of the audio collected in the home environment, the change of the energy magnitude of the sound at different times, and the relative position relationship of the sound signal at different times.

[0076] The target hidden Markov model refers to the hidden Markov model trained by the expectation maximization algorithm.

[0077] The preset voice feature library refers to the feature library that stores in advance the voice information of each user in the home environment and one or more wake-up words corresponding to waking up each intelligent device.

[0078] In the embodiment of the present invention, after obtaining the voice wake-up data in the home environment, by performing spectral analysis on the voice wake-up data, the background noise in the voice wake-up data is effectively removed, the voice features are accurately extracted, and the accuracy of the obtained initial human voice data is high. The target hidden Markov model is used to extract the human voice from the initial human voice data, considering the correlation between the front and back frames of the voice, and capturing the dynamic change law of the human voice by modeling the hidden state, so as to more accurately extract the human voice from the complex audio environment. By identifying the device type by comparing the target human voice data with the preset voice feature library, the intelligent device to be woken up by the voice wake-up data is determined, and device type data is generated. Finally, the corresponding intelligent device in the home environment is woken up according to the device type data, and home system wake-up data is obtained. By using spectral analysis and the target hidden Markov model to remove other noise interferences in the environment, the corresponding intelligent device is quickly woken up, improving the user experience, thereby solving the technical problem that the existing home system is prone to be interfered by environmental noise during the wake-up process, resulting in wake-up failure and poor user experience.

[0079] Please refer to Figure 2 , Figure 2 which is a flowchart of steps of another method for waking up a home system based on artificial intelligence provided in the second embodiment of the present invention.

[0080] Another method for waking up a home system based on artificial intelligence provided by the present invention includes:

[0081] Step 201, obtain voice wake-up data in the home environment, perform spectral analysis on the voice wake-up data, and generate initial human voice data.

[0082] Optionally, step 201 may include the following sub-steps S11-S13:

[0083] S11, perform short-time Fourier transform on the voice wake-up data to generate initial voice time-frequency data;

[0084] S12, filter the initial voice time-frequency data using a preset time-frequency filter to generate target voice time-frequency data;

[0085] S13, perform inverse short-time Fourier transform on the target voice time-frequency data to generate initial human voice data.

[0086] The preset time-frequency filter refers to a filter set according to the spectral characteristics of the user's voice and the voice that the intelligent device can receive and respond to. Among them, the type of the filter can be a low-pass filter, a high-pass filter, or a band-pass filter.

[0087] In the embodiment of the present invention, perform short-time Fourier transform on the voice wake-up data to convert the time-varying voice waveform data into a representation in the frequency domain, so that the frequency components of the voice can be clearly displayed. Input the initial voice time-frequency data obtained by the short-time Fourier transform into the preset time-frequency filter for filtering to remove noise in other frequency bands and obtain the target voice time-frequency data. Perform inverse short-time Fourier transform on the denoised target voice time-frequency data to convert it back to the time domain to obtain the denoised voice signal frame. After the inverse short-time Fourier transform, the obtained is the time-domain signal of each frame, and these signals need to be spliced and processed to restore the complete denoised voice signal to obtain the initial human voice data.

[0088] Further, S11 may include the following sub-steps S111-S113:

[0089] S111, filter out noise from the continuous-time signal in the voice wake-up data using the first short-time Fourier transform formula to generate first voice time-frequency data;

[0090] S112, filter out noise from the discrete-time signal in the voice wake-up data using the second short-time Fourier transform formula to generate second voice time-frequency data;

[0091] S113. Use the first audio time-frequency data and the second audio time-frequency data to construct initial audio time-frequency data.

[0092] The first short-time Fourier transform formula is:

[0093]

[0094] where X improved (τ, ω) represents the spectral value at time position τ and angular frequency ω; x(t) is the original continuous-time audio signal, that is, the continuous-time signal in the voice wake-up data; ω adaptive (t, τ) is the adaptive window function; e -jωt is the complex exponential function in the continuous signal, used to transform the time-domain signal to the frequency domain, where j is the imaginary unit, ω is the angular frequency, and t is the time variable.

[0095] The second short-time Fourier transform formula is:

[0096]

[0097] where X improved [m, k] is a two-dimensional array representing the spectral value at discrete time position m and discrete frequency index k; x[n] is the original discrete-time audio signal, that is, the discrete-time signal in the voice wake-up data, n represents the discrete-time index, corresponding to the sampling moment of the continuous-time signal; ω[n - m] is the discrete window function, used to window the discrete signal x[n]; is the complex exponential function in the discrete signal, used to transform the discrete time-domain signal to the discrete frequency domain, where j is the imaginary unit, is the frequency resolution, k is the discrete frequency index, n is the discrete time index, N is the number of points of the fast Fourier transform; c[m, k] is the noise compensation factor.

[0098] In the embodiment of the present invention, first, use the first short-time Fourier transform formula to filter the noise of the continuous-time signal in the voice wake-up data to obtain the first audio time-frequency data. Secondly, use the second short-time Fourier transform formula to filter the noise of the discrete-time signal in the voice wake-up data to obtain the second audio time-frequency data. After all the noise of the voice wake-up data is filtered, use the first audio time-frequency data and the second audio time-frequency data to construct the initial audio time-frequency data.

[0099] Further, S12 may include the following sub-steps S121 - S125:

[0100] S121. Calculate the average value using the noise spectral data corresponding to the home environment to generate an estimated energy distribution;

[0101] S122. Calculate the signal-to-noise ratio (SNR) estimates for each time-frequency point corresponding to the voice wake-up data using the voice time-frequency data and the estimated energy distribution respectively, generating multiple SNR estimates.

[0102] S123. Calculate the filter coefficients corresponding to the time-frequency points by using the SNR estimates and a preset target SNR respectively, generating the filter coefficients corresponding to the time-frequency points.

[0103] S124. Select the filter coefficients greater than a preset coefficient threshold to construct a time-frequency data set.

[0104] S125. Perform multiplication calculations on the time-frequency data in the time-frequency data set with the corresponding filter data respectively to generate the target voice time-frequency data.

[0105] The noise spectrum data refers to the silent segment or low-energy segment corresponding to the home environment.

[0106] The preset coefficient threshold refers to the critical value set based on actual needs for screening filter coefficients.

[0107] In the embodiments of the present invention, the average value is calculated using the noise spectrum data corresponding to the home environment to generate the estimated energy distribution. By finding the silent segment or low-energy segment in the voice signal, these regions are mainly composed of noise. The spectrum within these time periods is statistically averaged to obtain the estimated value of the noise spectrum, that is, the estimated energy distribution. The SNR estimates for each time-frequency point corresponding to the voice wake-up data are calculated using the voice time-frequency data and the estimated energy distribution respectively, generating multiple SNR estimates. Based on the noise energy estimation, the relative intensity of the voice signal at each time-frequency point is measured, and calculating multiple SNR estimates can comprehensively describe the signal quality of the voice wake-up data at different time-frequency points. The filter coefficients corresponding to the time-frequency points are calculated by using the SNR estimates and a preset target SNR respectively. By adjusting the filter coefficients, corresponding filtering processing can be performed on different time-frequency points according to the SNR situation of the signal, so as to achieve the purpose of enhancing the signal and suppressing the noise. And in the present invention, the filter coefficients greater than the preset coefficient threshold are selected to construct a time-frequency data set, screening out the time-frequency points that have a significant effect on signal enhancement and noise suppression, and removing those time-frequency points that may be more affected by noise or contribute less to the target voice, thereby improving the quality and effectiveness of the data. By performing multiplication calculations on the time-frequency data in the time-frequency data set with the corresponding filter data respectively, the target voice time-frequency data is generated, thereby highlighting the characteristics of the target voice, suppressing the noise, and obtaining a purer target voice time-frequency data.

[0108] Step 202. Perform voice extraction on the initial voice data using the target hidden Markov model to generate the target voice data.

[0109] Optionally, the training process of the target Hidden Markov Model may include the following sub-steps S21 - S23:

[0110] S21. Extract features from the voice data and non-voice data corresponding to the home environment using Mel Frequency Cepstral Coefficients to generate audio feature data;

[0111] S22. Initialize the initial Hidden Markov Model according to the preset initialization data to generate an intermediate Hidden Markov Model;

[0112] S23. Use the Expectation-Maximization algorithm to train the intermediate Hidden Markov Model to generate the target Hidden Markov Model.

[0113] The preset initialization data refers to the number of states and the meaning of states of the target Hidden Markov Model set based on actual needs, which can usually be divided into voice states and non-voice states, and can also be further subdivided. Randomly initialize the three core parameters of the target Hidden Markov Model: the initial state probability distribution, the state transition probability matrix, and the observation probability matrix.

[0114] The voice data and non-voice data corresponding to the home environment cover different places, speakers, speech contents, etc. in the home environment to ensure that the model has good generalization ability.

[0115] In the embodiment of the present invention, first, extract features from the voice data and non-voice data corresponding to the home environment using Mel Frequency Cepstral Coefficients to generate audio feature data. Then, initialize the initial Hidden Markov Model according to the preset initialization data to obtain an intermediate Hidden Markov Model. Finally, use the Expectation-Maximization algorithm to train the intermediate Hidden Markov Model to generate the target Hidden Markov Model. Continuously update the model parameters using the Expectation-Maximization algorithm to maximize the likelihood of the model for the training data. The specific process includes the expectation step, calculating the forward probability and backward probability of each training sample in each state, and then obtaining the expected observation value and state transition expectation of each state; the maximization step, updating the model parameters according to the results of the expectation step, that is, updating the initial state probability distribution, the state transition probability matrix, and the observation probability matrix, and repeating these two steps until the parameters converge or reach the preset number of iterations.

[0116] Optionally, step 202 may include the following sub-steps S31 - S36:

[0117] S31. Extract the features identical to the training features corresponding to the target Hidden Markov Model from the initial voice data to generate a voice feature sequence;

[0118] S32. Use the Viterbi algorithm through the target Hidden Markov Model to perform state estimation on the voice feature sequence to generate a state sequence;

[0119] S33. Extract the human voice segments in the status sequence to generate human voice data;

[0120] S34. Use spectral subtraction to remove the background data in the human voice data to generate initial human voice data;

[0121] S35. Use adaptive filtering to remove the reverberation in the initial human voice data to generate intermediate human voice data;

[0122] S36. Normalize the volume of the intermediate human voice data to generate target human voice data.

[0123] In the embodiment of the present invention, perform the same feature extraction operation on the initial human voice data to be processed as in the training stage to obtain the feature sequence of the initial human voice data. That is, extract the features in the initial human voice data that are the same as the training features corresponding to the target hidden Markov model to generate a human voice feature sequence. Use the Viterbi algorithm through the target hidden Markov model to perform state estimation on the human voice feature sequence to obtain a status sequence. The Viterbi algorithm is a dynamic programming algorithm. It starts from the initial moment, calculates the maximum cumulative probability in each state, and records the previous optimal state of each state. As time progresses, sequentially process the subsequent feature frames until the entire feature sequence of the initial human voice data is processed. Finally, obtain the most likely status sequence corresponding to the entire feature sequence by backtracking the optimal path.

[0124] According to the estimated status sequence, extract the audio part corresponding to the human voice status to generate human voice data. If specific human voice statuses are defined in the model, when the Viterbi algorithm estimates that the audio is in a certain human voice status during a certain time period, mark the audio of that time period as the human voice part. These audio segments marked as human voices can be merged to remove possible non-human voice segments in the middle. Use spectral subtraction to remove the background data in the human voice data to remove the remaining background noise and make the target human voice clearer to obtain the initial human voice data. Use adaptive filtering to remove the reverberation in the initial human voice data to improve the clarity and intelligibility of the sound to obtain the intermediate human voice data. Perform volume normalization on the extracted human voice to obtain the target human voice data to keep the volume of different parts consistent and avoid the situation of sudden volume changes.

[0125] Step 203. Perform device type recognition on the target human voice data and a preset voice feature library to generate device type data.

[0126] Optionally, step 203 may include the following sub-steps S41 - S43:

[0127] S41. Screen out the voiceprints matching the target human voice data from the preset voice feature library to generate candidate voiceprints;

[0128] S42. Use the improved cosine similarity algorithm to calculate the similarity between the target human voice data and the candidate voiceprint respectively, and generate the target cosine similarity;

[0129] S43. Based on the target cosine similarity, perform device type recognition to generate device type data.

[0130] In the embodiment of the present invention, the user voiceprint and the device wake-up word voiceprint that match the target human voice data are screened out from the preset voice feature library to obtain the candidate voiceprint. The improved cosine similarity algorithm used in the present invention is the Mahalanobis-cosine similarity. The Mahalanobis distance takes into account the correlation and variance between features. After calculating the Mahalanobis distance for the voiceprint feature vector, it is weighted and fused with the cosine similarity. For example, for a group of voiceprint feature vectors A and B, first calculate the Mahalanobis distance D M (A, B), and then calculate the cosine similarity Sim C (A, B). The final calculation formula for the target cosine similarity is Sim = ω1×D M (A, B)+ω2×Sim C (A, B), where Sim is the target cosine similarity; ω1 is the first weight; ω2 is the second weight, and the values of the weights are adjusted according to experiments, which can more accurately measure the voiceprint similarity. Finally, device type recognition is performed through the target cosine similarity to determine the intelligent device to be awakened, and device type data is obtained.

[0131] Optionally, step S43 may include the following sub-steps S431-S433:

[0132] S431. Judge whether the cosine similarity is greater than the preset similarity threshold. If so, execute step S432; if not, execute step S433.

[0133] S432. The voiceprint verification is passed, and device keyword recognition is performed on the target human voice data to generate device type data;

[0134] S433. The voiceprint verification fails, and a preset alarm is started to play preset voice prompt data.

[0135] The preset similarity threshold refers to the critical value used to judge whether the cosine similarity meets the requirements.

[0136] The preset alarm can be an external separate alarm or a built-in speaker of the intelligent device.

[0137] In an embodiment of the present invention, when the cosine similarity is greater than a preset similarity threshold, it indicates that the voiceprint verification of the user has passed, preventing malicious use by other users. Then, keyword recognition of the target human voice data is performed to determine the intelligent device to be awakened. When the cosine similarity is less than or equal to the preset similarity threshold, it indicates that the user of the voice wake-up data is not a legitimate user in this home environment. At this time, the voiceprint verification fails, and a preset alarm is activated to play preset voice prompt data. The actual needs and usage scenarios of the user are fully considered to avoid disturbing the user due to alarm prompts, while providing operation guidance, reflecting the user-friendly design of the system; enabling the device to quickly execute user instructions, all of which contribute to improving the user's satisfaction and trust in the smart home system.

[0138] Optionally, step S432 may include the following sub-steps S4321 - S4325:

[0139] S4321. Input the target human voice data into the convolutional layer for convolutional operation to generate an initial feature map;

[0140] S4322. Input the initial feature map into the pooling layer for downsampling to generate a target feature map;

[0141] S4323. Input the target feature map into the recurrent layer for capturing temporal data to generate sequence features;

[0142] S4324. Input the sequence features into the attention layer for feature weighting to generate feature vectors;

[0143] S4325. Input the feature vectors into the fully connected layer for device type recognition to generate device type data.

[0144] In an embodiment of the present invention, inputting the target human voice data into the convolutional layer for convolutional operation can automatically extract local features in the data to obtain an initial feature map. The human voice data contains rich acoustic features, such as information on the frequency, duration, and timbre of speech. The convolutional layer performs convolutional operations by sliding a convolutional kernel over the data, which can capture these local features, such as specific speech frequency band change patterns, the start and end features of pronunciation, etc. These local features are the key basis for subsequent device type recognition. Compared with directly using the original human voice data, the features extracted by convolution are more representative, which can significantly improve the accuracy of recognition. While extracting features, the convolutional operation can also reduce the dimension of the data. The dimension of the original human voice data may be relatively high, and the computational amount for processing is large. The convolutional layer reduces the dimension of the data through the operation of the convolutional kernel, reduces the subsequent computational amount while retaining the key features, and improves the operating efficiency of the system.

[0145] The initial feature map is input into the pooling layer for downsampling, further reducing the data dimension. Pooling operations (such as max pooling or average pooling) can compress the feature map without losing too much key information. This not only reduces the computational amount of the subsequent layers but also prevents the model from overfitting. Because in high-dimensional data, the model is prone to learning the noise and subtle variations in the data, resulting in overfitting. The pooling layer summarizes the features of local regions, enabling the model to focus on more macroscopic and representative features, improving the generalization ability of the model. The pooling operation makes the obtained target feature map more robust to local position changes. For example, in speech data, small changes in the speaker's pronunciation position or speaking speed may cause local changes in the feature map. The pooling layer can, to a certain extent, ignore these small changes by taking the maximum or average value of the local region, making the extracted features more stable and conducive to accurately identifying the device type subsequently.

[0146] Speech data has temporal characteristics, and the order of speaking contains important semantic information. Recurrent layers (such as RNN, LSTM, or GRU) can effectively capture this temporal data and generate sequence features. By performing recurrent calculations in the time dimension, the recurrent layer can remember the information of previous moments and use this information to process the data at the current moment. In device type recognition, this means that the sequential relationship between different parts of the voice command can be captured, such as the association between the previously mentioned "turn on" and the device name mentioned later, thereby more accurately understanding the user's intention and improving the accuracy of device type recognition. The lengths of voice commands from different users may vary, and the recurrent layer can handle this variable-length data well. It can dynamically adjust the calculation process according to the length of the input speech without the need to truncate or pad the data to a fixed length, improving the adaptability of the system to the speech habits of different users.

[0147] The attention layer performs feature weighting on the sequence features, can automatically learn the importance of different features, and assigns different weights to each feature according to the importance to obtain a feature vector. In speech data, certain features may be more critical for device type recognition, such as the pronunciation features of specific keywords. The attention layer highlights the role of these key features in the recognition process by assigning higher weights to them, while giving lower weights to less important features, so that the model pays more attention to key information and improves the recognition accuracy. The attention mechanism not only improves the performance of the model but also increases the interpretability of the model. By observing the weights assigned by the attention layer, it is possible to understand the key features that the model focuses on during the recognition process, which helps to analyze the decision-making process of the model and provides a basis for optimizing and improving the model.

[0148] The fully connected layer comprehensively processes the feature vectors output by the attention layer to achieve the recognition of device types. Each neuron in the fully connected layer is connected to all neurons in the previous layer, enabling full utilization of the feature information extracted from the previous layers. By performing weighted summation on these features and processing them through an activation function, the fully connected layer can map the feature vectors to different device type categories, ultimately generating device type data. This comprehensive processing method can fully explore the potential relationships between features, improving the accuracy and reliability of recognition. The fully connected layer can adapt to different device type recognition tasks by adjusting the weight parameters. Whether it is to recognize a few common devices or multiple different types of devices, the fully connected layer can learn appropriate weights through training to achieve accurate classification, with strong flexibility and adaptability.

[0149] Furthermore, the present invention can also set up a multi-level alarm mechanism. For example, in the case of the first failure, a gentle reminder voice such as "Voiceprint verification failed, please try again" is played through the built-in speaker of the smart device; in the case of two consecutive failures, in addition to the voice prompt, a reminder message is pushed to the user's mobile phone APP; if there are three or more failures, a high-decibel alarm is activated, and a text message notification containing the on-site location and abnormal situation is sent to the preset emergency contacts, enhancing the security protection ability.

[0150] Furthermore, the preset voice prompt data can be customized according to different scenarios and user needs. For families with the elderly or children, the voice prompt uses a more gentle and understandable expression; if it is detected that the voiceprint verification fails at night, the volume of the voice prompt is reduced to avoid disturbing the residents, and at the same time, it is reminded in a way such as flashing lights. In addition, the content of the voice prompt can include operation guidance, such as "Please approach the device and clearly say the wake-up word", to help users correctly complete the voiceprint verification.

[0151] Furthermore, the present invention can also implement an alarm linkage function. When the alarm is triggered due to the failure of voiceprint verification, the smart door lock at home is automatically locked to prevent unauthorized personnel from entering; the camera is turned on for recording, and the real-time video is transmitted to the user's mobile phone APP for the user to remotely view the situation at home, enhancing the home security.

[0152] Step 204: Wake up the corresponding smart device in the home environment according to the device type data to generate home system wake-up data.

[0153] The device type data includes descriptions of the device's category identifier (such as smart lamps, smart air conditioners, smart curtains, etc.), device model, functional features, and other aspects.

[0154] In the embodiment of the present invention, the smart device corresponding to the device type data is awakened and the corresponding operation is performed to obtain the home system wake-up data corresponding to the voice wake-up data.

[0155] Step 205: Divide the device response current values corresponding to the home system wake-up data according to a preset time period to generate multiple current value sets.

[0156] In the embodiments of the present invention, after the home system wakes up an intelligent device, it will collect the response current values during the device operation in real time, and these current values will change dynamically with factors such as the device operation state and workload. The "preset time period" is a time interval set in advance according to actual needs, such as 1 minute, 5 minutes, etc. The system will group the collected continuous device response current values according to this fixed time interval. For example, if the preset time period is 5 minutes, the system will group the current values collected within every 5 minutes into one group, and each group constitutes a "current value set". The purpose of such division is to facilitate subsequent targeted analysis of the device current situation in different time periods, such as comparing the power consumption stability of the device in different periods.

[0157] Step 206: Calculate the mean values of the current value sets respectively to generate multiple current means.

[0158] In the embodiments of the present invention, for each current value set generated in the previous step, the system will use a specific calculation method to add up all the current values in the set and then divide by the number of current values to obtain the average value of the set, that is, the "current mean". Calculating the current mean can eliminate the influence of short-term current fluctuations, more intuitively reflect the average level of the device current within a preset time period, help judge the overall trend of the device current during operation, and provide a basis for subsequent anomaly judgment.

[0159] Step 207: Compare the current means with a preset mean threshold respectively to generate current comparison data.

[0160] The preset mean threshold is a reference standard value set according to the current level during the normal operation of the device. This threshold takes into account factors such as the device model, power, and working mode, and has been determined during the device installation and commissioning or system setting stage.

[0161] In the embodiments of the present invention, each current mean is compared with the preset mean threshold one by one to determine whether the current mean is greater than, equal to, or less than the threshold. The comparison results will be recorded in a specific data format to form "current comparison data", for example, recorded as "the current mean is greater than the threshold", "the current mean is equal to the threshold", "the current mean is less than the threshold", so as to judge whether the current of the device is within the normal range in each time period.

[0162] Step 208: Compare the device temperature change data in the device response data with a preset temperature threshold to generate temperature comparison data.

[0163] The preset temperature threshold is a critical value set according to the design parameters, safety standards, etc. of the device, and is used to measure whether the device temperature is normal. The system compares the collected device temperature change data with the preset temperature threshold to determine whether the device temperature change exceeds the normal range, and records the comparison result as "temperature change data greater than the threshold", "temperature change data equal to the threshold", "temperature change data less than the threshold", etc., to form "temperature comparison data", so as to master the temperature operation status of the device.

[0164] In the embodiment of the present invention, during the operation of the intelligent device, its internal components will generate heat, resulting in changes in the device temperature. The system will collect the temperature data of the device in real time through a temperature sensor and calculate the temperature change data (such as the temperature rise amplitude per unit time, the difference between the current temperature and the initial temperature, etc.).

[0165] Step 209: When any one of the current comparison data and the temperature comparison data is greater than the preset alarm data, send the alarm data to the preset alarm device.

[0166] The preset alarm data is a key indicator that defines the trigger conditions for device abnormal alarms. It stipulates under what circumstances an alarm needs to be triggered, and its specific value range is set according to actual needs.

[0167] In the embodiment of the present invention, when "the current average value is greater than the preset average threshold" appears in the current comparison data, or "the device temperature change data is greater than the preset temperature threshold" appears in the temperature comparison data, as long as any one of these conditions is met, it means that the device operation has an abnormal situation, and there may be a fault or safety hazard. At this time, the system will immediately send "alarm data" containing device number, abnormal type (current abnormality or temperature abnormality), abnormal value, etc. to the "preset alarm device" (such as mobile phone, home alarm, property monitoring terminal, etc.) according to the pre-set alarm rules, to remind users or relevant management personnel to take measures to handle the device abnormal problem in time to avoid more serious consequences.

[0168] A method for waking up a home system based on artificial intelligence provided by the present invention, in the process of processing voice wake-up data, uses short-time Fourier transform, a preset time-frequency filter, and a complex noise filtering formula. First, the voice waveform data is converted to the frequency domain through short-time Fourier transform to clearly show the frequency components. Then, the preset time-frequency filter (such as a low-pass, high-pass, or band-pass filter) is used to filter according to the voice spectrum characteristics of the user's voice and the voice that the device can receive and respond to, removing noise in other frequency bands. When generating the initial voice time-frequency data, specific formulas are also used to filter noise for continuous and discrete time signals respectively. This multi-step and refined processing method can more effectively remove environmental noise, improve the quality of the voice signal, greatly reduce the interference of noise on the wake-up process, and reduce the situation of wake-up failure compared with traditional simple noise reduction methods.

[0169] The target hidden Markov model is used for human voice extraction. In the model training stage, Mel-frequency cepstral coefficients are used to extract features from human voice and non-human voice data in the home environment, covering various data such as different scenarios and speakers to ensure the generalization ability of the model. When extracting the human voice, the Viterbi algorithm is used to perform state estimation on the human voice feature sequence to accurately identify the human voice segment, and then the spectral subtraction method is combined to remove background data, the adaptive filtering method is used to remove reverberation, and volume normalization processing is performed to obtain clear and standard target human voice data. This series of operations can accurately separate the user's voice command from complex environmental sounds, and can improve the wake-up success rate and enhance the user experience even in a noisy environment.

[0170] In the process of device type recognition, the cosine similarity improvement algorithm (Mahalanobis-cosine similarity) is used to calculate the similarity between the target human voice data and the voiceprint. This algorithm comprehensively considers the correlation and variance between features and can measure the voiceprint similarity more accurately than the traditional cosine similarity algorithm. Voiceprint verification is performed by setting a preset similarity threshold. After the verification passes, device keyword recognition is performed. If the verification fails, a preset alarm is activated. And a multi-level alarm mechanism and alarm linkage function, as well as customizable preset voice prompt data, are also set. This not only improves the accuracy of device recognition, avoids false wake-up, but also enhances the security of the home system, provides more user-friendly services while ensuring the user's device usage, and improves the overall user experience.

[0171] The current value and temperature after waking up the device are monitored. The device response current value is divided according to a preset time period and the mean value is calculated and compared with the preset mean value threshold. At the same time, the device temperature change data is compared with the preset temperature threshold. When the current or temperature data exceeds the preset alarm data, alarm data is sent to the preset alarm device. This function can timely detect abnormal situations during the operation of the device, prevent device failures in advance, ensure the stable operation of the device, avoid affecting the user's use due to device failures, and improve the user experience indirectly.

[0172] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for waking up a home system based on artificial intelligence, characterized in that, It includes the following steps: Obtain voice wake-up data in the home environment, perform spectral analysis on the voice wake-up data, and generate initial human voice data; Use a target hidden Markov model to extract human voices from the initial human voice data and generate target human voice data; Perform device type recognition on the target human voice data and a preset voice feature library to generate device type data; Wake up the corresponding intelligent device in the home environment according to the device type data to generate home system wake-up data.

2. The wake-up method of a home system based on artificial intelligence according to claim 1, characterized in that, The performing spectral analysis on the voice wake-up data to generate initial human voice data includes: Perform short-time Fourier transform on the voice wake-up data to generate initial voice time-frequency data; Use a preset time-frequency filter to perform filtering processing on the initial voice time-frequency data to generate target voice time-frequency data; Perform inverse short-time Fourier transform on the target voice time-frequency data to generate initial human voice data.

3. The wake-up method of a home system based on artificial intelligence according to claim 2, wherein The performing short-time Fourier transform on the voice wake-up data to generate initial voice time-frequency data includes: Use a first short-time Fourier transform formula to filter out noise from the continuous time signal in the voice wake-up data to generate first voice time-frequency data; The first short-time Fourier transform formula is: Among them, X improved (τ, ω) represents the spectral value at the time position τ and the angular frequency ω; x(t) is the original continuous-time audio signal, that is, the continuous-time signal in the sound wake-up data; ω adaptive (t, τ) is an adaptive window function; e -jωt is the complex exponential function in the continuous signal, which is used to transform the time-domain signal to the frequency domain, where j is the imaginary unit, ω is the angular frequency, and t is the time variable; Use a second short-time Fourier transform formula to filter out noise from the discrete time signal in the voice wake-up data to generate second voice time-frequency data; The second short-time Fourier transform formula is: where X improved [m, k] is a two-dimensional array representing the spectral value at discrete time position m and discrete frequency index k; x[n] is the original discrete-time audio signal, i.e., the discrete-time signal in the sound wake-up data, where n represents the discrete-time index corresponding to the sampling moment of the continuous-time signal; ω[n - m] is a discrete window function used to window the discrete signal x[n]; is the complex exponential function in the discrete signal, used to transform the discrete-time domain signal to the discrete frequency domain, where j is the imaginary unit, is the frequency resolution, k is the discrete frequency index, n is the discrete-time index, N is the number of points of the fast Fourier transform; c[m, k] is the noise compensation factor; Use the first voice time-frequency data and the second voice time-frequency data to construct initial voice time-frequency data.

4. A method for waking up a home system based on artificial intelligence according to claim 2, characterized in that, The using a preset time-frequency filter to perform filtering processing on the initial voice time-frequency data to generate target voice time-frequency data includes: Calculate the average value of the noise spectrum data corresponding to the home environment to generate an energy distribution estimate value; Use the voice time-frequency data and the energy distribution estimate value to calculate the signal-to-noise ratio estimate values of each time-frequency point corresponding to the voice wake-up data respectively to generate multiple signal-to-noise ratio estimate values; Respectively use the signal-to-noise ratio estimate values and a preset target signal-to-noise ratio to calculate filter coefficients to generate filter coefficients corresponding to the time-frequency points; Select the filter coefficients greater than a preset coefficient threshold to construct a time-frequency data set; Perform multiplication calculation on the time-frequency data in the time-frequency data set and the corresponding filter data respectively to generate target voice time-frequency data.

5. A method for waking up a home system based on artificial intelligence according to claim 1, characterized in that, Before the step of using a target hidden Markov model to extract human voices from the initial human voice data and generate target human voice data, it further includes: Use Mel frequency cepstral coefficients to extract features from the human voice data and non-human voice data corresponding to the home environment to generate audio feature data; Initialize the model of the initial hidden Markov model according to preset initialization data to generate an intermediate hidden Markov model; Use the expectation-maximization algorithm to train the model of the intermediate hidden Markov model to generate a target hidden Markov model.

6. A method for waking up a home system based on artificial intelligence according to claim 1 or 5, characterized in that, The using a target hidden Markov model to extract human voices from the initial human voice data and generate target human voice data includes: Extract the features in the initial human voice data that are the same as the training features corresponding to the target hidden Markov model to generate a human voice feature sequence; Perform state estimation on the human voice feature sequence using the Viterbi algorithm through the target Hidden Markov Model to generate a state sequence; Extract the human voice segments from the state sequence to generate human voice data; Use spectral subtraction to remove background data from the human voice data to generate initial human voice data; Use adaptive filtering to remove reverberation from the initial human voice data to generate intermediate human voice data; Normalize the volume of the intermediate human voice data to generate target human voice data.

7. A method for waking up a home system based on artificial intelligence according to claim 1, characterized in that, The device type recognition of the target human voice data and a preset voice feature library to generate device type data includes: Screen out the voiceprint matching the target human voice data from the preset voice feature library to generate candidate voiceprints; Use a cosine similarity improvement algorithm to calculate the similarity between the target human voice data and the candidate voiceprints respectively to generate target cosine similarity; Perform device type recognition based on the target cosine similarity to generate device type data.

8. A method for waking up a home system based on artificial intelligence according to claim 7, characterized in that, The device type recognition based on the target cosine similarity to generate device type data includes: Judge whether the cosine similarity is greater than a preset similarity threshold; If so, the voiceprint verification passes, and perform device keyword recognition on the target human voice data to generate device type data; If not, the voiceprint verification fails, and start a preset alarm to play preset voice prompt data.

9. The wake-up method of a home system based on artificial intelligence according to claim 8, characterized in that, The device keyword recognition of the target human voice data to generate device type data includes: Input the target human voice data into the convolutional layer for convolution operation to generate an initial feature map; Input the initial feature map into the pooling layer for downsampling to generate a target feature map; Input the target feature map into the recurrent layer to capture temporal data to generate sequence features; Input the sequence features into the attention layer for feature weighting to generate feature vectors; Input the feature vectors into the fully connected layer for device type recognition to generate device type data.

10. A method for waking up a home system based on artificial intelligence according to claim 1, characterized in that, After the step of waking up the corresponding intelligent device in the home environment according to the device type data to generate home system wake-up data, it further includes: Divide the device response current value corresponding to the home system wake-up data according to a preset time period to generate multiple current value sets; Calculate the mean values of the current value sets respectively to generate multiple current means; Compare the current means with a preset mean threshold respectively to generate current comparison data; Compare the device temperature change data in the device response data with a preset temperature threshold to generate temperature comparison data; When any of the current comparison data and the temperature comparison data is greater than the preset alarm data, send alarm data to a preset alarm device.

Citation Information

Patent Citations

  • Voice wake-up method, acoustic model training method and related device

    CN115223555A

  • Voice control method and device, storage medium and electronic equipment

    CN116896488A

  • Electrical equipment data intelligent acquisition and transmission method based on Internet of Things

    CN118657517A