Air purifier voice control instruction analysis method based on semantic understanding
By building a multi-label control strategy network and synchronously collecting speech states, dynamic denoising and time-frequency feature extraction, and generating semantic coding vectors, the speech recognition problem of air purifiers in high-noise environments is solved, and adaptive control and intelligent improvement are achieved.
Patent Information
- Application Number
- CN202510780152.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing voice control methods for air purifiers have low recognition rates in high-noise environments, lack the ability to collaboratively process the equipment's operating status, make it difficult to accurately analyze complex control intentions, and the control response lacks perception of environmental parameters, resulting in a low level of intelligence.
A multi-label control strategy network is constructed. By synchronously collecting voice and state parameters, the noise feature library is dynamically selected for adaptive denoising, time-frequency mixed features are extracted, a convolutional-recursive neural network is used to generate semantic encoding vectors, and a control instruction queue is generated in combination with real-time environmental parameters.
It significantly improves the robustness of speech recognition in high-noise environments, enhances the precision of semantic modeling and the level of equipment intelligence, realizes adaptive control under multiple working conditions, and enhances the user interaction experience.
Smart Images

Figure CN120690191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent home appliance control technology, and in particular to a method for parsing voice control instructions for an air purifier based on semantic understanding. Background Art
[0002] With the development of smart home technology, voice interaction, as an important form of human-computer interaction, is increasingly used in smart devices such as air purifiers. Users can use voice to control the device's power on and off, adjust the wind speed, switch modes, and perform other operations, which improves the convenience of use and the naturalness of interaction. Voice control shows unique advantages, especially in scenarios where there is no need to touch the device, hands are occupied, or operation is performed from a distance. However, since air purifiers themselves generate background noise such as wind noise and motor noise during operation, the accuracy and stability of voice recognition become key factors affecting user experience.
[0003] Existing voice control methods for air purifiers mostly use fixed command recognition or general voice recognition models, lack the ability to coordinate processing with the operating status of the equipment, and cannot dynamically adapt to noise characteristics according to actual wind speed, mode or operating frequency, resulting in a significant decrease in recognition rate in high-noise scenarios. In addition, traditional methods mostly rely on static feature extraction and fail to effectively model the time-frequency variation characteristics and semantic context in voice signals, making it difficult to accurately analyze complex control intentions. The control response lacks perception of environmental parameters and cannot achieve scenario-based adaptive command generation, resulting in a low overall intelligence level. Summary of the Invention
[0004] The present invention provides a method for parsing air purifier voice control instructions based on semantic understanding, constructs a multi-label control strategy network, and outputs a control instruction queue that can dynamically adapt to the scene, thereby improving the accuracy, robustness and intelligence level of air purifier voice control.
[0005] The method for parsing voice control instructions for an air purifier based on semantic understanding includes the following steps:
[0006] S1, synchronous acquisition of voice and status parameters: real-time acquisition of user voice signals and synchronous acquisition of the operating status parameters of the air purifier, including the current wind speed level, operating mode code and motor operating frequency;
[0007] S2, device noise perception and adaptive denoising: Dynamically selects the corresponding noise feature library based on operating status parameters, constructs an adaptive filter to eliminate device noise from the original voice signal, and generates a denoised voice stream;
[0008] S3, time-frequency mixture feature extraction: extract time-frequency mixture features from the denoised speech stream, including device state-related formant shifts and dynamic fundamental frequency trajectories;
[0009] S4, semantic encoding vector generation: the time-frequency mixed features are input into the cascaded convolutional-recurrent neural network to generate a semantic encoding vector including device noise compensation information;
[0010] S5, spatiotemporal correlation analysis and control instruction generation: Based on the spatiotemporal correlation analysis of the semantic coding vector and the real-time environment parameters, a control instruction queue is generated.
[0011] Optionally, the synchronous collection of voice and state parameters in S1 includes:
[0012] S11, real-time collection and buffering of user voice signals: through the microphone array embedded in the air purifier body or its external module, with a sampling rate of R s Perform audio acquisition and use short-time Fourier transform (STFT) to divide the continuous speech stream into frames and store them into a frame vector set V speech ={v1,v2,...,v N}, where v1, v2, ..., v N are the frequency domain feature vectors of the 1st, 2nd, ..., Nth frames of speech signals, respectively, where N is the number of speech signal frames;
[0013] S12, air purifier operation status collection and synchronization timestamp binding: read the current operation status parameters of the air purifier in real time through the built-in sensor interface of the device, and construct the state parameter vector S device =[F w ,M c ,f m ], and bind the timestamp t s , and the speech frame set V speech The starting frame time t in v Align, if |t s -t v |≤Δt, the state parameter vector is considered to be synchronized with the speech data segment, where Δt is the maximum allowable synchronization time error threshold (100ms), and F w is the current wind speed level, M c Encode the operating mode, f m is the motor operating frequency;
[0014] S13, construct synchronous acquisition feature pairs and output: for each valid speech segment frame set Its corresponding state parameter vector Combine to form synchronous acquisition feature pairs Output synchronous acquisition feature sequence.
[0015] Optionally, the device noise perception and adaptive denoising in S2 includes:
[0016] S21, noise feature library matching and noise spectrum template extraction: According to the acquired equipment operating status parameter vector S device , retrieve the most matching noise spectrum template from the pre-built equipment noise feature library. The equipment noise feature library uses wind speed gear, operation mode and motor frequency as index fields, and pre-stores the background noise spectrum under the operating conditions. By calculating the Euclidean distance between the current state parameter vector and the sample in the equipment noise feature library, the most similar noise template is selected as the target noise spectrum estimation, which is recorded as N est (f);
[0017] S22, adaptive filter construction and speech denoising processing: the original speech frame vector sequence V speech Perform frame-level Fourier transform and based on the target noise spectrum N est (f) Construct an adaptive Wiener filter and apply the filter gain G(f) to the speech spectrum and then perform an inverse Fourier transform to obtain the denoised speech frame set V denoise , and concatenate to generate a time-series continuous denoised speech stream.
[0018] Optionally, the noise feature library matching and noise spectrum template extraction in S21 include:
[0019] S211, retrieve equipment noise feature library samples: suppose the pre-built equipment noise feature library is in, is the operating state parameter vector of the i-th sample, is the wind speed level of the i-th sample, Encode the operating mode of the i-th sample, is the motor frequency of the i-th sample, N (i) (f) is the corresponding equipment background noise spectrum template, and K is the historical noise sample record;
[0020] S212, calculating the Euclidean distance and matching the optimal template: calculating the Euclidean distance between the current state vector and the device noise feature library sample, and selecting the sample with the smallest distance as the optimal matching sample;
[0021] S213, output the target noise spectrum template: the noise spectrum template corresponding to the best matching sample is used as the current estimated value, recorded as
[0022] Optionally, the adaptive filter construction and speech denoising processing in S22 include:
[0023] S221, original speech frame spectrum conversion: the original speech signal x(t) is subjected to frame segmentation and windowing processing to obtain a frame vector sequence Among them, N is the total number of frames, and each frame is Fourier transformed to obtain the frequency domain speech signal X(j) (f);
[0024] S222, filter gain function construction: based on the obtained target noise spectrum N wst (f) Construct the frequency domain gain function of the Wiener filter for each frame to obtain the filter gain G (j) (f);
[0025] S223, denoising spectrum calculation and signal restoration: set the filter gain G (j) (f) Applying the denoised spectrum to each frame to obtain the denoised spectrum, and performing inverse Fourier transform on the denoised spectrum to restore it to the time domain frame signal;
[0026] S224, denoised speech stream splicing output: The time domain signals of all denoised frames are spliced together by overlapping and adding according to the original frame order to generate a time-series continuous denoised speech stream y(t).
[0027] Optionally, the time-frequency mixed feature extraction in S3 includes:
[0028] S31, formant frequency extraction and offset calculation: perform power spectrum estimation on each frame of denoised speech signal, extract its formant frequency using linear prediction cepstral coefficients (LPCC), and assume that the first P-order formant frequency extracted from the k-th frame is With standard static template frequency Compare and obtain the resonance peak offset vector;
[0029] S32, dynamic fundamental frequency trajectory estimation: using the autocorrelation function (ACF) method to extract the fundamental frequency of each frame of speech Dynamic trajectory modeling is performed based on a sliding time window, and the fundamental frequency of consecutive frames is generated to form a trajectory sequence f0;
[0030] S33, constructing a time-frequency hybrid feature vector: concatenate the formant offset vector and the fundamental frequency of each frame to construct the time-frequency hybrid feature vector z of the kth frame (k) , and summarize all frames to obtain the time-frequency mixed feature sequence for semantic coding .
[0031] Optionally, the semantic encoding vector generation in S4 includes:
[0032] S41, Convolutional feature extraction network construction and input: the extracted time-frequency mixed feature sequence Input to a one-dimensional convolutional neural network (1D-CNN), extract local feature representation through multiple convolution layers, and after M layers of convolution, the output is a convolution feature sequence G cnn =h (M) ;
[0033] S42, Recursive Feature Modeling and Time Series Context Fusion: Convolution output H cnn Input into the bidirectional gated recurrent unit Bi-GRU, encode the time dependency, obtain the context-aware semantic state vector sequence, and output the context-enhanced feature sequence H gru ;
[0034] S43, semantic encoding vector generation and global compression: the context-enhanced feature sequence is input into the temporal pooling unit to generate the global semantic encoding vector e sem .
[0035] Optionally, the recursive feature modeling and time series context fusion in S42 include:
[0036] S421, Convolution output dimension mapping and sequence preparation: Convolution network output feature vector of each frame Transformed into a vector that meets the Bi-GRU input requirements through linear mapping
[0037] S422, Bi-GRU network temporal modeling: Input the mapped vector into the Bi-GRU network, calculate the forward and backward hidden states respectively, extract the context information of each frame, and splice it into the context enhancement vector r (t) ;
[0038] S423, context enhancement sequence output: combine the context enhancement vectors of all frames into a sequence H gru .
[0039] Optionally, the spatiotemporal correlation analysis and control instruction generation in S5 includes:
[0040] S51, real-time environment parameter vector construction: obtain multiple sensor parameters in the air purifier working environment, and construct the spatiotemporal environment parameter vector e env =[T,H,P,D,t s ], where T is the current ambient temperature, H is the current air humidity, P is the current PM2.5 concentration, D is the current spatial noise intensity, t s is the sampling timestamp;
[0041] S52, spatiotemporal feature fusion and response matching: The obtained semantic encoding vector e sem and the environmental parameter vector e env Perform feature splicing to form a joint decision vector e joint =[e sem ;W E ·e env +b E ], where W E is the environmental parameter mapping weight matrix, b Eis the environmental parameter bias vector;
[0042] S53, control instruction queue generation: the joint decision vector e joint Input the control policy network, perform multi-label classification, and generate a control instruction queue.
[0043] Beneficial effects of the present invention:
[0044] The present invention realizes dynamic collaborative modeling of speech and device environment by introducing a synchronous acquisition mechanism of voice signals and device operating status. Combined with the adaptive noise filtering strategy driven by device status, it significantly improves the robustness of speech recognition in high-noise working environments. At the same time, the device noise template matching method driven by Euclidean distance can effectively adapt to noise changes in multiple working conditions and modes, effectively solving the problem of recognition errors caused by environmental background noise interference in traditional speech control.
[0045] The present invention constructs a multi-level neural network structure and inputs time-frequency features into a convolutional-recursive neural network model, thereby achieving deep abstraction of semantic intent in speech signals and fusion extraction of device status-related features. In particular, the introduction of resonance peak shift and fundamental frequency trajectory as significant features in dynamic speech streams effectively enhances the precision of semantic modeling. Further, a global pooling mechanism is used to generate semantic coding vectors, which enables device context perception.
[0046] The present invention uses a multi-label neural network to predict the control instruction queue. It can not only output multiple control actions at the same time, but also flexibly adjust the response strategy according to dynamic environmental data such as temperature, humidity, PM2.5, thereby realizing semantic-driven automatic control of the air purifier. It has high scalability and environmental adaptability, effectively improving the user interaction experience and the intelligence level of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 Schematic diagram of the analysis method according to an embodiment of the present invention;
[0049] Figure 2 A schematic diagram of generating a semantic coding vector according to an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art may also adopt other alternatives to implement some known technologies; and the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0051] like Figure 1-Figure 2 As shown, the method for parsing air purifier voice control instructions based on semantic understanding includes the following steps:
[0052] S1, synchronous acquisition of voice and status parameters: real-time acquisition of user voice signals and synchronous acquisition of the operating status parameters of the air purifier, including the current wind speed level, operating mode code and motor operating frequency;
[0053] S2, device noise perception and adaptive denoising: Dynamically selects the corresponding noise feature library based on operating status parameters, constructs an adaptive filter to eliminate device noise from the original voice signal, and generates a denoised voice stream;
[0054] S3, time-frequency mixture feature extraction: extract time-frequency mixture features from the denoised speech stream, including device state-related formant shifts and dynamic fundamental frequency trajectories;
[0055] S4, semantic encoding vector generation: the time-frequency mixed features are input into the cascaded convolutional-recurrent neural network to generate a semantic encoding vector including device noise compensation information;
[0056] S5, spatiotemporal correlation analysis and control instruction generation: Based on the spatiotemporal correlation analysis of the semantic coding vector and the real-time environment parameters, a control instruction queue is generated.
[0057] The synchronous acquisition of voice and status parameters in S1 includes:
[0058] S11, real-time collection and buffering of user voice signals: through the microphone array embedded in the air purifier body or its external module, with a sampling rate of R s Perform audio acquisition and use short-time Fourier transform (STFT) to divide the continuous speech stream into frames and store them into a frame vector set V speech ={v1,v2,...,v N}, where v1, v2, ..., v N are the frequency domain feature vectors of the 1st, 2nd, ..., Nth frames of speech signals, respectively. N is the number of speech signal frames, which depends on the length of the speech segment and the frame shift.
[0059] S12, air purifier operation status collection and synchronization timestamp binding: read the current operation status parameters of the air purifier in real time through the built-in sensor interface of the device, and construct the state parameter vector S device =[F w ,M c,f m ], and bind the timestamp t s , and the speech frame set V speech The starting frame time t in v Align, if |t s -t v |≤Δt, the state parameter vector is considered to be synchronized with the speech data segment, where Δt is the maximum allowable synchronization time error threshold (100ms), and F w is the current wind speed level, M c Encode the operating mode, f m is the motor operating frequency;
[0060] S13, construct synchronous acquisition feature pairs and output: for each valid speech segment frame set Its corresponding state parameter vector Combine to form synchronous acquisition feature pairs Output synchronous acquisition feature sequence.
[0061] Device noise perception and adaptive denoising in S2 include:
[0062] S21, noise feature library matching and noise spectrum template extraction: According to the acquired equipment operating status parameter vector S device , retrieve the most matching noise spectrum template from the pre-built equipment noise feature library. The equipment noise feature library uses wind speed gear, operation mode and motor frequency as index fields, and pre-stores the background noise spectrum under the operating conditions. By calculating the Euclidean distance between the current state parameter vector and the sample in the equipment noise feature library, the most similar noise template is selected as the target noise spectrum estimation, which is recorded as N est (f);
[0063] S22, adaptive filter construction and speech denoising processing: the original speech frame vector sequence V speech Perform frame-level Fourier transform and based on the target noise spectrum N est (f) Construct an adaptive Wiener filter and apply the filter gain G(f) to the speech spectrum and then perform an inverse Fourier transform to obtain the denoised speech frame set V denoise , and concatenate to generate a time-series continuous denoised speech stream.
[0064] Noise feature library matching and noise spectrum template extraction in S21 include:
[0065] S211, retrieve equipment noise feature library samples: suppose the pre-built equipment noise feature library is in, is the operating state parameter vector of the i-th sample, is the wind speed level of the i-th sample, Encode the operating mode of the i-th sample, is the motor frequency of the i-th sample, N (i) (f) is the corresponding equipment background noise spectrum template, and K is the historical noise sample record;
[0066] S212, calculate the Euclidean distance and match the optimal template: calculate the Euclidean distance between the current state vector and the device noise feature library sample, and select the sample with the smallest distance as the optimal matching sample, which is expressed as:
[0067]
[0068] i * =argmin i D i ;
[0069] Among them, D i is the Euclidean distance, i * is the index number of the best matching sample, indicating the one that is closest to the current state among all samples;
[0070] S213, output the target noise spectrum template: the noise spectrum template corresponding to the best matching sample is used as the current estimated value, recorded as
[0071] The adaptive filter construction and speech denoising processing in S22 include:
[0072] S221, original speech frame spectrum conversion: the original speech signal x(t) is subjected to frame segmentation and windowing processing to obtain a frame vector sequence Among them, N is the total number of frames, and each frame is Fourier transformed to obtain the frequency domain speech signal X (j) (f), expressed as:
[0073]
[0074] Among them, X (j) (f) is the complex spectrum speech signal at frequency f in the jth frame, x (j) (n) is the speech signal of the nth sampling point in the jth frame, L is the number of sampling points per frame, and n is the time domain sampling index;
[0075] S222, filter gain function construction: based on the obtained target noise spectrum N wst (f) Construct the frequency domain gain function of the Wiener filter for each frame to obtain the filter gain G (j) (f), expressed as:
[0076]
[0077] Among them, G (j)(f) is the filter gain of the jth frame at frequency f;
[0078] S223, denoising spectrum calculation and signal restoration: set the filter gain G (j) (f) is applied to each frame spectrum to obtain the denoised spectrum, and the denoised spectrum is inverse Fourier transformed to restore the time domain frame signal, which is expressed as:
[0079] Y (j) (f) = G (j) (f)·X (j) (f);
[0080]
[0081] Among them, Y (j) (f) is the denoised spectrum of the jth frame, y (j) (t) is the denoised time domain signal of the jth frame, is the inverse Fourier transform operation;
[0082] S224, denoised speech stream splicing output: The time domain signals of all denoised frames are spliced together by overlapping and adding according to the original frame order to generate a time-series continuous denoised speech stream y(t), which is expressed as:
[0083]
[0084] Time-frequency hybrid feature extraction in S3 includes:
[0085] S31, formant frequency extraction and offset calculation: perform power spectrum estimation on each frame of denoised speech signal, extract its formant frequency using linear prediction cepstral coefficients (LPCC), and assume that the first P-order formant frequency extracted from the k-th frame is With standard static template frequency By comparison, we get the resonance peak offset vector, which is expressed as:
[0086]
[0087] Where ΔF (k) is the formant offset vector of the kth frame;
[0088] S32, dynamic fundamental frequency trajectory estimation: using the autocorrelation function (ACF) method to extract the fundamental frequency of each frame of speech Dynamic trajectory modeling is performed based on the sliding time window, and the fundamental frequency on consecutive frames is generated to form a trajectory sequence f0, which is expressed as:
[0089]
[0090] Among them, R (k) (τ) is the autocorrelation value when the delay is τ, y(k) (n) is the speech signal of the nth sampling point in the kth frame, τ is the delay value range, and L' is the number of sampling points involved in the autocorrelation calculation in each frame;
[0091]
[0092] S33, constructing a time-frequency hybrid feature vector: concatenate the formant offset vector and the fundamental frequency of each frame to construct the time-frequency hybrid feature vector z of the kth frame (k) , and summarize all frames to obtain the time-frequency mixed feature sequence for semantic coding , expressed as:
[0093]
[0094]
[0095] Among them, z (1) ,z (2) ,...,z (K) are the time-frequency mixed feature vectors of the 1st, 2nd, ..., Kth frames respectively.
[0096] The semantic encoding vector generation in S4 includes:
[0097] S41, Convolutional feature extraction network construction and input: the extracted time-frequency mixed feature sequence Input to a one-dimensional convolutional neural network (1D-CNN), extract local feature representation through multiple convolution layers, and after M layers of convolution, the output is a convolution feature sequence G cnn =h (M) , expressed as:
[0098] h (l) =σ(W (l) *h (l-1) +b (l) );
[0099] Among them, h (l) is the output feature map of the lth layer, σ is the ReLU activation function, W (l) is the convolution kernel weight of the lth convolution layer, b (l) is the bias term of the lth layer, h (l-1) is the output feature map of the l-1 layer;
[0100] S42, Recursive Feature Modeling and Time Series Context Fusion: Convolution output H cnn Input into the bidirectional gated recurrent unit Bi-GRU, encode the time dependency, obtain the context-aware semantic state vector sequence, and output the context-enhanced feature sequence H gru ;
[0101] S43, semantic encoding vector generation and global compression: the context-enhanced feature sequence is input into the temporal pooling unit to generate the global semantic encoding vector e sem , expressed as:
[0102]
[0103] in, is the context-enhanced feature vector output by the Bi-GRU network at frame t, g comp is the device noise compensation vector.
[0104] Recursive feature modeling and time series context fusion in S42 include:
[0105] S421, Convolution output dimension mapping and sequence preparation: Convolution network output feature vector of each frame Transformed into a vector that meets the Bi-GRU input requirements through linear mapping Expressed as:
[0106]
[0107] in, is the Bi-GRU input vector after mapping, W R 、b R are mapping weights and bias terms respectively;
[0108] S422, Bi-GRU network temporal modeling: Input the mapped vector into the Bi-GRU network, calculate the forward and backward hidden states respectively, extract the context information of each frame, and splice it into the context enhancement vector r (t) , expressed as:
[0109]
[0110] in, is the hidden state output of the forward GRU of the t-th frame, is the hidden state output of GRU after the t-th frame, GRU fw , GRU bw They are forward GRU and backward GRU respectively. is the hidden state output of the forward GRU of the t-1th frame, is the hidden state output of GRU after the t+1th frame;
[0111]
[0112] GRU fw Expressed as:
[0113] Update Gate:
[0114] Reset the gate:
[0115] Candidate status:
[0116] Current hidden state:
[0117] GRU bw Expressed as:
[0118] Update Gate:
[0119] Reset the gate:
[0120] Candidate status:
[0121] Current hidden state:
[0122] in, is the Sigmoid activation function, is the updated gate vector of the t-th frame, is the reset gate vector of the t-th frame, is the candidate hidden state vector of the t-th frame, is the final forward hidden state output of the t-th frame, is the forward hidden state of the t-1th frame, are the weight matrices input to each gate, are the weight matrices from hidden state to each gate, are the bias vectors of each gate, is the updated gate vector (reverse) of the t-th frame, is the reset gate vector (reverse) of the t-th frame, is the candidate hidden state vector of the t-th frame (reverse), is the final reverse hidden state output of the t-th frame, is the reverse hidden state of the t+1th frame, are the weight matrices input to each gate (reverse), They are the weight matrices from hidden state to each gate (reverse), are the bias vectors (reverse) of each gate respectively;
[0123] S423, context enhancement sequence output: combine the context enhancement vectors of all frames into a sequence H gru , expressed as:
[0124] H gru ={r (1) ,r (2) ,...,r (K)};
[0125] Among them, r (1) ,r (2) ,...,r (K) are the final representations of the context information before and after fusion of the 1st, 2nd, ..., Kth frames respectively.
[0126] The spatiotemporal correlation analysis and control instruction generation in S5 include:
[0127] S51, real-time environment parameter vector construction: obtain multiple sensor parameters in the air purifier working environment, and construct the spatiotemporal environment parameter vector e env =[T,H,P,D,t s ], where T is the current ambient temperature, H is the current air humidity, P is the current PM2.5 concentration, D is the current spatial noise intensity, t s is the sampling timestamp;
[0128] S52, spatiotemporal feature fusion and response matching: The obtained semantic encoding vector e sem and the environmental parameter vector e env Perform feature splicing to form a joint decision vector e joint =[e sem ;W E ·e env +b E ], where W E is the environmental parameter mapping weight matrix, b E is the environmental parameter bias vector;
[0129] S53, control instruction queue generation: the joint decision vector e joint Input the control policy network, perform multi-label classification, and generate a control instruction queue, which is expressed as:
[0130] c=sigmoid(W3·σ2(W2·σ1(w1·e joint +b1)+b2)+b3);
[0131] Among them, c is the control instruction queue, W1, W2, and W3 are the weight matrices of the corresponding neural networks, b1, b2, and b3 are the bias terms of each layer, and σ1 and σ2 are the hidden layer activation functions.
[0132] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.
[0133] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. The air purifier voice control instruction parsing method based on semantic understanding is characterized by: The following steps are involved: S1, synchronous acquisition of voice and status parameters: real-time acquisition of user voice signals and synchronous acquisition of the operating status parameters of the air purifier, including the current wind speed level, operating mode code and motor operating frequency; S2, device noise perception and adaptation Denoising: Dynamically selects the corresponding noise feature library based on operating status parameters, constructs an adaptive filter to eliminate device noise from the original voice signal, and generates a denoised voice stream; S3, time-frequency mixture feature extraction: extract time-frequency mixture features from the denoised speech stream, including device state-related formant shifts and dynamic fundamental frequency trajectories; S4, semantic encoding vector generation: the time-frequency mixed features are input into the cascaded convolutional-recurrent neural network to generate a semantic encoding vector including device noise compensation information; S5, spatiotemporal correlation analysis and control instruction generation: Based on the spatiotemporal correlation analysis of the semantic coding vector and the real-time environment parameters, a control instruction queue is generated.
2. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 1, characterized in that: The synchronous acquisition of voice and state parameters in S1 includes: S11, real-time collection and buffering of user voice signals: through the microphone array embedded in the air purifier body or its external module, with a sampling rate of R s Perform audio acquisition and use short-time Fourier transform to divide the continuous speech stream into frames and store them into a frame vector set V speech ={v1,v2,...,v N }, where v1, v2, ..., v N are the frequency domain feature vectors of the 1st, 2nd, ..., Nth frames of speech signals, respectively, where N is the number of speech signal frames; S12, air purifier operation status collection and synchronization timestamp binding: read the current operation status parameters of the air purifier in real time through the built-in sensor interface of the device, and construct the state parameter vector S device =[F w ,M c ,f m ], and bind the timestamp t s , and the speech frame set V speech The starting frame time t in v Align, if |t s -t v |≤Δt, the state parameter vector is considered to be synchronized with the speech data segment, where Δt is the maximum allowable synchronization time error threshold, F w is the current wind speed level, M c Encode the operating mode, f m is the motor operating frequency; S13, construct synchronous acquisition feature pairs and output: for each valid speech segment frame set Its corresponding state parameter vector Combine to form synchronous acquisition feature pairs Output synchronous acquisition feature sequence.
3. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 2, characterized in that: The device noise perception and adaptive denoising in S2 include: S21, noise feature library matching and noise spectrum template extraction: According to the acquired equipment operating status parameter vector S device , retrieve the most matching noise spectrum template from the pre-built equipment noise feature library. The equipment noise feature library uses wind speed gear, operation mode and motor frequency as index fields, and pre-stores the background noise spectrum under the operating conditions. By calculating the Euclidean distance between the current state parameter vector and the sample in the equipment noise feature library, the most similar noise template is selected as the target noise spectrum estimation, which is recorded as N est (f); S22, adaptive filter construction and speech denoising processing: the original speech frame vector sequence V speech Perform frame-level Fourier transform and based on the target noise spectrum N est (f) Construct an adaptive Wiener filter and apply the filter gain G(f) to the speech spectrum and then perform an inverse Fourier transform to obtain the denoised speech frame set V denoise , and concatenate to generate a time-series continuous denoised speech stream.
4. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 3, characterized in that: The noise feature library matching and noise spectrum template extraction in S21 include: S211, retrieve equipment noise feature library samples: suppose the pre-built equipment noise feature library is in, is the operating state parameter vector of the i-th sample, is the wind speed level of the i-th sample, Encode the operating mode of the i-th sample, is the motor frequency of the i-th sample, N (i) (f) is the corresponding equipment background noise spectrum template, and K is the historical noise sample record; S212, calculating the Euclidean distance and matching the optimal template: calculating the Euclidean distance between the current state vector and the device noise feature library sample, and selecting the sample with the smallest distance as the optimal matching sample; S213, output the target noise spectrum template: the noise spectrum template corresponding to the best matching sample is used as the current estimated value, recorded as 5. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 4, characterized in that: The adaptive filter construction and speech denoising processing in S22 include: S221, original speech frame spectrum conversion: the original speech signal x(t) is subjected to frame segmentation and windowing processing to obtain a frame vector sequence Among them, N is the total number of frames, and each frame is Fourier transformed to obtain the frequency domain speech signal X (j) (f); S222, filter gain function construction: based on the obtained target noise spectrum N wst (f) Construct the frequency domain gain function of the Wiener filter for each frame to obtain the filter gain G (j) (f); S223, denoising spectrum calculation and signal restoration: set the filter gain G (j) (f) Applying the denoised spectrum to each frame to obtain the denoised spectrum, and performing inverse Fourier transform on the denoised spectrum to restore it to the time domain frame signal; S224, denoised speech stream splicing output: The time domain signals of all denoised frames are spliced together by overlapping and adding according to the original frame order to generate a time-series continuous denoised speech stream y(t).
6. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 1, characterized in that: The time-frequency mixed feature extraction in S3 includes: S31, formant frequency extraction and offset calculation: perform power spectrum estimation on each frame of denoised speech signal, extract its formant frequency using linear prediction cepstral coefficients, and set the first P-order formant frequency extracted from the k-th frame to be With standard static template frequency Compare and obtain the resonance peak offset vector; S32, dynamic fundamental frequency trajectory estimation: using the autocorrelation function method to extract the fundamental frequency of each frame of speech Dynamic trajectory modeling is performed based on a sliding time window, and the fundamental frequency of consecutive frames is generated to form a trajectory sequence f0; S33, constructing a time-frequency hybrid feature vector: concatenate the formant offset vector and the fundamental frequency of each frame to construct the time-frequency hybrid feature vector z of the kth frame (k) , and summarize all frames to obtain the time-frequency mixed feature sequence for semantic coding 7. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 6, characterized in that: The semantic encoding vector generation in S4 includes: S41, Convolutional feature extraction network construction and input: The extracted time-frequency mixed feature sequence z is input into the one-dimensional convolutional neural network, and the local feature representation is extracted through multiple convolutional layers. After M layers of convolution, the output is the convolution feature sequence H cnn =h (M) ; S42, Recursive Feature Modeling and Time Series Context Fusion: Convolution output H cnn Input into the bidirectional gated recurrent unit Bi-GRU, encode the time dependency, obtain the context-aware semantic state vector sequence, and output the context-enhanced feature sequence H gru ; S43, semantic encoding vector generation and global compression: the context-enhanced feature sequence is input into the temporal pooling unit to generate the global semantic encoding vector e sem .
8. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 7, characterized in that: The recursive feature modeling and time series context fusion in S42 include: S421, Convolution output dimension mapping and sequence preparation: Convolution network output feature vector of each frame Transformed into a vector that meets the Bi-GRU input requirements through linear mapping S422, Bi-GRU network temporal modeling: Input the mapped vector into the Bi-GRU network, calculate the forward and backward hidden states respectively, extract the context information of each frame, and splice it into the context enhancement vector r (t) ; S423, context enhancement sequence output: combine the context enhancement vectors of all frames into a sequence H gru .
9. The method for parsing air purifier voice control instructions based on semantic understanding according to claim 8, characterized in that: The spatiotemporal correlation analysis and control instruction generation in S5 include: S51, real-time environment parameter vector construction: obtain multiple sensor parameters in the air purifier working environment, and construct the spatiotemporal environment parameter vector e env =[T,H,P,D,t s ], where T is the current ambient temperature, H is the current air humidity, P is the current PM2.5 concentration, D is the current spatial noise intensity, t s is the sampling timestamp; S52, spatiotemporal feature fusion and response matching: The obtained semantic encoding vector e sem and the environmental parameter vector e env Perform feature splicing to form a joint decision vector e joint =[e sem ;W E ·e env +b E ], where W E is the environmental parameter mapping weight matrix, b E is the environmental parameter bias vector; S53, control instruction queue generation: the joint decision vector e joint Input the control policy network, perform multi-label classification, and generate a control instruction queue.