Model construction method, range hood control method and range hood

By combining time domain and frequency domain feature extraction methods, a stove ignition sound recognition model was constructed, which solved the problem of low recognition accuracy of traditional range hoods in complex kitchen environments and realized the intelligent and rapid startup of the range hood.

CN120472898APending Publication Date: 2025-08-12GUANGDONG VANWARD NEW ELECTRIC CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510663018.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The ignition sound recognition technology of traditional hood stoves is low in accuracy in complex kitchen environments and is easily disturbed by environmental noise, resulting in delayed automatic startup.

Method used

A stove ignition sound recognition model is constructed by combining time domain and frequency domain feature extraction methods. The time-frequency dual-stream structure and differentiable search structure are used to optimize the neural network for stove ignition sound recognition and automatic start-up.

Benefits of technology

The recognition accuracy and response speed of the stove ignition sound are improved, the delay of the automatic start-up of the range hood is reduced, and the intelligent and rapid start-up of the range hood is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472898A_ABST
    Figure CN120472898A_ABST
Patent Text Reader

Abstract

The invention relates to a model construction method, a range hood control method and a range hood, and can be applied to the technical field of range hoods. The method comprises the following steps: acquiring stove ignition audio data; performing time domain feature extraction processing on the stove ignition audio data to obtain time domain features of the stove ignition audio data; performing frequency domain feature extraction processing on the stove ignition audio data to obtain frequency domain features of the stove ignition audio data; performing feature fusion processing on the time domain feature and the frequency domain feature to obtain a time frequency feature of the stove ignition audio data; based on the time-frequency characteristics, constructing a stove ignition sound recognition model; the cooker ignition voice recognition model is used for being deployed into a range hood; the range hood is used for being started under the condition that the stove ignition sound associated with the range hood is recognized through the stove ignition sound recognition model. By adopting the method, the accuracy and precision of stove ignition voice recognition can be improved, and the delay of automatic starting of the range hood is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of range hoods, and in particular to a model building method, a range hood control method, and a range hood. Background Art

[0002] With the development of smart home technology, the intelligence level of kitchen appliances continues to improve, among which the automatic start function of range hoods is an important aspect of improving the level of intelligence.

[0003] In traditional technology, the automatic start-up of the range hood can be achieved by recognizing the sound of stove ignition. However, due to the complex and changeable noise in the kitchen environment, traditional sound recognition technology is easily interfered by environmental factors in actual application, resulting in a technical problem of low accuracy in stove ignition sound recognition. Summary of the Invention

[0004] Based on this, it is necessary to provide a model construction method, a range hood control method and a range hood that improve the accuracy of stove ignition sound recognition in order to address the above technical problems.

[0005] In a first aspect, the present application provides a model construction method. The method comprises:

[0006] Get the stove ignition audio data;

[0007] Performing time domain feature extraction processing on the cooker ignition audio data to obtain time domain features of the cooker ignition audio data;

[0008] Performing frequency domain feature extraction processing on the cooker ignition audio data to obtain frequency domain features of the cooker ignition audio data;

[0009] Performing feature fusion processing on the time domain features and the frequency domain features to obtain time-frequency features of the cooker ignition audio data;

[0010] Based on the time-frequency features, a stove ignition sound recognition model is constructed; the stove ignition sound recognition model is used to be deployed in a range hood; the range hood is used to start when the stove ignition sound associated with the range hood is recognized by the stove ignition sound recognition model.

[0011] The above-mentioned model construction method is conducive to comprehensively and accurately capturing the time domain characteristics and frequency domain characteristics of the stove ignition sound by extracting the stove ignition audio features from two dimensions: time domain and frequency domain; performing time domain feature extraction processing on the stove ignition audio data can effectively capture the instantaneous change characteristics of the ignition sound, while performing frequency domain feature extraction processing on the stove ignition audio data can accurately identify the frequency distribution characteristics of the ignition sound, thereby better distinguishing the ignition sound from the ambient noise; through feature fusion processing, richer and more accurate time-frequency features are obtained, which is conducive to improving the recognition accuracy and response speed of the stove ignition sound recognition model for the ignition sound, and can quickly and accurately identify the stove ignition event at the early stage of the sound, reduce the time required for the recognition process, and improve the accuracy and precision of the stove ignition sound recognition, thereby effectively reducing the delay in the automatic start-up of the range hood.

[0012] In one embodiment, the performing time domain feature extraction processing on the cooker ignition audio data to obtain the time domain features of the cooker ignition audio data includes:

[0013] Performing one-dimensional convolution processing on the cooker ignition audio data through a time-domain stream branching structure to obtain pulse features of the cooker ignition audio data;

[0014] The pulse features are subjected to time series variation feature recognition processing to obtain the time domain features.

[0015] In one embodiment, the performing frequency domain feature extraction processing on the cooker ignition audio data to obtain the frequency domain features of the cooker ignition audio data includes:

[0016] Performing Fourier transform processing on the cooker ignition audio data through a frequency domain stream branching structure to obtain a time-frequency spectrum of the cooker ignition audio data;

[0017] Performing frequency band feature extraction processing on the time-spectrum graph to obtain frequency band features of the time-spectrum graph;

[0018] Perform two-dimensional convolution processing on the frequency band features to obtain the frequency domain features.

[0019] In one embodiment, the performing feature fusion processing on the time domain features and the frequency domain features to obtain the time-frequency features of the cooker ignition audio data includes:

[0020] Determining the dynamic weight value of the time domain feature and the dynamic weight value of the frequency domain feature through an attention mechanism;

[0021] According to the dynamic weight value of the time domain feature and the dynamic weight value of the frequency domain feature, feature fusion processing is performed on the time domain feature and the frequency domain feature to obtain the time-frequency feature.

[0022] In one embodiment, constructing a stove ignition sound recognition model based on the time-frequency features includes:

[0023] Based on the differentiable search structure and the time-frequency features, the stove ignition sound recognition model is constructed.

[0024] In one embodiment, the stove ignition sound recognition model is constructed based on the differentiable search structure and the time-frequency features, including:

[0025] Based on the differentiable search structure and the time-frequency features, a basic stove ignition sound recognition model is constructed;

[0026] The basic stove ignition sound recognition model is dynamically pruned and compressed by a dynamic channel pruning unit to obtain the stove ignition sound recognition model.

[0027] In one embodiment, constructing a basic stove ignition sound recognition model based on the differentiable search structure and the time-frequency features includes:

[0028] Constructing a stove ignition sound recognition model to be trained based on the differentiable search structure and the time-frequency features;

[0029] Performing a double-layer gradient optimization alternating training process on the stove ignition sound recognition model to be trained by using the differentiable search structure to obtain a trained stove ignition sound recognition model;

[0030] In a case where the trained stove ignition sound recognition model satisfies a termination search mechanism, the trained stove ignition sound recognition model is used as the basic stove ignition sound recognition model.

[0031] In one embodiment, obtaining the cooker ignition audio data includes:

[0032] Get the raw ignition audio data of the stove;

[0033] Performing data enhancement processing on the original ignition audio data to obtain data-enhanced ignition audio data;

[0034] The data-enhanced ignition audio data is subjected to data preprocessing to obtain the cooker ignition audio data; the data preprocessing includes at least one of noise elimination processing, resampling processing, shearing processing and windowing processing.

[0035] In a second aspect, the present application further provides a range hood control method, comprising:

[0036] Collect current audio data;

[0037] Performing time domain feature extraction processing on the current audio data to obtain current time domain features of the current audio data;

[0038] Performing frequency domain feature extraction processing on the current audio data to obtain current frequency domain features of the current audio data;

[0039] Performing feature fusion processing on the current time domain feature and the current frequency domain feature to obtain current time-frequency features of the current audio data;

[0040] Inputting the current time-frequency features into a stove ignition sound recognition model deployed in the range hood to perform stove ignition sound recognition processing to obtain a stove ignition sound recognition result of the current audio data; the stove ignition sound recognition model is obtained by any of the methods described in the first aspect above;

[0041] When the cooker ignition sound recognition result indicates that the cooker ignition sound associated with the range hood is recognized, the range hood is started.

[0042] In a third aspect, the present application further provides a range hood, comprising a range hood body, a memory, and a processor, wherein the memory stores a computer program, and the processor implements the steps of any of the above-described methods when executing the computer program. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1 A schematic diagram of a flow chart of a model building method in one embodiment;

[0045] Figure 2 A flow chart of a range hood control method according to another embodiment;

[0046] Figure 3 Schematic diagram of the network structure of a waveform U network in one embodiment;

[0047] Figure 4 A schematic diagram of the ignition recognition model training process in one embodiment;

[0048] Figure 5 2. It is a structural diagram of a time-frequency dual-stream structure in one embodiment;

[0049] Figure 6Schematic diagram of the execution flow of a differentiable search structure in one embodiment;

[0050] Figure 7 A schematic diagram of an ignition recognition model application process in one embodiment;

[0051] Figure 8 A structural block diagram of a model building device in one embodiment;

[0052] Figure 9 is a structural block diagram of a range hood control device in another embodiment;

[0053] Figure 10 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0055] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0056] In an exemplary embodiment, Figure 1 As shown, a model building method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps:

[0057] Step S101: Acquire cooker ignition audio data.

[0058] Step S102 : performing time domain feature extraction processing on the cooker ignition audio data to obtain the time domain features of the cooker ignition audio data.

[0059] Step S103 , performing frequency domain feature extraction processing on the cooker ignition audio data to obtain frequency domain features of the cooker ignition audio data.

[0060] Step S104 : performing feature fusion processing on the time domain features and the frequency domain features to obtain the time-frequency features of the cooker ignition audio data.

[0061] Step S105: construct a stove ignition sound recognition model based on the time-frequency features; the stove ignition sound recognition model is used to be deployed in the range hood; the range hood is used to start when the stove ignition sound recognition model recognizes the stove ignition sound associated with the range hood.

[0062] Among them, the stove ignition audio data can be the sound information generated when the stove is ignited in a kitchen environment. For example, the stove ignition audio data can be the ignition sounds of various brands of stoves collected in a kitchen environment, including the "click" sound during ignition and audio files that have undergone data enhancement processing (mixing in kitchen environment noise, time stretching, and audio gain adjustment).

[0063] The time-domain stream branch structure may be a neural network structure for extracting time-domain features of audio signals. For example, the time-domain stream branch structure may be a network structure including a time-domain convolutional layer and a gated recurrent unit (GRU).

[0064] The time domain feature may be a feature that reflects the structure and change trend of the audio signal in the time dimension.

[0065] The frequency domain stream branch structure may be a neural network structure for extracting frequency domain features of an audio signal. For example, the frequency domain stream branch structure may be a network structure including fast Fourier transform (FFT) and Mel-frequency cepstral coefficients (MFCCs) calculations.

[0066] Among them, the frequency domain features can be features that reflect the characteristics of the audio signal in the frequency dimension. For example, the frequency domain features can be features extracted by FFT and MFCCs, which can reflect the frequency components and energy distribution of the stove ignition sound.

[0067] Among them, feature fusion processing can be the process of combining different features to form a more comprehensive feature representation. For example, feature fusion processing can be the process of automatically adapting the dynamic weight values of time domain and frequency domain features through a learnable attention mechanism to complete the dynamic adaptive fusion of time domain features and frequency domain features.

[0068] Among them, the time-frequency feature can be a feature that simultaneously contains the time domain and frequency domain information of the audio signal. For example, the time-frequency feature can be a comprehensive feature vector obtained by feature fusion processing that simultaneously contains the time domain features and frequency domain features of the stove ignition audio data.

[0069] The stove ignition sound recognition model may be a machine learning model for recognizing the stove ignition sound. For example, the stove ignition sound recognition model may be a lightweight model constructed using a neural architecture search (NAS) algorithm.

[0070] Optionally, the terminal acquires audio data from stove ignitions of various brands. To increase sample data volume and model robustness, the terminal performs data augmentation on the stove ignition audio data, including mixing in kitchen ambient noise, time stretching, and adjusting audio gain to better simulate real-world scenarios. The terminal then uses a Wave-U-Net architecture to denoise and enhance features of the stove ignition audio data, performing preprocessing operations such as noise removal, resampling, cropping, and windowing. The terminal uses a dual-stream time-frequency architecture to extract features from the stove ignition audio data. The time-domain branch captures short-duration pulse features through a time-domain convolutional layer and uses a gated recurrent unit (GRU) to extract dynamic temporal variation features, resulting in the time-domain features of the stove ignition audio data. The frequency-domain branch converts the time-domain signal into a time-spectrum graph using a fast Fourier transform (FFT), extracts frequency-domain features using Mel-frequency cepstral coefficients (MFCCs), and then performs two-dimensional convolution on the Mel-frequency spectrum to obtain the frequency-domain features of the stove ignition audio data. The terminal automatically adapts the dynamic weights of time-domain and frequency-domain features through a learnable attention mechanism, performing feature fusion processing on these features to ultimately derive the time-frequency characteristics of the cooker ignition audio data. Based on these time-frequency features, the terminal designs a search space and employs an optimized differentiable neural architecture search (NAS) algorithm to construct a cooker ignition sound recognition model. Dynamic channel pruning is then used to achieve model lightweighting, making the cooker ignition sound recognition model more suitable for deployment in the range hood's embedded hardware system. This allows the range hood to automatically activate when the cooker ignition sound recognition model identifies the associated cooker ignition sound, eliminating the need for manual activation and enhancing the range hood's intelligence.

[0071] In the above model construction method, by extracting the stove ignition audio features from the two dimensions of time domain and frequency domain, it is beneficial to comprehensively and accurately capture the time domain features and frequency domain features of the stove ignition sound; time domain feature extraction processing of the stove ignition audio data can effectively capture the instantaneous change characteristics of the ignition sound, while frequency domain feature extraction processing of the stove ignition audio data can accurately identify the frequency distribution characteristics of the ignition sound, thereby better distinguishing the ignition sound from the ambient noise; through feature fusion processing, richer and more accurate time-frequency features are obtained, which is beneficial to improve the recognition accuracy and response speed of the stove ignition sound recognition model for the ignition sound, and can quickly and accurately identify the stove ignition event at the early stage of the sound, reduce the time required for the recognition process, and improve the accuracy and precision of the stove ignition sound recognition, thereby effectively reducing the delay in the automatic start-up of the range hood.

[0072] In an exemplary embodiment, time domain feature extraction processing is performed on the stove ignition audio data to obtain the time domain features of the stove ignition audio data, specifically including the following contents: one-dimensional convolution processing is performed on the stove ignition audio data through a time domain stream branching structure to obtain the pulse features of the stove ignition audio data; and time series change feature recognition processing is performed on the pulse features to obtain the time domain features.

[0073] Among them, the one-dimensional convolution processing can be a processing process of performing a one-dimensional convolution operation on the audio data. For example, the one-dimensional convolution processing can be to process the stove ignition audio data through a time domain convolution layer (kernel_size=5, that is, the convolution kernel size=5) to capture short-time pulse features.

[0074] Among them, the pulse feature can be a feature of significant energy changes in an audio signal within a short period of time. For example, the pulse feature can be a short-term energy mutation feature in the stove ignition audio data captured by the time domain convolution layer. These features can reflect the instantaneous sound changes generated during the stove ignition process.

[0075] Among them, the timing change feature recognition processing can be a processing process for identifying the dynamic characteristics of the audio signal that change over time. For example, the timing change feature recognition processing can further process the pulse features through the gated recurrent unit (GRU) to capture the dynamic timing change characteristics during the stove ignition process.

[0076] Among them, the gated recurrent unit (GRU) can be a special recurrent neural network structure. For example, the gated recurrent unit (GRU) can be a recurrent neural network structure with 32 hidden units (hidden=32).

[0077] Optionally, the terminal performs one-dimensional convolution processing on the stove ignition audio data using a one-dimensional convolution layer within the time-domain stream branching structure. This convolution layer uses a one-dimensional convolution kernel with a kernel size of 5 to capture short-term pulse features in the stove ignition audio data, such as the "clicking" sound produced by the stove ignition and other instantaneous energy variation characteristics. After obtaining the pulse features, the terminal uses a gated recurrent unit (GRU) to identify temporal variation characteristics of the pulse features to obtain time-domain features. The GRU, through its internal update gate and reset gate mechanism, effectively captures the dynamic characteristics of the pulse features in the stove ignition audio data over time, thereby identifying the temporal variation patterns during the stove ignition process. The time-domain features obtained after processing by the GRU can fully express the dynamic variation characteristics of the stove ignition sound in the temporal dimension.

[0078] The technical solution provided in this embodiment performs one-dimensional convolution processing and time series change feature recognition processing on the stove ignition audio data through a time domain stream branching structure, which is beneficial to capturing the short-time pulse features and time series dynamic change features in the stove ignition sound, thereby facilitating obtaining more accurate time domain features.

[0079] In an exemplary embodiment, frequency domain feature extraction processing is performed on the stove ignition audio data to obtain the frequency domain features of the stove ignition audio data, specifically including the following contents: Fourier transform processing is performed on the stove ignition audio data through a frequency domain stream branching structure to obtain a time-frequency spectrum of the stove ignition audio data; frequency band feature extraction processing is performed on the time-frequency spectrum to obtain frequency band features of the time-frequency spectrum; and two-dimensional convolution processing is performed on the frequency band features to obtain frequency domain features.

[0080] Among them, the frequency domain stream branch structure can be a neural network structure used to extract frequency domain features of stove ignition audio data. For example, the frequency domain stream branch structure can be the frequency domain branch in the time-frequency dual stream structure, which is specifically used to extract feature information of audio data in the frequency dimension.

[0081] The Fourier transform process may be a process of converting the cooker ignition audio data in the time domain into frequency domain representation. For example, the Fourier transform process may be a signal conversion process from the time domain to the frequency domain implemented by fast Fourier transform (FFT).

[0082] Among them, the time-spectrogram can be a two-dimensional image representation of the energy distribution of the audio signal in the two dimensions of time and frequency. For example, the time-spectrogram can be a graph showing the energy distribution of each frequency component at different time points obtained by performing Fourier transform on the stove ignition audio data.

[0083] The frequency band feature may refer to energy distribution features of different frequency bands in the time-spectrogram. For example, the frequency band feature may be energy features on different frequency bands extracted by applying a filter bank to the time-spectrogram.

[0084] The two-dimensional convolution processing may be a convolution operation for extracting features from frequency band features. For example, the two-dimensional convolution processing may be a two-dimensional convolution operation for extracting features from frequency band features.

[0085] Optionally, the terminal performs Fourier transform processing on the stove ignition audio data through the fast Fourier transform (FFT) module in the frequency domain stream branch structure, converts the time domain audio signal into a frequency domain representation, and obtains a time-frequency spectrum of the stove ignition audio data, which intuitively shows the energy distribution characteristics of the stove ignition audio data on each frequency component; then, the terminal performs frequency band feature extraction processing on the time-frequency spectrum, specifically by applying a filter group to extract the frequency band features of the time-frequency spectrum, to obtain more representative frequency band features of the time-frequency spectrum; finally, the terminal performs two-dimensional convolution processing on the frequency band features, and by applying a two-dimensional convolution operation on the frequency domain features, obtains frequency domain features that can accurately characterize the frequency domain characteristics of the stove ignition audio data.

[0086] The technical solution provided in this embodiment is beneficial for comprehensively capturing the frequency domain characteristics of the stove ignition sound by performing Fourier transform processing, frequency band feature extraction processing and two-dimensional convolution processing on the stove ignition audio data, thereby facilitating the subsequent improvement of the model's recognition accuracy of the stove ignition sound.

[0087] In an exemplary embodiment, feature fusion processing is performed on time domain features and frequency domain features to obtain time-frequency features of stove ignition audio data, specifically including the following contents: determining the dynamic weight value of the time domain features and the dynamic weight value of the frequency domain features through an attention mechanism; and performing feature fusion processing on the time domain features and the frequency domain features according to the dynamic weight value of the time domain features and the dynamic weight value of the frequency domain features to obtain time-frequency features.

[0088] Among them, the attention mechanism can be a computational mechanism for automatically learning and assigning feature importance weights. For example, the attention mechanism can calculate the importance scores of time domain features and frequency domain features through trainable neural network parameters, thereby determining their weight contributions in the feature fusion process.

[0089] Among them, the dynamic weight value can be a feature importance coefficient that is adaptively adjusted according to the characteristics of the current input audio data. For example, the dynamic weight value can be a variable parameter calculated through the attention mechanism based on the actual characteristics of the stove ignition audio data, which can reflect the relative importance of time domain features and frequency domain features.

[0090] The dynamic weight value of the time domain feature may be a coefficient reflecting the importance of the time domain feature calculated through the attention mechanism.

[0091] Among them, the dynamic weight value of the frequency domain feature can be a coefficient reflecting the importance of the frequency domain feature calculated through the attention mechanism.

[0092] Optionally, the terminal constructs an attention network layer, which automatically determines the dynamic weight value α of the time domain feature and the dynamic weight value (1-α) of the frequency domain feature by evaluating the importance of the time domain feature and the frequency domain feature; then, the terminal performs weighted fusion processing on the time domain feature and the frequency domain feature according to the fusion formula Feature=αTimeFeat+(1-α)FreqFeat based on the calculated dynamic weight value α of the time domain feature and the dynamic weight value (1-α) of the frequency domain feature, where Feature represents the time-frequency feature, TimeFeat represents the time domain feature, and FreqFeat represents the frequency domain feature, to obtain the time-frequency feature that contains both time domain and frequency domain information.

[0093] The technical solution provided in this embodiment dynamically determines the weight values of time domain features and frequency domain features through the attention mechanism, which is conducive to adaptively adjusting the feature fusion ratio according to the specific characteristics of different stove ignition audio data, thereby making full use of the complementary information in the time domain and frequency domain, and improving the accuracy of subsequent stove ignition sound recognition.

[0094] In an exemplary embodiment, a stove ignition sound recognition model is constructed based on time-frequency features, which specifically includes the following contents: constructing a stove ignition sound recognition model based on a differentiable search structure and time-frequency features.

[0095] Among them, the differentiable search structure can be a differentiable computing framework for automatically searching and optimizing neural network architectures, which automatically searches for the optimal network structure through gradient descent.

[0096] Optionally, the terminal adopts an optimized differentiable architecture search architecture, realizes the continuity of the architecture parameter α through Gumbel Softmax (differentiable weight), executes all candidate operators in parallel, and automatically learns the operator combination with the optimal weighted fusion result through the gradient descent of α, thereby automatically searching and designing the optimal neural network architecture and network weight parameters for stove ignition sound recognition, and combines time-frequency features to construct a stove ignition sound recognition model.

[0097] The technical solution provided in this embodiment, by constructing a stove ignition sound recognition model based on a differentiable search structure and time-frequency features, is conducive to automatically searching and optimizing the neural network architecture that is most suitable for identifying stove ignition sounds, thereby reducing the complexity of manually designed models and improving the recognition accuracy and generalization ability of the model, so that the range hood can more accurately identify stove ignition sounds and realize the intelligent startup function.

[0098] In an exemplary embodiment, a stove ignition sound recognition model is constructed based on a differentiable search structure and time-frequency features, specifically including the following contents: constructing a basic stove ignition sound recognition model based on the differentiable search structure and time-frequency features; dynamically pruning and compressing the basic stove ignition sound recognition model through a dynamic channel pruning unit to obtain the stove ignition sound recognition model.

[0099] Among them, the basic stove ignition sound recognition model can be a neural network model with complete functions but not yet optimized and compressed, which is initially constructed through a differentiable search structure. For example, the basic stove ignition sound recognition model can be the original network architecture automatically searched by the differentiable search structure according to time-frequency features, which includes the complete number of network layers, channels and connection methods, but has not yet been compressed.

[0100] Among them, the dynamic channel pruning unit can be a model compression module used to reduce the complexity of the neural network. For example, the dynamic channel pruning unit can be a functional unit added to the forward structure of the neural network that can adjust the pruning ratio in real time according to the CPU utilization or battery power, thereby reducing the number of model parameters and calculations by cutting the number of unimportant feature channels.

[0101] Among them, dynamic pruning and compression processing can be a model optimization processing that dynamically adjusts the network structure scale according to the actual operating environment, which is used to reduce the number of model parameters and the amount of calculation.

[0102] Optionally, the terminal inputs the processed time-frequency features into a differentiable search space, which is composed of multiple unit microstructures, each unit can be regarded as a directed acyclic graph (DAG), in which the nodes represent the layers of the neural network and the edges represent the candidate operations in the search space; the terminal converts the discrete candidate operations into a continuous probability distribution through softmax (soft maximum) weighted averaging, so that the architecture parameter α is differentiable, and then the terminal performs feature screening and model training through the optimized and accelerated differentiable search architecture, greatly reduces the amount of calculation by replacing the second-order approximation calculation with the first-order approximation acceleration, and uses the Adam (adaptive moment estimation) optimizer to dynamically and adaptively adjust the learning rate to accelerate parameter convergence, thereby obtaining a basic stove ignition sound recognition model; subsequently, the terminal dynamically prunes and compresses the basic stove ignition sound recognition model through a dynamic channel pruning unit, reduces the number of model parameters and the amount of calculation by cutting the number of unimportant feature channels, and obtains a stove ignition sound recognition model.

[0103] The technical solution provided in this embodiment automatically constructs a basic model based on a differentiable search structure and combines it with a dynamic channel pruning unit for compression. This is beneficial for significantly reducing the number of model parameters and computational complexity while ensuring recognition accuracy, thereby reducing the difficulty of deploying the model in resource-constrained range hood embedded systems and improving the real-time performance and efficiency of stove ignition sound recognition.

[0104] In an exemplary embodiment, a basic stove ignition sound recognition model is constructed based on a differentiable search structure and time-frequency features, specifically including the following contents: constructing a stove ignition sound recognition model to be trained based on the differentiable search structure and time-frequency features; performing a double-layer gradient optimization alternating training process on the stove ignition sound recognition model to be trained through the differentiable search structure to obtain a trained stove ignition sound recognition model; when the trained stove ignition sound recognition model satisfies the termination search mechanism, using the trained stove ignition sound recognition model as the basic stove ignition sound recognition model.

[0105] The stove ignition sound recognition model to be trained may be a neural network model that has been initially constructed but has not yet been trained and optimized.

[0106] Among them, the dual-layer gradient optimization alternating training process can be a model training method that simultaneously updates network weights and architecture parameters through inner and outer layer optimization strategies.

[0107] Among them, the termination search mechanism can be a judgment criterion for terminating training early based on whether the model has reached the optimal state according to the changes in architecture parameters and the performance of the validation set.

[0108] The trained stove ignition sound recognition model may be a stove ignition sound recognition model obtained after training and optimization using a differentiable search structure.

[0109] Optionally, the terminal constructs a stove ignition sound recognition model to be trained based on a differentiable search structure and time-frequency features; then, the terminal performs a double-layer gradient optimization alternating training process on the stove ignition sound recognition model to be trained through the differentiable search structure, the inner layer optimizes the network weight w (minimizing the loss value of the training set), and the outer layer optimizes the architecture parameter α (minimizing the loss value of the validation set), and replaces the second-order approximation calculation with the first-order approximation acceleration, and uses the Adam optimizer to dynamically and adaptively adjust the learning rate to accelerate parameter convergence to obtain the trained stove ignition sound recognition model; when the trained stove ignition sound recognition model satisfies the termination search mechanism, the terminal uses the trained stove ignition sound recognition model as the basic stove ignition sound recognition model.

[0110] The technical solution provided in this embodiment automatically searches for the optimal network architecture and parameters through a differentiable search structure and a two-layer gradient optimization alternating training process, which helps reduce the workload of manually designing the network architecture. At the same time, a termination search mechanism is used to determine the model convergence state, thereby improving model training efficiency, avoiding overtraining, and realizing the automated optimization design of the stove ignition sound recognition model.

[0111] In an exemplary embodiment, obtaining cooker ignition audio data specifically includes the following: obtaining original cooker ignition audio data; performing data enhancement processing on the original ignition audio data to obtain data-enhanced ignition audio data; performing data preprocessing on the data-enhanced ignition audio data to obtain cooker ignition audio data; the data preprocessing includes at least one of noise elimination processing, resampling processing, cropping processing, and windowing processing.

[0112] The original ignition audio data may be unprocessed audio files from stoves of various brands.

[0113] Among them, data enhancement processing can be a series of operations that change the original audio data to increase the data volume and diversity. For example, data enhancement processing can be achieved by mixing in kitchen environment noise, time stretching, and adjusting the audio gain.

[0114] Among them, the ignition audio data after data enhancement can be a richer and more diverse stove ignition audio set obtained through data enhancement processing.

[0115] The noise cancellation process may be a process of removing unnecessary background noise from the audio.

[0116] The resampling process may be a process of changing the sampling rate of the audio data to match the system requirements.

[0117] The cutting process may be an operation of cutting out a valid portion of the audio data.

[0118] The windowing process may be a process of applying a window function during the audio analysis process to reduce spectrum leakage.

[0119] Optionally, the terminal obtains the original ignition audio data of the stove, which includes ignition audio files of various brands of stoves; then, the terminal performs data enhancement processing on the original ignition audio data to obtain data-enhanced ignition audio data. Specifically, the terminal increases the complexity and diversity of the original ignition audio data by mixing in kitchen environmental noise, time stretching, and adjusting the audio gain, so as to better simulate the noise in the real scene and improve the model's recognition ability of ignition audio in environmental noise; the terminal performs data preprocessing on the data-enhanced ignition audio data to obtain stove ignition audio data, wherein the data preprocessing includes noise elimination processing, resampling processing, cropping processing, and windowing processing.

[0120] The technical solution provided in this embodiment improves the quality and diversity of stove ignition audio data by combining data enhancement processing and multiple data preprocessing methods, thereby facilitating improved accuracy and robustness of subsequent stove ignition sound recognition models, enabling range hoods to more accurately detect stove ignition sounds.

[0121] In an exemplary embodiment, Figure 2 As shown, a range hood control method is provided. This embodiment uses the method applied to a range hood controller as an example. In this embodiment, the method includes the following steps:

[0122] Step S201: Collect current audio data.

[0123] Step S202: performing time domain feature extraction processing on the current audio data to obtain the current time domain features of the current audio data.

[0124] Step S203: Perform frequency domain feature extraction processing on the current audio data to obtain current frequency domain features of the current audio data.

[0125] Step S204: performing feature fusion processing on the current time domain features and the current frequency domain features to obtain the current time-frequency features of the current audio data.

[0126] Step S205: Input the current time-frequency features into the stove ignition sound recognition model deployed in the range hood for stove ignition sound recognition processing to obtain the stove ignition sound recognition result of the current audio data; the stove ignition sound recognition model is obtained by any of the above-mentioned model construction methods.

[0127] Step S206 , when the cooker ignition sound recognition result indicates that the cooker ignition sound associated with the range hood is recognized, the range hood is started.

[0128] The current audio data may be the ambient sound signal data collected by the range hood in real time.

[0129] The current time domain feature may be a time dimension feature extracted from the current audio data by the time domain stream branch structure.

[0130] Among them, the current frequency domain feature can be a frequency dimension feature extracted from the current audio data by the frequency domain stream branch structure.

[0131] Among them, the current time-frequency feature can be a comprehensive feature obtained by combining the current time domain feature and the current frequency domain feature through feature fusion processing. For example, the current time-frequency feature can be a feature vector containing time domain and frequency domain information obtained by weighted fusion of time domain features and frequency domain features under the attention mechanism.

[0132] Among them, the stove ignition sound recognition result can be the recognition judgment information output by the stove ignition sound recognition model after judging the current time-frequency features of the input. For example, the stove ignition sound recognition result can be the binary classification result output by the model indicating whether the stove ignition sound is detected.

[0133] Optionally, the range hood controller collects current audio data through a built-in microphone, and the current audio data includes real-time sounds in the kitchen environment; then, the range hood controller performs time domain feature extraction processing on the current audio data through a time domain stream branching structure to obtain current time domain features of the current audio data. At the same time, the range hood controller performs frequency domain feature extraction processing on the current audio data through a frequency domain stream branching structure to obtain current frequency domain features of the current audio data; then, the range hood controller performs feature fusion processing on the current time domain features and the current frequency domain features to obtain current time-frequency features of the current audio data. During feature fusion, the dynamic weight values of the time domain and frequency domain features can be automatically adapted through a learnable attention mechanism; the range hood controller inputs the current time-frequency features into a stove ignition sound recognition model deployed in the range hood to perform stove ignition sound recognition processing to obtain a stove ignition sound recognition result of the current audio data; when the stove ignition sound recognition result indicates that the stove ignition sound associated with the range hood is recognized, the range hood controller starts the range hood.

[0134] In the above range hood control method, by adopting the time domain stream branching structure and the frequency domain stream branching structure, the time domain features and frequency domain features of the audio data are synchronously extracted, and the complementary advantages of the two features are fully utilized through the feature fusion processing mechanism to construct a more comprehensive and accurate time-frequency feature, which is beneficial to improving the efficiency and accuracy of stove ignition sound recognition; in addition, the optimized stove ignition sound recognition model is directly deployed in the range hood for real-time processing, which is beneficial to realize the instant automatic start-up of the range hood, thereby reducing the delay of the automatic start-up of the range hood.

[0135] The following uses an application example to illustrate the model building method and range hood control method provided in this application. This application example uses the method applied to a terminal as an example.

[0136] With the advent of the intelligent era, household appliances are gradually transitioning from traditional models to intelligent ones. Currently, most range hoods on the market still require manual activation. Even those that use infrared technology to sense temperature changes suffer from drawbacks such as time delays. To further enhance the intelligence of range hoods, this application example utilizes a neural architecture search algorithm based on an accelerated optimized differentiable search architecture to propose a speech recognition method for intelligently identifying the sound of stove ignition. This technology enables intelligent activation of the range hood in conjunction with the range hood, eliminating the need for manual activation.

[0137] This application example proposes a speech recognition method for intelligent range hood recognition of stove ignition sounds based on the Neural Architecture Search (NAS) algorithm with a fast differentiable search architecture. The specific steps of the algorithm are as follows:

[0138] Step 1: Prepare ignition audio files of various brands of stoves. In order to increase the sample data volume and the robustness of the subsequent model, data enhancement processing is performed based on the existing ignition audio dataset before feature extraction. The specific approach is to mix in kitchen environmental noise, time stretch, and adjust the audio gain. The purpose is to increase the complexity and diversity of the data and better simulate the noise in real scenes to improve the model's ability to recognize ignition audio in environmental noise; then, the audio files are denoised, resampled (16kHz, covering the main frequency band of ignition sound 0-8kHz, where kHz represents kilohertz), cropped, and windowed. Among them, this algorithm uses the Wave-U-Net (waveform U network) structure to perform denoising and feature enhancement operations on audio data. Its network structure refers to Figure 3 , including: audio data input, encoder module, bottleneck layer, decoder module, output layer.

[0139] The Wave-U-Net implementation process is as follows: The original noisy speech waveform is first fed into the encoder as raw data. The encoder consists of three downsampling (DS) blocks, each of which includes a one-dimensional convolutional layer and a downsampling layer. As the audio data passes through multiple layers of encoding, the resolution of the feature map gradually decreases, but the level of feature abstraction gradually increases. The extracted feature data then enters a bottleneck layer consisting of two one-dimensional convolutional layers for further processing and integration to obtain a more representative feature representation. The enhanced features are then decoded according to the encoding process, followed by refinement and restoration of the features through a one-dimensional convolutional layer, gradually restoring the resolution of the feature map to its original size. After processing by the decoder, the output is a speech waveform with the same shape as the original input speech waveform. During the modeling process, the loss function uses the mean squared error (MSE) to measure the difference between the estimated useful firing signal and the clean noise signal, and the model parameters are trained by minimizing this loss function. By processing the original audio in time and frequency, the speech quality degradation caused by missing or inaccurate phase information in traditional frequency-domain methods is avoided.

[0140] Step 2: Feature extraction is performed on the preprocessed audio data using a dual-stream time-frequency architecture suitable for ignition audio. This architecture consists of two parallel neural network branches, one in the time domain and the other in the frequency domain. The audio signal undergoes a series of convolution and pooling operations in the time domain branch to obtain time-domain features. These features reflect the temporal structure and changing trends of the audio signal, such as the onset, end, and duration of the sound. The frequency domain branch extracts frequency-domain features using methods such as fast Fourier transforms (FFTs) and the calculation of Mel-frequency cepstral coefficients (MFCCs). The FFT transforms the time domain into the frequency domain, calculates the magnitude spectrum of the FFT, and obtains the amplitude of each frequency component. MFCCs are then used to extract frequency-domain features. The high-level feature representations obtained from the time and frequency domains are dynamically weighted and fused to form a complete feature vector that encompasses both the time and frequency domain features of the audio signal. This fused feature vector is then processed using convolutional or fully connected layers of the neural network to extract high-level semantic information from the audio signal, enabling rapid and accurate recognition of stove ignition audio.

[0141] Step 3: Design the search space. First, define the candidate set of basic operations, including 1D convolution (k=3), 1D convolution (k=5), dilated convolution (d=2), GRU unit (hidden=32), depthwise separable convolution, skip connection, max pooling layer (k=2), and temporal attention; define the supernet structure parameters, including the maximum network depth (6), the number of channel candidates (16, 64, 128), and the input feature type (raw audio waveform, MFCCs). Then, adopt an optimized differentiable architecture search architecture, achieve the continuity of the architecture parameter α through Gumbel Softmax, execute all candidate operators (convolution / pooling / GRU / attention) in parallel, and automatically learn the operator combination with the optimal weighted fusion result through gradient descent of α. At the same time, latency-aware dynamic pruning is introduced to estimate the inference time of the current model on the target hardware. When the cumulative delay exceeds the set threshold (target latency = 15.0 milliseconds), channel_prune is used to reduce the number of feature channels in real time according to a predetermined ratio (ratio = 0.5). This immediately reduces the amount of computation, improves network performance and efficiency, and accelerates model convergence, thereby obtaining a flexible and efficient ignition sound recognition model architecture.

[0142] Step 4: Divide the dataset. Divide all preprocessed audio data into training set, test set, and validation set in a ratio of 7:2:1. The validation set is used to adjust the model's hyperparameters during training, and the test set is used to evaluate the performance of the final model.

[0143] Step 5: Model training, load preprocessed data, initialize the search space network hyperparameters (maximum depth max_depth = 6, width range width_range = 16, resolution level resolution_levels = MFCCs) and the loss function. The loss function adopts the cross entropy loss function (Cross Entropy Loss). First, calculate the loss value of a single label category. The total cross entropy is the sum of the cross entropy of each category in the multi-label classification task. The calculation formula is as follows:

[0144]

[0145] where y i Indicates the corresponding true value category, p i Represents the corresponding category probability.

[0146] Iterative training is then performed (training cycle epoch=100, batch size batch_size=10). Through the optimized and accelerated differentiable search architecture, feature screening, model training, candidate operator combination optimization, and dynamic pruning and compression are performed, ultimately obtaining a lightweight model for speech recognition of stove ignition sounds.

[0147] Ignition recognition model training process reference Figure 4 , including: audio data input, data enhancement, Wave-U-Net denoising, resampling, cropping, windowing, and then divided into time domain convolution, GRU feature extraction and FFT, MFCCs processing, and then attention weighted dynamic feature fusion, and by dividing the data set, optimized and accelerated differentiable search is performed, followed by model training, pruning deployment, and recognition application.

[0148] Step 6: Deploy the trained model to the range hood embedded hardware system. Use the hardware data acquisition module to monitor the ignition sound in real time and pass it to the ignition audio recognition model after preprocessing. When the model recognizes the ignition sound of the stove in the kitchen, the range hood is intelligently started through program logic.

[0149] Compared with other speech recognition algorithms, the lightweight recognition model for stove ignition sound is established by using a neural architecture search algorithm. The optimal operation combination is automatically learned during training through differentiable weights (Gumbel Softmax). It can automatically search and design the optimal neural network architecture and network weight parameters for stove ignition sound recognition without the need for subjective human design of the network structure, which greatly improves the efficiency of model design. The optimization of the ignition recognition algorithm applied for includes: (1) introducing a time-frequency dual-stream structure based on the temporal characteristics of the ignition audio data, so that it focuses more on the extraction of feature information in the time domain and frequency domain of the ignition audio data, optimizes the feature extraction process, and improves the data feature mining capability; (2) optimizing the search strategy of the differentiable architecture, ensuring accuracy while greatly improving the operating efficiency. The specific approach is to adopt first-order approximation acceleration, dynamically adjust the learning rate parameters and design a termination search mechanism; (3) adding a dynamic channel pruning module to compress the number of network channels, accelerate the model convergence speed, dynamically compress the model, and achieve model lightweight; (4) providing a practical solution for the intelligent opening of the range hood.

[0150] 1. Time-frequency dual-stream structure:

[0151] A dual-stream parallel architecture for time-frequency separation is designed to process audio information separately. The time branch first performs a 1D convolution (kernel_size=5) on the preprocessed raw audio waveform to capture short-term pulse features (the "clicking" sound during ignition). A gated recurrent unit (GRU) is then used to capture the dynamic temporal variations of the ignition process. The frequency branch performs a short-time Fourier transform on the input audio signal to convert it into a time-spectrogram. A Mel filter bank is used to extract frequency band features. A 2D convolution is then performed on the Mel spectrum to obtain ignition features with a prominent energy distribution. Finally, a learnable attention mechanism automatically adapts the dynamic weights of the time and frequency domain features, achieving dynamic and adaptive feature fusion. The fusion formula is as follows: Feature = αTimeFeat + (1-α)FreqFeat. The weight α is derived from the attention mechanism's assessment of feature importance. This dual-stream architecture simultaneously captures the temporal and frequency characteristics of the audio signal, improving the model's accuracy in ignition sound recognition.

[0152] Time-frequency dual-stream structure reference Figure 5 , including: the preprocessed data is input into the time-frequency dual-stream structure, the time-frequency dual-stream structure includes a time domain stream branch and a frequency domain stream branch, the time domain stream branch includes time domain convolution and GRU feature extraction, the frequency domain stream branch includes FFT (Fast Fourier Transform) and MFCCs, and then attention weighted dynamic feature fusion is performed.

[0153] 2. Accelerated Differentiable Search Structure:

[0154] The search space of the differentiable search architecture consists of L unit microstructures, each of which can be viewed as a directed acyclic graph (DAG) with N nodes and M edges. The nodes represent the layers of the neural network, and the edges represent the candidate operations O in the search space. A network containing multiple operations on each edge is called a relaxed supernetwork. During the training process of convergence to the structure, all sub-operations are present and participate in the training. Finally, the discrete candidate operations are converted to a continuous probability distribution through softmax weighted averaging, making the architecture parameter α differentiable. The candidate operation on each edge in the search space is relaxed into a continuous probability distribution as follows:

[0155]

[0156] in is the architectural parameter of the kth candidate operation on edge (i, j), o k is a candidate operation (such as convolution, pooling). The network weight w and the architecture parameter α are the inner and outer double-layer optimization strategies, and the inner network weight Optimization is achieved by minimizing the loss value of the training set; the outer layer architecture parameter α is optimized by minimizing the loss value of the validation set, and w* is the solution of the inner layer optimization.

[0157] Due to the huge computational cost of inner layer optimization, the actual architecture parameters α and network weights w of the differentiable search architecture are optimized using the second-order approximation (implicit function theorem) to calculate the validation set loss function L val The gradient of the architecture parameter α ( ; As shown in the following formula:

[0158]

[0159] in is the learning rate of the network weight w, is the validation set loss function L val The gradient of the network weight w.

[0160] However, considering the Hessian matrix (the second-order partial derivative matrix of the function) The calculation of the second-order gradient results in too many calculation steps and time costs, and consumes too much resources. Considering that when the inner layer optimization converges quickly, w≈w*, the influence of the second-order term is small at this time, and the first-order approximation is sufficient. Therefore, this algorithm ignores the calculation of the second-order term on the original basis and directly uses a gradient to alternately update the network weight w and the architecture parameter α. The specific formula is as follows:

[0161]

[0162]

[0163]

[0164] By replacing second-order approximation calculations with first-order approximation acceleration, the amount of calculation is expected to be reduced by 30%-50%, greatly improving training efficiency.

[0165] Secondly, the network weights w and architecture parameters α of the differentiable search architecture are both optimized using SGD (stochastic gradient descent). The learning rate hyperparameters for updating w and α are fixed values, which can easily lead to unstable network convergence speed. This algorithm accelerates network convergence by dynamically and adaptively adjusting the learning rate. The specific approach is as follows: First, initialize the learning rate of the architecture parameter α to a large value ( = 0.01), Adam (adaptive moment estimation) was used to replace the original SGD optimizer. The learning rate was dynamically and adaptively adjusted according to the historical gradient to accelerate the parameter convergence speed. At the same time, Adam's momentum mechanism smoothed the gradient noise and alleviated the instability of the architecture parameters:

[0166]

[0167]

[0168]

[0169] Among them, m t is the first-order momentum, v t is the second-order momentum, and are the decay rates of momentum, which are set to 0.9 and 0.999 respectively.

[0170] Finally, the differentiable search architecture usually uses a fixed number of training times to obtain the optimal solution. Regardless of whether the optimal solution of the network architecture and parameters is obtained, the training is performed for the initial set number of iterations. This algorithm terminates the search early based on the performance saturation of the architecture parameters and the validation set to avoid the possibility of redundant calculations. ( The search is terminated by a zero-prevention constant (usually 1e-8), or by verifying that the loss value no longer decreases within 10 rounds.

[0171] According to the above operations, this algorithm uses first-order approximation to replace second-order approximation calculation, Adam adaptively adjusts the learning rate, and adds a termination search mechanism on the basis of the original differentiable search architecture, which reduces the consumption of computing resources and greatly improves the convergence speed.

[0172] Differentiable Search Structure (Differentiable Search Architecture) Execution Process Reference Figure 6 , including: search space definition and initialization, then candidate operation initialization, architecture parameter initialization, network weight initialization, followed by continuous relaxation and super network construction, two-layer gradient optimization alternating training, termination mechanism judgment, and architecture discretization to obtain the optimal model.

[0173] 3. Dynamic channel pruning to achieve model lightweighting:

[0174] A dynamic channel pruning module is added to the NAS forward neural network structure. The pruning module is activated according to CPU utilization or battery power, and the pruning ratio is adjusted in real time to reduce the number of channels, thereby reducing the number of parameters and data calculations, improving network performance and efficiency, accelerating model convergence, and achieving dynamic compression of complex models.

[0175] 4. Real-time monitoring and identification, intelligent start of range hood:

[0176] The NAS ignition sound lightweight recognition model is deployed in the range hood embedded system. When the ignition sound of the stove in the kitchen is collected, it undergoes the same data preprocessing operations as described above (denoising, resampling, cropping and windowing) and time-frequency dual-stream feature extraction. The feature data is then input into the model for prediction. After the embedded system obtains the ignition recognition positive and negative feedback data from the model, it executes the range hood startup program to achieve intelligent and rapid startup of the range hood.

[0177] Ignition Identification Model Application Process Reference Figure 7 , including: audio data collection at the beginning, Wave-U-Net denoising, resampling, cropping, windowing, then divided into time domain convolution, GRU feature extraction and FFT, MFCCs processing, then attention weighted dynamic feature fusion, followed by model prediction, to determine whether it is the ignition sound, if not, return to the audio data collection step, if so, start the smoke machine.

[0178] The technical solution provided in this application example is conducive to comprehensively and accurately capturing the time domain features and frequency domain features of the stove ignition sound by extracting the stove ignition audio features from both the time domain and frequency domain dimensions; performing time domain feature extraction processing on the stove ignition audio data can effectively capture the instantaneous change characteristics of the ignition sound, while performing frequency domain feature extraction processing on the stove ignition audio data can accurately identify the frequency distribution characteristics of the ignition sound, thereby better distinguishing the ignition sound from the ambient noise; through feature fusion processing, richer and more accurate time-frequency features are obtained, which is conducive to improving the recognition accuracy and response speed of the stove ignition sound recognition model for the ignition sound, and can quickly and accurately identify the stove ignition event at the initial stage of the sound, reduce the time required for the recognition process, and improve the accuracy and precision of the stove ignition sound recognition, thereby effectively reducing the delay in the automatic start-up of the range hood.

[0179] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0180] Based on the same inventive concept, embodiments of the present application further provide a model building device for implementing the aforementioned model building method and a range hood control device for implementing the aforementioned range hood control method. The implementation solutions provided by these devices are similar to those described in the aforementioned methods. Therefore, the specific limitations of one or more of the following embodiments of the model building device and range hood control device can be found in the above-described limitations of the model building method and range hood control method, and will not be further elaborated here.

[0181] In an exemplary embodiment, Figure 8 As shown, a model building device is provided, and the model building device 800 may include:

[0182] The data acquisition module 801 is used to acquire the cooker ignition audio data;

[0183] The first processing module 802 is configured to perform time domain feature extraction processing on the cooker ignition audio data to obtain time domain features of the cooker ignition audio data;

[0184] The second processing module 803 is used to perform frequency domain feature extraction processing on the cooker ignition audio data to obtain frequency domain features of the cooker ignition audio data;

[0185] The third processing module 804 is used to perform feature fusion processing on the time domain features and the frequency domain features to obtain the time-frequency features of the cooker ignition audio data;

[0186] The model building module 805 is used to build a stove ignition sound recognition model based on time-frequency features; the stove ignition sound recognition model is used to be deployed in the range hood; the range hood is used to start when the stove ignition sound associated with the range hood is recognized by the stove ignition sound recognition model.

[0187] In an exemplary embodiment, the first processing module 802 is also used to perform one-dimensional convolution processing on the stove ignition audio data through a time domain stream branching structure to obtain the pulse characteristics of the stove ignition audio data; and perform time series change feature recognition processing on the pulse characteristics to obtain time domain features.

[0188] In an exemplary embodiment, the second processing module 803 is also used to perform Fourier transform processing on the stove ignition audio data through the frequency domain stream branch structure to obtain a time-frequency spectrum of the stove ignition audio data; perform frequency band feature extraction processing on the time-frequency spectrum to obtain frequency band features of the time-frequency spectrum; and perform two-dimensional convolution processing on the frequency band features to obtain frequency domain features.

[0189] In an exemplary embodiment, the third processing module 804 is also used to determine the dynamic weight value of the time domain feature and the dynamic weight value of the frequency domain feature through the attention mechanism; based on the dynamic weight value of the time domain feature and the dynamic weight value of the frequency domain feature, the time domain feature and the frequency domain feature are subjected to feature fusion processing to obtain the time-frequency feature.

[0190] In an exemplary embodiment, the model building module 805 is further configured to build a stove ignition sound recognition model based on a differentiable search structure and time-frequency features.

[0191] In an exemplary embodiment, the model construction module 805 is also used to construct a basic stove ignition sound recognition model based on a differentiable search structure and time-frequency features; and to perform dynamic pruning and compression processing on the basic stove ignition sound recognition model through a dynamic channel pruning unit to obtain a stove ignition sound recognition model.

[0192] In an exemplary embodiment, the model construction module 805 is further used to construct a stove ignition sound recognition model to be trained based on a differentiable search structure and time-frequency features; through the differentiable search structure, the stove ignition sound recognition model to be trained is subjected to a double-layer gradient optimization alternating training process to obtain a trained stove ignition sound recognition model; when the trained stove ignition sound recognition model satisfies the termination search mechanism, the trained stove ignition sound recognition model is used as the basic stove ignition sound recognition model.

[0193] In an exemplary embodiment, the data acquisition module 801 is also used to obtain the original ignition audio data of the stove; perform data enhancement processing on the original ignition audio data to obtain data-enhanced ignition audio data; perform data preprocessing on the data-enhanced ignition audio data to obtain stove ignition audio data; data preprocessing includes at least one of noise elimination processing, resampling processing, shearing processing and windowing processing.

[0194] In an exemplary embodiment, Figure 9 As shown, a range hood control device is provided, and the range hood control device 900 may include:

[0195] Data acquisition module 901, used to collect current audio data;

[0196] The fourth processing module 902 is configured to perform time domain feature extraction processing on the current audio data to obtain current time domain features of the current audio data;

[0197] The fifth processing module 903 is used to perform frequency domain feature extraction processing on the current audio data to obtain the current frequency domain features of the current audio data;

[0198] The sixth processing module 904 is configured to perform feature fusion processing on the current time domain feature and the current frequency domain feature to obtain the current time-frequency feature of the current audio data;

[0199] The feature input module 905 is configured to input the current time-frequency features into a stove ignition sound recognition model deployed in the range hood for stove ignition sound recognition processing, thereby obtaining a stove ignition sound recognition result for the current audio data; the stove ignition sound recognition model is obtained by any of the above-mentioned model construction methods;

[0200] The range hood starting module 906 is configured to start the range hood when the stove ignition sound recognition result indicates that the stove ignition sound associated with the range hood is recognized.

[0201] Each module in the aforementioned model building device and range hood control device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor within a computer device in the form of hardware, or may be stored in a computer device memory in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0202] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 10 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, while the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means may be implemented via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a model building method and / or a range hood control method. The display unit of the computer device is used to produce a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0203] Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0204] In an exemplary embodiment, a range hood is provided, comprising a range hood body, a memory, and a processor. The memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0205] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0206] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0207] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0208] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0209] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A model building method, characterized in that: The method comprises: Get the stove ignition audio data; Performing time domain feature extraction processing on the cooker ignition audio data to obtain time domain features of the cooker ignition audio data; Performing frequency domain feature extraction processing on the cooker ignition audio data to obtain frequency domain features of the cooker ignition audio data; Performing feature fusion processing on the time domain features and the frequency domain features to obtain time-frequency features of the cooker ignition audio data; Based on the time-frequency features, a stove ignition sound recognition model is constructed; the stove ignition sound recognition model is used to be deployed in a range hood; the range hood is used to start when the stove ignition sound associated with the range hood is recognized by the stove ignition sound recognition model.

2. The method according to claim 1, characterized in that The performing time domain feature extraction processing on the cooker ignition audio data to obtain the time domain features of the cooker ignition audio data includes: Performing one-dimensional convolution processing on the cooker ignition audio data through a time-domain stream branching structure to obtain pulse features of the cooker ignition audio data; The pulse features are subjected to time series variation feature recognition processing to obtain the time domain features.

3. The method according to claim 1, characterized in that The performing frequency domain feature extraction processing on the cooker ignition audio data to obtain the frequency domain features of the cooker ignition audio data includes: Performing Fourier transform processing on the cooker ignition audio data through a frequency domain stream branching structure to obtain a time-frequency spectrum of the cooker ignition audio data; Performing frequency band feature extraction processing on the time-spectrum graph to obtain frequency band features of the time-spectrum graph; Perform two-dimensional convolution processing on the frequency band features to obtain the frequency domain features.

4. The method according to claim 1, wherein The performing feature fusion processing on the time domain features and the frequency domain features to obtain the time-frequency features of the cooker ignition audio data includes: Determining the dynamic weight value of the time domain feature and the dynamic weight value of the frequency domain feature through an attention mechanism; According to the dynamic weight value of the time domain feature and the dynamic weight value of the frequency domain feature, feature fusion processing is performed on the time domain feature and the frequency domain feature to obtain the time-frequency feature.

5. The method according to claim 1, wherein The method of constructing a stove ignition sound recognition model based on the time-frequency features includes: Based on the differentiable search structure and the time-frequency features, the stove ignition sound recognition model is constructed.

6. The method according to claim 5, characterized in that The stove ignition sound recognition model is constructed based on the differentiable search structure and the time-frequency features, including: Based on the differentiable search structure and the time-frequency features, a basic stove ignition sound recognition model is constructed; The basic stove ignition sound recognition model is dynamically pruned and compressed by a dynamic channel pruning unit to obtain the stove ignition sound recognition model.

7. The method according to claim 6, characterized in that The method of constructing a basic stove ignition sound recognition model based on the differentiable search structure and the time-frequency features includes: Constructing a stove ignition sound recognition model to be trained based on the differentiable search structure and the time-frequency features; Performing a double-layer gradient optimization alternating training process on the stove ignition sound recognition model to be trained by using the differentiable search structure to obtain a trained stove ignition sound recognition model; In a case where the trained stove ignition sound recognition model satisfies a termination search mechanism, the trained stove ignition sound recognition model is used as the basic stove ignition sound recognition model.

8. The method according to claim 1, characterized in that The method of obtaining the cooker ignition audio data includes: Get the raw ignition audio data of the stove; Performing data enhancement processing on the original ignition audio data to obtain data-enhanced ignition audio data; The data-enhanced ignition audio data is subjected to data preprocessing to obtain the cooker ignition audio data; the data preprocessing includes at least one of noise elimination processing, resampling processing, shearing processing and windowing processing.

9. A range hood control method, characterized in that: The method comprises: Collect current audio data; Performing time domain feature extraction processing on the current audio data to obtain current time domain features of the current audio data; Performing frequency domain feature extraction processing on the current audio data to obtain current frequency domain features of the current audio data; Performing feature fusion processing on the current time domain feature and the current frequency domain feature to obtain current time-frequency features of the current audio data; Inputting the current time-frequency features into a stove ignition sound recognition model deployed in the range hood to perform stove ignition sound recognition processing to obtain a stove ignition sound recognition result of the current audio data; the stove ignition sound recognition model is obtained by the method described in any one of claims 1 to 8; When the cooker ignition sound recognition result indicates that the cooker ignition sound associated with the range hood is recognized, the range hood is started.

10. A range hood, comprising a range hood body, a memory, and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.