Human-computer Interaction System and Method for Smart Refrigerator
By analyzing the environmental noise characteristics in real time and dynamically adjusting the noise reduction strategy, the human-computer interaction system of the smart refrigerator solves the problem of complex environmental noise affecting speech recognition, achieving more efficient voice signal processing and better user experience.
Patent Information
- Application Number
- CN202510362257.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-26
AI Technical Summary
When the human-computer interaction system of smart refrigerators faces complex environmental noise, the traditional noise reduction method is not effective, resulting in an increase in voice recognition error rate and affecting the user experience.
By analyzing the environmental noise characteristics in real time, dynamically adjusting the noise reduction strategy, using a microphone array to collect data for preliminary noise reduction processing, and then adaptive noise reduction is performed based on the environmental noise information to improve the clarity of the voice signal.
It effectively reduces the voice recognition error rate, improves the accuracy and user satisfaction of human-computer interaction in smart refrigerators, and ensures that users' commands can be effectively captured in noisy environments.
Smart Images

Figure CN119889322B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of smart refrigerators, and more specifically, to a human-computer interaction system and method for a smart refrigerator. Background Art
[0002] With the development of smart home technology, smart refrigerators, as an important part of the home Internet of Things, are gradually becoming an important bridge connecting users with home life. The human-computer interaction system is one of the key factors to improve the user experience, and speech recognition technology plays a crucial role in it. However, in practical applications, the diversity of environmental noise poses a huge challenge to speech recognition. For example, different types of noises such as the range hood, TV sound, and conversation sound in the kitchen will interfere with the clarity of the speech signal, resulting in an increase in the recognition error rate. Therefore, developing a human-computer interaction system that can effectively cope with complex environmental noise is crucial for improving the user experience of smart refrigerators.
[0003] Traditional noise reduction methods for human-computer interaction systems often rely on fixed algorithm models, such as spectral subtraction, Wiener filtering, etc. Although these methods perform well in specific scenarios, a single noise reduction algorithm cannot effectively handle all situations, and its effect is greatly reduced when facing changing environmental noise. In addition, traditional noise reduction algorithms for human-computer interaction systems usually work based on preset fixed parameters and are difficult to automatically adjust their behavior according to the changes in the noise characteristics of the actual scenario. For example, for spectral subtraction, it relies on the estimated noise power spectral density to subtract noise from the speech signal. When the noise type or intensity changes, if the parameters cannot be adjusted accordingly, it may lead to over-subtraction (speech distortion) or under-subtraction (noise residue), affecting the effectiveness of the human-computer interaction of smart refrigerators and reducing the user experience.
[0004] Therefore, an optimized human-computer interaction system for a smart refrigerator is desired. Summary of the Invention
[0005] This application provides a human-computer interaction system and method for a smart refrigerator, which can dynamically adjust the noise reduction strategy by real-time analyzing the environmental noise characteristics, thereby more effectively performing noise reduction processing on the user interaction instruction speech signal, making it more efficient and accurate to extract useful information from the original audio data, and improving the accuracy of the human-computer interaction of the smart refrigerator and the user satisfaction in a more intelligent way.
[0006] In a first aspect, a human-computer interaction system for a smart refrigerator is provided, including: a user interaction instruction voice acquisition module, configured to collect user interaction instruction voice signals using a microphone array built in the refrigerator after receiving a start voice interaction action.
[0007] The primary noise reduction module for the command voice signal is used to perform primary noise reduction on the user interaction command voice signal to obtain the user interaction command voice signal after primary noise reduction.
[0008] The environmental noise signal acquisition module is used to collect the environmental noise signal by using the microphone array.
[0009] The command voice signal adaptive noise reduction module is used to perform adaptive noise reduction on the user interaction command voice signal after primary noise reduction based on the environmental noise signal to obtain the user interaction command voice signal after secondary noise reduction.
[0010] The interactive command voice recognition module is used to send the user interaction command voice signal after secondary noise reduction to the voice recognition model to obtain the text description of the user interaction command.
[0011] The command parsing module is used to parse the text description of the user interaction command to obtain the parsing result of the command content.
[0012] The refrigerator operation control feedback module is used to control the refrigerator by calling the operation instruction based on the parsing result of the command content and feedback the operation result.
[0013] In a second aspect, a human-computer interaction method for an intelligent refrigerator is provided, including: after receiving the start voice interaction action, collecting the user interaction command voice signal by using the built-in microphone array of the refrigerator; performing primary noise reduction on the user interaction command voice signal to obtain the user interaction command voice signal after primary noise reduction; collecting the environmental noise signal by using the microphone array; performing adaptive noise reduction on the user interaction command voice signal after primary noise reduction based on the environmental noise signal to obtain the user interaction command voice signal after secondary noise reduction; sending the user interaction command voice signal after secondary noise reduction to the voice recognition model to obtain the text description of the user interaction command; parsing the text description of the user interaction command to obtain the parsing result of the command content; controlling the refrigerator by calling the operation instruction based on the parsing result of the command content and feedback the operation result.
[0014] A human-computer interaction system and method for an intelligent refrigerator provided by the present application preliminarily denoises the data collected by a microphone array, and then adaptively further purifies the voice signal of the user interaction instruction according to the simultaneously collected environmental noise information. This two-stage denoising process not only takes into account basic physical-level denoising, but also incorporates an understanding of context information to achieve a better voice clarity restoration effect. Finally, speech recognition technology and command parsing are used to output operation instructions for controlling the refrigerator. In this way, the denoising strategy can be dynamically adjusted by real-time analyzing the characteristics of environmental noise, so as to more effectively perform denoising processing on the voice signal of the user interaction instruction, making it more efficient and accurate to extract useful information from the original audio data, and improving the accuracy and user satisfaction of the human-computer interaction of the intelligent refrigerator in a more intelligent way. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present application and do not limit the present application.
[0016] Figure 1 It is a schematic block diagram of the human-computer interaction system of the intelligent refrigerator according to the embodiment of the present application.
[0017] Figure 2 It is a schematic diagram of the data flow of the human-computer interaction system of the intelligent refrigerator according to the embodiment of the present application.
[0018] Figure 3 It is a schematic block diagram of the adaptive noise reduction module for the command voice signal in the human-computer interaction system of the intelligent refrigerator according to the embodiment of the present application.
[0019] Figure 4 It is a schematic block diagram of the environmental noise-voice signal feature interaction processing unit in the human-computer interaction system of the intelligent refrigerator according to the embodiment of the present application.
[0020] Figure 5 It is a schematic flowchart of the human-computer interaction method of the intelligent refrigerator according to the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall also fall within the scope of protection of the present application.
[0022] In view of the above technical problems, in the technical solution of the present application, a human-computer interaction system for an intelligent refrigerator is proposed, asFigure 1 and Figure 2 As shown in Figure 2 , the human-computer interaction system of the smart refrigerator includes: a user interaction instruction voice acquisition module 10, configured to collect a user interaction instruction voice signal by using a microphone array built in the refrigerator after receiving a start voice interaction action; an instruction voice signal primary noise reduction module 20, configured to perform primary noise reduction on the user interaction instruction voice signal to obtain a primary noise-reduced user interaction instruction voice signal; an environmental noise signal acquisition module 30, configured to collect an environmental noise signal by using the microphone array; an instruction voice signal adaptive noise reduction module 40, configured to perform adaptive noise reduction on the primary noise-reduced user interaction instruction voice signal based on the environmental noise signal to obtain a secondary noise-reduced user interaction instruction voice signal; an interaction instruction voice recognition module 50, configured to send the secondary noise-reduced user interaction instruction voice signal to a voice recognition model to obtain a user interaction instruction text description; a command parsing module 60, configured to perform command parsing on the user interaction instruction text description to obtain a command content parsing result; and a refrigerator operation control feedback module 70, configured to control the refrigerator by invoking an operation instruction based on the command content parsing result and feedback an operation result.
[0023] Exemplarily, in the user interaction instruction voice acquisition module 10, after receiving a start voice interaction action, a user interaction instruction voice signal is collected by using a microphone array built in the refrigerator. It should be understood that by detecting a specific start voice interaction action (such as pressing a voice button or recognizing a wake-up word), the possibility of false triggering can be effectively reduced, ensuring that the system only starts recording when the user actually wants to interact with the refrigerator, thereby improving the response accuracy of the system and the user experience. At the same time, not continuously listening to all sounds in the environment can avoid unnecessary audio data collection and protect the privacy and security of users. Only when a start command is clearly received will the microphone be activated for recording, reducing interference with the user's private life. Continuous listening consumes a large amount of computing resources and power, while the on-demand activation method can reduce energy consumption while ensuring functionality, extend the service life of the device, and reserve more resources for other tasks.
[0024] In one embodiment, the activation of the voice interaction is that the user presses the voice button or says the wake-up word. Specifically, in the present application, high-quality microphone arrays are installed at different positions inside the refrigerator to cover as large a pickup range as possible and ensure that the user's voice commands can be clearly captured even in a noisy environment. At the same time, a low-power front-end processing module is implemented to monitor in real time whether a preset wake-up word is spoken. Once the wake-up word is detected, the main processor is immediately notified to prepare to receive subsequent instructions. At the same time, to optimize the user experience, the system also provides an intuitive feedback mechanism, such as a visual indicator or a prompt tone, to let the user know that their voice has been successfully received and is being processed. Additionally, in other examples of the present application, multiple activation methods are supported. In addition to voice wake-up, a physical button can also be set as an alternative solution to meet the habitual needs of different users. In summary, through the combination of a reasonable hardware layout and an efficient software algorithm, the acquisition process of the user interaction instruction voice signal after receiving the activation of the voice interaction is effectively realized, which not only improves the accuracy and efficiency of the interaction but also enhances the user's satisfaction.
[0025] Exemplarily, in the primary noise reduction module 20 of the instruction voice signal, the user interaction instruction voice signal is subjected to primary noise reduction to obtain a primary noise-reduced user interaction instruction voice signal. It should be understood that the unprocessed original voice signal may contain a large amount of redundant information. Directly performing complex processing on it (such as deep learning model analysis) not only increases the computational cost but may also cause delays. Primary noise reduction can filter out unnecessary noise at an early stage, reduce the workload of subsequent high-level processing modules, and thus speed up the response speed of the entire system. Primary noise reduction provides a basis for subsequent adaptive noise reduction. It can preliminarily purify the voice signal without relying on specific environmental characteristics, enabling the system to work in a more diverse environment rather than just being optimized for certain fixed scenarios.
[0026] In one embodiment, the primary noise reduction module for the instruction voice signal is configured to: perform primary noise reduction on the user interaction instruction voice signal using a frequency domain filter to remove the high-frequency noise of the user interaction instruction voice signal to obtain the user interaction instruction voice signal after primary noise reduction. Specifically, by converting the voice signal in the time domain to the frequency domain, the frequency domain filter can more easily identify and remove the noise concentrated in the high-frequency region. For example, the fast Fourier transform (FFT) is used to map the voice signal from the time axis to the frequency axis, and then a band-pass or low-pass filter is applied to eliminate the high-frequency part beyond the human speech frequency range. In other examples of this application, in addition to algorithmic improvements, a dedicated hardware accelerator (such as a DSP - digital signal processor) can be used to perform efficient real-time noise reduction operations. This can not only improve the processing speed but also ensure the smooth operation of the noise reduction program on resource-constrained devices. Although primary noise reduction mainly focuses on fixed noise reduction operations, a certain degree of adaptive mechanism can also be introduced. For example, the filter parameters are automatically adjusted based on the environmental sound samples in the initial few seconds to better adapt to different usage environments.
[0027] Exemplarily, in the environmental noise signal acquisition module 30, the microphone array is used to acquire the environmental noise signal. It should be understood that in the human-computer interaction scenario of a smart refrigerator, the home environment is often filled with various background noises, such as the sound of a range hood, a TV, conversations, etc. These noises can interfere with the user's voice commands, resulting in an increased error rate of the voice recognition system. Traditional fixed-parameter noise reduction methods are difficult to cope with this diverse and dynamically changing noise environment. By acquiring the environmental noise signal, the system can understand the current noise characteristics of the environment in real time and adjust the noise reduction strategy accordingly, ensuring that even in a noisy environment, the user's commands can be effectively captured, providing a smooth and responsive interaction experience, and thus improving user satisfaction.
[0028] In one embodiment, an array composed of multiple microphones is configured inside the smart refrigerator. These microphones are distributed at different positions of the refrigerator to cover as large a pickup range as possible and ensure that environmental noise can be captured from all angles. Such a design not only helps to obtain comprehensive noise information but also can enhance the sound collection ability in a specific direction through beamforming technology, further improving the signal-to-noise ratio. When a voice interaction action is initiated (for example, the user presses the voice button or says the wake-up word), the system will simultaneously activate all microphones to start collecting sound data. At this time, in addition to paying attention to the user's voice commands, the system will also specifically note the noise situation in the surrounding environment. The collected audio stream will be processed in two parts: one part is used for immediate voice command parsing; the other part is specifically used to analyze the environmental noise characteristics.
[0029] Exemplarily, in the instruction voice signal adaptive noise reduction module 40, based on the environmental noise signal, the primary noise-reduced user interaction instruction voice signal is adaptively noise-reduced to obtain a secondary noise-reduced user interaction instruction voice signal. It should be understood that in a home environment, the type and intensity of background noise vary with time and location, including various types of noise such as range hood noise, TV sound, and conversation sound. Traditional fixed-parameter noise reduction methods are difficult to cope with this diverse and dynamically changing noise environment, which may lead to over-reduction (voice distortion) or under-reduction (noise residue), affecting the effectiveness of the human-machine interaction of the smart refrigerator and reducing the user experience. Therefore, it is crucial to introduce adaptive noise reduction technology. It can dynamically adjust the noise reduction strategy according to the real-time collected environmental noise information to better adapt to the noise characteristics in different scenarios.
[0030] Specifically, in the human-machine interaction system of the above smart refrigerator, a two-stage noise reduction strategy combining a frequency domain filter and adaptive noise reduction technology is applied to perform noise reduction processing on the user interaction instruction voice signal to achieve a better noise reduction and enhancement effect of the user interaction instruction voice signal, which is beneficial to extracting useful interaction instructions and content semantics from the original audio data for corresponding control of the refrigerator. That is to say, the above smart refrigerator human-machine interaction system aims to solve the problem of speech recognition accuracy in complex environments by combining advanced signal processing technologies and artificial intelligence algorithms. Specifically, the system first performs preliminary noise reduction processing on the data collected by the microphone array, and then adaptively further purifies the user interaction instruction voice signal according to the simultaneously collected environmental noise information. This two-stage noise reduction process not only takes into account the basic physical-level noise reduction but also incorporates the understanding of context information to achieve a better speech clarity restoration effect. Finally, speech recognition technology and command parsing methods are used to output operation instructions for controlling the refrigerator. In this way, it is possible to dynamically adjust the noise reduction strategy by real-time analyzing the environmental noise characteristics, thereby more effectively performing noise reduction processing on the user interaction instruction voice signal, making it more efficient and accurate to extract useful information from the original audio data, and improving the accuracy of the human-machine interaction of the smart refrigerator and user satisfaction in a more intelligent way.
[0031] In particular, the process of adaptively denoising the primary denoised user interaction command voice signal based on the environmental noise signal is crucial. Through adaptive denoising technology, the system can more accurately estimate and remove the noise components in the background, thus retaining a clearer user command voice signal. This directly improves the accuracy of the subsequent speech recognition stage and reduces the misrecognition or non-recognition caused by noise interference. Moreover, the sound environment in home spaces such as kitchens is very complex and constantly changing, including various types of sounds such as the working sounds of household appliances and conversations. Traditional fixed-parameter denoising methods are difficult to cope with this diversity. The adaptive denoising method can dynamically adjust the denoising strategy according to the real-time collected environmental noise information, better adapting to the noise characteristics in different scenarios. Through this two-stage denoising process, it can ensure that even in a noisy environment, the user's commands can be effectively captured, providing a smooth and responsive interaction experience, thereby enhancing user satisfaction.
[0032] Specifically, in the process of adaptively denoising the primary denoised user interaction command voice signal based on the environmental noise signal, the technical concept of this application is to convert both the environmental noise signal and the primary denoised user interaction command voice signal into two-dimensional time-frequency diagrams, and then introduce image processing and analysis algorithms based on artificial intelligence and deep learning at the backend to analyze the environmental noise two-dimensional time-frequency diagram and the voice signal two-dimensional time-frequency diagram, so as to capture the time-frequency characteristics of the environmental noise and the voice signal, and perform dynamic memory feature-interaction and optimized expression on the time-frequency semantics of the interaction command voice signal based on the time-frequency semantics of the environmental noise, in order to use the time-frequency interaction semantics of the two to represent the time-frequency characteristics of the adaptively denoised voice signal, and then obtain the secondary denoised user interaction command voice signal through signal restoration. In this way, the environmental noise signal can be used to adaptively perform dynamic denoising processing on the user interaction command voice signal, so as to better adapt to the denoising requirements in different scenarios. Through this method, a more intelligent refrigerator human-computer interaction process can be realized to ensure that even in a noisy environment, the user's commands can be effectively captured, which in turn helps to improve the responsive interaction experience and enhance user satisfaction.
[0033] In one embodiment, as Figure 3As shown, the adaptive noise reduction module 40 for the instruction voice signal includes: an environmental noise signal conversion unit 41 for converting the environmental noise signal into an environmental noise two-dimensional time-frequency map; an instruction voice signal conversion unit 42 for converting the user interaction instruction voice signal after primary noise reduction into a voice signal two-dimensional time-frequency map; a sound signal time-frequency feature extraction unit 43 for respectively performing sound signal time-frequency feature extraction on the environmental noise two-dimensional time-frequency map and the voice signal two-dimensional time-frequency map to obtain an environmental noise time-frequency feature vector and a voice signal time-frequency feature vector; an environmental noise-voice signal feature interaction processing unit 44 for performing feature interaction processing based on dynamic memory on the environmental noise time-frequency feature vector and the voice signal time-frequency feature vector to obtain an adaptive noise reduction after voice signal time-frequency feature; a user interaction instruction generation unit 45 after noise reduction for generating the user interaction instruction voice signal after secondary noise reduction based on the adaptive noise reduction after voice signal time-frequency feature.
[0034] Exemplarily, in the environmental noise signal conversion unit 41, the environmental noise signal is converted into an environmental noise two-dimensional time-frequency map. It should be understood that environmental noise usually contains multiple frequency components, and these components change over time. By converting the environmental noise signal into a two-dimensional time-frequency map, the time variation (i.e., development over time) and frequency distribution of the noise can be captured simultaneously, thus comprehensively describing the dynamic characteristics of the noise. This is very important for distinguishing transient noise (such as a sudden loud noise) and persistent noise (such as the sound of a fan running). The two-dimensional time-frequency map provides an intuitive way to observe and analyze the patterns in the noise signal. Using a deep learning model to process these graphically represented data can extract meaningful feature vectors from complex background noise, thereby helping to design more effective noise reduction strategies.
[0035] In one embodiment, converting the environmental noise signal into a two-dimensional time-frequency map of environmental noise includes: First, the original audio data collected by the microphone array needs to be preliminarily processed, such as removing DC bias, normalization, etc., to ensure the consistency and stability of the input data. Then, apply the Short-Time Fourier Transform (STFT) or the Continuous Wavelet Transform (CWT). Both of these methods can decompose a long-time series signal into a series of spectral representations within shorter time periods. STFT is suitable for stationary noise sources, while CWT is more suitable for non-stationary signals because it can analyze the signal at different scales. Specifically, for STFT: Select an appropriate window length and overlap rate, divide the signal into multiple short segments, and then calculate the Fourier transform for each segment to obtain its corresponding spectrum. Finally, combine the spectra of all segments to form a two-dimensional matrix, where the rows represent positions on the time axis and the columns correspond to the respective frequency components. For CWT: For each point, convolve the signal with wavelet basis functions of different scales to generate a series of coefficient maps at different scales. These coefficient maps also form a two-dimensional structure, reflecting the performance of the signal at different times and frequencies. Finally, having obtained the results of the above transformations, they can be visualized as a color image, where the shade of the color represents the amplitude size. In this way, the so-called "heat map" or "pseudo-color map", that is, the two-dimensional time-frequency map of environmental noise, is formed.
[0036] Exemplarily, in the instruction voice signal conversion unit 42, the primary noise-reduced user interaction instruction voice signal is converted into a two-dimensional time-frequency map of the voice signal. It should be understood that the two-dimensional time-frequency map can highlight key features in the voice signal, such as the transitions between phonemes, short sound pulses, etc., which are crucial for understanding the voice content. By observing the frequency distribution in different time periods of the time-frequency map, it is easier to distinguish which parts are voice and which are unwanted noises, thus providing a basis for adaptive noise reduction. Modern deep learning techniques are good at extracting complex patterns from images. Therefore, converting the voice signal into a data format similar to an image (i.e., a two-dimensional time-frequency map) can make full use of existing image processing techniques and network architectures to achieve efficient voice processing tasks.
[0037] Exemplarily, in the sound signal time-frequency feature extraction unit 43, the time-frequency features of the environmental noise two-dimensional time-frequency map and the speech signal two-dimensional time-frequency map are respectively extracted to obtain the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector. It should be understood that for sound signals, the time-frequency map provides information about the variation of the sound signal frequency over time. At the same time, such signal time-frequency features can provide information about the differences between noise and interactive speech. Based on this, in order to learn and extract key features from the environmental noise two-dimensional time-frequency map and the speech signal two-dimensional time-frequency map that are helpful for distinguishing noise from useful speech signals, it is necessary to extract and capture the two-dimensional time-frequency map features of both. Based on this, in the technical solution of this application, the environmental noise two-dimensional time-frequency map and the speech signal two-dimensional time-frequency map are respectively input into the sound signal time-frequency feature extractor based on the GoogLeNet network to obtain the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector. Through the sound signal time-frequency feature extractor based on the GoogLeNet network, the time-frequency feature semantics and signal complex patterns in the environmental noise two-dimensional time-frequency map and the speech signal two-dimensional time-frequency map can be respectively captured. In particular, it is worth mentioning that the Inception module design in the GoogLeNet architecture allows different-scale receptive fields to be processed simultaneously in the same layer, which means that it can capture the local details and global structure information in the environmental noise two-dimensional time-frequency map and the speech signal two-dimensional time-frequency map. This is particularly useful for understanding subtle changes in sound signals, such as brief noise pulses or transitions between speech phonemes.
[0038] In one embodiment, the sound signal time-frequency feature extraction unit is configured to: respectively input the environmental noise two-dimensional time-frequency map and the speech signal two-dimensional time-frequency map into the sound signal time-frequency feature extractor based on the GoogLeNet network to obtain the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector. Here, one of the most significant features of GoogLeNet is the introduction of the Inception module. This module processes different-scale receptive fields simultaneously in one layer. By using convolutional kernels of sizes 1x1, 3x3, and 5x5 and 3x3 max-pooling operations in parallel, and then stitching their results together as the output of this layer. The advantage of doing this is that the network can automatically learn the optimal spatial aggregation method, and the network width is increased without significantly increasing the computational cost. In addition, in order to alleviate the problem of gradient disappearance that may occur during the training of deep networks, GoogLeNet adds two auxiliary classifier branches in the middle of the network. These branches are connected to the intermediate layers and participate in the loss function calculation together with the main classifier, providing additional supervision signals, which helps to accelerate the convergence process. Only the results of the main classifier are considered during the final prediction.
[0039] Unlike traditional CNNs that use a large number of fully connected layers to extract features, GoogLeNet adopts global average pooling layers to replace the functions of some fully connected layers. This not only reduces the number of parameters but also decreases the risk of overfitting, while improving the generalization ability of the model. GoogLeNet contains a total of 22 layers (more if the layers inside the Inception modules are counted), but its number of parameters is much less than that of other top models at that time. This is achieved through a carefully designed network topology, including techniques such as the aforementioned Inception modules and global average pooling.
[0040] Due to its efficiency and accuracy, GoogLeNet is widely used in various computer vision tasks, especially performing excellently in fields such as image classification and object detection. For example, in the human-computer interaction system of an intelligent refrigerator, GoogLeNet can be used to extract features from the two-dimensional time-frequency diagrams of environmental noise and speech signals. Specifically, when useful features need to be extracted from audio data, the two-dimensional time-frequency diagrams of environmental noise and speech signals can be respectively input into the time-frequency feature extractor of the sound signal based on the GoogLeNet network. GoogLeNet can capture the local details and global structure information in these images, thus helping to distinguish the key features of noise and useful speech signals.
[0041] Exemplarily, in the environmental noise-speech signal feature interaction processing unit 44, feature interaction processing based on dynamic memory is performed on the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector to obtain the time-frequency features of the speech signal after adaptive noise reduction. It should be understood that since the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector respectively contain two-dimensional time-frequency feature information about environmental noise and user interaction instruction speech signals. However, in actual interaction scenarios, environmental noise is complex and variable, and has different categories. In such a complex and variable environment, the relationship between noise and speech is non-linear and changes with time and frequency. In order to effectively interact and fuse the environmental noise time-frequency semantics and the interaction instruction speech signal time-frequency semantics so as to extract clear speech signals from noise, in the technical solution of this application, further feature interaction processing based on dynamic memory is performed on the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector to obtain the time-frequency features of the speech signal after adaptive noise reduction.
[0042] In one embodiment, as Figure 4As shown, the environmental noise-speech signal feature interaction processing unit 44 includes: a collaborative feature extraction subunit 441, configured to perform collaborative feature extraction on the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector to obtain a latent collaborative coding vector between the speech signal-environmental noise time-frequency features; a feature attention modulation optimization subunit 442, configured to perform feature attention modulation optimization on the speech signal time-frequency feature vector and the environmental noise time-frequency feature vector respectively based on the latent collaborative coding vector between the speech signal-environmental noise time-frequency features to obtain an optimized speech signal time-frequency feature vector and an optimized environmental noise time-frequency feature vector; a feature linear interaction processing subunit 443, configured to input the optimized speech signal time-frequency feature vector and the optimized environmental noise time-frequency feature vector into a feature linear interaction network to obtain an adaptively noise-reduced speech signal time-frequency feature vector as the time-frequency feature of the adaptively noise-reduced speech signal.
[0043] Exemplarily, in the collaborative feature extraction subunit 441, collaborative feature extraction is performed on the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector to obtain a latent collaborative coding vector between the speech signal-environmental noise time-frequency features. Specifically, this process can be expressed by the formula: ; where and are the speech signal time-frequency feature vector and the environmental noise time-frequency feature vector respectively, and are element-wise addition, element-wise subtraction, and element-wise multiplication respectively, represents concatenation processing, represents a point convolution layer, represents an activation function, is the latent collaborative coding vector between the speech signal-environmental noise time-frequency features.
[0044] That is, by performing collaborative feature extraction on the time-frequency feature vectors of environmental noise and speech signals, more complex interaction patterns between the two can be captured. This interaction pattern contains key clues on how to distinguish useful speech information from interfering noise, enabling subsequent noise reduction algorithms to more accurately remove background noise while retaining and enhancing the user's speech commands.
[0045] Exemplarily, in the feature attention modulation optimization subunit 442, based on the latent collaborative coding vector between the speech signal-environmental noise time-frequency features, feature attention modulation optimization is performed on the speech signal time-frequency feature vector and the environmental noise time-frequency feature vector respectively to obtain an optimized speech signal time-frequency feature vector and an optimized environmental noise time-frequency feature vector. Specifically, this process can be expressed by the formula: ; where is the implicit co-coding vector between the time-frequency features of the speech signal and the environmental noise, represents a 1×1 convolution operation, is the dynamic key vector of the speech signal - environmental noise time-frequency features, and respectively represent the speech signal time-frequency query weight matrix and the speech signal time-frequency query bias vector, and respectively represent the speech signal time-frequency value weight matrix and the speech signal time-frequency value bias vector, is matrix multiplication, and are respectively the speech signal time-frequency query feature vector and the speech signal time-frequency value feature vector, is vector transpose, is the length of, represents the normalized exponential function, is the optimized speech signal time-frequency feature vector, and respectively represent the environmental noise time-frequency query weight matrix and the environmental noise time-frequency query bias vector, and respectively represent the environmental noise time-frequency value weight matrix and the environmental noise time-frequency value bias vector, and are respectively the environmental noise time-frequency query feature vector and the environmental noise time-frequency value feature vector, is the length of, is the optimized environmental noise time-frequency feature vector.
[0046] In one embodiment, the feature attention modulation optimization subunit is configured to: write the implicit co-coding vector between the time-frequency features of the speech signal and the environmental noise into a dynamic memory unit to obtain a dynamic key vector of the speech signal - environmental noise time-frequency features; extract the dynamic key vector of the speech signal - environmental noise time-frequency features from the dynamic memory unit, and input the speech signal time-frequency feature vector and the dynamic key vector of the speech signal - environmental noise time-frequency features into a feature attention modulation module based on a first transformer structure to obtain the optimized speech signal time-frequency feature vector; extract the dynamic key vector of the speech signal - environmental noise time-frequency features from the dynamic memory unit, and input the environmental noise time-frequency feature vector and the dynamic key vector of the speech signal - environmental noise time-frequency features into a feature attention modulation module based on a second transformer structure to obtain the optimized environmental noise time-frequency feature vector.
[0047] Specifically, the process of feature interaction processing based on dynamic memory can effectively capture the implicit collaborative relationship between the time-frequency feature vector of the environmental noise and the time-frequency feature vector of the speech signal, and use the implicit collaborative relationship between the time-frequency semantics of the environmental noise and the time-frequency semantics of the speech signal as a dynamic key vector to construct the same attention space, so as to perform attention modulation optimization on the time-frequency feature vector of the environmental noise and the time-frequency feature vector of the speech signal in real time, thereby emphasizing the important information therein and suppressing noise and redundant information. The feature attention modulation module based on the Transformer structure adopts the self-attention mechanism, allowing the model to consider the influence of all other elements when processing each input element. This global perspective is particularly important for understanding long-range dependencies. Through feature modulation based on query attention, the expression of key information in the feature vector can be enhanced while reducing the influence of noise and irrelevant information. For example, at a specific time point, certain frequency components may be noise, while at another time point, they may be useful speech signals. The self-attention mechanism can help the model correctly allocate attention at different time points, so as to better separate noise and speech. This helps to highlight the useful parts of the speech signal and improve speech clarity.
[0048] Exemplarily, in the feature linear interaction processing sub-unit 443, the optimized time-frequency feature vector of the speech signal and the optimized time-frequency feature vector of the environmental noise are input into the feature linear interaction network to obtain the time-frequency feature vector of the speech signal after adaptive noise reduction as the time-frequency feature of the speech signal after adaptive noise reduction. Specifically, this process can be expressed by the formula: ; where is the optimized time-frequency feature vector of the speech signal, is the optimized time-frequency feature vector of the environmental noise, is the weight hyperparameter, is the time-frequency feature vector of the speech signal after adaptive noise reduction.
[0049] That is, feature alignment and linear fusion are performed on the optimized time-frequency feature vector of the environmental noise and the optimized time-frequency feature vector of the speech signal. During the fusion process, the weight is dynamically optimized by introducing a trainable weight hyperparameter, so as to use the time-frequency features of the environmental noise to optimize the expression of the time-frequency features of the speech signal, generating the time-frequency feature vector of the speech signal after adaptive noise reduction modulation optimization, providing high-quality input for subsequent speech recognition.
[0050] Exemplarily, in the post-noise reduction user interaction instruction generation unit 45, the post-secondary-noise reduction user interaction instruction voice signal is generated based on the time-frequency features of the adaptively post-noise reduction voice signal. It should be understood that the data processed by the feature extractor usually exists in the form of feature vectors. Although these feature vectors contain important acoustic information, they do not directly correspond to audible sound waveforms. In order to be understood and processed by the subsequent speech recognition model, these feature vectors must be converted back to the audio signal in the time domain. Therefore, in one embodiment, the post-noise reduction user interaction instruction generation unit is configured to: perform signal restoration on the time-frequency feature vector of the adaptively post-noise reduction voice signal to obtain the post-secondary-noise reduction user interaction instruction voice signal. By performing signal restoration on the time-frequency feature vector of the adaptively post-noise reduction voice signal, a clearer and lower-noise post-noise reduction user interaction instruction voice signal can be generated. This process helps to remove or weaken the influence of background noise, while retaining and enhancing the user's speech content components, thereby improving the quality of the final voice signal. In this way, the environmental noise signal can be used to adaptively perform dynamic noise reduction processing on the user interaction instruction voice signal to better meet the noise reduction requirements in different scenarios. Through this method, a more intelligent refrigerator human-computer interaction process can be realized to ensure that the user's commands can be effectively captured even in a noisy environment, which in turn helps to improve the responsive interaction experience and enhance the user's satisfaction.
[0051] In a specific embodiment, a decoder is used to perform signal restoration on the time-frequency feature vector of the adaptively post-noise reduction voice signal to obtain the post-secondary-noise reduction user interaction instruction voice signal. Specifically, first, an input layer is used to receive the time-frequency feature vector of the adaptively post-noise reduction voice signal as input. Then, a transposed convolutional layer is used to gradually restore the spatial resolution through transposed convolution operations and introduce new detailed information. These layers can help expand the spatial dimension of the input feature map while retaining important frequency information. Next, batch normalization and activation functions are applied after each transposed convolutional layer to accelerate the training process and stabilize the output distribution, and introduce non-linearity so that the model can learn more complex mapping relationships. Then, in order to better capture the long-term dependencies in the time series, several GRU layers are added on top of the transposed convolutional layer. GRU is a simplified version of LSTM (Long Short-Term Memory network), which can effectively process sequence data and has a lower computational cost. Finally, a fully connected layer is used to map the high-dimensional features back to the dimension of the original audio waveform to obtain the post-secondary-noise reduction user interaction instruction voice signal.
[0052] In the technical solution of this application, the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector respectively represent the time-frequency features of the environmental noise signal and the time-frequency features of the user interaction instruction speech signal after primary noise reduction. During the process of performing feature-interaction attention between them based on the dynamic memory mechanism, the randomness of the environmental noise signal will cause the fine-grained interaction structure in the dynamic interaction process to be unbalanced, which will affect the signal quality of the decoded signal restoration of the time-frequency feature vector of the speech signal after adaptive noise reduction.
[0053] Based on this, in a preferred embodiment of this application, signal restoration is performed on the time-frequency feature vector of the speech signal after adaptive noise reduction to obtain the user interaction instruction speech signal after secondary noise reduction, including: performing feature manifold optimization on the time-frequency feature vector of the speech signal after adaptive noise reduction to obtain an optimized time-frequency feature vector of the speech signal after adaptive noise reduction; performing signal restoration on the optimized time-frequency feature vector of the speech signal after adaptive noise reduction to obtain the user interaction instruction speech signal after secondary noise reduction.
[0054] Specifically, the process of the feature manifold optimization includes: calculating the L1 norm of the time-frequency feature vector of the speech signal after adaptive noise reduction to obtain the first speech signal time-frequency feature topology parameter, and calculating its L2 norm to obtain the second speech signal time-frequency feature topology parameter; ; where represents each eigenvalue of the time-frequency feature vector of the speech signal after adaptive noise reduction, represents the norm of the time-frequency feature vector of the speech signal after adaptive noise reduction, represents the first speech signal time-frequency feature topology parameter, represents the second speech signal time-frequency feature topology parameter.
[0055] Multiply the eigenvalues at each position of the time-frequency feature vector of the speech signal after adaptive noise reduction by the first speech signal time-frequency feature topology parameter and the second speech signal time-frequency feature topology parameter respectively to generate the corresponding first-order topology reference quantity of the speech signal time-frequency feature and the second-order topology reference quantity of the speech signal time-frequency feature, expressed as: ; where represents the eigenvalue corresponding to the first-order topology reference quantity of the speech signal time-frequency feature, represents the eigenvalue corresponding to the second-order topology reference quantity of the speech signal time-frequency feature.
[0056] Multiply the eigenvalues at each position of the time-frequency feature vector of the adaptively noise-reduced speech signal by the magnitude and square root of the time-frequency feature vector of the adaptively noise-reduced speech signal respectively to generate a first-order domain transformation parameter and a second-order domain transformation parameter of the speech signal time-frequency feature, expressed as: ; where represents the eigenvalue corresponding to the first-order domain transformation parameter of the speech signal time-frequency feature, represents the eigenvalue corresponding to the second-order domain transformation parameter of the speech signal time-frequency feature.
[0057] Divide the first-order topological reference quantity of the speech signal time-frequency feature by the offset between the first speech signal time-frequency feature topological parameter and the first-order domain transformation parameter of the speech signal time-frequency feature to obtain the first-order modulation coefficient of the speech signal time-frequency feature, expressed as: ; where represents the eigenvalue corresponding to the first-order modulation coefficient of the speech signal time-frequency feature.
[0058] Divide the second-order topological reference quantity of the speech signal time-frequency feature by the offset between the second speech signal time-frequency feature topological parameter and the second-order domain transformation parameter of the speech signal time-frequency feature to obtain the second-order modulation coefficient of the speech signal time-frequency feature, expressed as: ; where represents the eigenvalue corresponding to the second-order modulation coefficient of the speech signal time-frequency feature.
[0059] Calculate the weighted aggregation quantity between the first-order modulation coefficient and the second-order modulation coefficient of the speech signal time-frequency feature to obtain each eigenvalue of the optimized time-frequency feature vector of the adaptively noise-reduced speech signal, expressed as: ; where represents the vector composed of the first-order modulation coefficients of the speech signal time-frequency feature, represents the vector composed of the second-order modulation coefficients of the speech signal time-frequency feature, represents the weighted hyperparameter, represents addition by position, represents multiplication by position, represents the optimized time-frequency feature vector of the adaptively noise-reduced speech signal.
[0060] That is, for the geometric topological features of the attribute clusters of the time-frequency feature vectors of the speech signal after adaptive noise reduction in the multi-dimensional embedding domain, by taking the normalized spatial topological representation of the time-frequency feature vectors of the speech signal after adaptive noise reduction as the reference view, implementing scale-driven domain transformation on its respective attribute values, and completing the construction of the regional attention modulation mechanism. In this way, it is ensured that the time-frequency feature vectors of the speech signal after adaptive noise reduction maintain spatial mapping stability during the attribute field interaction process, thereby enhancing the stable convergence and extrapolation performance of its attribute clusters in fitting the generative form in the heterogeneous semantic topological environment, and ultimately improving the signal restoration quality of the user interaction command speech signal after secondary noise reduction.
[0061] Exemplarily, in the interactive command speech recognition module 50, the user interaction command speech signal after secondary noise reduction is sent to the speech recognition model to obtain the text description of the user interaction command. It should be understood that the speech signal after primary noise reduction and adaptive noise reduction has significantly reduced the interference of background noise, which makes the speech signal clearer, thereby improving the input quality of the speech recognition model. High-quality input helps reduce the possibility of misrecognition, ensuring that the system can accurately understand the user's intention. By recognizing the optimized speech signal, a smoother and faster response time can be provided, reducing delays or incorrect feedback caused by noise, and thus enhancing user satisfaction. Even in a noisy environment, users can obtain a stable and reliable interaction experience. Modern speech recognition technology is not limited to simple keyword matching, but can also parse complex natural language expressions. By providing clean audio input to the speech recognition model, the system can better capture the nuances in the user's command, such as multi-step instructions or multi-condition queries, etc., providing a more intelligent service for users.
[0062] In one embodiment, the speech recognition model is DeepSpeech. It should be understood that DeepSpeech is an end-to-end speech recognition system based on deep learning, which uses recurrent neural networks (RNNs) and convolutional neural networks (CNNs) to directly map the audio waveform to text. In a specific embodiment, select the appropriate installation method according to the target platform (such as installing the Python package via pip). Download and load the pre-trained DeepSpeech model file (in the `.pbmm` format). Use the provided Python API or other programming interfaces to pass the user interaction command speech signal after secondary noise reduction to the DeepSpeech model for recognition to obtain the text description of the user interaction command. Obtain the returned text description as the basis for command parsing.
[0063] Exemplarily, in the command parsing module 60, the user interaction instruction text description is parsed to obtain a command content parsing result. It should be understood that the text description generated by the speech recognition model only converts the user's speech into text form, but it does not directly tell the system which specific operations should be performed. Through command parsing, the true needs of the user can be deeply understood, including the information they want to query, the functions they set, or the actions they trigger, etc. Modern smart home devices can not only respond to simple keywords or phrases, but also understand and execute complex multi-step instructions or multi-condition queries. For example, "Please adjust the temperature to 4 degrees and tell me what food has the shortest remaining shelf life". Such composite instructions need to rely on advanced natural language processing techniques to be accurately parsed and decomposed into a series of executable operations. Through a strict command parsing process, the validity of the user input can be verified, preventing misoperations or illegal commands from being executed. In addition, it can also help filter out vague or incomplete requests, ensuring that only clear and reasonable commands are passed to the control module. Based on the user's historical behavior and personal preferences, command parsing can also provide customized suggestions and services. For example, recommend appropriate food storage methods according to the user's eating habits, or automatically adjust the working mode of the refrigerator to save energy.
[0064] In one embodiment, parsing the user interaction instruction text description to obtain a command content parsing result includes: First, use a word segmentation tool to split the continuous text into word or phrase units, which helps to understand the sentence structure at a finer granularity later. Apply techniques such as dependency syntax analysis, named entity recognition (NER), sentiment analysis, etc. to capture key information in the text, such as subject-verb-object relationships, specific nouns (such as food names), verb intentions, etc. Then, by training a machine learning model or a rule engine, identify the main intention of the user, such as querying, setting, or starting a certain function, etc. These models can be trained based on a large amount of labeled data to better adapt to different types of user inputs. Pre-define a set of possible command patterns and their corresponding execution logics to form a command template library. Each template contains specific slots for filling in the specific parameter values extracted from the user input. Match the text fragments after NLP processing with the command templates to find the template that best matches the current input, and fill in the relevant parameters in the corresponding positions. For example, for the command "Set the temperature to X degrees", X is the parameter that needs to be extracted.
[0065] In other embodiments, considering that the user may issue multiple related commands consecutively, the system needs to maintain a certain context memory to understand the connection between the previous and subsequent commands. This can be achieved by introducing a dialogue management mechanism to track the dialogue history and dynamically update the internal state. When encountering an input that is ambiguous or has multiple possible interpretations, the system should have the ability to ask for clarification, provide a list of options to the user, or request further explanation to ensure that the final parsed result is correct. Combining professional knowledge related to the refrigerator (such as food types, shelf life, etc.) can enhance the ability to understand specific terms. This can not only improve the parsing accuracy but also provide more professional and considerate services to users.
[0066] Exemplarily, in the refrigerator operation control feedback module 70, based on the command content parsing result, an operation instruction is called to control the refrigerator and feedback the operation result. It should be understood that the clear instruction obtained through command parsing can guide the system to accurately call the corresponding operation instruction to control the functions of the refrigerator. This not only improves the response accuracy of the system but also avoids operation errors caused by misunderstanding or incorrect parsing. When the user issues a voice instruction, they can immediately see or hear the specific response of the system to their request, enhancing the immediacy and intuitiveness of the interaction. This timely feedback makes the user feel that their needs are being taken seriously and that the system is working properly.
[0067] In one embodiment, the refrigerator operation control feedback module is configured to: based on the result of parsing the command content, call the operation instruction from the operation instruction library to control the refrigerator; and feedback the operation result through a voice feedback signal. Specifically, when constructing the operation instruction library, a set of standardized command formats are created for various functions of the refrigerator, including but not limited to power on / off, temperature setting, lighting control, food classification and shelf life management, etc. Each standard command corresponds to one or more API interfaces, which are responsible for actually executing specific tasks, ensuring that all possible operations have clear mapping relationships for subsequent calls. According to the parsing result, search for the operation instruction in the operation instruction library that best matches the current command description. If there are multiple possible choices, select the best match; if there is ambiguity, ask the user for further clarification. Once the specific operation instruction is determined, the corresponding API interface or function can be called to execute the command, such as adjusting the temperature setting, querying the inventory status, etc. During the execution of the operation, the system continuously monitors the status changes of the refrigerator to ensure the smooth completion of the operation. If problems occur (such as hardware failures, insufficient permissions, etc.), the abnormal situations will be captured in a timely manner and appropriate measures will be taken. Whether it is successful or failed, a concise and clear feedback message should be generated to inform the user. For successful operations, it can be simply confirmed; for failed situations, the reason should be explained and a solution should be provided. The feedback methods are diverse. It can be voice feedback: directly play voice messages through the built-in speaker, such as "The temperature has been adjusted to 4 degrees", "The food you mentioned cannot be found. Please check if the name is correct"; it can also be graphical interface feedback: update the display content on the refrigerator display screen to visually present the operation result, such as icon changes, text prompts, etc.; in addition, if the user is not near the refrigerator, notifications can also be sent through the mobile application or other remote terminals to keep the information synchronized. In this way, by controlling the refrigerator based on the result of parsing the command content and feedbacking the operation result, not only the user experience is improved, but also the adaptability and reliability of the system are enhanced.
[0068] In summary, the human-computer interaction system of the smart refrigerator according to the embodiments of the present application is elucidated. It uses the data collected by the microphone array for preliminary noise reduction processing, and then adaptively further purifies the voice signal of the user interaction instruction according to the simultaneously collected environmental noise information. This two-stage noise reduction process not only takes into account the basic physical-level noise reduction, but also incorporates the understanding of context information to achieve a better voice clarity restoration effect. Finally, voice recognition technology and command parsing methods are used to output operation instructions for controlling the refrigerator. In this way, it is possible to dynamically adjust the noise reduction strategy by real-time analyzing the characteristics of environmental noise, thereby more effectively performing noise reduction processing on the voice signal of the user interaction instruction, making it more efficient and accurate to extract useful information from the original audio data, and improving the accuracy and user satisfaction of the human-computer interaction of the smart refrigerator in a more intelligent way.
[0069] Figure 5 It is a schematic flowchart of the human-computer interaction method of the smart refrigerator according to the embodiments of the present application. As Figure 5 shown, the human-computer interaction method of the smart refrigerator includes: S1, after receiving the start voice interaction action, using the built-in microphone array of the refrigerator to collect the voice signal of the user interaction instruction; S2, performing primary noise reduction on the voice signal of the user interaction instruction to obtain the voice signal of the user interaction instruction after primary noise reduction; S3, using the microphone array to collect the environmental noise signal; S4, based on the environmental noise signal, performing adaptive noise reduction on the voice signal of the user interaction instruction after primary noise reduction to obtain the voice signal of the user interaction instruction after secondary noise reduction; S5, sending the voice signal of the user interaction instruction after secondary noise reduction to the voice recognition model to obtain the text description of the user interaction instruction; S6, performing command parsing on the text description of the user interaction instruction to obtain the command content parsing result; S7, based on the command content parsing result, calling the operation instruction to control the refrigerator and feedback the operation result.
[0070] Here, those skilled in the art can understand that the specific operations of each step in the above human-computer interaction method of the smart refrigerator have been introduced in detail in the description of the human-computer interaction system of the smart refrigerator above Figures 1 to 4 and therefore, the repeated description thereof will be omitted.
[0071] The embodiments of the present application also provide a computer program product, which includes computer program code. When the computer program code runs on a computer, it enables the computer to implement the methods in the above embodiments of the present application.
[0072] The embodiments of the present application also provide a computer-readable storage medium, which stores computer instructions. When the computer instructions run on a computer, it enables the computer to implement the methods in the above embodiments of the present application.
[0073] An embodiment of this application also provides a chip, including a circuit for executing the methods in the above various embodiments of this application.
[0074] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0075] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B can represent A or B; herein, "and / or" is an association relationship describing associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single (item) or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0076] In the embodiments of this application, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no restrictive effect on the position, order, priority, quantity, content, etc. of the described objects. In the embodiments of this application, the use of ordinal numbers and other prefix words for distinguishing described objects does not constitute a limitation on the described objects. The statement of the described objects refers to the description in the context of the claims or embodiments, and should not constitute an unnecessary limitation due to the use of such prefix words.
[0077] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0078] In each embodiment of this application, if there is no special explanation and logical conflict, the terms and / or descriptions between the embodiments are consistent and can be referred to each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0079] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0080] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.
[0081] As mentioned above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A human-computer interaction system for a smart refrigerator, characterized in that: include: A user interaction command voice collection module is used to collect user interaction command voice signals using a built-in microphone array of the refrigerator after receiving a voice interaction start action; A command voice signal primary noise reduction module, used for performing primary noise reduction on the user interaction command voice signal to obtain a user interaction command voice signal after primary noise reduction; An environmental noise signal acquisition module, used to acquire environmental noise signals using the microphone array; A command voice signal adaptive noise reduction module, used for adaptively reducing the noise of the user interaction command voice signal after primary noise reduction based on the environmental noise signal to obtain a user interaction command voice signal after secondary noise reduction; An interactive command speech recognition module, used for sending the user interactive command speech signal after the secondary noise reduction to a speech recognition model to obtain a text description of the user interactive command; A command parsing module, used for performing command parsing on the user interaction instruction text description to obtain a command content parsing result; A refrigerator operation control feedback module, for invoking an operation instruction to control the refrigerator and feeding back the operation result based on the command content analysis result; Wherein, the command voice signal adaptive noise reduction module includes: An environmental noise signal conversion unit, used to convert the environmental noise signal into a two-dimensional time-frequency diagram of environmental noise; A command voice signal conversion unit, used to convert the user interaction command voice signal after primary noise reduction into a two-dimensional time-frequency graph of the voice signal; A sound signal time-frequency feature extraction unit is used to extract the sound signal time-frequency features from the two-dimensional time-frequency graph of the environmental noise and the two-dimensional time-frequency graph of the speech signal to obtain a time-frequency feature vector of the environmental noise and a time-frequency feature vector of the speech signal; An environmental noise-speech signal feature interaction processing unit, used for performing feature interaction processing based on dynamic memory on the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector to obtain the time-frequency feature of the speech signal after adaptive noise reduction; A post-noise reduction user interaction instruction generating unit, configured to generate the secondary post-noise reduction user interaction instruction voice signal based on the time-frequency characteristics of the adaptive post-noise reduction voice signal; Wherein, the environmental noise-speech signal feature interaction processing unit includes: A collaborative feature extraction subunit, used for performing collaborative feature extraction on the environmental noise time-frequency feature vector and the speech signal time-frequency feature vector to obtain an implicit collaborative coding vector between the speech signal and environmental noise time-frequency features; A feature attention modulation optimization subunit is used to perform feature attention modulation optimization on the speech signal time-frequency feature vector and the ambient noise time-frequency feature vector based on the implicit collaborative coding vector between the speech signal and the ambient noise time-frequency feature to obtain an optimized speech signal time-frequency feature vector and an optimized ambient noise time-frequency feature vector; A feature linear interaction processing subunit, used for inputting the optimized speech signal time-frequency feature vector and the optimized ambient noise time-frequency feature vector into a feature linear interaction network to obtain the adaptive noise reduction speech signal time-frequency feature vector as the adaptive noise reduction speech signal time-frequency feature; Wherein, the feature attention modulation optimization subunit is used to: Writing the implicit collaborative coding vector between the speech signal and the environmental noise time-frequency features into the dynamic memory unit to obtain a speech signal and the environmental noise time-frequency feature dynamic key vector; Extract the speech signal-ambient noise time-frequency feature dynamic key vector from the dynamic memory unit, and input the speech signal time-frequency feature vector and the speech signal-ambient noise time-frequency feature dynamic key vector into a feature attention modulation module based on the first converter structure to obtain the optimized speech signal time-frequency feature vector; Extract the speech signal-ambient noise time-frequency feature dynamic key vector from the dynamic memory unit, and input the ambient noise time-frequency feature vector and the speech signal-ambient noise time-frequency feature dynamic key vector into a feature attention modulation module based on a second converter structure to obtain the optimized ambient noise time-frequency feature vector; The process of extracting collaborative features from the ambient noise time-frequency feature vector and the speech signal time-frequency feature vector to obtain an implicit collaborative coding vector between the speech signal and ambient noise time-frequency features can be expressed as follows: ;in, and are respectively the time-frequency feature vector of the speech signal and the time-frequency feature vector of the ambient noise, and They are positional addition, positional subtraction, and positional dot multiplication. Indicates cascade processing, represents the point convolution layer, express Activation function, It is the implicit collaborative coding vector between the time-frequency features of speech signal and environmental noise.
2. The human-computer interaction system of the smart refrigerator according to claim 1, characterized in that: The action of starting voice interaction is that the user presses a voice button or speaks a wake-up word.
3. The human-computer interaction system of the smart refrigerator according to claim 2, characterized in that: The command voice signal primary noise reduction module is used to: use a frequency domain filter to perform primary noise reduction on the user interaction command voice signal to remove high-frequency noise of the user interaction command voice signal to obtain the user interaction command voice signal after primary noise reduction.
4. The human-computer interaction system of the smart refrigerator according to claim 3, characterized in that: The sound signal time-frequency feature extraction unit is used to: input the two-dimensional time-frequency diagram of the environmental noise and the two-dimensional time-frequency diagram of the speech signal into a sound signal time-frequency feature extractor based on the GoogLeNet network to obtain the time-frequency feature vector of the environmental noise and the time-frequency feature vector of the speech signal.
5. The human-computer interaction system of the smart refrigerator according to claim 4, characterized in that: The post-noise reduction user interaction instruction generating unit is used to perform signal restoration on the time-frequency feature vector of the adaptive post-noise reduction speech signal to obtain the secondary post-noise reduction user interaction instruction speech signal.
6. The human-computer interaction system of the smart refrigerator according to claim 5, characterized in that: The refrigerator operation control feedback module is used to: Based on the command content parsing result, calling the operation instruction from the operation instruction library to control the refrigerator; The operation result is fed back via a voice feedback signal.
7. A human-computer interaction method for a smart refrigerator, using the human-computer interaction system for a smart refrigerator according to claim 1, characterized in that: include: After receiving the voice interaction start action, the microphone array built into the refrigerator is used to collect the user interaction command voice signal; Performing primary noise reduction on the user interaction instruction voice signal to obtain a user interaction instruction voice signal after primary noise reduction; Collecting environmental noise signals using the microphone array; Based on the environmental noise signal, adaptively denoising the primary noise-reduced user interaction instruction voice signal to obtain a secondary noise-reduced user interaction instruction voice signal; Sending the secondary noise-reduced user interaction command voice signal to a speech recognition model to obtain a user interaction command text description; Performing command parsing on the user interaction instruction text description to obtain a command content parsing result; Based on the analysis result of the command content, the operation instruction is called to control the refrigerator and the operation result is fed back.
Citation Information
Patent Citations
Voice recognition circuit, voice interaction device and household appliance
CN111091818A
Speech recognition method and device, equipment and medium
CN111640428A