Multi-information fusion voice interaction method and system for range hood

Through the multi-information interaction method of deep neural network and environmental sensor data fusion, the problem of low voice recognition accuracy of range hoods in noisy kitchen environments is solved, higher recognition accuracy and intelligent operation are achieved, and the user experience is improved.

CN120544569APending Publication Date: 2025-08-26XI AN JIAOTONG UNIV +2
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510754014.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-05-30
Filing Date
2025-06-06
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing voice interaction system of range hoods has low recognition accuracy in noisy kitchen environments and is difficult to effectively utilize multiple data resources, resulting in a high misjudgment rate.

Method used

A multi-information fusion voice interaction method is adopted. The mixed audio signal is processed through deep neural network noise reduction, and weighted fusion is performed with environmental sensor data to generate semantic features. The current status information is fed back through voice broadcast to dynamically adjust the operation of the equipment.

Benefits of technology

It significantly improves the accuracy of voice recognition and the intelligence level of device operation, reduces misjudgments, enhances user experience, reduces equipment wear and tear, and saves costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544569A_ABST
    Figure CN120544569A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of voice interaction, and discloses a multi-information fusion voice interaction method and system for a range hood, and the method comprises the following steps: carrying out the noise reduction of a mixed audio signal, and generating a user voice signal after noise reduction, the mixed audio signal comprising user voice and environment noise; based on a preset weighted fusion algorithm, performing weighted fusion on the de-noised user voice signal and environment sensor data to generate semantic features; converting the semantic features into a text instruction, and performing semantic analysis on the text instruction to generate a control instruction; a control signal is sent to the extractor hood according to a control instruction, the extractor hood is dynamically adjusted, the current state information of the extractor hood is obtained in real time after dynamic adjustment, the current state information is fed back to a user through voice broadcast, voice interaction with the user is carried out, and coping strategies of different times are set according to the situation that voice recognition fails. And interaction interruption caused by voice recognition failure is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of voice interaction and relates to a multi-information fusion voice interaction method and system for range hoods. Background Art

[0002] With the rapid development of artificial intelligence and voice interaction technologies, voice recognition technology has been widely applied in many fields and has gradually become a key component of smart home systems. In the home kitchen, range hoods are a common and critical appliance, and their intelligent development is of great significance for improving the user's cooking experience. By introducing voice interaction functions to range hoods, users can easily complete operations such as turning on the device, adjusting the fan speed, activating lighting, and setting timer tasks with simple voice commands, breaking away from the reliance on traditional physical buttons and greatly improving operational convenience and user experience.

[0003] However, existing technologies primarily rely on a single voice signal for interaction. This single data source ignores the rich multi-data resources in the kitchen environment, such as temperature, smoke concentration, and equipment operating status. The synergistic value inherent in this data has not been fully utilized. Although voice noise reduction technology can partially filter out background noise, the recognition accuracy of a single voice signal is still difficult to meet requirements in complex scenarios where range hoods operate at high frequencies and users operate multiple kitchen appliances simultaneously. For example, when a user issues a voice command to "increase the fan speed," it may be mistakenly interpreted as "turning off the lights" due to interference from ambient noise.

[0004] Therefore, improving speech recognition rate is a technical problem that needs to be urgently solved in existing technologies. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention aims to provide a multi-information fusion voice interaction method and system for range hoods.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions: The present invention provides a multi-information fusion voice interaction method for a range hood, comprising the following steps: performing noise reduction processing on a mixed audio signal to generate a noise-reduced user voice signal; based on a preset weighted fusion algorithm, performing weighted fusion on the noise-reduced user voice signal and environmental sensor data to generate semantic features; converting the semantic features into text instructions, and performing semantic analysis on the text instructions to generate control instructions.

[0007] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention discloses a multi-information fusion voice interaction method for range hoods, which performs noise reduction processing on mixed audio signals, effectively filters out background noise such as environmental mechanical sounds, electromagnetic interference, etc., generates a user voice signal with higher purity, significantly reduces the voice recognition error rate, and is suitable for voice interaction needs in noisy environments; based on a preset weighted fusion algorithm, the noise-reduced voice signal is dynamically weighted fused with environmental sensor data such as temperature, humidity, light, and equipment operating status, to generate semantic features containing environmental context information, breaking through the limitation of traditional voice recognition that only relies on voice signals, and combining with environmental parameters to perform more accurate semantic analysis, thereby improving the accuracy of command recognition and scene adaptability; the semantic features are converted into text commands, and the text commands are semantically analyzed to generate control commands, avoiding the risk of information loss in traditional multi-step conversion.

[0008] Furthermore, the dynamic adjustment also includes voice interaction feedback; the voice interaction feedback includes: obtaining the current status information of the range hood in real time, feeding back the current status information to the user through voice broadcast, and performing voice interaction with the user.

[0009] The present invention provides a multi-information fusion voice interaction method for range hoods. The method obtains the current status information of the range hood in real time, such as operating gear, wind speed, filter status, fault code, etc., and feeds it back to the user in the form of voice broadcast, so that the user can intuitively understand the operation status of the device without manually checking the device display screen, avoiding the user from repeating the operation due to no feedback after the operation, thereby improving the efficiency of human-computer interaction.

[0010] Furthermore, the preset weighted fusion algorithm is used to perform weighted fusion on the noise-reduced user voice signal and the environmental sensor data to generate semantic features, including:

[0011] in, is the enhanced semantic feature matrix after fusion, is the weight matrix, is element-wise multiplication, is the feature matrix of the speech signal, is the feature matrix of environmental sensor data, is a real matrix, is the number of time frames, The weighted fusion algorithm can dynamically adjust the weights according to the real-time changes in environmental sensor data, maintaining stable semantic parsing capabilities in different environments.

[0012] Furthermore, the weight matrix for:

[0013] in, is the normalized weight value, To process the feature matrix using the fully connected layer, is the feature matrix of environmental sensor data, is a real matrix, is the number of time frames, is the feature dimension.

[0014] The present invention provides a multi-information fusion voice interaction method for range hoods, which uses a weight matrix Dynamically adjust the feature matrix of the speech signal and the feature matrix of environmental sensor data The fusion ratio is adjusted to achieve the optimal feature extraction in different scenarios. For example, when the concentration of kitchen fume is high, the enhanced semantic feature matrix after fusion is It focuses more on environmental data, so as to accurately identify environmental information related to voice commands such as "oil fume concentration exceeds the standard".

[0015] Furthermore, the feature matrix of the speech signal for:

[0016] in, is a 1D convolution operation, is the time-frequency feature matrix of the speech signal after noise reduction, , is a real matrix, is the number of time frames, is the feature dimension.

[0017] Furthermore, the characteristic matrix of the environmental sensor data for:

[0018] in, is the fully connected layer, is the environmental sensor data matrix, , is the number of time frames, is the feature dimension.

[0019] Furthermore, the processing of the user voice signal after noise reduction also includes voice recognition, and when the voice recognition fails for the first time, a request to reissue the voice command is issued to the user.

[0020] The present invention provides a multi-information fusion voice interaction method for a range hood. When the first voice recognition fails, the method actively requests the user to reissue the voice command, which can effectively reduce the risk of misrecognition and thus ensure the accuracy of information transmission during the voice interaction process.

[0021] Furthermore, when the voice recognition fails for the second time, the operating parameters of the range hood are dynamically adjusted, and the operating parameters include reducing the fan speed or switching to a low noise mode.

[0022] This invention proposes a multi-information fusion voice interaction method for range hoods. This method dynamically adjusts range hood parameters when voice recognition fails twice, significantly improving the user experience. Reducing fan speed or switching to low-noise mode effectively reduces operating noise, allowing users to cook in a quiet atmosphere without being disturbed by noise. This also reduces fan load, minimizes equipment wear, extends the range hood's service life, and saves costs.

[0023] Furthermore, if the number of voice recognitions exceeds four times, the user is prompted to perform manual operation through voice.

[0024] The present invention provides a multi-information fusion voice interaction method for a range hood. When the number of voice recognitions exceeds four, the method prompts the user to perform manual operation through voice, thereby ensuring that the user's needs can be responded to in a timely manner and that tasks such as cooking can be carried out smoothly without being delayed indefinitely due to voice interaction problems.

[0025] The present invention also provides a multi-information fusion voice interaction system for range hoods, comprising: a voice noise reduction module: used to perform noise reduction processing on a mixed audio signal to generate a noise-reduced user voice signal; a multi-information fusion module: used to perform weighted fusion of the noise-reduced user voice signal and environmental sensor data based on a preset weighted fusion algorithm to generate semantic features; a voice recognition module: used to convert the semantic features into text instructions, and perform semantic analysis on the text instructions to generate control instructions. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flow chart of a multi-information fusion voice interaction method for range hoods according to the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] Example 1 The present invention provides a multi-information fusion voice interaction method for range hoods, such as Figure 1 As shown, the method includes the following steps: performing noise reduction processing on the mixed audio signal to generate a noise-reduced user voice signal; based on a preset weighted fusion algorithm, weightedly fusing the noise-reduced user voice signal with the environmental sensor data to generate semantic features; converting the semantic features into text instructions, and performing semantic analysis on the text instructions to generate control instructions.

[0029] The system sends control signals to the range hood based on control commands, dynamically adjusting the range hood. The mixed audio signal, consisting of user voice and ambient noise, is captured in real time through audio acquisition devices such as microphones installed on the range hood.

[0030] A deep neural network is used to perform noise reduction on mixed audio signals, removing background noise in strong noise environments, such as range hoods and kitchen noise, and generating a de-noised user voice signal. This de-noised voice signal is of higher quality and more conducive to subsequent voice feature extraction and semantic understanding.

[0031] Based on a preset weighted fusion algorithm, the noise-reduced user voice signal is weightedly fused with the environmental sensor data to generate semantic features, including:

[0032] in, is the enhanced semantic feature matrix after fusion, is the weight matrix, is element-wise multiplication, is the feature matrix of the speech signal, is the feature matrix of environmental sensor data, is a real matrix, is the number of time frames, is the feature dimension.

[0033]

[0034] in, is the normalized weight value, To process the feature matrix using the fully connected layer, is the feature matrix of environmental sensor data, is a real matrix, is the number of time frames, is the feature dimension.

[0035]

[0036] in, is a 1D convolution operation, is the time-frequency feature matrix of the speech signal after noise reduction, , is a real matrix, is the number of time frames, is the feature dimension.

[0037]

[0038] in, is the fully connected layer, is the environmental sensor data matrix, , is the number of time frames, is the feature dimension.

[0039] Perform short-time Fourier transform on the user voice signal after noise reduction and convert it into the time-frequency feature matrix of the voice signal ,in , is the number of time frames, is the feature dimension.

[0040] Use one-dimensional convolution operation to transform the time-frequency feature matrix of speech signal Processing is performed by sliding and calculating the convolution kernel on the time-frequency feature matrix to extract more advanced features of the speech signal and obtain the feature matrix of the speech signal .

[0041] Install a variety of environmental sensors on the range hood, such as temperature sensors, humidity sensors, and oil fume concentration sensors, to collect environmental sensor data in real time and obtain an environmental sensor data matrix , , is the time frame number, and 3 represents the three sensor data of temperature, humidity and oil smoke concentration.

[0042] Use a fully connected layer to process the environmental sensor data matrix The fully connected layer converts the environmental sensor data into a feature matrix of environmental sensor data that is consistent with the dimension of the speech signal feature matrix through linear transformation and nonlinear activation function processing of the input data. .

[0043] The feature matrix of environmental sensor data Input into another fully connected layer , further extracting high-level features related to environmental information. The function normalizes the output of the fully connected layer to obtain the weight matrix , The function can normalize the element values ​​in the weight matrix to between 0 and 1, and ensure that the sum of the elements in each row is 1, so that the weight matrix can reasonably distribute the weights of the speech signal and environmental sensor data in the fusion process.

[0044] Using weighted fusion algorithm

[0045] The weight matrix and the characteristic matrix of the speech signal Perform element-by-element multiplication to obtain the weighted features of the speech signal; and the feature matrix of environmental sensor data Perform element-by-element multiplication to obtain the weighted features of the environmental sensor data; finally, add the two weighted features to obtain the fused enhanced semantic feature matrix .

[0046] The semantic features are converted into text instructions, and the text instructions are semantically analyzed to generate control instructions.

[0047] Specifically: the fused enhanced semantic feature matrix The input is fed into a speech recognition model, such as one based on a recurrent neural network (RNN) or Transformer architecture, which converts the feature matrix into text instructions through analysis and transformation.

[0048] The generated text instructions are semantically analyzed, and natural language processing is used to understand the user's intentions and needs. Based on the results of the semantic analysis, the user's intentions are converted into specific control instructions, such as "increase the wind speed" and "turn on the lighting".

[0049] Based on the generated control command, the range hood's control system sends a control signal to the corresponding actuator. For example, if the control command is "increase air speed," the control system sends a control signal to the air speed control motor to increase the speed. The range hood's actuator then makes adjustments based on the received control signal, enabling dynamic adjustments to the range hood's air speed, lighting, and other functions.

[0050] The dynamic adjustment also includes voice interaction feedback; the voice interaction feedback includes: real-time acquisition of current status information of the range hood, the current status information is fed back to the user through voice broadcast, and voice interaction is performed with the user.

[0051] Specifically, during dynamic range hood adjustment, the system obtains real-time status information about the range hood, such as current wind speed, lighting conditions, and fume concentration. This information is then provided to the user via voice notification, enabling voice interaction. For example, when the range hood wind speed is adjusted, the system may announce, "The current wind speed has been adjusted to high," allowing the user to promptly understand the range hood's operating status.

[0052] The processing of the noise-reduced user voice signal also includes voice recognition. If the voice recognition fails for the first time, a request is issued to the user to reissue the voice command. If the recognition fails for the second time, the operating parameters of the range hood are dynamically adjusted, including reducing the fan speed or switching to low-noise mode. If the voice recognition fails more than four times, a voice prompt is provided to the user to perform manual operation.

[0053] Specifically, the speech recognition module performs the first recognition of the user's speech signal after noise reduction. It uses a feature extraction algorithm, such as the Mel-Frequency Cepstral Coefficient (MFCC) extraction algorithm, to extract the characteristics of the speech signal. These features are then input into the speech recognition model for matching and calculation, resulting in the text result of the first recognition.

[0054] The criteria for successful recognition can be set according to actual needs, for example, the confidence level of the recognition result exceeds a preset threshold such as 0.8, or the recognized text instructions have a high degree of match with the instructions in the preset instruction set, such as a match level exceeding 80%.

[0055] If the first speech recognition is successful, the recognized text instructions will be passed to the subsequent semantic analysis and instruction generation module, and subsequent operations will be carried out according to the normal process.

[0056] If the first attempt at voice recognition fails, the voice interaction system generates a corresponding voice prompt, such as "We didn't recognize your voice command. Please repeat your request." This prompt is played back to the user through the range hood's speakers, prompting them to repeat their voice command. The recognition count counter is incremented by 1, returning the value to 1.

[0057] After playing the voice prompt, wait for the user to issue a new voice command. Set a reasonable waiting time, such as 5 seconds. If no new voice signal is received within the waiting time, the user can be prompted again or the second recognition failure can be directly determined.

[0058] After receiving the voice signal re-sent by the user, a second recognition is performed, and the recognition process is the same as the first recognition.

[0059] If the second speech recognition is successful, the recognition result is weightedly fused with the environmental sensor data to generate semantic features.

[0060] If the second voice recognition attempt fails, the range hood's operating parameters are dynamically adjusted. Depending on the actual situation, the adjustment may involve reducing the fan speed or switching to low-noise mode. For example, if the range hood is currently operating at a high speed, the fan speed is reduced; if the current range hood operating mode is noisy, the range hood is switched to low-noise mode. At the same time, the recognition count counter is incremented by 1, bringing the value to 2.

[0061] The third and fourth voice recognition attempts are similar to the second, with only the number of recognition attempts updated. Specifically, the third and fourth voice recognition attempts are performed after the user re-issues the voice command. After each failed recognition attempt, a corresponding voice prompt (such as "We still haven't recognized your command. Please try again") is generated and played to the user. The recognition count counter is incremented by 1. After the third recognition attempt, the counter is 3, and after the fourth recognition attempt, the counter is 4.

[0062] When the value of the recognition times counter exceeds 4, it means that the voice recognition has failed multiple times, and the user is prompted to operate manually. A voice prompt content is generated, such as "Voice recognition has failed multiple times. It is recommended that you operate the range hood manually", and is played to the user through the speaker.

[0063] In summary, a multi-information fusion voice interaction method for range hoods uses a deep neural network to perform noise reduction on mixed audio signals, which can effectively remove background noise in a strong noise environment and generate higher-quality noise-reduced voice signals, laying a good foundation for subsequent voice feature extraction and semantic understanding, and improving the accuracy of voice interaction; based on a preset weighted fusion algorithm, the noise-reduced user voice signal is weightedly fused with the environmental sensor data, and the voice and environmental information are comprehensively considered to generate an enhanced semantic feature matrix, which helps to understand the user's intention more accurately, generate control instructions that are more in line with actual needs, and improve the intelligence level of voice interaction; according to the generated control The system can send control signals to the executive components of the range hood in real time to achieve dynamic adjustment of the range hood's wind speed, lighting and other functions to meet the user's usage needs in different scenarios. During the dynamic adjustment process, the system obtains the current status information of the range hood in real time and feeds it back to the user through voice broadcast, so that the user can understand the operating status of the range hood in time and enhance the interactive experience between the user and the range hood. In response to the failure of voice recognition, different response strategies are set, including requesting re-voice commands, dynamically adjusting the operating parameters of the range hood and prompting the user to operate manually, which improves the system's fault tolerance and user experience and avoids interaction interruptions due to voice recognition failure.

[0064] Example 2 A multi-information fusion voice interaction system for a range hood comprises a voice noise reduction module, a multi-information fusion module, and a voice recognition module.

[0065] Speech noise reduction module: used to perform noise reduction processing on the mixed audio signal to generate a noise-reduced user speech signal; multi-information fusion module: used to perform weighted fusion of the noise-reduced user speech signal and environmental sensor data based on a preset weighted fusion algorithm to generate semantic features; speech recognition module: used to convert the semantic features into text instructions, and perform semantic analysis on the text instructions to generate control instructions.

[0066] The multi-information fusion voice interaction system for range hoods provided by the present invention can implement method steps consistent with the above method, and therefore will not be described in detail.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A multi-information fusion voice interaction method for range hoods, characterized in that: The following steps are involved: Performing noise reduction processing on the mixed audio signal to generate a noise-reduced user voice signal; Based on a preset weighted fusion algorithm, the noise-reduced user voice signal is weightedly fused with the environmental sensor data to generate semantic features; The semantic features are converted into text instructions, and the text instructions are semantically analyzed to generate control instructions.

2. The multi-information fusion voice interaction method for a range hood according to claim 1, characterized in that: The dynamic adjustment also includes voice interaction feedback; The voice interaction feedback includes: obtaining the current status information of the range hood in real time, feeding back the current status information to the user through voice broadcast, and performing voice interaction with the user.

3. The multi-information fusion voice interaction method for range hoods according to claim 1, characterized in that: The method of weightedly fusing the noise-reduced user voice signal with the environmental sensor data based on a preset weighted fusion algorithm to generate semantic features includes: in, is the enhanced semantic feature matrix after fusion, is the weight matrix, is element-wise multiplication, is the feature matrix of the speech signal, is the feature matrix of environmental sensor data, is a real matrix, is the number of time frames, is the feature dimension.

4. The multi-information fusion voice interaction method for range hoods according to claim 3, characterized in that: The weight matrix for: in, is the normalized weight value, To process the feature matrix using the fully connected layer, is the feature matrix of environmental sensor data, is a real matrix, is the number of time frames, is the feature dimension.

5. The multi-information fusion voice interaction method for range hoods according to claim 3, characterized in that: The feature matrix of the speech signal for: in, is a 1D convolution operation, is the time-frequency feature matrix of the speech signal after noise reduction, , is a real matrix, is the number of time frames, is the feature dimension.

6. The multi-information fusion voice interaction method for range hoods according to claim 3, characterized in that: The feature matrix of the environmental sensor data for: in, is the fully connected layer, is the environmental sensor data matrix, , is the number of time frames, is the feature dimension.

7. The multi-information fusion voice interaction method for a range hood according to claim 1, characterized in that: The processing of the user voice signal after noise reduction further includes voice recognition. When the voice recognition fails for the first time, a request for reissuing the voice command is sent to the user.

8. The multi-information fusion voice interaction method for a range hood according to claim 7, characterized in that: When the voice recognition fails for the second time, the operating parameters of the range hood are dynamically adjusted, and the operating parameters include reducing the fan speed or switching to a low noise mode.

9. The multi-information fusion voice interaction method for a range hood according to claim 8, characterized in that: If the number of voice recognitions exceeds four, the user is prompted to perform manual operation through voice.

10. A multi-information fusion voice interaction system for a range hood, comprising: Speech noise reduction module: used to perform noise reduction processing on the mixed audio signal to generate a noise-reduced user speech signal; Multi-information fusion module: used to perform weighted fusion of the noise-reduced user voice signal and the environmental sensor data based on a preset weighted fusion algorithm to generate semantic features; Speech recognition module: used to convert the semantic features into text instructions, and perform semantic analysis on the text instructions to generate control instructions.

Citation Information

Patent Citations

  • Method and system for increasing voice recognition rate of kitchen ventilator

    CN106439967A

  • Intelligent extractor hood based on voice control

    CN111312221A

  • Multi-modal emotion recognition method based on state space model cross-modal interaction

    CN119128578A

  • Multi-modal data processing method, household appliance and control method and system thereof

    CN119598389A

  • Whole house intelligent scene generation system based on natural language

    CN120010282A