AI smart home control method and device based on embedded host

By preprocessing and enhancing voice commands in the home environment, the problem of low speech recognition accuracy is solved, and higher voice control accuracy and user experience are achieved.

CN120199239APending Publication Date: 2025-06-24SHENZHEN HENGTIAN ZHIXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643414.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In a home environment, the accuracy of speech recognition is low, resulting in misrecognition or non-recognition of instructions, affecting the user experience.

Method used

The AI ​​smart home control method based on embedded host is adopted, and the voice commands issued by users are preprocessed and enhanced, and the voice signal quality is improved. The specific steps include obtaining voice commands, performing short-time Fourier transform and complex spectrum enhancement, extracting voice features, and determining the instruction intention through a pre-trained voice recognition model, and finally generating a control signal to control the smart home device.

Benefits of technology

Effectively remove environmental noise, improve voice signal quality, reduce misidentification, improve the accuracy of voice control, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199239A_ABST
    Figure CN120199239A_ABST
Patent Text Reader

Abstract

The invention discloses an AI smart home control method and device based on an embedded host, and relates to the technical field of smart home control. The method comprises the steps of preprocessing a voice instruction sent by a user to obtain a first voice signal; performing denoising enhancement on the first voice signal, and performing feature extraction to obtain voice features; and taking the voice features as input of a pre-trained voice recognition model to obtain an instruction text, generating a control signal according to the instruction text, and sending the control signal to the corresponding smart home device to control the smart home device to execute a corresponding action. By preprocessing and enhancing the voice instruction sent by the user, the environmental noise can be effectively removed, and the quality of the voice signal is improved. Through the process, the voice recognition model can more accurately capture the instruction content, misrecognition caused by noise or environmental interference is reduced, and the accuracy of voice control is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart home control, and particularly to an AI smart home control method and device based on an embedded host. Background Art

[0002] With the rapid development of artificial intelligence, smart homes have gradually become an important part of modern families. The smart home system works in cooperation with various smart terminal devices through a central control unit embedded in the host, building a complete home automation network. This kind of embedded host usually adopts a high-performance processor and a dedicated AI chip, which are integrated inside a home gateway or a smart central control screen, and are connected to Internet of Things nodes such as lighting systems, temperature control devices, audio-visual entertainment devices, etc. in real time through multi-protocol communication modules such as Zigbee, Wi-Fi, and Bluetooth. In particular, the smart home control method based on voice recognition has received extensive attention and application due to its natural, convenient, and contactless characteristics. Users can easily control various devices in the home through voice commands, such as turning on or off the lights, adjusting the indoor temperature, playing music, controlling the curtains, etc. This interaction method converts natural language commands into structured control commands through the voice interaction middleware built into the embedded host, greatly improving the user experience and making home life more intelligent, efficient, and personalized.

[0003] However, although the smart home control method based on voice recognition has many advantages, it still faces many challenges in practical applications. Among them, the accuracy of voice recognition is one of the key factors affecting the user experience of smart homes. An ideal smart home voice control system should be able to quickly and accurately understand user commands and make corresponding responses. However, in a real home environment, such as the sound played by a TV or a stereo, the noise during kitchen cooking, the sound of a fan or an air conditioner running, these noises will cause the accuracy rate of voice recognition to decrease, resulting in misrecognition or non-recognition of commands and affecting the user experience. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem of low voice recognition accuracy mentioned in the above background art, and to propose an AI smart home control method and device based on an embedded host.

[0005] In the first aspect of the implementation of the present invention, an AI smart home control method based on an embedded host is provided. The method is applied to a home embedded host, and the method includes: Obtain a voice command issued by a user, and preprocess the voice command to obtain a first voice signal; Enhance the first voice signal to obtain a second voice signal, and extract features from the second voice signal to obtain voice features; Use the voice feature as the input of a pre-trained voice recognition model to obtain an instruction text, and determine an instruction intent according to the instruction text; Generate a control signal according to the instruction intent, and send the control signal to a corresponding smart home device to control the smart home device to perform a corresponding action.

[0006] Preferably, the enhancing the first voice signal to obtain a second voice signal includes: Perform a short-time Fourier transform on the first voice signal to obtain a complex spectrum including a real part and an imaginary part; and calculate an amplitude spectrum and a phase spectrum according to the complex spectrum; Use the spectrum of the real part, the spectrum of the imaginary part, the amplitude spectrum and the phase spectrum as the input of a pre-trained voice enhancement model to obtain an enhanced real part and an enhanced imaginary part; Perform waveform reconstruction according to the enhanced real part and the enhanced imaginary part to obtain a second voice signal.

[0007] Preferably, the voice enhancement model includes a first branch, a second branch and an integration layer; the first branch includes a first encoder, a first high-dimensional modeling module and a first decoder; the second branch includes a second encoder, a second high-dimensional modeling module, a second decoder and a third decoder; wherein: The first branch is configured to process the amplitude spectrum and the phase spectrum and output a ratio mask map; The second branch is configured to process the spectrum of the real part and the spectrum of the imaginary part and output a real part compensation map and an imaginary part compensation map; The integration layer is configured to obtain an enhanced real part and an enhanced imaginary part according to the amplitude spectrum, the phase spectrum, the real part compensation map and the imaginary part compensation map.

[0008] Preferably, the first encoder and the second encoder have the same structure but different weights; any one of the encoders includes two serially connected selective dilated convolution modules, and the calculation process of any one of the selective dilated convolution modules includes: ; wherein, X0 is the input of the selective dilated convolution module; f is an operation operator, the subscript DConv represents the dilated convolution operation, and the superscripts 1, 2, 5 represent the dilation rates; represents the sigmoid function; concat represents the concatenation operation; Attention represents the attention module; L1, R1, L2, R2, L3, R3, Y1, Y2, Y3 represent the feature maps output at each operation stage.

[0009] Preferably, the attention module adopts a convolutional block attention module CBAM.

[0010] Preferably, both the first high-dimensional modeling module and the second high-dimensional modeling module adopt a two-layer long short-term memory network structure.

[0011] Preferably, the first decoder, the second decoder, and the third decoder each include six upsampling layers for restoring the compressed feature map to its original size; where: The first decoder is connected to a convolutional block after the six upsampling layers for outputting a ratio mask map; The second decoder and the third decoder are respectively connected to a linear layer after the six upsampling layers for outputting the enhanced real part and imaginary part.

[0012] Preferably, the calculation process of the integration layer includes: ; where M is the magnitude spectrum; IRM is the ratio mask; EM is the enhanced magnitude spectrum; is the phase; and are respectively the real part and imaginary part of the first-stage enhancement; and are respectively the real part compensation and imaginary part compensation; ERP and EIP are respectively the finally enhanced real part and imaginary part.

[0013] Preferably, the voice feature adopts the FBank feature; the speech recognition model adopts the RNNT model.

[0014] In the second aspect of the implementation of the present invention, an AI smart home control device based on an embedded host is provided, and the device includes: A voice acquisition module for acquiring a voice command issued by a user and preprocessing the voice command to obtain a first voice signal; A voice enhancement module for enhancing the first voice signal to obtain a second voice signal and extracting features from the second voice signal to obtain voice features; An instruction recognition module for using the voice features as the input of a pre-trained speech recognition model to obtain an instruction text and determining an instruction intention according to the instruction text; An instruction execution module for generating a control signal according to the instruction intention and sending the control signal to a corresponding smart home device to control the smart home device to perform corresponding actions.

[0015] Advantages of the present invention: The present invention proposes an AI smart home control method based on an embedded host. The method includes: obtaining a voice command issued by a user, and preprocessing the voice command to obtain a first voice signal; enhancing the first voice signal to obtain a second voice signal, and extracting features from the second voice signal to obtain voice features; using the voice features as the input of a pre-trained speech recognition model to obtain an instruction text, and determining an instruction intention according to the instruction text; generating a control signal according to the instruction intention, and sending the control signal to a corresponding smart home device to control the smart home device to perform corresponding actions.

[0016] By preprocessing and enhancing the voice command issued by the user, environmental noise can be effectively removed and the quality of the voice signal can be improved. This process enables the speech recognition model to more accurately capture the command content, reduces misrecognition caused by noise or environmental interference, and improves the accuracy of voice control. Brief Description of the Drawings

[0017] The present invention will be further described below with reference to the accompanying drawings.

[0018] Figure 1 FIG. is a flowchart of an AI smart home control method based on an embedded host provided by an embodiment of the present invention; Figure 2 FIG. is a network architecture diagram of a voice enhancement model provided by an embodiment of the present invention; Figure 3 FIG. is a structural diagram of an AI smart home control device based on an embedded host provided by an embodiment of the present invention. Detailed Embodiments

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] An embodiment of the present invention provides an AI smart home control method based on an embedded host, which is applied to a home embedded host. Refer to Figure 1 , Figure 1 FIG. is a flowchart of an AI smart home control method based on an embedded host provided by an embodiment of the present invention. The method includes the following steps: S101, obtaining a voice command issued by a user, and preprocessing the voice command to obtain a first voice signal.

[0021] S102. Enhance the first voice signal to obtain a second voice signal, and extract features from the second voice signal to obtain voice features.

[0022] S103. Use the voice features as the input of a pre-trained voice recognition model to obtain an instruction text, and determine the instruction intention according to the instruction text.

[0023] S104. Generate a control signal according to the instruction intention, and send the control signal to the corresponding smart home device to control the smart home device to perform corresponding actions.

[0024] Among them, the embedded host provides local processing capabilities in this smart home control method, completing the whole process from voice signal acquisition, processing, feature extraction to instruction parsing and control signal generation. It can effectively support voice recognition and intelligent control tasks, and realize the control of devices by connecting with smart home devices.

[0025] Based on an AI smart home control method based on an embedded host provided by an embodiment of the present invention, by preprocessing and enhancing the voice commands issued by users, environmental noise can be effectively removed and the quality of voice signals can be improved. This process enables the voice recognition model to more accurately capture the instruction content, reduces misrecognition caused by noise or environmental interference, and improves the accuracy of voice control.

[0026] In one implementation, the preprocessing includes processing such as endpoint detection and pre-emphasis.

[0027] In one embodiment, enhancing the first voice signal to obtain a second voice signal includes: Step 1. Perform a short-time Fourier transform on the first voice signal to obtain a complex spectrum including a real part and an imaginary part; and calculate an amplitude spectrum and a phase spectrum according to the complex spectrum.

[0028] Step 2. Use the spectrum of the real part, the spectrum of the imaginary part, the amplitude spectrum and the phase spectrum as the input of a pre-trained voice enhancement model to obtain an enhanced real part and an enhanced imaginary part.

[0029] Step 3. Perform waveform reconstruction according to the enhanced real part and the enhanced imaginary part to obtain a second voice signal.

[0030] This embodiment combines short-time Fourier transform, complex spectrum enhancement and waveform reconstruction technologies, not only improving the accuracy of voice recognition, but also enhancing the anti-noise ability of the system, making AI smart home control more efficient, natural and reliable.

[0031] In one embodiment, refer to Figure 2 , Figure 2This is a network architecture diagram of a voice enhancement model provided by an embodiment of the present invention. The voice enhancement model includes a first branch, a second branch, and an integration layer; the first branch includes a first encoder, a first high-dimensional modeling module, and a first decoder connected in sequence; the second branch includes a second encoder, a second high-dimensional modeling module, a second decoder, and a third decoder connected in sequence; both the first branch and the second branch adopt a U-Net architecture, and skip connections are established between the decoder and the encoder. Among them: The first branch is used to process the magnitude spectrum (M) and the phase spectrum (θ), and output a ratio mask map IRM; The second branch is used to process the real part spectrum (RP) and the imaginary part spectrum (IP), and output a real part compensation map (CR) and an imaginary part compensation map (CI); The integration layer is used to obtain an enhanced real part (ERP) and an enhanced imaginary part (EIP) according to the magnitude spectrum, the phase spectrum, the real part compensation map, and the imaginary part compensation map.

[0032] In one implementation, the first encoder includes two serially connected selective dilated convolution modules (SDCM); the first high-dimensional modeling module includes a two-layer long short-term memory network (2-Layer LSTM); the first decoder includes six upsampling layers and a convolution block (CM). The calculation process of the first branch includes: Step 1, construct the magnitude spectrum and the phase map into a two-channel feature F1 as the input of the first encoder. The first encoder uses two serially connected selective dilated convolution modules to process the two-channel feature and outputs six feature maps. Specifically, the calculation process of a selective dilated convolution module includes: ; Among them, X0 is the input of the selective dilated convolution module; f is the operation operator, the subscript DConv represents the dilated convolution operation, and the superscripts 1, 2, and 5 represent the dilation rates; represents the sigmoid function; represents the element-wise multiplication operation; concat represents the concatenation operation; Attention represents the attention module, and the convolutional block attention module CBAM can be used; L1, R1, L2, R2, L3, R3, Y1, Y2, and Y3 represent the feature maps output at each operation stage. A selective dilated convolution module outputs three feature maps and makes a skip connection to the decoder. For distinction, the feature maps sequentially output by the first selective dilated convolution module are denoted as Y11, Y12, and Y13, and the feature maps sequentially output by the second selective dilated convolution module are denoted as Y21, Y22, and Y23.

[0033] The encoder uses dilated convolution (DConv) with multiple dilation rates (1, 2, 5) to expand the receptive field, which helps to capture voice features at different scales, so as to extract comprehensive local and global features.

[0034] In Step 2, the first high-dimensional modeling module flattens the feature map Y23 into a shape suitable for input to the LSTM, then processes it using a two-layer long short-term memory network to obtain high-dimensional features, and reshapes the high-dimensional features into multi-channel features suitable for processing by the next layer of the network.

[0035] Using two layers of LSTM for high-dimensional modeling can effectively model the temporal dependence in the speech signal. The LSTM can learn the speech structure information in the time dimension, reduce problems such as speech breaks and blurs, and improve the naturalness and coherence of the enhanced speech.

[0036] In Step 3, the first decoder uses six upsampling layers to restore the compressed feature map to its original size. Specifically, for each upsampling layer, the feature map output by the previous layer and the feature map of the encoder skip connection are concatenated as the input to this upsampling layer, and sub-pixel convolution (SPC, Sub-Pixel Convolution) is used to upsample the input feature map to output a feature map of a preset size. Traditional upsampling (such as transposed convolution) may cause interpolation errors, while SPC can generate a smoother speech signal by learning higher-precision feature points, reducing distortion. Combined with skip connections, artifacts or noise introduced during the speech enhancement process can be effectively avoided when restoring speech features.

[0037] In Step 4, the first decoder uses a convolutional block to process the feature map output by the last upsampling layer and outputs a ratio mask map. Specifically, the calculation process of the convolutional block CM includes: ; where Z is the feature map output by the last upsampling layer, that is, the input to this convolutional block; is a two-dimensional convolutional operation; and are activation functions; P1, P2, P3 are feature maps generated during the operation process; P4 is the desired ratio mask map.

[0038] In one implementation, the second encoder includes two serially connected selective dilated convolution modules, which have the same structure as the first encoder but different weights; the second high-dimensional modeling module includes a two-layer long short-term memory network; the second decoder and the third decoder both include six upsampling layers (with the same structure as the six upsampling layers of the first decoder but different weights) and a linear layer (Linear). The calculation process of the second branch includes: In Step 1, the real part spectrum (RP) and the imaginary part spectrum (IP) are constructed into a two-channel feature F2 as the input to the second encoder, and six feature maps are output.

[0039] In Step 2, the second high-dimensional modeling module processes the feature map Y23 output by the second selected dilated convolution module of the second encoder using two layers of long short-term memory networks to obtain high-dimensional features.

[0040] In Step 3, the second decoder uses six upsampling layers to restore the compressed feature map to its original size. The third decoder uses six upsampling layers to restore the compressed feature map to its original size.

[0041] In Step 4, the second decoder uses a linear layer to process the feature map output by the last upsampling layer and outputs the real part compensation (CR). The third decoder uses a linear layer to process the feature map output by the last upsampling layer and outputs the imaginary part compensation (CI).

[0042] The operation process of the second branch is similar to that of the first branch. The difference is that in the decoder part, feature reconstruction is performed with different objectives.

[0043] In one implementation, the calculation process of the integrated computing layer includes: ; where M is the magnitude spectrum; IRM is the ratio mask; EM is the enhanced magnitude spectrum; is the phase; and are the real and imaginary parts of the one-stage enhancement respectively; and are the real part compensation and imaginary part compensation respectively; ERP and EIP are the real and imaginary parts of the final enhancement respectively.

[0044] In one embodiment, the extracted speech features can be Fbank features (Filter Bank Features) or MFCC features (Mel-Frequency Cepstral Coefficients).

[0045] In one embodiment, the speech recognition model can adopt an RNNT (Recurrent Neural Network Transducer) model or a Conformer (CNN+Transformer) model.

[0046] In one embodiment, the speech recognition model converts the speech signal into text information, and then can match the text information with a standard instruction set through natural language technology to determine the user's intention, thereby generating a control signal.

[0047] An embodiment of the present invention provides an AI smart home control device based on an embedded host, which is applied to a home embedded host. Refer to Figure 3 , Figure 3 which is a structural diagram of an AI smart home control device based on an embedded host provided by an embodiment of the present invention. The device includes: A voice acquisition module, configured to acquire a voice command issued by a user and preprocess the voice command to obtain a first voice signal.

[0048] A voice enhancement module, configured to enhance the first voice signal to obtain a second voice signal and extract features of the second voice signal to obtain voice features.

[0049] An instruction recognition module, configured to use the voice features as the input of a pre-trained voice recognition model to obtain an instruction text, and determine an instruction intention according to the instruction text.

[0050] An instruction execution module, configured to generate a control signal according to the instruction intention and send the control signal to a corresponding smart home device to control the smart home device to perform a corresponding action.

[0051] Based on an AI smart home control method provided by an embodiment of the present invention, by preprocessing and enhancing a voice command issued by a user, environmental noise can be effectively removed and the quality of the voice signal can be improved. This process enables the voice recognition model to more accurately capture the instruction content, reduces misrecognition caused by noise or environmental interference, and improves the accuracy of voice control.

[0052] The above has described an embodiment of the present invention in detail, but the described content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.

Claims

1. An AI smart home control method based on an embedded host, characterized in that: The method is applied to a home embedded host, and the method comprises: Acquire a voice command issued by a user, and pre-process the voice command to obtain a first voice signal; enhancing the first speech signal to obtain a second speech signal, and extracting features from the second speech signal to obtain speech features; Using the speech feature as an input of a pre-trained speech recognition model to obtain a command text, and determining the command intent based on the command text; According to the instruction intention, a control signal is generated, and the control signal is sent to the corresponding smart home device to control the smart home device to perform a corresponding action.

2. The AI ​​smart home control method based on an embedded host according to claim 1, characterized in that: The enhancing the first speech signal to obtain a second speech signal comprises: Performing a short-time Fourier transform on the first speech signal to obtain a complex spectrum including a real part and an imaginary part; and calculating an amplitude spectrum and a phase spectrum based on the complex spectrum; Using the spectrum of the real part, the spectrum of the imaginary part, the amplitude spectrum and the phase spectrum as inputs of a pre-trained speech enhancement model to obtain enhanced real and imaginary parts; The waveform is reconstructed according to the enhanced real part and imaginary part to obtain a second speech signal.

3. The AI ​​smart home control method based on embedded host according to claim 2 is characterized in that: The speech enhancement model includes a first branch, a second branch and an integration layer; the first branch includes a first encoder, a first high-dimensional modeling module and a first decoder; the second branch includes a second encoder, a second high-dimensional modeling module, a second decoder and a third decoder; wherein: The first branch is used to process the amplitude spectrum and the phase spectrum and output a ratio masking map; The second branch is used to process the spectrum of the real part and the spectrum of the imaginary part, and output a real part compensation graph and an imaginary part compensation graph; The integration layer is used to obtain enhanced real and imaginary parts according to the amplitude spectrum, the phase spectrum, the real compensation map and the imaginary compensation map.

4. The AI ​​smart home control method based on embedded host according to claim 3 is characterized in that: The first encoder and the second encoder have the same structure and different weights; any encoder includes two selected hole convolution modules connected in series, and the calculation process of any selected hole convolution module includes: ; Among them, X0 is the input of the selected dilated convolution module; f is the operation operator, the subscript DConv represents the dilated convolution operation, and the superscripts 1, 2, and 5 represent the expansion rate; Represents the sigmoid function; concat represents the concatenation operation; Attention represents the attention module; L1, R1, L2, R2, L3, R3, Y1, Y2, and Y3 represent the feature maps output at each operation stage.

5. The AI ​​smart home control method based on embedded host according to claim 4 is characterized in that: The attention module adopts a convolutional block attention module CBAM.

6. The AI ​​smart home control method based on embedded host according to claim 3, characterized in that: Both the first high-dimensional modeling module and the second high-dimensional modeling module adopt a two-layer long short-term memory network structure.

7. The AI ​​smart home control method based on embedded host according to claim 3, characterized in that: The first decoder, the second decoder and the third decoder each include six upsampling layers, and the six upsampling layers are used to restore the compressed feature map to its original size; wherein: The first decoder is connected with a convolution block after the six upsampling layers to output a ratio mask map; The second decoder and the third decoder are respectively connected to a linear layer after the six upsampling layers to output enhanced real and imaginary parts.

8. The AI ​​smart home control method based on embedded host according to claim 3 is characterized in that: The calculation process of the integration layer includes: ; Where, M is the amplitude spectrum; IRM is the ratio masking; EM is the enhanced amplitude spectrum; is the phase; and are the real and imaginary parts of one-stage enhancement, respectively; and are real part compensation and imaginary part compensation respectively; ERP and EIP are the real and imaginary parts of the final enhancement respectively.

9. The AI ​​smart home control method based on embedded host according to claim 3, characterized in that: The speech feature adopts FBank feature; the speech recognition model adopts RNNT model.

10. An AI smart home control device based on an embedded host, characterized in that: The device is applied to a home embedded host, and the device comprises: A voice acquisition module, used to obtain a voice command issued by a user, and pre-process the voice command to obtain a first voice signal; A speech enhancement module, used for enhancing the first speech signal to obtain a second speech signal, and performing feature extraction on the second speech signal to obtain speech features; An instruction recognition module is used to use the speech features as input to a pre-trained speech recognition model to obtain an instruction text, and determine the instruction intent based on the instruction text; The instruction execution module is used to generate a control signal according to the instruction intention, and send the control signal to the corresponding smart home device to control the smart home device to perform a corresponding action.