Infrared AI smart glasses for controlling smart home based on voice recognition technology
Through infrared AI smart glasses based on voice recognition technology, convenient smart home control is achieved, solving the shortcomings of smart glasses in the existing technology in device control, and improving the flexibility of user experience and device access.
Patent Information
- Application Number
- CN202510430038.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing smart glasses lack efficient voice recognition and precise device control capabilities in smart home control. Mobile APP control is not convenient for use when both hands are busy, and smart speakers cannot accurately locate user demand scenarios.
Design an infrared AI smart glasses based on speech recognition technology, including microphone arrays, AI processing chips, wireless communication modules, speakers and software systems, using Mel frequency cepspectral coefficient feature extraction, semantic understanding of Transformer architecture and multi-protocol wireless communication to achieve accurate speech recognition and device control.
Users can easily control smart home devices through voice, improve operation efficiency, reduce error operation rate, support access to smart home devices of various brands and types, and provide free selection and matching.
Smart Images

Figure CN120299456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart home control, and specifically to an infrared AI smart glasses for controlling smart home based on voice recognition technology. Background Art
[0002] With the rapid development of the Internet of Things and artificial intelligence technologies, smart homes have gradually entered people's lives. Currently, the control methods of smart homes mainly include mobile APP control, smart speaker voice control, etc. However, mobile APP control requires manual operation, which is inconvenient to use when both hands are busy; although smart speakers can achieve voice control, there are problems such as fixed positions and inability to accurately locate user demand scenarios.
[0003] Wearable devices have developed rapidly in recent years. As an important branch of them, smart glasses have the characteristics of portability and availability at any time. However, most of the existing smart glasses focus on functions such as information display, photo taking and video recording, and are less applied in smart home control, and lack efficient voice recognition and accurate device control capabilities. Therefore, developing an AI smart glasses for controlling smart home based on voice recognition can make up for the deficiencies of existing smart home control methods and improve users' smart home control experience, which has important practical significance. Summary of the Invention
[0004] The purpose of the present invention is to provide an infrared AI smart glasses for controlling smart home based on voice recognition technology to solve the problems raised in the above background art.
[0005] In view of the above problems, the technical solution proposed by the present invention is:
[0006] An infrared AI smart glasses for controlling smart home based on voice recognition technology, including a glasses body, a hardware system and a software system arranged on the glasses body. The hardware system includes a microphone array, an AI processing chip, a wireless communication module, a speaker, and a power supply module. The software system integrates a voice recognition algorithm, a semantic understanding module, a device control module, and a voice prompt module on the AI processing chip.
[0007] As a preferred technical solution of the present invention, the microphone array is used to collect user voice signals, the AI processing chip performs real-time processing and analysis on the voice signals, the wireless communication module realizes data mutual transmission between the glasses body and smart home devices, the speaker is used to receive and play the voice prompt information processed by the AI processing chip, and the power supply module supplies power to the microphone array, the AI processing chip, the wireless communication module, and the speaker;
[0008] Specifically, the microphone array is installed at one of the temple arms of the glasses body and close to the user's ear, the speaker is installed at the other temple arm of the glasses body and close to the user's ear, the AI processing chip and the wireless communication module are installed inside the frame of the glasses body, the power module is installed inside the temple arm of the glasses body, and the microphone array adopts a multi-microphone combination method.
[0009] As a preferred technical solution of the present invention, the speech recognition algorithm preprocesses, extracts features, trains a model, and makes inferences on the speech signal to obtain a speech recognition result;
[0010] Specifically, in the speech recognition algorithm, the preprocessing of the speech signal includes frame segmentation and windowing operations, and the extracted speech feature is the Mel-frequency cepstral coefficient, and its calculation steps are as follows:
[0011] Calculate the power spectrum P j (k) of the j-th frame signal:
[0012] where P j (k) is the power spectrum value of the j-th frame signal at the k-th frequency point, j is the frame index, starting from 0; k is the frequency point index, starting from 0, N is the frame length, xj(n) is the value of the n-th sample point in the j-th frame of speech signal, n is the index of the sample point within this frame, ranging from 0 to N - 1, e is the natural constant, and i is the imaginary unit.
[0013] Map the power spectrum to the Mel-frequency domain through the Mel filter bank to obtain the output S j (m) of the m-th Mel filter:
[0014] where H m (k) is the frequency response of the m-th Mel filter, M is the number of Mel filters, m is the Mel filter index, N is the frame length, P j (k) is the power spectrum value of the j-th frame signal at the k-th frequency point, k is the frequency point index, and j is the frame index, starting from 0.
[0015] Take the logarithm of S j (m) and perform a discrete cosine transform to obtain the Mel-frequency cepstral coefficient feature c j (l):
[0016]
[0017] where c j(l) is the l-th Mel-frequency cepstral coefficient eigenvalue of the j-th frame signal, where l is the index of the Mel-frequency cepstral coefficient feature, L is the dimension of the Mel-frequency cepstral coefficient feature, M is the number of Mel filters, and S j (m) is the output value of the j-th frame signal after passing through the m-th Mel filter, where j is the index of the frame and m is the index of the Mel filter.
[0018] Specifically, in the speech recognition algorithm, the speech recognition algorithm is trained using labeled speech data, and the cross-entropy loss function is used to measure the difference between the model prediction result and the true label. The cross-entropy loss function is:
[0019] where T is the number of training samples, C is the number of categories, y jr is the r-th true label of the h-th sample, is the predicted probability of the r-th category of the h-th sample by the model, where h is the index of the sample and r is the index of the true label.
[0020] As a preferred technical solution of the present invention, the semantic understanding module performs keyword matching, grammar analysis, and intention judgment on the speech recognition result to generate a control instruction;
[0021] Specifically, in the semantic understanding module, the pre-trained language model BERT based on the Transformer architecture is used for semantic understanding. The speech recognition result is input into the language model BERT to obtain a semantic representation, and then the user intention is classified and judged through a fully connected layer. Let the input text be w, the semantic representation v is obtained through the language model BERT, and the classification result z is obtained through the fully connected layer. The formula is:
[0022] v = BERT(w)
[0023] z = softmax(W·v + b)
[0024] where W is the weight matrix of the fully connected layer and b is the bias vector.
[0025] As a preferred technical solution of the present invention, the device control module sends a control signal to the corresponding smart home device through the wireless communication module according to the control instruction. The specific steps for the device control module to execute the control instruction include:
[0026] S1. Receive the control instruction transmitted by the semantic understanding module, where the control instruction includes the device name, operation type, and operation parameters;
[0027] S2. Query in the internally maintained device information database according to the parsed device name. The device information database stores the communication protocols, device addresses, and operation instruction sets supported by various smart home devices.
[0028] S3. Generate a control signal according to the queried device information, the parsed operation type, and parameters in accordance with the corresponding communication protocol rules.
[0029] S4. Send the generated control signal to the target smart home device through the wireless communication module of the glasses body.
[0030] S5. The voice prompt module receives the operation result feedback information sent after the smart home device executes the operation.
[0031] As a preferred technical solution of the present invention, the voice prompt module receives the operation result feedback information of the smart home device and plays it to the user through the speaker. The specific steps of the voice prompt module working include:
[0032] S6. The voice prompt module receives the operation result feedback information through the wireless communication module, parses the feedback information, and extracts the key information therein.
[0033] S7. The voice prompt module matches according to the parsed operation result in the internally preset prompt text library. The prompt text library is classified and stored according to different device types, operation types, and operation results.
[0034] S8. After the voice prompt module obtains the matching prompt text, it calls the speech synthesis engine to perform text-to-speech conversion.
[0035] S9. The voice prompt module transmits the generated voice signal to the speaker for playback.
[0036] On the other hand, the present invention provides a method for an infrared AI smart glasses to control smart home based on voice recognition technology, including the following steps:
[0037] Step 1: Wear the smart glasses. The microphone array collects the user's voice signal. The microphone array is installed at one of the temple arms of the glasses body and close to the user's ear, and can accurately capture the voice in a complex environment. The collected voice signal is transmitted to the AI processing chip. In the chip, the speech recognition algorithm preprocesses the voice signal, and then extracts the Mel frequency cepstral coefficients as speech features.
[0038] Step 2: The AI processing chip uses the trained speech recognition model to perform inference on the preprocessed speech features. After the model inference, the speech recognition result is obtained and input into the semantic understanding module. The semantic understanding module performs semantic understanding through rule matching and the pre-trained language model BERT based on the Transformer architecture;
[0039] Step 3: The device control module receives the control instructions transmitted by the semantic understanding module and parses out the device name. Subsequently, it queries the corresponding device information in the device information database maintained internally, and generates a control signal according to the queried device information, the parsed operation type, and parameters in accordance with the corresponding communication protocol rules. Finally, through the wireless communication module of the glasses body, the generated control signal is sent to the target smart home device;
[0040] Step 4: After the smart home device receives the control signal and executes the operation, it will send operation result feedback information to the glasses body. The voice prompt module receives this feedback information through the wireless communication module, parses the feedback information, extracts the key information therein, and then matches it in the preset prompt text library internally. After obtaining the matching prompt text, it calls the speech synthesis engine to perform text-to-speech conversion, and transmits the generated voice signal to the speaker of the smart glasses for playback. The speaker is installed at the other temple of the glasses body and close to the user's ear, so as to feedback information such as device operation results and system prompts to the user in the form of voice.
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] First, users can conveniently control smart home devices through voice, and the operation is more convenient and efficient, improving the experience of smart home control;
[0043] Second, the speech recognition and semantic understanding technologies can accurately identify user instructions, reduce the misoperation rate, and ensure that the smart home devices operate according to the user's expectations;
[0044] Third, the multi-protocol wireless communication module supports the access of various brands and types of smart home devices. Users can freely select and match smart home devices without being restricted by the device brand. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a system block diagram of an infrared AI smart glasses for controlling smart homes based on speech recognition technology disclosed in an embodiment of the present invention;
[0046] Figure 2 It is a three-dimensional structure schematic diagram of an infrared AI smart glasses for controlling smart homes based on speech recognition technology disclosed in an embodiment of the present invention;
[0047] Figure 3Device control flowchart of an infrared AI smart glasses for controlling smart home based on speech recognition technology disclosed in the embodiments of the present invention;
[0048] Figure 4 Speech prompt flowchart of an infrared AI smart glasses for controlling smart home based on speech recognition technology disclosed in the embodiments of the present invention.
[0049] In the figure: 100, glasses body; 200, microphone array; 300, speaker. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] Please refer to Figure 1 - Figure 4 , the present invention provides a technical solution: an infrared AI smart glasses for controlling smart home based on speech recognition technology, including a glasses body (100), a hardware system and a software system provided on the glasses body (100), the hardware system includes a microphone array (200), an AI processing chip, a wireless communication module, a speaker (300), and a power module, and the software system integrates a speech recognition algorithm, a semantic understanding module, a device control module, and a speech prompt module on the AI processing chip.
[0052] As an embodiment of the present invention, further, the microphone array (200) is used to collect user voice signals, the AI processing chip performs real-time processing and analysis on the voice signals, the wireless communication module realizes data transmission between the glasses body (100) and the smart home device, the speaker (300) is used to receive and play the voice prompt information processed by the AI processing chip, and the power module supplies power to the microphone array (200), the AI processing chip, the wireless communication module, and the speaker (300);
[0053] Specifically, the microphone array (200) is installed at one of the temple arms of the glasses body (100) and close to the user's ear, the speaker (300) is installed at the other temple arm of the glasses body (100) and close to the user's ear, the AI processing chip and the wireless communication module are installed inside the frame of the glasses body (100), the power module is installed inside the temple arm of the glasses body (100), and the microphone array (200) adopts a multi-microphone combination method.
[0054] As an embodiment of the present invention, further, the speech recognition algorithm preprocesses, extracts features, trains a model, and performs inference on a speech signal to obtain a speech recognition result;
[0055] Specifically, in the speech recognition algorithm, the preprocessing performed on the speech signal includes frame division and windowing operations, and the extracted speech feature is the Mel-frequency cepstral coefficient. Its calculation steps are as follows:
[0056] Calculate the power spectrum P j (k) of the j-th frame signal:
[0057] where P j (k) is the power spectrum value of the j-th frame signal at the k-th frequency point, j is the index of the frame, starting from 0; k is the index of the frequency point, starting from 0, N is the frame length, xj(n) is the value of the n-th sample point in the j-th frame of the speech signal, n is the index of the sample point within the frame, ranging from 0 to N - 1, e is the natural constant, and i is the imaginary unit.
[0058] Map the power spectrum to the Mel-frequency domain through a Mel filter bank to obtain the output S j (m) of the m-th Mel filter:
[0059] where H m (k) is the frequency response of the m-th Mel filter, M is the number of Mel filters, m is the index of the Mel filter, N is the frame length, P j (k) is the power spectrum value of the j-th frame signal at the k-th frequency point, k is the index of the frequency point, and j is the index of the frame, starting from 0.
[0060] Take the logarithm of S j (m) and perform a discrete cosine transform to obtain the Mel-frequency cepstral coefficient feature c j (l):
[0061]
[0062] where c j (l) is the l-th Mel-frequency cepstral coefficient feature value of the j-th frame signal, l is the index of the Mel-frequency cepstral coefficient feature, L is the dimension of the Mel-frequency cepstral coefficient feature, M is the number of Mel filters, and S j (m) is the output value of the j-th frame signal after passing through the m-th Mel filter, j is the index of the frame, and m is the index of the Mel filter.
[0063] Specifically, in the speech recognition algorithm, the speech recognition algorithm is trained using labeled speech data, and the cross-entropy loss function is used to measure the difference between the model prediction result and the true label. The cross-entropy loss function is as follows:
[0064] where T is the number of training samples, C is the number of categories, yj r is the r-th true label of the h-th sample, is the predicted probability of the r-th category of the h-th sample by the model, h is the index of the sample, and r is the index of the true label.
[0065] As an embodiment of the present invention, further, the semantic understanding module performs keyword matching, grammar analysis, and intention judgment on the speech recognition result to generate a control instruction;
[0066] Specifically, in the semantic understanding module, the pre-trained language model BERT based on the Transformer architecture is used for semantic understanding. The speech recognition result is input into the language model BERT to obtain a semantic representation, and then the user intention is classified and judged through a fully connected layer. Let the input text be w, the semantic representation v is obtained through the language model BERT, and the classification result z is obtained through the fully connected layer. The formula is:
[0067] v = BERT(w)
[0068] z = softmax(W·v + b)
[0069] where W is the weight matrix of the fully connected layer and b is the bias vector.
[0070] As an embodiment of the present invention, further, the device control module sends a control signal to the corresponding smart home device through the wireless communication module according to the control instruction. The specific steps for the device control module to execute the control instruction include:
[0071] S1. Receive the control instruction transmitted by the semantic understanding module. The control instruction includes the device name, operation type, and operation parameters;
[0072] S2. Query according to the parsed device name in the device information database maintained internally. The device information database stores the communication protocols, device addresses, and operation instruction sets supported by various smart home devices;
[0073] S3. Generate a control signal according to the queried device information and the parsed operation type and parameters according to the corresponding communication protocol rules;
[0074] S4. Send the generated control signal to the target smart home device through the wireless communication module of the glasses body (100);
[0075] S5. The voice prompt module receives the operation result feedback information sent after the smart home device executes an operation.
[0076] As an embodiment of the present invention, further, the voice prompt module receives the operation result feedback information of the smart home device and plays it to the user through the loudspeaker (300). The specific steps of the voice prompt module working include:
[0077] S6. The voice prompt module receives the operation result feedback information through the wireless communication module, analyzes the feedback information, and extracts the key information therein.
[0078] S7. The voice prompt module matches according to the analyzed operation result in the internally preset prompt text library, and the prompt text library is classified and stored according to different device types, operation types, and operation results.
[0079] S8. After the voice prompt module obtains the matched prompt text, it calls the speech synthesis engine to perform text-to-speech conversion.
[0080] S9. The voice prompt module transmits the generated voice signal to the loudspeaker (300) for playing.
[0081] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0082] Please refer to Figure 1 - Figure 4 , the present invention provides a technical solution: an infrared AI smart glasses for controlling smart home based on voice recognition technology, including the following steps:
[0083] As an embodiment of the present invention, further, a user walks into the living room, at which time the TV in the living room is playing a program, and there is a certain amount of noise in the environment. However, the microphone array 200 installed on the temple of the glasses body 100 near the ear, with its multi-microphone combination and precise sound pickup capability, clearly captures the user's voice command: "Turn on the lights in the living room and adjust the air conditioner to 26 degrees". The collected voice signal is quickly transmitted to the AI processing chip. Within the chip, the speech recognition algorithm immediately pre-processes the voice signal, first divides it into frames, then performs a windowing operation, and then calculates the Mel-frequency cepstral coefficients as speech features. For example, the power spectrum of a frame signal is calculated, the power distribution of the frame at different frequency points is determined, and then the power spectrum is mapped to the Mel frequency domain through a Mel filter group, and finally the logarithm is taken and a discrete cosine transform is performed to complete the feature extraction.
[0084] As an embodiment of the present invention, further, the AI processing chip uses a pre-trained speech recognition model to infer the extracted speech features. The model is based on an end-to-end deep neural network structure, combined with a convolutional neural network and a long short-term memory network. It can accurately recognize the speech content under the training of a large amount of labeled speech data. After the model infers the speech recognition result, it is input into the semantic understanding module. On the one hand, the semantic understanding module performs preliminary matching through pre-set rules, such as "if the text contains 'turn on' and 'lamp in the living room', a control instruction to turn on the living room light is generated"; on the other hand, the pre-trained language model BERT based on the Transformer architecture is used to input the text after speech recognition into it to obtain a semantic representation, and then classified through the fully connected layer, accurately judging that the user's intention is to control the living room light and air conditioning equipment, and generating corresponding control instructions, which include living room lights, air conditioners, operation types: turn on, adjust temperature and operation parameters; 26 degrees.
[0085] As an embodiment of the present invention, further, after the device control module receives the control instruction from the semantic understanding module, it queries the internally maintained device information database according to the parsed device names "living room lamp" and "air conditioner". The database stores the communication protocols, device addresses and operation instruction sets supported by various types of smart home devices. Based on the queried information and the parsed operation types and parameters, a control signal is generated according to the corresponding communication protocol rules. For example, for the living room lamp, a control signal for turning on the light is generated according to the Wi-Fi protocol; for the air conditioner, a control signal for adjusting the temperature to 26 degrees is generated according to the Bluetooth protocol. Subsequently, the control signals are sent to the corresponding target smart home devices through the wireless communication module inside the frame of the glasses body 100.
[0086] As an embodiment of the present invention, further, after the living room lamp receives the control signal, it lights up immediately, and the air conditioner also starts to adjust the temperature to 26 degrees. After the operation is completed, these two smart home devices respectively send operation result feedback information to the glasses body 100. The voice prompt module receives these feedback information through the wireless communication module, and parses them to extract key information, such as "the living room lamp has been turned on" and "the air conditioner has been adjusted to 26 degrees". Then, according to the parsed operation results, the voice prompt module matches in the internally preset prompt text library to find the corresponding prompt text. Then it calls the voice synthesis engine to convert the prompt text into a voice signal, which is transmitted to the speaker 300 installed near the ear on the other temple of the glasses body 100 for playback. The user hears the voice prompt of "the living room lamp has been turned on, and the air conditioner has been adjusted to 26 degrees", and confirms that the device has operated according to his own instructions.
Claims
1. An infrared AI smart glasses for controlling smart home based on speech recognition technology, characterized in that, It includes a glasses body (100), a hardware system and a software system provided on the glasses body (100); The hardware system includes a microphone array (200), an AI processing chip, a wireless communication module, a speaker (300), and a power module. The microphone array (200) is used to collect user voice signals. The AI processing chip performs real-time processing and analysis on the voice signals. The wireless communication module realizes data transmission between the glasses body (100) and smart home devices. The speaker (300) is used to receive and play the voice prompt information processed by the AI processing chip. The power module supplies power to the microphone array (200), the AI processing chip, the wireless communication module, and the speaker (300); The software system integrates a voice recognition algorithm, a semantic understanding module, a device control module, and a voice prompt module on the AI processing chip. The voice recognition algorithm performs preprocessing, feature extraction, model training, and inference on the voice signals to obtain a voice recognition result. The semantic understanding module performs keyword matching, grammar analysis, and intention judgment on the voice recognition result to generate a control instruction. The device control module sends a control signal to the corresponding smart home device through the wireless communication module according to the control instruction. The voice prompt module receives the operation result feedback information of the smart home device and plays it to the user through the speaker (300).
2. The infrared AI smart glasses for controlling smart home based on voice recognition technology according to claim 1, characterized in that, In the voice recognition algorithm, the preprocessing performed on the voice signals includes frame division and windowing operations, and the extracted voice feature is the Mel Frequency Cepstral Coefficient.
3. The infrared AI smart glasses for controlling smart home based on speech recognition technology according to claim 1, wherein, In the voice recognition algorithm, the voice recognition algorithm is trained using labeled voice data, and the cross-entropy loss function is used to measure the difference between the model prediction result and the true label.
4. An infrared AI smart glasses for controlling smart home based on voice recognition technology according to claim 1, wherein In the semantic understanding module, the pre-trained language model BERT based on the Transformer architecture is used for semantic understanding. The voice recognition result is input into the language model BERT to obtain a semantic representation, and then the user intention is classified and judged through a fully connected layer.
5. The infrared AI smart glasses for controlling smart home based on voice recognition technology according to claim 1, characterized in that, The specific steps for the device control module to execute the control instruction include: S1. Receive the control instruction transmitted by the semantic understanding module. The control instruction includes the device name, operation type, and operation parameters; S2. Query according to the parsed device name in the device information database maintained internally. The device information database stores the communication protocols, device addresses, and operation instruction sets supported by various smart home devices; S3. Generate a control signal according to the queried device information and the parsed operation type and parameters according to the corresponding communication protocol rules; S4. Send the generated control signal to the target smart home device through the wireless communication module of the glasses body (100); S5. The voice prompt module receives the operation result feedback information sent after the smart home device executes the operation.
6. The infrared AI smart glasses for controlling smart home based on speech recognition technology according to claim 5, characterized in that, The specific steps for the voice prompt module to work include: S6. The voice prompt module receives the operation result feedback information through the wireless communication module, parses the feedback information, and extracts the key information therein; S7. The voice prompt module matches according to the parsed operation result in a pre-set prompt text library inside, and the prompt text library is classified and stored according to different device types, operation types, and operation results; S8. After the voice prompt module obtains the matched prompt text, it calls a voice synthesis engine to perform text-to-speech conversion; S9. The voice prompt module transmits the generated voice signal to the speaker (300) for playback.
7. The infrared AI smart glasses for controlling smart home based on voice recognition technology according to claim 1, characterized in that, The microphone array (200) adopts a multi-microphone combination method.
8. An infrared AI smart glasses for controlling smart home based on speech recognition technology according to claim 1, characterized in that, The microphone array (200) is installed at one of the temple arms of the glasses body (100) and close to the user's ear, the speaker (300) is installed at the other temple arm of the glasses body (100) and close to the user's ear, the AI processing chip and the wireless communication module are installed inside the frame of the glasses body (100), and the power module is installed inside the temple arm of the glasses body (100).
Citation Information
Cited By
Intelligent glasses voiceprint anti-counterfeiting recognition method and system based on NFC (Near Field Communication)
CN121054002A