Smart home control system based on multi-mode interaction
By integrating voice and image recognition modules and combining them with multi-module optimization algorithms in the control center, a multimodal interactive smart home system has been realized. This solves the limitations of the single interaction method in traditional systems, improves user experience and system intelligence, and provides a safe, comfortable, and energy-efficient living environment.
Patent Information
- Application Number
- CN202511920886.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-02-13
AI Technical Summary
Traditional smart home systems rely on a single interaction method, which limits the diversity of user experience and the adaptability of the system. How to efficiently integrate and process multimodal data to achieve precise control decisions is a technical problem that urgently needs to be solved.
The system employs a voice recognition module and an image recognition module for synchronous processing. Combined with the control algorithm of the control center and multiple modules (environmental perception, prediction, self-learning, safety detection, energy-saving optimization, and remote control), it optimizes multimodal interactive decision-making and achieves synchronous processing of user voice commands and gestures or facial expressions.
It enriches the interaction methods, enhances the flexibility and intelligence of the user experience, and provides a safe, comfortable, energy-saving, and personalized living environment.
Smart Images

Figure CN121523086A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of smart home control, in particular to a smart home control system based on multi-modal interaction. BACKGROUND
[0002] With the vigorous development of Internet of Things technology, smart home systems have gradually become an indispensable part of modern families. However, traditional smart home systems often rely only on a single interaction mode, such as voice or touch screen control, which greatly limits the diversity of user experience and the adaptability of the system. In order to overcome this limitation, multi-modal interaction technology has emerged, which combines voice, image and other interaction modes to provide users with a more natural and intuitive interaction experience. However, how to efficiently integrate and process data from different modalities and make accurate control decisions based on them is still a technical problem that needs to be solved for current smart home systems. Therefore, a smart home control system based on multi-modal interaction is proposed. SUMMARY
[0003] The purpose of the present application is to provide a smart home control system based on multi-modal interaction to solve the problems raised in the background art.
[0004] To achieve the above purpose, the present application provides the following technical solution: a smart home control system based on multi-modal interaction, comprising a voice recognition module, an image recognition module, a smart home device interface, a control center and a control instruction output module, the voice recognition module, the image recognition module and the smart home device interface are connected to the control center respectively, and the control center is connected to the control instruction output module; by integrating the voice recognition module and the image recognition module, the synchronous processing of user voice instructions and gestures or facial expressions is realized, thereby greatly enriching the interaction mode and improving the flexibility of user experience; The voice recognition module is used to receive user voice instructions and convert them into text information; The image recognition module is used to capture user gestures or facial expressions and convert them into corresponding control signals; The smart home device interface is used to communicate with various smart home devices; The control center is used to process signals from the voice recognition module and the image recognition module, and generate control instructions through a control algorithm; The control instruction output module is used to send the control instructions generated by the control center to the corresponding smart home devices to perform corresponding operations.
[0005] As a preferred embodiment, the control center comprises a data processing unit and a control algorithm. The data processing unit is configured to process signals from the voice recognition module and the image recognition module. The control algorithm is configured to generate control instructions for multi-modal interaction decision-making and to optimize the multi-modal interaction decision-making, and the multi-modal interaction decision-making is made according to the following formula:
[0006] wherein, represents a decision result, represents a confidence level of the voice recognition module, represents a confidence level of the image recognition module, represents a user preference weight, , and are weighting coefficients, and satisfy + + = 1. The multi-modal interaction decision-making is optimized according to the following formula:
[0007] wherein, is a slope parameter, is a threshold parameter, used to adjust the influence of the confidence level of the voice recognition module on the final decision result.
[0008] As a preferred embodiment, the image recognition module uses a deep learning algorithm to recognize gestures or facial expressions, and converts the recognition result into a control signal. The training process of the deep learning algorithm includes the following formula:
[0009] wherein, represents an updated weight, represents a weight before updating, is a learning rate, is a gradient of a loss function with respect to the weight.
[0010] As a preferred embodiment, the control center includes an environment perception module, a prediction module, a self-learning module, a safety detection module, an energy-saving optimization module, and a remote control interface. The environment perception module is configured to monitor environmental parameters in real time, and adjust the user preference weight in the control algorithm according to the environmental parameters . The prediction module is configured to predict user behavior according to historical interaction data of the user, and to adjust the state of the smart home device in advance. The self-learning module is configured to automatically adjust the weighting coefficients , and ; The security detection module is configured to detect abnormal behavior and take corresponding security measures. The energy-saving optimization module is configured to automatically adjust the operation mode of the device according to the user's living habits and device usage to achieve energy-saving effect. The remote control interface is configured to allow the user to remotely send control instructions to the control center through a mobile device.
[0011] Preferably, the environment perception module includes temperature sensors, humidity sensors, and light sensors, and the environmental parameter adjustment formula is as follows:
[0012] wherein, is the initial user preference weight, and are the current temperature, humidity, and light intensity, respectively, and are the ideal environmental parameters set by the user, and are the environmental parameter adjustment coefficients.
[0013] Preferably, the prediction module uses the following formula for behavior prediction:
[0014] wherein, is the behavior prediction result at time , is the weight of the th historical behavior, is the functional representation of the th historical behavior, is the time lag of the th historical behavior.
[0015] Preferably, the self-learning module uses the following formula to adjust the weighting coefficients:
[0016] wherein, and are the updated weighting coefficients, and are the weighting coefficients before updating, and are the adjustment amounts.
[0017] Preferably, the security detection module uses the following formula for abnormal behavior detection:
[0018] wherein, is an abnormal score, is the score of the th detection index, is the weight of the th detection index, is the number of detection indexes.
[0019] As preferred, the energy saving optimization module adopts the following formula for energy saving optimization:
[0020] wherein, is an energy saving effect score, is the energy consumption of the th device, is the weight of the th device, is the number of devices.
[0021] As preferred, the remote control interface adopts the following formula for instruction verification:
[0022] wherein, is a remote instruction verification result, is the score of the th verification index, is the weight of the th verification index, is the number of verification indexes.
[0023] Compared with the prior art, the above technical scheme has the following technical effects: the voice recognition module and the image recognition module are fused, the synchronous processing of the user voice instruction and the gesture or facial expression is realized, the interaction mode is greatly enriched, and the flexibility of the user experience is improved; the control center optimizes the multi-modal interaction decision by means of the control algorithm combined with the user preference weight, the intelligent level of the system and the user satisfaction are improved; and through the setting of the environment perception module, the prediction module, the self-learning module, the safety detection module, the energy saving optimization module and the remote control interface, the practicability and the convenience of the system are further improved, and a safer, more comfortable, more energy-saving and more personalized living environment can be provided for the user. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only constitute some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0025] Figure 1 System diagram of the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0027] It should be understood that the structures, proportions, sizes, etc. shown in the drawings of the present specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and do not define the limiting conditions for the implementation of the present application, so they do not have technical substantive significance. Any modification of structure, change of proportion relationship or adjustment of size, which does not affect the effects and purposes that can be achieved by the present application, should still fall within the scope of the technical content disclosed by the present application.
[0028] EMBODIMENT Please refer to Figure 1 The present application provides a technical solution: an intelligent home control system based on multi-modal interaction, including a voice recognition module, an image recognition module, an intelligent home device interface, a control center and a control instruction output module. The voice recognition module, the image recognition module and the intelligent home device interface are respectively connected to the control center, and the control center is connected to the control instruction output module. By integrating the voice recognition module and the image recognition module, the synchronous processing of user voice instructions and gestures or facial expressions is realized, thereby greatly enriching the interaction mode and improving the flexibility of user experience. The voice recognition module is used to receive user voice instructions and convert them into text information. The image recognition module is used to capture user gestures or facial expressions and convert them into corresponding control signals. The intelligent home device interface is used to communicate with various intelligent home devices. The control center is used to process signals from the voice recognition module and the image recognition module, and generate control instructions through a control algorithm. The control instruction output module is configured to send the control instruction generated by the control center to a corresponding smart home device to perform a corresponding operation.
[0029] The control center comprises a data processing unit and a control algorithm; the control center optimizes the multi-modal interaction decision by combining the user preference weight through the control algorithm, thereby improving the intelligent level of the system and the user satisfaction; The data processing unit is configured to process signals from the voice recognition module and the image recognition module. The control algorithm is configured to generate control instructions for multi-modal interaction decision and optimization of the multi-modal interaction decision, and adopts the following formula for multi-modal interaction decision:
[0030] wherein, represents the decision result, represents the confidence of the voice recognition module, represents the confidence of the image recognition module, represents the user preference weight, , and are weighting coefficients, and satisfy + + =1. The following formula is adopted to optimize the multi-modal interaction decision:
[0031] wherein, is a slope parameter, is a threshold parameter for adjusting the influence of the confidence of the voice recognition module on the final decision result.
[0032] The image recognition module adopts a deep learning algorithm to recognize gestures or facial expressions and converts the recognition result into a control signal. The training process of the deep learning algorithm comprises the following formula:
[0033] wherein, represents the updated weight, represents the weight before updating, is a learning rate, is the gradient of the loss function with respect to the weight.
[0034] The control center includes an environmental perception module, a prediction module, a self-learning module, a safety detection module, an energy-saving optimization module, and a remote control interface. The inclusion of these modules further enhances the system's practicality and convenience, providing users with a safer, more comfortable, energy-efficient, and personalized living environment. The environmental perception module is used to monitor environmental parameters in real time and adjust the user preference weights in the control algorithm based on these parameters. ; The prediction module is used to predict user behavior based on historical user interaction data and adjust the status of smart home devices in advance. The self-learning module is used to automatically adjust the weighting coefficients in the control algorithm based on user feedback. , and ; The security detection module is used to detect abnormal behavior and take corresponding security measures. The energy-saving optimization module is used to automatically adjust the equipment operating mode according to the user's lifestyle and equipment usage to achieve energy-saving effects; The remote control interface is used by users to send control commands to the control center remotely via mobile devices.
[0035] The environmental sensing module includes a temperature sensor, a humidity sensor, and a light sensor. The adjustment formula based on environmental parameters is as follows:
[0036] in, As initial user preference weights, and These are the current temperature, humidity, and light intensity, respectively. and Ideal environmental parameters set by the user. and These are adjustment coefficients for environmental parameters.
[0037] The prediction module uses the following formula to predict behavior:
[0038] in, For time Behavioral prediction results For the first The weight of each historical behavior, For the first A function representation of historical behavior, For the first The time lag of a historical action.
[0039] The self-learning module adjusts the weighting coefficients using the following formula:
[0040] wherein, and are the updated weighting coefficients, and are the weighting coefficients before updating, and are the adjustment amounts.
[0041] The security detection module detects abnormal behavior using the following formula:
[0042] wherein, is the abnormal score, is the score of the th detection indicator, is the weight of the th detection indicator, is the number of detection indicators.
[0043] The energy-saving optimization module optimizes energy saving using the following formula:
[0044] wherein, is the energy-saving effect score, is the energy consumption of the th device, is the weight of the th device, is the number of devices.
[0045] The remote control interface verifies the instructions using the following formula:
[0046] wherein, is the remote instruction verification result, is the score of the th verification indicator, is the weight of the th verification indicator, is the number of verification indicators.
[0047] In summary, by fusing the voice recognition module and the image recognition module, the user voice instruction and the hand gesture or facial expression are simultaneously processed, so that the interaction mode is greatly enriched and the flexibility of user experience is improved; the control center optimizes the multi-modal interaction decision by the control algorithm combined with the user preference weight, so that the intelligent level of the system and the user satisfaction are improved; moreover, the system further improves the practicability and convenience of the system through the environment perception module, the prediction module, the self-learning module, the safety detection module, the energy saving optimization module and the remote control interface, so as to provide a safer, more comfortable, more energy-saving and more personalized living environment for the user.
[0048] It will be appreciated by persons skilled in the art that features of various embodiments and / or claims of the present application can be combined or / and integrated, even if such combinations or integrations are not expressly disclosed in the present application. In particular, features of various embodiments and / or claims of the present application can be combined or / and integrated in any manner, without departing from the spirit and teaching of the present application. All such combinations and / or integrations are within the scope of the present application.
Claims
1. A smart home control system based on multimodal interaction, characterized in that: It includes a voice recognition module, an image recognition module, a smart home device interface, a control center, and a control command output module. The voice recognition module, image recognition module, and smart home device interface are respectively connected to the control center, and the control center is connected to the control command output module. The speech recognition module is used to receive the user's voice commands and convert them into text information; The image recognition module is used to capture the user's gestures or facial expressions and convert them into corresponding control signals; The smart home device interface is used for communication connection with various smart home devices; The control center is used to process signals from the speech recognition module and the image recognition module, and to generate control commands through control algorithms; The control command output module is used to send the control commands generated by the control center to the corresponding smart home devices to perform the corresponding operations.
2. The smart home control system based on multimodal interaction according to claim 1, characterized in that: The control center includes a data processing unit and a control algorithm; The data processing unit is used to process signals from the speech recognition module and the image recognition module; The control algorithm is used to generate control commands for multimodal interaction decision-making and to optimize multimodal interaction decision-making. The following formula is used for multimodal interaction decision-making: ; in, Represents the decision-making outcome. This represents the confidence level of the speech recognition module. Represents the confidence level of the image recognition module. Represents user preference weights. , and These are weighting coefficients, and they satisfy... + + =1; The following formula is used to optimize multimodal interaction decision-making: ; in, For the slope parameter, This is a threshold parameter used to adjust the impact of the confidence level of the speech recognition module on the final decision result.
3. The smart home control system based on multimodal interaction according to claim 1, characterized in that: The image recognition module uses a deep learning algorithm to recognize gestures or facial expressions and converts the recognition results into control signals. The training process of the deep learning algorithm includes the following formula: ; in, Represents the updated weights. Represents the weight before the update. For learning rate, This is the gradient of the loss function with respect to the weights.
4. The smart home control system based on multimodal interaction according to claim 1, characterized in that: The control center includes an environmental perception module, a prediction module, a self-learning module, a safety detection module, an energy-saving optimization module, and a remote control interface; The environmental sensing module is used to monitor environmental parameters in real time and adjust the user preference weights in the control algorithm based on the environmental parameters. ; The prediction module is used to predict user behavior based on historical user interaction data and adjust the status of smart home devices in advance. The self-learning module is used to automatically adjust the weighting coefficients in the control algorithm based on user feedback. , and ; The security detection module is used to detect abnormal behavior and take corresponding security measures; The energy-saving optimization module is used to automatically adjust the equipment operating mode according to the user's living habits and equipment usage to achieve energy-saving effects; The remote control interface is used by users to remotely send control commands to the control center via mobile devices.
5. The smart home control system based on multimodal interaction according to claim 4, characterized in that: The environmental sensing module includes a temperature sensor, a humidity sensor, and a light sensor. The adjustment formula based on environmental parameters is as follows: ; in, As initial user preference weights, and These are the current temperature, humidity, and light intensity, respectively. and Ideal environmental parameters set by the user. and These are adjustment coefficients for environmental parameters.
6. The smart home control system based on multimodal interaction according to claim 4, characterized in that: The prediction module uses the following formula to predict behavior: ; in, For time behavioral prediction results For the first The weight of each historical behavior For the first A function representation of historical behavior, For the first The time lag of a historical action.
7. The smart home control system based on multimodal interaction according to claim 4, characterized in that: The self-learning module adjusts the weighting coefficients using the following formula: ; in, and These are the updated weighting coefficients. and These are the weighting coefficients before the update. and For adjustment purposes.
8. The smart home control system based on multimodal interaction according to claim 4, characterized in that: The security detection module uses the following formula to detect abnormal behavior: ; in, For abnormal scoring, For the first The score of each detection indicator, For the first The weight of each detection indicator, This refers to the number of indicators detected.
9. The smart home control system based on multimodal interaction according to claim 4, characterized in that: The energy-saving optimization module uses the following formula for energy-saving optimization: ; in, To score the energy-saving performance, For the first Energy consumption of each device For the first The weight of each device This refers to the number of devices.
10. The smart home control system based on multimodal interaction according to claim 4, characterized in that: The remote control interface uses the following formula for command verification: ; in, Verify the results of remote commands. For the first The scores of each verification metric, For the first The weights of each verification metric, To verify the number of indicators.