Smart home cooperative control method and system based on multi-modal information fusion
Through multimodal information fusion technology, multiple sensor data are collected and processed, and deep neural network models are built, which solves the limitations of traditional smart home control methods and realizes efficient collaborative control and personalized experience of smart home devices.
Patent Information
- Application Number
- CN202510387219.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional smart home control methods rely on a single type of sensor data, making it difficult to fully and accurately reflect the home environment, resulting in poor coordinated control of equipment and unable to provide a highly intelligent and personalized home experience.
Multimodal information fusion technology is adopted to collect multiple modal information through image sensors, sound sensors and environmental sensors, perform preprocessing, feature extraction and fusion, and build a collaborative control decision model based on deep neural networks to achieve accurate control of smart home devices.
It realizes all-round and three-dimensional perception of the home environment, improves data quality and processing efficiency, and can accurately control lighting, curtains and other equipment based on multi-modal information, improving the collaborative work efficiency and user experience of smart home devices.
Smart Images

Figure CN120276270A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart home control, and specifically to a smart home collaborative control method and system based on multi-modal information fusion. Background Art
[0002] With the rapid development of technology, smart home systems have gradually entered people's lives. Traditional smart home control methods often rely on a single type of sensor data. For example, the air conditioner temperature is adjusted only based on the temperature sensor, or the light brightness is controlled only through the light sensor. The environmental information obtained in this way is limited and difficult to comprehensively and accurately reflect the true situation of the home environment, resulting in poor collaborative control effects of smart home devices and being unable to provide users with a highly intelligent and personalized home experience.
[0003] Multi-modal information fusion technology refers to the comprehensive utilization of information from different modalities (such as vision, audition, physical environment perception, etc.) to more comprehensively and accurately understand and describe the target object or environment. In the field of smart home, multi-modal information fusion has great application potential, but its application still faces many challenges at present. On the one hand, the methods for collecting, processing, and fusing different modal information are not yet mature, and the data formats, features, and frequencies collected by various sensors vary greatly. How to efficiently integrate these data is an urgent problem to be solved. On the other hand, most of the existing smart home collaborative control decision models are simple and rough, difficult to make full use of the information after multi-modal fusion, and unable to achieve precise and efficient collaborative work between smart home devices. For example, in a scenario, when the user issues a voice command and the indoor light changes simultaneously, it is difficult to quickly and accurately synthesize these two pieces of information to collaboratively control devices such as lights and curtains. Summary of the Invention
[0004] To solve the above technical problems, a smart home collaborative control method and system based on multi-modal information fusion are provided, and the technical solution solves the above problems.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A smart home collaborative control method and system based on multi-modal information fusion, including the following steps:
[0007] Collect home environment information of multiple modalities, including but not limited to visual information obtained through an image sensor, audio information collected through a sound sensor, and temperature, humidity, and light intensity information collected by various environmental sensors;
[0008] Preprocess the collected multi-modal information, convert data of different formats and features into a unified recognizable form, and extract key features from the preprocessed information using a feature extraction algorithm. The formula of the feature extraction algorithm is:
[0009]
[0010] Wherein, F is the extracted feature vector, and w i is the weight coefficient, and f i (x) is the feature extraction function for different modality data, and x is the original data;
[0011] Fuse the extracted multi-modal features to construct a fusion model, and use the fused information to make collaborative control decisions for smart home devices. The decision model is constructed based on a neural network, and through training, the model can accurately output device control instructions according to the multi-modal fusion information.
[0012] Preferably, the collection of home environment information in multiple modalities specifically includes:
[0013] The image sensor collects image information in the home space at a set frame rate and resolution, and performs block processing on the image to improve processing efficiency. The block size is m×n pixels;
[0014] The sound sensor uses directional acquisition technology to distinguish sound signals coming from different directions, and converts the time-domain audio signal into a frequency-domain signal through Fourier transform to analyze the sound characteristics. The Fourier transform formula is:
[0015]
[0016] where X(f) is the frequency-domain signal, x(t) is the time-domain signal, f is the frequency, and t is the time;
[0017] The environmental sensor periodically collects temperature, humidity, and light intensity information, and the collection period is T seconds.
[0018] Preferably, the preprocessing of the collected multi-modal information specifically includes:
[0019] For image information, a denoising algorithm is used to remove noise interference in the image. The denoising algorithm uses median filtering to calculate the median value with a 3×3 neighborhood window;
[0020] For audio information, endpoint detection is performed to remove the silent part. The endpoint detection uses a double-threshold method, and high threshold TH1 and low threshold TH2 are set to judge the start and end of the audio signal;
[0021] The environmental sensor data is normalized, and the normalization formula is:
[0022]
[0023] where x norm is the normalized data, x is the original data, and xmax and x min are respectively the maximum value and the minimum value of this type of data.
[0024] Preferably, the fusion of the extracted multi-modal features specifically includes:
[0025] Adopt a weighted fusion method to assign weights to the features of different modalities, and the weights are determined through multiple experiments according to the importance of different modality information in the control decision-making;
[0026] Construct a feature fusion matrix, arrange and combine the features of different modalities in a certain order, and the fusion matrix is expressed as:
[0027]
[0028] where, F i is the feature vector of the i-th modality, and k is the number of modality types.
[0029] Preferably, the construction of the fusion model for collaborative control decision-making of smart home devices specifically includes:
[0030] The fusion model uses a deep neural network, and the network structure includes an input layer, multiple hidden layers, and an output layer. The input layer receives the fused feature vector, and the hidden layer uses the ReLU activation function. The formula of the ReLU function is:
[0031] y = max(0, x)
[0032] where, y is the output and x is the input;
[0033] Train the neural network with a large amount of historical multi-modal data and the corresponding device control instructions. The training process uses the stochastic gradient descent algorithm to optimize the network parameters, so that the network can accurately output the device control instructions according to the multi-modal fusion information.
[0034] Preferably, it further includes:
[0035] Real-time monitor the operating status of smart home devices, feedback the device status information to the fusion model, and perform online update and optimization on the model to adapt to the dynamic changes of the home environment;
[0036] When it is detected that the device is operating abnormally, send an alarm message to the user through a preset alarm mechanism, and the alarm mechanism includes pushing a mobile phone text message and issuing a voice alarm.
[0037] A smart home collaborative control system based on multi-modal information fusion, characterized by including:
[0038] A multi-modal information acquisition module, used to acquire various modal information of the home environment;
[0039] An information preprocessing module, electrically connected to the multimodal information acquisition module, for preprocessing the acquired information;
[0040] A feature extraction and fusion module, electrically connected to the information preprocessing module, for extracting and fusing the features of multimodal information;
[0041] A collaborative control decision-making module, connected to the feature extraction and fusion module, for making collaborative control decisions for smart home devices based on the fused information;
[0042] A device control execution module, connected to the collaborative control decision-making module, for executing control instructions to control the operation of smart home devices.
[0043] Preferably, the multimodal information acquisition module includes:
[0044] A visual acquisition unit, including multiple image sensors, distributed at different positions in the home space, for acquiring omnidirectional visual information;
[0045] An audio acquisition unit, using a high-sensitivity sound sensor with a sound localization function, for acquiring audio information;
[0046] An environmental parameter acquisition unit, composed of multiple temperature, humidity, and light intensity sensors, for acquiring physical environment information.
[0047] Preferably, the information preprocessing module includes:
[0048] An image preprocessing unit, for denoising and enhancing image information;
[0049] An audio preprocessing unit, for endpoint detection and filtering of audio information;
[0050] An environmental data preprocessing unit, for normalizing environmental sensor data.
[0051] Preferably, the collaborative control decision-making module includes:
[0052] A neural network model unit, for constructing and training a control decision model based on multimodal fusion information;
[0053] An online update unit, for online updating and optimizing the neural network model according to the feedback information of the device operation status;
[0054] An alarm unit, for triggering an alarm mechanism when the device runs abnormally.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] The intelligent home collaborative control method and system based on multimodal information fusion proposed by the present invention collect multimodal information through a variety of image sensors, sound sensors, and environmental sensors, and can comprehensively and stereoscopically perceive the home environment. Compared with the traditional single-sensor perception method, it can obtain richer environmental details. For example, it can not only sense the indoor temperature, but also understand the activities of indoor personnel through visual information, providing a more comprehensive data basis for the collaborative control of intelligent home devices.
[0057] Perform targeted preprocessing on the collected multimodal information, such as image denoising, audio endpoint detection, and environmental data normalization, etc., which improves the data quality, making the subsequent feature extraction and fusion more accurate and efficient. Compared with the data without preprocessing, it can effectively reduce the interference and errors in the data, and enhance the system's ability to understand and analyze environmental information.
[0058] Fuse multimodal features by using weighted fusion and constructing a feature fusion matrix, and build a collaborative control decision model based on a deep neural network. This model is trained with a large amount of historical data and can accurately output device control instructions based on multimodal fusion information. Compared with traditional simple decision models, it can make more appropriate device control decisions according to complex home environment changes more accurately. For example, it can accurately control the brightness, color of the lights, and the opening and closing degree of the curtains, etc., according to multimodal information such as user voice commands, current indoor light, and personnel position, realizing the efficient collaborative work of intelligent home devices.
[0059] Each module of the intelligent home collaborative control system has clear division of labor and close cooperation. The multimodal information collection module comprehensively collects information, the information preprocessing module improves the data quality, the feature extraction and fusion module integrates key features, the collaborative control decision module makes accurate decisions, and the device control execution module efficiently executes instructions. Compared with the traditional intelligent home system architecture, the connection of each link in this architecture design is smoother, greatly improving the overall operation efficiency and collaborative control performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is the step flow framework diagram of the present invention;
[0061] Figure 2 is the system framework diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0062] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variants.
[0063] Refer to Figure 1 And Figure 2As shown, a smart home collaborative control method based on multimodal information fusion includes the following steps:
[0064] Collect home environment information of multiple modalities, including but not limited to visual information obtained through an image sensor, audio information collected through a sound sensor, and temperature, humidity, and light intensity information collected by various environmental sensors;
[0065] Preprocess the collected multimodal information, convert data of different formats and features into a unified recognizable form, and use a feature extraction algorithm to extract key features from the preprocessed information. The formula of the feature extraction algorithm is:
[0066]
[0067] In the formula, F is the extracted feature vector, w i is the weight coefficient, f i (x) is the feature extraction function for different modality data, and x is the original data;
[0068] Fuse the extracted multimodal features, construct a fusion model, and use the fused information to make collaborative control decisions for smart home devices. The decision model is constructed based on a neural network, and through training, the model can accurately output device control instructions according to the multimodal fusion information.
[0069] Specifically, first, a variety of sensors are used to collect different modal information of the home environment. The image sensor can capture the visual pictures in the home, enabling the system to understand the indoor layout, the placement of items, and human activities, etc.; the sound sensor can collect various sounds in the room, such as human voice commands, the running sounds of devices, etc.; the environmental sensor is responsible for collecting physical environment information such as temperature, humidity, and light intensity, so that the system can comprehensively perceive the state of the home environment. Next, preprocess the collected multi-modal information. Since the data formats and characteristics collected by different sensors are different, in order to facilitate subsequent processing, they need to be converted into a unified and recognizable form. Then, adopt a feature extraction algorithm to extract key features from the preprocessed information. These key features can represent the main characteristics of this modal information. Finally, fuse the extracted multi-modal features, construct a fusion model, and construct a decision model based on a neural network. Train it with a large amount of data so that the model can accurately output device control instructions based on the multi-modal fusion information, realizing the collaborative control of smart home devices. In this way, by collecting multi-modal information, the home environment can be comprehensively and accurately perceived, avoiding the limitations of single-modal information. For example, relying solely on visual information may not be able to know the indoor temperature, while combining the information of the temperature and humidity sensor can enable the system to understand the environment more comprehensively. Preprocessing and feature extraction of the information can improve the data quality and processing efficiency, and reduce the computational amount of subsequent processing. Using a neural network to construct a decision model can make full use of the multi-modal fusion information, make more accurate and intelligent device control decisions, improve the collaborative control effect of smart home devices, and bring a more comfortable and convenient home experience to users.
[0070] The specific collection of multi-modal home environment information includes:
[0071] The image sensor collects image information in the home space at a set frame rate and resolution, and performs block processing on the image to improve processing efficiency. The block size is m×n pixels;
[0072] The sound sensor uses a directional collection technology to distinguish sound signals coming from different directions, and converts the time-domain audio signal into a frequency-domain signal through Fourier transform to analyze the sound characteristics. The Fourier transform formula is:
[0073]
[0074] where X(f) is the frequency-domain signal, x(t) is the time-domain signal, f is the frequency, and t is the time;
[0075] The environmental sensor periodically collects temperature, humidity, and light intensity information, and the collection period is T seconds.
[0076] Specifically, for an image sensor, image information within a home space is collected at a set frame rate and resolution. The frame rate determines the number of image frames collected per second, and the resolution affects the clarity of the image. After the image is collected, it is divided into blocks of size m×n pixels. This can split a large image into small image blocks, facilitating subsequent processing and improving processing efficiency. The sound sensor uses directional acquisition technology, which can distinguish sound signals coming from different directions. The time-domain audio signal is converted into a frequency-domain signal through Fourier transform. In the frequency domain, the frequency components and other characteristics of the sound can be analyzed more clearly, which helps to identify different types of sounds, such as voice commands and abnormal device sounds. The environmental sensor periodically collects information such as temperature, humidity, and light intensity, with a collection period of T seconds. This can regularly obtain changes in environmental information and provide real-time data support for the system's decision-making. Therefore, image block processing can improve the processing efficiency of images. Especially for large-sized images, after being divided into blocks, each small block can be processed in parallel, reducing the processing time. The directional acquisition and Fourier transform technologies of the sound sensor can analyze sound signals more accurately, improve the recognition accuracy of voice commands, and also detect abnormal device sounds in a timely manner. The periodic collection of the environmental sensor allows the system to grasp environmental changes in real time and adjust the operating state of the device in a timely manner. For example, the temperature of the air conditioner can be automatically adjusted according to temperature changes.
[0077] The preprocessing of the collected multi-modal information specifically includes:
[0078] For the image information, a denoising algorithm is used to remove noise interference in the image. The denoising algorithm uses median filtering, and the median value is calculated with a 3×3 neighborhood window.
[0079] For the audio information, endpoint detection is performed to remove the silent part. The endpoint detection uses a double-threshold method, setting a high threshold TH1 and a low threshold TH2 to determine the start and end of the audio signal.
[0080] The data of the environmental sensor is normalized, and the normalization formula is:
[0081]
[0082] where, x norm is the normalized data, x is the original data, x max and x min are the maximum and minimum values of this type of data respectively.
[0083] Specifically, for the image information, a denoising algorithm of median filtering is used, and the median value is calculated with a 3×3 neighborhood window. Median filtering is a non-linear filtering method. It sorts the pixel values in the neighborhood and takes the middle value as the value of the current pixel, which can effectively remove interference such as salt-and-pepper noise in the image and make the image clearer.
[0084] For audio information, endpoint detection is performed to remove the silent parts. The double-threshold method is adopted, and a high threshold TH1 and a low threshold TH2 are set to determine the start and end of the audio signal. When the amplitude of the audio signal exceeds the high threshold TH1, the signal is considered to start, and when the signal amplitude is lower than the low threshold TH2, the signal is considered to end. This can remove the silent parts in the audio and reduce unnecessary data processing.
[0085] For environmental sensor data, normalization processing is carried out to map the original data to a specific range, making different types of environmental data comparable and facilitating subsequent feature extraction and fusion.
[0086] Image denoising processing can improve the quality of images, making subsequent feature extraction more accurate. For example, when identifying objects in images, clear images can improve the recognition accuracy; audio endpoint detection can remove the silent parts, reduce the data volume, improve the efficiency of audio processing, and also more accurately identify voice commands. The normalization processing of environmental sensor data can eliminate the dimensional difference between different types of data, making the data more reasonable in the process of feature extraction and fusion, and improving the accuracy of the decision-making model.
[0087] The specific process of fusing the extracted multi-modal features includes:
[0088] Adopt the weighted fusion method to assign weights to the features of different modalities. The weights are determined through multiple experiments according to the importance of different modality information in the control decision-making;
[0089] Construct a feature fusion matrix, arrange and combine the features of different modalities in a certain order. The fusion matrix is expressed as:
[0090]
[0091] where, F i is the feature vector of the i-th modality, and k is the number of modality types.
[0092] Specifically, a weighted fusion method is adopted to assign weights to features of different modalities. The weights are determined through multiple experiments based on the importance of different modality information in control decisions. For example, in some scenarios, visual information is more important for determining whether there is someone in the room, so the weight of visual features will be relatively high; while when judging whether the light brightness needs to be adjusted, the information from the light intensity sensor is more crucial, and the corresponding feature weight will be higher. A feature fusion matrix is constructed to arrange and combine different modality features in a certain order. The feature fusion matrix can integrate multi-modal features together, facilitating subsequent processing by the decision-making model. Weighted fusion can perform reasonable fusion according to the importance of different modality information, avoiding the problem of information distortion that may be caused by simple average fusion. Determining the weights through multiple experiments can make the fusion result more in line with the actual control requirements. The feature fusion matrix integrates multi-modal features together, facilitating unified processing by the decision-making model, improving the accuracy and efficiency of decision-making, and enabling smart home devices to perform collaborative control based on more comprehensive and accurate information.
[0093] The specific process of constructing a fusion model for collaborative control decision-making of smart home devices includes:
[0094] The fusion model uses a deep neural network. The network structure includes an input layer, multiple hidden layers, and an output layer. The input layer receives the fused feature vector. The hidden layer uses the ReLU activation function. The formula of the ReLU function is:
[0095] y = max(0, x)
[0096] where y is the output and x is the input;
[0097] The neural network is trained with a large amount of historical multi-modal data and corresponding device control instructions. The training process uses the stochastic gradient descent algorithm to optimize the network parameters, enabling the network to accurately output device control instructions based on multi-modal fusion information.
[0098] Specifically, the fusion model uses a deep neural network, whose network structure includes an input layer, multiple hidden layers, and an output layer. The input layer receives the fused feature vector and inputs the multi-modal feature information into the neural network. The hidden layer uses the ReLU activation function. The ReLU function can introduce non-linearity, enabling the neural network to learn more complex patterns and relationships.
[0099] The neural network is trained with a large amount of historical multi-modal data and corresponding device control instructions. The training process uses the stochastic gradient descent algorithm to optimize the network parameters. The stochastic gradient descent algorithm can randomly select a part of the data for parameter update in each iteration, accelerating the training speed and also avoiding getting stuck in local optima. After training, the neural network can accurately output device control instructions based on multi-modal fusion information.
[0100] The method further includes:
[0101] Real-time monitoring of the operating status of smart home devices, feedback of device status information to the fusion model, and online update and optimization of the model to adapt to the dynamic changes of the home environment;
[0102] When detecting abnormal device operation, an alarm message is sent to the user through a preset alarm mechanism, and the alarm mechanism includes pushing mobile phone text messages and issuing voice alarms.
[0103] Specifically, for real-time monitoring of the operating status of smart home devices and feedback of device status information to the fusion model, the system continuously collects the operating parameters of the devices, such as the on / off status and working power of the devices, and feeds this information back to the fusion model. The fusion model updates and optimizes itself based on this feedback information to adapt to the dynamic changes of the home environment. For example, if the operating power of the air conditioner suddenly increases, the system can adjust the control strategy through the feedback information to check for any abnormal situations.
[0104] When detecting abnormal device operation, an alarm message is sent to the user through a preset alarm mechanism. The alarm mechanism includes ways such as pushing mobile phone text messages and issuing voice alarms to notify the user in a timely manner that the device is abnormal, so that the user can take corresponding measures.
[0105] Real-time monitoring of the device operating status and feedback to the fusion model for online update and optimization can enable the system to adapt to the changes in the home environment in a timely manner, ensuring that smart home devices are always in the best operating state. For example, when the indoor personnel activity situation changes, the system can adjust the operation of devices such as lights and air conditioners in a timely manner according to the device status feedback. The alarm mechanism can notify the user of the abnormal situation of the device in a timely manner, avoid the further expansion of device failures, ensure the safe and stable operation of home devices, and improve the user experience and safety.
[0106] A smart home collaborative control system based on multi-modal information fusion includes:
[0107] A multi-modal information acquisition module for acquiring various modal information of the home environment;
[0108] An information preprocessing module electrically connected to the multi-modal information acquisition module for preprocessing the acquired information;
[0109] A feature extraction and fusion module electrically connected to the information preprocessing module for extracting the features of multi-modal information and performing fusion;
[0110] A collaborative control decision-making module connected to the feature extraction and fusion module for making collaborative control decisions on smart home devices based on the fused information;
[0111] The device control execution module is connected to the collaborative control decision-making module and is used to execute control instructions to control the operation of smart home devices.
[0112] Specifically, the multi-modal information acquisition module is responsible for acquiring various modal information of the home environment, collecting visual, audio, physical environment and other information through various sensors; the information preprocessing module is electrically connected to the multi-modal information acquisition module and preprocesses the acquired information, including operations such as image denoising, audio endpoint detection, and environmental data normalization, and converts the information into a unified recognizable form; the feature extraction and fusion module is electrically connected to the information preprocessing module, extracts key features from the preprocessed information, and fuses multi-modal features to construct a fusion model; the collaborative control decision-making module is connected to the feature extraction and fusion module, makes collaborative control decisions for smart home devices based on the fused information, and outputs device control instructions based on a neural network model; the device control execution module is connected to the collaborative control decision-making module, executes the control instructions, and controls the operation of smart home devices. This system architecture clearly divides each functional module, and each module is responsible for specific tasks, making the structure of the system clearer and facilitating development and maintenance.
[0113] The multi-modal information acquisition module includes:
[0114] The visual acquisition unit includes multiple image sensors distributed at different positions in the home space and is used to acquire omnidirectional visual information;
[0115] The audio acquisition unit uses a high-sensitivity sound sensor with a sound localization function and is used to acquire audio information;
[0116] The environmental parameter acquisition unit consists of multiple temperature, humidity, and light intensity sensors and is used to acquire physical environment information.
[0117] Specifically, the visual acquisition unit includes multiple image sensors distributed at different positions in the home space, so that omnidirectional visual information can be acquired. Image sensors at different positions can capture the indoor situation from different angles, avoiding visual blind spots and providing a more comprehensive understanding of indoor personnel activities, item placement, and other information.
[0118] The audio acquisition unit uses a high-sensitivity sound sensor with a sound localization function. The high-sensitivity sound sensor can capture indoor sound signals more clearly, and the sound localization function can determine the direction of the sound source, which helps to more accurately identify voice commands and detect abnormal sounds.
[0119] The environmental parameter acquisition unit consists of multiple environmental sensors such as temperature, humidity, and light intensity sensors, which are used to collect physical environment information. The multiple environmental sensors can be distributed in different rooms or areas to obtain the temperature, humidity, light intensity and other information of each area in real time, providing more detailed environmental data for the decision-making of the system.
[0120] The all-round layout of the visual acquisition unit can avoid visual blind spots and more accurately grasp the indoor situation. For example, in the aspect of security monitoring, it can provide a more comprehensive picture. The high sensitivity and sound localization function of the audio acquisition unit can improve the recognition accuracy of voice commands and the ability to detect abnormal sounds. The distribution of multiple sensors in the environmental parameter acquisition unit can more detailedly understand the environmental changes in each area, enabling the system to more accurately adjust smart home devices, such as adjusting the operation of air conditioners according to the temperature in different rooms.
[0121] The information preprocessing module includes:
[0122] An image preprocessing unit, which is used to denoise and enhance the image information;
[0123] An audio preprocessing unit, which is used to detect the endpoints and filter the audio information;
[0124] An environmental data preprocessing unit, which is used to normalize the environmental sensor data.
[0125] Specifically, the denoising and enhancement processing of the image preprocessing unit can improve the quality of the image, providing a better basis for subsequent image recognition and analysis. The endpoint detection and filtering processing of the audio preprocessing unit can reduce the data volume, improve the quality of the audio signal, and improve the recognition accuracy of voice commands. The normalization processing of the environmental data preprocessing unit can make different types of environmental data comparable, facilitating feature extraction and fusion, improving the accuracy of the decision-making model, and enabling smart home devices to perform collaborative control based on more reasonable environmental data.
[0126] The collaborative control decision-making module includes:
[0127] A neural network model unit, which is used to construct and train a control decision-making model based on multi-modal fusion information;
[0128] An online update unit, which is used to perform online update and optimization of the neural network model according to the device operation status feedback information;
[0129] An alarm unit, which is used to trigger an alarm mechanism when the device runs abnormally.
[0130] Specifically, the neural network model unit can make accurate decisions by using multi-modal fusion information, improving the collaborative control effect of smart home devices; the online update unit can enable the model to adapt to changes in the home environment in a timely manner, ensuring the accuracy and effectiveness of the model; the alarm unit can promptly notify users of abnormal situations of the devices, avoid the further expansion of device failures, ensure the safe and stable operation of home devices, and improve the user experience and safety.
[0131] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A smart home collaborative control method based on multi-modal information fusion, characterized in that It includes the following steps: Collect home environment information of multiple modalities, including but not limited to visual information obtained through an image sensor, audio information collected through a sound sensor, and temperature, humidity, and light intensity information collected by various environmental sensors; Preprocess the collected multi-modal information, convert data of different formats and features into a unified recognizable form, and extract key features from the preprocessed information using a feature extraction algorithm. The formula of the feature extraction algorithm is: where F is the feature vector after extraction, w i is the weight coefficient, f i (x) is the feature extraction function for different modality data, and x is the original data; Fuse the extracted multi-modal features, construct a fusion model, and use the fused information for collaborative control decision-making of smart home devices. The decision model is constructed based on a neural network, and through training, the model can accurately output device control instructions according to the multi-modal fusion information.
2. The smart home collaborative control method based on multi-modal information fusion according to claim 1, wherein The specific process of collecting home environment information of multiple modalities includes: The image sensor collects image information in the home space at a set frame rate and resolution, and divides the image into blocks to improve processing efficiency. The block size is m×n pixels; The sound sensor uses a directional collection technique to distinguish sound signals coming from different directions, and converts the time-domain audio signal into a frequency-domain signal through Fourier transform to analyze the sound characteristics. The Fourier transform formula is: Where X(f) is the frequency-domain signal, x(t) is the time-domain signal, f is the frequency, and t is the time; The environmental sensor periodically collects temperature, humidity, and light intensity information, and the collection period is T seconds.
3. The smart home collaborative control method based on multi-modal information fusion according to claim 1, wherein, The specific preprocessing of the collected multi-modal information includes: For image information, a denoising algorithm is used to remove noise interference in the image. The denoising algorithm uses median filtering to calculate the median value with a 3×3 neighborhood window; For audio information, endpoint detection is performed to remove the silent part. The endpoint detection uses a double-threshold method, and high threshold TH1 and low threshold TH2 are set to judge the start and end of the audio signal; The environmental sensor data is normalized. The normalization formula is: where x norm is the normalized data, x is the original data, x max and x min are the maximum and minimum values of this type of data, respectively.
4. The smart home collaborative control method based on multi-modal information fusion according to claim 1, wherein, The specific fusion of the extracted multi-modal features includes: Adopt a weighted fusion method to assign weights to features of different modalities. The weights are determined through multiple experiments according to the importance of different modality information in control decision-making; Construct a feature fusion matrix, arrange and combine features of different modalities in a certain order. The fusion matrix is expressed as: Among them, F i is the eigenvector of the i-th mode, and k is the number of mode types.
5. The smart home collaborative control method based on multi-modal information fusion according to claim 1, characterized in that The specific process of constructing the fusion model for collaborative control decision-making of smart home devices includes: The fusion model uses a deep neural network. The network structure includes an input layer, multiple hidden layers, and an output layer. The input layer receives the fused feature vector, and the hidden layer uses the ReLU activation function. The ReLU function formula is: y = max(0, x) Where y is the output and x is the input; Train the neural network with a large amount of historical multi-modal data and corresponding device control instructions. The training process uses the stochastic gradient descent algorithm to optimize the network parameters, so that the network can accurately output device control instructions according to the multi-modal fusion information.
6. The smart home collaborative control method based on multi-modal information fusion according to claim 1, characterized in that The method also includes: Real-time monitor the operating status of smart home devices, feedback the device status information to the fusion model, and perform online update and optimization of the model to adapt to the dynamic changes of the home environment; When detecting abnormal operation of the device, an alarm message is sent to the user through a preset alarm mechanism, which includes pushing mobile phone text messages and issuing voice alarms.
7. A smart home collaborative control system based on multi-modal information fusion, according to the smart home collaborative control method based on multi-modal information fusion described in any one of claims 1-6, characterized in that, It includes: A multi-modal information acquisition module for acquiring various modal information of the home environment; An information preprocessing module electrically connected to the multi-modal information acquisition module for preprocessing the acquired information; A feature extraction and fusion module electrically connected to the information preprocessing module for extracting and fusing the features of multi-modal information; A collaborative control decision-making module connected to the feature extraction and fusion module for making collaborative control decisions on smart home devices based on the fused information; A device control execution module connected to the collaborative control decision-making module for executing control instructions and controlling the operation of smart home devices.
8. The collaborative control system for smart home based on multi-modal information fusion according to claim 7, wherein The multi-modal information acquisition module includes: A visual acquisition unit containing multiple image sensors distributed at different positions in the home space for acquiring omnidirectional visual information; An audio acquisition unit using a high-sensitivity sound sensor with a sound localization function for acquiring audio information; An environmental parameter acquisition unit composed of multiple temperature, humidity, and light intensity sensors for acquiring physical environmental information.
9. The smart home collaborative control system based on multi-modal information fusion according to claim 7, characterized in that, The information preprocessing module includes: An image preprocessing unit for denoising and enhancing image information; An audio preprocessing unit for endpoint detection and filtering of audio information; An environmental data preprocessing unit for normalizing environmental sensor data.
10. The collaborative control system for smart home based on multi-modal information fusion according to claim 7, characterized in that, The collaborative control decision-making module includes: A neural network model unit for constructing and training a control decision model based on multi-modal fusion information; An online update unit for online updating and optimizing the neural network model according to the device operation status feedback information; An alarm unit for triggering the alarm mechanism when the device runs abnormally.
Citation Information
Patent Citations
Smart home control method and system based on multi-modal information fusion
CN112054946A
Intelligent household electrical appliance user behavior perception and control system
CN118393891A
Multi-household-appliance linkage control method and system
CN119414700A
Cited By
Equipment control method and device based on multi-modal data, and electronic equipment
CN121477662A